Semantic Impedance Matching in Multi-Agent Language Model Pipelines: Measuring, predicting, and correcting information loss at agent boundaries
Abstract
When a signal crosses a boundary between transmission media of mismatched impedance, part of it reflects and transmitted power is lost, the phenomenon the Smith Chart was created to tame in 1936. This study tests whether an analogous framework describes information loss when context is handed off between large language model agents in a pipeline. Relay chains of four agents pass synthetic incident briefings under a fixed word budget, and the survival of twenty programmatically generated, unguessable atomic facts is measured at every handoff using fully deterministic grading with no model judging at any step. The study comprises 998 relay chains and 3,992 model calls across five models; the main grid is four conditions by 30 tickets by two replicates at fixed temperature on four open-weight models spanning 3B to 120B parameters. Results. (1) Pipelines of identically instructed agents attenuate gently, while pipelines whose agents hold mismatched representational conventions collapse at the first mismatched boundary, and the loss is selective rather than random: business-salient facts survive while all sixteen technical identifiers vanish in every chain. (2) The central control isolates mismatch from obedience. Where adjacent hops merely alternate output format, a JSON object versus narrative prose, with no instruction to omit anything and the same preservation goal in both, transmission still falls significantly on all four models (0.05 to 0.26 absolute, every p below 1e-4). Representational mismatch destroys information independently of any instruction to discard it, although it accounts for part rather than all of the larger persona effect. (3) A one-paragraph matching network, a mandatory structured fact ledger emitted at each boundary and paid for out of the same word budget, recovers mismatched pipelines by up to 0.83 absolute, restores them to the matched baseline rather than beyond it, and inverts below a capability floor, where the smallest model ends worse than its own baseline because it cannot execute the protocol. (4) Heterogeneous chains alternating two model families show that measured loss is receiver-relative: an apparent 29-point deficit was the grading harness failing to read one model's Unicode output dialect, exposed when the other model recovered the invisible facts. Reported correction. An earlier version of this work advanced, as a headline finding, that mismatch loss grows with model capability. Under the larger replicated design the trend is non-monotone and, with only four model points, cannot reach significance even in principle (exact permutation floor p = 0.083; observed p = 0.33). It is restated in the paper as an open question requiring at least six models, with the power analysis given in full. This deposit includes the paper, the complete analysis and experiment code, the seeded payload generator, and every stored briefing at every hop, so that all reported numbers and figures can be regenerated without re-running inference. Disclosure. This study was designed, executed, analyzed, and drafted in an interactive session with Claude (Anthropic) acting as a research assistant under the author's direction. The author set the research question, approved the design, and is responsible for the content and its claims.
// Source
Authors: Gustavo Carreno
Institutions: University of Colorado Boulder, University of Colorado System