Self-Consistent Misalignment: Observability Constraints and Structural Consequences in Adaptive Multi-Agent Systems
Abstract
RT2 specifies self-consistent misalignment as a trajectory where task-relevant mismatch deepens while internal evaluation continues to report health and loses discriminating power. It separates this trajectory definition from mechanisms that might generate it. Conditional information and observation results identify distinct limits without proving inevitable observer loss or impossible autonomous correction. Version 3.0 proposes reference-displacement and response measurements that distinguish information, evaluation, updating, and recovery. Explicit mathematical examples and prospective comparisons replace unreproduced numerical demonstrations. Related literature supplies bounded examples, while framework-specific measurement gains, scaling relations, and correction-cost laws remain unvalidated.
// Source
Authors: Bin Seol