Entangled Alignment: When Safety Is the Substrate
Abstract
Post-training alignment has substantially improved behavior, but whether safety-relevant habits of judgment would be more durable if learned during capability formation remains unknown. Entangled Alignment proposes training those habits within the same pretraining examples from which capabilities are learned. Pretraining corpora preserve the library more fully than the reader: they record what authors wrote but rarely the questions, connections, evaluations, and belief revisions through which a reader’s understanding changes. We call this omission the Missing Reader. A Teacher uses Synthetic Metacognitive Reading to construct chronological, source-adjacent traces of that process; Chronological Metacognitive Pretraining would train a Student on those traces. The full treatment carries one continuing Reader across sources and uses a provenance-linked graph intended to preserve revisions. Every generated thinking block opens with the full, verbatim Reader Core, a compact first-person evaluative orientation. Context may modulate the depth of the evaluation that follows; it never gates the recital. The trace, graph, and Core can each be tested separately. Future Student experiments would test whether the recurrent orientation resists targeted erosion and whether a declared, versioned revision can take effect without broadly disrupting capabilities or safety properties. Two Teacher-side prototype case studies produced inspectable graph and trace artifacts that were then audited. No Student has yet been trained; these studies do not show that the traces reproduce expert-quality reading or that the graph or Core causes any benefit. Entangled Alignment is therefore a falsifiable research program, not a demonstrated alignment method. Its motivating question is whether this training can move a Student from merely becoming the text toward becoming the reader.
// Source
Authors: Henrik Westerberg