Grounds, Frames, and Standing: The Absorbable and Irreducible Classes of Human Intervention in Agentic Work
Abstract
Who drives successful execution: the AI or the human? Binary attribution is not just reductive but also limited as an answer, because human intervention comes in different classes with varying degrees of absorbability. The split of AI-human contributions can remain stable while the nature and effect of human interventions change underneath it, as better models take over some classes, while still requiring human input for others. This paper derives where that distinction sits and why, as well as which specific classes fall on each side. The thirteen mechanisms of human intervention catalogued in the companion paper sort into three durability classes by what they supply, and each class has an absorption route of its own. Grounds, considerations bearing on entitlement to conclusions, live in the world, so retrieval, wider context, and stronger verification reach them, at a rate set by the domain's verifier availability rather than by capability alone. Frames, replacements of the hypothesis space, require a position outside the current framing, so the class yields to a second reasoner with decorrelated priors. What absorbs the class is therefore the existence of such reasoners, a population fact rather than a capability fact. Under monoculture that route closes, and because a frame cannot be requested by name, remaining frame supply turns on a decision to invoke a source, which is threshold-setting and relocates the burden onto the third class. Standing is the partition's irreducible class, absorbed by neither capability growth nor architecture. It is what a stop signal, a demand for judgment, and a re-anchoring to purpose transfer while supplying no domain content of their own: a stop's timing and sign reveal the threshold's value, which a record can teach, while the authority to set and revise that threshold does not reduce to a fact, so a stop an agent issues to itself is a decision rather than a stop. The impossibility is not a mathematical one, since a greatest-fixed-point reading of the conferral relation, one that admits circles of mutual conferral with no originator, would validate self-certifying cycles, and nothing in the mathematics forbids it. Institutions enforce the grounded reading instead, which makes standing's non-self-conferrability a governance fact with a repeal path rather than a theorem. The requirement the class leaves is for a principal, not for a human, a role architecture can relocate but not eliminate. The empirical results reported in this paper qualify the partition rather than support it. A pre-registered four-condition test of the supply condition gave the same thirty-four decision points to four models spanning two generations, under identical briefs, each asked for eight candidate reframings: the human's actual reframe was present in 2.9 to 11.8 percent of sets, the four conditions do not differ at a permuted p of 0.496, and twenty-eight of the thirty-four items were covered by none of them while two were covered by three of four, where independent misses at these same marginals predict a tenth of one. Adding models of one lineage did not add frames. That is the monoculture case Section 2.1 anticipates rather than a refutation of the route, but it removes the reading on which a better or merely different model absorbs the class. Two generators from other vendors were then run under the same rubric and covered none of a different twenty-eight, the items surviving every pre-registered gate across all six conditions, which is not the uncovered set above because five of its members were covered, so the union over six conditions equals the union over four and nothing was being reached elsewhere; the registered test of whether they reach different frames needs one non-Anthropic hit to divide and returned no result at zero. A separate limit sits under all of these figures and is not a property of any model: three of the six frames-family classes act on the collaboration rather than on the problem, the generation rubric asks only about the problem, and those three classes hold roughly a third of the items and had never once been covered. Two tasks were then run over the same material and they disagree in one direction. Asked whether the human's move is present among eight candidates, raters credit it on a handful of items. Asked instead whether each credited turn and the one candidate they themselves named make the same suggestion, with the search step removed, none of the eight pairs so tested is unanimously the same move and six are unanimously not, against raters who caught seven or eight of eight constructed mismatches and accepted eight of eight known-same pairs. Membership among eight therefore inflates credits relative to direct comparison, and every rate quoted here rests on the more permissive task. What the stricter one shows is sharper than a rate: the model's nearest candidate is reliably identifiable, raters agree on which one it is, and it is still not the move the human made. The tempting reading of that gap, that these generators supply the move which surfaces a choice and miss the one that exercises it, is refuted by its own test. Coding candidates alone, with human turns interleaved unlabelled and anchors gating five admitted and genuinely distinct draws, puts the surfacing share of candidates at 0.500, interval 0.313 to 0.687 against a pre-registered threshold of 0.70, and the humans beside them at 0.579. Candidates skew toward surfacing less rather than more, and the offset the reading rested on is a selection effect of the eight pairs. Every rate here is judge-dependent as well: a panel from a third lineage, reading the same packet bytes, credited three of the fifty-six probes the two added conditions comprise where the first panel credited none, and six raters on the condition holding those credits returned six, three, one, five, seven and zero on byte-identical input. Every reliability coefficient below is an agreement among language model instances rather than among people, and no human second-rater run has been completed. Rater, coder, judge and reader name one role throughout, and the bound covers all four. That holds for every panel reported here and is stated once, at this point, rather than repeated beside each figure. None of the four therefore carries a human default here: a figure like the 0.828 that follows bounds how consistently one model family applies a rubric, not how well that rubric survives contact with an independent human reader. One bound sits under every count above rather than beside them. This practice issues interventions through two channels, typing them and writing them into a version-controlled rule store so they need not be typed again, and only the first is coded. Over the two months for which both are datable, 1,094 interventions were typed against 277 rules persisted, a ratio that moved from 0.11 to 0.30 within that window, so about one intervention in five went where no count here looks and its share was climbing. That also supplies a competing explanation for the grounds class thinning out, since a procedural fact is written down once and never again and the flow of such rules decays as the enumerable set fills, whatever capability is doing. Read on the rule store instead, which is the longest channel the practice has, the compositional prediction is not confirmed: 141 durable rules sampled at up to 18 per month and coded blind at a Fleiss kappa of 0.828 put the grounds share rising from 14 to 24 percent at p = 0.169 where the prediction is a decline, on an instrument whose sampling frame is itself in question, so what the record supports is that absorption has not been shown rather than that it cannot be. Two further instruments qualify the partition from directions the channel bound does not reach. A pre-registered attempt to recover the frames-standing boundary from how interventions are phrased did not find it, at Cramér's V of 0.201 and a session-permuted p of 0.1114, and the attempt's own registered controls undercut most of what that null would otherwise be worth: a keyword classifier matched the coding, and the direction test passed one of three. Blind coding of 82 accounts of what the agent failed to reach, at Krippendorff's alpha of 0.920 and 0.900, puts 29 of them, the largest single category, outside all three classes: nothing was missing and nothing was misorganised, and the assistant's own default produced the failure. The partition classifies interventions rather than failures and is not obliged to cover that, but the mechanisms which correct it are the zero-content ones this paper argues absorb last. A third channel sits under the counts alongside the rule store and damages them differently, putting attribution rather than coverage in question. Section 7 reports that work supplied into one session and relayed by an agent into another arrives as that agent's articulation, and that the reported contribution split moves with the number of hops whether or not a true split exists. Under incomplete contracting, capability improves inference from a fixed information set without enlarging it, so expected loss from a misset delegation threshold converges not to zero but to a strictly positive floor set by the residual dispersion of the principal's threshold, its conditional variance under squared-error loss and its conditional mean absolute deviation under absolute loss. The floor rests on either of two premises, non-enumerability that persists under learning or preference construction in the threshold itself, so the framework's sharpest falsifier is the failure of both premises together. The two are not interchangeable under challenge: the first is the target of a known formal attack on incomplete-contract theory and the second is not reached by it, so the weight sits on preference construction wherever that attack is granted. The mechanism list transfers across substrates while its couplings do not. A minor
// Source
Authors: Jongsun Suh