Three instrument measurements behind a data specification: within-person autocorrelation, the power it implies, and the ~1,500-observation threshold
Abstract
METHODS / INSTRUMENT REPORT. Version v0.1. This deposit repays an obligation recorded in the co-measurement gap note (v1.2, 10.5281/zenodo.21767213, section 6), which states three figures and labels them claims rather than evidence because they were held only in a private repository. The three: (1) a median within-person AR(1) of 0.282 across 1,719 single-item series from four acquired ESM datasets (225 persons); (2) power under 11% to detect a two-state generator below a 6 SD separation, on a benchmark matched to that autocorrelation, the project's earlier benchmark having been mis-specified at AR(1) 0.81-0.88; (3) the ~1,500-observations-per-person threshold for the persistence route. Figure 3 did not survive its own re-derivation, and this deposit reports that instead of the figure. The central result, in the scope it actually has: WITHIN THE MODEL CLASS THE GATE ASSUMES - a two-state switching autoregression - the observed autocorrelation is incompatible with the regime in which the gate can fire. Sweeping the true dwell time of the generating process, the gate first fires at a dwell of about 3, and a dwell of 3 produces a marginal AR(1) of 0.49. The acquired data's marginal is 0.282. Switching between separated states adds autocorrelation of its own, so dwell and marginal move together, and the observed value pins the dwell at about 2 - inside the region where the gate returns 0 of 100 across separations of 3, 4 and 5 SD and across 600 AND 1,500 observations. No gate constant is consulted anywhere in that chain. So ~1,500 observations per person is not the price of the persistence route on data of this shape; no N is. That was the only achievable-looking data specification the project had produced. The scope limit is load-bearing and is stated rather than buried. The relation between dwell and marginal autocorrelation is a property of that model class, which is the class the instrument already assumes. If the real process lies outside it - a continuously drifting nonlinear system, a fold with non-Markovian dwell times, a slowly moving attractor - the mapping does not transfer, and the observed marginal implies nothing about a dwell because there is no dwell to imply. Whether the class fits at all is a prior question this report cannot answer, since the instrument under test is also the instrument that would answer it. A third round of critique asked instead whether simulating a two-state process to test a two-state detector is circular; as put it is not, since that is what a power analysis is and the claim about the data rests on one measured statistic. The sharper version, above, is the one that holds. The mechanism first offered for the zero was asserted and was wrong, and the correction is recorded in place rather than substituted. The gate's verdict is a conjunction - a BIC margin AND a minimum implied dwell - and the calibration recorded only the conjunction, yet the text explained the zero by naming one clause. Decomposed at 100 reps per cell, the two clauses turn out to be mutually exclusive at dwell 2: 0 of 600 replications satisfy both, where independence would predict 13 to 20 per cell. When the fit recovers the true fast-switching structure it wins on BIC and reports an implied dwell near 2, failing the dwell clause; when it fits a slow-switching alternative it passes the dwell clause and no longer beats the AR(1) null. At one cell the BIC clause fails 99% of the time and the dwell constant does nothing at all. Nor is the zero an artefact of the adopted minimum-dwell constant, and that was measured rather than argued. The threshold is applied after the fit, so the full sensitivity curve comes free from the same replications: power is 0% at every threshold at or above 2.5, and opens only at 1.5 - below the process's own dwell, i.e. at a value that would admit as states exactly what the constant exists to exclude. An earlier version of this text answered that objection by saying that lowering the threshold would be drift. That answered a different question. Refusing to move a threshold in order to rescue a result is a discipline rule; measuring how an instrument behaves as a function of its own threshold is instrument characterisation, which is what a methods report is for. The first does not excuse skipping the second, and this document had used it to. One escape route was proposed and measured: a lower within-state autocorrelation with a longer dwell, said to reach the same marginal. At a 4 SD separation it does not - within-state 0.05 with dwell 5 gives a marginal of 0.52, 0.05 with dwell 10 gives 0.66, 0.15 with dwell 5 gives 0.57, none reproduces 0.282, and all three give the gate 100% power. A conditional this text had added earlier was therefore over-stated and is corrected. Further limits: all of the dwell work is at a 4 SD separation and 600 observations, so a low-separation, long-dwell, low-within-state corner is not excluded by measurement. That corner is also where the gate is weakest, which is an argument and is marked as one. Where the gate does work it is stronger than the superseded 10-reps-per-cell table said: at an adequate dwell, a 4 SD separation at 600 observations is already 100%, and 3 SD at 1,500 observations is 100% rather than 60%. False alarms are 0 of 100 in all twelve calibration cells. Two cells therefore contradict the old table in the theory's favour, which is why the disposal rule - the numbers with a derivation attached win - was frozen in the script before the run rather than chosen after seeing the output. The old table cannot simply be called wrong either: its generator is unknown, because no script in the source repository ever produced it, an absence checked against the full git history rather than asserted. Why a number was reported without a committed derivation is answered rather than merely recorded: it was produced in an interactive session and transcribed, while this project's own convention that analysis code lives in the repository was skipped for that one run. It establishes nothing about whether a felt signal tracks viability, and nothing about the theory that motivated the work. Every measurement here is about the instrument. Competing interest, stated because a reader is entitled to know it rather than infer it: the author is an independent researcher with no institutional affiliation, the work was unfunded, and the author originated the framework whose measurement route this report constrains. Three rounds of AUTOMATED CRITIQUE preceded deposit and this was not peer review - a language model was asked to attack the draft, with no independent replication, no accountable reviewer and no editorial process. A critic that cannot be wrong in public is not a reviewer. Round 2's central premise, that the conclusion was an artefact of one adopted constant, was refuted by the measurement it prompted. A stopping rule is declared rather than left to fatigue: yield fell from four live points to two to one, and no fourth round was run, because the remaining gaps close with data or with simulation outside the assumed model class and neither is produced by further critique. Three further declarations. Two separation figures previously circulated by this project, 1.92 SD and 3.70 SD, are RETIRED and must not be quoted in either unit. The four datasets are the same four used in this project's published negative head-to-head (10.5281/zenodo.21422215), so this is not an independent sample. And person-level derived files are withheld because the datasets were acquired under redistribution restrictions; the code and the fully synthetic simulation outputs are included, so a reader holding the same four datasets can re-run every number reported. Reading guidance, because the document's format has a cost. Corrections are stacked in date order and superseded text is kept in place rather than deleted, so the deposit opens with a flat index - section 0.0 - listing the claims it currently makes and where each is measured, the claims it has withdrawn and what replaced them (including the author's own wrong mechanism and his own over-stated caveat), and what it does not establish. Nothing beneath that index was altered to produce it. One rule survives the layering: quote section 4.5, not section 4.1, which reports the superseded table and is kept only for the record.
// Source
Authors: Hiroaki Aizawa