The AI Automation Claims Scorecard: a motive-neutral referee for self-reported vendor figures and stated safety red-lines
Abstract
An open, reproducible, motive-neutral referee for the claims that frontier AI labs and large adopters make about how much of their code an AI now writes, and for the recursion red-lines those labs state in their own safety frameworks. Version 0.1.1 (title refined to The AI Automation Claims Scorecard to reflect empirical audit of self-reported claims vs execution benchmarks). Twelve public claims (Anthropic, OpenAI, Microsoft, Google, Meta, Salesforce, Amazon, Snap) are scored not on whether the headline percentage is true but on whether it is checkable, using four evidentiary tests that apply to any public statistic: whether the denominator is defined, whether the figure is externally audited, whether an outsider could reproduce it, and whether an independent source has checked it. Across the twelve, one defines a rigorous denominator with a further six partial, none is externally audited, none is externally reproducible, and two carry an independent check of the specific figure. The composite score is an unweighted convenience index (mean 0.50 out of 4), not a validated metric; the individually checkable per-dimension counts are the finding, and the index is the same 0.50 for current claims and for forward-looking ones taken separately. Every claim is traced to a primary document held and hash-verified, listed with its address, retrieval date and SHA-256 in an accompanying source register, and for the six transcript sources with the timestamp at which the quoted words occur. Read from source, four of the circulating figures prove narrower than the versions in general use: one changes meaning inside the interview it comes from, one is scoped to some projects rather than all repositories, one is about the listener's codebase rather than the speaker's company, and one drops a qualifier. A fifth property, whether a claim isolates the research-loop subset that safety frameworks are written about, is reported but deliberately not scored, since a public statement about general coding automation has no duty to break it out; three of the twelve isolate it, correcting version 0, which reported that column as empty. A second layer records the labs' own published red-lines for automated research and development and dates their status, together with their stated arrival dates, the one datable observed instance of an AI improving its own training (Google DeepMind's AlphaEvolve), and the first independent cross-lab instrument, METR's entity-based pilot of February to April 2026. No lab reports its recursion red-line crossed. A dependency-free Python script (reproduce_v0.1.py) recomputes the scored table, the aggregate findings and the register, verifies every held source against its recorded hash, and exits with an error if any row is unsourced or unanchored. This is a claims referee, not a capability benchmark: it complements the measurement agenda set out by Chan et al. (arXiv 2603.03992) by grading the disclosures that already exist against it, rather than proposing new metrics. Corrections to version 0, including three counts that were wrong at the date of that deposit, are recorded in full in the accompanying corrections register. AI disclosure: the research is the author's; this text was drafted with AI assistance and reviewed by the author. Conflict of interest: models from Anthropic and Google assisted with retrieval, calculation and drafting, and an OpenAI model assisted with the pre-publication red-team review; all three are graded parties here. Anthropic holds the highest-scoring row, is the most vocal party on recursive self-improvement and is named by Snap as the vendor behind its own figure; Google holds a row of its own; and OpenAI is a graded party. No row was softened or sharpened for any vendor, and every claim is anchored to a held primary so a reader can check the graded text against the source without trusting any party or the author. Independent analysis and open-science documentation only, not investment advice.
// Source
Authors: N Milton