Black Box and Beyond: Auditing Large Language Models, AI Agents, and AI Infrastructure Under the UK and EU Regulatory Frameworks
Abstract
Regulators in the European Union and the United Kingdom now require organisations to make high risk AI systems auditable, and neither has said what an audit record must be. This paper supplies the missing content. It argues first that a record is evidence only where its modification is detectable and where it can be retrieved and checked by a party the deployer does not control, that these two properties are what separate an audit log from a log, and that no binding instrument in either jurisdiction requires either, although both have been standardised engineering for decades. It argues second that the single word auditability carries three capacities of sharply different achievability, and that one of them, the reconstruction of why a model produced a given output, is unavailable for current large language models and agents as a limit of the architecture and not of infrastructure investment, so an obligation at that level will be met by a substitute, and the likeliest substitute, a logged chain-ofthought trace, is a model output correlated with the computation and not a record of it. It argues third that the technical content of the record will be settled inside the window that widened when the high risk obligations were deferred to December 2027 and August 2028, and that standardisation responds to whatever installed base exists by the time it reports. The paper therefore states a specification in the open, as technical, organisational and regulatory layers built with primitives standardised for decades, and examines one instantiation with its limits stated.
// Source
Authors: Julius Osi Abu
Institutions: Aptevo Therapeutics (United Kingdom)