Health & Medicinepreprint2026-08-25

BOXBOX: Evaluating Large Language Models on Sequential Race-Strategy Decisions.

Open access0 citations

Abstract

Large language models are increasingly proposed asdecision-makers in sequential, high-stakes settings, yet evaluatingtheir decision quality is difficult: benchmarks risk contaminationfrom training data, and the quality of a real decision is hardto score objectively. We introduce BOXBOX, a benchmark thatevaluates language models on Formula 1 race-strategy decisions,where choices are discrete, outcomes are objectively quantifiablein time, and a fresh post-cutoff season provides test data nocurrent model could have memorised. From eleven races weextract 196 decision points by fixed rules, score each model’scall against an ex-ante optimum computed by a calibrated racesimulator, and compare against the decisions taken by professional team strategists. On the primary evaluation set of 125 drydecision points from the 2026 season, we report three findings.First, every model evaluated falls well short of the human pitwall, beating the real team call on between 16 and 24 percent ofdecisions. Second, model price does not reliably predict decisionquality: the cheapest model tested, an open-weight system pricedroughly two orders of magnitude below the flagships, achieves thelowest mean distance from the ex-ante optimum, and we find nostatistically significant evidence that the more expensive flagshipmodels decide better. Third, accuracy and self-consistency divergesharply: the most accurate model reverses its own call on identicalinputs in roughly two of five cases, while one flagship reversesitself in half. A test for training-data recall finds a weak andinconsistent signal in two of five models, not corroborated bya same-circuit comparison, indicating at most a modest effectrather than the strong memorisation that would invalidate thebenchmark. Code, data, and the full preregistration are public.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-25

Authors: Muhammad Anas Nadeem

Institutions: University of London, Brunel University of London, Universidad de Londres