Author
David Antonio Tomé
0 works0 citations
Recent research
- AI & ComputingOpen access
An x-ray for AI: a deterministic, model-free grounding audit of an LLM abstention benchmark
An external, deterministic, model-free grounding audit of the Rafe and Das (2026) helmet-abstention benchmark (1,464 emergency-department injury narratives). A model-free gate re-reads each note against a sealed schema and returns a three-valued verdict; it reproduces the benchma...
- AI & ComputingOpen access
Kohli (2026, arXiv:2605.29800) showed that nine frontier LLM judges deliver only ~2.2 effective independent votes because their errors are correlated, and that the correlation is mostly not explained by item difficulty. This deposit contains a pre-registered replication on consum...