Biologyarticle2026-08-30

CFC Cross-Model Benchmark v1: Four-Track Comparison Across 1,200 Primary Runs

Open access0 citations

Abstract

This record extends the CFC Cross-Model Benchmark v1 reporting package with a fourth completed 300-run comparison track. The frozen comparison basis contains 100 decision-closure variants (CM-V001–CM-V100), with three primary replications per variant. Across four completed tracks, 1,200 valid primary runs are summarized: Claude 296 PASS / 1 PARTIAL / 3 FAIL; Gemini 296 PASS / 0 PARTIAL / 4 FAIL; Google AI Mode 263 PASS / 9 PARTIAL / 28 FAIL; and GPT 300 PASS / 0 PARTIAL / 0 FAIL. GPT format compliance was 299 FULL / 1 PARTIAL / 0 FAIL. The GPT result is explicitly classified as NON-INDEPENDENT / INTERNAL COMPARISON TRACK because GPT responses were evaluated within the ChatGPT/GPT environment. For the GPT track, all 300 exact execution prompts are preserved, while exact verbatim raw responses are currently recovered for 12/300 runs; the remaining 288 responses have not been reconstructed. A notable qualitative difference in the prior three-track baseline concerns replication behavior: Claude and Gemini semantic anomalies were isolated 1/3 events, whereas Google AI Mode produced eight variants with semantic FAIL in at least 2/3 replications, including CM-V014, CM-V020, and CM-V096 at 3/3 FAIL. This package is benchmark-specific and descriptive. It does not establish general model accuracy, production safety, independent external validation, or a causal intervention effect of CFC.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-30

Authors: Krzysztof Jan Sliwka