AI & Computingpreprint2026-08-23

Independent Sessions, Shared Evidence: Auditing Evidence Convergence in LLM Web Research

Open access0 citations

Abstract

Independent model sessions are often treated as independent checks, but shared access to the public web can correlate the evidence they retrieve. We audit this dependence using 600 primary research runs from two consumer LLM web-research systems evaluating the same 300 frozen SEO/GEO claims. Matched pairs shared at least one exact retained URL in 67.3% of cases, but their mean exact-URL Jaccard overlap was only 0.223, indicating partial rather than near-identical evidence sets. In 10,000 system-scope-preserving claim permutations, exact-URL overlap averaged 13.1% and mean exact-URL Jaccard averaged 0.047; no permutation reached the matched observation. Domain overlap was more saturated, with 90.9% matched overlap versus a 47.6% scope-preserving null mean. A post-freeze provenance join showed that 158/205 matched pairs with normalized-URL overlap shared at least one source linked to the claim’s original discovery. The remaining 47 residual pairs converged on 43 distinct normalized URLs outside that recorded discovery-linked set. Verdict agreement was 294/300 (98.0%), but 290/300 pairs were SUPPORTED/SUPPORTED, so the verdict result is strongly prevalence-conditioned. We present the study as a method for auditing evidence dependence, verdict convergence, and surfaced process telemetry in multi-model web research, not as an accuracy benchmark or a measurement of correlated factual error. The main missing baseline is within-system test-retest repeatability, which we leave for a pre-registered follow-up rather than retrofitting post-hoc repeats into the frozen experiment.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-23

Authors: Mika Sipilä

Institutions: Rochester Institute of Technology - Dubai