Society & Economicsarticle2026-07-31

A Top-Down Method for Estimating the Market Value of Creative Work to AI Systems: Evidence from the Reddit-Google Licence

Open access0 citations

Abstract

Google reportedly pays Reddit \$60 million a year for data access covering both model pre-training and the grounding of live answers. This paper uses that disclosed price, together with a top-down valuation method emerging from the data-economics literature, to estimate the share of AI-system revenue attributable to the creative work and human-authored content those systems are built on. The method runs in four stages: the product's revenue pool, the share of model value attributable to training data, the share of that attributable to the relevant content class, and the share of the class attributable to the specific corpus. Run forward, the chain prices what the Reddit archive could be worth for training. Run backwards from the disclosed fee, it returns the valuation of creative work the deal implies. Because the contractual split of the fee between training and live access is not public, the implied share is reported as a range of 7 to 24\%. A Monte Carlo simulation across 200,000 runs, treating that split as unknown, returns a median of 10.5\%, with nine runs in ten between 4.6 and 19.6\%. That median converges with the 10% anchor derived independently from the scaling literature. Applied to the sector, a 10% share values the creative-work pool across OpenAI and Anthropic alone at about $7.2 billion a year, against content-licensing payments estimated at roughly \$2.2 billion. The paper closes with the allocation question: how an agreed pool should be divided among those who contributed to it.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-07-31

Authors: Conor Roche

Institutions: Institute for New Economic Thinking