Society & Economicsarticle2026-08-07

Recognized but Not Machine-Readable: An Evidence-Tiered Audit of Sahitya Akademi Awardees in Wikidata

Open access0 citations

Abstract

Wikidata is a major open source of structured, machine-readable knowledge, used across Wikimedia's own projects and increasingly drawn on by search engines, question-answering systems, and knowledge-grounded AI tools. A writer's absence from it can shape whether they are found or correctly described wherever those systems are used. The stakes are sharpened in India, where the Sahitya Akademi, India's National Academy of Letters, maintains parallel literary-award categories across 24 recognized languages and treats them as formally equal. Digital infrastructure offers no such courtesy, so print parity guarantees no digital parity. This study puts that unevenness to a direct test: are Sahitya Akademi Award recipients verifiably present in Wikidata, language by language? Using the 2024 award cohort (Bengali's 2024 award remained undeclared at data collection, so its 2023 award stands in), two reviewers assessed all 24 writers under a six-tier protocol distinguishing three situations: a search that found nothing, a candidate found but never solidly confirmed, and an absence that independent evidence outside Wikidata actually supports. Twelve records were confirmed and three were probable; for nine, 37.5 percent, the award-winning writer had no plausible Wikidata candidate at all (Bodo, Kashmiri, Maithili, Manipuri, Marathi, Nepali, Rajasthani, Sindhi, and Urdu), and independent non-Wikimedia sources still confirm the award for eight of those nine writers, arguing for genuine coverage gaps over search failure rather than any judgment about the languages themselves. A supplementary four-year analysis of the 2020–2023 cohorts finds this same no-candidate outcome recurring for different award-winning writers across six of those languages (Maithili, Manipuri, Marathi, Nepali, Rajasthani, and Urdu), alongside a discovery process that fails in both directions, missing writers already present under a different name form and returning confident but wrong matches for colliding aliases. Read through Susan Leigh Star's infrastructure theory and the emerging field of Indian Postcolonial Digital Humanities, these findings do more than document an absence. Infrastructure ordinarily becomes visible mainly at points of breakdown, and this audit deliberately searches at that fringe; what it finds there locates the unevenness in the representational pipeline connecting Indian-language literary production to structured knowledge infrastructure, rather than in any judgment about which writers or literatures deserve to be found, though language-specific factors likely still play some role. Beyond a reusable audit protocol, the findings support a concrete institutional response: structured, openly licensed Sahitya Akademi award data, following models Wikimedia's own GLAM collaborations already provide elsewhere, would let print recognition reach the infrastructures that increasingly mediate whether a writer is found at all

// Source

View paper (DOI)Open access versionOpenAlexKnowledge Commons (Lakehead University)Published 2026-08-07

Authors: S Anas Ahmad

Institutions: Aligarh Muslim University