Beyond the gold standard: evaluating AI for subject indexing of Swedish LGBTQ + fiction
Abstract
Purpose This study evaluates the potential of large language models (LLMs) for subject indexing of Swedish LGBTQ + fiction using the dedicated QLIT thesaurus, based on Homosaurus. It examines both the quantitative performance and qualitative characteristics of AI-generated subject terms, with particular attention to thesaurus compliance and indexing policy. Design/methodology/approach Four general-purpose AI tools (ChatGPT, Claude, Gemini, and DeepSeek) were evaluated on a collection of 14 historical LGBTQ + literary works represented by 16 texts. Subject terms were generated from both metadata records and full texts and compared with existing indexing in the Queerlit database of Swedish LGBTQ + fiction. The metadata subset produced 518 AI-generated subject terms in total, while the full-text subset produced 590 terms. The analysis was based on a coding scheme, which was refined through several rounds of analysis until complete agreement was achieved. Findings Performance was substantially higher for metadata than for full texts, with average F1 scores of 44% and 12%, respectively. Across both datasets, fully incorrect terms constituted the largest category of generated outputs (37.5% for metadata and 56.1% for full texts), while fully correct terms accounted for only 30.0 and 8.3%, respectively. Qualitative analysis revealed recurring problems, including the assignment of broader concepts rather than, or in addition to, more specific concepts, failure to follow thesaurus definitions and indexing policies, confusion between general themes and LGBTQ + -specific themes, and the application of contemporary LGBTQ + concepts to historical texts. Poetry proved particularly challenging because themes were often implicit and open to interpretation. Although a small proportion of generated terms were judged potentially useful despite not appearing in existing indexing, no AI-generated terms were classified as fully correct additions beyond the current Queerlit metadata. Research limitations/implications The study focuses on a relatively specific collection of Swedish historical LGBTQ + fiction and a specialized controlled vocabulary. The findings nevertheless demonstrate the importance of complementing quantitative measures with qualitative evaluation and suggest that assessment of automated subject indexing should consider thesaurus definitions, indexing policies, and interpretive aspects of literary aboutness in addition to agreement with existing metadata. Practical implications The results indicate that current general-purpose LLMs are unlikely to be suitable as semi-automated indexing tools for historical LGBTQ + fiction without substantial human oversight. While AI systems may assist in identifying candidate concepts and additional access points, effective indexing continues to depend on expert knowledge of controlled vocabularies and indexing policy. Originality/value This study contributes to research on AI-assisted subject indexing by combining quantitative and qualitative evaluation of LLM-generated subject terms in a specialized LGBTQ + knowledge organization context. It demonstrates the limitations of gold-standard-based evaluation for fiction indexing and highlights the role of human expertise in applying controlled vocabularies to historically and culturally complex literary materials.
// Source
Authors: Koraljka Golub, Olof Falk, Siska Humlesjö
Institutions: University of Gothenburg, University of Borås, University of Arts and Industrial Design Linz