AI & Computingpreprint2026-08-02

IIfSLM: An Intelligence Index for Small Language Models — A Contamination-Resistant, Community-Driven Benchmark Suite for the 0.5B–3B Parameter Range

Open access0 citations

Abstract

We introduce IIfSLM (Intelligence Index for Small Language Models), an open benchmark suite targeting language models in the 0.5B–3B parameter range. IIfSLM addresses two gaps in existing small-model evaluation practice: the absence of a benchmark suite specifically scoped to this parameter range, and the well-documented risk that widely-used benchmarks (GSM8K, HumanEval, ARC-Challenge) have their exact phrasing present in large-scale web-scraped pretraining corpora, allowing verbatim memorization to be mistaken for reasoning capability. IIfSLM's core methodology is to rephrase question surface form using a large language model while always preserving the original, human-verified reference answer, so that correctness never depends on the rephrasing model's own reliability. We describe the dataset construction pipeline, the domain-specific safeguards used to prevent rephrasing from silently corrupting a question's answer key, our evaluation methodology, and an important finding regarding the sensitivity of measured accuracy to generation token budget. We report initial results for one model, Atomight-V2.5-1.7B, on one of three released domains, and describe IIfSLM as an open, community-scored leaderboard accepting further submissions.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-02

Authors: NovatasticRoScript