Kriyam TamperFlow: Benchmarking Document Tampering Detection under Realistic Document Processing Pipelines
Abstract
Kriyam TamperFlow is a benchmark for evaluating document tampering detection models under realistic document processing conditions. The benchmark comprises 1,050 source documents across six broad document categories, including medical, financial, invoice, application, educational, and administrative documents. The dataset contains 350 authentic and 700 tampered documents, with tampering types including copy-move, splicing, text replacement, and inpainting. Each source document is evaluated across three processing tiers (C0, C2, and C4), resulting in 3,150 evaluation images. The benchmark provides per-region bounding-box annotations for tampered regions and evaluates both tampered-region localization and document-level detection. The benchmark is designed to study how document processing operations such as JPEG recompression, printing, scanning, and photocopy simulation affect the forensic signals used by existing document tampering detection models. Five representative state-of-the-art detectors are evaluated across the three processing tiers using Region-F1, Doc-AUC, Doc-AUPRC, FPR@TPR, and compression robustness metrics. The results highlight a substantial gap between localization performance and document-level ranking robustness, demonstrating the need for evaluation protocols that better reflect real-world document workflows. The dataset, evaluation code, prediction format, and benchmark resources are released to support reproducible research in document forensics and trustworthy document intelligence.
// Source
Authors: Avishek Jana, Swati Kumari
Institutions: Googol Technology (China)