Author

Boss Ohkubo

0 works0 citationsORCID

Recent research

  • AI & ComputingOpen access

    Who checks the answer key? Ground-truth defects in eight professional-practice benchmarks

    Benchmarks are known to contain incorrect ground truth, and models are already used to find it. That work concerns sparse annotation errors on individual instances, found statistically or by re-annotation. We report a class it does not reach: systematic errors in the program that...

    Zenodo (CERN European Organization for Nuclear Research)2026-08-110 citationsDOI
  • AI & ComputingOpen access

    Who checks the answer key? Ground-truth defects in eight professional-practice benchmarks

    Benchmarks are known to contain incorrect ground truth, and models are already used to find it. That work concerns sparse annotation errors on individual instances, found statistically or by re-annotation. We report a class it does not reach: systematic errors in the program that...

    Zenodo (CERN European Organization for Nuclear Research)2026-08-090 citationsDOI