Safety Benchmark Gains Do Not Guarantee Safety Transfer : A Comprehensive Study of Fine-Tuning Small Language Model Safety Guards for High-Compliance and General Safety Domains
Abstract
Benchmark gains do not guarantee transfer. A safety guard — a small model that labels each request safe or unsafe before an assistant acts — is often selected by whichever fine-tune scores best on a public benchmark. That score mixes two different outcomes: learning the sources represented in fine-tuning and transferring to sources the guard never saw. We separate them by comparing every tuned guard only with its own pre-tuning checkpoint, on identical rows and at matched false-alarm budgets. Fine-tuning raises represented-source ranking by +0.32 macro-AP while held-out transfer moves −0.06: positive for the weakest base, negative for the strongest. Read at an equal alarm budget the reversal is sharper: at each base guard’s own false-alarm rate, tuned-guard transfer recall falls 0.517 → 0.217 and is worse on all four checkpoints. That is an ROC-point comparison at a common budget, not a deployable threshold (Appendix B.6). The claim is therefore not that fine-tuning always harms transfer; it is that a represented-benchmark gain, by itself, does not establish transfer. Because every deployment sentence here is about an alarm budget while the headline metric averages over the whole ranking, we re-read the same rows over FPR [0, 0.05]. No cell changes sign, so the direction holds — but macro-AP understates both halves of the trade: the transfer cost is −0.059 on macro-AP against −0.174 on partial AUC, and the represented gain +0.323 against +0.686 (Section 3.6). Two extensions and one remedy bound that result. A directional, non-confirmatory extension to six released purpose-built guards shows the same specialization pattern; KL-regularized SFT retains transfer only by giving back represented gain, and fails its registered non-inferiority margin. Averaging a base with its own adapter recovers most of the lost transfer (+0.076 vs. SFT) for one extra inference pass. Against a hosted frontier model, the ranking reverses by traffic regime: hosted leads by +0.109 [+0.077, +0.139] recall on unfamiliar prompts, while the tuned panel leads by +0.083 [+0.013, +0.157] on sources represented in its training manifest — a post-hoc aggregate over 3 purposively chosen corpora, which resampling the source set widens to [−0.019, +0.220]. The resulting workflow is simple: compare every tune with its own base, at a matched alarm budget, on represented and held-out sources; then apply a domain-specific gate and recalibrate. A frozen, dual-labeled mortgage benchmark illustrates the looks-safe-but-non-compliant stratum a general safety score cannot identify.
// Source
Authors: Reza Rahimi
Institutions: Namazi Hospital