Climate & Environmentpreprint2026-08-04

Easy-Negative Inflation in Wildfire-Risk Modeling: Honest Evaluation, Seasonal Recalibration, and a Public Prospective Benchmark from an Operational System

Open access0 citations

Abstract

We report on the evaluation and operation of Tutela Ignis, a publicly deployed machine-learning wildfire-risk system for mainland Portugal that issues daily ignition-risk scores on a 1 km grid (88,527 continental cells) from a gradient-boosted ensemble over 63 tabular features drawn from more than a dozen open data sources. Rather than a new state-of-the-art number, our contributions are methodological and operational, grounded in measurements that most offline studies cannot make. First, we show that our own headline metric was inflated: with negatives sampled from non-fire-prone grid locations, out-of-time AUC reaches 0.98–0.99 by answering the susceptibility question ("is this a place that burns?") rather than the occurrence question ("will it burn now?"). Under a hard-negative design in which negatives share the positives' locations, honest skill is AUC ≈ 0.83, squarely in the published 0.7–0.85 band for date-specific fire occurrence. We document how naive attempts to harden negatives silently re-inflate the metric unless both the spatial domain and the seasonal date distribution are matched. Second, we show that stratified (information-content) and out-of-time (distribution-shift) evaluations disagree structurally, that the stratified meter is blind to time-invariant covariates by construction, and that class-balanced validation corrupts operating-point selection through three faces of one prevalence trap. Third, we describe a deployed seasonal-recalibration architecture (per-month isotonic calibration plus a seasonal model selector) that recovers a spring recall collapse caused by summer-dominated training data (+2.6% F1, +3.4 points of spring F1 operationally), and show that three simpler alternatives (per-month thresholds, training-window sliding, wider routing) all underperform it. Fourth, we report a bank of controlled negative results (vegetation indices, live fuel moisture, drought indices, human-ignition layers, sample re-weighting) suggesting the model sits at the skill ceiling of the available tabular sources. Finally, we introduce what we believe is the first continuously updated, fully public prospective benchmark of an ML fire-risk product against official forecasts: since 19 May 2026, every morning's national prediction is archived before any fire occurs and compared daily against real ignitions (ICNF), the official Portuguese municipal fire-danger product (IPMA RCM), and the European FWI (EFFIS). Over the first 71 days (3,661 ignitions), at each official product's own ignition coverage our 1 km map requires 1.8–3.4× less flagged area; on fires that grew beyond 1 ha the gap narrows to 1.3–2.6×. The benchmark also shows that fire-danger products are below random at ignition-location coverage per unit area, sharpening the distinction between the two tasks, and that our served scores rank well but overstate per-cell-day event frequency by 30–100×, a gap we publish and label rather than hide. All benchmark data are public and update nightly.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-04

Authors: Martim Ramos

Institutions: Neural Signals (United States)