AI & Computingpreprint2026-08-26

Tri-Modal Calibration for Road Network Conflation: Four-State Labelling, Dirichlet Multi-Class Probabilities, and Conformal Coverage Guarantees

Open access0 citations

Abstract

Extended version. This is the full-length (13-page) version of a poster paper accepted to the 34th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26), Riverside, CA, USA, 3–6 November 2026. A 4-page version appears in the SIGSPATIAL '26 proceedings; this deposit carries the proofs, ablations, and replication results the page limit excluded. Partial-overlap road-network conflation matches admit a calibration-free geometric address: a shape-dispersion ratio bounded by sharp Cauchy–Schwarz inequalities and a Mean Closest-Point Distance (MCPD) dominated by the directional Hausdorff distance. We use this geometric core to expose a calibration failure mode in production pipelines whose per-match scores are read as probabilities. On 992 hand-labelled TomTom→KSJ matches from a Niigata three-city polygon under a four-state schema (wrong, partial, full, uncertain) at partial-overlap threshold τ = 0.60, the class-conditional score is tri-modal: the [0.70, 0.85) confused band is 4.7% wrong, 36.8% partial, 58.5% full. The strict-vs-lenient Expected Calibration Error (ECE) gap is 0.245 vs 0.019 under a PAVA-isotonic baseline. Beyond the geometric core (C0), four calibration results follow. (C1) The strict-vs-lenient ECE decomposition makes the failure mode measurable on a single corpus. (C2) First Dirichlet multi-class calibration of conflation scores (novelty scope delimited in the Related Work section): 3-class class-wise ECE 0.018, 25% below the uncalibrated baseline 0.024. (C3) Post-hoc PAVA-isotonic on the scalar (5-fold CV strict ECE 0.017 ± 0.004, 200 seeds) matches a six-feature direct-LR head (0.014 ± 0.003): calibrate the scalar, don't replace it. (C4) First distribution-free conflation coverage guarantee via split conformal prediction (same scope), mean empirical coverage 0.898 ± 0.026 at α = 0.10. Against a buffer-baseline of Hootenanny's match/miss/review algorithm on the same corpus, strict ECE is 4.6× smaller; an OSM→KSJ replication (Niigata, Osaka; N = 494) shows the C1 gap recurs (≥ 8×). The Spark + Sedona implementation reaches 51.5% KSJ coverage on 728,621 links in 97 minutes on a twenty-thread workstation. Data attribution. The evaluated target networks are open-licensed: OpenStreetMap and Overture under ODbL, and the KSJ road network under the MLIT National Land Numerical Information Download Site Content Terms of Use (Government Standard Terms of Use compliant, CC BY 4.0-compatible); use of KSJ data is not endorsed by the Ministry. The TomTom probe feed is used under a commercial licence that permits academic publication of derived results on the licensed three-city polygon but not redistribution of the probe-segment table. v2 (2026-08-26). This version repins every seed-dependent statistic to the 200-seed reporting protocol (previously an unpinned 10-seed set had leaked into several scripts): Table-1 binary ECEs 0.017 ± 0.004 and 0.014 ± 0.003 with Nadeau–Bengio p = 0.791; Hootenanny gaps 4.6× / 3.3×; schema-collapse range [0.008, 0.017]; τ-sweep endpoint 0.293; shape-dispersion AUC lift 0.044 ± 0.024; Venn–Abers 0.021 ± 0.003; Mondrian class-conditional coverage 0.898 ± 0.028. It also corrects the inter-rater audit corpus identity (a separate TomTom→KSJ sample, not the open-data corpus), the MCPD sample-count equation, the symmetric shape-dispersion definition, and the CCS concept ids. The 4-page ACM poster (DOI 10.1145/3841645.3843021) reports the corrected values.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-26

Authors: Genaro Peque

Institutions: Capgemini (Netherlands)