Health & Medicinepreprint2026-08-13

SpaceFlight-Clinical-AI-Benchmarks: An Edge-Native Evaluation Framework for NASA ACHS and Deep Space Operations

Open access0 citations

Abstract

The integration of Large Language Models (LLMs) into the NASA Autonomous Crew Healthcare System (ACHS) presents unprecedented opportunities for deep space medical support. However, standard LLM evaluation benchmarks fail to account for the unique operational constraints of spaceflight, namely extreme communication latency, edge-native hardware limitations, and the severe physiological hazards of the aerospace environment. In this paper, we introduce the SpaceFlight-Clinical-AI-Benchmarks, the first open-source, offline evaluation framework designed specifically for aerospace medicine. We present a novel evaluation methodology featuring a hard-gate safety multiplier that strictly penalizes AI models for recommending contraindications in high-stakes scenarios such as microgravity-induced arrhythmias and acute radiation sickness. Our pipeline demonstrates that by evaluating models exclusively in zero-connectivity edge environments, we can accurately measure their viability for autonomous clinical decision support during deep space transit.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-13

Authors: Mohammad Al Kharabsheh