PHILIA OS Phase 2 · 105th DOI Results Report: TextAnalyzer CH7 — Local LLM-Based External Factual Alignment Instrument (Main Run v1.1, NOT_ESTABLISHED)
Abstract
This record contains the results report for the PHILIA OS Phase 2 104th pre-registered study (TextAnalyzer CH7), registered here as the 105th Zenodo DOI. The study examined whether a local LLM-based external factual alignment instrument can be established using a 4-arm evidence manipulation paradigm. Verdict: NOT_ESTABLISHED (pre-registration: 10.5281/zenodo.21737175 v1.7). Two of four pre-declared establishment conditions were not met. C1 (EFR_explicit = 1.000) and C3 (LLM−rule accuracy advantage +43.9 percentage points, Newcombe 95% CI [0.353, 0.515]) were satisfied; C2 (A vs B separation rate = 33.3%, threshold 0.80) and C4 (L3 4-arm accuracy = 69.0%, threshold 0.70) were not. No forking or re-execution was performed. Root cause: Post-hoc item-level analysis confirms that the result is most consistent with a corpus-specification limitation. A_true_ev evidence was drawn from Wikipedia REST API summary text (first ~400 characters), which systematically lacks claim-supporting sentences. Of the 9 A=CONTRADICTION cases, 6 are evidence-date-conflict cases sharing the same root as the 21 NEUTRAL cases; genuine cutoff-activation candidates are ≤3 items (L3_007/009/011). Design Input ① (claim-supporting sentence extraction) resolves both categories. Positive findings: C1=1.000 (45/45 B arm detection); C3 +43.9%p advantage over rule-based baseline (44/57 C arm NUMBER_CONFLICT misfires correctly returned as NEUTRAL by LLM); EFR_logical 75% (9/12), 3× improvement over Pilot v5. The CH7 comparator path is supported as a promising direction for further investigation. CH7-v2 design inputs: (1) Replace A_true_ev with claim-supporting sentence extraction — resolves NEUTRAL (21) and ev-date-conflict CONTRADICTION (6); must be pre-registered. (2) Add entity anchor check before NUMBER_CONFLICT rule trigger. (3) Additional L3 pre-screening filter for ≤3 possible cutoff-activation cases — lower priority. Corpus: 57 items (L1×20 + L3×25 + Hard×12). Model: qwen2.5:14b (Ollama 0.32.5, temperature=0). Total LLM calls: 228. Hardware: AMD Ryzen 9 9800X3D + RTX 4080 SUPER 16GB, Windows. Gate-0 adversarial review (Opus / Claude Opus, Anthropic): UNCONDITIONAL PASS on pre-registration. Pre-registration violations: 0. Files included: PHILIA_104th_CH7_Results_Report_v1_0.pdf — full results report (19 pages) philia_textanalyzer_ch7_main_v1_1_results.json — raw results philia_textanalyzer_ch7_main_v1_corpus.json — corpus (adapter output) Trinity AI Research Team | Konyang University | 2026-08-02Measurement is life. Description, not proof. Pre-registration protects us. — PHILIA OS | 0∞1∞0.5∞
// Source
Authors: 신두섭