Under review · Under review
Under review (ARR), 2026
We introduce Cross-Model Semantic Entropy (CMSE), a training-free, plug-and-play method that quantifies disagreement across K heterogeneous LLMs in meaning-space rather than token-space. Using NLI-based bidirectional entailment to cluster answers from different models into semantic equivalence classes, we compute Shannon entropy over the cluster distribution. Our selective cascade (S-CMSE) routes each question through a quality-ordered pilot subset of models and escalates to the full ensemble only when entropy exceeds a threshold. S-CMSE outperforms majority voting by +4.6pp on TruthfulQA (95% CI: [0.820, 0.870] vs. [0.772, 0.827]) while calling only 2.0 models per question on average, a 50% reduction in API cost. CMSE entropy is a reliable uncertainty signal (AUROC = 0.765), and an oracle upper bound analysis reveals a 6.6pp exploitable gap above the best single model, providing a principled motivation for cross-model ensembling.