Skip to main content
Intermediate-Evidence Substitution: A Benchmark Design Methodology for Isolating Numerical Derivation from Protocol Execution in Deterministic Climate Verification

Intermediate-Evidence Substitution: A Benchmark Design Methodology for Isolating Numerical Derivation from Protocol Execution in Deterministic Climate Verification

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Supplementary Files

Authors

Aaditya Anand Thokal 

Abstract

We introduce \textbf{Intermediate-Evidence Substitution (IES)}, a benchmark design methodology that changes evidence granularity while preserving task semantics: decisive aggregate statistics in benchmark evidence are replaced with the raw intermediate values from which they are derived, while task labels, decision protocols, thresholds, and evaluation criteria remain unchanged. IES functions as a diagnostic probe\,---\,any change in model accuracy between the original and transformed evidence reveals how model performance depends on information representation, whether that change is a decrease, an increase, or neither. We instantiate IES on \textbf{ClimateTwinBench}, a 240-question deterministic hypothesis-verification benchmark built over India Meteorological Department (IMD) gridded climate data, spanning six verification categories. Comparing the original evidence design (V1) to its IES-transformed counterpart (V2), we observe statistically significant accuracy changes for Claude Sonnet~4.6 (100.00\%~$\rightarrow$~96.25\%, McNemar $p=0.008$, $\chi^2=7.111$) and Gemini~3.5 (32.50\%~$\rightarrow$~45.42\%, $p=0.008$, $\chi^2=7.087$), while the change for DeepSeek Chat is not statistically significant ($p=0.248$). ChatGPT improves markedly (+25.83~pp, $p\approx4.8\times10^{-11}$); all such outcomes\,---\,in either direction\,---\,are informative empirical results of the diagnostic. To separate numerical derivation from protocol execution, we conduct a structured derivation evaluation in which five models explicitly output predicted intermediate metrics, followed by oracle protocol recovery and cascade analysis distinguishing absorbed from propagated derivation errors. Claude Sonnet~4.6, Claude Opus~4.6, and Gemini~3.5 achieve 100.00\% oracle-recovered protocol accuracy; DeepSeek Chat achieves 99.58\%, with its single propagated failure attributable to connected-component counting in the Spatial Coherence category. GPT-OSS-120B exhibits failures in both metric derivation (4.17\%) and oracle-recovered protocol accuracy (42.92\%), preventing the clean separation between error sources that characterizes the frontier models. For the evaluated frontier models on ClimateTwinBench, numerical derivation\,---\,not protocol execution\,---\,is the dominant source of residual error, a finding that has direct implications for where benchmark evaluation effort should be directed.

DOI

https://doi.org/10.31223/X53B7F

Subjects

Computer Sciences, Oceanography and Atmospheric Sciences and Meteorology, Physical Sciences and Mathematics

Keywords

Benchmark Design, Evaluation Methodology, Numerical Reasoning, Climate NLP, LLM Evaluation, Benchmarking, Deterministic Verification, Intermediate-Evidence Substitution, ClimateTwinBench

Dates

Published: 2026-07-29 10:28

Last Updated: 2026-07-30 05:23

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
None.

Data Availability:
All code, data, and evaluation scripts are publicly available at: https://github.com/aadityat23/climateTwinIndia

Metrics

Views: 16

Downloads: 3