STATION 02 · SCIENTIFIC TRANSPARENCY

Evaluation Methodology &
Dual-Tier OLS Projection Architecture

NeurionForge Eval is built on a single uncompromising principle: the whole product is trust. We openly disclose our calibration data, regression weights, standard errors, and cryptographic task hashes.

⚖️ The Golden Rule of Credibility: Measured vs. Projected

A Projected score must never render identically to a Measured one. Measured runs reflect empirical ground truth executed directly on our custom suites with BYOK or local tunnels. Projected scores reflect an Ordinary Least Squares (OLS) prediction with explicit 95% Confidence Intervals (±SE).

1. Ordinary Least Squares (OLS) Projection Engine

We fit a multi-variable linear model per category using our 21-model calibration dataset:
ProjectedScore = b + (w₁·GPQA) + (w₂·SWE-bench) + (w₃·AIME) + (w₄·MMLU-Pro) + (w₅·NormElo)

Reasoning Regression Model

Reasoning balances GPQA Diamond and AIME scientific logic with MMLU-Pro knowledge coverage.

Goodness of Fit
R² = 0.89
95% Conf. Interval
±4.6%
w₁ (GPQA Diamond)
0.32
w₂ (SWE-bench)
0.22
w₃ (AIME 2024)
0.24
w₄ (MMLU-Pro)
0.16
w₅ (Arena Elo)
0.06
b (Intercept)
+4.2

2. Ground-Truth Calibration Dataset (21 Open LLMs)

Public benchmark scores (freely verifiable) alongside our empirical measured runs:

ModelProviderGPQASWE-benchAIMEMMLU-ProArena EloMeas. ReasonMeas. Code
Qwen 2.5 72B InstructQwen / Alibaba49.5%41.9%50%62.8%128385.2%87%
Llama 3.3 70B InstructMeta50.7%39.6%43.3%58%126082.1%83.4%
DeepSeek R1DeepSeek71.5%49.2%79.8%84%136293.6%92.4%
DeepSeek V3DeepSeek59.1%42%39.2%67.1%132888.4%89.2%
Qwen 2.5 Coder 32BQwen / Alibaba44.8%46.5%41.2%59%124080.4%90.5%
Mistral Large 2411Mistral48%38%35%55.4%124879.5%81.2%
Llama 3.1 405B InstructMeta51.1%41.5%45%63.5%128586%87.5%
Codestral 2501Mistral42%44.2%38%52%122076%88%
Llama 3.1 8B InstructMeta32.8%18.7%22%42%114568.2%66%
Qwen 2.5 14B InstructQwen / Alibaba40.2%29.5%33%51.5%120575%77.2%

Showing 10 of 21 calibration models. View full calibration dataset in source repo under lib/projection/calibrationData.ts.

3. Cryptographic Task Integrity (SHA-256)

Adopted from BridgeBench transparency standards: benchmark tasks are open and public, while canonical reference answers remain cryptographically hashed until task pack retirement. Any retroactive modification to task prompts or answers causes an immediate SHA-256 mismatch.

Security Task Hashes (sample)

sec-sqli-01 (sql_injection):
b7e452a832d20739c9431e2170ba37937510d9cebcfeaebe34a7eb06bc8dc6aa
sec-sqli-02 (sql_injection):
5f8a0328990ec3beaa79b76c8cb81a2ef1d596bf7cfd08226e6d1c448fa5629c
sec-sqli-03 (sql_injection):
9742a3cf283d6a978280f971b31a316ec1be87fa9d0bb5697669bb15993a405e
sec-sqli-04 (sql_injection):
d831ac8be87c0827ea80f089691b10ec87c88002bb5e8e89cbff7415442be772

Hallucination Task Hashes (sample)

hal-cite-01 (invented_citation):
6d0b67bf6fc0f15cfa959a4bb3a4ff6ae4a070eb374526df6d639b70bb122754
hal-cite-02 (invented_citation):
235b2e67cbdb14a4c585c5dc96ce1723bb2e1858c142eefad2488a0333246eb3
hal-cite-03 (invented_citation):
f697ca6faecf577cf38198f3ce820063ceec43194a73752e2553b924df072eb0
hal-cite-04 (invented_citation):
ad6a287a2a07c57ce4d983416824f9f743c35b6bbf202a0a2569e5d48380e227

4. Append-Only Run Journal & Grader Versioning

Every TestRun records an immutable scoringVersion (e.g. v1). When evaluation algorithms or rule heuristics improve, past empirical runs are preserved in their original state. This prevents retrospective revisionism and maintains continuous scientific reproducibility.