STATION 02 · SCIENTIFIC TRANSPARENCY
Evaluation Methodology &
Dual-Tier OLS Projection Architecture
NeurionForge Eval is built on a single uncompromising principle: the whole product is trust. We openly disclose our calibration data, regression weights, standard errors, and cryptographic task hashes.
⚖️ The Golden Rule of Credibility: Measured vs. Projected
A Projected score must never render identically to a Measured one. Measured runs reflect empirical ground truth executed directly on our custom suites with BYOK or local tunnels. Projected scores reflect an Ordinary Least Squares (OLS) prediction with explicit 95% Confidence Intervals (±SE).
1. Ordinary Least Squares (OLS) Projection Engine
We fit a multi-variable linear model per category using our 21-model calibration dataset:ProjectedScore = b + (w₁·GPQA) + (w₂·SWE-bench) + (w₃·AIME) + (w₄·MMLU-Pro) + (w₅·NormElo)
Reasoning Regression Model
Reasoning balances GPQA Diamond and AIME scientific logic with MMLU-Pro knowledge coverage.
2. Ground-Truth Calibration Dataset (21 Open LLMs)
Public benchmark scores (freely verifiable) alongside our empirical measured runs:
| Model | Provider | GPQA | SWE-bench | AIME | MMLU-Pro | Arena Elo | Meas. Reason | Meas. Code |
|---|---|---|---|---|---|---|---|---|
| Qwen 2.5 72B Instruct | Qwen / Alibaba | 49.5% | 41.9% | 50% | 62.8% | 1283 | 85.2% | 87% |
| Llama 3.3 70B Instruct | Meta | 50.7% | 39.6% | 43.3% | 58% | 1260 | 82.1% | 83.4% |
| DeepSeek R1 | DeepSeek | 71.5% | 49.2% | 79.8% | 84% | 1362 | 93.6% | 92.4% |
| DeepSeek V3 | DeepSeek | 59.1% | 42% | 39.2% | 67.1% | 1328 | 88.4% | 89.2% |
| Qwen 2.5 Coder 32B | Qwen / Alibaba | 44.8% | 46.5% | 41.2% | 59% | 1240 | 80.4% | 90.5% |
| Mistral Large 2411 | Mistral | 48% | 38% | 35% | 55.4% | 1248 | 79.5% | 81.2% |
| Llama 3.1 405B Instruct | Meta | 51.1% | 41.5% | 45% | 63.5% | 1285 | 86% | 87.5% |
| Codestral 2501 | Mistral | 42% | 44.2% | 38% | 52% | 1220 | 76% | 88% |
| Llama 3.1 8B Instruct | Meta | 32.8% | 18.7% | 22% | 42% | 1145 | 68.2% | 66% |
| Qwen 2.5 14B Instruct | Qwen / Alibaba | 40.2% | 29.5% | 33% | 51.5% | 1205 | 75% | 77.2% |
Showing 10 of 21 calibration models. View full calibration dataset in source repo under lib/projection/calibrationData.ts.
3. Cryptographic Task Integrity (SHA-256)
Adopted from BridgeBench transparency standards: benchmark tasks are open and public, while canonical reference answers remain cryptographically hashed until task pack retirement. Any retroactive modification to task prompts or answers causes an immediate SHA-256 mismatch.
Security Task Hashes (sample)
Hallucination Task Hashes (sample)
4. Append-Only Run Journal & Grader Versioning
Every TestRun records an immutable scoringVersion (e.g. v1). When evaluation algorithms or rule heuristics improve, past empirical runs are preserved in their original state. This prevents retrospective revisionism and maintains continuous scientific reproducibility.