STATION 02 — LLM BENCHMARK & STATISTICAL PROJECTION

We Built
The Tests.
See How
Frontier
Stacks Up.

Empirically measured scores on open-weight calibration LLMs. Statistically projected scores with published 95% Confidence Intervals for expensive closed frontier models.

ANATOMY OF A RUN

  1. 1
    Select model

    Provider API (OpenAI, Anthropic, Gemini, OpenRouter) or local Ollama / vLLM endpoint.

  2. 2
    Bring your key (BYOK)

    API keys stay in your browser session only. Zero server-side key storage.

  3. 3
    7 Locked Benchmarks

    Reasoning, Coding, Math, Security, Hallucination, Speed, and Cost.

  4. 4
    Live Streaming Runner

    Real-time SSE token stream with deterministic rule graders and blind judge fallback.

  5. 5
    Transparent Evidence

    Measured runs feed our calibration set, improving OLS frontier projections.

FRONTIER DEX · 7 BENCHMARK CATEGORIES

The Leaderboard

Click any category to sort. Measured scores represent ground truth; Projected scores include 95% Confidence Intervals.

Measured (Empirical)Projected (OLS ±SE)
RankModelProviderCompositeReasoningCodingMathSecurityHallucinationSpeedCost / 1MEvidence
Rank 01GPT-5 (Preview / o1-max)OpenAI94.7%94.1%96.3%95.2%92.4%95.5%38 tps$15.00
~94.7% ±4.1%
Rank 02Claude 3.7 Sonnet (Hybrid)Anthropic93.6%93.2%95.8%92.5%91.8%94.8%48 tps$9.00
~93.6% ±4.2%
Rank 03Gemini 2.5 ProGoogle92.8%92.5%93.4%94.0%90.5%93.6%56 tps$3.12
~92.8% ±4.3%
Rank 04DeepSeek R1DeepSeek92.4%93.6%92.4%95.8%88.0%92.0%34 tps$2.19
92.4% Measured
Rank 05Claude 3.5 SonnetAnthropic91.2%90.8%94.5%89.2%90.0%91.5%68 tps$9.00
~91.2% ±4.4%
Rank 06DeepSeek V3DeepSeek87.9%88.4%89.2%86.0%84.0%88.0%62 tps$0.28
87.9% Measured
Rank 07Llama 3.1 405B InstructMeta86.1%86.0%87.5%83.2%88.0%86.0%22 tps$2.00
86.1% Measured
Rank 08Qwen 2.5 72B InstructAlibaba / Qwen85.4%85.2%87.0%82.4%84.0%88.0%46.5 tps$0.38
85.4% Measured
Rank 09Qwen 2.5 Coder 32BAlibaba / Qwen84.2%80.4%90.5%78.0%84.0%80.0%78 tps$0.20
84.2% Measured
Rank 10Llama 3.3 70B InstructMeta81.8%82.1%83.4%79.6%80.0%84.0%52 tps$0.35
81.8% Measured
Rank 11Codestral 2501Mistral80.9%76.0%88.0%74.5%80.0%78.0%92 tps$0.60
80.9% Measured
Rank 12Mistral Large 2411Mistral79.9%79.5%81.2%76.8%80.0%82.0%45 tps$4.00
79.9% Measured
Rank 13Phi-4 (14B)Microsoft79.0%78.4%79.5%81.0%76.0%80.0%72 tps$0.15
79.0% Measured
Rank 14Llama 3.1 8B InstructMeta67.2%68.2%66.0%64.0%68.0%72.0%115 tps$0.05
67.2% Measured
TestYour Model Endpoint · Test any Ollama or BYOK providerLaunch Instant Benchmark Run →BYOK
Showing 15 models across 7 benchmark dimensionsInspect OLS Regression Formulas & Calibration Dataset →

METHODOLOGY · TRANSPARENT BY DEFAULT

How Projection Works

No black box — every formula, weight, and standard error is published and verifiable.

1

Calibrate

We run our 7 custom test batteries on 21 open-weight models we can afford — forming the empirical calibration ground truth.

2

Correlate (OLS)

Those models have published external benchmark scores (GPQA, SWE-bench, AIME, MMLU-Pro). We fit an explainable multi-variable linear model.

3

Project & Bound

Frontier models' public scores enter the formula. Out comes a Projected Score with published 95% Confidence Intervals (±SE) — never false precision.

View Full Methodology & Interactive Formula Breakdown →

BENCHMARK SUITE · 7 LOCKED DIMENSIONS

Final Evaluation Dimensions

Focused on what matters to builders: correctness, security, truthfulness, and economics.

ReasoningCodingMathematicsSecurityHallucinationSpeed (tok/s)Cost ($/1M)

ROADMAP

The Build Path

Phased development roadmap aligned with open scientific benchmarking standards.

Phase 0Done

Foundations

  • Monorepo + Next.js shell
  • Prisma / Supabase schema
  • 5 Provider clients (BYOK)
Phase 1Done

Core Loop (Measured)

  • Reasoning / Coding / Math suites
  • Real-time SSE execution engine
  • Live grading & model registry
Phase 2Live

7 Benchmarks + OLS Projection

  • Security & Hallucination suites
  • Speed & Cost computed columns
  • Ordinary Least Squares 95% CI
Phase 3Planned

Custom Evaluations

  • Dataset upload (JSONL/CSV)
  • LLM-as-a-Judge engine
  • Shareable report links (/custom/[shareId])
Phase 4Planned

Local Connector & Coverage

  • @neurionforge/eval-connector CLI
  • Auto-detect Ollama/vLLM/LM Studio
  • 2-file provider expansion pattern
Phase 5Planned

Community Prompt Library

  • Public prompt submission & voting
  • Ranked prompt search across 7 tests
  • One-click Custom Eval seeding
>_Dual-tier projection: Ordinary Least Squares (OLS) with 95% CI bounds
Next.js 16•TypeScript•Prisma ORM•Supabase Postgres•Zero-Key Server Storage