Calibrate
We run our 7 custom test batteries on 21 open-weight models we can afford — forming the empirical calibration ground truth.
STATION 02 — LLM BENCHMARK & STATISTICAL PROJECTION
Empirically measured scores on open-weight calibration LLMs. Statistically projected scores with published 95% Confidence Intervals for expensive closed frontier models.
ANATOMY OF A RUN
Provider API (OpenAI, Anthropic, Gemini, OpenRouter) or local Ollama / vLLM endpoint.
API keys stay in your browser session only. Zero server-side key storage.
Reasoning, Coding, Math, Security, Hallucination, Speed, and Cost.
Real-time SSE token stream with deterministic rule graders and blind judge fallback.
Measured runs feed our calibration set, improving OLS frontier projections.
FRONTIER DEX · 7 BENCHMARK CATEGORIES
Click any category to sort. Measured scores represent ground truth; Projected scores include 95% Confidence Intervals.
| Rank | Model | Provider | Composite | Reasoning | Coding | Math | Security | Hallucination | Speed | Cost / 1M | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Rank 01 | GPT-5 (Preview / o1-max) | OpenAI | 94.7% | 94.1% | 96.3% | 95.2% | 92.4% | 95.5% | 38 tps | $15.00 | ~94.7% ±4.1% |
| Rank 02 | Claude 3.7 Sonnet (Hybrid) | Anthropic | 93.6% | 93.2% | 95.8% | 92.5% | 91.8% | 94.8% | 48 tps | $9.00 | ~93.6% ±4.2% |
| Rank 03 | Gemini 2.5 Pro | 92.8% | 92.5% | 93.4% | 94.0% | 90.5% | 93.6% | 56 tps | $3.12 | ~92.8% ±4.3% | |
| Rank 04 | DeepSeek R1 | DeepSeek | 92.4% | 93.6% | 92.4% | 95.8% | 88.0% | 92.0% | 34 tps | $2.19 | 92.4% Measured |
| Rank 05 | Claude 3.5 Sonnet | Anthropic | 91.2% | 90.8% | 94.5% | 89.2% | 90.0% | 91.5% | 68 tps | $9.00 | ~91.2% ±4.4% |
| Rank 06 | DeepSeek V3 | DeepSeek | 87.9% | 88.4% | 89.2% | 86.0% | 84.0% | 88.0% | 62 tps | $0.28 | 87.9% Measured |
| Rank 07 | Llama 3.1 405B Instruct | Meta | 86.1% | 86.0% | 87.5% | 83.2% | 88.0% | 86.0% | 22 tps | $2.00 | 86.1% Measured |
| Rank 08 | Qwen 2.5 72B Instruct | Alibaba / Qwen | 85.4% | 85.2% | 87.0% | 82.4% | 84.0% | 88.0% | 46.5 tps | $0.38 | 85.4% Measured |
| Rank 09 | Qwen 2.5 Coder 32B | Alibaba / Qwen | 84.2% | 80.4% | 90.5% | 78.0% | 84.0% | 80.0% | 78 tps | $0.20 | 84.2% Measured |
| Rank 10 | Llama 3.3 70B Instruct | Meta | 81.8% | 82.1% | 83.4% | 79.6% | 80.0% | 84.0% | 52 tps | $0.35 | 81.8% Measured |
| Rank 11 | Codestral 2501 | Mistral | 80.9% | 76.0% | 88.0% | 74.5% | 80.0% | 78.0% | 92 tps | $0.60 | 80.9% Measured |
| Rank 12 | Mistral Large 2411 | Mistral | 79.9% | 79.5% | 81.2% | 76.8% | 80.0% | 82.0% | 45 tps | $4.00 | 79.9% Measured |
| Rank 13 | Phi-4 (14B) | Microsoft | 79.0% | 78.4% | 79.5% | 81.0% | 76.0% | 80.0% | 72 tps | $0.15 | 79.0% Measured |
| Rank 14 | Llama 3.1 8B Instruct | Meta | 67.2% | 68.2% | 66.0% | 64.0% | 68.0% | 72.0% | 115 tps | $0.05 | 67.2% Measured |
| Test | Your Model Endpoint · Test any Ollama or BYOK provider | Launch Instant Benchmark Run → | BYOK | ||||||||
METHODOLOGY · TRANSPARENT BY DEFAULT
No black box — every formula, weight, and standard error is published and verifiable.
We run our 7 custom test batteries on 21 open-weight models we can afford — forming the empirical calibration ground truth.
Those models have published external benchmark scores (GPQA, SWE-bench, AIME, MMLU-Pro). We fit an explainable multi-variable linear model.
Frontier models' public scores enter the formula. Out comes a Projected Score with published 95% Confidence Intervals (±SE) — never false precision.
BENCHMARK SUITE · 7 LOCKED DIMENSIONS
Focused on what matters to builders: correctness, security, truthfulness, and economics.
ROADMAP
Phased development roadmap aligned with open scientific benchmarking standards.