FORGE is built to investigate this question. It generates every problem at runtime from a seed, grades answers computationally, and produces a reproducible score that anyone can verify or challenge. Current results should be treated as preliminary.
Whether large language models do something that resembles reasoning over internalized structure, or succeed primarily through pattern-matching against training data, is not settled. Researchers disagree. The evidence is mixed. FORGE does not resolve this — it tries to produce cleaner evidence than what currently exists.
What we hypothesize: A model's performance profile across difficulty tiers is informative about whether it has internalized mathematical structure or is primarily retrieving from training distribution. We do not claim FORGE scores prove either interpretation.
Mechanistic interpretability research suggests models develop structured internal representations during training. But several papers have also documented that model accuracy collapses — rather than degrades gracefully — when pushed slightly beyond training distribution. Both observations are real. FORGE is designed to produce data that bears on which pattern dominates, and at what difficulty level.
Falsifiability: The hypothesis would be evidence against if models scoring highly on hard tiers suffer catastrophic degradation when mathematical structure is held constant but surface framing is inverted. We have not run this experiment systematically. It is a direction for future work.
FORGE generates every problem procedurally from a seed using SHA-256 derived RNGs. No static question bank. Same seed produces identical problems every time, on any machine. FORGE uses procedural generation to make verbatim contamination statistically unlikely at normal evaluation scales. Individual category problem spaces vary significantly in size. Structural similarity to training data cannot be excluded.
All evaluation uses deterministic computational verification rather than LLM judges. This eliminates LLM judge bias but introduces its own edge cases, documented per category. The grading logic is inspectable and consistent — the same input always produces the same grade. Known grading limitations are flagged in the category status table.
| Engine | Used for |
|---|---|
| SymPy | Symbolic equivalence, differential equations, logic reduction |
| NumPy | Matrix operations, DFT, numerical precision |
| python-chess | Game state validation, legal move verification |
Every category is assessed against three criteria: problem space size (too large for feasible exhaustive contamination), grader self-consistency at 100%, and empirically validated difficulty scaling. Categories that fail any criterion are flagged.
| Category | Status | Reason |
|---|---|---|
| Polynomial Roots | FLAGGED | Problem space of 62.8M. Exhaustive contamination feasible on consumer hardware |
| RSA Arithmetic | FLAGGED | Tests computational limits, not reasoning. Under review for redesign |
| Boolean Minimization | RESEARCH | Known grader bug at difficulty 4-5 |
| Chess Mate-in-N | RESEARCH | Generation pool constraints at high difficulty |
| Algebra Groups | RESEARCH | Extremely slow generation at high difficulty |
| Quantum Amplitudes | RESEARCH | Answer parser fragile on non-standard notation |
| Shannon Entropy | RESEARCH | Borderline tolerance cases in cross-model testing |
| Jordan Normal Form | RESEARCH | Degenerate eigenvalue ambiguity at difficulty 5 |
| Formal Grammars | RESEARCH | String matching edge cases in grading |
| Algorithmic Trace | RESEARCH | Output formatting edge cases in grading |
15 categories are [FORGE CERTIFIED]. 8 are [RESEARCH]. 2 are [FLAGGED] (scores excluded from main FORGE score). Full details in the README.
A simple mean accuracy conflates performance at easy tiers (where training data overlap is likely) with performance at hard tiers (where it is less likely). FORGE weights harder tiers exponentially more, so a model that collapses at difficulty 3 scores differently from one that holds through difficulty 5.
The Cliff Index is the first difficulty tier where accuracy drops more than 30% from the previous tier. It is an observation about where performance degrades, not a proof of what causes the degradation. It is the most diagnostically useful single number FORGE produces.
| Metric | What it measures | What it does not prove |
|---|---|---|
| FORGE Score | Weighted accuracy across all categories and tiers | That the model has an internal world-model |
| Interpolation Score | Accuracy on easy/medium tiers | That these problems were in training data |
| Extrapolation Score | Accuracy on hard/expert tiers | That these problems were absent from training data |
| Cliff Index | First tier where accuracy drops >30% | The cause of the drop |
A model with high interpolation and low extrapolation scores is consistent with retrieval-dominant behavior. A model that maintains accuracy across tiers is consistent with generalization. Neither interpretation is proven by the score alone.
AI labs face a quiet incentive problem. A model is released. It scores well on public benchmarks. The scores become marketing. Then, quietly, the model is updated. The published scores no longer reflect what users are actually running.
Fixed benchmarks cannot distinguish between "the model got smarter" and "the model was tuned toward this specific benchmark distribution."
Because every FORGE run is governed by a published seed and every question is generated deterministically at runtime, a score is reproducible. Anyone can rerun the same seed against the same endpoint months later and compare.
FORGE does not accuse anyone of nerfing. It makes the question more tractable.
FORGE accepts seeds with a 256-bit hash space (SHA-256). The number of genuinely distinct questions per category varies. Question set uniqueness at the 25-category level is substantially larger than any individual category space.
| Metric | Value |
|---|---|
| Seed input space | 2256 ≈ 1.2 × 1077 |
| Time to exhaust (1T hashes/sec) | ~3.7 × 1051 years |
| Age of universe | 1.4 × 1010 years |
Note: The above describes the seed space. Individual category problem spaces vary significantly. See Category Status for details.
These are the specific problems FORGE was designed to address, and the tradeoffs involved in each design decision.
| Problem | Design decision | Tradeoff |
|---|---|---|
| Static question banks contaminate pre-training | Procedural generation from seed at runtime | Structural similarity to training data cannot be excluded. Individual category problem spaces vary in size |
| Multiple-choice formats are gameable | Free-form answers graded computationally | Prompt sensitivity affects scores; format non-compliance penalizes correct answers |
| LLM judges introduce hallucination and inconsistency | SymPy, NumPy, python-chess for grading | Eliminates LLM judge bias but introduces grader-specific edge cases, documented per category |
| Fixed benchmarks cannot detect model version changes | Seeded, reproducible runs with published timestamps | Interpreting score changes requires controlled conditions we do not enforce |
| Human-curated hard problems are expensive to scale | Parametric difficulty via mathematical parameters | Parametric difficulty may not map cleanly to human difficulty perception |
During evaluation on seed 42, qwen/qwen3-coder-flash produced confident multi-step entropy calculations using incorrect log2 values. Example: log2(107) computed as 6.7175 versus correct 6.7415 (error −0.024). The model used log2(a/b) = log2(a) − log2(b) with systematically wrong log2 values for integers 103, 107, 221, 309, 321, 332, 347. The reasoning trace showed no uncertainty at the point of error. FORGE's deterministic grader flagged all 10 entropy questions as incorrect. Raw response (first failure, truncated): "log2(50/321) = log2(50) - log2(321) = 5.6439 - 8.3219 = -2.6780" (actual: -2.6826).
On seed 17017656696087371159 (nano mode, 5 questions, Shannon Entropy difficulty 3), 8 frontier models were evaluated. Five passed (DeepSeek V4 Pro, GPT-5.5, Claude Opus 4.7, Nemotron 550B, MiniMax M3). Three failed on the same entropy question (expected 2.7037): Claude Opus 4.8 extracted 2.7038, Claude Opus 4.6 extracted 2.7036 — both within 4th-decimal rounding error of the correct value. Claude Sonnet 4.6 extracted 2.7056, a larger deviation traced to an incorrect intermediate logarithm: ln(0.072650) computed as −2.6439 versus correct −2.6226. Notably, Opus 4.7 passes while 4.6 and 4.8 fail — a non-monotonic pattern across versions. Sonnet 4.6 is the only model with a demonstrably wrong logarithm value; the other two failures are borderline tolerance cases.
Note: The Researcher Tool requires the
Python backend running locally. It is a frontend UI that connects to
localhost:7860 — without the backend installed and running, the page will
show errors when you try to run a benchmark.
| Mode | Questions | Time | Est. Cost |
|---|---|---|---|
| nano | 5 | ~30s | <$0.01 |
| quick | 250 | ~5 min | ~$0.10 |
| standard | 2,500 | ~1 hr | ~$1.00 |
| full | 10,000 | ~4 hr | ~$4.00 |
Cost estimation settings — Adjust to match your model's pricing. Defaults reflect typical OpenRouter rates (e.g., GPT-4o: $5/$15, GPT-4o-mini: $0.15/$0.60, Llama-3.1-405B: $3/$9).