Active Research Phase — FORGE is not production ready. Current results are preliminary. Some categories have known limitations. See Category Status.
Open-source · MIT License · Active research phase

When a model solves a problem
it has never seen, is that
understanding — or a very good guess?

FORGE is built to investigate this question. It generates every problem at runtime from a seed, grades answers computationally, and produces a reproducible score that anyone can verify or challenge. Current results should be treated as preliminary.

Run FORGE View on GitHub
01 — The Hypothesis

A contested question, taken seriously.

Whether large language models do something that resembles reasoning over internalized structure, or succeed primarily through pattern-matching against training data, is not settled. Researchers disagree. The evidence is mixed. FORGE does not resolve this — it tries to produce cleaner evidence than what currently exists.

What we hypothesize: A model's performance profile across difficulty tiers is informative about whether it has internalized mathematical structure or is primarily retrieving from training distribution. We do not claim FORGE scores prove either interpretation.

Mechanistic interpretability research suggests models develop structured internal representations during training. But several papers have also documented that model accuracy collapses — rather than degrades gracefully — when pushed slightly beyond training distribution. Both observations are real. FORGE is designed to produce data that bears on which pattern dominates, and at what difficulty level.

Falsifiability: The hypothesis would be evidence against if models scoring highly on hard tiers suffer catastrophic degradation when mathematical structure is held constant but surface framing is inverted. We have not run this experiment systematically. It is a direction for future work.

02 — How it works

25 categories. Every question generated at runtime.

FORGE generates every problem procedurally from a seed using SHA-256 derived RNGs. No static question bank. Same seed produces identical problems every time, on any machine. FORGE uses procedural generation to make verbatim contamination statistically unlikely at normal evaluation scales. Individual category problem spaces vary significantly in size. Structural similarity to training data cannot be excluded.

25
Categories
5
Difficulty tiers

Category registry

01Arithmetic Chain ✓ CERTIFIED
02Polynomial Roots ⚠ FLAGGED
03Matrix Determinants ✓ CERTIFIED
04Jordan Normal Form ◎ RESEARCH
05RLC Circuits ✓ CERTIFIED
06Boolean Minimization ◎ RESEARCH
07Chess Mate-in-N ◎ RESEARCH
08Shortest Path ✓ CERTIFIED
09RSA Arithmetic ⚠ FLAGGED
10Stars and Bars ✓ CERTIFIED
11Shannon Entropy ◎ RESEARCH
12DFT ✓ CERTIFIED
13Nim States ✓ CERTIFIED
14Group Orders ◎ RESEARCH
15Linear Systems ✓ CERTIFIED
16Modular Exponentiation ✓ CERTIFIED
17Vector Divergence ✓ CERTIFIED
18Polygon Area ✓ CERTIFIED
19Financial Math ✓ CERTIFIED
20Bayesian Updating ✓ CERTIFIED
21Taylor Series ✓ CERTIFIED
22Diophantine Equations ✓ CERTIFIED
23Formal Grammars ◎ RESEARCH
24Quantum States ◎ RESEARCH
25Algorithmic Trace ◎ RESEARCH
✓ CERTIFIED = passes all validation criteria  ·  ◎ RESEARCH = known limitations, interpret with caution  ·  ⚠ FLAGGED = fundamental issue under investigation, scores excluded

Computational grading

All evaluation uses deterministic computational verification rather than LLM judges. This eliminates LLM judge bias but introduces its own edge cases, documented per category. The grading logic is inspectable and consistent — the same input always produces the same grade. Known grading limitations are flagged in the category status table.

EngineUsed for
SymPySymbolic equivalence, differential equations, logic reduction
NumPyMatrix operations, DFT, numerical precision
python-chessGame state validation, legal move verification
Category Status

Which categories are verified, and which are not.

Every category is assessed against three criteria: problem space size (too large for feasible exhaustive contamination), grader self-consistency at 100%, and empirically validated difficulty scaling. Categories that fail any criterion are flagged.

CategoryStatusReason
Polynomial RootsFLAGGEDProblem space of 62.8M. Exhaustive contamination feasible on consumer hardware
RSA ArithmeticFLAGGEDTests computational limits, not reasoning. Under review for redesign
Boolean MinimizationRESEARCHKnown grader bug at difficulty 4-5
Chess Mate-in-NRESEARCHGeneration pool constraints at high difficulty
Algebra GroupsRESEARCHExtremely slow generation at high difficulty
Quantum AmplitudesRESEARCHAnswer parser fragile on non-standard notation
Shannon EntropyRESEARCHBorderline tolerance cases in cross-model testing
Jordan Normal FormRESEARCHDegenerate eigenvalue ambiguity at difficulty 5
Formal GrammarsRESEARCHString matching edge cases in grading
Algorithmic TraceRESEARCHOutput formatting edge cases in grading

15 categories are [FORGE CERTIFIED]. 8 are [RESEARCH]. 2 are [FLAGGED] (scores excluded from main FORGE score). Full details in the README.

03 — Scoring

The Complexity Cliff Index

A simple mean accuracy conflates performance at easy tiers (where training data overlap is likely) with performance at hard tiers (where it is less likely). FORGE weights harder tiers exponentially more, so a model that collapses at difficulty 3 scores differently from one that holds through difficulty 5.

C_c = Σ(αd · Sc,d) / Σ(αd) α = 1.5

The Cliff Index is the first difficulty tier where accuracy drops more than 30% from the previous tier. It is an observation about where performance degrades, not a proof of what causes the degradation. It is the most diagnostically useful single number FORGE produces.

MetricWhat it measuresWhat it does not prove
FORGE ScoreWeighted accuracy across all categories and tiersThat the model has an internal world-model
Interpolation ScoreAccuracy on easy/medium tiersThat these problems were in training data
Extrapolation ScoreAccuracy on hard/expert tiersThat these problems were absent from training data
Cliff IndexFirst tier where accuracy drops >30%The cause of the drop

A model with high interpolation and low extrapolation scores is consistent with retrieval-dominant behavior. A model that maintains accuracy across tiers is consistent with generalization. Neither interpretation is proven by the score alone.

04 — The Nerfed Model Problem

Seeded runs are reproducible by anyone.

AI labs face a quiet incentive problem. A model is released. It scores well on public benchmarks. The scores become marketing. Then, quietly, the model is updated. The published scores no longer reflect what users are actually running.

Fixed benchmarks cannot distinguish between "the model got smarter" and "the model was tuned toward this specific benchmark distribution."

Because every FORGE run is governed by a published seed and every question is generated deterministically at runtime, a score is reproducible. Anyone can rerun the same seed against the same endpoint months later and compare.

1 Lab publishes: Model X, Seed 42, FORGE Score 0.847, date 2025-01-15
2 Six months later, anyone reruns seed 42 against the same endpoint
3 If the score has shifted, the shift is documented
4 Whether the shift reflects capability change or distribution tuning requires further investigation

FORGE does not accuse anyone of nerfing. It makes the question more tractable.

Seed security

FORGE accepts seeds with a 256-bit hash space (SHA-256). The number of genuinely distinct questions per category varies. Question set uniqueness at the 25-category level is substantially larger than any individual category space.

MetricValue
Seed input space2256 ≈ 1.2 × 1077
Time to exhaust (1T hashes/sec)~3.7 × 1051 years
Age of universe1.4 × 1010 years

Note: The above describes the seed space. Individual category problem spaces vary significantly. See Category Status for details.

05 — Design decisions

Why each choice was made

These are the specific problems FORGE was designed to address, and the tradeoffs involved in each design decision.

ProblemDesign decisionTradeoff
Static question banks contaminate pre-trainingProcedural generation from seed at runtimeStructural similarity to training data cannot be excluded. Individual category problem spaces vary in size
Multiple-choice formats are gameableFree-form answers graded computationallyPrompt sensitivity affects scores; format non-compliance penalizes correct answers
LLM judges introduce hallucination and inconsistencySymPy, NumPy, python-chess for gradingEliminates LLM judge bias but introduces grader-specific edge cases, documented per category
Fixed benchmarks cannot detect model version changesSeeded, reproducible runs with published timestampsInterpreting score changes requires controlled conditions we do not enforce
Human-curated hard problems are expensive to scaleParametric difficulty via mathematical parametersParametric difficulty may not map cleanly to human difficulty perception
Observed Behaviors

Empirical findings from FORGE runs

Shannon Entropy — log2 hallucination

During evaluation on seed 42, qwen/qwen3-coder-flash produced confident multi-step entropy calculations using incorrect log2 values. Example: log2(107) computed as 6.7175 versus correct 6.7415 (error −0.024). The model used log2(a/b) = log2(a) − log2(b) with systematically wrong log2 values for integers 103, 107, 221, 309, 321, 332, 347. The reasoning trace showed no uncertainty at the point of error. FORGE's deterministic grader flagged all 10 entropy questions as incorrect. Raw response (first failure, truncated): "log2(50/321) = log2(50) - log2(321) = 5.6439 - 8.3219 = -2.6780" (actual: -2.6826).

Shannon Entropy — cross-model non-monotonic failure

On seed 17017656696087371159 (nano mode, 5 questions, Shannon Entropy difficulty 3), 8 frontier models were evaluated. Five passed (DeepSeek V4 Pro, GPT-5.5, Claude Opus 4.7, Nemotron 550B, MiniMax M3). Three failed on the same entropy question (expected 2.7037): Claude Opus 4.8 extracted 2.7038, Claude Opus 4.6 extracted 2.7036 — both within 4th-decimal rounding error of the correct value. Claude Sonnet 4.6 extracted 2.7056, a larger deviation traced to an incorrect intermediate logarithm: ln(0.072650) computed as −2.6439 versus correct −2.6226. Notably, Opus 4.7 passes while 4.6 and 4.8 fail — a non-monotonic pattern across versions. Sonnet 4.6 is the only model with a demonstrably wrong logarithm value; the other two failures are borderline tolerance cases.

06 — Get started

Run FORGE in under a minute.

Note: The Researcher Tool requires the Python backend running locally. It is a frontend UI that connects to localhost:7860 — without the backend installed and running, the page will show errors when you try to run a benchmark.

Install

# Clone and install
$ git clone https://github.com/mohamedhossammohamed/FORGE.git
$ cd FORGE && pip install -r requirements.txt

Run via CLI

# Quick evaluation (~5 min, 250 questions)
$ python -m forge.cli run \
  --mode quick \
  --model qwen/qwen3-235b-a22b-thinking-2507 \
  --api-base https://openrouter.ai/api/v1 \
  --api-key sk-... \
  --seed 42

Run via web UI

# Launch the researcher tool
$ python -m forge.server
# Opens at http://localhost:7860/researcher.html

Run modes

ModeQuestionsTimeEst. Cost
nano5~30s<$0.01
quick250~5 min~$0.10
standard2,500~1 hr~$1.00
full10,000~4 hr~$4.00

Cost estimation settings — Adjust to match your model's pricing. Defaults reflect typical OpenRouter rates (e.g., GPT-4o: $5/$15, GPT-4o-mini: $0.15/$0.60, Llama-3.1-405B: $3/$9).