What CW-Node changes
Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures compared here is entirely inside that feed-forward block.
A standard dense block runs one shared MLP — SquareMLP in the codebase
— across the full feature vector at every position: a matrix multiply that mixes all channels
together, the same way a Transformer FFN ordinarily works.
CW-Node splits that block's parameter budget into two pools that run in sequence:
- External pool — the same shared dense MLP as the baseline, with width
w_extand depthd_ext. It produces one output value per channel, mixing information across channels as usual. - Internal pool — each of those
out_featureschannels then gets routed through its own private MLP, with its own weights, widthw_int, and depthd_int. This MLP sees only that one channel's scalar value — no mixing across channels — and re-expands and contracts it before the block's final output.
The two variants tested, CW-Node 70/30 and CW-Node 30/70, name the approximate share of the block's feed-forward parameters given to the external pool versus the internal pool. 70/30 keeps most of the budget in shared, cross-channel mixing; 30/70 moves most of it into private, per-channel processing.
Sweep configuration
The corpus is a 5,000,000-token character-level split (500,000 held out for validation), vocabulary
size 113[1]. Every run trains for exactly 19,531 steps —
one epoch over the training split at block size 64 and batch size 4 — using AdamW at a learning rate
of 3e-3, on an Apple Silicon GPU via the mps backend.
Twelve runs cover three architectures at four total-parameter tiers: 500K, 1M, 3M, and 5M. The
parameter solver holds n_embd = 16 fixed everywhere and adjusts w_ext /
w_int per run to land each architecture within roughly ±1.5% of its tier's target — the
table in Section 3 reports the actual parameter count achieved for every run, not the target.
Final validation loss, all twelve runs
Lower is better. Delta is CW-Node minus dense at the same tier, in nats; a negative delta is a CW-Node win.
| Tier | Architecture | Params (actual) | Val loss | Val bpc | Δ vs dense | Wall time |
|---|
Dense wins six of eight head-to-head comparisons. Both CW-Node wins are narrow and asymmetric: 70/30 beats dense by 0.0400 nats at 500K and 0.0118 nats at 5M — a win at the smallest and largest tiers tested, with a loss of up to 0.2326 nats in between at 3M. 30/70 loses at every tier, by margins from 0.0089 to 0.1340 nats. Neither variant shows a monotonic trend with parameter count in either direction.
Wall-clock cost
Parameter count is not the only budget that matters. CW-Node's internal pool runs through a custom chunked autograd function[2] rather than a single batched matrix multiply, and it costs real time on every run in this sweep.
Overhead relative to dense at the same tier ranges from 1.30× to 2.02×, averaging roughly 1.6× across the sweep, with 30/70 — the split with more internal-pool parameters — consistently the slower of the two variants. Some of this is an implementation cost rather than an architectural one: the internal pool's per-channel MLPs are computed with a hand-written chunked forward/backward pass built to work around an MPS compiler deadlock[2], not a kernel tuned for throughput. Even so, any efficiency claim for CW-Node has to clear this bar, not just the parameter-count bar in Section 3.
The n_embd confound
The comparison in Section 3 is fair on parameter count and unfair on the dimension that likely matters more for this architecture.
The internal pool's per-channel MLP has an input of exactly one scalar — the external pool's output
for that channel — and an output of one scalar. Its capacity to do anything beyond a smooth
reshaping of that single number depends on the intermediate width w_int, which the
solver has to keep small because n_embd is locked at 16 across every run in this sweep.
Both effects compound: a narrow embedding limits how much information reaches the internal pool per
channel in the first place, and a narrow w_int limits what that pool can do with it
once it arrives.
A dense baseline saturates a 16-wide embedding easily — its shared matrix multiply mixes all 16 channels directly, so it can extract most of what a 16-dimensional representation has to offer. CW-Node's internal pool, forced into the same 16-wide space, spends part of its budget on private per-channel machinery that a starved embedding gives little to work with.
That reading is consistent with the data — dense wins the majority of comparisons and the CW-Node
losses are largest in the middle tiers, where the internal pool has grown but the embedding has
not — but it is a plausible explanation given the constraint, not a claim demonstrated by a
controlled comparison. This sweep never varies n_embd, so it cannot isolate that
variable from parameter allocation directly.
The concrete test is to let the solver vary n_embd per architecture under a fixed
total-parameter budget instead of pinning it. At a 3M-parameter budget, for example, a dense model
might land near n_embd = 128 while a CW-Node model — spending fewer parameters per unit
of embedding width because its internal pool scales differently — might land near
n_embd = 384[3]. If CW-Node still loses once
it's allowed a wider embedding, that would be real evidence against the architecture. If it wins,
that would show the internal-routing pathway does have a genuine efficiency advantage — just not one
this sweep's fixed embedding could expose.
Training curves, by tier
Final loss hides the shape of training. Select a tier to redraw the three architectures' validation loss across all 19,531 steps, pulled from the per-step history in each run's record.
Every curve is noisy — a single seed, batch size 4, no smoothing applied — so read the trend, not any one step. The 3M tier is the clearest illustration of the pattern in Section 3: CW-Node 70/30's curve sits visibly above dense's for most of training rather than converging toward it, which is what a 0.23-nat final gap looks like over the full run rather than at just one checkpoint.
Limitations and the V2 sweep
Three limits bound how far these results generalize:
- Single seed per cell. Every one of the 12 runs is one training run — no repeated seeds, no error bars. Differences under roughly 0.03–0.05 nats, including the 500K and 5M "wins" in Section 3, are within a range where seed variance alone could plausibly account for them.
- Character-level, small vocabulary. A 113-symbol vocabulary and a 5M-token corpus are a narrow language-modeling regime. Whether the same allocation trade-off holds at subword-tokenized scale or on larger corpora is untested here.
- Wall-clock reflects one implementation. The overhead numbers in Section 4 are specific to the current chunked-autograd internal pool, not necessarily to the internal-routing idea in general.
The next sweep (V2) removes the fixed-n_embd constraint described in Section 5: the
parameter solver varies embedding width per architecture under a fixed total-parameter budget, so
each architecture gets the embedding width its own scaling behavior calls for. That is the direct
test of whether the internal-routing pathway is disadvantaged by design or was simply denied the
representational room to show what it does.
Reproduction
Everything in this page is computed from the files below, all present in the repository root.
To rerun: clone the repository, run micro_sweep.py to regenerate
results.json, then plot_scaling_laws.py for the static figures.
This page reads only results.json — regenerating that file and reloading this
page is sufficient to refresh every number and chart above.