CW-Node Research — Micro Scaling-Law Sweep, V1 github.com/mohamedhossammohamed/cw-node-research ↗

The n_embd Confound

A twelve-run, parameter-matched sweep of dense and connection-weighted node transformers from 500K to 5M parameters — and why the sweep can't yet separate a routing effect from an embedding-width bottleneck.

Abstract

CW-Node splits each transformer feed-forward block into two pools: an external pool, a standard dense MLP shared across all channels, and an internal pool, a private small MLP owned by each output channel individually. This sweep asks whether shifting parameter budget from the external pool to the internal one buys parameter efficiency over a dense-only baseline, holding total parameters fixed across four tiers. The dense baseline wins six of eight matched comparisons against the two CW-Node splits. Both CW-Node wins are under 0.04 nats and sit at opposite ends of the parameter range. A structural constraint in this sweep — every configuration is locked to an embedding width of 16 — means the internal pool never has enough per-channel bandwidth to do useful work, so the sweep cannot yet distinguish a real routing disadvantage from an embedding-width bottleneck. Section 5 lays out that argument against the run data directly, and Section 7 specifies the follow-up sweep it implies.

parameter efficiencyscaling lawsfeed-forward routingcharacter-level LMembedding width
01 Section 1

What CW-Node changes

Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures compared here is entirely inside that feed-forward block.

A standard dense block runs one shared MLP — SquareMLP in the codebase — across the full feature vector at every position: a matrix multiply that mixes all channels together, the same way a Transformer FFN ordinarily works.

CW-Node splits that block's parameter budget into two pools that run in sequence:

  • External pool — the same shared dense MLP as the baseline, with width w_ext and depth d_ext. It produces one output value per channel, mixing information across channels as usual.
  • Internal pool — each of those out_features channels then gets routed through its own private MLP, with its own weights, width w_int, and depth d_int. This MLP sees only that one channel's scalar value — no mixing across channels — and re-expands and contracts it before the block's final output.

The two variants tested, CW-Node 70/30 and CW-Node 30/70, name the approximate share of the block's feed-forward parameters given to the external pool versus the internal pool. 70/30 keeps most of the budget in shared, cross-channel mixing; 30/70 moves most of it into private, per-channel processing.

02 Section 2

Sweep configuration

The corpus is a 5,000,000-token character-level split (500,000 held out for validation), vocabulary size 113[1]. Every run trains for exactly 19,531 steps — one epoch over the training split at block size 64 and batch size 4 — using AdamW at a learning rate of 3e-3, on an Apple Silicon GPU via the mps backend.

Twelve runs cover three architectures at four total-parameter tiers: 500K, 1M, 3M, and 5M. The parameter solver holds n_embd = 16 fixed everywhere and adjusts w_ext / w_int per run to land each architecture within roughly ±1.5% of its tier's target — the table in Section 3 reports the actual parameter count achieved for every run, not the target.

Layer depth changes with tier — 2 layers at 500K and 1M, 3 layers at 3M and 5M — because the solver adds depth rather than width once width alone can't reach the target within the embedding constraint below.
03 Section 3

Final validation loss, all twelve runs

Lower is better. Delta is CW-Node minus dense at the same tier, in nats; a negative delta is a CW-Node win.

final_val_loss / final_val_bpc / wall_time_s — read directly from results.json
Tier Architecture Params (actual) Val loss Val bpc Δ vs dense Wall time
Scaling-law frontier — params vs. validation loss
Dense CW-Node 70/30 CW-Node 30/70
x-axis: actual total parameters, log scale · y-axis: final validation loss (nats) · computed live from the 12 run records embedded in this page

Dense wins six of eight head-to-head comparisons. Both CW-Node wins are narrow and asymmetric: 70/30 beats dense by 0.0400 nats at 500K and 0.0118 nats at 5M — a win at the smallest and largest tiers tested, with a loss of up to 0.2326 nats in between at 3M. 30/70 loses at every tier, by margins from 0.0089 to 0.1340 nats. Neither variant shows a monotonic trend with parameter count in either direction.

04 Section 4

Wall-clock cost

Parameter count is not the only budget that matters. CW-Node's internal pool runs through a custom chunked autograd function[2] rather than a single batched matrix multiply, and it costs real time on every run in this sweep.

Training wall time by architecture and tier
bar length is proportional to wall_time_s within each tier group; dense is always the shortest bar

Overhead relative to dense at the same tier ranges from 1.30× to 2.02×, averaging roughly 1.6× across the sweep, with 30/70 — the split with more internal-pool parameters — consistently the slower of the two variants. Some of this is an implementation cost rather than an architectural one: the internal pool's per-channel MLPs are computed with a hand-written chunked forward/backward pass built to work around an MPS compiler deadlock[2], not a kernel tuned for throughput. Even so, any efficiency claim for CW-Node has to clear this bar, not just the parameter-count bar in Section 3.

05 Section 5

The n_embd confound

The comparison in Section 3 is fair on parameter count and unfair on the dimension that likely matters more for this architecture.

The internal pool's per-channel MLP has an input of exactly one scalar — the external pool's output for that channel — and an output of one scalar. Its capacity to do anything beyond a smooth reshaping of that single number depends on the intermediate width w_int, which the solver has to keep small because n_embd is locked at 16 across every run in this sweep. Both effects compound: a narrow embedding limits how much information reaches the internal pool per channel in the first place, and a narrow w_int limits what that pool can do with it once it arrives.

Why this matters for the results above

A dense baseline saturates a 16-wide embedding easily — its shared matrix multiply mixes all 16 channels directly, so it can extract most of what a 16-dimensional representation has to offer. CW-Node's internal pool, forced into the same 16-wide space, spends part of its budget on private per-channel machinery that a starved embedding gives little to work with.

That reading is consistent with the data — dense wins the majority of comparisons and the CW-Node losses are largest in the middle tiers, where the internal pool has grown but the embedding has not — but it is a plausible explanation given the constraint, not a claim demonstrated by a controlled comparison. This sweep never varies n_embd, so it cannot isolate that variable from parameter allocation directly.

The concrete test is to let the solver vary n_embd per architecture under a fixed total-parameter budget instead of pinning it. At a 3M-parameter budget, for example, a dense model might land near n_embd = 128 while a CW-Node model — spending fewer parameters per unit of embedding width because its internal pool scales differently — might land near n_embd = 384[3]. If CW-Node still loses once it's allowed a wider embedding, that would be real evidence against the architecture. If it wins, that would show the internal-routing pathway does have a genuine efficiency advantage — just not one this sweep's fixed embedding could expose.

06 Section 6

Training curves, by tier

Final loss hides the shape of training. Select a tier to redraw the three architectures' validation loss across all 19,531 steps, pulled from the per-step history in each run's record.

Validation loss vs. training step
Dense CW-Node 70/30 CW-Node 30/70

Every curve is noisy — a single seed, batch size 4, no smoothing applied — so read the trend, not any one step. The 3M tier is the clearest illustration of the pattern in Section 3: CW-Node 70/30's curve sits visibly above dense's for most of training rather than converging toward it, which is what a 0.23-nat final gap looks like over the full run rather than at just one checkpoint.

07 Section 7

Limitations and the V2 sweep

Three limits bound how far these results generalize:

  • Single seed per cell. Every one of the 12 runs is one training run — no repeated seeds, no error bars. Differences under roughly 0.03–0.05 nats, including the 500K and 5M "wins" in Section 3, are within a range where seed variance alone could plausibly account for them.
  • Character-level, small vocabulary. A 113-symbol vocabulary and a 5M-token corpus are a narrow language-modeling regime. Whether the same allocation trade-off holds at subword-tokenized scale or on larger corpora is untested here.
  • Wall-clock reflects one implementation. The overhead numbers in Section 4 are specific to the current chunked-autograd internal pool, not necessarily to the internal-routing idea in general.

The next sweep (V2) removes the fixed-n_embd constraint described in Section 5: the parameter solver varies embedding width per architecture under a fixed total-parameter budget, so each architecture gets the embedding width its own scaling behavior calls for. That is the direct test of whether the internal-routing pathway is disadvantaged by design or was simply denied the representational room to show what it does.

08 Section 8

Reproduction

Everything in this page is computed from the files below, all present in the repository root.

cw_node.py — CWNodeLayer, SquareMLP, and the chunked internal-node autograd function (Section 1, 4)
extract_micro_dataset.py — builds the 5M/500K token split from the source corpus (Section 2)
micro_sweep.py — parameter solver and experiment runner for all 12 configurations
results.json — full per-step training/validation history for every run — the source for every chart on this page
plot_scaling_laws.py — the repository's own matplotlib plotting script
scaling_law_frontier.png, learning_curves_3M.png — static renders from plot_scaling_laws.py
micro_train.bin, micro_val.bin, meta.json — the tokenized dataset and character vocabulary

To rerun: clone the repository, run micro_sweep.py to regenerate results.json, then plot_scaling_laws.py for the static figures. This page reads only results.json — regenerating that file and reloading this page is sufficient to refresh every number and chart above.

— Notes
  1. Vocabulary and tokenizer are defined in meta.json (113 characters, simple stoi/itos character mapping). Training and validation splits are micro_train.bin (10,000,000 bytes = 5,000,000 uint16 tokens) and micro_val.bin (1,000,000 bytes = 500,000 tokens).
  2. See CWNodeAutogradFunction and internal_node_forward in cw_node.py. The forward pass processes the internal pool in fixed-size chunks (default 128) under torch.no_grad() and recomputes chunk-by-chunk for the backward pass — a workaround for an MPS graph-compiler deadlock noted in the source, not a throughput-optimized kernel.
  3. The n_embd = 128 / n_embd = 384 figures are the repository's own illustrative example of what an unconstrained solver might select at a 3M-parameter budget — a hypothesis motivating the V2 sweep, not a measured result. No run in this dataset varies n_embd.