mHC: Manifold-Constrained Hyper-Connections
mHC: Manifold-Constrained Hyper-Connections
Cross-references: Architecture · Schema Reference · Training Safety
This spec documents the Frankenstein integration of mHC (Manifold-Constrained Hyper-Connections), arXiv:2512.24880 (DeepSeek-AI).
What problem does mHC solve?
A normal residual connection (the x + F(x) pattern in every transformer
layer) is what lets gradients flow cleanly through a deep network. When you
expand this to an n-stream residual (keeping n copies of the hidden
state instead of one), you gain representational power but lose the
identity-mapping property: the network can now distort or lose information
between layers, which causes unstable gradients and poor scaling.
mHC fixes this by constraining the within-stream mixing matrix to the Birkhoff polytope (doubly-stochastic matrices). This restores the identity/conservation property, so you get the benefits of a wider stream without the gradient-instability cost. In plain terms: it lets the model use a wider, more expressive residual while guaranteeing the math stays stable.
Overview
mHC replaces the standard residual connection with an n-stream residual.
The residual stream width is expanded by a factor n, so the stream carried
across layers is x ∈ R^{n×C} while each layer’s internal function F
(attention and FFN) still operates at dimension C. Three learnable mappings
read, write and mix the stream each layer:
H[pre] ∈ R^{1×n}— aggregates then·C-dim stream into theC-dim layer input.H[post] ∈ R^{1×n}— maps the layer output back onto the stream.H[res] ∈ R^{n×n}— mixes features within the residual stream.
Unconstrained Hyper-Connections (HC) lose the identity-mapping property of the
residual connection, causing unstable gradients and restricted scalability. mHC
constrains ``H[res]`` to the Birkhoff polytope (the set of doubly stochastic
matrices) via the Sinkhorn-Knopp projection, which restores the identity-mapping
/conservation property: the composite product Π_l H_l[res] stays doubly
stochastic across all depth, spectral norm ‖H[res]‖₂ ≤ 1 (non-expansive), and
H[res] x becomes a convex combination of the stream features.
When n = 1 the doubly-stochastic condition degenerates to scalar 1, exactly
recovering the identity mapping.
Mathematical formulation
Per layer l, with stream x_l ∈ R^{n×C}:
x̃_l = vec(x_l) ∈ R^{1×nC}
H̃ = (1/r)·(α ⊙ (x̃_l φ_l)) + b_l # r = ‖x̃_l‖₂ / √(nC); α gating (init 0.01)
H_l[pre] = σ(H̃_pre) # non-negative
H_l[post] = 2σ(H̃_post) # non-negative, range [0, 2]
H_l[res] = SinkhornKnopp(exp(H̃_res)) # doubly stochastic
Fpre = H_l[pre] @ x_l # [1, C]
x_{l+1} = H_l[res] @ x_l + H_l[post]ᵀ ⊗ F(Fpre, W_l)
φ_l ∈ R^{nC × (n²+2n)}— learned linear projection (full precision).b_l ∈ R^{1 × (n²+2n)}— learned bias.α_pre, α_post, α_res— learnable scalar gates, initialised small (0.01).Sinkhorn-Knopp: starting from
exp(H̃_res), alternate row and column normalisations (defaultt_max = 20rounds) to converge to a doubly stochastic matrix. The backward pass recomputes the iteration on-chip and differentiates through it (exact Jacobian-vector product).
Implementation (Frankenstein)
New module src/model/mhc.py:
SinkhornKnoppFunction(torch.autograd.Function)— differentiable projection.ManifoldHyperConnections(nn.Module)— holdsφ_l(proj),b_l(bias) and the three gating scalars; exposesfpre,recombineandmappings.
Wiring in src/model/frankenstein_model.py:
HybridLayergainsmhc_attnandmhc_ffn(one module per layer function)._forward_dense_mhcruns attention then FFN as layer functions over the shared(B, S, n, C)stream.FrankensteinTransformerexpands theC-dim embedding to(B, S, n, C)viamhc_in_projand collapses back viamhc_out_projbefore the head.mhc_checkpointoptionally applies gradient checkpointing per layer to mitigate the ~``n``× activation-memory increase of the n-stream residual.
Config reference
The model.mhc sub-object (hierarchical schema) or flat keys:
YAML ( |
Flat key |
Type |
Default |
Meaning |
|---|---|---|---|---|
|
|
bool |
|
Enable the mHC n-stream residual. |
|
|
int ≥ 1 |
|
Stream expansion factor |
|
|
int ≥ 1 |
|
Sinkhorn-Knopp normalisation rounds. |
|
|
float > 0 |
|
Initial value of the gating scalars |
|
|
bool |
|
Gradient checkpointing on mHC layers. |
|
|
bool |
|
Keep |
Example config: configs/examples/es_arch_mhc_adamw.yaml.
Enabling mHC in your config
The nested model.mhc block and its flat equivalent are equivalent. To turn
mHC on with a wider stream and checkpointing:
model:
mhc:
enabled: true
expansion_rate: 4 # stream width = 4 × hidden_size
sinkhorn_iters: 20
gating_init: 0.01
checkpoint: true # trade compute for lower activation memory
Choosing expansion_rate
expansion_rate (n) is the width multiplier of the residual stream. Larger
n gives more representational capacity but multiplies activation memory by
roughly n× (this is why mhc_checkpoint exists). As a rule of thumb:
Goal |
Suggested |
|---|---|
Minimal overhead, near-standard residual |
|
Balanced capacity vs. memory (paper default) |
|
Maximum expressiveness (large memory budget) |
|
When n = 1, mHC degenerates to the plain identity residual and adds no
benefit. If you are memory-constrained, start at n = 2 with checkpoint: true.
Constraints
mHC is incompatible with ``use_mixture_of_depths`` (MoD token routing operates on a single
C-dim stream, conflicting with the n-stream residual). AValueErroris raised if both are enabled.The
φ_lprojection stays full-precision under BitNet by default (full_prec_under_bitnet: true) to avoid ternary-quantisation noise on the small mHC coefficients. Set it tofalseto useBitLinear.
Hyperloop mode (arXiv:2604.21254)
Setting model.mhc.hyperloop: true replaces the per-sublayer wiring above with
the Hyperloop Transformer (Zeitoun, Torroba-Hennigen & Kim, MIT): the layer
stack is partitioned middle-cycle into a begin block (runs once), a looped
middle block, and an end block (runs once), and hyper-connections fire once
per loop iteration with per-loop parameters {W_l, b_l, α_l, e_l}
instead of shared per-sublayer modules:
x → begin → copy×n → { read(H^pre) → middle → +e_l → update(H^res, H^post) } × num_loops → avg → end → head
y_t^{(l+1)} = H_l^res · y_t^{(l)} + H_l^post ⊗ ( F(H_l^pre · y_t^{(l)}) + e_l )
H_l^pre = σ(α_pre · (W_pre z_t) + b_pre) # z_t = RMSNorm(flatten(y_t^{(l)}))
H_l^post = 2σ(α_post · (W_post z_t) + b_post)
H_l^res = diag(σ(α_res · (W_res z_t) + b_res)) # "diagonal" (paper default)
Why it matters (paper results): matches depth-matched Transformers at ~50%
fewer parameters (136M Hyperloop PPL 14.40 vs 238M Transformer 14.65 on
FineWeb-Edu), survives INT4/GPTQ quantization, and — because the loop-level
H^res uses a diagonal sigmoid and fires once per loop — adds only ~5%
throughput overhead in plain PyTorch (per-layer mHC loses ~33%).
Deltas vs per-layer mHC implemented above:
Hyper-connections fire once per loop, not after every sub-layer — the paper’s Table 5 ablation also shows every-loop placement is the most performant.
The n-stream residual is created by copying the begin-block output and collapsed by averaging — no learned
mhc_in_proj/mhc_out_proj.Parameters (
W/b/α/e) are per loop, relaxing strict weight sharing so looped representations can deviate across iterations (the paper’s Table 8 cosine-similarity analysis supports this as the source of the gains).H^resdefaults to a diagonal sigmoid (diagonal), the paper’s best ablation (14.40 vs 14.59 Sinkhorn / 14.61 identity, Table 6);"sinkhorn"and"identity"remain available.
New module src/model/hyperloop.py:
HyperloopConnections— per-loopprojs/biases/alphas_*/loop_pos_embswithexpand(copy),collapse(average),mappings,apply_fpre(stream → C) andapply_update(writeF(...) + e_l, mix withH^res).
Wiring:
FrankensteinEncoder.__init__buildsself.hyperloopinstead ofmhc_in_proj/mhc_out_projwhenmhc_hyperloopis set;forwardruns the begin slice, copies the stream, then per loop: readH^pre→ middle layers →apply_update(mixH^res, write withH^post, adde_l) — and finally averages the streams before the end slice.HybridLayerdisables its per-layermhc_attn/mhc_ffnunder Hyperloop (layers run the standard C-dim path), which restores ``use_mixture_of_depths`` compatibility — MoD token routing lives inside middle-block layers and never sees the n-stream residual.FrankensteinViTwires the same loop-level path in_run_encoder.
Hyperloop config reference
YAML ( |
Flat key |
Type |
Default |
Meaning |
|---|---|---|---|---|
|
|
bool |
|
Enable loop-level hyper-connections over the middle-cycle looped block. |
|
|
int ≥ 0 |
|
Begin-block layers (run once before the loop). |
|
|
int ≥ 0 |
|
End-block layers (run once after the loop). |
|
|
enum |
|
|
Shared with per-layer mHC: enabled, expansion_rate (n), gating_init,
checkpoint, full_prec_under_bitnet, sinkhorn_iters (only used by the
sinkhorn parameterization). hyperloop requires enabled: true and
dims.num_loops >= 2. Parameter budget per loop: 3n²C + 3n + 3 + C
(diagonal) — e.g. the paper’s 136M model (C=1024, n=4, 3 loops) spends only
~150K extra parameters vs the plain looped Transformer.
Paper middle-cycle presets (paper uses ~25% / 50% / 25% parameter split over begin/middle/end; unrolled depth = begin + middle×loops + end):
Paper model |
Begin |
Middle (×3) |
End |
|---|---|---|---|
136M (depth 16) |
2L |
4L |
2L |
580M (depth 18) |
3L |
4L |
3L |
990M (depth 38) |
4L |
10L |
4L |
Example config: configs/examples/hyperloop_mhc_adamw.yaml (scaled-down
2L → 2L(x3) → 2L middle cycle).
Hyperloop constraints
Requires
use_mhc: true; plainmhc.enabledwithouthyperloopkeeps the per-sublayer wiring (the two modes are mutually exclusive at runtime).Requires
num_loops >= 2(the paper’s models use 3).hyperloop_begin_layers + hyperloop_end_layers < num_layers(the looped middle block must contain at least one layer).Incompatible with AttnRes residuals (
residuals.type: full_attn/block_attn) — AttnRes applies depth-wise attention per logical layer on the n-stream residual, while Hyperloop only mixes at loop boundaries.Unlike per-layer mHC, Mixture-of-Depths is compatible (middle-block layers run the standard C-dim path).