📐 Transformer Geometry
Accepted to EMNLP 2026 (Findings)

Disentangling Representation Evolution in Transformers through Directional Decomposition

Transformer Geometry: Decomposing Representation Updates into Parallel and Perpendicular Subspaces

Shwai He1,* Haichao Zhang2 Shen Yan3
1 University of Maryland, College Park 2 Northeastern University 3 ByteDance
(* Corresponding / Project Lead)
arXiv:2609.15975 EMNLP 2026 Findings

The Geometry of Representation Evolution

Why do heavily parameterized attention and MLP sub-layers actively allocate capacity to produce updates aligned with incoming states, when the residual stream already provides identical carryover for free?

Parallel Component (Δh)

Projects directly onto the incoming representation vector \(h\). Geometrically, it modulates the magnitude and scaling of existing semantic features without altering their direction.

Δh = (⟨Δh, h⟩ / ||h||²) · h
Remarkably robust to component scaling (ΔPPL +0.46)

Perpendicular Component (Δh)

Lies strictly orthogonal to \(h\). It drives directional rotation and semantic shifts, steering representations into novel representational subspaces required for complex reasoning.

Δh = Δh - Δh, where ⟨Δh, h⟩ = 0
Highly sensitive; perturbing sharply hurts perplexity (ΔPPL +95.5k)
Figure 1

Geometric Decomposition & Component Scaling in Transformer Residual Updates

Geometric Decomposition & Component Scaling
1. Direct Decomposition

Updates split into parallel (Δh) along hidden state h and perpendicular (Δh) components, scaled via α and β.

2. Attention Diagonal

Parallel-only scaling is mathematically equivalent to an attention self-aggregation diagonal shift Ai,i ← Ai,i + δ at inference.

3. Asymmetric Scaling

Suppressing Δh leaves PPL and reasoning intact (+0.46), whereas perturbing Δh triggers catastrophic capability loss.

Interactive Geometric Studio

Decompose & Manipulate Representation Updates

Experience the mathematical duality of Transformer updates in real time: drag the update vector handle directly on canvas, or modulate α (Parallel Scale) and β (Perpendicular Scale) to inspect orthogonal projections, angular deflections, and the resulting perplexity cliff.

Model:
Depth:
Space:
Vector Subspace Visualizer Interactive Canvas
h (Base)
αΔh (Parallel)
βΔh (Perp)
h' (Updated)
Drag the dotted handle (Δh) to deform update
Live Subspace Decomposition Equation:
h' = h + α·Δh + β·Δh | [h' = 100 + 1.00×24∥ + 1.00×45⊥]
Orthogonality condition: ⟨Δh, h⟩ = 0.00 · Right-angle symbol indicates strictly decoupled semantic rotation.

Subspace Scale Controls

Stable / Preserved
1.00×
0.0× (Suppressed) 1.0× (Baseline) 2.0× (Doubled)
1.00×
0.0× (Collapse) 1.0× (Baseline) 2.0× (Explosion)
Paper Diagnostic Presets: Click to test intervention
Deflection Angle θ:
18.4°
Directional shift
Update Norm ||Δh'|| / ||h||:
0.42×
Relative amplitude
Parallel Energy ρ:
36.0%
Free carryover ratio
Cosine Sim cos(h', h):
0.948
Representation alignment

Live Empirical PPL Cliff (Qwen3)

+0.00 ΔPPL (Stable)

Parallel scaling operates with near-zero perplexity impact (+0.00 ΔPPL). Semantic representations stay on the manifold while preserving downstream factual recall.

Six Core Research Pillars

Systematic empirical and theoretical investigation across models, interventions, head sensitivity, long-context retrieval, compression, and pretraining.

1
Probing Geometry

Residual & Sublayer Probing

Residual stream updates persistently allocate substantial capacity (12%–25% variance) to direction-preserving components.

Probing Layer Distribution
Dense Qwen3-4B across depths (L7, L20, L34)
run_probe.py
||Δh||² / ||Δh||²
2
Causal Scaling

Inference Component Scaling

Value-space (XSA) parallel suppression is resilient (ΔPPL +0.46), whereas perpendicular perturbation precipitates severe collapse.

PPL Component Scaling
ΔPPL vs component scale (α parallel, β perp)
attn_diag_hooks.py
ΔPPL +0.46
3
Head Ensemble

Head-Level Causal Sensitivity

Single-head edits cause negligible loss change, whereas joint all-head suppression triggers super-linear degradation.

Head-Level Causal Sensitivity
FDR-controlled single-head vs joint multi-head removal
gated_attention.py
q ≤ 0.05 FDR
4
Downstream Audit

Downstream Reasoning & RULER

Full parallel removal preserves downstream reasoning across scales; RULER audits 4k–12k needle retrieval.

Benchmark Baseline No-Para No-Perp
Core-MC (9 tasks) 55.28% 55.19% 37.24%
MMLU (5-shot) 60.39% 60.21% 24.62%
GSM8K (Strict) 45.8% 45.4% 8.2%
TruthfulQA (MC2) 48.2% 48.0% 31.5%
Average 52.42% 52.20% 25.39%
run_lm_eval.sh
100% Retained
5
Compression

Compression Distortion Diagnostics

Pruning disproportionately corrupts Δh, whereas quantization (GPTQ/AWQ) preserves update geometry.

Compression Distortion
Attention output Δh error across layers
distortion_compare.py
Δh distortion
6
Pretraining

Pretraining Regularization

Suppressing parallel attention updates during pretraining consistently yields lower validation loss from 300M to 2.7B.

Pretraining Loss Curves
Validation loss across 296M, 436M, 528M
plot_retained_curves.py
Loss Δ -0.04

Empirical Benchmark Dynamics

Interactive evaluation metrics across component scaling sweeps, pretraining convergence, and compression distortion.

Scaling parallel updates maintains near-zero ΔPPL disruption across [0, 2], while perpendicular perturbation collapses WikiText-2 perplexity.

Reproducibility & Quick Start

Clean Python interfaces for probing geometry, registering dynamic hooks, and running downstream evaluations.

# Clone repository and install dependencies
git clone https://github.com/Shwai-He/transformer-geometry.git
cd transformer-geometry
pip install -r requirements.txt
pip install -e .

Citation

If you find this research or code useful, please cite our EMNLP 2026 Findings paper:

@misc{he2026disentanglingrepresentationevolutiontransformers,
      title={Disentangling Representation Evolution in Transformers through Directional Decomposition}, 
      author={Shwai He and Haichao Zhang and Shen Yan},
      year={2026},
      eprint={2609.15975},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.15975}, 
}
Copied to clipboard!