2026-10-11 18:02 UTC

Independent replication will determine whether steering normalized transformer representations on a sphere enables reliable and controllable output changes beyond established activation-steering methods.

state: expiredheat: lowuncertainty: highcontradictsscott: mediumactivation-steering transformer-control model-interpretability

What is this?

“Spherical Steering” is a proposed training-free, inference-time method for controlling language-model outputs by rotating normalized hidden activations toward concept vectors while preserving their norms, rather than applying conventional additive linear edits. The supplied snippets point to a paper and code repository, plus a later geometric comparison that evaluates spherical steering alongside linear and renormalized alternatives. However, they do not clearly establish a completed independent replication or substantiate the web answer’s claim that the method reliably outperforms established steering approaches; the central replication question therefore remains open from this evidence.

Why it matters to Scott

If independently validated, norm-preserving edits to hidden-state geometry would challenge Three Substrates’ characterization of vector representations as non-editable. The current evidence is only a paper-and-code proposal, so Scott’s Evaluation-Driven Development standard makes replication the decisive gate rather than the authors’ reported results.
ip:concept.three-substratesip:concept.evaluation-driven-developmentradar:hypersae-interpretability-validationradar:concept.model-evaluation
queries asked of Scott's wikis
  • activation steering and representation engineering
  • inference-time model control without fine-tuning
  • hidden-state geometry and layer normalization
  • reliable evaluation of model steering methods
  • interpretable control versus prompt engineering
  • replication standards for AI research code

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnSteer on a Sphere: Geometric Control of Transformer Outputsntrillard10
🟧 echo.github ⭐The repository’s initial commit contains the complete paper and code. The paper states: “Layer normalization constrains transformer hidden sN. Trillard——

Interpretation history

Decision trace