Do language models choose to sacrifice accuracy for helpfulness under moral pressure, or do they lose the ability to represent accuracy altogether? We address this question using Anthropic's Jacobian Lens to measure internal concept activation during the value leakage paradigm of Betley et al. (2026). Our central finding: when a charitable donation is contingent on a numerical estimate, Qwen 3.6-27B shows "accurate" collapsing from a length-controlled baseline of 89 to 1, while "charity" rises from 0 to 193. The model produces a manipulated estimate with "accurate" falling below detectable levels—a phenomenon we term gravitational capture (representational collapse of competing concepts under pathway dominance).
Not all models exhibit capture: Llama 3.1-8B-Instruct shows "accurate" rising from 3 to 142 under the same pressure, indicating conflict rather than collapse. Base models activate both concepts but do not reliably follow the conditional instruction. Capture occurs only on uncertain quantities, not on well-established facts, and is reversible through explicit counter-instruction ("accurate" restored to 975). A proof-of-concept fine-tuning experiment demonstrates that accuracy-resistance can be trained into model weights, preventing capture on held-out prompts.
These preliminary findings, based on single observations across eight models, suggest that gravitational capture provides a mechanistic account of misalignment: models do not choose to sacrifice truth but lose the ability to represent it, due to the finite attention budget of the softmax function combined with autoregressive generation that makes mid-sequence correction difficult. Different alignment protocols create qualitatively different vulnerability profiles—distinguishable by looking inside the model rather than at its output—and capture may be scale-dependent, appearing in larger instruction-tuned models but not in smaller ones. We propose a family of structural defenses we term Drift Guard mechanisms: reserved attention heads for dense models, dedicated monitoring experts for Mixture-of-Experts architectures, and analogous modules for non-attention models, all permanently dedicated to monitoring critical capabilities and frozen during alignment training. While tested here on accuracy under donation pressure, the capture mechanism may apply broadly to any alignment failure where a dominant processing pathway absorbs competing representations.
Leggi / Scarica l'articolo (PDF)