The published observation
Arditi and colleagues studied 13 open chat models up to 72B parameters. In their experiments, a dominant direction in residual-stream activations mediated much of the tested refusal behavior. Intervening on that direction changed refusal rates. This is an empirical finding for the evaluated models and prompts, not a universal switch shared by all LLMs.
A direction is not a neuron or a fixed layer
The direction is a vector in an activation space. Its location, extraction and effect depend on the model and experimental setup. There is no universal Layer 16 at which every model decides to refuse. A two-dimensional picture can illustrate a projection, but cannot show the whole learned mechanism.
What changes and what remains uncertain
Removing one measured direction can suppress a class of refusals. Other representations, training effects, system instructions or external checks can still affect behavior. A low refusal rate on one test set does not establish that every future request will be answered or that all other capabilities remain unchanged.
How to assess the evidence
Keep model, prompts, baseline and intervention fixed and report refusal rates together with unrelated capability evaluations. Include different prompt types and repeated trials. Inspect examples as well as aggregate scores. Compare the unmodified checkpoint and the modified checkpoint under identical inference settings.
Recovery is an additional training experiment
Fine-tuning may change side effects of an intervention, but DPO is not a universal healing step with guaranteed recovery. Any follow-up training needs its own data, objective and evaluation. Being able to reload the original checkpoint is different from demonstrating that an edited model retained the original behavior.
Conceptual projection, not a model execution
An activation vector has a component along a chosen direction and a component perpendicular to it. A projection can change the first component. This geometric description says nothing by itself about which behavior the direction represents.
h = h∥ + h⊥