Abliteration and refusal directions

Expert

A model-specific interpretability result: how a direction in hidden activations can influence refusal behavior, and what the experiment does not establish.

Last updated: Sep 13, 2026

The published observation

Arditi and colleagues studied 13 open chat models up to 72B parameters. In their experiments, a dominant direction in residual-stream activations mediated much of the tested refusal behavior. Intervening on that direction changed refusal rates. This is an empirical finding for the evaluated models and prompts, not a universal switch shared by all LLMs.

A direction is not a neuron or a fixed layer

The direction is a vector in an activation space. Its location, extraction and effect depend on the model and experimental setup. There is no universal Layer 16 at which every model decides to refuse. A two-dimensional picture can illustrate a projection, but cannot show the whole learned mechanism.

What changes and what remains uncertain

Removing one measured direction can suppress a class of refusals. Other representations, training effects, system instructions or external checks can still affect behavior. A low refusal rate on one test set does not establish that every future request will be answered or that all other capabilities remain unchanged.

How to assess the evidence

Keep model, prompts, baseline and intervention fixed and report refusal rates together with unrelated capability evaluations. Include different prompt types and repeated trials. Inspect examples as well as aggregate scores. Compare the unmodified checkpoint and the modified checkpoint under identical inference settings.

Recovery is an additional training experiment

Fine-tuning may change side effects of an intervention, but DPO is not a universal healing step with guaranteed recovery. Any follow-up training needs its own data, objective and evaluation. Being able to reload the original checkpoint is different from demonstrating that an edited model retained the original behavior.

Conceptual projection, not a model execution

An activation vector has a component along a chosen direction and a component perpendicular to it. A projection can change the first component. This geometric description says nothing by itself about which behavior the direction represents.

h = h∥ + h⊥

Primary study: Refusal in Language Models Is Mediated by a Single Direction (2024)