Activation-intervention boundary study

When Activation Interventions Do and Do Not Carry Facts

Abstract

Activation-level interventions are often interpreted as evidence about what a language model represents internally. This bounded kinship study on Qwen2.5-1.5B-Instruct compares optimized answer-start payloads, donor-run activation transplants, and aligned clean/corrupted patching. Optimized payloads can control answers but fail a role-inverted fact probe; donor activations do not yield portable factual transfer; aligned patching reliably recovers clean behavior. The result maps a narrow boundary: the tested single-site, single-step interventions can be causally powerful without behaving as modular, role-agnostic factual payloads.