← Back to glossary

Steering (Interventions)

A set of interventions used to shift model behavior at inference time without performing a complete retraining run. Control may operate on context, decoding probabilities, or internal activations; activation steering, for example, adds or removes a vector direction at a layer to promote an observed attribute.

Steering changes a distribution without creating a logical guarantee about the output. Excessive strength may degrade coherence, interfere with neighboring capabilities, or work only in the domain where the direction was estimated. Every intervention needs a baseline, strength control, task-level evals, and rollback.