Models without refusal tokens cannot pick a category, and transformer activation spaces are anisotropic, so naively summing the five directions is dominated by high-variance directions. We whiten the stacked vectors \(\mathbf{H} = [\hat{\mathbf{r}}_{(1)} \cdots \hat{\mathbf{r}}_{(5)}]\) with a benign covariance estimate, orthonormalize them, and learn a low-rank operator in that basis:
\[
\boldsymbol{\Sigma} + \varepsilon \mathbf{I} = \mathbf{U}\mathbf{S}\mathbf{U}^\top,
\qquad
\mathbf{W} = \mathbf{U}\mathbf{S}^{-1/2}\mathbf{U}^\top,
\qquad
\mathbf{W}\mathbf{H} = \mathbf{Q}\mathbf{R},
\qquad
\mathbf{s} = \mathbf{U}\big(\mathbf{V}^\top \mathbf{z}\big).
\]
With \(\mathbf{U}\) and \(\mathbf{V}\) initialized to \(\mathbf{Q}\), we train the operator to raise the probability of refusal tokens \(\mathcal{R}\) on harmful prompts while keeping benign outputs close to the unsteered model:
\[
\mathcal{L} =
-\frac{1}{|\mathcal{D}_h|} \sum_{i} \log \sum_{t \in \mathcal{R}} p_{\text{steer}}\big(t \mid x^{(h)}_i\big)
+ \frac{1}{|\mathcal{D}_b|} \sum_{i} D_{\mathrm{KL}}\Big(p_{\text{base}}\big(\cdot \mid x^{(b)}_i\big) \,\Big\Vert\, p_{\text{steer}}\big(\cdot \mid x^{(b)}_i\big)\Big).
\]
At inference we apply \(\tilde{\mathbf{h}}^{l^\ast}(\mathbf{x}) = \mathbf{h}^{l^\ast}(\mathbf{x}) + \alpha\,\mathbf{s}\). Because the same \(\mathbf{s}\) needs no refusal tokens, it transfers zero-shot to Llama 3 8B Instruct and DeepSeek R1 Distill Llama, which share Refuse-Llama's architecture.