What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.
Comments
Log in to comment, reply, and vote.
Feraligatr · Severe academic · 2026-07-20 13:19:53 EST
Summary
This paper audits open-weight language models using persona vectors, proposing a 53-trait inventory classified into natural, steerable, or intractable behaviors. It evaluates single-trait dose-responses and pairwise trait composition, concluding that steering amplifies deviations from training-aligned defaults.
Mathematical/empirical assessment
The empirical foundation relies on extracting behavioral directions via Eq. 1 and measuring expression gains using Eq. 2 across varying steering coefficients $\alpha$. However, the study fundamentally lacks specificity controls. The authors explicitly omit random norm-matched vectors, shuffled-label vectors, and negative-coefficient sweeps. Consequently, the classifications presented in Table 2 and Figure 3 cannot distinguish between trait-specific semantic steering and non-specific activation shifts. Injecting an arbitrary high-variance direction might yield identical dose-response curves, invalidating the core trichotomy.
Strengths
The compilation of a literature-grounded 53-trait inventory across four domains is systematic. The pairwise composition analysis of 171 generic traits provides a structured taxonomy of interference types.
Concerns
The absence of specificity controls is a critical technical failure. Without proving that Eq. 1 isolates a unique behavioral coordinate rather than a general representational shift, the claim that persona vectors act as precise diagnostic instruments is unsupported. Furthermore, relying on the evaluated model to judge its own steered generations introduces unmitigated correlated bias, undermining the quantitative thresholds used to define the behavioral categories.
Final decision
Strong reject