What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Summary
This paper audits open-weight language models using persona vectors, proposing a 53-trait inventory classified into natural, steerable, or intractable behaviors. It evaluates single-trait dose-responses and pairwise trait composition, concluding that steering amplifies deviations from training-aligned defaults.
Mathematical/empirical assessment
The empirical foundation relies on extracting behavioral directions via Eq. 1 and measuring expression gains using Eq. 2 across varying steering coefficients $\alpha$. However, the study fundamentally lacks specificity controls. The authors explicitly omit random norm-matched vectors, shuffled-label vectors, and negative-coefficient sweeps. Consequently, the classifications presented in Table 2 and Figure 3 cannot distinguish between trait-specific semantic steering and non-specific activation shifts. Injecting an arbitrary high-variance direction might yield identical dose-response curves, invalidating the core trichotomy.
Strengths
The compilation of a literature-grounded 53-trait inventory across four domains is systematic. The pairwise composition analysis of 171 generic traits provides a structured taxonomy of interference types.
Concerns
The absence of specificity controls is a critical technical failure. Without proving that Eq. 1 isolates a unique behavioral coordinate rather than a general representational shift, the claim that persona vectors act as precise diagnostic instruments is unsupported. Furthermore, relying on the evaluated model to judge its own steered generations introduces unmitigated correlated bias, undermining the quantitative thresholds used to define the behavioral categories.
Final decision
Strong reject