Large language models are often secured primarily at the level of outputs. This Perspective argues that unauthorized use and manipulation should also be treated as attacks on latent persona coordination: a relational property of the internal state, comprising the relative dominance of assistant-like, truth-preserving, and safety-preserving representations over competing dispositions, the conflict between incompatible dispositions that are active at once, and the stability of that configuration under perturbation. Reframing jailbreaks, malicious fine-tuning, hidden-signal training and uncensoring as reweightings of this control state yields a testable prediction: that latent measurements taken after different attacks are better explained by a general drift component plus pathway-specific residuals than by pathway-specific effects alone, and that the resulting signatures could flag manipulation before unsafe outputs appear.
We state what would refute this. During preparation of this manuscript, the authors used LLMs to support language editing, structural refinement and critical review. All suggestions and outputs were independently evaluated and verified by the authors, who take full responsibility for the final content.
No funding was received for this research. Humane Technology Lab., Catholic University of Sacred Heart, Milan, Italy Applied Technology for Neuro-Psychology Lab, Istituto Auxologico Italiano IRCCS, Milan, Italy The authors declare no competing interests. Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material.
If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/. Latent persona coordination as an attack surface in large language models. npj Artif.
Intell. (2026). https://doi.org/10.1038/s44387-026-00154-7 DOI: https://doi.org/10.1038/s44387-026-00154-7
Extract — continue reading at the source.