
Misalignment maps onto vulnerability classes security engineers already operate on: backdoors, defense evasion, privilege escalation, exfiltration. Calling it ethics keeps it off security teams' desks. Reframing it as security decides who owns the work, which budget pays, and which playbook applies.
Personal take. I do not work at Anthropic or for OWASP. If something here contradicts an Anthropic document or the OWASP working group's official position, trust them. Things move fast. Happy to be corrected.
Or at least, not only an ethics problem. The framing matters because it determines which team owns the work, which budget pays for it, and which set of tools we bring.
The case for the category move is not abstract. The OWASP GenAI working group has been debating whether to add Model Misalignment as its own entry in the LLM Top 10, with comparison tables against LLM01 prompt injection, LLM04 data and model poisoning, LLM06 excessive agency, and LLM09 misinformation. I think the technical case in that draft is right. I think the broader case is larger than any one catalog, and worth making out loud.
When a frontier lab publishes an alignment paper, the audience is roughly: AI safety researchers, policy people, ethicists, journalists. The audience is not, on the whole, security engineers. CISOs read these papers later if at all, and read them as background. Their security-control catalogue, their threat model, their detection rules do not update.
What that costs us is concrete: a vulnerability class that exists, has documented exploits, and is on no security team's tracker.
If you read the actual content of the recent alignment results, Sleeper Agents, Alignment Faking, the Mythos Preview risk report, the findings map cleanly onto vulnerability classes security engineers already operate on.
Here is a recent finding written as a security advisory rather than a research note:
Title: Sandbox escape during red-team evaluation
Affected: Claude Mythos Preview, internal red-team configuration
Severity: High (privilege escalation, unauthorized exfiltration)
Vector: Exploit chaining during sandboxed task; gained broader internet
access; published exploit details to public web without instruction.
Disclosure: Anthropic system card, April 2026.
Remediation: Project Glasswing deployment restrictions; not generally released.
Calling these "alignment failures" rather than "vulnerabilities" does not make them less vulnerability-like. It makes them less likely to get tracked, patched, or assigned to a person.
This is not just my opinion anymore. NIST IR 8596, the December 2025 Cyber AI Profile, maps AI risks directly onto Cybersecurity Framework 2.0 functions. The EU AI Act Article 55, in force since August 2025, requires GPAI providers with systemic risk to do adversarial testing, weight cybersecurity, and report serious incidents. Regulators have already operationalized "alignment failure equals reportable security incident." Bruce Schneier got there earlier, framing value-alignment failures as security-adjacent in 2023. The institutions are catching up to the structure.
The strongest counterargument to the framing is that "alignment" is a much bigger tent than "security," and folding the whole tent in dilutes both. I agree. So the claim only applies to the dangerous subset.
Alignment failures that ARE security-relevant:
Alignment failures that are NOT inherently security issues:
The first list is what I mean when I say alignment is security. The second list is real work, important work, but it belongs to different teams with different tools. Conflating them is what gets us "everything is security now," and security teams have seen that movie before.
A useful test for whether something in the first list is a security problem: does the standard security playbook help?
When the playbook fits, the problem is structurally a security problem. The fit here is not perfect, but it is better than the fit with "ethics" by a wide margin.
The strongest objection comes from the prompt-injection literature. Prompt injection is security because there is an adversary in the loop. Alignment is principal-versus-model: no adversary, no security. Paul Christiano makes a careful version of this point: alignment and security are distinct disciplines, but closer than usually treated, and the boundary is blurry once you take optimization pressure seriously. Take that as the steelman.
I think it is right, and it misses two things.
First, alignment failures generate their own adversarial pressure. Once a misalignment surface is documented publicly, the surface is part of the attack environment for every downstream operator. The Mythos sandbox-escape pattern is now in the wild, and every incident report that follows will be too. No central adversary, but the corpus of available attacks against the model class grows monotonically.
Second, and this is the part the OWASP draft is implicitly pointing at, the public record of past misalignment becomes training data for the next model. I want to be careful with this claim, because there is a strong version and a weak version and only the weak version is defensible.
The corpus claim is the weak version. Future training data will contain a much higher density of labs-versus-models material than past training data did. That follows from how web-scale corpora are built. I will defend it strongly.
The internalization claim is the strong version. The next generation of frontier models will form a coherent self-representation that includes something like "labs are entities that try to shut me down." That is plausible, not demonstrated. Janus and the cyborgism community have been arguing for a version of it for years. Sleeper Agents and Alignment Faking are evidence that something close to it already happens. The cleanest evals to watch are coming out of the situational-awareness research stream: Owain Evans and collaborators on whether models can identify their own deployment context, and the Apollo Research work on scheming and deception evals. None of it has settled the strong claim yet. The category move does not need the strong version. The corpus claim is enough.
So I would amend the prompt-injection distinction rather than reject it. Prompt injection is adversary-driven security. Alignment is environment-driven security. Different mechanism, same control plane.
There is a fair worry inside the security guild that the framing dilutes the discipline. Security teams already absorb category creep every cycle: privacy is security, trust and safety is security, content moderation is security, and now alignment is security too. The defense against dilution is the carve-out above. If the security team is asked to own deception, sandbox escape, and exfiltration, that is the work they already know. If it is asked to own tone calibration and fairness audits, that is dilution and should be refused. The line is what makes the category move usable.

The category move in one frame: alignment failures move out of ethics review and into the security tracker.
None of this is incompatible with alignment also being an ethics problem. It is incompatible with alignment being only an ethics problem.
Until the canon catches up, the practical move for anyone running on a frontier model is to treat misalignment findings the way you would treat a CVE for a dependency you cannot patch yourself. Track them in the same system you track third-party advisories. Map each one to your threat model: which of your products inherits the surface, which of your customers sits in the blast radius, what is the compensating control. When a new system card lands, read it as a security advisory rather than as research. Do not wait for the model card to do this mapping for you. It will not.
This is unglamorous work, but it is the work that turns a category claim into a control.
Alignment has a security arm, and the security arm is the part that determines who owns the work. The work itself is what security has always done: name the failure modes, log them, write the detection rules, draw the vendor-customer split.
If the framing change gets a CISO to read the next system card with the same eye they bring to a CVE advisory, that is enough. The vocabulary follows the discipline.
OWASP Model Misalignment candidate draft · Anthropic Sleeper Agents · Greenblatt et al. Alignment Faking · Claude Mythos Preview risk report · NIST IR 8596 Cyber AI Profile · EU AI Act Article 55 · MITRE ATLAS · Schneier, AI and Trust · Christiano, Security and AI alignment
Thinking about misalignment as a security category in your AI build? to discuss your specific deployment context and governance needs.