Why it is different
A model is a new kind of component, with a new kind of attack surface
It reads untrusted text as instructions, it can be steered by content it retrieves, it may hold real permissions through tools, and it was built from data and weights you did not produce. Classic application security still applies, and it is no longer enough.
Prompt injection
Instructions hidden in a web page, a PDF or an email override the system prompt and redirect the model. With tools attached, the model acts on them.
Data leakage
System prompts, retrieved documents, other users’ records or training data surface in responses, through direct extraction or side channels.
Excessive agency
Agents granted broad tool permissions take actions nobody intended: sending mail, changing records, spending money, calling internal APIs.
Supply-chain compromise
Poisoned fine-tuning data, backdoored weights, malicious model files that execute on load, and dependencies nobody reviewed.
Model theft and abuse
Unmetered endpoints extracted or used as a free compute resource; guardrails bypassed to produce content that creates legal exposure.
Silent degradation
No evaluation harness, so a prompt tweak or a provider model update changes safety behaviour and nobody notices until a customer does.