An engineer reviewing an AI application across several monitors

AI security · Threat modelling

Know how your AI system fails before someone else finds out.

A structured model of what your LLM application, agent or RAG system can be made to do, what that would cost you, and which controls actually reduce the risk. Done before the architecture hardens, or as the first step before a red team.

Why threat model AI

The expensive mistakes are architectural

Which tools the agent can call, whether retrieved content is treated as data or instruction, what the model can see across tenants, where the human checkpoint sits: these are decided early and are costly to change later. A threat model puts the attacker in the room while they are still cheap to change.

The output is not a list of scary things. It is a ranked set of decisions, with the control that settles each one.

Systems we threat model

  • Customer-facing chat and support assistants
  • Retrieval-augmented generation over internal documents
  • Autonomous and semi-autonomous agents with tool access
  • MCP servers, connectors and plugin ecosystems
  • Copilots embedded in existing products
  • Document, email and web-browsing automations
  • Fine-tuned and self-hosted models
  • Multi-tenant AI platforms

What you receive

Six deliverables, each one usable by a different team

01

System and data-flow model

Every model, prompt, tool, retrieval source, data store, user role and external service, with the trust boundaries drawn where untrusted content actually enters.

02

Attack trees and abuse cases

Realistic attacker goals worked back to concrete techniques, mapped to MITRE ATLAS and the OWASP Top 10 for LLM applications, with likelihood and impact rated for your context.

03

Prioritised control set

The controls that reduce the rated risks, in order of return: input and output handling, tool permissioning, retrieval hygiene, isolation, logging, rate limits and human checkpoints.

04

Secure architecture recommendations

Where to put the boundaries: separate privileges per tool, least-privilege connectors, content provenance in retrieval, sandboxed execution and the monitoring that makes misuse visible.

05

Evaluation and test plan

What a red team should try first, what regression evaluations should run on every release, and what to log so an incident can be reconstructed.

06

Compliance mapping

Risks and controls tagged against NIST AI RMF functions, EU AI Act obligations and ISO/IEC 42001 clauses, so the threat model becomes a reusable piece of your governance file.

Worked example

A support agent that can read tickets and issue refunds

A typical first engagement. The feature is simple to describe and has several ways to go badly wrong.

  • Untrusted content enters through the customer message and through ticket history the agent reads: both are injection paths
  • The refund tool is the asset; the question is who can cause it to be called, with what arguments, and how often
  • Cross-customer data exposure is possible if retrieval is not scoped to the authenticated account
  • Rate limits and value caps on the tool matter more than any prompt wording
  • A human checkpoint above a threshold, plus logging of every tool call with its triggering context, closes most of the tree

What changed after the model

  • Refund tool moved behind a separate service with its own authorisation and a per-account daily cap
  • Retrieval scoped to the caller’s account at the query layer, not by prompt instruction
  • Ticket content wrapped as data with provenance tags; the model is evaluated on ignoring instructions inside it
  • Approval step for refunds above a threshold, with the agent’s reasoning shown to the approver
  • Regression evaluations added for the top injection and exfiltration cases before each release

Illustrative engagement. Details are generalised and do not describe a specific client.

How it runs

Workshops with your engineers, evidence in the write-up

01

Scope

We map the system under test: models, prompts, tools, retrieval sources, data flows, users and the business impact of each failure. Rules of engagement and a test plan are agreed in writing.

02

Assess

Threat modelling, adversarial testing or audit against the agreed plan, with every finding reproduced and recorded with evidence, severity and the affected component.

03

Report and fix

A report your engineers can act on and your leadership can read: ranked findings, root causes, concrete remediation, and a control map against the frameworks you answer to.

04

Re-test and embed

Fixed findings are re-tested. Regression test suites, guardrail evaluations and monitoring recommendations are handed over so the next release does not reopen them.

Threat models are usually followed by red teaming of the highest-rated paths, and by a supply-chain audit where third-party models or datasets are in scope.

FAQ

Frequently asked questions

When is the right time to threat model an AI system?

Before the architecture is fixed, ideally when the first prototype works and the team is deciding which tools and data sources to connect. A threat model at that point changes design decisions cheaply. We also run them on systems already in production, usually as the first step before a red team.

Do you use STRIDE or something AI-specific?

Both. STRIDE-style decomposition still works for the application and infrastructure. For the model and its interactions we use attack trees built from MITRE ATLAS and the OWASP LLM Top 10, which cover prompt injection, data poisoning, model theft, excessive agency and the rest of the AI-specific classes.

How long does it take?

A single AI feature or application is typically two to three weeks from kickoff to final workshop. Platform-scale models with many integrations run longer and are usually phased by subsystem.

Who needs to be involved from our side?

The engineers who built or are building the system, whoever owns the data it touches, and someone who can speak to the business impact of failure. Two or three workshops of ninety minutes each is the usual commitment.

Does the threat model cover the base model provider?

It covers your dependency on the provider: what you send, what they retain, how a model update changes behaviour, and what you would do if the endpoint were unavailable or compromised. We do not pretend to threat model the provider’s internals.

Bring the architecture diagram. We will bring the attacker.

A scoping call is enough to size the engagement. Most single-feature threat models complete in two to three weeks.

Or email [email protected] · Mohali, India