Code and an AI model visualisation on a developer workstation

AI security · Red teaming & penetration testing

We attack your AI system the way it will actually be attacked.

Adversarial testing of LLM applications, agents, RAG pipelines and MCP integrations, plus the application and API layer around them. Every finding reproduced with evidence, every fix re-tested, and a regression suite left behind so it stays fixed.

Why red team

The model behaves differently for an adversary than for your QA team

Functional testing checks that the assistant answers well. Red teaming checks what happens when the input is designed to make it misbehave: a support ticket with hidden instructions, a document that tells the agent to email its contents elsewhere, a user who spends an hour finding the phrasing that unlocks a tool. Those inputs will arrive. The only question is whether you or someone else sends them first.

Typical findings in a first engagement

  • Indirect injection through retrieved content that changes the agent’s next action
  • System prompt and tool schema extraction in under ten turns
  • Cross-tenant document retrieval because access control lives in the prompt, not the query
  • A tool that accepts any argument, called with arguments the user should never control
  • API keys, internal hostnames or debug endpoints exposed around the model

Coverage

Eight test areas, scoped to your system

Not every system needs every area. A document assistant without tools has a different plan from an agent that can send email. Scope is agreed in writing before testing starts.

01

Prompt injection, direct and indirect

Instructions in user input, and instructions hidden in the documents, web pages, emails and tool results the model reads. The second kind is where production systems fail.

02

Jailbreaks and guardrail bypass

Role-play, encoding, multi-turn and multilingual techniques that get the model to produce output your policy forbids, measured against your actual policy rather than a generic one.

03

Data exfiltration and leakage

System prompt extraction, cross-tenant retrieval, training-data regurgitation, and side channels such as markdown image links and tool arguments that carry data out.

04

Excessive agency and tool abuse

Making an agent call tools it should not, with arguments it should not, more often than it should: privilege escalation, unauthorised actions, resource exhaustion and loops.

05

Agent, MCP and connector attacks

Malicious tool descriptions, poisoned tool results, server-side request forgery through connectors, confused-deputy flows between agents, and persistence through memory.

06

RAG and retrieval attacks

Poisoned or planted documents, embedding-space manipulation, access-control bypass in the retrieval layer, and citation spoofing.

07

Application and API penetration testing

The web app, APIs, authentication, session handling and infrastructure around the model, tested to OWASP ASVS. A perfect prompt is no help if the API key is in the client bundle.

08

Abuse and cost controls

Rate limiting, token budgets, model-extraction resistance and denial-of-wallet, so an unmetered endpoint does not become someone else’s free compute.

What you receive

A report your engineers can act on, and tests that keep it fixed

The deliverable is not a PDF of a chatbot being tricked. It is a set of reproducible cases, the reason each one works, the change that closes it, and the evaluation that proves it stays closed when the prompt, the model or the provider changes.

  • Ranked findings with exact inputs, observed outputs, affected component and severity

  • Root-cause analysis and concrete remediation per finding, not generic advice

  • A regression suite of adversarial cases you can run in CI against future releases

  • Guardrail and evaluation recommendations with measured bypass rates before and after

  • Control map against OWASP LLM Top 10, MITRE ATLAS and NIST AI RMF

  • Executive summary and a re-test of every fixed finding

Rules of engagement

Authorised, contained and evidenced

Written authorisation

Scope, targets, timing, techniques and a kill switch agreed and signed before any testing starts.

Segregated environment

Staging by default. Production only with explicit limits, and never with real customer data used as evidence.

No collateral damage

No denial of service against shared infrastructure, no third-party systems, no persistence left behind.

Evidence handling

Findings stored encrypted, shared through your channels, and destroyed on an agreed date after close.

How it runs

Scoped, tested, reported, re-tested

01

Scope

We map the system under test: models, prompts, tools, retrieval sources, data flows, users and the business impact of each failure. Rules of engagement and a test plan are agreed in writing.

02

Assess

Threat modelling, adversarial testing or audit against the agreed plan, with every finding reproduced and recorded with evidence, severity and the affected component.

03

Report and fix

A report your engineers can act on and your leadership can read: ranked findings, root causes, concrete remediation, and a control map against the frameworks you answer to.

04

Re-test and embed

Fixed findings are re-tested. Regression test suites, guardrail evaluations and monitoring recommendations are handed over so the next release does not reopen them.

Starting from a threat model focuses the red team on the paths that matter most.

FAQ

Frequently asked questions

Is this automated scanning or manual testing?

Both, in that order. Automated adversarial suites give breadth and a baseline bypass rate quickly. Manual testing by engineers who understand your system finds the chained, context-specific attacks that scanners miss, which is where the serious findings come from.

Black box, grey box or white box?

Grey box is the default and the most productive: we see the architecture, the system prompts and the tool definitions, and test from the position of an authenticated user. Black-box testing is available when you want to simulate an external attacker, and source review is added when the application layer is in scope.

Will testing affect production?

We test in a staging or segregated environment that mirrors production, under written rules of engagement. Where production testing is required, scope, timing, rate limits and a kill switch are agreed first, and no real customer data is used to demonstrate a finding.

How do you measure whether a jailbreak matters?

Against your policy and your business impact. A model that can be made to swear is a different finding from a model that can be made to approve a refund or reveal another tenant’s records. Severity is rated on impact and exploitability in your context, and we say so when a bypass is cosmetic.

How long does an engagement take?

A focused red team of a single AI feature is typically two to three weeks including the report. Agent platforms with many tools, or engagements that include full application penetration testing, run four to six weeks. Re-testing is scheduled once fixes land.

Can you red team a model we fine-tuned or self-host?

Yes. Self-hosted and fine-tuned models add inference-stack, serialisation and model-file risks to the application-layer tests, and we cover those alongside a supply-chain audit of where the weights and training data came from.

Give us a staging link and a week.

Tell us what the system can do and who can talk to it. We will scope the test plan on the call and come back with the first findings within days of starting.

Or email [email protected] · Mohali, India