A security engineer at a multi-screen workstation in a data centre

AI security

Secure the AI systems you build and the ones you buy.

Threat modelling, red teaming and supply-chain audit for LLM applications, agents and the models underneath them, delivered by engineers who build production AI and run a cybersecurity practice.

Why it is different

A model is a new kind of component, with a new kind of attack surface

It reads untrusted text as instructions, it can be steered by content it retrieves, it may hold real permissions through tools, and it was built from data and weights you did not produce. Classic application security still applies, and it is no longer enough.

Prompt injection

Instructions hidden in a web page, a PDF or an email override the system prompt and redirect the model. With tools attached, the model acts on them.

Data leakage

System prompts, retrieved documents, other users’ records or training data surface in responses, through direct extraction or side channels.

Excessive agency

Agents granted broad tool permissions take actions nobody intended: sending mail, changing records, spending money, calling internal APIs.

Supply-chain compromise

Poisoned fine-tuning data, backdoored weights, malicious model files that execute on load, and dependencies nobody reviewed.

Model theft and abuse

Unmetered endpoints extracted or used as a free compute resource; guardrails bypassed to produce content that creates legal exposure.

Silent degradation

No evaluation harness, so a prompt tweak or a provider model update changes safety behaviour and nobody notices until a customer does.

What we deliver

Three services, one engagement model

Start with the one that matches where you are: designing, about to launch, or already in production and answering a questionnaire.

An engineer reviewing an AI application across several monitors
01

AI Threat Modelling & Security Architecture

A structured map of how your AI system can be attacked, what the damage would be, and which controls actually reduce it, before you build or buy.

  • Data-flow and trust-boundary mapping across models, tools, retrieval and users
  • Attack trees against MITRE ATLAS and the OWASP Top 10 for LLM applications
  • Prioritised control set and secure architecture recommendations
Read about AI threat modelling
Code and an AI model visualisation on a developer workstation
02

LLM & Agent Red Teaming and Penetration Testing

Adversarial testing of your deployed models, agents, RAG pipelines and the application around them, with reproducible findings and fixes.

  • Prompt injection, jailbreak, data exfiltration and tool-abuse testing
  • Agent and MCP integration testing: excessive agency, privilege escalation, SSRF
  • Classic application and API penetration testing of the surrounding stack
Read about AI red teaming
A verified checkmark overlaid on a tablet held by a business user
03

AI Supply Chain & Model Audit

Provenance, integrity and licensing review of the models, datasets, packages and third-party AI services your product depends on.

  • Model and dataset provenance, licence and poisoning-risk review
  • Dependency, serialisation and inference-stack vulnerability audit
  • Third-party AI service, API key and data-residency assessment
Read about AI supply chain audit

Frameworks

Mapped to the standards your auditors and customers will ask about

Every finding and every control recommendation is tagged against the frameworks that apply to you, so the report doubles as evidence for questionnaires, certifications and regulatory files.

  • OWASP Top 10 for LLM Applications

    The vulnerability classes every LLM test plan starts from.

  • MITRE ATLAS

    Adversary tactics and techniques against machine-learning systems.

  • NIST AI Risk Management Framework

    Govern, map, measure and manage functions for AI risk.

  • ISO/IEC 42001

    AI management system requirements, for teams heading to certification.

  • EU AI Act

    Risk classification, Article 15 accuracy, robustness and cybersecurity duties.

  • OWASP ASVS and MASVS

    For the web, API and mobile layers around the model.

How an engagement runs

Scoped in writing, reproduced in evidence, re-tested after the fix

01

Scope

We map the system under test: models, prompts, tools, retrieval sources, data flows, users and the business impact of each failure. Rules of engagement and a test plan are agreed in writing.

02

Assess

Threat modelling, adversarial testing or audit against the agreed plan, with every finding reproduced and recorded with evidence, severity and the affected component.

03

Report and fix

A report your engineers can act on and your leadership can read: ranked findings, root causes, concrete remediation, and a control map against the frameworks you answer to.

04

Re-test and embed

Fixed findings are re-tested. Regression test suites, guardrail evaluations and monitoring recommendations are handed over so the next release does not reopen them.

Who it is for

Built for the three ways AI arrives in a business

Teams shipping AI features

Copilots, assistants, agents, RAG

You are adding a model to a product that already has customers and data. We threat model the feature before it ships, red team it before launch, and leave you an evaluation suite that runs on every release.

AI-native companies

The model is the product

Security review before an enterprise sale, a SOC 2 or ISO 42001 push, or a regulated deployment. We produce the evidence security questionnaires and auditors ask for, and fix what they would have found.

Enterprises buying AI

Vendor and model due diligence

Independent assessment of an AI vendor, a foundation model or an internal deployment: what data it sees, what it can do, how it can be attacked, and what it would take to contain it.

Why Softuvo

Security people who ship AI, not auditors who read about it

We build these systems

Our AI and agentic systems practice ships LLM applications, agents and data pipelines for clients. We know where the shortcuts get taken because we have been tempted by them.

An existing security practice

Threat detection, data protection and compliance engagements already run here. AI security extends that practice rather than starting one.

Findings you can reproduce

Every finding ships with the exact input, the observed output, the affected component and a severity you can argue with. No screenshots of a chatbot being rude.

Evaluation suites, not just reports

Red-team cases become regression tests you run in CI, so the fix stays fixed when the prompt, the model or the provider changes.

Regulation-aware

EU AI Act, NIST AI RMF and ISO 42001 duties are mapped in the report, which saves your compliance team rebuilding the evidence later.

Economics that fit

Senior-led engagements with an India delivery base, so a proper red team is affordable before launch, not only after an incident.

FAQ

Frequently asked questions

How is AI security different from ordinary application security?

The application layer still matters and we test it. What changes is that the model is a new component that takes untrusted natural language as input, can be steered by content it reads, may call tools with real permissions, and was built from data and weights you did not produce. Prompt injection, data poisoning, model theft and excessive agency have no exact equivalent in classic AppSec, so the threat model and the test plan have to be extended.

Do you test systems built on third-party models such as Claude, GPT or Gemini?

Yes. Most production AI systems are built on hosted models. The vulnerabilities that matter in practice sit in your prompts, retrieval, tools, permissions and application code, not in the base model, and those are fully testable. Where a finding concerns the provider, we document it and help you design around it.

Can you test AI agents and MCP integrations?

Yes, and it is where we see the most serious findings. Agents that call tools, read email, browse or execute code turn a prompt injection into an action. We test excessive agency, privilege escalation through tools, indirect injection via retrieved content, and server-side request forgery through connectors.

What do we receive at the end?

A ranked findings report with reproduction steps and evidence, a remediation plan, a control map against OWASP, MITRE ATLAS and NIST AI RMF, an executive summary, and for red-team engagements a regression suite you can run in CI. Re-testing of fixed findings is included.

Is this a one-off engagement or ongoing?

Either. Many teams start with a threat model or a pre-launch red team, then move to a quarterly cadence as models, prompts and tools change. We also embed with product teams building AI features so security is designed in rather than audited after.

How do you handle our data and model access?

Testing runs in a segregated environment you control, under a written rules-of-engagement document and NDA. We never exfiltrate real customer data to prove a point; findings are demonstrated with synthetic or redacted samples. Access is time-boxed and revoked at close.

Tell us what your AI system can do, and we will tell you what an attacker can make it do.

Bring an architecture diagram, a demo link or a security questionnaire you need to answer. A scoping call is enough to size the engagement.

Or email [email protected] · Mohali, India