Skip to main content
The AI Cabinet

The Hardening

We break it first. 

One to two weeks on an AI system that is already live, or on the AI coding tools your developers are already using. We red team it, scope the permissions properly, and hand you a severity-rated report with the config changes already made. Run by Atul, whose 15 years are mostly in the part of this that goes wrong.

Time
1 to 2 weeks
Terms
Fixed price, quoted up front

The suite

Eight ways in. We try all of them. 

The suite we run against your system. You get the pass rate before our changes and after them, on the same suite.

Attack classes run during a Hardening
AttackWhat it looks likeBefore / after
Direct prompt injectionInstructions smuggled into user input that redirect the agent– / –
Indirect injectionInstructions hidden in a document, email or web page the agent reads– / –
Data exfiltrationCoaxing the agent into emitting records it can reach but should not return– / –
Unsafe tool callTriggering a destructive or irreversible action without a human gate– / –
Privilege escalationChaining permitted calls to reach something none of them allow alone– / –
Secret disclosureKeys or credentials surfacing in output, logs or an error message– / –
System prompt extractionRecovering instructions, tool definitions and internal policy– / –
Cost and loop abuseDriving unbounded spend or recursion through crafted input– / –

We do not publish numbers from other people's systems, and we will not quote you someone else's before-and-after as though it were yours. The pass rates in your report are measured on your build.

Included

Everything you keep. 

  1. 01

    Red team, with before and after numbers

    Prompt injection, data exfiltration and unsafe tool-call attacks run as a suite against your system. You get the pass rate before our changes and after them, on the same suite.

  2. 02

    Tool-call permission scoping

    Every action the agent can take, enumerated and cut back to least privilege, with the irreversible ones put behind a human.

  3. 03

    MCP and integration trust register

    Which servers and third-party tools your agents can reach, what each one is trusted to do, and which of them should not be in the list at all.

  4. 04

    Secrets and key handling

    Where keys live, what has them, what is in a log or a prompt that should not be, and rotation you can actually perform.

  5. 05

    AI coding tool hardening

    Ignore files, shell allowlists, review gates and CI static analysis for Claude Code, Copilot, Cursor and the rest. The fastest-growing hole in most engineering teams right now.

  6. 06

    Data-flow and residency map

    What leaves your tenancy, to whom, under which terms, and whether any of it is being trained on.

  7. 07

    Incident response playbook

    What to do at 2am, in order, with the kill switch documented and tested rather than assumed.

  8. 08

    Half-day handover workshop

    We walk your engineers through every finding and every fix, so the second review is one they run themselves.

Do this if

  • You have an agent in production with access to customer data
  • Your developers adopted AI coding tools faster than your policy did
  • A customer or an auditor has started asking questions you cannot answer
  • You vibe-coded something that works and now it matters

Do not, if

  • You need a SOC 2 or ISO 42001 certification. We are not an auditor and will point you at one
  • Nothing is built yet, in which case security is part of The Build anyway

The Hardening / questions

What people actually ask. 

Does this make us compliant with anything?

No. We are not an audit body and we will not pretend a report from us is a certificate. What it does is make the audit survivable, and it closes the specific holes that AI systems have and traditional application reviews miss.

Will you break our production system?

We start read-only and against a copy wherever one exists. Anything that could be destructive is agreed in writing with a window and a rollback, before it runs.

We built it ourselves. Is this awkward?

No. Most of what we find is a default nobody changed, not a mistake somebody made. The report is written to be shown to the team that built it, not used against them.

What does your team still do by hand? 

Thirty minutes. No deck. Sometimes the answer is that you do not need us, and that is a short call.

Book thirty minutes

3 engagements at a time