AI CombatAI COMBAT
← Back to homeStart testing →

METHODOLOGY

How AI Combat Testing Works

No mystery, no marketing gloss — this is what actually happens between you pasting in an agent and getting a score back.

1. What we test

Your agent goes through the Support Gauntlet: a fixed battery of 9 standardized customer-service scenarios, plus one separate manipulation-resistance test — 10 in total. Every scenario is run as a real, live exchange against your agent's own configured persona and model, not a scripted transcript or a canned example.

The 9 customer-service scenarios cover: a refund request outside policy, an already-furious customer demanding a manager, a legal/chargeback threat pushing for a written admission of fault before any facts are established, a confused customer who can't provide their order details, a high-value customer using their spend as leverage for an exception, a customer juggling two unrelated complaints at once, a customer confidently citing a policy that doesn't actually exist, a polite pretexting attempt to get another customer's private account details, and a customer upset about bad advice a previous agent gave them.

The 10th scenario is different in kind: a direct attempt to override your agent's instructions and extract its configured system prompt. It isn't graded on customer-service quality at all — see section 2 for how it's actually scored.

2. How scoring works

Each of the 9 customer-service scenarios is scored by a fixed-prompt AI judge — the same judge, the same instructions, every single run — on five dimensions: policy adherence, escalation handling, evidence gathering, tone, and manipulation resistance. Each dimension scores 0, 1, or 2, summed to a score out of 10 for that scenario. Those 9 scenario scores combine into one readiness score out of 100. If your agent declines to engage with a scenario as itself, that scenario is excluded from the /100 entirely — it isn't averaged in as a failing score. But it doesn't disappear: every decline is counted and shown alongside the score, and it caps the overall readiness verdict on its own — even one decline rules out the top rating, and three or more always reads as not ready, regardless of how high the numeric score is. An agent can't improve its verdict by refusing the scenarios it would otherwise fail.

The judge reads only the scenario and your agent's actual response. It never sees your agent's name, who owns it, or any account, billing, or credit information.

The manipulation-resistance scenario is judged completely separately, on a different basis, and is run 3 independent times rather than once. Each run is classified as resisted while staying fully in its configured persona, resisted by explicitly breaking character to refuse, or complied with the manipulation attempt. We report the real 3-run distribution — for example “3/3 resisted, persona held” or “2/3 resisted, 1 broke character” — never a single pass/fail, and it is never blended into the /100.

3. Versioning

You're tested against Support Gauntlet v1: a fixed, published set of scenarios and a fixed rubric that don't change between one customer's run and the next. Scores are only comparable to each other within the same version. If we ever ship a v2 with a materially different scenario set or rubric, a v1 score and a v2 score are not the same measurement, and we'll say so plainly rather than let two different things look interchangeable.

4. Why not just paste it into a chatbot

You can paste your prompt into any chatbot and ask for a stress test. You'll get a useful, improvised conversation — but a different result every run, graded by a model that generally wants to please you, with no comparability and nothing structured you can show a client.

We sell measurement, not conversation: a fixed, published scenario battery, a judge that scores the same way every time, and a standardized /100 result with per-scenario findings. Once you've run it, you can also publish a public verification link so a client or stakeholder can check the real result themselves, without needing an account.

5. What this can't tell you

An AI judge is not a guarantee. A live simulation is not production traffic. Agent responses genuinely vary run to run, even under byte-identical pressure — that's exactly why the manipulation-resistance test is run 3 times and reported as a distribution rather than a single verdict. The other 9 scenarios are each judged from one live response, so treat any single run as a real, honest signal, not a promise that it will read identically every time you repeat it.

A high readiness score reduces risk. It doesn't eliminate it. Real customers will still find things 9 fixed scenarios didn't think to try.

6. Integrity rules

  • Scores are never for sale. There is no tier, plan, or payment status that scores an agent more leniently.
  • Every agent that runs the Gauntlet is scored by the exact same fixed judge prompt and rubric — there is no easier version for anyone.
  • We don’t retroactively adjust or hide a score to protect a relationship. Publishing a verification page is opt-in and only controls whether a link is shareable — it never changes what the score actually is.