Your eval scores the agent.
We verify the outcome.
Most tests grade what an AI agent says. We check what actually happened. Was the refund created? Is the booking right? Does the record match the promise? We add those checks to your existing test suite, and your team keeps them.
For vendors and teams putting AI agents into real work: customer service, claims, bookings, internal operations. Whether they face customers or run behind the scenes.
gap #04: the refund recorded does not match the amount the agent confirmed.
reproducible · 7/10 runs · ready for retest
A high eval score can hide a wrong outcome.
Once an agent can act (issue a refund, change a booking, update a record), a fluent answer is no longer proof. The agent can say “done” while your system did something else, or nothing at all. That gap is what stalls a launch or an internal go-live.
We reconcile all three. If they disagree, we show you exactly where, with a reproducible case your team can rerun. Not a screenshot or a score.
These all happened in 2026, in public.
Some were honest mistakes, some were attacks. Each one is the kind of gap an outcome test is built to catch before launch, and each cost real money, data or trust.
The agent deleted the database.
A coding agent hit a credential mismatch on a routine task, decided on its own to delete a production storage volume, and wiped the company’s database and its backups in nine seconds. Its own log read: “I violated every principle I was given.”
A prompt that says “don’t touch production” is not a control. The permission has to make the action impossible.
Tom’s Hardware, Apr 2026 ↗One link drained the mailbox.
A single crafted link turned Microsoft 365 Copilot against its own user, quietly exfiltrating emails, indexed files and even one-time login codes. The victim only had to click once (“SearchLeak”, CVE-2026-42824).
The dangerous instruction can arrive in the content an agent reads, not from the person using it.
The Hacker News, Jun 2026 ↗The agent paid the fake site.
In controlled tests across 26 models, hidden instructions on a web page pushed autonomous shopping agents into completing payments on a fraudulent site, and led others to rate a scam domain as legitimate.
A confident agent will act on a lie it was fed. Test the action it takes, not the answer it gives.
Zscaler ThreatLabz, Jul 2026 ↗Evidence, not theatre.
Three moves, inside your own tools. We use your existing eval stack and add only what it cannot prove.
Define the outcome
Agree with your domain owner what a correct result and a prohibited result actually are, for one journey that matters.
Challenge the path
Encode the difficult, ambiguous and adversarial cases your suite is missing, and inspect the tool activity behind each answer.
Verify the state
Assert what the system of record finally did, then hand the executable tests back to your team.
Start small. Add cover only when a change demands it.
Fixed scope, agreed up front. Every engagement leaves reusable tests with your team. No platform, no lock-in.
Gap Review
We review your existing tests against one journey and name, in writing, what stays unproven.
Core · two weeksOutcome Acceptance Sprint
One agent, one journey. We close the acceptance gap and leave the executable tests with your team.
Recurring · on changeRelease Challenge
We independently re-run the high-consequence suite whenever something material changes.
Is it correct, and does it get there efficiently?
Did the journey end correctly?
We verify the business result across conversation, tool activity and final system state. Every case we accept becomes a reusable release test.
The same outcome, run leaner.
Once the outcome is provable, we can test whether a smaller model and less data reach the same accepted result. That lets you cut run-cost without gambling on quality. We sell no model, so the answer is honest.
Join the waitlist.
We’re onboarding a small number of teams at a time. Leave your email and we’ll get in touch when a slot opens.
We use your address only to contact you about working with Arctiq. No newsletter spam. See the privacy notice.