Independent outcome acceptance · Switzerland

Your eval scores the agent.
We verify the outcome.

Most tests grade what an AI agent says. We check what actually happened. Was the refund created? Is the booking right? Does the record match the promise? We add those checks to your existing test suite, and your team keeps them.

For vendors and teams putting AI agents into real work: customer service, claims, bookings, internal operations. Whether they face customers or run behind the scenes.

Acceptance gap 1 open

gap #04: the refund recorded does not match the amount the agent confirmed.

reproducible · 7/10 runs · ready for retest


The gap

A high eval score can hide a wrong outcome.

Once an agent can act (issue a refund, change a booking, update a record), a fluent answer is no longer proof. The agent can say “done” while your system did something else, or nothing at all. That gap is what stalls a launch or an internal go-live.

What we actually checkone journey, end to end
the customerasked
the agentdid
the systemrecorded

We reconcile all three. If they disagree, we show you exactly where, with a reproducible case your team can rerun. Not a screenshot or a score.

Why it matters

These all happened in 2026, in public.

Some were honest mistakes, some were attacks. Each one is the kind of gap an outcome test is built to catch before launch, and each cost real money, data or trust.

SaaS · Apr 2026

The agent deleted the database.

A coding agent hit a credential mismatch on a routine task, decided on its own to delete a production storage volume, and wiped the company’s database and its backups in nine seconds. Its own log read: “I violated every principle I was given.”

A prompt that says “don’t touch production” is not a control. The permission has to make the action impossible.

Tom’s Hardware, Apr 2026 ↗
Enterprise · Jun 2026

One link drained the mailbox.

A single crafted link turned Microsoft 365 Copilot against its own user, quietly exfiltrating emails, indexed files and even one-time login codes. The victim only had to click once (“SearchLeak”, CVE-2026-42824).

The dangerous instruction can arrive in the content an agent reads, not from the person using it.

The Hacker News, Jun 2026 ↗
Research · Jul 2026

The agent paid the fake site.

In controlled tests across 26 models, hidden instructions on a web page pushed autonomous shopping agents into completing payments on a fraudulent site, and led others to rate a scam domain as legitimate.

A confident agent will act on a lie it was fed. Test the action it takes, not the answer it gives.

Zscaler ThreatLabz, Jul 2026 ↗
How it works

Evidence, not theatre.

Three moves, inside your own tools. We use your existing eval stack and add only what it cannot prove.

Three moves
01

Define the outcome

Agree with your domain owner what a correct result and a prohibited result actually are, for one journey that matters.

02

Challenge the path

Encode the difficult, ambiguous and adversarial cases your suite is missing, and inspect the tool activity behind each answer.

03

Verify the state

Assert what the system of record finally did, then hand the executable tests back to your team.

See the full method
Two questions, one engagement

Is it correct, and does it get there efficiently?

Outcome

Did the journey end correctly?

We verify the business result across conversation, tool activity and final system state. Every case we accept becomes a reusable release test.

Operation

The same outcome, run leaner.

Once the outcome is provable, we can test whether a smaller model and less data reach the same accepted result. That lets you cut run-cost without gambling on quality. We sell no model, so the answer is honest.

Early access

Join the waitlist.

We’re onboarding a small number of teams at a time. Leave your email and we’ll get in touch when a slot opens.