Start small. Add cover only when a real change demands it.
Fixed scope, inside your own tools, for agents that face customers or run inside your operation. Every engagement leaves reusable tests with your team.
Eval-to-Outcome Gap Review
We review your existing test set against one journey and the release decision, and name what stays unproven.
- a written acceptance-gap statement
- a go / no-go on a full sprint
- best when the gap isn’t yet scoped
Outcome Acceptance Sprint
One agent, one journey. We close the acceptance gap and leave the tests with your team.
- eval-gap map & outcome contract
- executable cases + final-state assertions
- reproducible evidence, ranked by impact
- customer-owned tests + one retest
Managed Release Challenge
We independently re-run the high-consequence suite when something material changes: a model, a prompt, a policy, a tool.
- change-triggered reruns
- production failures added as tests
- a release delta each cycle
Fixed scope, agreed up front. You keep the tests either way. No lock-in, no automatic renewal. Pricing is set per engagement after a short scoping call.
We are honest about when to call us.
- an agent, customer-facing or internal, goes live or changes within 90 days
- it retrieves account data or triggers a business workflow
- an enterprise customer needs acceptance evidence
- a staging path and an accountable owner exist
- it only answers low-consequence FAQs
- a human approves every material action and QA is enough
- there is no named release decision
- you need certification or a generic pentest
Every engagement leaves an asset behind, not just a report.
Whatever the scope, you walk away with executable tests your own team can rerun on the next release. Nothing is locked to us.
Illustrative. ○ observed · ● verified · ■ final state · • caveat.
What teams ask before they call us.
“We already have evaluation tools.”
Good. We expect that, and we use them. We add value only where an important business outcome, exception or final-state check remains unproven. If nothing does, we say so.
“The vendor should test its own agent.”
It should. We come in where supplier and enterprise need independent evidence, or share an acceptance question neither side can settle alone.
“We can’t share production data.”
We default to staging, test identities, synthetic or minimised data, and customer-controlled access. If the required evidence can’t be reached safely, we narrow the scope and state that limit in the report.
“Can you certify the agent?”
No. We report what was tested and what was observed, with the remaining uncertainty stated. Legal classification and certification are outside our scope, on purpose.
We onboard a small number of teams at a time.
Join the waitlist and we’ll get in touch when a slot opens.