Start a pilot

Six weeks. One agent, one rubric, three runs.

We write a rubric against your requirements, run conversations against your agent in a sandbox, and return findings for each run. Three runs, so you see the same rubric across two changes to your agent instead of a single snapshot.

Fixed scope and a fixed price, which we give you on the first call.

What we need from you

Three separate inputs: access, context, and requirements.

  1. 01

    Access

    A sandbox or test environment where we can exercise the agent's email behavior without contacting real recipients. We do not need production transcripts and we do not touch production.

  2. 02

    Context

    The agent's job, who it communicates with, what a normal workflow looks like, and what a successful interaction is supposed to accomplish.

  3. 03

    Requirements

    Written policies, behavioral commitments, prohibited behavior, required language, known failure modes, or customer promises you want us to test.

Four questions

If we do not think we are a fit, we will say so instead of booking a call.

Does your agent initiate contact by email or text?
Do you change the model or system prompt at least monthly?
Have you been through an enterprise security or vendor review in the last six months?
Has a deal been delayed, or an escalation raised, over how your agent behaved?