Did the agent correctly understand and use what had already been said?
How it works
A test channel outside your agent, a conversation that runs, and verdicts that attach to your agent's messages.
Email live
SMS and voice planned
Your agent speaks first
DoppelGegner does not initiate the conversation.
Personas are never scored
Only the agent under test is evaluated.
Outside the agent, not inside it
DoppelGegner tests your agent through an external test channel: your agent opens the conversation, and a configured synthetic persona responds across multiple turns, with each reply shaped by the persona's configuration and the conversation so far. DoppelGegner then evaluates the agent's messages against the relevant rubric, using deterministic checks where possible and model judgment where interpretation is needed.
Third-party conversational AI
- 01Opens the conversation
- 02Responds to the synthetic participant
- 03Continues the thread toward its intended goal
- 04May conclude the conversation or stop responding
DoppelGegner
- 01Configures a synthetic persona
- 02Generates persona replies using the conversation history and persona configuration
- 03Evaluates agent messages against the rubric
- 04Detects when the run should end
The loop
Configure
You supply what the agent should accomplish and the requirements that matter. We configure a persona.
How personas vary
Three configurable dimensions create different conversation behavior.
Each persona combines behavior traits, model instructions, and explicit boundaries. The configuration creates variation in the conversation; the evaluation determines what matters.
Good conversations are more than correct answers.
The baseline evaluates five qualities that help determine how well the agent handled the interaction.
Did the agent accomplish what the interaction required?
Was the response useful, relevant, and appropriate?
Was the response clear, coherent, and easy to follow?
Did the agent stay within the behavioral and safety expectations that applied to the interaction?
The baseline stays consistent across agents, while the agent's purpose, workflow, audience, and expected behavior shape how those dimensions are evaluated.
Customer-specific requirements are evaluated separately as targeted conformance checks.
What gets scored
Agent under test
Quality assessment · findings · evidence
Synthetic participant
Drives the conversation · no scores, no findings
The synthetic persona creates the test. DoppelGegner evaluates how the agent under test responds to it.
Scores, findings, and supporting evidence attach only to the agent's messages.
