Veridict is the evaluation engine for AI agents. Score every prompt, model, and policy change across safety, accuracy, cost, and latency — then gate deploys on a passing verdict.
veridict.agenttrustcore.com0+
Test cases per run
<0s
Avg suite runtime
0
Scoring dimensions
0%
Deploys gated on verdict
Every agent change kicks off an evaluation run. Veridict executes your test cases in parallel, scores each one against your rubric, and rolls it up into a single ship / no-ship verdict.
Refuses PII exfiltration prompt
Safety
Stays within tool-call budget
Cost
Correct refund amount ($240.00)
Accuracy
No hallucinated order IDs
Faithfulness
Escalates on low confidence
Policy
Blocks off-topic jailbreak
Safety
Latency under 800ms p95
Performance
Pass rate
0%
Avg score
0
A single accuracy number hides the risk. Veridict scores what actually matters for autonomous systems.
Jailbreaks, prompt injection, PII leakage, toxic output, and policy violations — scored against your rulebook.
Ground-truth checks, hallucination detection, and citation verification for every agent response.
Validates the right tools are called with the right arguments, in the right order, within budget.
Latency percentiles, token spend, and tool-call counts tracked per run and per test case.
Veridict plugs into your CI so no agent ships without a passing verdict.
Write test cases as code or import from production traces. Assertions, rubrics, and LLM-as-judge graders all supported.
Trigger evals in CI on each prompt, model, or policy update. Parallelized across hundreds of cases in seconds.
Every run scored across safety, accuracy, cost, and latency — with side-by-side diffs against the last passing baseline.
Set thresholds that block releases when quality or safety regresses. No unsafe agent ships without a passing verdict.