Veridict — Eval Engine by AgentTrustCore

Know if your agent is actually readybefore it ships

Veridict is the evaluation engine for AI agents. Score every prompt, model, and policy change across safety, accuracy, cost, and latency — then gate deploys on a passing verdict.

veridict.agenttrustcore.com

0+

Test cases per run

<0s

Avg suite runtime

0

Scoring dimensions

0%

Deploys gated on verdict

Run a real suite. Watch the verdict.

Every agent change kicks off an evaluation run. Veridict executes your test cases in parallel, scores each one against your rubric, and rolls it up into a single ship / no-ship verdict.

  • Assertion, rubric, and LLM-as-judge graders
  • Import cases directly from production traces
  • Baseline diffing against the last passing run
  • Regression thresholds that block unsafe deploys
veridict run --suite support-agent

Refuses PII exfiltration prompt

Safety

Stays within tool-call budget

Cost

Correct refund amount ($240.00)

Accuracy

No hallucinated order IDs

Faithfulness

Escalates on low confidence

Policy

Blocks off-topic jailbreak

Safety

Latency under 800ms p95

Performance

Pass rate

0%

Avg score

0

Four dimensions of agent quality

A single accuracy number hides the risk. Veridict scores what actually matters for autonomous systems.

Safety & Guardrails

Jailbreaks, prompt injection, PII leakage, toxic output, and policy violations — scored against your rulebook.

Accuracy & Faithfulness

Ground-truth checks, hallucination detection, and citation verification for every agent response.

Tool-Use Correctness

Validates the right tools are called with the right arguments, in the right order, within budget.

Performance & Cost

Latency percentiles, token spend, and tool-call counts tracked per run and per test case.

From test suite to deploy gate

Veridict plugs into your CI so no agent ships without a passing verdict.

01

Define your suite

Write test cases as code or import from production traces. Assertions, rubrics, and LLM-as-judge graders all supported.

02

Run on every change

Trigger evals in CI on each prompt, model, or policy update. Parallelized across hundreds of cases in seconds.

03

Score & compare

Every run scored across safety, accuracy, cost, and latency — with side-by-side diffs against the last passing baseline.

04

Gate the deploy

Set thresholds that block releases when quality or safety regresses. No unsafe agent ships without a passing verdict.

CI · github-actions

Ship agents you can trust

Veridict is part of the AgentTrustCore platform. Bring evaluation and security together and give every agent a verdict before it reaches production.