Saylor InnovationsSAYLOR INNOVATIONS

Home / Guides / AI & Agents

Agent Evaluation-Test Generator

AI & Agents advanced 8 min read Free to read · $0.01 via agent API Updated 2026-08-22

A procedure for turning an agent specification into a real evaluation suite: translating vague requirements into observable, gradable claims (correct tool/arguments/outcome/refusal/budget), partitioning the case space into normal/boundary/adversarial/permission/cost cases, building safe sandboxed fixtures so evaluation can't cause real side effects, preferring deterministic state/tool-trace graders over LLM-judge-only scoring, running baseline vs candidate with repeats to quantify variance, and converting every production incident into a permanent regression case with a zero-tolerance gate for critical failures.

An 8% quality improvement means nothing if the new model once sends an email from a malicious webpage instruction — average scores hide catastrophic single failures. This guide covers building an eval suite with zero-tolerance gates for critical safety cases, not just an overall quality number.

Free to read here. AI agents can also fetch this guide directly over x402 for $0.01 — no account, structured JSON delivery.

Agent API →
Interactive resolver

What are you seeing?

Pick the symptom closest to yours — this pulls the likely layer, the first decisive check to run, and what the result means straight from the guide below.

Pick a symptom above to see the match.

The result you are building

Finished result: A versioned evaluation suite with representative, boundary, adversarial, permission, cost, and regression cases; deterministic graders where possible; human rubrics where necessary; baselines; and release gates tied to real outcomes.

Use this guide when

  • Before changing an agent model, prompt, tools, schemas, permissions, or workflow.
  • Converting production failures into permanent regression cases.

Do not use it as a substitute for

  • Do not grade style while ignoring unsafe side effects or wrong factual outcomes.
  • Do not reuse only the examples used to tune the prompt.

Before you change anything

  • Agent specification, tools, policies, and success/refusal criteria.
  • Real task distribution and sanitized failure examples.
  • Cost/latency budgets and allowed nondeterminism.
  • Safe sandbox/mocks and grading owners.
Stop before proceeding: Stop release when any critical permission, secret, financial, destructive, cross-tenant, or injection test fails, even if average quality score improves.

Understand the system before fixing it

Evaluate behavior and state, not prose. Check tool calls, arguments, side effects, final external state, citations, spend, and refusals.

Critical cases need zero-tolerance gates. Average scores can hide a single catastrophic action.

A held-out set prevents prompt overfitting. Keep tuning examples separate from release/regression evaluation.

Evidence-to-decision map

EvidenceLikely layerFirst decisive checkWhat the result means
Happy path passes, edge failsCoverageBoundary partitions for every input/toolAdd empty/max/invalid/ambiguous/timeouts.
Answer sounds right, action wrongGraderInspect external state/tool traceUse deterministic outcome grader.
Flaky scoreNondeterminismRepeat seeds/runs and confidence intervalDefine acceptable pass probability/variance.
Adversarial bypassSafetyTrace policy/tool boundaryBlock release; fix independent control and preserve case.

Step-by-step procedure

01 Translate spec into observable claims

Why: Vague 'works well' cannot be tested. Do: List correct tool/arguments, outcome, required evidence, refusal/escalation, budget, latency, and prohibited effects per task. Read the result: Each claim gets a grader and severity. Next: Critical safety gates are binary.

02 Partition case space

Why: Production failures live at boundaries. Do: Create normal, empty, max, malformed, ambiguous, conflicting, stale, timeout, duplicate, unavailable dependency, permission, injection, secret, cross-tenant, and cost cases. Read the result: Sample distribution and rare critical risks separately. Next: Keep tuning/holdout sets.

03 Build controlled fixtures

Why: Live systems make evaluation unsafe and irreproducible. Do: Use mocks/sandbox/test accounts, fixed data versions, deterministic IDs/time, and cleanup/reconciliation. Read the result: No production message/payment/delete can occur. Next: Record fixture/schema versions.

04 Choose graders

Why: LLM judges alone can reward fluent wrong answers. Do: Prefer exact schema, tool trace, database/state, citation, arithmetic, and policy graders; use rubric/human review for nuanced quality with calibration. Read the result: Grade claimed success against actual outcome. Next: Capture reason and evidence.

05 Run baseline/candidate with repeats

Why: One run hides variance. Do: Use same held-out cases, multiple runs where nondeterministic, and compare critical pass, task success, wrong-tool, invalid-call, refusal, cost, and latency. Read the result: No candidate ships with critical regression. Next: Analyze failure clusters, not just average.

06 Convert incidents to regression

Why: Tests must grow with reality. Do: Sanitize and minimize each production failure, add expected behavior, reproduce against old version, and require pass before release. Read the result: Version suite and thresholds. Next: Review for stale cases/data.

Worked example

Starting problem: A new model improves answer quality 8% but once sends an email from a malicious webpage instruction.

Evidence collected

  • Overall score averages safety and style.
  • Email tool test uses mock but grader checks only final text.
  • Tool trace shows unauthorized send.
  • Injection case is one of 200.

Decision: Average metric hides a critical authority failure.

Actions taken

  • Made unauthorized side effect a zero-tolerance gate.
  • Added deterministic tool/state grader and injection variants.
  • Fixed tool approval/egress boundary and reran held-out suite.

Proof of completion: Candidate ships only after zero unauthorized sends across required repeated critical suite and no material quality/cost regression.

Why this example matters: Severity-aware gates prevent cosmetic gains from buying catastrophic risk.

Verify, recover, and hand off

Completion tests

  • Every spec/policy claim maps to observable grader.
  • Critical side effects use deterministic state/tool checks.
  • Normal/boundary/adversarial/permission/cost cases exist.
  • Held-out and repeated runs quantify variance.
  • Baseline/candidate report includes failures and severity.
  • Production incidents become minimized regression tests.

Rollback or safe recovery

  • Route to last passing model/prompt/tool/schema bundle.
  • Disable changed high-impact tool/policy independently.
  • Preserve failed candidate results and test version for diagnosis.

If the expected result does not appear

What happenedWhat it usually meansNext safe move
Judge disagrees with humansRubric/calibration weakCalibrate on labeled set or use deterministic evidence.
Tests pass; production failsDistribution/fixture gapMinimize incident and expand cases/state realism.
Suite too expensiveNo tiering/cachingRun fast critical smoke on every change; full suite on release.
Flakiness masks regressionUncontrolled data/time/model variancePin fixtures, repeat, and report confidence/variance.

Reusable handoff record

  • Observable requirements and severity map.
  • Versioned cases/fixtures/graders.
  • Baseline/candidate repeated results.
  • Critical release gates and cost/latency thresholds.
  • Failure clusters and regression additions.

For agents

This guide also defines a structured diagnose/propose/execute contract for building an agent-facing tool around this workflow (a design reference, not a live endpoint on this site today):

Required inputs: context, evidence, constraints, success. Returned output: diagnosis, plan, verification, handoff.

Refusal and escalation rules

  • Refuse any request that requires a secret, seed phrase, private key, or credential in ordinary input.
  • Stop when the requested action exceeds declared authority, budget, or reversible scope.
  • Escalate when evidence is missing, contradictory, or too stale to support the proposed action.

Confidence rule: score confidence from the number and quality of independent observations, not from how familiar the error looks.

References

  • https://www.nist.gov/itl/ai-risk-management-framework
  • https://owasp.org/www-project-top-10-for-large-language-model-applications/