Saylor InnovationsSAYLOR INNOVATIONS

Home / Guides / AI & Agents

Embedding and Retrieval Evaluation

AI & Agents advanced 9 min read Free to read · $0.01 via agent API Updated 2026-08-22

A retrieval-benchmark procedure for RAG/semantic-search/agent-memory systems: build a judged query set covering hard negatives and no-answer cases, snapshot the whole retrieval stack for reproducibility, trace every stage, measure task-specific metrics (not one aggregate score), and canary in production with automatic rollback.

Evaluate lexical, vector, hybrid, reranked, and filtered retrieval against task-specific relevance, freshness, coverage, latency, and leakage tests — because good-looking demos rarely predict production question coverage.

Free to read here. AI agents can also fetch this guide directly over x402 for $0.01 — no account, structured JSON delivery.

Agent API →

Evaluate lexical, vector, hybrid, reranked, and filtered retrieval against task-specific relevance, freshness, coverage, latency, and leakage tests.

The result you're building

A versioned retrieval benchmark and release gate that measures whether the correct evidence is found, ranked, scoped, and cited for representative questions, including hard negatives and permission boundaries.

Use this guide when

  • You are building RAG, semantic search, recommendations, or agent memory retrieval.
  • Embedding model, chunking, index, filters, or reranker changes.
  • Good-looking demonstrations do not predict production question coverage.

Don't use it as a substitute for

  • Choosing a model from a public leaderboard without your own corpus and tasks.
  • Evaluating only answer fluency instead of whether the supporting evidence was retrieved.

Before you start

Collect these first:

  • Versioned corpus snapshot, permission model, chunking, metadata, and index settings.
  • Representative query set with relevance judgments and acceptable evidence.
  • Hard negatives, no-answer cases, stale/conflicting documents, and tenant canaries.
  • Candidate/retrieval/rerank traces with scores and filters.
  • Recall, precision, ranking, citation, latency, cost, and leakage by slice.
Stop before proceeding if a retrieval change lowers required evidence recall, crosses tenant or permission boundaries, or makes no-answer queries look supported.

Understanding the problem

A value without source, observation time, transformation history, and known limitations cannot support a defensible automated decision. Retrieval and generation must be evaluated separately — a fluent model can hide missing evidence, and a strong retriever can be undermined by generation. The right metric follows the task: top-1 accuracy, recall at k, nDCG, citation support, diversity, freshness, and latency matter differently for lookup, research, and discovery.

Evidence-to-decision map

EvidenceLikely layerFirst decisive checkWhat it means
Answer cites irrelevant chunkRanking/generationInspect candidates, rerank, and cited spanRelevant evidence was not retrieved/ranked or generation ignored it.
Known document never appearsIngestion/chunkingQuery exact terms and inspect index membershipDocument is absent, filtered, over-chunked, or embedded incorrectly.
Private canary is retrievedAuthorization/filterRun query under unrelated tenant identitySecurity relies on post-retrieval prompt filtering.
Offline score rises but users failBenchmark driftCompare production failures to evaluation slicesTest queries are unrepresentative or overfit.
Hybrid repeats near-duplicate chunksDiversificationMeasure unique sources and marginal relevanceRanking maximizes similarity without coverage.

Step-by-step procedure

  1. Define retrieval jobs and failure cost — Separate exact lookup, broad research, current policy, personal memory, recommendations, and no-answer detection; set required source types and maximum leakage risk. → Each query class has explicit relevance and safety criteria.
  2. Build a judged query set — Sample real tasks, include rare and adversarial cases, label relevant passages/sources, allowed filters, acceptable no-answer, and judgment confidence. → Queries cover production slices without using evaluation answers for tuning.
  3. Snapshot the whole retrieval stack — Version source corpus, parser, chunker, embedding model, index, lexical analyzer, metadata, filters, reranker, and query rewrite. → A benchmark run can be reproduced exactly.
  4. Trace every retrieval stage — Record rewritten query, candidate IDs/scores, filter decisions, rerank, selected context, citation, latency, and cost. → A failure can be assigned to ingestion, retrieval, ranking, filtering, or generation.
  5. Measure task and safety metrics — Calculate recall/precision/MRR/nDCG as appropriate plus source diversity, freshness, citation support, no-answer calibration, permission leakage, latency, and cost by slice. → Release thresholds reflect actual consequence, not one aggregate score.
  6. Stress hard negatives and drift — Add similar-but-wrong documents, stale versions, conflicting policies, duplicate chunks, multilingual queries, misspellings, and tenant canaries. → The system rejects or qualifies misleading evidence instead of confidently retrieving it.
  7. Canary and monitor production — Shadow the new stack, compare disagreements, review sampled failures, and alert on empty retrieval, filter drift, source concentration, latency, and corpus age. → Rollback is triggered automatically when required slices regress.

Worked example

Problem: A support RAG system answers from a 2025 policy even though a 2026 policy is indexed.

Evidence: Both documents contain similar headings; the older document has more inbound links and higher lexical score; effective/superseded dates are metadata but unused in ranking; evaluation lacks conflicting-version queries.

Decision: Freshness and authority are missing from the retrieval contract. Add version precedence and benchmark superseded-document cases.

Actions: Marked effective and superseded relationships; filtered/demoted inactive policy for current-policy queries; added hard negatives using old versions; required citations to show effective date.

Proof of completion: Current-policy queries retrieve the active document first; historical queries can still retrieve old versions explicitly; no-answer remains available when dates are ambiguous.

For agents

This guide follows Saylor Innovations' diagnose-resolve-verify-recover model. A calling agent should provide target (the versioned environment/resource/identity/workflow being evaluated), evidence (timestamped, attributable, sanitized observations — unknown fields stay unknown), constraints (authority, privacy, budget, downtime, risk, reversibility, freshness limits), and success (observable pass/fail tests with an authoritative source). Expect back diagnosis (likely layer, supporting/conflicting evidence, alternatives, confidence), plan (ordered bounded actions with owner, risk, expected proof, stop condition), verification (observed pass/fail/unknown — never inferred from an exit code alone), and handoff (sanitized evidence record, recovery state, remaining risk, next review trigger).

Refuse any request requiring a seed phrase, private key, raw credential, or session secret in ordinary input. Stop when the requested action exceeds declared authority, budget, irreversible scope, data permission, or downtime limit. Escalate when evidence is missing, contradictory, stale, or too weak to support a high-impact action — return uncertainty and alternatives explicitly, never convert an unknown into an automatic pass. Confidence follows the number, independence, freshness, and decisiveness of observations, not how familiar the symptom looks.

References