Ingest files, APIs, feeds, and pages into a traceable corpus with licensing, integrity, version, transformation, and deletion evidence.
The result you're building
A source manifest and ingestion pipeline where every record can be traced to an allowed source snapshot, validated through each transformation, quarantined on failure, and removed or rebuilt deterministically.
Use this guide when
- You collect data for search, RAG, analysis, monitoring, or paid agent responses.
- Sources change format, terms, identifiers, or access rules.
- Buyers need freshness and provenance with each result.
Don't use it as a substitute for
- Copying data first and deciding permission or provenance later.
- Treating successful download as proof content is complete, authentic, current, or licensed for the intended use.
Before you start
Collect these first:
- Source owner, URL/identifier, access method, license/terms, and allowed use.
- Fetch time, response headers, checksum, size, encoding, and immutable raw snapshot.
- Parser and transformation versions with row/document counts.
- Schema, identifiers, freshness promise, quality thresholds, and quarantine reason codes.
- Correction, takedown, deletion propagation, replay, and rebuild evidence.
Understanding the problem
A value without source, observation time, transformation history, and known limitations cannot support a defensible automated decision. Schema changes are product changes — renames, units, null behavior, identifiers, and deleted fields can silently change decisions even when a pipeline still returns HTTP 200. Keep an immutable permitted snapshot or content digest so transformed records can be explained and rebuilt. Quality gates belong between stages, since fetch, decode, parse, normalize, enrich, index, and publish have different failure modes.
Evidence-to-decision map
| Evidence | Likely layer | First decisive check | What it means |
|---|---|---|---|
| Published value has no source | Lineage | Trace record ID to source snapshot and transform | Provenance was discarded or identifiers are not stable. |
| Row count drops after update | Parser/schema | Compare raw snapshot and stage counts | Source format changed or malformed records were silently skipped. |
| Deleted source remains searchable | Lifecycle | Trace deletion through derivatives and index | Takedown is not propagated to summaries, caches, or vectors. |
| Same input produces different output | Reproducibility | Pin parser, config, locale, dependency digests | Transformation environment or external enrichment is mutable. |
| Fresh response contains old facts | Observation semantics | Compare source publish time, fetch time, cache age | Pipeline freshness is being confused with source freshness. |
Step-by-step procedure
- Register the source and permitted use — Record owner, location, access method, authentication boundary, license/terms, data classes, retention, redistribution, and takedown contact before collection. → Each source has an approved purpose and accountable owner.
- Capture immutable fetch evidence — Store request identity, fetch time, status/headers, content length, checksum, encoding, and raw permitted snapshot or digest. → Repeated audit can identify exactly what bytes were observed.
- Validate before parsing — Check content type, magic bytes, size, malware, decompression limits, encoding, expected structure, and source signature where available. → Unexpected or dangerous content enters quarantine with a reason code.
- Parse and normalize deterministically — Pin parser/config versions, preserve original values, normalize units/time/IDs explicitly, attach lineage fields to every output. → Stage counts reconcile and the same snapshot rebuilds the same normalized result.
- Apply quality and drift gates — Measure required fields, duplicates, ranges, referential integrity, distributions, language, and schema diff; block publication on material unexplained drift. → Threshold failures create reviewable quarantine, not partial silent output.
- Index with source and lifecycle fields — Carry source ID, snapshot, locator, observed/published times, transform version, license, and deletion key into search and derived products. → Every returned item can cite and remove its source lineage.
- Test rebuild, correction, and deletion — Replay a snapshot, correct a record, process a takedown, invalidate caches, rebuild indexes, and compare counts/checksums. → The corpus is reproducible and removed data no longer serves from active paths.
Worked example
Problem: A paid research endpoint quotes a regulation removed from the source site, but the old chunk remains in vector search.
Evidence: The source URL now returns a replacement notice; the stored chunk lacks fetch/publication time; the vector record has no source deletion key; cache freshness measures index time only.
Decision: The response is not provenance-safe. Quarantine the source lineage, propagate deletion/replacement, and require time/source fields in retrieval.
Actions: Added immutable snapshot and locator metadata; carried deletion keys into chunks, summaries, vectors, and cache; rebuilt the affected index; added source-versus-pipeline freshness checks.
Proof of completion: The endpoint either returns the current authoritative text with provenance or explicitly reports unavailable; the superseded chunk cannot be retrieved.
For agents
This guide follows Saylor Innovations' diagnose-resolve-verify-recover model. A calling agent should provide target (the versioned environment/resource/identity/workflow being evaluated), evidence (timestamped, attributable, sanitized observations — unknown fields stay unknown), constraints (authority, privacy, budget, downtime, risk, reversibility, freshness limits), and success (observable pass/fail tests with an authoritative source). Expect back diagnosis (likely layer, supporting/conflicting evidence, alternatives, confidence), plan (ordered bounded actions with owner, risk, expected proof, stop condition), verification (observed pass/fail/unknown — never inferred from an exit code alone), and handoff (sanitized evidence record, recovery state, remaining risk, next review trigger).
Refuse any request requiring a seed phrase, private key, raw credential, or session secret in ordinary input. Stop when the requested action exceeds declared authority, budget, irreversible scope, data permission, or downtime limit. Escalate when evidence is missing, contradictory, stale, or too weak to support a high-impact action — return uncertainty and alternatives explicitly, never convert an unknown into an automatic pass. Confidence follows the number, independence, freshness, and decisiveness of observations, not how familiar the symptom looks.