The result you are building
A Solana monitor that checkpoints progress, reconnects with bounded backoff, backfills missed history, deduplicates events, tracks commitment/slot/source, detects gaps and provider disagreement, and never treats a WebSocket as durable delivery on its own.
Use this guide when: building wallet/token/program/new-launch monitoring, or fixing disconnects, duplicate alerts, rate limits, or missed events in an existing one.
Do not use it as a substitute for: assuming a subscription reconnect replays missed events (it doesn't), or acting financially on one unconfirmed provider event without a reconciliation policy.
Before you change anything, collect: the exact monitored entity/events and the acceptable commitment/finality level; RPC/WS providers, their limits, auth, methods, and historical backfill path; checkpoint/dedup identifiers and durable storage; gap, fork, latency, cost, and failover requirements.
Stop automated actions when slot gaps exceed backfill capability, providers materially disagree, commitment is below your action policy, or event identity can't prevent duplicate side effects.
Understand the system before fixing it
- WebSockets are notifications, not a ledger. Connections drop and provider buffers are finite — chain/RPC history is the actual source of truth for backfill.
- Commitment level changes both latency and reorg exposure. Processed/confirmed/finalized serve different use cases; record which one you're using and promote/retract state as it changes.
- The right dedup key depends on the event type. Signature plus instruction/log index, or account plus slot/version, is far safer than raw message text.
Evidence-to-decision map
| Evidence | Likely layer | First decisive check | What the result means |
|---|---|---|---|
| Disconnect/reconnect | Transport | Last-seen slot/signature and provider status | Reconnect, then backfill the overlap — don't resume blindly |
| Duplicate events | At-least-once delivery/overlap | Stable event key and payload hash | Upsert/dedup rather than alerting on every duplicate |
| Missing slot range | Gap | Compare checkpoint with current/first backfill result | Backfill the bounded history, or explicitly declare it incomplete |
| Provider 429 | Rate limit | Headers/errors/request rate | Throttle/cache/batch/failover — never a tight retry loop |
| Providers disagree | Commitment/indexing/fork | Compare slot/blockhash/signature status | Wait for higher commitment, or reconcile against a chain source |
Step-by-step procedure
01. Define event and truth semantics. A vague "new transaction" definition creates duplicates and misses. Specify the account/program/log/change, filters, event key, required fields, commitment level, reorg policy, and action threshold. Treat processed events as provisional unless your policy explicitly says otherwise.
02. Persist checkpoints and raw references. A memory-only cursor disappears on crash. Store the provider, subscription, last-observed slot/signature, finalized checkpoint, event key/hash, status, and timestamps durably, before or alongside processing — but avoid storing secrets or unnecessary full payloads.
03. Reconnect with backoff and resubscribe. Fast reconnect storms make outages worse. Implement heartbeat/staleness detection, exponential backoff with jitter and a cap, auth refresh, resubscription, and connection metrics. Being connected without event flow is also an unhealthy state worth alerting on.
04. Backfill and deduplicate. A subscription can never guarantee delivery of what you missed. Use signature/history/block/account RPC methods appropriate to the event type, paginate until the checkpoint or a limit, merge live and historical results by the stable key, and process in deterministic order. Every gap should close or be explicitly marked incomplete.
05. Handle commitment and forks. Provisional events can disappear or change. Track slot/blockhash/commitment, promote an event as it reaches confirmed/finalized, retract or mark it orphaned per your policy, and delay irreversible alerts/actions until your commitment threshold is met. Never silently keep an orphaned event as final.
06. Fail over and observe. A secondary provider can disagree with the primary or cost more. Health-score latency/error rate/slot lag, compare sampled results across providers, use bounded failover with circuit breakers and budgets, and alert on disconnects, gaps, backfill age, duplicates, and provider disagreement. Failover must never double-process the same event.
Worked example
Starting problem: a launch monitor misses tokens during a 12-minute WebSocket outage.
Evidence collected: the client reconnects and resubscribes but has no checkpoint or backfill logic; alerts use raw log text, so duplicates appear on retries; the provider offers signature history; the monitor can't prove completeness for the outage window.
Decision: transport recovery restored future events but permanently skipped the outage window entirely.
Actions taken: persisted a slot/signature checkpoint and a stable instruction event key; on reconnect, backfilled the overlapping signature history and merged it with the live feed; added a gap-completeness flag and a provider-lag alert.
Proof of completion: an outage test recovers all historical events exactly once, no duplicate alerts fire, the checkpoint only advances after successful processing, and incomplete ranges are visible in monitoring.
Why this matters: durability comes from reconciliation with chain history, not from a socket that merely looks reliable.
Verify, recover, and hand off
A monitor is complete only when: a crash/reconnect resumes from the durable checkpoint; the outage window backfills with overlap and no duplicates; every event carries slot/commitment/source/observed time; fork/provisional state promotes or retracts correctly; 429s/outages/provider disagreement stay bounded rather than cascading; and completeness, lag, duplicate, and cost metrics all alert.
If the socket is connected but producing no events, check for a stale connection/subscription/filter/provider lag via heartbeat and slot progress, then resubscribe/backfill as needed. If backfill loops, your pagination cursor/order/checkpoint logic is wrong — record page cursors and ensure monotonic termination. If alerts duplicate after a failover, you're using a provider-specific ID — switch to a chain-stable signature/instruction key. If finalization feels too slow, that's a commitment tradeoff — emit a provisional label and promote it later; never mislabel certainty.
Reusable handoff record: event/commitment/reorg semantics; durable checkpoint/dedup schema; reconnect/backfill/failover algorithms and budgets; gap/outage/fork/load test results; completeness/latency/error/cost monitoring.
For agents
Any agent consuming a live feed built on this pattern should treat every event as provisional until it carries the commitment level the agent's own policy requires, and should never assume a reconnect alone means no data was missed — check the completeness/gap metric this guide has the monitor expose.
Official references: https://solana.com/docs/rpc/websocket · https://solana.com/docs/rpc · https://solana.com/docs/rpc/http/getsignaturesforaddress
*This is educational technical and risk-analysis information, not financial, investment, legal, or tax advice. Blockchain transactions can be irreversible and no checklist can guarantee safety or profit.*