Incident Response In Depth
Production Solana products fail in ways that look similar to users (stuck wallets, empty balances, failed sends) but require opposite first moves from engineers.
Search across all documentation pages
Production Solana products fail in ways that look similar to users (stuck wallets, empty balances, failed sends) but require opposite first moves from engineers.
A rate-limited RPC provider is not a vault exploit. A Custom program error after a bad upgrade is not a blockhash expiry. Treating every red dashboard as "retry harder" burns minutes and sometimes funds.
This page is the section umbrella: how to classify incidents, gather evidence, contain damage, mitigate without rewriting history, and close the loop with prevention and blameless RCA for Solana programs, dApps, and backends.
Solana products straddle immutable chain state and mutable off-chain systems.
On-chain: programs, accounts, mints, vault PDAs, and upgrade authority live on a cluster under Agave consensus. Once a transaction is finalized, the ledger does not offer an "undo" API. Mitigation is pause, freeze, redeploy, migrate, or compensate.
Off-chain: RPC providers, indexers, frontends, backends, wallets, and ops tooling decide what users can see and submit. Most "outages" users report start here.
An incident is any unexpected degradation of safety, funds control, correctness, or availability that needs coordinated response beyond normal debugging. Severity usually tracks user funds at risk, duration, and blast radius:
| Sev | Typical signal | Example |
|---|---|---|
| 1 | Active fund loss or critical pause needed | Exploit drain, compromised authority |
| 2 | Major product path broken, funds not actively draining | Bad deploy spike of program errors |
| 3 | Partial degradation with workaround | Single-region RPC lag, one instruction failing |
| 4 | Localized or cosmetic | Explorer mismatch, non-critical UI |
Roles beat heroics. Name an incident commander (decides sequence), a comms lead (users and stakeholders get facts, not speculation), and domain leads (program, client, infra). One channel for decisions; one timeline in UTC with signatures and slots.
Evidence is the Solana-specific glue between classes of failure. Capture early:
Without that chain, people argue anecdotes while the attacker or the 429 storm continues.
The section's sibling pages each own one face of the machine:
This page keeps the whole loop in view.
Incidents enter through alerts (TVL velocity, error rate, slot lag, 429s), user reports, whitehat messages, or explorer anomalies. First job is classify, not fix:
Symptom
|
v
Evidence pack (sig, slot, program, RPC, funds?)
|
+-- Tx / sim / Custom / CU / accounts --> failed-tx path
+-- 429, lag, WS drop, multi-user blank UI --> RPC path
+-- Unexpected outflows, CPI abuse, authority txs --> on-chain path
+-- Post-deploy error spike, wrong program id --> deploy/client pathFailed transactions leave logs, simulation output, and Anchor or native error codes. Simulate with @solana/kit 7.0.0 (simulateTransaction, logs, unitsConsumed) before rebroadcast. Refresh blockhash; map Custom: N to the program enum; separate user error from program bug from compute budget. Blind retry with the same stale message multiplies noise and fees.
RPC and infrastructure outages present as empty balances, "account not found," stuck confirms, or mass WebSocket disconnects. Primary actions: health-check providers (slot lag, genesis hash, error budgets), fail over reads and writes, open circuit breakers so retries do not cascade, and avoid declaring an on-chain exploit when the chain is fine and your edge is not.
On-chain incidents (exploits, authority compromise, critical logic bugs) demand pre-built pause paths, multisig or cold upgrade authority, and evidence preservation. Draining transactions cannot be reversed; containment is stop the bleed, snapshot addresses, and prepare a verified patch or authority rotation. Assume active exploitation until logs prove otherwise.
Bad deploys and client mismatches sit between program and frontend: wrong cluster env, feature flag on for a broken instruction, or upgrade that passes tests but fails under mainnet account shapes. Frontend kill-switches are often faster than program rollback; both may be required.
Solana "rollback" is a product metaphor, not a ledger rewind. Practical levers:
| Lever | Speed | Scope | Notes |
|---|---|---|---|
| Client feature flag / kill switch | Fast | New user flows | Stops new bad paths |
| Multi-provider RPC failover | Fast | Availability | Does not fix program bugs |
| Program pause / mode flag | Medium | Mutating ix | Must exist before crisis |
| Authority rotation | Medium | Control plane | Break-glass multisig |
Redeploy prior verified .so | Medium-slow | Program logic | Verifiable build hash |
| State migration / compensation | Slow | Balances, claims | Needs design and comms |
Order of operations under pressure:
Upgrade authority and pause admin are production features. If they live only on a hot laptop with no runbook, you do not have incident response; you have hope.
Internal bridge: UTC timestamps, signature hashes, program ids, decisions and owners. External: short factual updates (impact, mitigation status, user actions). Do not publish unconfirmed exploit theories or private keys in "debug" screenshots. For Sev-1, plan a customer-facing summary within a fixed window after containment (sibling RCA page details the template).
Build for response, not only for happy path:
GlobalConfig patterns and native equivalents).Teams shipping only feature work without these controls eventually invent them during an incident, at maximum cost.
| Question | If yes | Primary playbook |
|---|---|---|
| Are funds leaving unexpectedly? | Sev-1 contain | On-chain incident response |
| Are many users failing on one provider only? | Failover | RPC and infrastructure outages |
| Does simulation show Custom / CU / accounts? | Decode | Debugging failed transactions |
| Did error rate jump right after deploy? | Mitigate | Rollback and mitigation |
| Is the path flaky under load only? | Fees / landing / RPC | Failed txs + RPC pages |
Run the matrix in order; parallelize only after classification so eng and security do not thrash each other.
When the bridge ends, work is not done. A blameless postmortem separates proximate cause from contributing factors, records what went well, and assigns action items with owners and due dates. Prevention checklists convert those items into Tier-1 deploy gates, Tier-2 30-day fixes, and Tier-3 quarterly drills (pause drills, synthetic landing, authority monitoring).
Incident response maturity is measurable: mean time to classify, mean time to pause or failover, fraction of Sev-1s with completed action items, and drill pass rate. Stack versions matter for drills too: re-validate scripts against Agave 4.1.1, Solana CLI 3.0.10, Anchor 0.32.1, Rust 1.91.1, and @solana/kit 7.0.0 after toolchain bumps so runbooks do not rot.
Whitehat reports, law enforcement, insurers, and auditors may need the same evidence pack your engineers use. Preserve logs, signatures, and authority history; do not "clean up" explorers or rotate keys without recording what was compromised. Program security design (account checks, CPI hygiene) reduces incident frequency; incident response reduces severity when design fails. Both are required.
It is the coordinated practice of classifying chain and off-chain failures, containing fund and UX damage with prebuilt levers, mitigating forward without ledger rewind, and turning each event into durable prevention.
A failed transaction has a signature or simulation result and a program or runtime error class; an outage is often multi-user unavailability caused by RPC, network, or dependent infrastructure without a single program root cause.
Cluster, symptom, signatures and slots, program id, RPC provider, whether funds moved, and recent deploys or authority changes; then assign commander, eng, and comms roles.
Pause when funds or critical state are at risk from on-chain logic or exploitation; fail over RPC when evidence points to provider health, rate limits, or slot lag with healthy on-chain balances.
You can redeploy a prior verified binary to the same program id if you hold upgrade authority; you cannot erase finalized account changes that already executed under the bad version.
Simulation surfaces logs and compute usage without spending mainnet fees on blind retries, and it separates preflight-detectable bugs from landing and confirmation problems.
They let you route around 429s, regional failures, and single-vendor maintenance; health checks and circuit breakers prevent cascading retries that make outages worse.
Pre-deployed pause, break-glass admin keys or multisig, upgrade authority procedure, evidence capture for exploiter addresses, comms templates, and a path to verified patch under audit pressure.
Draft the timeline while memory is fresh (same day when possible); complete blameless RCA with action items after containment, and publish customer-facing facts on a fixed Sev-1 SLA.
It explains system and process conditions that made the failure likely, without shaming individuals, and still assigns concrete fixes with owners and dates.
Every meaningful Sev should leave monitors, CI gates, or drills on a checklist so the next release proves the gap is closed before users rediscover it.
Client and backend flags disable broken instruction paths faster than on-chain upgrades; pair them with on-chain pause when funds can still move through other clients.
Kit 7.0.0 improves typed RPC, simulation, and send paths, but classification, commitment, failover, and pause authority remain operational and program design concerns.
Pin scripts and examples to Agave 4.1.1, Solana CLI 3.0.10, Anchor 0.32.1, Rust 1.91.1, and @solana/kit 7.0.0, and re-drill after upgrades.
Use Debugging Failed Transactions and RPC & Infrastructure Outages for the two most common non-exploit paths; open On-Chain Incident Response and Rollback & Mitigation for fund-risk events; close the loop with Prevention Checklists and Blameless Post-Mortems & RCA.
Stack versions: This page was written for Agave 4.1.1, Solana CLI 3.0.10, Anchor 0.32.1, Rust 1.91.1, and @solana/kit 7.0.0.
Reviewed by Chris St. John·Last updated Jul 15, 2026