Engineering8 min read

The 23-item validator infrastructure checklist we use before we call it 'done'

Keys, monitoring, runbooks, upgrades, backups, security. The full checklist, with what we check, what we reject, and the order in which we build.

By the RanarTech engineering team · Published 2026-09-22

Before we call any validator deployment 'done,' we run through this checklist. It is the synthesis of running validators across EVM and Cosmos chains for foundations, exchanges, and DAOs. Use it as a starting point; adapt the items that are chain-specific.

Keys and secrets

1. Validator private key in HSM, never on disk. 2. Mnemonic (if applicable) in sealed envelope with documented dual custody. 3. Remote attestation enabled on the HSM. 4. No fallback to a plaintext key under any circumstance; the runbook for an HSM failure is to migrate, not bypass.

Double-sign protection

5. Single signing key reference per data center with process-level locking. 6. CI test that attempts to start a second validator instance and confirms refusal. 7. Sentinel job that detects any second instance and pages immediately. 8. Chain-level slashing protection (if the chain supports it) enabled and verified.

Monitoring and alerts

9. Attestation success rate with paged alerts below 99%. 10. Block proposal latency with paged alerts above the SLO. 11. Disk usage with paged alerts above 70%. 12. Peer count with paged alerts below the chain-recommended minimum. 13. Sync state with paged alerts on divergence.

Network and sentries

14. Sentry architecture: at least three public-facing nodes between the validator and the internet. 15. DDoS protection on the sentries (Cloudflare Spectrum or equivalent). 16. Rate limiting on sentry RPC endpoints. 17. RPC redundancy for the validator client (own sentry plus paid provider plus public fallback).

Upgrades

18. Staging network used for every client upgrade before production. 19. Documented rollback procedure with the previous client version. 20. Consensus-breaking version transitions reviewed by two engineers. 21. Upgrade windows scheduled outside high-activity periods.

Backup and recovery

22. Snapshot of validator state backed up to a region-independent location with at least 30-day retention. 23. Documented recovery procedure with a target RTO of one hour for a single validator, and a tested full-region recovery for multi-validator setups.

What we reject

We will not approve a validator deployment that uses a hot key on disk, that lacks sentry architecture, that has monitoring without paging, or that has no runbook for client upgrades. The cost of getting these wrong is denominated in the asset, and we have seen teams lose five and six figures in a single incident.

FAQ

How long does this take to implement?

Two to three weeks for a single validator with sentry architecture and a full monitoring stack. The cost is dominated by the HSM and the monitoring infrastructure; the engineering time is mostly monitoring configuration.

What is the minimum we can ship with?

HSM-backed keys, double-sign protection, sentry architecture, basic monitoring with paging, and an upgrade runbook. Everything else is hardening. We will not sign off on a production validator without those five.

More from RanarTech Insights

Have a project like this?

Discovery call within 48 hours. NDA-friendly. Most engagements kick off within 1-2 weeks.

[email protected]