Backup & disaster recovery
นี้ trading/financial app: database holds trading accounts copy profiles prop-firm challenges audit chains และ Data Protection key ring losing มัน loses money และ breaks regulatory/audit obligations back มัน up และ prove restore works
Targets
| Metric | Target | Meaning |
|---|---|---|
| RPO (max data loss) | ≤ 5 min | Use point-in-time recovery (continuous WAL) ไม่ใช่ เพียงnightly dumps |
| RTO (max downtime) | ≤ 1 h | Time ไป restore + re-point app ที่ restored database |
| Backup retention | ≥ 35 days | Covers late-discovered corruption + monthly audit windows |
| Restore drill | monthly | untested backup ไม่ใช่ backup |
What must backed up
- Postgres database — ทั้งหมด app data (single logical database
appdb) - Data Protection key ring — persisted in database
(
PersistKeysToDbContext<DataContext>) และ PFX-encrypted ผ่านApp:DataProtectionCertBase64มัน rides along ใน DB backup แต่ protecting certificate + password ของมัน (App:DataProtectionCertPassword) secrets stored outside DB — back พวกเขา up ใน secrets manager ของคุณ ไม่มี cert คุณ cannot decrypt secrets (cTID passwords Open API tokens node secrets AI key) หลัง restore
Managed Postgres (recommended)
ทั้ง cloud IaC paths provision managed Postgres ด้วย built-in PITR — enable + verify retention:
- Azure (
deploy/azure/main.bicepFlexible Server): setbackup.backupRetentionDays(≥ 35) และgeoRedundantBackupที่ compliance requires มัน restore ด้วย Point-in-time restore ไป new server จากนั้น update app ของappdbconnection string - AWS (
deploy/awsRDS Postgres Terraform): setbackup_retention_period(≥ 35) และbackup_window; keep automated backups + optional cross-region copy restore ด้วย RestoreDBInstanceToPointInTime จากนั้น repoint app
Managed PITR gives ≤ 5 min RPO ด้วย ไม่มี app changes — app just needs new connection string (และ existing retrying execution strategy ดู scaling.md tolerates cutover blip)
Self-hosted Postgres
- Continuous archiving (PITR): enable WAL archiving (
archive_mode=onarchive_commandไป object storage) + periodicpg_basebackuprestore = restore base backup + replay WAL ไป target time นี้ คือ what meets RPO target - Logical dumps (secondary): nightly
pg_dump -Fc appdbไป off-box storage สำหรับ portability / partial restores ไม่ sufficient alone สำหรับ RPO target - encrypt backups ที่ rest; store off database host
Restore drill (run monthly)
- restore latest backup (PITR ไป "now − 10 min") ไป scratch database ไม่ production
- point throwaway app instance (หรือ psql session) ที่มัน
- verify schema:
dotnet ef migrations listshows ไม่มี pending migrations app starts และ becomes/health-ready - Verify audit chain intact และ unbroken ผ่าน
IAuditTrailVerifier(tamper-evidentAuditChainInterceptorchain) — broken chain หลัง restore means corruption หรือ tampering - confirm secret decryption works (เช่น Open API authorization decrypts) — proves Data Protection cert + password were restored correctly
- record drill result (time taken vs RTO) และ destroy scratch database
automate steps 1–4 ใน CI ที่ environment allows (restore seeded backup ไป Testcontainer
run dotnet ef migrations list + audit-chain verify) ดังนั้น broken-backup regression caught
ก่อน คุณ need มัน
After real restore
- restore DB (PITR ไป just ก่อน incident)
- ensure Data Protection cert + password เป็น same ones ใน use ก่อน incident
- repoint app
appdbconnection string; roll replicas - startup runs migrations ภายใต้ advisory lock (ดู scaling.md) — safe ด้วย N replicas
- copy/prop-firm supervisors reclaim leases ของพวกเขา และ resync จาก broker (cTrader source ของ truth) ดังนั้น open positions reconverge อัตโนมัติ — nothing trusted จาก stale local state
- verify audit chain + spot-check recent trading data