Záloha a disaster recovery
Toto je trading/finanční app: databáze drží trading účty, copy profily, prop-firm challenges, audit chains, a Data Protection key ring. Ztráta jí znamená ztrátu peněz a porušení regulatorních/audit povinností. Zálohujte ji, a dokažte že restore funguje.
Cíle
| Metrika | Cíl | Význam |
|---|---|---|
| RPO (max data loss) | ≤ 5 min | Použijte point-in-time recovery (continuous WAL), ne jen noční dumpy. |
| RTO (max downtime) | ≤ 1 h | Čas na obnovení + přesměrování app na obnovenou databázi. |
| Backup retention | ≥ 35 days | Pokrývá pozdně-objevenou korupci + měsíční audit windows. |
| Restore drill | měsíčně | Netestovaná záloha není záloha. |
Co musí být zálohováno
- Postgres databáze — veškerá app data (single logical database
appdb). - Data Protection key ring — persisted in databázi
(
PersistKeysToDbContext<DataContext>) a PFX-encrypted viaApp:DataProtectionCertBase64. Ride along in DB backup, but the protecting certificate + its password (App:DataProtectionCertPassword) are secrets stored outside the DB — zálohujte je ve svém secrets manageru. Bez certu nemůžete dešifrovat tajemství (cTID hesla, Open API tokeny, node tajemství, AI klíč) po obnovení.
Managed Postgres (doporučeno)
Obě cloud IaC cesty provisionují managed Postgres s built-in PITR — enable + verify retention:
- Azure (
deploy/azure/main.bicep, Flexible Server): setbackup.backupRetentionDays(≥ 35) andgeoRedundantBackupwhere compliance requires it. Restore with Point-in-time restore to a new server, then update app'sappdbconnection string. - AWS (
deploy/aws, RDS Postgres, Terraform): setbackup_retention_period(≥ 35) andbackup_window; keep automated backups + optional cross-region copy. Restore with RestoreDBInstanceToPointInTime, then repoint the app.
Managed PITR dává ≤ 5 min RPO bez app změn — app just needs the new connection string (and the existing retrying execution strategy, viz scaling.md, tolerates the cutover blip).
Self-hosted Postgres
- Continuous archiving (PITR): enable WAL archiving (
archive_mode=on,archive_commandto object storage) + a periodicpg_basebackup. Restore = restore base backup + replay WAL to the target time. Toto is what meets the RPO target. - Logical dumps (secondary): nightly
pg_dump -Fc appdbto off-box storage for portability / partial restores. Not sufficient alone for the RPO target. - Encrypt backups at rest; store off the database host.
Restore drill (spouštět měsíčně)
- Obnovte poslední zálohu (PITR to "now − 10 min") into a scratch database, not production.
- Namiřte throwaway app instance (nebo psql session) na ni.
- Ověřte schema:
dotnet ef migrations listshows no pending migrations, app starts and becomes/health-ready. - Ověřte audit chain is intact and unbroken via
IAuditTrailVerifier(the tamper-evidentAuditChainInterceptorchain) — broken chain after restore znamená korupci nebo tampering. - Potvrďte secret decryption works (e.g. an Open API authorization decrypts) — proves the Data Protection cert + password were restored correctly.
- Record the drill result (time taken vs RTO) and destroy the scratch database.
Automatizujte kroky 1–4 in CI where the environment allows (restore a seeded backup into a Testcontainer,
run dotnet ef migrations list + the audit-chain verify) takže a broken-backup regression is caught
before you need it.
Po reálném obnovení
- Obnovte DB (PITR to just before the incident).
- Ensure the Data Protection cert + password are the same ones in use before the incident.
- Repoint app
appdbconnection string; roll the replicas. - Startup runs migrations under the advisory lock (viz scaling.md) — safe with N replicas.
- Copy/prop-firm supervisoři reclaim their leases and resync from the broker (cTrader is the source of truth), takže open positions reconverge automatically — nothing is trusted from stale local state.
- Ověřte audit chain + spot-check recent trading data.