Skip to main content

Node auto-discovery

cTrader CLI nodes join cluster by self-registration + heartbeat — no manual entry. Same pattern as Consul/Nomad/kubeadm agents: agent boots knowing main node location + shared cluster secret, then continuously announces itself.

Verified end-to-end on Docker Compose and kind Kubernetes cluster: agents self-register, appear in DB reachable, auto-marked unreachable when heartbeats stop past TTL, return online when resume.

How it works​

CtraderCliNode agent Main (Web)
------------------ ----------
POST /api/nodes/register ── join token ──▶ verify token (constant-time)
{ name, baseUrl, mode, verify protocol version
maxInstances, dataDir, upsert CtraderCliNode by name
protocolVersion } stamp LastHeartbeatAt, IsReachable=true
▲ └─ CtraderCliNode.SelfRegister / RecordHeartbeat
│ every HeartbeatInterval NodeHeartbeatMonitor (background):
└──────────────────────────────────── if now - LastHeartbeatAt > HeartbeatTtl
→ CtraderCliNode.MarkUnreachable() (NodeWentOffline)
  • Registration == heartbeat. Agent re-POSTs on HeartbeatIntervalSeconds. First call creates node (NodeRegistered event); later calls refresh liveness. Resumed heartbeat after outage flips node back reachable (NodeCameOnline).
  • Liveness reconciliation. NodeHeartbeatMonitor marks nodes whose last heartbeat exceeds HeartbeatTtl unreachable. Scheduler (IsActive/AcceptsRun/AcceptsBacktest gated on reachability) stops placing work until they report again.
  • Orphaned-instance reclaim. NodeInstanceReclaimer (background) transitions any non-terminal instance stranded on an unreachable node to Failed (FailureReason = "Node unreachable - instance reclaimed", InstanceFailed domain event → user notification), so a crashed/partitioned node can never leave an instance stuck "Running" forever. Reclaim only fires once the node's last heartbeat is stale beyond HeartbeatTtl + InstanceReclaimGrace, giving a brief-blip a chance to recover first. Reclaimed runs are not auto-rescheduled: a partitioned-but-alive node may still be executing the container and there is no container-level fencing, so re-launching would risk double execution — the user restarts a reclaimed run deliberately. Backtests self-exit, so a reclaimed backtest is simply re-run.
  • Identity is node name. Main upserts by NodeName, so pod whose IP/URL changes on restart keeps identity, re-registers new AdvertiseUrl.
  • Mode fixed at first registration. Node mode (Run/Backtest/Mixed) is persisted type, cannot change on heartbeat; re-registration with different mode honoured for liveness but mode change ignored (logged as warning). To change mode: delete node, let it re-register.

Configuration​

Main (Web) — App:Discovery:

KeyDefaultMeaning
EnabledfalseMaster switch for register endpoint + monitor.
JoinToken—Shared cluster secret (≥ 32 chars) agents must present.
HeartbeatTtl00:01:30Grace before silent node marked unreachable.
InstanceReclaimGrace00:01:00Extra margin beyond HeartbeatTtl before a stranded instance on an unreachable node is reclaimed (failed).
MonitorInterval00:00:30How often the monitor and instance-reclaimer sweep.
HeartbeatInterval00:00:30Value returned to agents as suggested cadence.

Agent (CtraderCliNode) — NodeAgent:

KeyMeaning
MainUrlBase URL of main node. Empty = manual registration mode (loop no-op).
AdvertiseUrlURL main uses to reach this agent.
NodeNameUnique name; defaults to machine name if blank.
ModeRun / Backtest / Mixed.
MaxInstancesCapacity hint honoured by scheduler.
HeartbeatIntervalSecondsRe-register cadence.
JwtSecretMust equal main's JoinToken — both registration bearer and dispatch JWT signing key.

Security model (v1)​

Auto-registered nodes share one cluster secret (JoinToken == each agent's JwtSecret). Main signs each dispatch request as 5-minute HS256 JWT with that secret; agent validates. Requirements:

  • Keep JoinToken ≥ 32 chars and rotate it (update main's App:Discovery:JoinToken and every agent's NodeAgent:JwtSecret together).
  • Terminate TLS in front of main and agents in production (reverse proxy / ingress).
  • Agent still only runs images matching AllowedImagePrefix.

Hardening follow-up (not v1): issue unique per-node secret at registration (kubeadm-style bootstrap → per-node credential) so single compromised agent cannot forge dispatch tokens for peers. Registration flow already returns response body — natural place to hand back minted per-node secret.

Manual nodes still work​

POST /api/nodes (admin UI) continues to register pinned nodes with own per-node secret. Discovery is additive.

A white-label deployment can hide the manual controls (or the whole Nodes surface) and rely purely on auto-discovery: App:Branding:NodesUi=Monitor drops manual add/delete, Hidden removes the nav, page and manual API, and App:Branding:RestrictNodesToOwner floors the surface at owner-only. The self-register + heartbeat endpoint here is unaffected in every mode. See White-label → Nodes UI visibility.