Devnet had TURN running on ZERO nodes for namespace anchat-test after a node replacement, with every node reporting failed=0 and six hours of logs containing no matching lines. Four defects, one outage. #173 (root cause) — namespace_cluster_nodes accumulated rows for permanently dead nodes; removeClusterNodeAssignment existed but was never called, and RepairCluster is add-only. With 4 members (2 corpses) the WebRTC reconciler computed 2*live > members => 2*2 > 4 => false: a permanent 50/50 deadlock that could never resolve. Adds pruneStaleClusterNodes, wired into RepairCluster and the 60s reconcile loop, keyed off dns_nodes staleness. Also explains why the ring health monitor never fired: startDNSHeartbeat flips a silent node to inactive at 120s, but getRingNeighbors only probes active nodes, so the node leaves every observer's set before the monitor's own ~120s threshold and its miss count is discarded. The DNS sweep almost always wins that race. The prune is independent of it. Second bug found: an unpruned corpse made RepairCluster count it as active, so a replacement was never triggered. #170 — viable and live member sets came from two separate rqlite queries, so live ⊆ viable was incidental, not structural; combined with webrtcReconcileQuorumOK(live, 0) returning true, an empty viable set passed quorum and deallocated every role while allocating nothing back. Now one merged query split in Go, plus an explicit empty-set guard. #171 — regression in the prior fix: past the grace window both numerator and denominator derive from the same liveness signal, so a lone node always had "quorum" and would strip every other node's roles onto itself. Reachable via our own serial rolling-restart runbook. Adds webrtcReconcileMajorityHeld (viable >= (raw+1)/2) and a 5m startup grace. #172 — shortfall log moved after the plan and conditioned on the viable set, stale invariant comments corrected, grace-boundary and end-to-end tests added. Verified on devnet 0.122.100: TURN 0 -> 2 nodes, SFU 1 -> 3, every allocation on a live node, both corpses gone from namespace_cluster_nodes. 183 tests, go vet and -race clean. Testnet untouched. Build note: vault/build.zig.zon declares minimum_zig_version 0.15.2; Zig 0.16 removed GeneralPurposeAllocator and process.argsAlloc. Build with zig@0.15. Co-Authored-By: Claude <noreply@anthropic.com>
24 KiB
Node Replacement Runbook (Nameserver VPS Swap)
How to replace one Orama nameserver VPS with a new machine without losing
platform Raft quorum — and without breaking namespaces like anchat-test.
Written from the devnet cutover on 2026-08-03 (replaced old ns3 51.38.128.56
with Contabo 169.58.118.206). Use the same process for testnet.
Related docs:
- DEV_DEPLOY.md — install, join, rolling upgrades, recover-raft
- DEVNET_INSTALL.md — install command examples
- CLEAN_NODE.md — full wipe of a VPS after it leaves the cluster
- NAMESERVER_SETUP.md — NS / glue DNS
- COMMON_PROBLEMS.md — WireGuard / Olric / vault issues
Inventory file: core/scripts/nodes.conf
Golden rules (do not skip)
- Join the new node first. Never clean/remove the old node first.
- Only replace one platform voter at a time. Quorum on 3 voters = 2.
- Prefer replacing a follower, not the Raft leader.
- Platform health ≠ namespace health. Namespaces have their own RQLite/Olric/gateway.
- Do not point namespace DNS at a node that is not running that namespace.
- Drive ops with
oramaCLI / documented recovery. Avoid parallel restarts. - Match binary version to the live cluster (
/opt/orama/manifest.json). - Wait for the new node to fully catch up (applied index,
voter: trueon both sides) before removing anyone. - Platform join ≠ serverless-ready. A new node starts with an empty IPFS blob store. Function WASM CIDs live in IPFS, not only in RQLite. Until you backfill/pin every active function WASM onto every peer that may serve invokes (including the new node), cold nodes will 15s-timeout on
ipfs catand deploys can 504 on upload — see bugboard #167 /docs/COMMON_PROBLEMS.md.
Mental model
┌─────────────────────────────────────┐
Public clients ──►│ Caddy :443 → platform gateway │ (each nameserver)
│ reverse_proxy → localhost:6001 │
└──────────────┬──────────────────────┘
│ proxies to
▼
┌─────────────────────────────────────┐
│ Namespace gateways (:10004, …) │
│ Namespace RQLite / Olric / SFU │ ← separate Raft/cluster
└─────────────────────────────────────┘
Platform RQLite (port 5001 / raft 7001) = cluster membership, DNS DB, routing
Namespace RQLite (e.g. 10000/10001) = that namespace's app data
Replacing a nameserver VPS touches:
| Layer | What changes |
|---|---|
| Platform Raft | Add new voter, remove old voter |
| WireGuard | New peer 10.0.0.x |
Zone DNS (ns1/ns2/ns3, apex) |
A records for the NS slot |
| IPFS / IPFS-Cluster | New peer joins with an empty local Kubo repo — backfill function WASM (see below) |
| Each namespace hosted on the old IP | Must rebalance / recover — not automatic enough |
| Registrar glue (optional) | Parent-zone glue for nsN if you control it |
Preflight checklist
1. Identify the three nameservers
# From laptop
cat core/scripts/nodes.conf | grep '^devnet\|^testnet'
# Live Raft (run on any healthy node)
curl -sS http://127.0.0.1:5001/nodes | python3 -m json.tool
curl -sS http://127.0.0.1:5001/status | python3 -c \
"import sys,json; r=json.load(sys.stdin)['store']['raft']; print(r.get('state'), r.get('num_peers'))"
Pick:
- Victim = follower (not
leader: true) - Join hub = current leader or any healthy nameserver public IP
- New VPS = fresh Ubuntu, root/ubuntu SSH, enough RAM/disk (devnet boxes were ~12 vCPU / 45 GB / 290 GB; Contabo was smaller but worked)
2. Confirm env health
curl -sS https://orama-<env>.network/v1/health # or your gateway_url
# expect status healthy
# For critical namespaces:
curl -sS https://ns-<name>.orama-<env>.network/v1/health
3. List which namespaces run on the victim
On each nameserver:
systemctl list-units 'orama-namespace-*' --no-pager --state=running
In platform RQLite (auth: orama + /opt/orama/.orama/secrets/rqlite-password):
-- namespace_cluster_nodes joined to dns_nodes
SELECT ncn.role, ncn.status, dn.ip_address, dn.internal_ip, dn.status
FROM namespace_cluster_nodes ncn
JOIN namespace_clusters nc ON nc.id = ncn.namespace_cluster_id
LEFT JOIN dns_nodes dn ON dn.id = ncn.node_id
WHERE nc.namespace_name = 'anchat-test'; -- repeat per namespace
If the victim is a namespace RQLite voter, removing it without recovery can leave the namespace at 1/3 voters → no leader. Plan namespace recovery before Raft-remove.
4. Match install binary version
# On an existing node
cat /opt/orama/manifest.json # version, e.g. 0.122.99
# Locally you need the same archive, e.g.
# /tmp/orama-0.122.99-linux-amd64.tar.gz
# or: orama build (then use the produced archive)
5. Secrets / SSH
- Prefer SSH key on the new VPS (
debros-nodesor rootwallet vault). orama node setupneeds unlocked rootwallet; manual path below does not.
Phase A — Join new node (old node stays fully up)
A1. Invite token (on an existing installed node)
sudo /opt/orama/bin/orama node invite --expiry 2h
Save:
--token …--ca-fingerprint …(if printed)- Join URL hint (HTTPS domain or
http://<hub-public-ip>)
Tokens are single-use. Generate one per join.
A2. Bootstrap new VPS
# On new VPS as root
mkdir -p /opt/orama
tar -xzf /tmp/orama-<version>-linux-amd64.tar.gz -C /opt/orama
install -m 755 /opt/orama/bin/orama /usr/local/bin/orama
# Stop Docker if present (port fights with IPFS)
systemctl stop docker docker.socket 2>/dev/null || true
systemctl disable docker docker.socket 2>/dev/null || true
A3. Install as nameserver joining the cluster
sudo orama node install \
--join http://<HUB_PUBLIC_IP> \
--token <TOKEN> \
--ca-fingerprint <FP> \
--vps-ip <NEW_PUBLIC_IP> \
--domain <base-domain> \
--base-domain <base-domain> \
--nameserver \
--environment <devnet|testnet> \
--ssh-user ubuntu
Notes:
- Prefer
http://<hub-ip>if DNS/TLS is flaky during cutover (docs allow this). - Never join via
:6001(blocked by UFW). - Installer may warn that the base domain does not yet resolve to the new IP — expected until DNS update.
- Assigned WG IP example:
10.0.0.17.
A4. Wait until fully synced (mandatory gate)
On leader:
curl -sS http://127.0.0.1:5001/nodes | python3 -m json.tool
# NEW wg address must show: voter true, reachable true
On new node:
curl -sS http://127.0.0.1:5001/status | python3 -c "
import sys,json
r=json.load(sys.stdin)['store']['raft']
print('state', r.get('state'), 'voter', r.get('voter'),
'peers', r.get('num_peers'), 'applied', r.get('applied_index'))
"
# Need: voter True, peers >= 3 (for 4-node temp), applied_index catching leader
Also:
# From leader
ping -c 2 10.0.0.<new>
systemctl is-active orama-node coredns caddy
curl -sS http://127.0.0.1:6001/health # or /v1/health
Do not continue until the new node is a healthy platform voter with a non-zero applied index. First minutes often show voter: false / empty /nodes while snapshotting — wait (can take several minutes on large DBs).
Temporary topology: 4 platform voters. Quorum = 3. Still safe.
Phase B — DNS (platform zone)
Auth for RQLite:
PASS=$(sudo cat /opt/orama/.orama/secrets/rqlite-password)
AUTH="orama:$PASS"
B1. Inspect current NS / apex records
curl -sS -u "$AUTH" -G 'http://127.0.0.1:5001/db/query' \
--data-urlencode "q=SELECT id,fqdn,value,is_active FROM dns_records WHERE fqdn IN (
'ns1.orama-devnet.network.','ns2.orama-devnet.network.','ns3.orama-devnet.network.',
'orama-devnet.network.','*.orama-devnet.network.'
) ORDER BY fqdn,value"
(Adjust domain for testnet.)
Also:
curl -sS -u "$AUTH" -G 'http://127.0.0.1:5001/db/query' \
--data-urlencode "q=SELECT * FROM dns_nameservers"
B2. Point the replaced NS slot at the new public IP
Example: replacing ns3:
-- ns3 A → new IP
UPDATE dns_records SET value='<NEW_PUBLIC_IP>', updated_at=CURRENT_TIMESTAMP
WHERE fqdn='ns3.<domain>.' AND record_type='A' AND value='<OLD_PUBLIC_IP>';
-- remove accidental dual A records on ns3
DELETE FROM dns_records WHERE fqdn='ns3.<domain>.' AND value NOT IN ('<NEW_PUBLIC_IP>');
-- apex + wildcard: replace old IP with new
UPDATE dns_records SET value='<NEW_PUBLIC_IP>', updated_at=CURRENT_TIMESTAMP
WHERE value='<OLD_PUBLIC_IP>'
AND fqdn IN ('<domain>.','*.<domain>.','push.<domain>.')
AND is_active=1;
-- dns_nameservers table
UPDATE dns_nameservers
SET ip_address='<NEW_PUBLIC_IP>',
node_id='<NEW_LIBP2P_NODE_ID>',
updated_at=CURRENT_TIMESTAMP
WHERE hostname='ns3';
Execute via:
curl -sS -u "$AUTH" -X POST 'http://127.0.0.1:5001/db/execute?pretty' \
-H 'Content-Type: application/json' \
-d '["<SQL>"]'
Verify authoritative (not only public cache):
dig +short A ns3.<domain> @<ns1-public-ip>
# expect NEW_PUBLIC_IP only
B3. Registrar glue (if applicable)
If the domain registry has glue for ns3.<domain> → old IP, update to new IP.
Zone DNS alone is not always enough for resolvers that only have glue.
Phase C — Namespace safety (the part that bit us on devnet)
Problem
Namespace DNS A records (ns-<namespace>.…) were bulk-updated when we rewrote every row with value=<old_ip>. That made clients hit the new public IP, which was a platform nameserver but did not run orama-namespace-*@<ns> units → TLS timeouts + platform circuit breakers (all upstream circuits are open).
Also, namespace RQLite still listed dead voters (10.0.0.6, 10.0.0.11). After removing the last peer, membership became 1/3 → Candidate, no leader.
Safe DNS rule for namespaces
Only advertise IPs that currently run that namespace’s gateway.
-- After cutover: keep only live gateway IP(s) for a namespace
UPDATE dns_records SET is_active=0, updated_at=CURRENT_TIMESTAMP
WHERE fqdn LIKE '%anchat-test%' AND value != '<LIVE_GATEWAY_PUBLIC_IP>' AND is_active=1;
UPDATE dns_records SET is_active=1, updated_at=CURRENT_TIMESTAMP
WHERE fqdn IN (
'ns-anchat-test.<domain>.',
'*.ns-anchat-test.<domain>.'
) AND value='<LIVE_GATEWAY_PUBLIC_IP>';
Do not blindly map old IP → new IP for all namespace rows unless the new node is already hosting those services.
Ideal path (HA preserved)
- Join new platform node (Phase A) — done.
- Before removing old node: rebalance each namespace so the new node (or another survivor) hosts gateway/rqlite/olric, or ensure ≥2 live namespace voters remain after remove.
- Only then remove old platform voter.
- Update namespace DNS to the live set.
If automatic cluster recovery does not reassign in time, use manual recovery below.
Emergency: namespace RQLite lost quorum (single survivor)
On the only live namespace host (example ports 10000 HTTP / 10001 raft — check namespace_cluster_nodes):
NS=anchat-test
NODE_ID=$(grep ^NODE_ID= /opt/orama/.orama/data/namespaces/$NS/rqlite.env | tail -1 | cut -d= -f2)
DATA=/opt/orama/.orama/data/namespaces/$NS/rqlite/$NODE_ID
RAFT=$DATA/raft
ADV=$(grep RAFT_ADV_ADDR /opt/orama/.orama/data/namespaces/$NS/rqlite.env | cut -d= -f2)
# e.g. ADV=10.0.0.2:10001
sudo systemctl stop orama-namespace-gateway@$NS
sudo systemctl stop orama-namespace-rqlite@$NS
# backup first
sudo cp -a "$RAFT/peers.info" "$RAFT/peers.info.bak-$(date +%Y%m%d)"
# rqlite recovery: peers.json is consumed at startup then removed
echo "[{\"id\":\"$ADV\",\"address\":\"$ADV\",\"non_voter\":false}]" | sudo tee "$RAFT/peers.json"
sudo chown orama:orama "$RAFT/peers.json"
sudo systemctl start orama-namespace-rqlite@$NS
# wait until Leader
curl -sS http://127.0.0.1:10000/status | python3 -c \
"import sys,json; print(json.load(sys.stdin)['store']['raft'].get('state'))"
Then fix Olric + gateway configs so they do not dial dead WG IPs:
# configs/olric-<node>.yaml — single node
memberlist:
peers: []
# configs/gateway-<node>.yaml
olric_servers:
- 10.0.0.<live>:10002
sudo systemctl restart orama-namespace-olric@$NS
sudo systemctl restart orama-namespace-gateway@$NS
curl -sS http://127.0.0.1:10004/v1/health # expect healthy, rqlite ok, olric ok
Mark old assignment rows stopped in platform DB:
UPDATE namespace_cluster_nodes
SET status='stopped', updated_at=CURRENT_TIMESTAMP,
error_message='node replaced <date>'
WHERE node_id='<OLD_LIBP2P_ID>' AND status='running';
Later: re-provision HA (add second/third namespace peers) so you are not single-node forever.
Required: IPFS function-WASM backfill (bugboard #167)
Why: Namespace gateways load function code with POST http://localhost:4501/api/v0/cat?arg=<wasm_cid>. Metadata (function name → CID) is in namespace RQLite and is fine after replace. The bytes are in Kubo. A replaced VPS has a nearly empty repo (repo/stat shows tens of objects vs thousands on old peers). The first invoke of each function on that node (or any cold peer after restart) can hang until the IPFS deadline (~15s → function 100% errors). The same IPFS layer hanging on add surfaces as orama function deploy 504 (proxy budget 30s) while other invokes still succeed.
Gateway /v1/health ipfs: ok only means the daemon answers — not that every registered WASM is local.
When to run: After the new node is a platform voter and IPFS + IPFS-Cluster are up on all nameservers. Re-run after any full IPFS repo wipe.
Steps (run from any host that can SSH to all nameservers; example uses anchat-test ports):
# 1) On a live namespace RQLite host (e.g. ns that runs orama-namespace-rqlite@anchat-test)
curl -sS -G 'http://127.0.0.1:10000/db/query?level=none' \
--data-urlencode "q=SELECT DISTINCT wasm_cid FROM functions WHERE status='active' AND wasm_cid IS NOT NULL AND wasm_cid != ''" \
| python3 -c 'import sys,json; v=json.load(sys.stdin)["results"][0].get("values")or[]; open("/tmp/cids.txt","w").write("\n".join(r[0] for r in v)+"\n"); print(len(v),"cids")'
# 2) Copy /tmp/cids.txt to EVERY nameserver, then on EACH node:
# Local pin (bitswap from peers that already hold the blocks) — this is the critical step.
while IFS= read -r cid; do
[ -z "$cid" ] && continue
curl -sS -m 180 -X POST "http://127.0.0.1:4501/api/v0/pin/add?arg=${cid}&recursive=true" >/dev/null \
|| echo "FAIL $cid"
done < /tmp/cids.txt
# Optional: also ask cluster to pin everywhere (RF=-1). Useful but not sufficient alone
# if a peer stays "unpinned" in peer_map — still do local pin/add above.
# curl -sS -X POST "http://127.0.0.1:9094/pins/${cid}?replication-factor-min=-1&replication-factor-max=-1"
# 3) Verify on EACH node (including the new one)
curl -sS -X POST http://127.0.0.1:4501/api/v0/repo/stat # new node repo size should jump (MB→100s MB)
# Hot CID from a real function (example from #167):
curl -sS -m 20 -o /dev/null -w "%{http_code} %{size_download} %{time_total}\n" \
-X POST "http://127.0.0.1:4501/api/v0/cat?arg=<HOT_WASM_CID>"
# Expect http=200, size ~1MB+, time well under 1s after backfill.
# 4) Upload path smoke test (same size class as AnChat deploys)
dd if=/dev/urandom of=/tmp/big.bin bs=1024 count=1200 status=none
curl -sS -m 60 -X POST -F file=@/tmp/big.bin http://127.0.0.1:4501/api/v0/add
Done for IPFS only when:
- Local pin audit: almost all active
wasm_cids returnKeyson every nameserver (investigate any remaining 500s — often a dead/orphan CID; redeploy that function). - Hot function CID
catis fast on the new node, not only on survivors. - ~1.2 MB
ipfs addsucceeds quickly on all peers. https://ns-<namespace>.…/v1/healthstays healthy and AnChat canorama function deploy+ invoke a hot function.
Do not mark a cutover complete because platform Raft is 3/3 alone.
Circuit breakers
Platform gateway tracks ns:<ip> breakers. Dead backends open circuits → HTTP 503
namespace gateway unavailable: all upstream circuits are open.
Fix: correct DNS + live gateways; wait or restart orama-node one follower at a time to clear in-memory breakers (never restart all voters at once).
Phase D — Remove old node from platform Raft
Only when:
- New node is synced voter
- Zone NS DNS points at new IP
- Namespace DNS only lists live gateway IPs
- You accept namespace HA state (rebalanced or recovered)
On platform leader:
PASS=$(sudo cat /opt/orama/.orama/secrets/rqlite-password)
AUTH="orama:$PASS"
# Confirm 4 voters, all reachable
curl -sS -u "$AUTH" http://127.0.0.1:5001/nodes | python3 -m json.tool
# Remove old WG raft id, e.g. 10.0.0.6:7001
curl -sS -u "$AUTH" -X DELETE http://127.0.0.1:5001/remove \
-H 'Content-Type: application/json' \
-d '{"id":"10.0.0.6:7001"}'
sleep 3
curl -sS -u "$AUTH" http://127.0.0.1:5001/nodes | python3 -m json.tool
# expect exactly 3 voters, all reachable; leader still elected
Mark old dns_nodes inactive:
UPDATE dns_nodes SET status='inactive', updated_at=CURRENT_TIMESTAMP
WHERE id='<OLD_LIBP2P_ID>';
Phase E — Clean the old VPS
Only after platform remove succeeds and remaining cluster is healthy.
Follow CLEAN_NODE.md on the old box (stop services, tear down WG, wipe /opt/orama, reset UFW to SSH-only).
Optional: repurpose the cleaned box (e.g. new jarvis operator host).
Phase F — Inventory & SSH
- Update
core/scripts/nodes.conf— victim IP → new IP for that role. - Update
~/.ssh/confighost aliases (and keep a break-glass alias for any leftover public workloads). - Optional: store SSH key in rootwallet vault for the new host.
Verification matrix (must all pass)
# Platform Raft = 3
curl -sS http://127.0.0.1:5001/nodes # 3 voters, all reachable
# Platform public
curl -sS https://orama-<env>.network/v1/health
# status: healthy — rqlite, olric, ipfs, vault, wireguard ok
# NS glue / zone
dig +short A ns1.<domain> @8.8.8.8
dig +short A ns2.<domain> @8.8.8.8
dig +short A ns3.<domain> @8.8.8.8
dig +short A ns3.<domain> @<ns1-ip> # authoritative truth
# Each critical namespace
curl -sS https://ns-<name>.orama-<env>.network/v1/health
# 200 healthy; rqlite ok; olric ok preferred
# Namespace must not resolve to a node without orama-namespace-gateway@<name>
dig +short A ns-<name>.orama-<env>.network @<ns1-ip>
# Optional: hit each A record with --resolve and confirm 200
What we did on devnet (2026-08-03) — reference
| Item | Value |
|---|---|
| Env | orama-devnet.network |
| Kept | ns1 storm 57.129.7.232 (10.0.0.1, leader) |
| Kept | ns2 wolverine 57.131.41.160 (10.0.0.2) |
| New ns3 | Contabo 169.58.118.206 (10.0.0.17) — SSH alias magneto |
| Old ns3 | 51.38.128.56 cleaned → new jarvis operator host |
| Binary | 0.122.99 archive, join via http://57.129.7.232 |
| Mistake | Bulk-rewrote namespace DNS to new IP before namespace services existed; lost namespace RQLite quorum |
| Fix | DNS → only wolverine for anchat-test; single-node rqlite peers.json recovery; olric/gateway peers → local only |
Post-fixover (until rebalanced):
- Platform: 3/3 healthy
anchat-test: healthy but single-node HA on wolverine only (namespace gateway/rqlite not on Contabo)- IPFS follow-up (bugboard #167, same day): Contabo joined with an empty blob store → cold WASM
cattimeouts + deploy 504s. Mitigation: localpin/addof 120/121 activewasm_cids on all three nameservers (Contabo repo ~empty → ~150 MB). Orphan CIDQmQeDv5K…(group-key-fetch-sharesv2) unpinnable cluster-wide — redeploy that function. Hot CIDQmNayszfin…cat~5–10 ms on all peers; ~1.2 MBadd~15–20 ms.
Testnet tomorrow — condensed checklist
Current testnet lines in nodes.conf (verify live before starting):
testnet|ubuntu@51.195.109.238|nameserver-ns1 # ironman
testnet|ubuntu@57.131.41.159|nameserver-ns1 # thor (role label may need cleanup)
testnet|ubuntu@51.38.130.69|nameserver-ns1 # hulk
curltestnet health + list namespaces + which nodes host them- Pick follower victim; note WG IP + libp2p id + public IP
- Same binary version as testnet
- Invite + install join on new VPS (
--environment testnet, correct--base-domain) - Wait until new node
voter:true+ applied_index caught up - Update
nsN/ apex /dns_nameserversin testnet DB - Namespace pass: either rebalance services onto new node, or keep DNS only on remaining live gateways
- IPFS WASM backfill on all nameservers (export CIDs from each critical namespace RQLite; local
pin/add; verifycat+ ~1.2 MBadd) — see section above - Platform
DELETE /removeold raft id - Clean old VPS (CLEAN_NODE.md)
- Update
nodes.conf+ RootWallet SSH vault entryIP/ubuntu+authorized_keys - Verification matrix (platform + every critical
ns-*health 10×) - If any namespace RQLite is Candidate → single-node recovery before declaring done
- Schedule HA rebalance (restore 3 namespace peers) if you recovered to 1 node
- AnChat (or app owner):
function deploy+ invoke a hot function (e.g. receipt path) after backfill
Explicit anti-patterns
| Don't | Why |
|---|---|
| Clean old node first | Can lose platform or namespace quorum |
Restart all orama-node at once |
Platform Raft split |
Map all DNS old_ip → new_ip blindly |
Clients hit empty Caddy/gateway |
| Assume install complete when process exits | May still be snapshotting |
Ignore olric unavailable on namespace health |
Often dead peer lists after topology change |
| Skip namespace rqlite leader check | App writes/auth can 503 while /v1/health still looks “ok” under weak reads |
| Skip IPFS WASM backfill after join | Cold node: function invoke 15s timeouts; deploy 504 while “ipfs: ok” |
| Trust cluster RF=-1 alone without local pin check | Peer_map can show unpinned on the new node; Kubo repo stays empty |
| Declare done when only platform Raft is 3/3 | Serverless still broken for cold CIDs |
Recovery quick refs
| Symptom | Action |
|---|---|
| Platform no leader / Candidate | DEV_DEPLOY.md orama node recover-raft --env … --leader <ip> |
| New node never becomes voter | Check WG ping, logs, re-invite + reinstall if partial |
| Namespace health 503 circuit open | Fix DNS to live gateways; restart one platform gateway |
Namespace rqlite leader not found |
Single-node peers.json recovery on survivor |
| Dual A on nsN | Delete extra rows; dig authoritative |
Function WASM cat 15s timeout after replace/restart |
Export active wasm_cids; local pin/add on every node (section above) |
function deploy 504 / proxy budget while invokes work |
Same IPFS layer — check add latency on each peer; fix blob path before blaming function.yaml |
New node repo/stat still tiny after hours |
Backfill never ran or bitswap blocked — re-run pin/add; check WG + swarm peers |
Done criteria
You are done only when all of these are true:
- Exactly 3 platform voters, all reachable, stable leader
ns1/ns2/ns3resolve correctly (authoritative dig)- Platform
/v1/health→ healthy - Every critical namespace
/v1/health→ 200 with rqlite ok (and olric ok if required) - No active DNS A records for decommissioned public IPs on those namespaces
- IPFS: every nameserver has pinned (or can fast-
cat) the active function WASM set; new node repo size is not empty - RootWallet:
IP/ubuntuSSH vault entry exists and is inauthorized_keyson the new VPS nodes.conf+ SSH match reality- Old VPS cleaned or intentionally repurposed
If (4) fails, the swap is not finished — fix namespace layer before walking away.