emma-hermes/docs/profile-isolation/restart-recovery-runbook.md

9.7 KiB

Gateway / API Server / OpenWebUI Restart & Recovery Runbook

Purpose

Use this runbook when code changes affecting gateway, api_server, managed downloads, or the OpenWebUI runs proxy need to go live safely.

This document answers four operator questions:

  1. Do these changes need a restart?
  2. How do we restart without causing restart loops or half-broken state?
  3. What should we verify before and after restart?
  4. How do we diagnose and recover when restart fails or keeps flapping?

Does this work require restart?

Yes for the current workstream.

These changes modify live Python code paths:

  • gateway/platforms/api_server.py
  • hermes_cli/managed_downloads.py
  • /Users/hermes/.hermes/scripts/openwebui_runs_proxy.py

If the gateway and/or the OpenWebUI runs proxy are already running, their existing Python processes will keep serving the old code until they are restarted.

Live findings from this session

Verified during runtime inspection:

  • http://127.0.0.1:8642/health returned 200
  • http://127.0.0.1:8653/health returned 200
  • http://127.0.0.1:8646/line/webhook/health returned 200
  • launchctl print gui/$(id -u)/ai.hermes.gateway showed the active LaunchAgent still running:
    • gateway run --replace
    • runs = 41
    • forks = 2400
    • last terminating signal = Killed: 9

Interpretation:

  • the service is currently healthy
  • but the long-lived LaunchAgent configuration is risky because --replace can create self-replacement churn in a supervised service
  • this should be corrected before or as part of the next controlled reload

Critical launchd rule learned from this rollout

When the LaunchAgent plist itself changes, launchctl kickstart -k <label> is not enough to guarantee the live job adopts the new ProgramArguments.

Use these rules instead:

  • use kickstart -k when you only want to restart the already-loaded job definition
  • use bootout ...plist + bootstrap ...plist when you changed the plist and need launchd to load the new definition

Real evidence from this rollout:

  • after removing --replace from ~/Library/LaunchAgents/ai.hermes.gateway.plist, a plain kickstart -k gui/$(id -u)/ai.hermes.gateway still left the live arguments on gateway run --replace
  • only after launchctl bootout gui/$(id -u) ~/Library/LaunchAgents/ai.hermes.gateway.plist plus launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/ai.hermes.gateway.plist did launchctl print show live arguments updated to plain gateway run

Safe rollout sequence

Phase A — preflight (do not restart yet)

  1. Confirm modified code compiles.
  2. Run targeted tests for:
    • tests/gateway/test_api_server.py
    • tests/gateway/test_session_api.py
    • tests/gateway/test_openwebui_runs_proxy_handoff.py
  3. Confirm current health endpoints before touching launchd:
    • 8642/health
    • 8653/health
    • 8646/line/webhook/health
  4. Record current listeners / PIDs:
    • port 8642
    • port 8653
  5. Confirm whether the gateway LaunchAgent still contains --replace.

Completion criteria:

  • tests pass
  • current runtime is healthy
  • we know whether launchd config itself needs correction

Phase B — fix the launchd steady-state configuration

For a long-lived LaunchAgent, prefer:

  • gateway run

Avoid in steady-state launchd config:

  • gateway run --replace

Reason:

  • --replace is appropriate for explicit operator takeover flows
  • but in a persistent RunAtLoad + KeepAlive LaunchAgent it increases the risk of self-replacement races, SIGTERM churn, and restart flapping

Phase C — reload in a controlled order

  1. Update the LaunchAgent plist if it still contains --replace.
  2. If the plist changed, reload the LaunchAgent definition with bootout + bootstrap instead of relying on kickstart -k.
  3. Verify the new gateway process command line no longer includes --replace.
  4. Re-check 8642/health before treating OpenWebUI as recovered.
  5. Re-check 8653/health and run a real OpenWebUI proxy smoke request.

Completion criteria:

  • gateway healthy on 8642
  • runs proxy healthy on 8653
  • live command line matches expected startup arguments

Phase C1 — exact commands for plist-changing reloads

Gateway:

uid=$(id -u)
launchctl bootout gui/${uid} ~/Library/LaunchAgents/ai.hermes.gateway.plist
launchctl bootstrap gui/${uid} ~/Library/LaunchAgents/ai.hermes.gateway.plist

OpenWebUI runs proxy:

uid=$(id -u)
launchctl bootout gui/${uid} ~/Library/LaunchAgents/ai.hermes.openwebui-runs-proxy.plist
launchctl bootstrap gui/${uid} ~/Library/LaunchAgents/ai.hermes.openwebui-runs-proxy.plist

Recommended health wait loops:

for i in {1..30}; do
  if curl -fsS http://127.0.0.1:8642/health; then
    echo "gateway healthy on try ${i}"
    break
  fi
  sleep 1
done

for i in {1..30}; do
  if curl -fsS http://127.0.0.1:8653/health; then
    echo "proxy healthy on try ${i}"
    break
  fi
  sleep 1
done

Observed in this real rollout:

  • gateway recovered on try 13
  • runs proxy recovered on try 2

Post-restart verification checklist

Core health

  • 8642/health returns 200
  • 8653/health returns 200
  • 8646/line/webhook/health returns 200
  • lsof shows one expected listener on 8642
  • lsof shows one expected listener on 8653

Launchd / process state

  • launchctl print gui/$(id -u)/ai.hermes.gateway shows gateway run without --replace
  • PID changed only once during the controlled restart
  • no immediate follow-up SIGTERM / Killed:9 churn appears in logs

Functional checks

  • non-streaming /v1/responses still rewrites local artifact paths to managed download URLs
  • streaming /v1/responses no longer leaks /Users/... or C:\... in live deltas
  • /v1/chat/completions streaming still emits normal assistant text and keepalives
  • OpenWebUI proxy rewrites local file paths to clickable managed download links
  • download_link_mode=private/public still propagates correctly

Failure modes and how to diagnose them

1. Service does not come back at all

Symptoms:

  • 8642/health fails
  • no listener on 8642
  • launchd job exists but no stable PID

Check first:

  • gateway.error.log
  • gateway.log
  • launchctl print gui/$(id -u)/ai.hermes.gateway
  • command line / working directory / venv path in the plist

Most likely causes:

  • syntax/import error in changed Python code
  • bad venv path or stale interpreter path
  • environment mismatch
  • port bind failure

2. Service comes back, then keeps restarting

Symptoms:

  • runs / forks climb quickly in launchctl print
  • repeated Received SIGTERM
  • last terminating signal = Killed: 9 or similar

Check first:

  • whether the LaunchAgent still contains --replace
  • whether another helper job is repeatedly kickstarting the service
  • whether a watchdog or restart helper is racing the main LaunchAgent

Most likely causes:

  • supervised launchd service using --replace
  • duplicate restart authority (helper + launchd + watchdog)
  • repeated crash on startup with KeepAlive immediately reviving it

3. Gateway is healthy but OpenWebUI still broken

Symptoms:

  • 8642/health is healthy
  • 8653 is unhealthy or returns bad stream behavior

Check first:

  • runs proxy process / listener on 8653
  • runs proxy logs
  • whether the proxy is still serving old code
  • whether it can reach upstream 8642

Most likely causes:

  • proxy not restarted after code change
  • upstream base URL mismatch
  • missing dependency / import in the proxy process
  • old proxy process still running

4. Health endpoints are fine but behavior is stale

Symptoms:

  • routes work, but old behavior remains
  • new download-link mode / streaming rewrite behavior not visible

Check first:

  • live PID start times
  • whether the file on disk matches the process that launchd is running
  • whether only one process instance exists per port

Most likely causes:

  • service not actually restarted
  • edited one file, but launchd runs another path
  • stale long-lived proxy/gateway process
  • the job was kickstart -k restarted but the changed plist definition was never reloaded, so launchd kept the old arguments
  1. Confirm health endpoints and listeners.
  2. Inspect launchd runtime state (launchctl print).
  3. If the plist changed, confirm whether the job was truly reloaded with bootout + bootstrap rather than merely kickstart -k restarted.
  4. Inspect recent gateway stdout/stderr logs with timestamps around the restart.
  5. Verify the exact source file on disk contains the expected fix.
  6. If needed, stop auto-restart pressure and run the gateway/proxy once in foreground with the same venv and env vars.
  7. Only after the foreground process is healthy, restore launchd supervision.

Anti-flap rules

  • Do not combine multiple independent restart authorities unless one is clearly primary.
  • Do not leave --replace inside a steady-state launchd job.
  • Do not restart blindly when health endpoints are already good; first confirm whether the issue is stale behavior vs runtime outage.
  • Do not declare success from static code inspection alone; always verify the live listener and health endpoint.

Rollback rule

If restart introduces instability:

  1. stop the restart loop source
  2. revert to the last known-good commit / script version
  3. start the service once in foreground or via a controlled launchd reload
  4. verify health and behavior before restoring automatic supervision

Current recommendation for this environment

Before the next production restart of the gateway:

  • remove --replace from ~/Library/LaunchAgents/ai.hermes.gateway.plist
  • keep the gateway in RunAtLoad + KeepAlive, but with plain gateway run
  • restart in a controlled sequence
  • verify both 8642 and 8653
  • smoke-test streaming and managed-download rewriting immediately after restart