Newer
Older
navi-1 / deploy / UPDATE.md
@Eugene Sukhodolskiy Eugene Sukhodolskiy 11 hours ago 7 KB deploy: runbook for updating a running instance

Updating a running Navi instance

Runbook for the agent (or human) updating a deployed Navi. Written after a release where the prod frontend had been rebuilt but the fixes were still missing — the checks below exist to make that failure mode impossible to miss.

What an update actually is

  • Code comes from GitBucket. The frontend bundle (webclient/dist/) is tracked in git and served by the API server, so an update is git pull + systemctl restart navi. There is no node/npm on the server and no separate frontend deploy step.
  • deploy/install.sh is only needed when pyproject.toml gained dependencies — it re-runs pip install -e .. It is idempotent and safe to re-run any time.
  • Database changes are additive DDL executed at startup (CREATE TABLE IF NOT EXISTS). No migration step, no DB rollback needed.
  • The API server must be restarted for backend changes. Config files under mcp_servers.d/ are re-read per request, but table creation happens at startup, so restart anyway.

Step 0 — find out what prod runs

systemctl cat navi | grep -E 'WorkingDirectory|ExecStart'   # where the clone is
cd <clone-dir>
git rev-parse --abbrev-ref HEAD      # master or deploy?
git log -1 --oneline                 # the currently deployed commit
git status --short                   # must be clean before pulling

Two shapes, and the update differs:

  • prod tracks master — .env is untracked on the server. Update is a plain fast-forward git pull.
  • prod tracks deploy — .env is tracked in that branch (it is the branch's only diff from master). The branch must be rebased onto master beforehand, on the development machine, and never by checking out deploy in the main development worktree: master does not track .env, so that checkout silently deletes or clobbers it (this has happened twice).

Preflight (on the prod host, both shapes)

cp .env ~/navi-env-$(date +%F-%H%M).bak   # back up; never copy anything ONTO .env
git rev-parse --short HEAD                # record for rollback
git status --short                        # modified mcp_servers.d/*.json or .env?

If git status shows local edits to mcp_servers.d/*.json (these files are tracked and contain credentials), decide explicitly whether the server keeps its local values or takes the repo's — otherwise the pull will refuse or a reset will overwrite them. Optional DB safety net:

sudo docker exec navi-postgres pg_dump -U navi navi | gzip > ~/navi-$(date +%F).sql.gz

Dev-side prep (before touching prod)

  1. cd webclient && npm run test && npm run build — the dist must be rebuilt and committed together with the source. A code change without a rebuilt dist leaves prod serving the previous UI while git log looks updated; this is exactly how the notificationsSupported fix stayed off prod, because the commit carrying it had never been pushed.
  2. cd .. && uv run pytest tests -q.
  3. Commit source + webclient/dist together, then push master.
  4. If prod is on deploy: in a separate worktree (e.g. /tmp/navi-deploy), back up .env, git rebase master, confirm git diff --stat master -- shows only .env, then git push --force-with-lease origin deploy.

Update (prod host)

cd <clone-dir>
cp .env <backup>                 # again, right before touching git
git fetch origin
git pull --ff-only               # on deploy: git reset --hard origin/deploy
bash deploy/install.sh           # ONLY if pyproject.toml changed since the deployed commit
sudo systemctl restart navi
systemctl status navi --no-pager
journalctl -u navi -n 50 --no-pager

Verify — do not skip

curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8099/health      # 200

# 1. The served bundle is the freshly built one: compare with the repo's index.html
curl -s https://navi.gnexus.space/ | grep -o 'assets/index-[^"]*\.js'
grep -o 'assets/index-[^"]*\.js' webclient/dist/index.html

# 2. The bundle really contains the new code
js=$(curl -s https://navi.gnexus.space/ | grep -o 'assets/index-[^"]*\.js' | head -1)
curl -s "https://navi.gnexus.space/$js" > /tmp/prod.js
grep -c enableBackgroundMode /tmp/prod.js   # >= 1  Android JS-bridge push shipped
grep -c navi_push_android    /tmp/prod.js   # >= 1
grep -c notificationsSupported /tmp/prod.js # == 0  the reload-reset bug
grep -c accepts_user_key     /tmp/prod.js   # >= 1  new MCP panel

# 3. Web push still configured (503 => NAVI_PUSH_VAPID_* missing from .env)
curl -s -o /dev/null -w '%{http_code}\n' https://navi.gnexus.space/push/vapid-key   # 200

# 4. New MCP route present (401 = auth on and route exists; 404 = not deployed)
curl -s -o /dev/null -w '%{http_code}\n' https://navi.gnexus.space/mcp-keys

# 5. Table created at startup
sudo docker exec -it navi-postgres psql -U navi -d navi -c '\d mcp_user_keys'

Then in the UI, after logging in: Settings → MCP must list the servers (key input on the ones that declare a user_key slot), and Settings → Notifications must show a toggle instead of "this browser does not support web push".

Rollback

git reset --hard <commit-recorded-in-preflight>
sudo systemctl restart navi

The frontend lives in the same tree, so it rolls back with the code. Schema changes are additive only. Restore .env from the backup if it was edited.

Traps worth remembering

  • dist not rebuilt or not committed → old UI on prod, new code in git. The server never builds the frontend.
  • vendor/gnexus-ui-kit/dist is gitignored (.gitignore:12-14) — a fresh clone cannot run npm run build. Production is unaffected while webclient/dist is committed, but any pipeline that builds the frontend on a clean clone needs that snapshot committed first.
  • Browser / service-worker cache: dist/sw.js is stamped with a build digest and calls skipWaiting() + clients.claim(), so one extra reload is enough; an installed PWA may need to be closed and reopened.
  • Android app: no APK reinstall needed. The JS bridge lives in the installed APK and the UI it loads comes from the server (WebView cacheMode = LOAD_NO_CACHE). Reopen the app after the deploy.
  • .env trap: on deploy it is tracked; git checkout master deletes it. Only ever move the branch toward deploy, and back it up first.
  • Secrets in the tree: mcp_servers.d/*.json is tracked and carries the shared gcr_… credential in plaintext. A pull can overwrite the prod copy — diff before accepting.
  • Config slot must exist on prod too: the MCP personal-key input appears only for servers whose prod-side config declares user_key. If the prod copy of mcp_servers.d/gnexus-creds.json is not updated, the panel shows the row read-only.

This release, specifically (e68c33e → 10385c6)

  • Android WebView push through the NaviAndroid JS bridge, with a foreground service holding the WebSocket while minimized; the notificationsSupported ReferenceError is fixed, so the web-push toggle no longer resets after a reload.
  • MCP settings tab lists every connected server instead of only the ones with a key slot; gnexus-creds declares user_key: {header: Authorization, prefix: "Bearer "}.
  • Backend surface touched: /mcp-keys only. No new Python dependencies, so deploy/install.sh is not required for this one.