# Updating a running Navi instance

Runbook for the agent (or human) updating a deployed Navi. Written after a
release where the prod frontend had been rebuilt but the fixes were still
missing — the checks below exist to make that failure mode impossible to miss.

## What an update actually is

- Code comes from GitBucket. The frontend bundle (`webclient/dist/`) is
  **tracked in git** and served by the API server, so an update is
  `git pull` + `systemctl restart navi`. **There is no node/npm on the server
  and no separate frontend deploy step.**
- `deploy/install.sh` is only needed when `pyproject.toml` gained
  dependencies — it re-runs `pip install -e .`. It is idempotent and safe to
  re-run any time.
- Database changes are additive DDL executed at startup
  (`CREATE TABLE IF NOT EXISTS`). No migration step, no DB rollback needed.
- The API server must be restarted for backend changes. Config files under
  `mcp_servers.d/` are re-read per request, but table creation happens at
  startup, so restart anyway.

## Step 0 — find out what prod runs

```bash
systemctl cat navi | grep -E 'WorkingDirectory|ExecStart'   # where the clone is
cd <clone-dir>
git rev-parse --abbrev-ref HEAD      # master or deploy?
git log -1 --oneline                 # the currently deployed commit
git status --short                   # must be clean before pulling
```

Two shapes, and the update differs:

- **prod tracks `master`** — `.env` is untracked on the server. Update is a
  plain fast-forward `git pull`.
- **prod tracks `deploy`** — `.env` **is tracked in that branch** (it is the
  branch's only diff from master). The branch must be rebased onto master
  *beforehand, on the development machine*, and never by checking out `deploy`
  in the main development worktree: master does not track `.env`, so that
  checkout silently deletes or clobbers it (this has happened twice).

## Preflight (on the prod host, both shapes)

```bash
cp .env ~/navi-env-$(date +%F-%H%M).bak   # back up; never copy anything ONTO .env
git rev-parse --short HEAD                # record for rollback
git status --short                        # modified mcp_servers.d/*.json or .env?
```

If `git status` shows local edits to `mcp_servers.d/*.json` (these files are
tracked and contain credentials), decide explicitly whether the server keeps
its local values or takes the repo's — otherwise the pull will refuse or a
reset will overwrite them. Optional DB safety net:

```bash
sudo docker exec navi-postgres pg_dump -U navi navi | gzip > ~/navi-$(date +%F).sql.gz
```

## Dev-side prep (before touching prod)

1. `cd webclient && npm run test && npm run build` — **the dist must be
   rebuilt and committed together with the source.** A code change without a
   rebuilt dist leaves prod serving the previous UI while `git log` looks
   updated; this is exactly how the `notificationsSupported` fix stayed off
   prod, because the commit carrying it had never been pushed.
2. `cd .. && uv run pytest tests -q`.
3. Commit source + `webclient/dist` together, then push master.
4. If prod is on `deploy`: in a **separate worktree** (e.g. `/tmp/navi-deploy`),
   back up `.env`, `git rebase master`, confirm `git diff --stat master --`
   shows only `.env`, then `git push --force-with-lease origin deploy`.

## Update (prod host)

```bash
cd <clone-dir>
cp .env <backup>                 # again, right before touching git
git fetch origin
git pull --ff-only               # on deploy: git reset --hard origin/deploy
bash deploy/install.sh           # ONLY if pyproject.toml changed since the deployed commit
sudo systemctl restart navi
systemctl status navi --no-pager
journalctl -u navi -n 50 --no-pager
```

## Verify — do not skip

```bash
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8099/health      # 200

# 1. The served bundle is the freshly built one: compare with the repo's index.html
curl -s https://navi.gnexus.space/ | grep -o 'assets/index-[^"]*\.js'
grep -o 'assets/index-[^"]*\.js' webclient/dist/index.html

# 2. The bundle really contains the new code
js=$(curl -s https://navi.gnexus.space/ | grep -o 'assets/index-[^"]*\.js' | head -1)
curl -s "https://navi.gnexus.space/$js" > /tmp/prod.js
grep -c enableBackgroundMode /tmp/prod.js   # >= 1  Android JS-bridge push shipped
grep -c navi_push_android    /tmp/prod.js   # >= 1
grep -c notificationsSupported /tmp/prod.js # == 0  the reload-reset bug
grep -c accepts_user_key     /tmp/prod.js   # >= 1  new MCP panel

# 3. Web push still configured (503 => NAVI_PUSH_VAPID_* missing from .env)
curl -s -o /dev/null -w '%{http_code}\n' https://navi.gnexus.space/push/vapid-key   # 200

# 4. New MCP route present (401 = auth on and route exists; 404 = not deployed)
curl -s -o /dev/null -w '%{http_code}\n' https://navi.gnexus.space/mcp-keys

# 5. Table created at startup
sudo docker exec -it navi-postgres psql -U navi -d navi -c '\d mcp_user_keys'
```

Then in the UI, after logging in: **Settings → MCP** must list the servers
(key input on the ones that declare a `user_key` slot), and
**Settings → Notifications** must show a toggle instead of "this browser does
not support web push".

## Rollback

```bash
git reset --hard <commit-recorded-in-preflight>
sudo systemctl restart navi
```

The frontend lives in the same tree, so it rolls back with the code. Schema
changes are additive only. Restore `.env` from the backup if it was edited.

## Traps worth remembering

- **dist not rebuilt or not committed** → old UI on prod, new code in git.
  The server never builds the frontend.
- **`vendor/gnexus-ui-kit/dist` is gitignored** (`.gitignore:12-14`) — a fresh
  clone cannot run `npm run build`. Production is unaffected while
  `webclient/dist` is committed, but any pipeline that builds the frontend on
  a clean clone needs that snapshot committed first.
- **Browser / service-worker cache**: `dist/sw.js` is stamped with a build
  digest and calls `skipWaiting()` + `clients.claim()`, so one extra reload is
  enough; an installed PWA may need to be closed and reopened.
- **Android app**: no APK reinstall needed. The JS bridge lives in the
  installed APK and the UI it loads comes from the server
  (WebView `cacheMode = LOAD_NO_CACHE`). Reopen the app after the deploy.
- **`.env` trap**: on `deploy` it is tracked; `git checkout master` deletes it.
  Only ever move the branch toward `deploy`, and back it up first.
- **Secrets in the tree**: `mcp_servers.d/*.json` is tracked and carries the
  shared `gcr_…` credential in plaintext. A pull can overwrite the prod copy —
  diff before accepting.
- **Config slot must exist on prod too**: the MCP personal-key input appears
  only for servers whose prod-side config declares `user_key`. If the prod
  copy of `mcp_servers.d/gnexus-creds.json` is not updated, the panel shows
  the row read-only.

## This release, specifically (e68c33e → 10385c6)

- Android WebView push through the `NaviAndroid` JS bridge, with a foreground
  service holding the WebSocket while minimized; the `notificationsSupported`
  ReferenceError is fixed, so the web-push toggle no longer resets after a
  reload.
- MCP settings tab lists every connected server instead of only the ones with
  a key slot; `gnexus-creds` declares
  `user_key: {header: Authorization, prefix: "Bearer "}`.
- Backend surface touched: `/mcp-keys` only. No new Python dependencies, so
  `deploy/install.sh` is not required for this one.
