| 2026-10-08 |

subagents, secrets: an honest result contract, and two leaks closed
...
A pass over the sub-agent system, plus the two findings from the previous one.
Leaks:
- `redact_args` / `is_sensitive_key` (navi/tools/_internal/redact.py) mask
credential-shaped keys at every log site that prints tool arguments — by key
name, for all tools, so a future MCP tool is covered without declaring
anything. ssh_exec passwords were sitting in the prod journal in plaintext.
- `mcp_servers.d/*.json` are tracked in git and carried the scaffolding
machine's absolute paths, so navi-3d/navi-web silently vanished from the
toolset off that machine. They now hold project-relative paths, resolved in
`McpClient.open_transport` — not at load time, because create_mcp_server
round-trips every config file through save_mcp_servers.
Sub-agent result contract:
- `run_ephemeral` returns a `SubAgentOutcome` (text + status: ok, timeout,
max_iterations, thinking_stall, user_stop, context_overflow) instead of
`(text, bool)`. Every exit routes through `SubAgentRunner._finish`, so token
accounting is never dropped and a partial run always carries a progress
report — the stop-after-stream path used to return a bare string.
- `spawn_agent` renders its header from the status. Before this, every short
run reached the parent as "hit iteration limit", and the parent repeats that
diagnosis to the user — a timeout was reported as an iteration limit.
- `ContextTooLargeError` no longer kills a sub-agent: the parent compresses its
own context, a sub-agent has no session to compress into, so it stops with
`context_overflow` and hands back what it did rather than surfacing as
"Sub-agent failed: …" with the whole run discarded.
- The result is capped at 8000 chars (head and tail kept) so a long run cannot
flood the parent's context, and tool arguments in the progress report go
through `redact_args`.
- The navi-3d calling convention left the runner for the server's own tool
schemas — a generic runner should not know one MCP server's path shapes.
- Validation moved inside spawn_agent's try: a bad argument is a clean
ToolResult the model can read, not an exception in the parent's loop.
Tests: 1636 passed, 1 skipped. New: SubAgentOutcome paths and headers, result
truncation, context overflow, clean argument failures.
Eugene Sukhodolskiy
committed
6 hours ago
|

profiles: tool_developer folds into developer
...
The profile was a duplicate on every axis we could measure. 22 of its 26
native tools were already developer's; the four it alone held —
reload_tools, create_mcp_server, test_mcp_tool, mcp_status — are 2.9 KB of
schema. Its model chain was the same seven models. Of its 14 KB prompt,
about 8 KB was copied verbatim from developer's (the whole `## Orchestration
model` block and everything from `## Editing policy` down), and most of the
remainder restated manuals/create_mcp_server.md, which already carried the
same ten-step workflow in more detail. It was not a specialisation, it was a
snapshot: `git log -S '"reload_tools"'` shows the tool lived in developer
until 61fa370 rewrote that profile around MCP and cut it off.
What kept it alive was a premise that no longer holds — that Navi's own
capabilities would be written as in-repo tools. They are MCP servers now,
and an MCP server is not a file in this repository with a life of its own:
it is an isolated process registered from mcp_servers.d/. So there is no
reason left for a profile whose only distinct feature is a toolset a general
developer profile can hold, and every reason to stop maintaining a second
prompt that drifts against the first.
- developer: + reload_tools, create_mcp_server, test_mcp_tool, mcp_status
(24 → 28 native). Its sub-agent gains tool_manual, which is what it
actually needed to reach the manual while writing a server — the previous
tool_developer sub-agent had it, developer's did not.
- server_admin: + reload_tools only (24 → 25). Adding a third-party MCP
server is something this profile does as often as developer does.
Deliberately not to its sub-agent: reconnecting the MCP manager is a
process-wide operation belonging to the main agent.
- The prompt and the manual took on what the deleted profile knew and
create_mcp_server.md did not: reload_tools before the first test_mcp_tool
(a freshly registered server is not connected, so the test fails and the
iteration is wasted), the smoke test read by exit code — 124 means timeout
killed a server still running, 0 means it exited on its own, usually a
main() without parentheses — absolute command/cwd, mcp_status as discovery
only, and the steps that stay inline instead of going to a sub-agent.
mcp_status and test_mcp_tool were built without an MCP manager, and their
fallback did `from navi.api.deps import _mcp_manager` — a name that does not
exist, so a live call raised ImportError rather than the intended "MCP
manager not available". The tools always passed a manager in tests, which is
why nothing caught it. Both now receive the manager at construction and fall
back to the live one lazily.
Sessions and profile_overrides are reassigned before the restart: agent.py
resolves the session's profile without a guard, so a deleted profile turns
every session that referenced it into an uncaught ProfileNotFound. Nothing
in the test suite pins the profile inventory, and profiles are read once at
import time — reload_tools does not re-read them — so this ships as a
restart, and the restart is also what makes it take effect.
Eugene Sukhodolskiy
committed
7 hours ago
|

tools, mcp: a description is a hook, not a manual
...
With 24 tools in a profile the schemas are the largest fixed block of every
request, and they are paid for whether or not the tool is called. Eleven of
them carried a manual's worth of prose in their `description` — examples,
error tables, "common mistakes", the mechanics of the file areas. Reading
them against the manuals first, the manuals already held all of it, and
richer: this was duplication sitting in the one place the model cannot avoid
reading, not knowledge that had nowhere else to live. So the descriptions
became selection hooks — what the tool is for, when to pick it over its
neighbour, and the one rule that cannot wait — and the detail moved to
tool_manual("<tool>"), which costs nothing until the tool is actually in
play.
filesystem, spawn_agent, schedule_recall, manage_recall, scratchpad, plan,
share_file, content_publish, reflect, todo and peer. Descriptions drop from
18.4 KB to 11.5 KB, those tools' schemas from 25.5 KB to 16.8 KB, and a
profile's whole native toolset from 38.3 KB to 29.6 KB (server_admin,
developer; secretary 33.7 → 25.0, navi_code 32.9 → 25.9, tool_developer
37.9 → 29.2, modeler_3d 33.4 → 24.8). Stated as a budget rather than a byte
count: roughly 4k tokens of the model's window per request, back.
Two things were deliberately not shortened. The filesystem edit ladder
(`edit` → `edit_lines` → `smart_edit`, last resort, costs an LLM call) and
todo's mandatory `validation` on `done` both exist to prevent an extra
round trip; a shorter description there would trade tokens for turns. What
todo lost is the prose around the rule, not the rule. `enabled.json` was
left alone too: weather, gmail and get_current_datetime are opt-in and the
DB shows 12, 16 and 40 calls, so they are opt-in tools that are actually
used.
peer had no manual at all, which is how this surfaced: its description was
the only documentation the tool had. manuals/peer.md is now written from
the tool and its /peer route — the ask/status/list actions, the fact that
the hive is only a phone book and the question travels peer-to-peer, the
one-concurrent-answer semaphore, the two independent recursion guards, and
the six error codes.
The MCP server instructions leave the system prompt the same way. They were
8 KB of always-present prose per request, most of it workflow detail the
model only needs when it is about to use that server. Each server now
contributes one line — a new optional `summary` field in mcp_servers.d/*.json,
defaulting to the first sentence of its `instructions` — and
tool_manual("<server>") returns the full text plus the tool names the server
declares. server_admin's MCP block goes from 8.3 KB to 1.3 KB, secretary's
from 8.1 to 1.2. tool_manual learned to answer for a server (a tool name
still wins over a server of the same name, and dash/underscore spelling is
tolerated), and the index names the servers alongside the tools.
One judgement call worth recording: gnexus-book's and navi-web's
instructions end in an absolute "NEVER bypass these tools with filesystem,
terminal or code_exec" rule. Moving that on demand would have been a
behavioural regression dressed as a token saving, so those two `summary`
fields carry the prohibition verbatim alongside the hook, and it stays
always visible.
docs/tools.md gains a section on the split (a description is paid for every
request, a manual only when the tool is used), which is where the next
person will look before trimming a description back into a manual.
Full suite green (1577 passed, 1 skipped). No frontend change and no new
dependency, so the deploy is a pull and a restart.
Eugene Sukhodolskiy
committed
8 hours ago
|

tool_manual: a manual the agent cannot reach is not a manual
...
tool_manual asked the filesystem for the exact string the model sent, then
fell back to registry.get(). The executor already tolerated the ways models
mangle MCP names; the tool that documents them did not, so the agent calling
mcp__navi-3d__compile_scad got a schema dump while manuals/compile_scad.md sat
there unread. Resolving a name now happens in one place — resolve_tool moved
out of the executor into navi/core/tool_utils.py — so a name that works when
the tool is *called* also works when it is *asked about*. A miss suggests the
closest names instead of "not found", and a tool that exists but is not enabled
for this profile says exactly that: documenting it is still the useful answer,
but the agent must not go on to call it. params["tool_name"] was also a crash
waiting to happen — KeyError with the key omitted, AttributeError on null or a
number — and now anything that is not a usable string means "no name given",
which returns the index of manuals grouped by source.
The generated manual now renders the whole schema. An array of objects showed
as "(array, optional)" with its item shape invisible, so a model that could not
see the fields guessed them; objects, array items, oneOf/anyOf branches, enums
and defaults are all spelled out now, depth-capped. The header names the source
(native / mcp__<server>__) and says the text is a parameter contract, not a
curated manual — otherwise a thin schema reads as "this tool is simple".
Then the manuals themselves. A manual is filed under the tool's own name, and
five of the sixteen were not: write_tool.md and write_mcp_server.md documented
tools that do not exist anywhere (the tools are reload_tools and
create_mcp_server), model_3d.md documented compile_scad, render_3d.md documented
render_stl, and write_context_provider.md documented no tool at all — it is a
guide, and now lives in manuals/guides/. The prompts pointed at the same wrong
names (tool_developer's system prompt, persona_navi_code, docs/context_providers),
so fixing the files without the callers would have moved the breakage rather
than removed it. On HEAD the probe was blunt: tool_manual("mcp__navi-3d__compile_scad")
returned None and the agent got the schema; now it returns the manual.
Nothing was enforcing any of this, which is how it drifted. tests/unit/tools/
test_manual_drift.py reads the real repository tree: every manuals/*.md must be
named after a tool that exists (built-ins and tools/*.py by their name
assignments, MCP tools by @mcp.tool(name=…) in the server sources and by the
mcp_servers.d groups, whose file name is the server name), every manual must be
reachable by both its bare name and its full MCP spelling, a guide must never
shadow a tool's name, and every manuals/<x>.md or tool_manual("x") cited in a
doc or a prompt must resolve. The next rename fails the suite instead of
quietly wasting a manual.
Finally, manuals for the built-ins the agent uses most — ssh_exec, todo, memory,
plan, filesystem, notify, code_exec — written from the current schemas and the
code paths that produce the errors, so the gotchas are the real ones (todo's
`done` validation is enforced; memory's `list` returns categories, not facts;
filesystem's `delete` removes a tree with no prompt). reload_tools.md absorbs
the file format from the deleted write_tool.md, which is where a self-extension
recipe belongs now.
1556 passed, 1 skipped. ruff delta zero.
Eugene Sukhodolskiy
committed
9 hours ago
|

agent: a profile switch now survives the turn that made it
...
switch_profile repoints the session with one narrow UPDATE, and the run's own
save() — which lands right after every tool batch, the switch's own batch
included — wrote its stale in-memory profile back over it. The post-turn
reload then read the row it had just clobbered, saw no change, and left the
turn bound to the old profile's tools: the agent called exactly the tools
switch_profile had just advertised ("Newly available: ssh_exec, terminal, …")
and got "tool not found" for every one of them. Prod shows both halves — the
session that reported two switches kept the profile it started with, and
agent.profile_reloaded appears zero times in the whole boot.
The reload now compares against the profile the run is bound to
(bound_profile_id), which set_profile cannot move, instead of against
session.profile_id, which the store either rewrites or mutates to the new
value — making the old guard false in both branches by construction. And
save() leaves the profile_id column alone on conflict: set_profile owns it
after creation, the same way name/created_at are owned elsewhere.
A tool the live tool_map does not hold is now logged (tool.not_found, with
profile and live tool count). It previously left no trace at all in the
journal, which is why this took a database dig to find.
Three tests, each verified to fail without its fix: the next LLM call is
offered the new profile's tools, save() does not rewrite profile_id, and a
missing tool is logged.
Eugene Sukhodolskiy
committed
10 hours ago
|

tools: list_tools defaults to the profile the run executes as
...
Asking "what tools do I have?" required the agent to know its own profile id
— and to be right about it after a switch_profile had moved the session. The
run knows: run_stream now publishes the active profile id (re-published when
switch_profile reloads the run) alongside the model, SubAgentRunner publishes
its own rather than inheriting the parent's, and ToolContext carries it into
every execute() call.
list_tools reads that when profile_id is omitted, so the bare call answers the
question, and it says which profile it assumed ("current profile") so the
answer cannot be mistaken for another profile's.
scope='agent'|'subagent' selects which of the profile's two toolsets to list —
the set spawn_agent would hand a sub-agent running that profile, which is a
different list and until now had no way to be inspected.
Item D of the list_tools plan, kept separate from the accuracy/compactness work.
Eugene Sukhodolskiy
committed
10 hours ago
|
| 2026-10-07 |

reload: pick up the new code, and drop MCP tools that are gone
...
reload_tools reported success while three separate things kept it from
doing what it says.
The bytecode cache is keyed on (mtime in whole seconds, file size), so a
tool edited to the same length inside the same second as its previous
load re-ran the OLD code — the reload was real, the new version was not
live. The loader now compiles the source itself instead of consulting
__pycache__ (importlib.invalidate_caches() does not help here).
MCP registrations only ever grew: register_mcp_tools called
register_external, and unregister_external was used in one place, so a
server removed from the config or a tool a server stopped exposing
stayed in the registry and failed only when the model called it. Reload
now clears external tools and rebuilds them.
enabled.json naming a tool that failed to load (the gmail/html2text case)
was visible only as a log line. It is now part of the report.
The reload itself moves to navi/core/reload.py, one implementation shared
by the tool and the admin route, so the two cannot leave different
toolsets behind. list_tools.py also read enabled.json through its own
cwd-relative path, which made it disagree with the real toolset, and the
class-based loader rejected execute(self, params, ctx=None) — the shape
every built-in uses.
Eugene Sukhodolskiy
committed
1 day ago
|
switch_profile: run the switch alone, then the batch on the new tools
...
A tool called in the same batch as switch_profile was still dispatched against
the old tool_map and died with "tool 'X' not found" — reload_tools did, twice
in one session. The batch is now split: the switch runs first, its
ProfileSwitched event is watched for the target profile, the tools are
re-resolved from it, and the remaining calls run against the new set.
The turn's own profile binding is deliberately left alone — the
end-of-iteration reload in run_stream() is what rebinds profile/llm/schemas,
and pre-setting session.profile_id here would make it skip that step.
Eugene Sukhodolskiy
committed
1 day ago
|
switch_profile: deliver profile_switched, report the new toolset
...
The tool took its sink as `ctx.event_sink if ctx else current_event_sink.get()`,
but the agent loop builds tool_ctx with event_sink=None — the ContextVar is the
real channel — so the branch always picked None, the event was silently
dropped, and the profile badge in the header never moved. plan.py already had
the fall-through (`if ctx and ctx.event_sink else current_event_sink.get()`);
this is the same fix.
The result also said nothing about what the switch changed, so the model kept
calling tools the target profile does not have (reload_tools after leaving
tool_developer). It now names the gained and lost tools, and the timing claim
is corrected: the new prompt and tools are in force from the next step of the
same turn, not "from the next message".
Eugene Sukhodolskiy
committed
1 day ago
|

Allocate message sequence numbers in the DB, not in memory
...
save() numbered new messages from Session.db_next_sequence, a counter read
when the session was loaded. Any second writer of the same session handed
out the numbers it still believed were free, and the two inserts collided on
UNIQUE(session_id, sequence_number) — in production, when switch_profile
loaded the session mid-run and saved it back. The turn died with
"Internal error: duplicate key value violates unique constraint".
The range is now claimed inside save()'s transaction with a single
UPDATE ... RETURNING, so concurrent writers serialize on the session row,
and GREATEST() seeds the pre-counter sessions whose next_sequence is still 0
instead of relying on a racy max()+1 fallback in memory.
switch_profile itself no longer saves a session at all: it repoints the row
through a narrow set_profile() UPDATE, which keeps it out of the running
turn's way. It runs mid-turn on a session the turn still holds, so saving a
second copy from there was the collision in the first place.
Eugene Sukhodolskiy
committed
1 day ago
|
mcp keys W2: per-user client clones, KeyResolver wiring, user_id in call path
Eugene Sukhodolskiy
committed
1 day ago
|
| 2026-10-06 |
synapse reactions W3: reaction runner, notify tool, gateway hookup
...
- navi/synapse/reactions.py: fire-and-forget reaction run — per-user
reactions_enabled gate, one-shot dispatcher meta-pass (visible profiles
only, hidden dispatcher, secretary fallback), special=True session
headed by the resolved profile, headless agent run via the orchestrator
run registry, finalise: app push per completion_notify +
low-level reaction.finished/failed event to Synapse
- navi/synapse/outbound.py: source-side emit (syn_* key from env,
silent when not configured)
- navi/tools/notify.py: notify tool — message + level (info/warning/
intervention), both legs per push_target, honest delivered/skipped result
- push/service.py: notify_custom without cooldown
- gateway: verified + non-duplicate delivery schedules the reaction
- profiles: notify added to native tools (7 profiles)
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W2: hidden dispatcher profile + synapse_instructions tool
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W1: special sessions flag + per-user reaction settings and instruction history
Eugene Sukhodolskiy
committed
2 days ago
|
webclient + navi: session title click-to-rename; empty rename regenerates via LLM
Eugene Sukhodolskiy
committed
2 days ago
|
navi: session naming folds agent's final replies into context
Eugene Sukhodolskiy
committed
2 days ago
|
navi: session delete covers pending (not-yet-persisted) sessions
...
Lazy persistence keeps never-messaged sessions out of the DB, so the bare
DELETE FROM sessions reported 0 rows for a very real session and the API
answered 404 'Session not found' — the webclient's list cleanup and the
welcome-screen redirect silently never ran.
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-10-05 |

compressor: single-summary invariant — old summaries always fold into the fresh one
...
partition_messages now strips is_summary messages out of the turn structure
before grouping: a summary is already-compressed history, not a user request.
Previously its role=user made it a turn head, so midturn compression kept the
old summary verbatim beside the fresh one (head kept + new summary added),
and adaptive swap resurrected summary turns whose text pattern-matches the
importance heuristics. This piled stale summary copies into the context
(one session accumulated 6 copies = 120k of 239k chars) and made compression
no-op for whole stretches before a lucky pass finally shrank it.
Also folds summaries on the early-exit path so two stale summaries
consolidate via the meta-summary, and fixes the demotion loop in
compress_and_save_session to cover session.context too: after a reload,
context-only rows (is_display=False summaries) were absent from
session.messages and never demoted, so is_context=True persisted in the DB
and the copies survived into the loaded context.
Eugene Sukhodolskiy
committed
3 days ago
|

background tasks: review fixes B1-B11 batch
...
- stop mid-batch cancels in-flight tools (B1) and keeps real results
of already-finished ones, mixing them with synthetic stopped notes
in call order (B2)
- queued messages: headless drain publishes session_sync (B4), the
run's teardown broadcasts session_sync to other sockets but not the
owner socket (prevents double reload) (B5)
- task_update notifications are chained per task so late running
updates can't overtake the terminal one (B7; client drops late
re-flicker of a terminal task chip) (B8)
- client: ui_component results reference the owning message via
card.parentMsg instead of a stale msg reference (B6)
- tasks cancel reports the real outcome after a bounded wait instead
of an optimistic 'cancelled' (B9)
- session delete cancels its running background jobs and drops
pending result notes (B10)
- subagent tool loop routes background:true calls through
ToolExecutor._maybe_background like the main loop (B11)
- tests for all of the above; docs/tasks.md stop/cancel semantics
Eugene Sukhodolskiy
committed
3 days ago
|

local mode isolation: anonymous user is a NULL-owner user, not an alias
...
Local mode (NAVI_AUTH_ENABLED=false) used an anonymous admin whose admin
role widened every session listing to ALL users' chats — acceptable only
while assuming a private database, but a real leak on any shared one.
- session listings (list_all/list_page/count_all/search_list): new
scoping rule — non-admin with user_id=None sees only user_id IS NULL
rows; named owner unchanged; admin unchanged
- session create/list routes: owner id is None in local mode (sessions
are persisted ownerless), admin listing flag resolved only when auth
is enabled
- check_session_access: local mode allows only NULL-owner sessions —
a user-owned session by id is 403
- debug/admin endpoints (require_admin) stay reachable in local mode
- prod DB: 30 stray user_id='anonymous' rows reassigned to NULL owner
Tests: store scoping unit tests, local-mode integration tests (sidebar
filtering + access by id), updated check_session_access unit tests.
Full pytest 1239 passed, 1 skipped.
Eugene Sukhodolskiy
committed
3 days ago
|

chat history pagination: paged load instead of full-session fetch
...
Backend:
- PgSessionStore.get_meta — light session head (sessions row + COUNT),
no message payload
- PgSessionStore.get_history_page — page of display history (archive ∪
hot UNION), newest-first page, oldest-first result, limit+1 peek for
has_more; each message stamped metadata.display_index (ROW_NUMBER-1
over the full display history) so client ids stay stable across
paged loads; dangling-tool-call repair keeps real ranks by
sequence-number lookup
- GET /sessions/{id}/meta and GET /sessions/{id}/messages/history
(before_seq cursor, limit ≤200)
Webclient:
- loadSession/reloadSession fetch meta + newest 200-item page
(Promise.all), archive state from the page itself
- buildMessageList ids (h_*) and rawIndices use the global
display_index — feedback keys and search-jump stay stable when
history is paged
- scroll-up loads older pages through the same history endpoint
(archive table alone misses sessions whose threshold never moved)
- search-jump pulls pages until the target index is covered (bounded)
Tests: store unit tests (ranks, has_more peek, placeholder shift) +
vitest updates; full pytest 1231 passed, vitest 94 passed.
Eugene Sukhodolskiy
committed
3 days ago
|
| 2026-09-26 |
context: task-note messages survive build()'s system filter
...
E2E caught a real delivery bug: the completion note was drained and
persisted (agent.task_notes_drained logged) but build() filtered out all
system-role history from session.context, so the LLM never saw it — the
agent could only get detached results via tasks check. Notes now carry
metadata source=task_note (kept by build) and is_display=False (clients
already show live task_update events). Verified live: agent reads the
note verbatim without calling any tool.
Eugene Sukhodolskiy
committed
12 days ago
|
e2e fixes: tasks tool in profiles, bg timeout lift, stepIcon fix, queue semantics docs
...
E2E findings addressed:
- tasks tool was registered but not in any profile's tools.agent.native —
added to all six profiles (agent could not check/wait/cancel bg tasks)
- detached terminal/code_exec/ssh_exec runs without explicit timeout are
lifted to 300s: foreground defaults (20/30/60s) marked long commands
'completed' with partial output while the process still ran
- ToolCard.vue: define stepIcon(status) — template referenced it but the
function was missing (render crash on task_update step)
- message_queued reachability documented: WS read loop is sequential, the
queue path is only reachable from a second socket/headless recall
Eugene Sukhodolskiy
committed
12 days ago
|
container: runtime-import PushService — master could not boot since the PWA commit
...
PushService was only imported under TYPE_CHECKING; create_container raised
NameError at startup. Caught by the first real boot (e2e) — the running
production container predates the PWA commit, so master was silently
un-launchable.
Eugene Sukhodolskiy
committed
12 days ago
|
agent parallelism: background tools, bg spawn_agent, parallel tool batches, message queue
...
- backgroundable tools (terminal/ssh_exec/peer/spawn_agent/code_exec) detach
via TaskManager with per-session/global/spawn caps, rate limit and TTL
- completion delivery both ways: task_update event (out-of-band) + pending
notes injected into the next turn; tasks tool (list/check/wait/cancel)
- parallel tool-call batches (profile-level gate, off by default) with
ToolStarted up-front, tagged event mux, call-order results, one save,
plus dangling tool_call repair on session load
- user message queue instead of busy error: message_queued frame,
back-to-back drain on the same socket, headless fallback on disconnect
- pin mcp<2 (v2 renames FastMCP with breaking API changes)
- docs: tasks.md (new), websocket/api/agent/config updates, tasks manual,
spawn_agent background param, persona contract section
Eugene Sukhodolskiy
committed
12 days ago
|

PWA: installable webclient, offline shell, web push
...
Installability:
- public/manifest.webmanifest (standalone, theme #16161E) + PNG icons
generated from logo.svg (regular + maskable, served via /images mount)
- index.html: manifest link, theme-color, apple-touch-icon
Offline shell:
- hand-rolled sw.js (no workbox): navigation = network-first (3s race)
with cached-shell fallback + background refresh (a stale cached shell
would 404 on entry chunks after a deploy); /assets/* cache-first
(content-hashed); /images/* cache-first capped; /api,/ws,/auth,/push,
/content pass-through
- vite closeBundle plugin stamps __NAVI_BUILD_VERSION__ (digest of
index.html + asset names) into dist/sw.js; sw.js served no-store so
every deploy reactivates the SW and activation evicts old caches
- SW registration in main.js, PROD only (dev HMR untouched)
- OfflineBanner (useOnline composable) over the app shell
Web push (VAPID, pywebpush):
- navi/push/ package: push_subscriptions table (postgres, boot-time DDL),
PushSubscriptionStore, PushService (async fan-out, to_thread sends,
404/410 prunes dead endpoints, per-session cooldown)
- routes: GET /push/vapid-key, POST/DELETE /push/subscribe (auth-gated)
- trigger in orchestrator run_agent + run_recall: push on StreamEnd when
no WebSocket client watches the session; fire-and-forget, never
disturbs the run; anonymous fallback only when auth is off
- client: usePush composable + Notifications settings panel (enable/
disable via PushManager.subscribe with the server VAPID key)
- notification click focuses the app at /#<session_id> (hash routing
opens the right chat); payload body is a markdown-stripped <=140-char
preview
NAVIVAPID keys empty = push fully disabled (graceful, like other optional
integrations). dist/ artifacts committed per repo convention.
Tests: pytest push store/service/routes/trigger (+23), vitest usePush
(83 webclient tests green). Full suite 1143 passed.
Eugene Sukhodolskiy
committed
12 days ago
|

planning: fix LLM-output failure modes that killed every plan
...
Production planning_logs showed three ways a plan silently dies:
1. glm-5.3-flash wraps the whole answer in a structured envelope
(response:unknown{value:...}<tool_call|>) — the planner only knew the
gemma4 thought<channel|> artifact, so the wrapped analysis/plan failed
to parse. The stripper now unwraps response:<type>{value:...} envelopes,
including a truncated (unbalanced) one.
2. Empty content with spent completion tokens: glm on ollama-cloud puts
the whole answer in the thinking channel on some calls. Both planning
phases now fall back to thinking when content is empty; the debug log
records which channel served (source: content|thinking).
3. phase3_timeout at the shared 120s LLM_COMPLETE_TIMEOUT: planning now
has its own PLANNING_LLM_TIMEOUT_SEC (default 240).
Unit tests cover the envelope unwrap (balanced/truncated/inner braces),
the thinking fallback in both phases, and the clean-fail path.
Eugene Sukhodolskiy
committed
12 days ago
|
| 2026-09-10 |
swarm stage 2: peer-to-peer channel — /peer endpoints + peer tool
...
Server side (navi/api/routes/peer.py, port 8099):
- GET /peer/hello — open discovery ping (name, uuid prefix, version)
- GET /peer/status — PSK; identity, uptime, machine facts, hive view
- POST /peer/ask — PSK; one-shot agent run under PEER_ASK_PROFILE
(default server_admin) answers, one concurrent ask at a time
- loop guard: answering agent runs without the peer tool (deterministic
recursion cut) + own-uuid asks refused with 409
- verify_swarm_key in navi/swarm.py accepts .swarm-key.previous
(rotation window, constant-time)
Client side (navi/tools/peer.py):
- peer tool: list (hive book, stale cache fallback marked), status,
ask by swarm name; asks go direct peer-to-peer, hive never on the
message path
- audit events on both sides (peer.ask_sent/received/answered/failed)
peer tool added to server_admin, navi_code, developer profiles.
Eugene Sukhodolskiy
committed
29 days ago
|
| 2026-09-09 |
swarm stage 1: instance identity, PSK, hive registry, announce loop
...
- navi/identity.py: adjective-animal names + uuid in instance.json
(generated at install, renameable by hand)
- navi/swarm.py: HiveAnnouncer - periodic POST /announce to the hive
with non-blocking reachability tracking (transitional logs, /health
export, hive_status context provider reports outages to the agent)
- hive/: standalone FastAPI address book (SQLite, port 8087, PSK via
X-Swarm-Key with .swarm-key.previous rotation window, TTL online
status, host from client IP, port from payload). Not installed or
started by default - run manually on the main server.
- ports moved to 8099 (API) / 8098 (UI MCP)
- install.sh: PYTHONIOENCODING=utf-8 in the systemd unit, generates
instance.json and .swarm-key on fresh installs
Eugene Sukhodolskiy
committed
29 days ago
|

compress: route every trigger path through compress_and_save_session
...
- CompressionWorker keeps the real-token gate (ctx.context_tokens) but now
delegates to compress_and_save_session(reason="postturn") — it used to
call compress_context directly with its own save logic and silently
no-op'ed on any summarizer failure, leaving the session over the
threshold until the next turn's gates
- pre-turn and /compact gates use real_baseline_estimate instead of the
chars//3 heuristic (mid-turn already did), so code-heavy contexts
compress in time instead of tripping ContextTooLargeError
- summarizer input keeps head + tail instead of head-only truncation,
which silently dropped the newest summarized messages closest to the
keep window
- critical tool results over the 4000-char budget keep head+tail halves
instead of collapsing to the 800-char non-critical preview (a 3999-char
read survived whole, a 4001-char one lost 98%)
- images count 500 tokens each, not 500 per message carrying them
- _plan_compression is the single partition decision shared by
compress_context and would_compress, so the dry-run prediction cannot
drift from the real path; the trigger reason is now logged
(preturn/midturn/postturn/forced)
Eugene Sukhodolskiy
committed
29 days ago
|