| 2026-10-08 |

memory+llm: three failures the prod logs gave up
...
All three came out of a survey of the last days of journalctl on prod,
and each one silently destroyed something the user had already paid for.
memory_facts: the no-embedding INSERT bound $13 while its column list
had 12 entries, so every fact written while the embedding backend was
down died on PostgresSyntaxError and the extraction was lost — 47
embed failures in three days, most of them landing in this branch.
Embedding input now gets clipped instead of 400'd away. A 400 takes the
whole embedding with it and recall drops to ILIKE over everything; a
reaction-session prompt (a full event envelope inlined) did that 41
times. settings.embedding_max_chars (6000, 0 disables) caps the input
at the model's window, so recall still works on the head of the text.
Message strips NUL bytes at the model boundary. PostgreSQL text cannot
hold one, and a NUL arriving in a tool result (reading a binary file)
made the whole session_messages INSERT fail with
CharacterNotInRepertoireError — the turn died and the user lost it.
Three times on 2026-10-07. As a Message validator it covers every
writer downstream: session store, kv store, memory extraction.
Each new test was checked to fail with its fix reverted.
Eugene Sukhodolskiy
committed
9 hours ago
|

agent: a profile switch now survives the turn that made it
...
switch_profile repoints the session with one narrow UPDATE, and the run's own
save() — which lands right after every tool batch, the switch's own batch
included — wrote its stale in-memory profile back over it. The post-turn
reload then read the row it had just clobbered, saw no change, and left the
turn bound to the old profile's tools: the agent called exactly the tools
switch_profile had just advertised ("Newly available: ssh_exec, terminal, …")
and got "tool not found" for every one of them. Prod shows both halves — the
session that reported two switches kept the profile it started with, and
agent.profile_reloaded appears zero times in the whole boot.
The reload now compares against the profile the run is bound to
(bound_profile_id), which set_profile cannot move, instead of against
session.profile_id, which the store either rewrites or mutates to the new
value — making the old guard false in both branches by construction. And
save() leaves the profile_id column alone on conflict: set_profile owns it
after creation, the same way name/created_at are owned elsewhere.
A tool the live tool_map does not hold is now logged (tool.not_found, with
profile and live tool count). It previously left no trace at all in the
journal, which is why this took a database dig to find.
Three tests, each verified to fail without its fix: the next LLM call is
offered the new profile's tools, save() does not rewrite profile_id, and a
missing tool is logged.
Eugene Sukhodolskiy
committed
9 hours ago
|
mcp: point navi_ui at the port its server actually binds
...
navi_ui is navi's own MCP server (navi/mcp/ui_server), started in-process on
NAVI_UI_MCP_PORT — 8098. The config still carried the pre-8099/8098 default,
localhost:8001, where nothing has listened for a long time: on prod the connect
fails on every start and render_component only ever appeared as an unregistered
phantom in list_tools.
The URL now names the loopback literal rather than "localhost", which can
resolve to ::1 while the server binds IPv4 only.
Three tests assert the file against the live FastMCP object (host, port, path,
transport) instead of a second copy of the setting — with the old URL two of
them fail, so the drift that produced this cannot come back unnoticed.
Eugene Sukhodolskiy
committed
10 hours ago
|

tools: list_tools defaults to the profile the run executes as
...
Asking "what tools do I have?" required the agent to know its own profile id
— and to be right about it after a switch_profile had moved the session. The
run knows: run_stream now publishes the active profile id (re-published when
switch_profile reloads the run) alongside the model, SubAgentRunner publishes
its own rather than inheriting the parent's, and ToolContext carries it into
every execute() call.
list_tools reads that when profile_id is omitted, so the bare call answers the
question, and it says which profile it assumed ("current profile") so the
answer cannot be mistaken for another profile's.
scope='agent'|'subagent' selects which of the profile's two toolsets to list —
the set spawn_agent would hand a sub-agent running that profile, which is a
different list and until now had no way to be inspected.
Item D of the list_tools plan, kept separate from the accuracy/compactness work.
Eugene Sukhodolskiy
committed
10 hours ago
|

tools: list_tools tells the truth and costs a tenth of the context
...
The tool read the profile config, not the registry: a tool the config declares
but nothing registered — mcp__navi_ui__render_component, whose server never
connects — was advertised as callable, and the agent walked into "tool not
found" (28 such warnings on prod). It now resolves every name against the live
registry and reports the rest separately as "Not registered", so a phantom
entry reads as a broken server, not as a tool to try.
It also returned every description it could find. For server_admin that was
125 tools and 35 316 bytes per call — ~9k tokens to answer "do I have anything
for ssh". Names are now grouped by source (native, then one section per MCP
server) with descriptions behind verbose, and query filters by substring over
names and descriptions: the ssh question costs 88 bytes instead of 35 KB, and
a plain listing drops to 3.7 KB (9.5x smaller).
Tests cover the phantom-tool case, query matching by name and by description,
and the size property that makes names-only the default.
Eugene Sukhodolskiy
committed
10 hours ago
|
| 2026-10-07 |
admin: POST /admin/tools/reload
...
reload_tools is granted to one profile only (tool_developer), so an admin
whose profile lacks it cannot reload at all — the tool answers "not
found" instead of reloading. The route calls the same reload_all() the
tool does, behind require_admin, for whoever is logged in as an admin.
An admin may read: ok, the tools loaded, the registry total, per-file
errors, names in enabled.json nothing answers to, and the MCP/providers
summary.
Eugene Sukhodolskiy
committed
1 day ago
|

reload: pick up the new code, and drop MCP tools that are gone
...
reload_tools reported success while three separate things kept it from
doing what it says.
The bytecode cache is keyed on (mtime in whole seconds, file size), so a
tool edited to the same length inside the same second as its previous
load re-ran the OLD code — the reload was real, the new version was not
live. The loader now compiles the source itself instead of consulting
__pycache__ (importlib.invalidate_caches() does not help here).
MCP registrations only ever grew: register_mcp_tools called
register_external, and unregister_external was used in one place, so a
server removed from the config or a tool a server stopped exposing
stayed in the registry and failed only when the model called it. Reload
now clears external tools and rebuilds them.
enabled.json naming a tool that failed to load (the gmail/html2text case)
was visible only as a log line. It is now part of the report.
The reload itself moves to navi/core/reload.py, one implementation shared
by the tool and the admin route, so the two cannot leave different
toolsets behind. list_tools.py also read enabled.json through its own
cwd-relative path, which made it disagree with the real toolset, and the
class-based loader rejected execute(self, params, ctx=None) — the shape
every built-in uses.
Eugene Sukhodolskiy
committed
1 day ago
|
switch_profile: run the switch alone, then the batch on the new tools
...
A tool called in the same batch as switch_profile was still dispatched against
the old tool_map and died with "tool 'X' not found" — reload_tools did, twice
in one session. The batch is now split: the switch runs first, its
ProfileSwitched event is watched for the target profile, the tools are
re-resolved from it, and the remaining calls run against the new set.
The turn's own profile binding is deliberately left alone — the
end-of-iteration reload in run_stream() is what rebinds profile/llm/schemas,
and pre-setting session.profile_id here would make it skip that step.
Eugene Sukhodolskiy
committed
1 day ago
|
switch_profile: deliver profile_switched, report the new toolset
...
The tool took its sink as `ctx.event_sink if ctx else current_event_sink.get()`,
but the agent loop builds tool_ctx with event_sink=None — the ContextVar is the
real channel — so the branch always picked None, the event was silently
dropped, and the profile badge in the header never moved. plan.py already had
the fall-through (`if ctx and ctx.event_sink else current_event_sink.get()`);
this is the same fix.
The result also said nothing about what the switch changed, so the model kept
calling tools the target profile does not have (reload_tools after leaving
tool_developer). It now names the gained and lost tools, and the timing claim
is corrected: the new prompt and tools are in force from the next step of the
same turn, not "from the next message".
Eugene Sukhodolskiy
committed
1 day ago
|

Allocate message sequence numbers in the DB, not in memory
...
save() numbered new messages from Session.db_next_sequence, a counter read
when the session was loaded. Any second writer of the same session handed
out the numbers it still believed were free, and the two inserts collided on
UNIQUE(session_id, sequence_number) — in production, when switch_profile
loaded the session mid-run and saved it back. The turn died with
"Internal error: duplicate key value violates unique constraint".
The range is now claimed inside save()'s transaction with a single
UPDATE ... RETURNING, so concurrent writers serialize on the session row,
and GREATEST() seeds the pre-counter sessions whose next_sequence is still 0
instead of relying on a racy max()+1 fallback in memory.
switch_profile itself no longer saves a session at all: it repoints the row
through a narrow set_profile() UPDATE, which keeps it out of the running
turn's way. It runs mid-turn on a session the turn still holds, so saving a
second copy from there was the collision in the first place.
Eugene Sukhodolskiy
committed
1 day ago
|
Await the session pool in notify and three neighbours
...
PgSessionStore._get_pool() is async, and four call sites passed its coroutine
straight into a store constructor, so the first query inside died with
"'coroutine' object has no attribute 'fetchrow'": notify always failed, the
reaction runner never got past reading its settings, synapse_instructions was
unusable, and the BYOK resolver caught the AttributeError and silently fell
back to the default credential.
The tests missed it because each one mocked the store or the pool provider
away; the new ones run on a fake asyncpg pool with nothing faked below the tool.
Eugene Sukhodolskiy
committed
1 day ago
|

MCP settings tab: list every connected server, slot only where declared
...
The tab was empty on every install: it listed only servers whose config
declares a `user_key` slot, and no config declared one — which read as
"no MCP servers connected" even though five are wired to profiles.
- GET /mcp-keys now returns every server referenced by at least one
profile, keyed ones first, with `accepts_user_key`, the slot location
(null when there is none) and the profile ids that connect it. The
per-user key store is skipped entirely when nothing has a slot.
- gnexus-creds declares `user_key: {header: Authorization, prefix:
"Bearer "}` — it is the one server carrying a shared credential, so its
personal-key field is now real: users with a key run under their own,
users without one fall back to the shared default.
- The panel lists all servers (transport + profiles), dims the keyless
rows, and shows a key input only for slotted ones, spelling out the
shared-key fallback.
docs/api.md and docs/mcp.md updated; backend 1367 passed, webclient 148.
Eugene Sukhodolskiy
committed
1 day ago
|
mcp keys: tool-management tools aligned with BYOK (test_mcp_tool runs on user key, mcp_status marks key slots)
Eugene Sukhodolskiy
committed
1 day ago
|
mcp keys W2: per-user client clones, KeyResolver wiring, user_id in call path
Eugene Sukhodolskiy
committed
1 day ago
|
mcp keys W1: user_key config slot, encrypted keystore, /mcp-keys REST
Eugene Sukhodolskiy
committed
1 day ago
|
| 2026-10-06 |
synapse settings: source_ready flag in GET/PUT response
...
Tells the UI whether Synapse-linked options are usable (source key
configured); not persisted, computed per request.
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W3: reaction runner, notify tool, gateway hookup
...
- navi/synapse/reactions.py: fire-and-forget reaction run — per-user
reactions_enabled gate, one-shot dispatcher meta-pass (visible profiles
only, hidden dispatcher, secretary fallback), special=True session
headed by the resolved profile, headless agent run via the orchestrator
run registry, finalise: app push per completion_notify +
low-level reaction.finished/failed event to Synapse
- navi/synapse/outbound.py: source-side emit (syn_* key from env,
silent when not configured)
- navi/tools/notify.py: notify tool — message + level (info/warning/
intervention), both legs per push_target, honest delivered/skipped result
- push/service.py: notify_custom without cooldown
- gateway: verified + non-duplicate delivery schedules the reaction
- profiles: notify added to native tools (7 profiles)
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W2: hidden dispatcher profile + synapse_instructions tool
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W1: special sessions flag + per-user reaction settings and instruction history
Eugene Sukhodolskiy
committed
2 days ago
|

handbook alignment: profile.updated mirror, health contract, Synapse gateway
...
Handbook gaps closed (10-platform auth/health/notifications):
- webhooks: handle user.profile_updated — mirror the gnexus-auth profile
into navi_users (name, contact fields, avatar_url — new column, boot
migration); role/permissions keep their dedicated events
- auth: avatar_url now flows through the login upsert and the API-token
resolution path
- /health: status is now the aggregate of the sub-checks (degraded embed
or hive no longer reports plain ok); embed probe cached for 10 s; body
grew the machine-readable checks{} map, legacy embed/hive payloads kept
- Synapse: s2s delivery gateway — POST /webhooks/synapse (both spellings),
per-user delivery secrets (synapse_targets, Fernet-encrypted, shown
once), verification via gnexus-synapse v0.1.2 client lib, idempotent
event ids; deliveries are verified and logged for now — navi generates
no outbound events by decision; web-push stays as the local browser
channel
- webclient settings: Synapse deliveries panel (add/revoke targets)
Eugene Sukhodolskiy
committed
2 days ago
|
webclient + navi: session title click-to-rename; empty rename regenerates via LLM
Eugene Sukhodolskiy
committed
2 days ago
|
navi: session naming folds agent's final replies into context
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-10-05 |

compressor: single-summary invariant — old summaries always fold into the fresh one
...
partition_messages now strips is_summary messages out of the turn structure
before grouping: a summary is already-compressed history, not a user request.
Previously its role=user made it a turn head, so midturn compression kept the
old summary verbatim beside the fresh one (head kept + new summary added),
and adaptive swap resurrected summary turns whose text pattern-matches the
importance heuristics. This piled stale summary copies into the context
(one session accumulated 6 copies = 120k of 239k chars) and made compression
no-op for whole stretches before a lucky pass finally shrank it.
Also folds summaries on the early-exit path so two stale summaries
consolidate via the meta-summary, and fixes the demotion loop in
compress_and_save_session to cover session.context too: after a reload,
context-only rows (is_display=False summaries) were absent from
session.messages and never demoted, so is_context=True persisted in the DB
and the copies survived into the loaded context.
Eugene Sukhodolskiy
committed
3 days ago
|

webclient: Backgrounds tab in artifacts panel + task toasts
...
- new Backgrounds tab: task list (running first, badges, per-tool icons)
with a detail view (status, timestamps, sub-agent tokens, result/progress)
- GET /sessions/{id}/tasks snapshot endpoint: task_update events are not
replayed on reconnect, so the client fetches the task list on session
load/reload; live task_update entries are merged on top (chat.fetchTasks)
- terminal task_update no longer removes the entry — it marks it finished
so the tab shows recent completions; _terminalTaskIds still blocks
resurrection by a late running update
- toasts for background task start / finish (info / success / error) in
the WS dispatch; silent for other sessions and unknown terminal ids
- persistent background-tasks chip removed (replaced by the tab); only
the queued-messages chip stays
- fix invisible status text in terminal detail rows: filled .status-*
backgrounds now scope to .terminal-status-badge only
Eugene Sukhodolskiy
committed
3 days ago
|

background tasks: review fixes B1-B11 batch
...
- stop mid-batch cancels in-flight tools (B1) and keeps real results
of already-finished ones, mixing them with synthetic stopped notes
in call order (B2)
- queued messages: headless drain publishes session_sync (B4), the
run's teardown broadcasts session_sync to other sockets but not the
owner socket (prevents double reload) (B5)
- task_update notifications are chained per task so late running
updates can't overtake the terminal one (B7; client drops late
re-flicker of a terminal task chip) (B8)
- client: ui_component results reference the owning message via
card.parentMsg instead of a stale msg reference (B6)
- tasks cancel reports the real outcome after a bounded wait instead
of an optimistic 'cancelled' (B9)
- session delete cancels its running background jobs and drops
pending result notes (B10)
- subagent tool loop routes background:true calls through
ToolExecutor._maybe_background like the main loop (B11)
- tests for all of the above; docs/tasks.md stop/cancel semantics
Eugene Sukhodolskiy
committed
3 days ago
|

local mode isolation: anonymous user is a NULL-owner user, not an alias
...
Local mode (NAVI_AUTH_ENABLED=false) used an anonymous admin whose admin
role widened every session listing to ALL users' chats — acceptable only
while assuming a private database, but a real leak on any shared one.
- session listings (list_all/list_page/count_all/search_list): new
scoping rule — non-admin with user_id=None sees only user_id IS NULL
rows; named owner unchanged; admin unchanged
- session create/list routes: owner id is None in local mode (sessions
are persisted ownerless), admin listing flag resolved only when auth
is enabled
- check_session_access: local mode allows only NULL-owner sessions —
a user-owned session by id is 403
- debug/admin endpoints (require_admin) stay reachable in local mode
- prod DB: 30 stray user_id='anonymous' rows reassigned to NULL owner
Tests: store scoping unit tests, local-mode integration tests (sidebar
filtering + access by id), updated check_session_access unit tests.
Full pytest 1239 passed, 1 skipped.
Eugene Sukhodolskiy
committed
3 days ago
|

chat history pagination: paged load instead of full-session fetch
...
Backend:
- PgSessionStore.get_meta — light session head (sessions row + COUNT),
no message payload
- PgSessionStore.get_history_page — page of display history (archive ∪
hot UNION), newest-first page, oldest-first result, limit+1 peek for
has_more; each message stamped metadata.display_index (ROW_NUMBER-1
over the full display history) so client ids stay stable across
paged loads; dangling-tool-call repair keeps real ranks by
sequence-number lookup
- GET /sessions/{id}/meta and GET /sessions/{id}/messages/history
(before_seq cursor, limit ≤200)
Webclient:
- loadSession/reloadSession fetch meta + newest 200-item page
(Promise.all), archive state from the page itself
- buildMessageList ids (h_*) and rawIndices use the global
display_index — feedback keys and search-jump stay stable when
history is paged
- scroll-up loads older pages through the same history endpoint
(archive table alone misses sessions whose threshold never moved)
- search-jump pulls pages until the target index is covered (bounded)
Tests: store unit tests (ranks, has_more peek, placeholder shift) +
vitest updates; full pytest 1231 passed, vitest 94 passed.
Eugene Sukhodolskiy
committed
3 days ago
|
| 2026-09-26 |
context: task-note messages survive build()'s system filter
...
E2E caught a real delivery bug: the completion note was drained and
persisted (agent.task_notes_drained logged) but build() filtered out all
system-role history from session.context, so the LLM never saw it — the
agent could only get detached results via tasks check. Notes now carry
metadata source=task_note (kept by build) and is_display=False (clients
already show live task_update events). Verified live: agent reads the
note verbatim without calling any tool.
Eugene Sukhodolskiy
committed
12 days ago
|
e2e fixes: tasks tool in profiles, bg timeout lift, stepIcon fix, queue semantics docs
...
E2E findings addressed:
- tasks tool was registered but not in any profile's tools.agent.native —
added to all six profiles (agent could not check/wait/cancel bg tasks)
- detached terminal/code_exec/ssh_exec runs without explicit timeout are
lifted to 300s: foreground defaults (20/30/60s) marked long commands
'completed' with partial output while the process still ran
- ToolCard.vue: define stepIcon(status) — template referenced it but the
function was missing (render crash on task_update step)
- message_queued reachability documented: WS read loop is sequential, the
queue path is only reachable from a second socket/headless recall
Eugene Sukhodolskiy
committed
12 days ago
|
agent parallelism: background tools, bg spawn_agent, parallel tool batches, message queue
...
- backgroundable tools (terminal/ssh_exec/peer/spawn_agent/code_exec) detach
via TaskManager with per-session/global/spawn caps, rate limit and TTL
- completion delivery both ways: task_update event (out-of-band) + pending
notes injected into the next turn; tasks tool (list/check/wait/cancel)
- parallel tool-call batches (profile-level gate, off by default) with
ToolStarted up-front, tagged event mux, call-order results, one save,
plus dangling tool_call repair on session load
- user message queue instead of busy error: message_queued frame,
back-to-back drain on the same socket, headless fallback on disconnect
- pin mcp<2 (v2 renames FastMCP with breaking API changes)
- docs: tasks.md (new), websocket/api/agent/config updates, tasks manual,
spawn_agent background param, persona contract section
Eugene Sukhodolskiy
committed
12 days ago
|