| 2026-10-10 |

synapse: a reaction continues its conversation, and the request becomes a record
...
Every Synapse event used to open its own service session, so two messages of
one Telegram chat landed in two unrelated chats. The dispatcher now returns a
`session_key` — derived from the payload according to the user's routing
document — and a later event carrying the same key joins the session that
opened it, within `reaction_session_ttl_minutes` of idle time (0 restores the
old always-fresh behaviour). A session that is still running is never joined:
`create_run` overwrites `state.run`, so a second run would orphan the first
one's subscribers — a busy thread gets a fresh session instead. When the
dispatcher names a different profile for the event, the session's profile is
switched through `set_profile`, never through `save()`, which does not rewrite
`name` or `profile_id`: the thread keeps its name and its history.
What the user reads in the transcript is now only the request. The event
record is a system message (`metadata.source == "synapse_event"`) drawn by a
new SynapseEventNotice.vue in the visual language of the recall badge, and the
framing plus the JSON envelope become a hidden user turn written once, on the
first event of a thread — not on every one. The instructions from Settings
reach the model through a new `current_reaction_instructions` ContextVar and
the `[Reaction session]` block of the system prompt, so they appear in no
transcript at all. The stored system record still has to reach the model, so
ContextBuilder's allowlist becomes `_LLM_SYSTEM_SOURCES = {task_note,
synapse_event}`.
The dispatcher gets a document of its own — how to route, rather than what to
do — stored beside the reaction instructions, versioned separately, and edited
in its own field in the settings panel. `synapse_instruction_versions` gains a
`doc` discriminator, so existing rows keep `'reaction'` and their history
verbatim. The new columns are added by `_MIGRATE`, and the index over `doc`
lives there too: `_DDL` runs first, so an index over a column the migration has
not added yet kills the whole batch on an existing database — which is what
the prod database is. A guard test keeps it that way.
Tests: 1776 passed, 1 skipped (backend), 242 passed (webclient, 28 files).
Eugene Sukhodolskiy
committed
13 hours ago
|

webclient: the small UI components in en, uk and ru
...
The sixth and last area of the interface translation. The components in
components/ui/ now build their text through t(): the login screen, the welcome
screen, the confirmation dialog, the offline banner, the image lightbox, the
selection toolbar, the card grid and the form, plus the sidebar's close button,
which is one of the strings that moved.
24 new slugs in all three dictionaries, and one moved:
- 24 in ui.* — the login and welcome screens (the heading, the tagline they
both show, the GNEXUS button), the confirmation dialog, the offline banner,
the lightbox (its heading, its close button, the failed-to-load message), the
selection toolbar's Reply and its tooltip, the card grid's three section
headings and its "+{n} more", and the form's submit button, its submitted
note and its ten validation messages;
- common.close, moved out of sidebar.*: the sidebar's close button and the
lightbox's say the same word, so it sits in the vocabulary area now.
Four details worth knowing:
- OfflineBanner's sentence was hard-coded Russian inside an otherwise English
file — for a reader in English or Ukrainian it was the one Russian sentence
in the app. It is a slug now, and the Russian value is that same sentence
verbatim, so nothing changes for the reader who could already read it.
- The form's validation messages interpolate: "{field} is required" carries the
field's own label, and the numeric limits carry {n}. The two length messages
are dictionary plurals now, so English says "Minimum length is 1 character"
where it used to say "1 characters", and Russian and Ukrainian get the three
forms a numeral needs — "1 символ", "2 символа", "5 символов".
- The card grid's "+4 more" is one slug with a count, not a sentence glued
together in the template.
- ui.tagline is written once and read by both the login and the welcome screen,
which show the same words.
The English values are verbatim the literals the code hard-coded, so the
English interface is unchanged and no test needed editing.
Checked live in Chromium against the built bundle on the dev backend, all three
languages: the login screen (reached with no account at all, its language
coming from navi.locale), the offline banner with the context put offline under
it, the selection toolbar over a real selection in a message, and the
confirmation dialog opened from a session row's Delete. Two further passes open
a session whose tool output carries a card grid and one that carries images, for
the card grid's section headings and the lightbox's heading and close button.
The form is the one component no saved session renders — nothing in the dev
database carries a form payload — so its messages were printed through t()
instead, including the plural ones: "Поле «Email» обязательно", "Минимальная
длина — 2 символа", "Мінімальна довжина — 5 символів", "Minimum length is 1
character".
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the chat area in en, uk and ru
...
The fifth area of the interface translation. The chat screen now builds its
text through t(): the header (the rename tooltip, the sidebar toggle, the
artifacts toggle, and the recall banner with its Cancel and Skip next
buttons), the list around it (the connection-lost banner and its Retry, and —
unreachable for now — the empty state), the input bar (the sending notice, the
paperclip, the placeholder, Send and Stop generation), the attachment strip
(its alt text and its remove button), the message list (Loading older
messages… and the scroll-to-bottom tooltip) and the queued-messages chip. Also
four toasts the WebSocket composable raises — background task started /
finished, MCP server connected / disconnected — and the compression notice's
own pending text.
26 new slugs in all three dictionaries:
- 20 in chat.* — the header and recall banner, the input bar, the attachment
strip, the queued chip and the empty state;
- 4 in toast.* — the two background-task and the two MCP toasts;
- 2 in common.* — task and toolsCount.
Four details worth knowing:
- The recall banner used to read the server's call_type enum out loud
("Scheduled once recall at …"). The word is now a slug
(chat.recallOnce / chat.recallRecurring / chat.recallImmediate), so it agrees
with the sentence around it — Russian and Ukrainian both change the
adjective's gender and the preposition, which a bare enum cannot.
- The banner's time was formatted with the browser's locale; it now formats
against locale.value, so the date follows the interface language rather than
the machine.
- The queued chip counted with `count === 1 ? '' : 's'`. It is now a dictionary
plural driven by {n}, which is what Russian and Ukrainian need — "1
сообщение", "2 сообщения", "5 сообщений".
- toolsCount moved from messages.* to common.*: the tool-count badge in the
assistant message and the MCP-connected toast both render it, so the same
word lives in the vocabulary area instead of twice under two areas.
docs/i18n.md records the move.
The English values are verbatim the literals the code hard-coded, so the
English interface is unchanged and no test needed editing.
Checked live in Chromium against the built bundle on the dev backend, all
three languages: the header title, the input placeholder, both input-row
buttons, the recall banner and its buttons, twelve message rows and the
attachment's alt and remove — read fresh for each language by picking it in
Settings first — with a scan of the page's text, titles, aria-labels, alt texts
and placeholders for a slug that leaked into the interface as the pass
condition.
Two states a saved session cannot be put into are covered by a throwaway mount
instead, and the file was removed again: the streaming input bar (Sending…,
Stop generation) and the queued chip, plus the recall words and the toast
titles, printed in Russian — "1 сообщение в очереди" / "2 сообщения" /
"5 сообщений".
The connection-lost banner never appears in this harness: with no socket token
the server refuses the socket and the client never exhausts its retries inside
the test's window, on the current build and on HEAD's alike. It is covered by
the existing tests/unit/composables/useWebSocket.test.js and by reading
ChatArea.vue; the same goes for chat.nothingOpen / chat.startConversation,
which App.vue's WelcomeScreen shadows before ChatArea can render its own empty
state — the branch is dead today and keeps its slugs for when it is not.
Eugene Sukhodolskiy
committed
1 day ago
|
| 2026-10-09 |

webclient: the message area in en, uk and ru
...
The fourth area of the interface translation. Everything a conversation is made
of now builds its text through t(): the thinking card and the status line it
replaces while the model reasons, the plan card, the tool cards (both the plain
one and its sub-agent steps), the published-content card, the UI-component card
and the compression notice, plus the two new chat bubbles' copy button and the
"Scheduled recall" mark on an injected user message.
25 new slugs in all three dictionaries:
- 23 in messages.* — the cards' own headers and labels (Thinking, Plan,
Arguments, Show all / Hide older tools, Useful / Not useful, Open file,
Content, Unknown UI component: {name}, Context summary, Compressing context…,
Context compressed), the running and elapsed labels (Running {tool}…,
Preparing next action…, running, starting) and the two counted ones that a
compact JSON preview renders, "{n} items" and "{n} keys";
- 2 in common.* — Copy and Copied!, which the assistant and the user message
both need.
Three details worth knowing:
- The English values are verbatim the literals the code hard-coded, with one
exception: "{n} tools" and "({before} → {after} messages)" are now dictionary
plurals, so English says "1 tool" and "1 message" where it used to say
"1 tools" and "1 messages". No test asserted either string.
- The braces around the two compact-JSON counts stay in the template
(`[${t('messages.jsonItems', { n })}]`), so a translator never has to keep a
bracket it cannot see the other half of.
- The message list's ContentCard shows the same published-file widget as the
artifacts drawer, so it reads the six artifacts.* slugs of that widget rather
than carrying six duplicates under messages.*; and its two tool-output labels
— Result and Live output — are the words the tool card already uses, so they
sit in common.* and the drawer's copies moved there too. docs/i18n.md records
both.
Checked live in Chromium against the built bundle on the dev backend: five
saved sessions chosen so that between them they render every component of the
area (thinking cards, plan cards, 55 tool cards, published-content cards,
UI-component cards and a compression notice), six loads per language, with
every collapsed block expanded first and a scan for a slug leaking into the
rendered text as the pass condition. The sub-agent step cannot be reached from
a saved session — the runtime does not persist tool steps — so SubagentStep was
rendered on its own in a throwaway mount instead: "выполняется", "Аргументы",
"Результат", "[3 элемента]".
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the sidebar in en, uk and ru
...
The third area of the interface translation. AppSidebar, SessionList and
SessionItem now build their text through t(): New Chat, the profile strip and
its filter buttons, the list label (Conversations / Service sessions, and
"30 results" while searching), the empty and loading states, the per-session
icons' tooltips (pinned, scheduled recall, service session), the row actions
(pin / unpin / cancel recall / skip next / delete and its confirm dialog), and
the footer identity (Local mode / Log out / Log in / Settings / Admin).
Also the shared time labels in composables/useTime.js — "just now" and
"5 min ago" — which the messages area renders too; they go to common.*
because both screens say the same words. The label is a cached string, not a
template that re-renders, so the composable now watches locale and rewrites
itself when the language changes. The absolute fallbacks format against
locale.value instead of the browser locale.
31 new slugs: 28 in sidebar.*, plus common.delete, common.minutesAgo (plural)
and confirm.deleteSession. English values are verbatim the old literals, so
the English interface is unchanged and no test needed editing.
Checked live in Chromium against the built bundle on the dev backend: the
sidebar was read in all three languages both in the plain list and with the
search field open, including every title / aria-label / placeholder / alt, with
a slug-leak scan as the pass condition.
One English string stays put on purpose: GnSearchField's aria-label
="Clear search" is baked into the UI kit component, not passed in, so it is
English on every gnexus client; docs/i18n.md now says so.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the artifacts drawer in en, uk and ru
...
The second area of the interface translation. Every user-visible string in
ArtifactsPanel.vue now goes through t(): the four tab chips, the empty states,
the action labels (open preview / open raw file / download / copy link), the
pull-to-refresh captions, the background-task badges and ages, the terminal
detail rows and the live-output header.
60 new slugs, in all three dictionaries:
- 39 in artifacts.*;
- 11 in the shared col.* vocabulary (Tool, Status, Started, Finished, PID,
Command, CWD, Uptime, Background, Reason, Sub-agent tokens) — the drawer is
where these table headers first appear;
- 10 in common.* — Download, Yes, No, N/A, the tool-output labels Result and
Live output, and the four the sidebar needs as well: the three
pull-to-refresh captions and "just now". These are the same words on both
screens, so they sit in the shared vocabulary area and the drawer keeps only
what is about the drawer.
Four details worth knowing:
- The English values are verbatim the literals the code hard-coded, so the
English interface is byte-identical and the existing tests stay untouched.
- Counted strings use dictionary plurals: '{n} links in session' is
one/few/many in ru and uk, one/other in en.
- Times (formatTime / formatDate) and ages now format against locale.value
instead of the browser locale, so a Russian interface shows 16:56 where
English shows 04:56 PM.
- Unit suffixes in formatUptime (s/m/h) are left alone: they are symbols
inside a number, not words.
Checked live in Chromium against the built bundle on the dev backend: the
drawer was driven through all four tabs in all three languages on two real
sessions — one with published STL files and a viewed-images group, one with
eight extracted links — with a slug-leak scan as the pass condition, and the
time/age formatting difference confirmed in the rendering.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the settings screen in en, uk and ru
...
Settings is the first area off hard-coded English: 122 strings for the shell and
its five tabs — App, Account, Notifications, Synapse and MCP — plus the 43 the
shared vocabulary contributes (common.*, col.*, confirm.*, toast.*).
The English dictionary carries the very literals the templates used, so the
interface a test reads is byte-for-byte what it was and the 56 assertions on
English UI text stay green; only a slug that is genuinely counted had to be
restructured into a plural. In ru and uk the counted strings get the three
Slavic forms driven by n. Product names are never translated: Navi, Synapse,
MCP, Ollama, Tool call, Subagent and Token read the same in all three
languages, and the settings screen keeps one settings.* area with
panel-prefixed slugs rather than a slug per panel, as the other gnexus clients
do.
Strings carrying markup are split around it instead of hiding tags in the
dictionary: the install steps keep their inline icons and <code>http://</code>
in the template, with the words of each sentence on either side. Four lists of
labels — the settings tabs, two table headers, and the reaction options — are
now computed rather than module-level constants, since a list built at import
keeps the language that was active then and ignores a switch. Dates in the
panels go through the interface language as well, not the browser's.
Checked in Chromium against the built bundle: all three languages on all five
tabs, with a scan for a slug leaking into the rendered text as the pass
condition, and the account left with no override afterwards.
Eugene Sukhodolskiy
committed
1 day ago
|

i18n: en/uk/ru, and the interface language of the account
...
The client hard-coded English in every template. It now goes through t() from
src/i18n/ — a locale ref, {param} interpolation, plural dictionaries driven by
n, and flat per-area dictionaries in src/i18n/messages/ carrying the same slugs
in all three languages. vue-i18n is deliberately not pulled in: the platform
handbook rules it out, and what it would buy — plural forms, a reactive locale,
an English fallback — is about eighty lines here. A slug nobody carries falls
back to English, so a screen already translated and one still speaking
hard-coded English can coexist while the work goes area by area.
The language comes from locale_effective in GET /auth/me: the choice made in
Settings (navi_users.locale_override, a new column) over the gnexus-auth
account locale (navi_users.locale, a straight mirror the user.profile_updated
webhook keeps writing — the override sits in its own column precisely so that
mirror cannot wipe it), over en. PATCH /auth/me writes the override, and null
means "follow the account" again. With NAVI_AUTH_ENABLED=false there is no
account to follow, so the choice is kept in localStorage under navi.locale.
The picker is the first block of Settings → Account. Each language is named in
its own language, which is why those three strings are identical in all three
dictionaries. Lists of labels are computed rather than module-level constants:
one built at import would keep the language that was active then and ignore a
switch.
Checked in Chromium against the built bundle: choosing Русский rewrites
<html lang> and the document title, survives a reload, and Auto brings the
account's own language back.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: an App section, and a link to any settings section
...
Settings gains an App section: how to install Navi in the browser — a real
Install button where the browser offers beforeinstallprompt, per-platform
steps where it cannot, since Safari and Firefox never expose that API and iOS
goes through the share sheet — and the Android APK, with its version, size and
checksum, behind a download the app hands to the system browser. The PWA
listener is armed from main.js because beforeinstallprompt fires once, during
startup, long before the panel is ever mounted.
App is now the first tab and every section is linkable: #settings/app,
#settings/mcp, and so on. Parsing moved into utils/hash.js, which the three
places in App.vue that route on the hash now share — before it, #settings/app
was read as the id of a session and opened an empty chat. Switching tabs
rewrites the address with replaceState, so Back leaves Settings instead of
walking back through the tabs, and a bare or unknown #settings normalises to
the first tab.
Checked in Chromium against both the dev server and the built bundle: each
hash opens its section, no /api/sessions/ request is ever made for one, a tab
click leaves history.length alone, and at 390 and 320px the strip keeps five
tabs with the active one scrolled into view and no sideways overflow.
Eugene Sukhodolskiy
committed
1 day ago
|

profiles: give the restricted profiles image_view
...
The 3D designer renders PNG previews with mcp__navi-3d__render_stl and then
had no way to look at them: image_view was withheld from all three
restricted profiles, and designer_3d's system prompt said so outright
("You cannot look at the render yourself"). It shipped geometry it had
never seen, with only lint/compile output behind it.
Add image_view to all three rather than to designer_3d alone. The three
sets are kept identical precisely because switch_profile lets a session
move between them, so a restricted session can reach image_view from any
of the three anyway; putting it in one profile would have made the
restriction look like a boundary it is not.
designer_3d now has to call it on every preview the renderer returned
(one angle is not enough), fix what it sees, and treat the visual check
as an addition to the mechanical ones rather than a replacement. The
prompt also asks for session paths only — image_view itself accepts any
image path or URL, so that part is a rule, not a check. Sub-agents keep
the tool out of their set: they gather facts, they never render.
docs/profiles.md loses the image_view omission row and gains the grant,
its raster-only limit and the accepted residual risk.
Eugene Sukhodolskiy
committed
1 day ago
|

profiles: restricted profiles for ordinary users, and a role gate that holds
...
`is_admin_only` was checked in one place out of nine and was not read from
config.json at all, so all seven profiles were reachable by every account with
role `user`. The flag now lives in config.json — the file is the baseline, a
`profile_overrides` row still wins on top of it — and one predicate,
`admin_only_blocked`, is the single place the rule is expressed. The nine
surfaces that list, switch to, spawn or resolve a profile all consult it:
`POST /sessions`, the WebSocket, switch_profile, list_profiles, the system
prompt's "Available profiles" block, spawn_agent and the Synapse reaction
runner. The prompt cache is now keyed by (profile, role), so a user's prompt
can never be served an admin's profile list.
The seven existing profiles (developer, discuss, dispatcher, modeler_3d,
navi_code, secretary, server_admin) are marked admin-only. Three new ones take
their place for ordinary users: assistant, designer_3d and coder. They share one
native tool set — ssh_exec, peer, reload_tools, create_mcp_server, test_mcp_tool,
image_view and gmail are withheld — and differ only in system prompt, model and
MCP groups. navi-web's raw `request` group, and the whole of gnexus-creds and
tgclient, are withheld too.
MCP per-user keys gain the missing half of the rule: a server that declares a
`user_key` slot is refused to anyone but an admin who has no personal key, and
is left out of their tool list entirely, instead of quietly falling back to the
owner's credential and appearing as a tool that cannot work. The refusal names
the server and points at Settings.
Also closes `GET /agents/prompts`, which served every profile's system prompt to
anyone, with no user dependency at all.
The accepted residual risk is written down in docs/profiles.md: the working
directory is a convention, not a sandbox.
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-10-08 |

tools: a non-admin may reach their own session directory
...
A non-admin was confined to user_data/<user_id>/ alone, which made four
tools reject the directory the user's own uploads land in: share_file
refused the file it was asked to send back, filesystem could not read it,
and terminal/code_exec could not be pointed at it. In a prod session the
agent worked around the refusal by copying the uploaded file into its
sandbox with code_exec — the very bypass the injected security policy
forbids — and the copy collided with a same-named file, so the user got a
link to Project_1.mp3 instead of their own upload, plus an 8.6 MB duplicate.
navi/tools/_internal/areas.py now names the two roots a user owns, and the
four call sites share it. Relative paths still resolve into the sandbox,
and each tool keeps its own way of refusing: terminal returns
sandbox_violation, code_exec silently falls back to the sandbox root. The
share_file refusal lists both roots instead of "outside user sandbox", so
the agent retries in the right one instead of routing around it.
The session directory belongs to the same user: its id arrives from the
runtime context, never from tool arguments, and only the session's owner
can open it.
Eugene Sukhodolskiy
committed
2 days ago
|

mcp: credentials leave the tree, and sub-agents lose the task board
...
Five tracked configs (gnexus-creds, gntodo, hard-panel, synapse, tgclient)
carried live bearer tokens — in git, in every clone, and handed to the admin
client by GET /admin/mcp/config. They now name a variable, and the value lives
in the service .env (chmod 600, untracked).
Substitution happens once, at transport open, next to the project-relative
path resolution: that is the single place a transport is built, so every
stored and serialised config stays placeholder-only and save_mcp_servers —
reached by create_mcp_server and PUT /admin/mcp/config — cannot write a secret
back into a tracked file. A variable that is not set refuses the connection
instead of dropping the header, which would fail open: a server may answer an
unauthenticated request as an anonymous user rather than returning 401.
Sub-agents also drop gntodo, gnexus-creds, tgclient and synapse from their
scopes, and the runner injects a scope anchor unconditionally — a sub-agent
asked to echo one word had picked a project off the task board and researched
it for forty iterations.
Docs: docs/mcp.md#secrets, a migration recipe in deploy/UPDATE.md, the variable
names in .env.example and deploy/env.template, and a regression guard over the
tracked configs.
The old values stay in git history; rotating all five at their sources is the
remediation, not a follow-up.
Eugene Sukhodolskiy
committed
2 days ago
|

subagents, secrets: an honest result contract, and two leaks closed
...
A pass over the sub-agent system, plus the two findings from the previous one.
Leaks:
- `redact_args` / `is_sensitive_key` (navi/tools/_internal/redact.py) mask
credential-shaped keys at every log site that prints tool arguments — by key
name, for all tools, so a future MCP tool is covered without declaring
anything. ssh_exec passwords were sitting in the prod journal in plaintext.
- `mcp_servers.d/*.json` are tracked in git and carried the scaffolding
machine's absolute paths, so navi-3d/navi-web silently vanished from the
toolset off that machine. They now hold project-relative paths, resolved in
`McpClient.open_transport` — not at load time, because create_mcp_server
round-trips every config file through save_mcp_servers.
Sub-agent result contract:
- `run_ephemeral` returns a `SubAgentOutcome` (text + status: ok, timeout,
max_iterations, thinking_stall, user_stop, context_overflow) instead of
`(text, bool)`. Every exit routes through `SubAgentRunner._finish`, so token
accounting is never dropped and a partial run always carries a progress
report — the stop-after-stream path used to return a bare string.
- `spawn_agent` renders its header from the status. Before this, every short
run reached the parent as "hit iteration limit", and the parent repeats that
diagnosis to the user — a timeout was reported as an iteration limit.
- `ContextTooLargeError` no longer kills a sub-agent: the parent compresses its
own context, a sub-agent has no session to compress into, so it stops with
`context_overflow` and hands back what it did rather than surfacing as
"Sub-agent failed: …" with the whole run discarded.
- The result is capped at 8000 chars (head and tail kept) so a long run cannot
flood the parent's context, and tool arguments in the progress report go
through `redact_args`.
- The navi-3d calling convention left the runner for the server's own tool
schemas — a generic runner should not know one MCP server's path shapes.
- Validation moved inside spawn_agent's try: a bad argument is a clean
ToolResult the model can read, not an exception in the parent's loop.
Tests: 1636 passed, 1 skipped. New: SubAgentOutcome paths and headers, result
truncation, context overflow, clean argument failures.
Eugene Sukhodolskiy
committed
2 days ago
|

profiles: tool_developer folds into developer
...
The profile was a duplicate on every axis we could measure. 22 of its 26
native tools were already developer's; the four it alone held —
reload_tools, create_mcp_server, test_mcp_tool, mcp_status — are 2.9 KB of
schema. Its model chain was the same seven models. Of its 14 KB prompt,
about 8 KB was copied verbatim from developer's (the whole `## Orchestration
model` block and everything from `## Editing policy` down), and most of the
remainder restated manuals/create_mcp_server.md, which already carried the
same ten-step workflow in more detail. It was not a specialisation, it was a
snapshot: `git log -S '"reload_tools"'` shows the tool lived in developer
until 61fa370 rewrote that profile around MCP and cut it off.
What kept it alive was a premise that no longer holds — that Navi's own
capabilities would be written as in-repo tools. They are MCP servers now,
and an MCP server is not a file in this repository with a life of its own:
it is an isolated process registered from mcp_servers.d/. So there is no
reason left for a profile whose only distinct feature is a toolset a general
developer profile can hold, and every reason to stop maintaining a second
prompt that drifts against the first.
- developer: + reload_tools, create_mcp_server, test_mcp_tool, mcp_status
(24 → 28 native). Its sub-agent gains tool_manual, which is what it
actually needed to reach the manual while writing a server — the previous
tool_developer sub-agent had it, developer's did not.
- server_admin: + reload_tools only (24 → 25). Adding a third-party MCP
server is something this profile does as often as developer does.
Deliberately not to its sub-agent: reconnecting the MCP manager is a
process-wide operation belonging to the main agent.
- The prompt and the manual took on what the deleted profile knew and
create_mcp_server.md did not: reload_tools before the first test_mcp_tool
(a freshly registered server is not connected, so the test fails and the
iteration is wasted), the smoke test read by exit code — 124 means timeout
killed a server still running, 0 means it exited on its own, usually a
main() without parentheses — absolute command/cwd, mcp_status as discovery
only, and the steps that stay inline instead of going to a sub-agent.
mcp_status and test_mcp_tool were built without an MCP manager, and their
fallback did `from navi.api.deps import _mcp_manager` — a name that does not
exist, so a live call raised ImportError rather than the intended "MCP
manager not available". The tools always passed a manager in tests, which is
why nothing caught it. Both now receive the manager at construction and fall
back to the live one lazily.
Sessions and profile_overrides are reassigned before the restart: agent.py
resolves the session's profile without a guard, so a deleted profile turns
every session that referenced it into an uncaught ProfileNotFound. Nothing
in the test suite pins the profile inventory, and profiles are read once at
import time — reload_tools does not re-read them — so this ships as a
restart, and the restart is also what makes it take effect.
Eugene Sukhodolskiy
committed
2 days ago
|

tools, mcp: a description is a hook, not a manual
...
With 24 tools in a profile the schemas are the largest fixed block of every
request, and they are paid for whether or not the tool is called. Eleven of
them carried a manual's worth of prose in their `description` — examples,
error tables, "common mistakes", the mechanics of the file areas. Reading
them against the manuals first, the manuals already held all of it, and
richer: this was duplication sitting in the one place the model cannot avoid
reading, not knowledge that had nowhere else to live. So the descriptions
became selection hooks — what the tool is for, when to pick it over its
neighbour, and the one rule that cannot wait — and the detail moved to
tool_manual("<tool>"), which costs nothing until the tool is actually in
play.
filesystem, spawn_agent, schedule_recall, manage_recall, scratchpad, plan,
share_file, content_publish, reflect, todo and peer. Descriptions drop from
18.4 KB to 11.5 KB, those tools' schemas from 25.5 KB to 16.8 KB, and a
profile's whole native toolset from 38.3 KB to 29.6 KB (server_admin,
developer; secretary 33.7 → 25.0, navi_code 32.9 → 25.9, tool_developer
37.9 → 29.2, modeler_3d 33.4 → 24.8). Stated as a budget rather than a byte
count: roughly 4k tokens of the model's window per request, back.
Two things were deliberately not shortened. The filesystem edit ladder
(`edit` → `edit_lines` → `smart_edit`, last resort, costs an LLM call) and
todo's mandatory `validation` on `done` both exist to prevent an extra
round trip; a shorter description there would trade tokens for turns. What
todo lost is the prose around the rule, not the rule. `enabled.json` was
left alone too: weather, gmail and get_current_datetime are opt-in and the
DB shows 12, 16 and 40 calls, so they are opt-in tools that are actually
used.
peer had no manual at all, which is how this surfaced: its description was
the only documentation the tool had. manuals/peer.md is now written from
the tool and its /peer route — the ask/status/list actions, the fact that
the hive is only a phone book and the question travels peer-to-peer, the
one-concurrent-answer semaphore, the two independent recursion guards, and
the six error codes.
The MCP server instructions leave the system prompt the same way. They were
8 KB of always-present prose per request, most of it workflow detail the
model only needs when it is about to use that server. Each server now
contributes one line — a new optional `summary` field in mcp_servers.d/*.json,
defaulting to the first sentence of its `instructions` — and
tool_manual("<server>") returns the full text plus the tool names the server
declares. server_admin's MCP block goes from 8.3 KB to 1.3 KB, secretary's
from 8.1 to 1.2. tool_manual learned to answer for a server (a tool name
still wins over a server of the same name, and dash/underscore spelling is
tolerated), and the index names the servers alongside the tools.
One judgement call worth recording: gnexus-book's and navi-web's
instructions end in an absolute "NEVER bypass these tools with filesystem,
terminal or code_exec" rule. Moving that on demand would have been a
behavioural regression dressed as a token saving, so those two `summary`
fields carry the prohibition verbatim alongside the hook, and it stays
always visible.
docs/tools.md gains a section on the split (a description is paid for every
request, a manual only when the tool is used), which is where the next
person will look before trimming a description back into a manual.
Full suite green (1577 passed, 1 skipped). No frontend change and no new
dependency, so the deploy is a pull and a restart.
Eugene Sukhodolskiy
committed
2 days ago
|

tool_manual: a manual the agent cannot reach is not a manual
...
tool_manual asked the filesystem for the exact string the model sent, then
fell back to registry.get(). The executor already tolerated the ways models
mangle MCP names; the tool that documents them did not, so the agent calling
mcp__navi-3d__compile_scad got a schema dump while manuals/compile_scad.md sat
there unread. Resolving a name now happens in one place — resolve_tool moved
out of the executor into navi/core/tool_utils.py — so a name that works when
the tool is *called* also works when it is *asked about*. A miss suggests the
closest names instead of "not found", and a tool that exists but is not enabled
for this profile says exactly that: documenting it is still the useful answer,
but the agent must not go on to call it. params["tool_name"] was also a crash
waiting to happen — KeyError with the key omitted, AttributeError on null or a
number — and now anything that is not a usable string means "no name given",
which returns the index of manuals grouped by source.
The generated manual now renders the whole schema. An array of objects showed
as "(array, optional)" with its item shape invisible, so a model that could not
see the fields guessed them; objects, array items, oneOf/anyOf branches, enums
and defaults are all spelled out now, depth-capped. The header names the source
(native / mcp__<server>__) and says the text is a parameter contract, not a
curated manual — otherwise a thin schema reads as "this tool is simple".
Then the manuals themselves. A manual is filed under the tool's own name, and
five of the sixteen were not: write_tool.md and write_mcp_server.md documented
tools that do not exist anywhere (the tools are reload_tools and
create_mcp_server), model_3d.md documented compile_scad, render_3d.md documented
render_stl, and write_context_provider.md documented no tool at all — it is a
guide, and now lives in manuals/guides/. The prompts pointed at the same wrong
names (tool_developer's system prompt, persona_navi_code, docs/context_providers),
so fixing the files without the callers would have moved the breakage rather
than removed it. On HEAD the probe was blunt: tool_manual("mcp__navi-3d__compile_scad")
returned None and the agent got the schema; now it returns the manual.
Nothing was enforcing any of this, which is how it drifted. tests/unit/tools/
test_manual_drift.py reads the real repository tree: every manuals/*.md must be
named after a tool that exists (built-ins and tools/*.py by their name
assignments, MCP tools by @mcp.tool(name=…) in the server sources and by the
mcp_servers.d groups, whose file name is the server name), every manual must be
reachable by both its bare name and its full MCP spelling, a guide must never
shadow a tool's name, and every manuals/<x>.md or tool_manual("x") cited in a
doc or a prompt must resolve. The next rename fails the suite instead of
quietly wasting a manual.
Finally, manuals for the built-ins the agent uses most — ssh_exec, todo, memory,
plan, filesystem, notify, code_exec — written from the current schemas and the
code paths that produce the errors, so the gotchas are the real ones (todo's
`done` validation is enforced; memory's `list` returns categories, not facts;
filesystem's `delete` removes a tree with no prompt). reload_tools.md absorbs
the file format from the deleted write_tool.md, which is where a self-extension
recipe belongs now.
1556 passed, 1 skipped. ruff delta zero.
Eugene Sukhodolskiy
committed
2 days ago
|

agent: a profile switch now survives the turn that made it
...
switch_profile repoints the session with one narrow UPDATE, and the run's own
save() — which lands right after every tool batch, the switch's own batch
included — wrote its stale in-memory profile back over it. The post-turn
reload then read the row it had just clobbered, saw no change, and left the
turn bound to the old profile's tools: the agent called exactly the tools
switch_profile had just advertised ("Newly available: ssh_exec, terminal, …")
and got "tool not found" for every one of them. Prod shows both halves — the
session that reported two switches kept the profile it started with, and
agent.profile_reloaded appears zero times in the whole boot.
The reload now compares against the profile the run is bound to
(bound_profile_id), which set_profile cannot move, instead of against
session.profile_id, which the store either rewrites or mutates to the new
value — making the old guard false in both branches by construction. And
save() leaves the profile_id column alone on conflict: set_profile owns it
after creation, the same way name/created_at are owned elsewhere.
A tool the live tool_map does not hold is now logged (tool.not_found, with
profile and live tool count). It previously left no trace at all in the
journal, which is why this took a database dig to find.
Three tests, each verified to fail without its fix: the next LLM call is
offered the new profile's tools, save() does not rewrite profile_id, and a
missing tool is logged.
Eugene Sukhodolskiy
committed
2 days ago
|

tools: list_tools defaults to the profile the run executes as
...
Asking "what tools do I have?" required the agent to know its own profile id
— and to be right about it after a switch_profile had moved the session. The
run knows: run_stream now publishes the active profile id (re-published when
switch_profile reloads the run) alongside the model, SubAgentRunner publishes
its own rather than inheriting the parent's, and ToolContext carries it into
every execute() call.
list_tools reads that when profile_id is omitted, so the bare call answers the
question, and it says which profile it assumed ("current profile") so the
answer cannot be mistaken for another profile's.
scope='agent'|'subagent' selects which of the profile's two toolsets to list —
the set spawn_agent would hand a sub-agent running that profile, which is a
different list and until now had no way to be inspected.
Item D of the list_tools plan, kept separate from the accuracy/compactness work.
Eugene Sukhodolskiy
committed
2 days ago
|

tools: list_tools tells the truth and costs a tenth of the context
...
The tool read the profile config, not the registry: a tool the config declares
but nothing registered — mcp__navi_ui__render_component, whose server never
connects — was advertised as callable, and the agent walked into "tool not
found" (28 such warnings on prod). It now resolves every name against the live
registry and reports the rest separately as "Not registered", so a phantom
entry reads as a broken server, not as a tool to try.
It also returned every description it could find. For server_admin that was
125 tools and 35 316 bytes per call — ~9k tokens to answer "do I have anything
for ssh". Names are now grouped by source (native, then one section per MCP
server) with descriptions behind verbose, and query filters by substring over
names and descriptions: the ssh question costs 88 bytes instead of 35 KB, and
a plain listing drops to 3.7 KB (9.5x smaller).
Tests cover the phantom-tool case, query matching by name and by description,
and the size property that makes names-only the default.
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-10-07 |
config: capture the live server configuration
...
Profiles gain the gntodo / gnexus-book / gnexus-creds scopes for agent and
subagent runs, refreshed model lists, and write access where the live setup
has it. MCP server configs pick up their user_key slots, gnexus-book learns
delete_pending_change, and the hard-panel / synapse / tgclient servers join
the tree.
root
committed
3 days ago
|
webclient: serve the PWA artwork past the images cache
...
/images/* is served cache-first from IMAGES_CACHE, which activate deliberately
keeps across builds. That is right for content whose URL is unique and wrong
for the app's own artwork: /images/icon-*, /images/logo-icon* and
/images/apple-splash/* keep their URLs while their contents change with the
logo, so a browser that had once loaded an icon would keep showing the old one
even after a deploy — and the apple-touch-icon is linked from index.html, so it
does travel through the page and the service worker.
Those paths now go network-first with a cached fallback (offline still works);
every other /images/ request is untouched.
Tests: frontend 148 passed; backend 1367 passed, 1 skipped.
Eugene Sukhodolskiy
committed
3 days ago
|

webclient: draw the PWA icons at full size again
...
The maskable icons and the apple-touch-icon carried the mark at 21% of the
canvas — the artwork scaled down and pasted in the centre — so on a home
screen the logo read as a fragment of itself. The mark takes 73.4% of the
canvas in logo.svg and in the launcher tile of the Android app icon; that is
the proportion all five files use now.
scripts/gen_pwa_icons.py redraws the mark from logo.svg's geometry with
Pillow (already a project dependency, no SVG rasterizer needed) and writes the
whole set, so the scale lives in one constant; --check measures what is on
disk and reports SUSPECT if it drifts again. Two runs produce identical bytes.
The regenerated icon-192/512 are geometrically identical to the rsvg-rendered
ones they replace: ink bbox 376x376 with 68px insets at 50% coverage, and the
same ink mass across the stroke. Only the antialiasing bytes differ.
dist/ rebuilt (it carries a copy of public/ and is served by the backend).
Tests: frontend 148 passed; backend 1367 passed, 1 skipped.
Eugene Sukhodolskiy
committed
3 days ago
|

MCP settings tab: list every connected server, slot only where declared
...
The tab was empty on every install: it listed only servers whose config
declares a `user_key` slot, and no config declared one — which read as
"no MCP servers connected" even though five are wired to profiles.
- GET /mcp-keys now returns every server referenced by at least one
profile, keyed ones first, with `accepts_user_key`, the slot location
(null when there is none) and the profile ids that connect it. The
per-user key store is skipped entirely when nothing has a slot.
- gnexus-creds declares `user_key: {header: Authorization, prefix:
"Bearer "}` — it is the one server carrying a shared credential, so its
personal-key field is now real: users with a key run under their own,
users without one fall back to the shared default.
- The panel lists all servers (transport + profiles), dims the keyless
rows, and shows a key input only for slotted ones, spelling out the
shared-key fallback.
docs/api.md and docs/mcp.md updated; backend 1367 passed, webclient 148.
Eugene Sukhodolskiy
committed
3 days ago
|
docs: fix stale VAPID env names (NAVI_VAPID_* -> NAVI_PUSH_VAPID_*, per config.py)
Eugene Sukhodolskiy
committed
3 days ago
|
mcp keys: tool-management tools aligned with BYOK (test_mcp_tool runs on user key, mcp_status marks key slots)
Eugene Sukhodolskiy
committed
3 days ago
|
docs: mcp.md (per-user MCP keys BYOK), /mcp-keys endpoints, user_key config
Eugene Sukhodolskiy
committed
3 days ago
|
| 2026-10-06 |
docs: synapse integration guide, special sessions filter, new endpoints
...
- docs/synapse.md: full guide — gateway, reaction runner pipeline,
notify tool, per-user settings, service sessions in the UI
- index/tools/config/api: entry point + built-in tools + env vars +
GET /sessions special param + synapse REST/webhook endpoints
Eugene Sukhodolskiy
committed
4 days ago
|
| 2026-10-05 |

webclient: Backgrounds tab in artifacts panel + task toasts
...
- new Backgrounds tab: task list (running first, badges, per-tool icons)
with a detail view (status, timestamps, sub-agent tokens, result/progress)
- GET /sessions/{id}/tasks snapshot endpoint: task_update events are not
replayed on reconnect, so the client fetches the task list on session
load/reload; live task_update entries are merged on top (chat.fetchTasks)
- terminal task_update no longer removes the entry — it marks it finished
so the tab shows recent completions; _terminalTaskIds still blocks
resurrection by a late running update
- toasts for background task start / finish (info / success / error) in
the WS dispatch; silent for other sessions and unknown terminal ids
- persistent background-tasks chip removed (replaced by the tab); only
the queued-messages chip stays
- fix invisible status text in terminal detail rows: filled .status-*
backgrounds now scope to .terminal-status-badge only
Eugene Sukhodolskiy
committed
5 days ago
|

background tasks: review fixes B1-B11 batch
...
- stop mid-batch cancels in-flight tools (B1) and keeps real results
of already-finished ones, mixing them with synthetic stopped notes
in call order (B2)
- queued messages: headless drain publishes session_sync (B4), the
run's teardown broadcasts session_sync to other sockets but not the
owner socket (prevents double reload) (B5)
- task_update notifications are chained per task so late running
updates can't overtake the terminal one (B7; client drops late
re-flicker of a terminal task chip) (B8)
- client: ui_component results reference the owning message via
card.parentMsg instead of a stale msg reference (B6)
- tasks cancel reports the real outcome after a bounded wait instead
of an optimistic 'cancelled' (B9)
- session delete cancels its running background jobs and drops
pending result notes (B10)
- subagent tool loop routes background:true calls through
ToolExecutor._maybe_background like the main loop (B11)
- tests for all of the above; docs/tasks.md stop/cancel semantics
Eugene Sukhodolskiy
committed
5 days ago
|