| 2026-10-10 |

webclient: the small UI components in en, uk and ru
...
The sixth and last area of the interface translation. The components in
components/ui/ now build their text through t(): the login screen, the welcome
screen, the confirmation dialog, the offline banner, the image lightbox, the
selection toolbar, the card grid and the form, plus the sidebar's close button,
which is one of the strings that moved.
24 new slugs in all three dictionaries, and one moved:
- 24 in ui.* — the login and welcome screens (the heading, the tagline they
both show, the GNEXUS button), the confirmation dialog, the offline banner,
the lightbox (its heading, its close button, the failed-to-load message), the
selection toolbar's Reply and its tooltip, the card grid's three section
headings and its "+{n} more", and the form's submit button, its submitted
note and its ten validation messages;
- common.close, moved out of sidebar.*: the sidebar's close button and the
lightbox's say the same word, so it sits in the vocabulary area now.
Four details worth knowing:
- OfflineBanner's sentence was hard-coded Russian inside an otherwise English
file — for a reader in English or Ukrainian it was the one Russian sentence
in the app. It is a slug now, and the Russian value is that same sentence
verbatim, so nothing changes for the reader who could already read it.
- The form's validation messages interpolate: "{field} is required" carries the
field's own label, and the numeric limits carry {n}. The two length messages
are dictionary plurals now, so English says "Minimum length is 1 character"
where it used to say "1 characters", and Russian and Ukrainian get the three
forms a numeral needs — "1 символ", "2 символа", "5 символов".
- The card grid's "+4 more" is one slug with a count, not a sentence glued
together in the template.
- ui.tagline is written once and read by both the login and the welcome screen,
which show the same words.
The English values are verbatim the literals the code hard-coded, so the
English interface is unchanged and no test needed editing.
Checked live in Chromium against the built bundle on the dev backend, all three
languages: the login screen (reached with no account at all, its language
coming from navi.locale), the offline banner with the context put offline under
it, the selection toolbar over a real selection in a message, and the
confirmation dialog opened from a session row's Delete. Two further passes open
a session whose tool output carries a card grid and one that carries images, for
the card grid's section headings and the lightbox's heading and close button.
The form is the one component no saved session renders — nothing in the dev
database carries a form payload — so its messages were printed through t()
instead, including the plural ones: "Поле «Email» обязательно", "Минимальная
длина — 2 символа", "Мінімальна довжина — 5 символів", "Minimum length is 1
character".
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the chat area in en, uk and ru
...
The fifth area of the interface translation. The chat screen now builds its
text through t(): the header (the rename tooltip, the sidebar toggle, the
artifacts toggle, and the recall banner with its Cancel and Skip next
buttons), the list around it (the connection-lost banner and its Retry, and —
unreachable for now — the empty state), the input bar (the sending notice, the
paperclip, the placeholder, Send and Stop generation), the attachment strip
(its alt text and its remove button), the message list (Loading older
messages… and the scroll-to-bottom tooltip) and the queued-messages chip. Also
four toasts the WebSocket composable raises — background task started /
finished, MCP server connected / disconnected — and the compression notice's
own pending text.
26 new slugs in all three dictionaries:
- 20 in chat.* — the header and recall banner, the input bar, the attachment
strip, the queued chip and the empty state;
- 4 in toast.* — the two background-task and the two MCP toasts;
- 2 in common.* — task and toolsCount.
Four details worth knowing:
- The recall banner used to read the server's call_type enum out loud
("Scheduled once recall at …"). The word is now a slug
(chat.recallOnce / chat.recallRecurring / chat.recallImmediate), so it agrees
with the sentence around it — Russian and Ukrainian both change the
adjective's gender and the preposition, which a bare enum cannot.
- The banner's time was formatted with the browser's locale; it now formats
against locale.value, so the date follows the interface language rather than
the machine.
- The queued chip counted with `count === 1 ? '' : 's'`. It is now a dictionary
plural driven by {n}, which is what Russian and Ukrainian need — "1
сообщение", "2 сообщения", "5 сообщений".
- toolsCount moved from messages.* to common.*: the tool-count badge in the
assistant message and the MCP-connected toast both render it, so the same
word lives in the vocabulary area instead of twice under two areas.
docs/i18n.md records the move.
The English values are verbatim the literals the code hard-coded, so the
English interface is unchanged and no test needed editing.
Checked live in Chromium against the built bundle on the dev backend, all
three languages: the header title, the input placeholder, both input-row
buttons, the recall banner and its buttons, twelve message rows and the
attachment's alt and remove — read fresh for each language by picking it in
Settings first — with a scan of the page's text, titles, aria-labels, alt texts
and placeholders for a slug that leaked into the interface as the pass
condition.
Two states a saved session cannot be put into are covered by a throwaway mount
instead, and the file was removed again: the streaming input bar (Sending…,
Stop generation) and the queued chip, plus the recall words and the toast
titles, printed in Russian — "1 сообщение в очереди" / "2 сообщения" /
"5 сообщений".
The connection-lost banner never appears in this harness: with no socket token
the server refuses the socket and the client never exhausts its retries inside
the test's window, on the current build and on HEAD's alike. It is covered by
the existing tests/unit/composables/useWebSocket.test.js and by reading
ChatArea.vue; the same goes for chat.nothingOpen / chat.startConversation,
which App.vue's WelcomeScreen shadows before ChatArea can render its own empty
state — the branch is dead today and keeps its slugs for when it is not.
Eugene Sukhodolskiy
committed
1 day ago
|
| 2026-10-09 |

webclient: the message area in en, uk and ru
...
The fourth area of the interface translation. Everything a conversation is made
of now builds its text through t(): the thinking card and the status line it
replaces while the model reasons, the plan card, the tool cards (both the plain
one and its sub-agent steps), the published-content card, the UI-component card
and the compression notice, plus the two new chat bubbles' copy button and the
"Scheduled recall" mark on an injected user message.
25 new slugs in all three dictionaries:
- 23 in messages.* — the cards' own headers and labels (Thinking, Plan,
Arguments, Show all / Hide older tools, Useful / Not useful, Open file,
Content, Unknown UI component: {name}, Context summary, Compressing context…,
Context compressed), the running and elapsed labels (Running {tool}…,
Preparing next action…, running, starting) and the two counted ones that a
compact JSON preview renders, "{n} items" and "{n} keys";
- 2 in common.* — Copy and Copied!, which the assistant and the user message
both need.
Three details worth knowing:
- The English values are verbatim the literals the code hard-coded, with one
exception: "{n} tools" and "({before} → {after} messages)" are now dictionary
plurals, so English says "1 tool" and "1 message" where it used to say
"1 tools" and "1 messages". No test asserted either string.
- The braces around the two compact-JSON counts stay in the template
(`[${t('messages.jsonItems', { n })}]`), so a translator never has to keep a
bracket it cannot see the other half of.
- The message list's ContentCard shows the same published-file widget as the
artifacts drawer, so it reads the six artifacts.* slugs of that widget rather
than carrying six duplicates under messages.*; and its two tool-output labels
— Result and Live output — are the words the tool card already uses, so they
sit in common.* and the drawer's copies moved there too. docs/i18n.md records
both.
Checked live in Chromium against the built bundle on the dev backend: five
saved sessions chosen so that between them they render every component of the
area (thinking cards, plan cards, 55 tool cards, published-content cards,
UI-component cards and a compression notice), six loads per language, with
every collapsed block expanded first and a scan for a slug leaking into the
rendered text as the pass condition. The sub-agent step cannot be reached from
a saved session — the runtime does not persist tool steps — so SubagentStep was
rendered on its own in a throwaway mount instead: "выполняется", "Аргументы",
"Результат", "[3 элемента]".
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the sidebar in en, uk and ru
...
The third area of the interface translation. AppSidebar, SessionList and
SessionItem now build their text through t(): New Chat, the profile strip and
its filter buttons, the list label (Conversations / Service sessions, and
"30 results" while searching), the empty and loading states, the per-session
icons' tooltips (pinned, scheduled recall, service session), the row actions
(pin / unpin / cancel recall / skip next / delete and its confirm dialog), and
the footer identity (Local mode / Log out / Log in / Settings / Admin).
Also the shared time labels in composables/useTime.js — "just now" and
"5 min ago" — which the messages area renders too; they go to common.*
because both screens say the same words. The label is a cached string, not a
template that re-renders, so the composable now watches locale and rewrites
itself when the language changes. The absolute fallbacks format against
locale.value instead of the browser locale.
31 new slugs: 28 in sidebar.*, plus common.delete, common.minutesAgo (plural)
and confirm.deleteSession. English values are verbatim the old literals, so
the English interface is unchanged and no test needed editing.
Checked live in Chromium against the built bundle on the dev backend: the
sidebar was read in all three languages both in the plain list and with the
search field open, including every title / aria-label / placeholder / alt, with
a slug-leak scan as the pass condition.
One English string stays put on purpose: GnSearchField's aria-label
="Clear search" is baked into the UI kit component, not passed in, so it is
English on every gnexus client; docs/i18n.md now says so.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the artifacts drawer in en, uk and ru
...
The second area of the interface translation. Every user-visible string in
ArtifactsPanel.vue now goes through t(): the four tab chips, the empty states,
the action labels (open preview / open raw file / download / copy link), the
pull-to-refresh captions, the background-task badges and ages, the terminal
detail rows and the live-output header.
60 new slugs, in all three dictionaries:
- 39 in artifacts.*;
- 11 in the shared col.* vocabulary (Tool, Status, Started, Finished, PID,
Command, CWD, Uptime, Background, Reason, Sub-agent tokens) — the drawer is
where these table headers first appear;
- 10 in common.* — Download, Yes, No, N/A, the tool-output labels Result and
Live output, and the four the sidebar needs as well: the three
pull-to-refresh captions and "just now". These are the same words on both
screens, so they sit in the shared vocabulary area and the drawer keeps only
what is about the drawer.
Four details worth knowing:
- The English values are verbatim the literals the code hard-coded, so the
English interface is byte-identical and the existing tests stay untouched.
- Counted strings use dictionary plurals: '{n} links in session' is
one/few/many in ru and uk, one/other in en.
- Times (formatTime / formatDate) and ages now format against locale.value
instead of the browser locale, so a Russian interface shows 16:56 where
English shows 04:56 PM.
- Unit suffixes in formatUptime (s/m/h) are left alone: they are symbols
inside a number, not words.
Checked live in Chromium against the built bundle on the dev backend: the
drawer was driven through all four tabs in all three languages on two real
sessions — one with published STL files and a viewed-images group, one with
eight extracted links — with a slug-leak scan as the pass condition, and the
time/age formatting difference confirmed in the rendering.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: the settings screen in en, uk and ru
...
Settings is the first area off hard-coded English: 122 strings for the shell and
its five tabs — App, Account, Notifications, Synapse and MCP — plus the 43 the
shared vocabulary contributes (common.*, col.*, confirm.*, toast.*).
The English dictionary carries the very literals the templates used, so the
interface a test reads is byte-for-byte what it was and the 56 assertions on
English UI text stay green; only a slug that is genuinely counted had to be
restructured into a plural. In ru and uk the counted strings get the three
Slavic forms driven by n. Product names are never translated: Navi, Synapse,
MCP, Ollama, Tool call, Subagent and Token read the same in all three
languages, and the settings screen keeps one settings.* area with
panel-prefixed slugs rather than a slug per panel, as the other gnexus clients
do.
Strings carrying markup are split around it instead of hiding tags in the
dictionary: the install steps keep their inline icons and <code>http://</code>
in the template, with the words of each sentence on either side. Four lists of
labels — the settings tabs, two table headers, and the reaction options — are
now computed rather than module-level constants, since a list built at import
keeps the language that was active then and ignores a switch. Dates in the
panels go through the interface language as well, not the browser's.
Checked in Chromium against the built bundle: all three languages on all five
tabs, with a scan for a slug leaking into the rendered text as the pass
condition, and the account left with no override afterwards.
Eugene Sukhodolskiy
committed
1 day ago
|

i18n: en/uk/ru, and the interface language of the account
...
The client hard-coded English in every template. It now goes through t() from
src/i18n/ — a locale ref, {param} interpolation, plural dictionaries driven by
n, and flat per-area dictionaries in src/i18n/messages/ carrying the same slugs
in all three languages. vue-i18n is deliberately not pulled in: the platform
handbook rules it out, and what it would buy — plural forms, a reactive locale,
an English fallback — is about eighty lines here. A slug nobody carries falls
back to English, so a screen already translated and one still speaking
hard-coded English can coexist while the work goes area by area.
The language comes from locale_effective in GET /auth/me: the choice made in
Settings (navi_users.locale_override, a new column) over the gnexus-auth
account locale (navi_users.locale, a straight mirror the user.profile_updated
webhook keeps writing — the override sits in its own column precisely so that
mirror cannot wipe it), over en. PATCH /auth/me writes the override, and null
means "follow the account" again. With NAVI_AUTH_ENABLED=false there is no
account to follow, so the choice is kept in localStorage under navi.locale.
The picker is the first block of Settings → Account. Each language is named in
its own language, which is why those three strings are identical in all three
dictionaries. Lists of labels are computed rather than module-level constants:
one built at import would keep the language that was active then and ignore a
switch.
Checked in Chromium against the built bundle: choosing Русский rewrites
<html lang> and the document title, survives a reload, and Auto brings the
account's own language back.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: an App section, and a link to any settings section
...
Settings gains an App section: how to install Navi in the browser — a real
Install button where the browser offers beforeinstallprompt, per-platform
steps where it cannot, since Safari and Firefox never expose that API and iOS
goes through the share sheet — and the Android APK, with its version, size and
checksum, behind a download the app hands to the system browser. The PWA
listener is armed from main.js because beforeinstallprompt fires once, during
startup, long before the panel is ever mounted.
App is now the first tab and every section is linkable: #settings/app,
#settings/mcp, and so on. Parsing moved into utils/hash.js, which the three
places in App.vue that route on the hash now share — before it, #settings/app
was read as the id of a session and opened an empty chat. Switching tabs
rewrites the address with replaceState, so Back leaves Settings instead of
walking back through the tabs, and a bare or unknown #settings normalises to
the first tab.
Checked in Chromium against both the dev server and the built bundle: each
hash opens its section, no /api/sessions/ request is ever made for one, a tab
click leaves history.length alone, and at 390 and 320px the strip keeps five
tabs with the active one scrolled into view and no sideways overflow.
Eugene Sukhodolskiy
committed
1 day ago
|
android-client: a release-signed APK, published through the server
...
Release builds had no signingConfig, so assembleRelease produced an APK that
Android refuses to install. Sign with a key kept out of the repo: the keystore
lives in ~/.navi-android/, and android-client/keystore.properties (gitignored,
600) points at it and carries the passwords. A clone without that file still
builds — it just gets an unsigned APK, which is the right thing to fail into
for someone who has no business signing a release.
tools/publish-apk.sh builds the release APK into webclient/public/download/
and writes a JSON sidecar (version, size, sha256, build time) for the settings
page to read. The server mounts that directory at /download, so the APK ships
with a plain git pull like everything else. dist/download is ignored: vite
copies public/ into dist/ on every build, and serving the copy out of public/
means one binary in git instead of two.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: round the PWA icons
...
The icon artwork is a full-bleed dark tile, so every surface that shows it
without a mask of its own (the desktop install, the taskbar) drew a square.
Give the alpha channel a rounded-rectangle mask at 22.37% of the side — the
iOS/Android icon radius — sampled 4x so the corner reads as a curve rather
than a staircase.
The maskable pair is a byte-for-byte copy of the plain pair, so it is rounded
the same way; the transparent corners fall well inside any mask a launcher
applies. The apple-touch-icon is left alone: iOS masks it itself.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: keep the artifacts tab icons on a phone
...
The artifacts panel has two squeeze ladders that did not know about each
other. At <=980px the mobile block hides the tab label, leaving icon +
count; a container query at <=360px separately hides the icon, because
the docked panel has no room for icon + label + count. The query
measures the panel, and on a phone the panel is the viewport — so on a
360px-wide screen (the most common Android width) both fired and every
tab became a bare counter: "Artifacts 0 Links 0 Files 0
Backgrounds 0". Between 361 and 980px the icons were fine, which is why
it only showed up on phones.
Wrap the icon-hiding step in @media (min-width: 981px), so it describes
the docked panel only, where the label is still attached and dropping
the icon still buys something. The panel is otherwise untouched.
Measured in Chromium against the built stylesheet and the component's
real scoped markup, old versus new side by side: at 320px and 360px the
icons go 0/4 -> 4/4 with no row overflow; 390, 430 and 900px are
unchanged; and every docked width tested (300, 340, 370, 450, 537px)
reports byte-identical values before and after. The pre-existing
overflow of the docked row between roughly 300 and 450px is untouched.
Eugene Sukhodolskiy
committed
1 day ago
|

profiles: give the restricted profiles image_view
...
The 3D designer renders PNG previews with mcp__navi-3d__render_stl and then
had no way to look at them: image_view was withheld from all three
restricted profiles, and designer_3d's system prompt said so outright
("You cannot look at the render yourself"). It shipped geometry it had
never seen, with only lint/compile output behind it.
Add image_view to all three rather than to designer_3d alone. The three
sets are kept identical precisely because switch_profile lets a session
move between them, so a restricted session can reach image_view from any
of the three anyway; putting it in one profile would have made the
restriction look like a boundary it is not.
designer_3d now has to call it on every preview the renderer returned
(one angle is not enough), fix what it sees, and treat the visual check
as an addition to the mechanical ones rather than a replacement. The
prompt also asks for session paths only — image_view itself accepts any
image path or URL, so that part is a rule, not a check. Sub-agents keep
the tool out of their set: they gather facts, they never render.
docs/profiles.md loses the image_view omission row and gains the grant,
its raster-only limit and the accepted residual risk.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: animate <details> cards in every browser, not just Chromium
...
Open/close of the tool cards jumped in Firefox and Safari: the previous
attempt leaned on interpolate-size/::details-content, and interpolate-size
shipped in Chromium alone, so the @supports block degraded to a hard
attribute flip everywhere else.
Drive the element's own height from JS instead. Nothing there is
engine-specific: open animates to scrollHeight, close to the summary
height, overflow is clipped for the duration and the inline styles come
off on transitionend with a timer as a backstop. The element keeps its
native semantics — keyboard, screen readers, find-in-page — because the
open attribute is still what changes; the toggle event still fires, so
ContentCard's :open/@toggle pairing is untouched. prefers-reduced-motion
settles in a single frame.
A newly inserted card grows its grid row from 0fr to 1fr, which
interpolates everywhere and resolves to whatever the card measures, so
the card below slides instead of being displaced.
Measured in Chromium 153 and Firefox 155 against the real component
markup and the built bundle: 33 -> 586px over 18 distinct heights in
both, and the same for insertion, the nested Arguments block and the
plan card. The scroll behaviour of MessageList was left alone.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: a stale session list no longer erases a session just created
...
Starting a session from the welcome screen sets selectedProfileId and
then awaits the create POST. The profile watcher fires on the microtask
that the await yields, so the list refetch is issued while the POST is
still in flight — it returns a page that predates the new session, and
fetchSessions replaced the whole array, so the placeholder createSession
had just inserted was gone. The session only came back on a manual
refresh. Reproduced in a store test.
sessions.js now remembers a session it created along with the id of the
last list request that had already started at that moment, and a page
that started after the creation is still believed. So the rescue is
bounded by one round trip — it cannot turn into a ghost row: a request
that left once the session existed reports it missing and it goes, and
deleteSession clears the mark outright. Search and special-list pages
are never padded this way, since they answer a different question.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: let the profile card description run to its full length
...
The three-line clamp I added with the card-height fix cut the text of
eight of the nine profiles — every description ended in an ellipsis, so
the grid told you less about each profile the more there was to say.
Measured in a browser against the built CSS: 8 of 9 clipped at 1440px,
9 of 9 at 390px. With the clamp gone, none are clipped at 1440, 1100,
900, 768, 600 or 390px.
`grid-auto-rows: 1fr` already keeps the cards in a row level, so the
clamp bought nothing it was not paying for: the row is as tall as the
longest description (230px on desktop, 205px on a phone), and the
welcome pane scrolls when the grid outgrows it.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: one icon per profile, and one height for every card
...
Six of the ten profiles rendered the same robot, because the map in
WelcomeScreen only knew four ids. The map moves to utils/profileIcons.js,
covers all ten — including the close pairs (`coder`/`developer`, the two
3D profiles) — and falls back to a hashed glyph from a pool rather than
one robot, so custom profiles differ from each other and stay stable
across renders. Covered by five unit tests.
Cards were ragged because each row took the height of its own content:
`grid-auto-rows: 1fr` makes every row as tall as the tallest card. A
fixed pixel height was the obvious alternative and is worse — "Server
Administrator" wraps on a phone. A three-line clamp on the desktop
description keeps one wordy profile from inflating every row.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: let the welcome pane scroll, and pin the mobile menu toggle
...
The profile grid grows with the profile list, and on a phone it is taller than
the viewport once there are more than a couple of profiles. Nothing could
scroll: `#app` and `.app-main` both hide overflow, and `.welcome-screen` had no
overflow of its own, so the cards were clipped rather than pushed down.
The pane now scrolls, and its vertical centering is `safe center` — with a plain
`center` the overflow goes past the flex start edge and cannot be scrolled back
to. The sidebar toggle becomes `position: fixed` so it does not scroll away with
the content it sits above.
Eugene Sukhodolskiy
committed
2 days ago
|
webclient: rebuild dist for the role-aware MCP keys panel
...
`webclient/dist` is tracked and served straight by the API server — there is no
node on the prod host — so the source change alone would have left prod showing
the old copy, telling an ordinary user that the shared key from the config is in
use for a server that is simply not connected for them.
Eugene Sukhodolskiy
committed
2 days ago
|

profiles: restricted profiles for ordinary users, and a role gate that holds
...
`is_admin_only` was checked in one place out of nine and was not read from
config.json at all, so all seven profiles were reachable by every account with
role `user`. The flag now lives in config.json — the file is the baseline, a
`profile_overrides` row still wins on top of it — and one predicate,
`admin_only_blocked`, is the single place the rule is expressed. The nine
surfaces that list, switch to, spawn or resolve a profile all consult it:
`POST /sessions`, the WebSocket, switch_profile, list_profiles, the system
prompt's "Available profiles" block, spawn_agent and the Synapse reaction
runner. The prompt cache is now keyed by (profile, role), so a user's prompt
can never be served an admin's profile list.
The seven existing profiles (developer, discuss, dispatcher, modeler_3d,
navi_code, secretary, server_admin) are marked admin-only. Three new ones take
their place for ordinary users: assistant, designer_3d and coder. They share one
native tool set — ssh_exec, peer, reload_tools, create_mcp_server, test_mcp_tool,
image_view and gmail are withheld — and differ only in system prompt, model and
MCP groups. navi-web's raw `request` group, and the whole of gnexus-creds and
tgclient, are withheld too.
MCP per-user keys gain the missing half of the rule: a server that declares a
`user_key` slot is refused to anyone but an admin who has no personal key, and
is left out of their tool list entirely, instead of quietly falling back to the
owner's credential and appearing as a tool that cannot work. The refusal names
the server and points at Settings.
Also closes `GET /agents/prompts`, which served every profile's system prompt to
anyone, with no user dependency at all.
The accepted residual risk is written down in docs/profiles.md: the working
directory is a convention, not a sandbox.
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-10-08 |

tools: a non-admin may reach their own session directory
...
A non-admin was confined to user_data/<user_id>/ alone, which made four
tools reject the directory the user's own uploads land in: share_file
refused the file it was asked to send back, filesystem could not read it,
and terminal/code_exec could not be pointed at it. In a prod session the
agent worked around the refusal by copying the uploaded file into its
sandbox with code_exec — the very bypass the injected security policy
forbids — and the copy collided with a same-named file, so the user got a
link to Project_1.mp3 instead of their own upload, plus an 8.6 MB duplicate.
navi/tools/_internal/areas.py now names the two roots a user owns, and the
four call sites share it. Relative paths still resolve into the sandbox,
and each tool keeps its own way of refusing: terminal returns
sandbox_violation, code_exec silently falls back to the sandbox root. The
share_file refusal lists both roots instead of "outside user sandbox", so
the agent retries in the right one instead of routing around it.
The session directory belongs to the same user: its id arrives from the
runtime context, never from tool arguments, and only the session's owner
can open it.
Eugene Sukhodolskiy
committed
2 days ago
|
persona: the name rule, and a template that actually loads the persona
...
persona.txt gains a third strict rule: the first name is Navi — Latin even
inside Russian text, «Нави» known but not spoken by default — and the
instance codename is a second name, Latin and capitalised, never
transliterated. The model had only ever seen the Latin name plus the
identity block's lowercase codename, which is where "Нэви" came from.
None of this reached a server: deploy/env.template never set
NAVI_PERSONA_FILE, so an install from the template ran with
settings.navi_persona empty and every rule in the file inert. The variable
joins the template, where .env.example already had it.
Eugene Sukhodolskiy
committed
2 days ago
|

mcp: credentials leave the tree, and sub-agents lose the task board
...
Five tracked configs (gnexus-creds, gntodo, hard-panel, synapse, tgclient)
carried live bearer tokens — in git, in every clone, and handed to the admin
client by GET /admin/mcp/config. They now name a variable, and the value lives
in the service .env (chmod 600, untracked).
Substitution happens once, at transport open, next to the project-relative
path resolution: that is the single place a transport is built, so every
stored and serialised config stays placeholder-only and save_mcp_servers —
reached by create_mcp_server and PUT /admin/mcp/config — cannot write a secret
back into a tracked file. A variable that is not set refuses the connection
instead of dropping the header, which would fail open: a server may answer an
unauthenticated request as an anonymous user rather than returning 401.
Sub-agents also drop gntodo, gnexus-creds, tgclient and synapse from their
scopes, and the runner injects a scope anchor unconditionally — a sub-agent
asked to echo one word had picked a project off the task board and researched
it for forty iterations.
Docs: docs/mcp.md#secrets, a migration recipe in deploy/UPDATE.md, the variable
names in .env.example and deploy/env.template, and a regression guard over the
tracked configs.
The old values stay in git history; rotating all five at their sources is the
remediation, not a follow-up.
Eugene Sukhodolskiy
committed
2 days ago
|

subagents, secrets: an honest result contract, and two leaks closed
...
A pass over the sub-agent system, plus the two findings from the previous one.
Leaks:
- `redact_args` / `is_sensitive_key` (navi/tools/_internal/redact.py) mask
credential-shaped keys at every log site that prints tool arguments — by key
name, for all tools, so a future MCP tool is covered without declaring
anything. ssh_exec passwords were sitting in the prod journal in plaintext.
- `mcp_servers.d/*.json` are tracked in git and carried the scaffolding
machine's absolute paths, so navi-3d/navi-web silently vanished from the
toolset off that machine. They now hold project-relative paths, resolved in
`McpClient.open_transport` — not at load time, because create_mcp_server
round-trips every config file through save_mcp_servers.
Sub-agent result contract:
- `run_ephemeral` returns a `SubAgentOutcome` (text + status: ok, timeout,
max_iterations, thinking_stall, user_stop, context_overflow) instead of
`(text, bool)`. Every exit routes through `SubAgentRunner._finish`, so token
accounting is never dropped and a partial run always carries a progress
report — the stop-after-stream path used to return a bare string.
- `spawn_agent` renders its header from the status. Before this, every short
run reached the parent as "hit iteration limit", and the parent repeats that
diagnosis to the user — a timeout was reported as an iteration limit.
- `ContextTooLargeError` no longer kills a sub-agent: the parent compresses its
own context, a sub-agent has no session to compress into, so it stops with
`context_overflow` and hands back what it did rather than surfacing as
"Sub-agent failed: …" with the whole run discarded.
- The result is capped at 8000 chars (head and tail kept) so a long run cannot
flood the parent's context, and tool arguments in the progress report go
through `redact_args`.
- The navi-3d calling convention left the runner for the server's own tool
schemas — a generic runner should not know one MCP server's path shapes.
- Validation moved inside spawn_agent's try: a bad argument is a clean
ToolResult the model can read, not an exception in the parent's loop.
Tests: 1636 passed, 1 skipped. New: SubAgentOutcome paths and headers, result
truncation, context overflow, clean argument failures.
Eugene Sukhodolskiy
committed
2 days ago
|

profiles: tool_developer folds into developer
...
The profile was a duplicate on every axis we could measure. 22 of its 26
native tools were already developer's; the four it alone held —
reload_tools, create_mcp_server, test_mcp_tool, mcp_status — are 2.9 KB of
schema. Its model chain was the same seven models. Of its 14 KB prompt,
about 8 KB was copied verbatim from developer's (the whole `## Orchestration
model` block and everything from `## Editing policy` down), and most of the
remainder restated manuals/create_mcp_server.md, which already carried the
same ten-step workflow in more detail. It was not a specialisation, it was a
snapshot: `git log -S '"reload_tools"'` shows the tool lived in developer
until 61fa370 rewrote that profile around MCP and cut it off.
What kept it alive was a premise that no longer holds — that Navi's own
capabilities would be written as in-repo tools. They are MCP servers now,
and an MCP server is not a file in this repository with a life of its own:
it is an isolated process registered from mcp_servers.d/. So there is no
reason left for a profile whose only distinct feature is a toolset a general
developer profile can hold, and every reason to stop maintaining a second
prompt that drifts against the first.
- developer: + reload_tools, create_mcp_server, test_mcp_tool, mcp_status
(24 → 28 native). Its sub-agent gains tool_manual, which is what it
actually needed to reach the manual while writing a server — the previous
tool_developer sub-agent had it, developer's did not.
- server_admin: + reload_tools only (24 → 25). Adding a third-party MCP
server is something this profile does as often as developer does.
Deliberately not to its sub-agent: reconnecting the MCP manager is a
process-wide operation belonging to the main agent.
- The prompt and the manual took on what the deleted profile knew and
create_mcp_server.md did not: reload_tools before the first test_mcp_tool
(a freshly registered server is not connected, so the test fails and the
iteration is wasted), the smoke test read by exit code — 124 means timeout
killed a server still running, 0 means it exited on its own, usually a
main() without parentheses — absolute command/cwd, mcp_status as discovery
only, and the steps that stay inline instead of going to a sub-agent.
mcp_status and test_mcp_tool were built without an MCP manager, and their
fallback did `from navi.api.deps import _mcp_manager` — a name that does not
exist, so a live call raised ImportError rather than the intended "MCP
manager not available". The tools always passed a manager in tests, which is
why nothing caught it. Both now receive the manager at construction and fall
back to the live one lazily.
Sessions and profile_overrides are reassigned before the restart: agent.py
resolves the session's profile without a guard, so a deleted profile turns
every session that referenced it into an uncaught ProfileNotFound. Nothing
in the test suite pins the profile inventory, and profiles are read once at
import time — reload_tools does not re-read them — so this ships as a
restart, and the restart is also what makes it take effect.
Eugene Sukhodolskiy
committed
2 days ago
|

tools, mcp: a description is a hook, not a manual
...
With 24 tools in a profile the schemas are the largest fixed block of every
request, and they are paid for whether or not the tool is called. Eleven of
them carried a manual's worth of prose in their `description` — examples,
error tables, "common mistakes", the mechanics of the file areas. Reading
them against the manuals first, the manuals already held all of it, and
richer: this was duplication sitting in the one place the model cannot avoid
reading, not knowledge that had nowhere else to live. So the descriptions
became selection hooks — what the tool is for, when to pick it over its
neighbour, and the one rule that cannot wait — and the detail moved to
tool_manual("<tool>"), which costs nothing until the tool is actually in
play.
filesystem, spawn_agent, schedule_recall, manage_recall, scratchpad, plan,
share_file, content_publish, reflect, todo and peer. Descriptions drop from
18.4 KB to 11.5 KB, those tools' schemas from 25.5 KB to 16.8 KB, and a
profile's whole native toolset from 38.3 KB to 29.6 KB (server_admin,
developer; secretary 33.7 → 25.0, navi_code 32.9 → 25.9, tool_developer
37.9 → 29.2, modeler_3d 33.4 → 24.8). Stated as a budget rather than a byte
count: roughly 4k tokens of the model's window per request, back.
Two things were deliberately not shortened. The filesystem edit ladder
(`edit` → `edit_lines` → `smart_edit`, last resort, costs an LLM call) and
todo's mandatory `validation` on `done` both exist to prevent an extra
round trip; a shorter description there would trade tokens for turns. What
todo lost is the prose around the rule, not the rule. `enabled.json` was
left alone too: weather, gmail and get_current_datetime are opt-in and the
DB shows 12, 16 and 40 calls, so they are opt-in tools that are actually
used.
peer had no manual at all, which is how this surfaced: its description was
the only documentation the tool had. manuals/peer.md is now written from
the tool and its /peer route — the ask/status/list actions, the fact that
the hive is only a phone book and the question travels peer-to-peer, the
one-concurrent-answer semaphore, the two independent recursion guards, and
the six error codes.
The MCP server instructions leave the system prompt the same way. They were
8 KB of always-present prose per request, most of it workflow detail the
model only needs when it is about to use that server. Each server now
contributes one line — a new optional `summary` field in mcp_servers.d/*.json,
defaulting to the first sentence of its `instructions` — and
tool_manual("<server>") returns the full text plus the tool names the server
declares. server_admin's MCP block goes from 8.3 KB to 1.3 KB, secretary's
from 8.1 to 1.2. tool_manual learned to answer for a server (a tool name
still wins over a server of the same name, and dash/underscore spelling is
tolerated), and the index names the servers alongside the tools.
One judgement call worth recording: gnexus-book's and navi-web's
instructions end in an absolute "NEVER bypass these tools with filesystem,
terminal or code_exec" rule. Moving that on demand would have been a
behavioural regression dressed as a token saving, so those two `summary`
fields carry the prohibition verbatim alongside the hook, and it stays
always visible.
docs/tools.md gains a section on the split (a description is paid for every
request, a manual only when the tool is used), which is where the next
person will look before trimming a description back into a manual.
Full suite green (1577 passed, 1 skipped). No frontend change and no new
dependency, so the deploy is a pull and a restart.
Eugene Sukhodolskiy
committed
2 days ago
|

tool_manual: a manual the agent cannot reach is not a manual
...
tool_manual asked the filesystem for the exact string the model sent, then
fell back to registry.get(). The executor already tolerated the ways models
mangle MCP names; the tool that documents them did not, so the agent calling
mcp__navi-3d__compile_scad got a schema dump while manuals/compile_scad.md sat
there unread. Resolving a name now happens in one place — resolve_tool moved
out of the executor into navi/core/tool_utils.py — so a name that works when
the tool is *called* also works when it is *asked about*. A miss suggests the
closest names instead of "not found", and a tool that exists but is not enabled
for this profile says exactly that: documenting it is still the useful answer,
but the agent must not go on to call it. params["tool_name"] was also a crash
waiting to happen — KeyError with the key omitted, AttributeError on null or a
number — and now anything that is not a usable string means "no name given",
which returns the index of manuals grouped by source.
The generated manual now renders the whole schema. An array of objects showed
as "(array, optional)" with its item shape invisible, so a model that could not
see the fields guessed them; objects, array items, oneOf/anyOf branches, enums
and defaults are all spelled out now, depth-capped. The header names the source
(native / mcp__<server>__) and says the text is a parameter contract, not a
curated manual — otherwise a thin schema reads as "this tool is simple".
Then the manuals themselves. A manual is filed under the tool's own name, and
five of the sixteen were not: write_tool.md and write_mcp_server.md documented
tools that do not exist anywhere (the tools are reload_tools and
create_mcp_server), model_3d.md documented compile_scad, render_3d.md documented
render_stl, and write_context_provider.md documented no tool at all — it is a
guide, and now lives in manuals/guides/. The prompts pointed at the same wrong
names (tool_developer's system prompt, persona_navi_code, docs/context_providers),
so fixing the files without the callers would have moved the breakage rather
than removed it. On HEAD the probe was blunt: tool_manual("mcp__navi-3d__compile_scad")
returned None and the agent got the schema; now it returns the manual.
Nothing was enforcing any of this, which is how it drifted. tests/unit/tools/
test_manual_drift.py reads the real repository tree: every manuals/*.md must be
named after a tool that exists (built-ins and tools/*.py by their name
assignments, MCP tools by @mcp.tool(name=…) in the server sources and by the
mcp_servers.d groups, whose file name is the server name), every manual must be
reachable by both its bare name and its full MCP spelling, a guide must never
shadow a tool's name, and every manuals/<x>.md or tool_manual("x") cited in a
doc or a prompt must resolve. The next rename fails the suite instead of
quietly wasting a manual.
Finally, manuals for the built-ins the agent uses most — ssh_exec, todo, memory,
plan, filesystem, notify, code_exec — written from the current schemas and the
code paths that produce the errors, so the gotchas are the real ones (todo's
`done` validation is enforced; memory's `list` returns categories, not facts;
filesystem's `delete` removes a tree with no prompt). reload_tools.md absorbs
the file format from the deleted write_tool.md, which is where a self-extension
recipe belongs now.
1556 passed, 1 skipped. ruff delta zero.
Eugene Sukhodolskiy
committed
2 days ago
|

startup, llm, profiles, todo: five more the prod logs gave up
...
The same 5-day sweep of prod's journal as 5e51f87, continued. Every one of
these failed silently: no error in the UI, just a line in the log or a hole
in the database.
complete()/embed() retry a connection error before falling back.
stream_complete has always given its first chunk two attempts, but the
non-streaming path made exactly one per (server, model), and with a single
configured server (ollama.com) one ReadTimeout raised "All backends
exhausted: ReadTimeout" outright. That is the whole story of 2026-10-08,
when 27 calls died that way in a day — 23 of them memory summarisation, the
rest planning. Attempts now come from settings.llm_complete_retries (2) with
llm_retry_backoff_sec (2.0) between them; a missing model is still not
retried (a 404 is an answer, not a hiccup), and a lone server is still never
blacklisted — blacklisting it would block the next request for _TTL.
Startup waits for Postgres. create_container() opens the pool eagerly and
the retry loop for the DDL tables sits below it, so a host that comes up
before docker takes the whole lifespan with it. On 2026-10-07 15:19:40 the
agent ran `sudo reboot` on itself at the user's request ("Перезагрузи
себя"); the host booted 15:20:03, navi started at 15:20:18, hit
ConnectionRefusedError on the pool and logged "Application startup failed.
Exiting." Connection errors now retry for ~30s; a genuinely broken config
(a missing DATABASE_URL) still fails fast.
The profile loader reports its tally. A profile dropped by an unreadable
config.json is skipped by design so it cannot take the server down, but on
that same 15:20 boot four of them — modeler_3d, navi_code, secretary,
server_admin — were dropped with `Extra data: line 125 column 1`, and the
only trace was one error line each: the UI simply showed four profiles
fewer.
todo recovers a dropped `op`. Of the 3385 todo calls stored on prod, 73
arrived with the discriminator flattened away ({"index": 1, "status":
"in_progress"}, {"action": "view"}, {"": "add", "tasks": [...]}) and got a
bare "Unknown op: None" — a wasted round trip for an intent the remaining
arguments state plainly. The error message now names the five ops and what
each needs.
Message.created_at is stamped at creation. It defaulted to None and only
some construction sites filled it, so 85% of stored rows had no timestamp at
all (every tool message, 14 168 of 14 168) and no query could slice history
by time or measure a pause. Loading a row passes the column through
explicitly, so NULL rows stay NULL instead of being restamped with "now".
Every new test was checked to fail with its fix reverted. Full suite green
(1466 passed, 1 skipped). No frontend or dependency change, so the deploy is
a pull and a restart.
Eugene Sukhodolskiy
committed
2 days ago
|

memory+llm: three failures the prod logs gave up
...
All three came out of a survey of the last days of journalctl on prod,
and each one silently destroyed something the user had already paid for.
memory_facts: the no-embedding INSERT bound $13 while its column list
had 12 entries, so every fact written while the embedding backend was
down died on PostgresSyntaxError and the extraction was lost — 47
embed failures in three days, most of them landing in this branch.
Embedding input now gets clipped instead of 400'd away. A 400 takes the
whole embedding with it and recall drops to ILIKE over everything; a
reaction-session prompt (a full event envelope inlined) did that 41
times. settings.embedding_max_chars (6000, 0 disables) caps the input
at the model's window, so recall still works on the head of the text.
Message strips NUL bytes at the model boundary. PostgreSQL text cannot
hold one, and a NUL arriving in a tool result (reading a binary file)
made the whole session_messages INSERT fail with
CharacterNotInRepertoireError — the turn died and the user lost it.
Three times on 2026-10-07. As a Message validator it covers every
writer downstream: session store, kv store, memory extraction.
Each new test was checked to fail with its fix reverted.
Eugene Sukhodolskiy
committed
2 days ago
|

agent: a profile switch now survives the turn that made it
...
switch_profile repoints the session with one narrow UPDATE, and the run's own
save() — which lands right after every tool batch, the switch's own batch
included — wrote its stale in-memory profile back over it. The post-turn
reload then read the row it had just clobbered, saw no change, and left the
turn bound to the old profile's tools: the agent called exactly the tools
switch_profile had just advertised ("Newly available: ssh_exec, terminal, …")
and got "tool not found" for every one of them. Prod shows both halves — the
session that reported two switches kept the profile it started with, and
agent.profile_reloaded appears zero times in the whole boot.
The reload now compares against the profile the run is bound to
(bound_profile_id), which set_profile cannot move, instead of against
session.profile_id, which the store either rewrites or mutates to the new
value — making the old guard false in both branches by construction. And
save() leaves the profile_id column alone on conflict: set_profile owns it
after creation, the same way name/created_at are owned elsewhere.
A tool the live tool_map does not hold is now logged (tool.not_found, with
profile and live tool count). It previously left no trace at all in the
journal, which is why this took a database dig to find.
Three tests, each verified to fail without its fix: the next LLM call is
offered the new profile's tools, save() does not rewrite profile_id, and a
missing tool is logged.
Eugene Sukhodolskiy
committed
2 days ago
|
mcp: point navi_ui at the port its server actually binds
...
navi_ui is navi's own MCP server (navi/mcp/ui_server), started in-process on
NAVI_UI_MCP_PORT — 8098. The config still carried the pre-8099/8098 default,
localhost:8001, where nothing has listened for a long time: on prod the connect
fails on every start and render_component only ever appeared as an unregistered
phantom in list_tools.
The URL now names the loopback literal rather than "localhost", which can
resolve to ::1 while the server binds IPv4 only.
Three tests assert the file against the live FastMCP object (host, port, path,
transport) instead of a second copy of the setting — with the old URL two of
them fail, so the drift that produced this cannot come back unnoticed.
Eugene Sukhodolskiy
committed
2 days ago
|