| 2026-10-09 |

webclient: animate <details> cards in every browser, not just Chromium
...
Open/close of the tool cards jumped in Firefox and Safari: the previous
attempt leaned on interpolate-size/::details-content, and interpolate-size
shipped in Chromium alone, so the @supports block degraded to a hard
attribute flip everywhere else.
Drive the element's own height from JS instead. Nothing there is
engine-specific: open animates to scrollHeight, close to the summary
height, overflow is clipped for the duration and the inline styles come
off on transitionend with a timer as a backstop. The element keeps its
native semantics — keyboard, screen readers, find-in-page — because the
open attribute is still what changes; the toggle event still fires, so
ContentCard's :open/@toggle pairing is untouched. prefers-reduced-motion
settles in a single frame.
A newly inserted card grows its grid row from 0fr to 1fr, which
interpolates everywhere and resolves to whatever the card measures, so
the card below slides instead of being displaced.
Measured in Chromium 153 and Firefox 155 against the real component
markup and the built bundle: 33 -> 586px over 18 distinct heights in
both, and the same for insertion, the nested Arguments block and the
plan card. The scroll behaviour of MessageList was left alone.
Eugene Sukhodolskiy
committed
13 hours ago
|

webclient: a stale session list no longer erases a session just created
...
Starting a session from the welcome screen sets selectedProfileId and
then awaits the create POST. The profile watcher fires on the microtask
that the await yields, so the list refetch is issued while the POST is
still in flight — it returns a page that predates the new session, and
fetchSessions replaced the whole array, so the placeholder createSession
had just inserted was gone. The session only came back on a manual
refresh. Reproduced in a store test.
sessions.js now remembers a session it created along with the id of the
last list request that had already started at that moment, and a page
that started after the creation is still believed. So the rescue is
bounded by one round trip — it cannot turn into a ghost row: a request
that left once the session existed reports it missing and it goes, and
deleteSession clears the mark outright. Search and special-list pages
are never padded this way, since they answer a different question.
Eugene Sukhodolskiy
committed
14 hours ago
|
webclient: let the profile card description run to its full length
...
The three-line clamp I added with the card-height fix cut the text of
eight of the nine profiles — every description ended in an ellipsis, so
the grid told you less about each profile the more there was to say.
Measured in a browser against the built CSS: 8 of 9 clipped at 1440px,
9 of 9 at 390px. With the clamp gone, none are clipped at 1440, 1100,
900, 768, 600 or 390px.
`grid-auto-rows: 1fr` already keeps the cards in a row level, so the
clamp bought nothing it was not paying for: the row is as tall as the
longest description (230px on desktop, 205px on a phone), and the
welcome pane scrolls when the grid outgrows it.
Eugene Sukhodolskiy
committed
14 hours ago
|
webclient: one icon per profile, and one height for every card
...
Six of the ten profiles rendered the same robot, because the map in
WelcomeScreen only knew four ids. The map moves to utils/profileIcons.js,
covers all ten — including the close pairs (`coder`/`developer`, the two
3D profiles) — and falls back to a hashed glyph from a pool rather than
one robot, so custom profiles differ from each other and stay stable
across renders. Covered by five unit tests.
Cards were ragged because each row took the height of its own content:
`grid-auto-rows: 1fr` makes every row as tall as the tallest card. A
fixed pixel height was the obvious alternative and is worse — "Server
Administrator" wraps on a phone. A three-line clamp on the desktop
description keeps one wordy profile from inflating every row.
Eugene Sukhodolskiy
committed
14 hours ago
|
webclient: let the welcome pane scroll, and pin the mobile menu toggle
...
The profile grid grows with the profile list, and on a phone it is taller than
the viewport once there are more than a couple of profiles. Nothing could
scroll: `#app` and `.app-main` both hide overflow, and `.welcome-screen` had no
overflow of its own, so the cards were clipped rather than pushed down.
The pane now scrolls, and its vertical centering is `safe center` — with a plain
`center` the overflow goes past the flex start edge and cannot be scrolled back
to. The sidebar toggle becomes `position: fixed` so it does not scroll away with
the content it sits above.
Eugene Sukhodolskiy
committed
23 hours ago
|
webclient: rebuild dist for the role-aware MCP keys panel
...
`webclient/dist` is tracked and served straight by the API server — there is no
node on the prod host — so the source change alone would have left prod showing
the old copy, telling an ordinary user that the shared key from the config is in
use for a server that is simply not connected for them.
Eugene Sukhodolskiy
committed
23 hours ago
|

profiles: restricted profiles for ordinary users, and a role gate that holds
...
`is_admin_only` was checked in one place out of nine and was not read from
config.json at all, so all seven profiles were reachable by every account with
role `user`. The flag now lives in config.json — the file is the baseline, a
`profile_overrides` row still wins on top of it — and one predicate,
`admin_only_blocked`, is the single place the rule is expressed. The nine
surfaces that list, switch to, spawn or resolve a profile all consult it:
`POST /sessions`, the WebSocket, switch_profile, list_profiles, the system
prompt's "Available profiles" block, spawn_agent and the Synapse reaction
runner. The prompt cache is now keyed by (profile, role), so a user's prompt
can never be served an admin's profile list.
The seven existing profiles (developer, discuss, dispatcher, modeler_3d,
navi_code, secretary, server_admin) are marked admin-only. Three new ones take
their place for ordinary users: assistant, designer_3d and coder. They share one
native tool set — ssh_exec, peer, reload_tools, create_mcp_server, test_mcp_tool,
image_view and gmail are withheld — and differ only in system prompt, model and
MCP groups. navi-web's raw `request` group, and the whole of gnexus-creds and
tgclient, are withheld too.
MCP per-user keys gain the missing half of the rule: a server that declares a
`user_key` slot is refused to anyone but an admin who has no personal key, and
is left out of their tool list entirely, instead of quietly falling back to the
owner's credential and appearing as a tool that cannot work. The refusal names
the server and points at Settings.
Also closes `GET /agents/prompts`, which served every profile's system prompt to
anyone, with no user dependency at all.
The accepted residual risk is written down in docs/profiles.md: the working
directory is a convention, not a sandbox.
Eugene Sukhodolskiy
committed
23 hours ago
|
| 2026-10-08 |

tools: a non-admin may reach their own session directory
...
A non-admin was confined to user_data/<user_id>/ alone, which made four
tools reject the directory the user's own uploads land in: share_file
refused the file it was asked to send back, filesystem could not read it,
and terminal/code_exec could not be pointed at it. In a prod session the
agent worked around the refusal by copying the uploaded file into its
sandbox with code_exec — the very bypass the injected security policy
forbids — and the copy collided with a same-named file, so the user got a
link to Project_1.mp3 instead of their own upload, plus an 8.6 MB duplicate.
navi/tools/_internal/areas.py now names the two roots a user owns, and the
four call sites share it. Relative paths still resolve into the sandbox,
and each tool keeps its own way of refusing: terminal returns
sandbox_violation, code_exec silently falls back to the sandbox root. The
share_file refusal lists both roots instead of "outside user sandbox", so
the agent retries in the right one instead of routing around it.
The session directory belongs to the same user: its id arrives from the
runtime context, never from tool arguments, and only the session's owner
can open it.
Eugene Sukhodolskiy
committed
1 day ago
|
persona: the name rule, and a template that actually loads the persona
...
persona.txt gains a third strict rule: the first name is Navi — Latin even
inside Russian text, «Нави» known but not spoken by default — and the
instance codename is a second name, Latin and capitalised, never
transliterated. The model had only ever seen the Latin name plus the
identity block's lowercase codename, which is where "Нэви" came from.
None of this reached a server: deploy/env.template never set
NAVI_PERSONA_FILE, so an install from the template ran with
settings.navi_persona empty and every rule in the file inert. The variable
joins the template, where .env.example already had it.
Eugene Sukhodolskiy
committed
1 day ago
|

mcp: credentials leave the tree, and sub-agents lose the task board
...
Five tracked configs (gnexus-creds, gntodo, hard-panel, synapse, tgclient)
carried live bearer tokens — in git, in every clone, and handed to the admin
client by GET /admin/mcp/config. They now name a variable, and the value lives
in the service .env (chmod 600, untracked).
Substitution happens once, at transport open, next to the project-relative
path resolution: that is the single place a transport is built, so every
stored and serialised config stays placeholder-only and save_mcp_servers —
reached by create_mcp_server and PUT /admin/mcp/config — cannot write a secret
back into a tracked file. A variable that is not set refuses the connection
instead of dropping the header, which would fail open: a server may answer an
unauthenticated request as an anonymous user rather than returning 401.
Sub-agents also drop gntodo, gnexus-creds, tgclient and synapse from their
scopes, and the runner injects a scope anchor unconditionally — a sub-agent
asked to echo one word had picked a project off the task board and researched
it for forty iterations.
Docs: docs/mcp.md#secrets, a migration recipe in deploy/UPDATE.md, the variable
names in .env.example and deploy/env.template, and a regression guard over the
tracked configs.
The old values stay in git history; rotating all five at their sources is the
remediation, not a follow-up.
Eugene Sukhodolskiy
committed
1 day ago
|

subagents, secrets: an honest result contract, and two leaks closed
...
A pass over the sub-agent system, plus the two findings from the previous one.
Leaks:
- `redact_args` / `is_sensitive_key` (navi/tools/_internal/redact.py) mask
credential-shaped keys at every log site that prints tool arguments — by key
name, for all tools, so a future MCP tool is covered without declaring
anything. ssh_exec passwords were sitting in the prod journal in plaintext.
- `mcp_servers.d/*.json` are tracked in git and carried the scaffolding
machine's absolute paths, so navi-3d/navi-web silently vanished from the
toolset off that machine. They now hold project-relative paths, resolved in
`McpClient.open_transport` — not at load time, because create_mcp_server
round-trips every config file through save_mcp_servers.
Sub-agent result contract:
- `run_ephemeral` returns a `SubAgentOutcome` (text + status: ok, timeout,
max_iterations, thinking_stall, user_stop, context_overflow) instead of
`(text, bool)`. Every exit routes through `SubAgentRunner._finish`, so token
accounting is never dropped and a partial run always carries a progress
report — the stop-after-stream path used to return a bare string.
- `spawn_agent` renders its header from the status. Before this, every short
run reached the parent as "hit iteration limit", and the parent repeats that
diagnosis to the user — a timeout was reported as an iteration limit.
- `ContextTooLargeError` no longer kills a sub-agent: the parent compresses its
own context, a sub-agent has no session to compress into, so it stops with
`context_overflow` and hands back what it did rather than surfacing as
"Sub-agent failed: …" with the whole run discarded.
- The result is capped at 8000 chars (head and tail kept) so a long run cannot
flood the parent's context, and tool arguments in the progress report go
through `redact_args`.
- The navi-3d calling convention left the runner for the server's own tool
schemas — a generic runner should not know one MCP server's path shapes.
- Validation moved inside spawn_agent's try: a bad argument is a clean
ToolResult the model can read, not an exception in the parent's loop.
Tests: 1636 passed, 1 skipped. New: SubAgentOutcome paths and headers, result
truncation, context overflow, clean argument failures.
Eugene Sukhodolskiy
committed
1 day ago
|

profiles: tool_developer folds into developer
...
The profile was a duplicate on every axis we could measure. 22 of its 26
native tools were already developer's; the four it alone held —
reload_tools, create_mcp_server, test_mcp_tool, mcp_status — are 2.9 KB of
schema. Its model chain was the same seven models. Of its 14 KB prompt,
about 8 KB was copied verbatim from developer's (the whole `## Orchestration
model` block and everything from `## Editing policy` down), and most of the
remainder restated manuals/create_mcp_server.md, which already carried the
same ten-step workflow in more detail. It was not a specialisation, it was a
snapshot: `git log -S '"reload_tools"'` shows the tool lived in developer
until 61fa370 rewrote that profile around MCP and cut it off.
What kept it alive was a premise that no longer holds — that Navi's own
capabilities would be written as in-repo tools. They are MCP servers now,
and an MCP server is not a file in this repository with a life of its own:
it is an isolated process registered from mcp_servers.d/. So there is no
reason left for a profile whose only distinct feature is a toolset a general
developer profile can hold, and every reason to stop maintaining a second
prompt that drifts against the first.
- developer: + reload_tools, create_mcp_server, test_mcp_tool, mcp_status
(24 → 28 native). Its sub-agent gains tool_manual, which is what it
actually needed to reach the manual while writing a server — the previous
tool_developer sub-agent had it, developer's did not.
- server_admin: + reload_tools only (24 → 25). Adding a third-party MCP
server is something this profile does as often as developer does.
Deliberately not to its sub-agent: reconnecting the MCP manager is a
process-wide operation belonging to the main agent.
- The prompt and the manual took on what the deleted profile knew and
create_mcp_server.md did not: reload_tools before the first test_mcp_tool
(a freshly registered server is not connected, so the test fails and the
iteration is wasted), the smoke test read by exit code — 124 means timeout
killed a server still running, 0 means it exited on its own, usually a
main() without parentheses — absolute command/cwd, mcp_status as discovery
only, and the steps that stay inline instead of going to a sub-agent.
mcp_status and test_mcp_tool were built without an MCP manager, and their
fallback did `from navi.api.deps import _mcp_manager` — a name that does not
exist, so a live call raised ImportError rather than the intended "MCP
manager not available". The tools always passed a manager in tests, which is
why nothing caught it. Both now receive the manager at construction and fall
back to the live one lazily.
Sessions and profile_overrides are reassigned before the restart: agent.py
resolves the session's profile without a guard, so a deleted profile turns
every session that referenced it into an uncaught ProfileNotFound. Nothing
in the test suite pins the profile inventory, and profiles are read once at
import time — reload_tools does not re-read them — so this ships as a
restart, and the restart is also what makes it take effect.
Eugene Sukhodolskiy
committed
1 day ago
|

tools, mcp: a description is a hook, not a manual
...
With 24 tools in a profile the schemas are the largest fixed block of every
request, and they are paid for whether or not the tool is called. Eleven of
them carried a manual's worth of prose in their `description` — examples,
error tables, "common mistakes", the mechanics of the file areas. Reading
them against the manuals first, the manuals already held all of it, and
richer: this was duplication sitting in the one place the model cannot avoid
reading, not knowledge that had nowhere else to live. So the descriptions
became selection hooks — what the tool is for, when to pick it over its
neighbour, and the one rule that cannot wait — and the detail moved to
tool_manual("<tool>"), which costs nothing until the tool is actually in
play.
filesystem, spawn_agent, schedule_recall, manage_recall, scratchpad, plan,
share_file, content_publish, reflect, todo and peer. Descriptions drop from
18.4 KB to 11.5 KB, those tools' schemas from 25.5 KB to 16.8 KB, and a
profile's whole native toolset from 38.3 KB to 29.6 KB (server_admin,
developer; secretary 33.7 → 25.0, navi_code 32.9 → 25.9, tool_developer
37.9 → 29.2, modeler_3d 33.4 → 24.8). Stated as a budget rather than a byte
count: roughly 4k tokens of the model's window per request, back.
Two things were deliberately not shortened. The filesystem edit ladder
(`edit` → `edit_lines` → `smart_edit`, last resort, costs an LLM call) and
todo's mandatory `validation` on `done` both exist to prevent an extra
round trip; a shorter description there would trade tokens for turns. What
todo lost is the prose around the rule, not the rule. `enabled.json` was
left alone too: weather, gmail and get_current_datetime are opt-in and the
DB shows 12, 16 and 40 calls, so they are opt-in tools that are actually
used.
peer had no manual at all, which is how this surfaced: its description was
the only documentation the tool had. manuals/peer.md is now written from
the tool and its /peer route — the ask/status/list actions, the fact that
the hive is only a phone book and the question travels peer-to-peer, the
one-concurrent-answer semaphore, the two independent recursion guards, and
the six error codes.
The MCP server instructions leave the system prompt the same way. They were
8 KB of always-present prose per request, most of it workflow detail the
model only needs when it is about to use that server. Each server now
contributes one line — a new optional `summary` field in mcp_servers.d/*.json,
defaulting to the first sentence of its `instructions` — and
tool_manual("<server>") returns the full text plus the tool names the server
declares. server_admin's MCP block goes from 8.3 KB to 1.3 KB, secretary's
from 8.1 to 1.2. tool_manual learned to answer for a server (a tool name
still wins over a server of the same name, and dash/underscore spelling is
tolerated), and the index names the servers alongside the tools.
One judgement call worth recording: gnexus-book's and navi-web's
instructions end in an absolute "NEVER bypass these tools with filesystem,
terminal or code_exec" rule. Moving that on demand would have been a
behavioural regression dressed as a token saving, so those two `summary`
fields carry the prohibition verbatim alongside the hook, and it stays
always visible.
docs/tools.md gains a section on the split (a description is paid for every
request, a manual only when the tool is used), which is where the next
person will look before trimming a description back into a manual.
Full suite green (1577 passed, 1 skipped). No frontend change and no new
dependency, so the deploy is a pull and a restart.
Eugene Sukhodolskiy
committed
1 day ago
|

tool_manual: a manual the agent cannot reach is not a manual
...
tool_manual asked the filesystem for the exact string the model sent, then
fell back to registry.get(). The executor already tolerated the ways models
mangle MCP names; the tool that documents them did not, so the agent calling
mcp__navi-3d__compile_scad got a schema dump while manuals/compile_scad.md sat
there unread. Resolving a name now happens in one place — resolve_tool moved
out of the executor into navi/core/tool_utils.py — so a name that works when
the tool is *called* also works when it is *asked about*. A miss suggests the
closest names instead of "not found", and a tool that exists but is not enabled
for this profile says exactly that: documenting it is still the useful answer,
but the agent must not go on to call it. params["tool_name"] was also a crash
waiting to happen — KeyError with the key omitted, AttributeError on null or a
number — and now anything that is not a usable string means "no name given",
which returns the index of manuals grouped by source.
The generated manual now renders the whole schema. An array of objects showed
as "(array, optional)" with its item shape invisible, so a model that could not
see the fields guessed them; objects, array items, oneOf/anyOf branches, enums
and defaults are all spelled out now, depth-capped. The header names the source
(native / mcp__<server>__) and says the text is a parameter contract, not a
curated manual — otherwise a thin schema reads as "this tool is simple".
Then the manuals themselves. A manual is filed under the tool's own name, and
five of the sixteen were not: write_tool.md and write_mcp_server.md documented
tools that do not exist anywhere (the tools are reload_tools and
create_mcp_server), model_3d.md documented compile_scad, render_3d.md documented
render_stl, and write_context_provider.md documented no tool at all — it is a
guide, and now lives in manuals/guides/. The prompts pointed at the same wrong
names (tool_developer's system prompt, persona_navi_code, docs/context_providers),
so fixing the files without the callers would have moved the breakage rather
than removed it. On HEAD the probe was blunt: tool_manual("mcp__navi-3d__compile_scad")
returned None and the agent got the schema; now it returns the manual.
Nothing was enforcing any of this, which is how it drifted. tests/unit/tools/
test_manual_drift.py reads the real repository tree: every manuals/*.md must be
named after a tool that exists (built-ins and tools/*.py by their name
assignments, MCP tools by @mcp.tool(name=…) in the server sources and by the
mcp_servers.d groups, whose file name is the server name), every manual must be
reachable by both its bare name and its full MCP spelling, a guide must never
shadow a tool's name, and every manuals/<x>.md or tool_manual("x") cited in a
doc or a prompt must resolve. The next rename fails the suite instead of
quietly wasting a manual.
Finally, manuals for the built-ins the agent uses most — ssh_exec, todo, memory,
plan, filesystem, notify, code_exec — written from the current schemas and the
code paths that produce the errors, so the gotchas are the real ones (todo's
`done` validation is enforced; memory's `list` returns categories, not facts;
filesystem's `delete` removes a tree with no prompt). reload_tools.md absorbs
the file format from the deleted write_tool.md, which is where a self-extension
recipe belongs now.
1556 passed, 1 skipped. ruff delta zero.
Eugene Sukhodolskiy
committed
1 day ago
|

startup, llm, profiles, todo: five more the prod logs gave up
...
The same 5-day sweep of prod's journal as 5e51f87, continued. Every one of
these failed silently: no error in the UI, just a line in the log or a hole
in the database.
complete()/embed() retry a connection error before falling back.
stream_complete has always given its first chunk two attempts, but the
non-streaming path made exactly one per (server, model), and with a single
configured server (ollama.com) one ReadTimeout raised "All backends
exhausted: ReadTimeout" outright. That is the whole story of 2026-10-08,
when 27 calls died that way in a day — 23 of them memory summarisation, the
rest planning. Attempts now come from settings.llm_complete_retries (2) with
llm_retry_backoff_sec (2.0) between them; a missing model is still not
retried (a 404 is an answer, not a hiccup), and a lone server is still never
blacklisted — blacklisting it would block the next request for _TTL.
Startup waits for Postgres. create_container() opens the pool eagerly and
the retry loop for the DDL tables sits below it, so a host that comes up
before docker takes the whole lifespan with it. On 2026-10-07 15:19:40 the
agent ran `sudo reboot` on itself at the user's request ("Перезагрузи
себя"); the host booted 15:20:03, navi started at 15:20:18, hit
ConnectionRefusedError on the pool and logged "Application startup failed.
Exiting." Connection errors now retry for ~30s; a genuinely broken config
(a missing DATABASE_URL) still fails fast.
The profile loader reports its tally. A profile dropped by an unreadable
config.json is skipped by design so it cannot take the server down, but on
that same 15:20 boot four of them — modeler_3d, navi_code, secretary,
server_admin — were dropped with `Extra data: line 125 column 1`, and the
only trace was one error line each: the UI simply showed four profiles
fewer.
todo recovers a dropped `op`. Of the 3385 todo calls stored on prod, 73
arrived with the discriminator flattened away ({"index": 1, "status":
"in_progress"}, {"action": "view"}, {"": "add", "tasks": [...]}) and got a
bare "Unknown op: None" — a wasted round trip for an intent the remaining
arguments state plainly. The error message now names the five ops and what
each needs.
Message.created_at is stamped at creation. It defaulted to None and only
some construction sites filled it, so 85% of stored rows had no timestamp at
all (every tool message, 14 168 of 14 168) and no query could slice history
by time or measure a pause. Loading a row passes the column through
explicitly, so NULL rows stay NULL instead of being restamped with "now".
Every new test was checked to fail with its fix reverted. Full suite green
(1466 passed, 1 skipped). No frontend or dependency change, so the deploy is
a pull and a restart.
Eugene Sukhodolskiy
committed
1 day ago
|

memory+llm: three failures the prod logs gave up
...
All three came out of a survey of the last days of journalctl on prod,
and each one silently destroyed something the user had already paid for.
memory_facts: the no-embedding INSERT bound $13 while its column list
had 12 entries, so every fact written while the embedding backend was
down died on PostgresSyntaxError and the extraction was lost — 47
embed failures in three days, most of them landing in this branch.
Embedding input now gets clipped instead of 400'd away. A 400 takes the
whole embedding with it and recall drops to ILIKE over everything; a
reaction-session prompt (a full event envelope inlined) did that 41
times. settings.embedding_max_chars (6000, 0 disables) caps the input
at the model's window, so recall still works on the head of the text.
Message strips NUL bytes at the model boundary. PostgreSQL text cannot
hold one, and a NUL arriving in a tool result (reading a binary file)
made the whole session_messages INSERT fail with
CharacterNotInRepertoireError — the turn died and the user lost it.
Three times on 2026-10-07. As a Message validator it covers every
writer downstream: session store, kv store, memory extraction.
Each new test was checked to fail with its fix reverted.
Eugene Sukhodolskiy
committed
1 day ago
|

agent: a profile switch now survives the turn that made it
...
switch_profile repoints the session with one narrow UPDATE, and the run's own
save() — which lands right after every tool batch, the switch's own batch
included — wrote its stale in-memory profile back over it. The post-turn
reload then read the row it had just clobbered, saw no change, and left the
turn bound to the old profile's tools: the agent called exactly the tools
switch_profile had just advertised ("Newly available: ssh_exec, terminal, …")
and got "tool not found" for every one of them. Prod shows both halves — the
session that reported two switches kept the profile it started with, and
agent.profile_reloaded appears zero times in the whole boot.
The reload now compares against the profile the run is bound to
(bound_profile_id), which set_profile cannot move, instead of against
session.profile_id, which the store either rewrites or mutates to the new
value — making the old guard false in both branches by construction. And
save() leaves the profile_id column alone on conflict: set_profile owns it
after creation, the same way name/created_at are owned elsewhere.
A tool the live tool_map does not hold is now logged (tool.not_found, with
profile and live tool count). It previously left no trace at all in the
journal, which is why this took a database dig to find.
Three tests, each verified to fail without its fix: the next LLM call is
offered the new profile's tools, save() does not rewrite profile_id, and a
missing tool is logged.
Eugene Sukhodolskiy
committed
1 day ago
|
mcp: point navi_ui at the port its server actually binds
...
navi_ui is navi's own MCP server (navi/mcp/ui_server), started in-process on
NAVI_UI_MCP_PORT — 8098. The config still carried the pre-8099/8098 default,
localhost:8001, where nothing has listened for a long time: on prod the connect
fails on every start and render_component only ever appeared as an unregistered
phantom in list_tools.
The URL now names the loopback literal rather than "localhost", which can
resolve to ::1 while the server binds IPv4 only.
Three tests assert the file against the live FastMCP object (host, port, path,
transport) instead of a second copy of the setting — with the old URL two of
them fail, so the drift that produced this cannot come back unnoticed.
Eugene Sukhodolskiy
committed
1 day ago
|

tools: list_tools defaults to the profile the run executes as
...
Asking "what tools do I have?" required the agent to know its own profile id
— and to be right about it after a switch_profile had moved the session. The
run knows: run_stream now publishes the active profile id (re-published when
switch_profile reloads the run) alongside the model, SubAgentRunner publishes
its own rather than inheriting the parent's, and ToolContext carries it into
every execute() call.
list_tools reads that when profile_id is omitted, so the bare call answers the
question, and it says which profile it assumed ("current profile") so the
answer cannot be mistaken for another profile's.
scope='agent'|'subagent' selects which of the profile's two toolsets to list —
the set spawn_agent would hand a sub-agent running that profile, which is a
different list and until now had no way to be inspected.
Item D of the list_tools plan, kept separate from the accuracy/compactness work.
Eugene Sukhodolskiy
committed
1 day ago
|

tools: list_tools tells the truth and costs a tenth of the context
...
The tool read the profile config, not the registry: a tool the config declares
but nothing registered — mcp__navi_ui__render_component, whose server never
connects — was advertised as callable, and the agent walked into "tool not
found" (28 such warnings on prod). It now resolves every name against the live
registry and reports the rest separately as "Not registered", so a phantom
entry reads as a broken server, not as a tool to try.
It also returned every description it could find. For server_admin that was
125 tools and 35 316 bytes per call — ~9k tokens to answer "do I have anything
for ssh". Names are now grouped by source (native, then one section per MCP
server) with descriptions behind verbose, and query filters by substring over
names and descriptions: the ssh question costs 88 bytes instead of 35 KB, and
a plain listing drops to 3.7 KB (9.5x smaller).
Tests cover the phantom-tool case, query matching by name and by description,
and the size property that makes names-only the default.
Eugene Sukhodolskiy
committed
1 day ago
|
| 2026-10-07 |
config: let server_admin manage synapse hub keys
...
server_admin keeps the whole synapse server now — read, write and admin — so
it can issue and revoke hub API keys and change hub settings without a
detour into the admin panel. All 35 registered tools resolve.
secretary stays on read; the admin group is still withheld from every other
profile.
Eugene Sukhodolskiy
committed
2 days ago
|
config: give server_admin and secretary the synapse MCP tools
...
The synapse server was connecting and registering all 35 of its tools, but
no profile listed it under tools.*.mcp, so build_tool_list never handed a
single one to an agent — the hub was up and unreachable at the same time.
server_admin takes read + write (32 tools: events, deliveries, rules,
sources, targets, routing); secretary takes read only (15). The admin group
— key_issue, key_revoke, settings_put — is deliberately left out of both:
issuing and revoking hub keys stays a manual operation.
Eugene Sukhodolskiy
committed
2 days ago
|
config: run every profile on glm-5.3-flash:cloud first
...
Four chains still led with gemma4:31b-cloud, so those sessions were resolved
onto gemma4 even though glm-5.3-flash:cloud is the instance default. glm now
leads every profile.
The rest of each chain stays behind it as a fallback — that is what carried
today's 16:57 run through a ReadTimeout against ollama.com instead of failing
it. developer, navi_code, discuss and modeler_3d already led with glm and are
untouched.
Eugene Sukhodolskiy
committed
2 days ago
|
config: give tgclient the shared key it needs to handshake
...
The server declares a user_key slot and no shared credential, so the
startup handshake went out unauthenticated, got a 401, and the client
was marked disconnected — which kept its tools out of the registry
entirely, groups or no groups. The shared key now backs the handshake
and stays the fallback: a personal key from /mcp-keys still wins.
Eugene Sukhodolskiy
committed
2 days ago
|
webclient: hold the chat at the bottom when a stream ends
...
The list was pinned per streaming delta, but the last things to land arrive
after that final pin: the stats/rating footer, which renders only once
msg.done is set, and the copy buttons attached to code blocks after render.
Nothing re-clamped afterwards — the length watcher never fires (the message
stays in the array) and the landing loop only runs when a session opens — so
the view was left short of the bottom, looking like it had scrolled up.
Re-clamp through the same settle window the landing uses when streaming goes
true -> false, unless the user has scrolled up or a session is loading.
This addresses the late-layout half of the problem; the row is still remounted
when msg.id becomes h_<n>, which is the other source of a jump.
Eugene Sukhodolskiy
committed
2 days ago
|
config: tool groups for tgclient and synapse
...
Both servers declare no groups, so every profile asking for "tgclient":
["read", "write"] resolved to nothing: resolve_group reads the static
config, returned [], and no mcp__tgclient__* name ever reached the agent.
Worse, the server's own instructions did reach it — they are selected by
server name, not by group — so the model was told about tools it did not
have. synapse is not exposed by any profile, so its groups change nothing
today; without them, exposing it would repeat the same failure.
read — observe only.
write — changes to sources, types, targets and routing rules.
admin — hub-wide knobs, not routing: issuing and revoking a source's API
key, and settings overrides on top of .env. The same fence
gnexus-book puts around its service-operator tools.
Eugene Sukhodolskiy
committed
2 days ago
|
deps: declare html2text, which tools/gmail.py imports
...
The import worked on the server only because the package had been
installed into the venv by hand; a fresh sync would have pruned it and
tools/gmail.py would have stopped loading while tools/enabled.json kept
naming gmail. Reload now reports that drift, but the fix is to declare
the dependency. Lock gains html2text and nothing else.
Eugene Sukhodolskiy
committed
2 days ago
|
webclient: reload tools from a button in the MCP tab
...
Settings → MCP gains a Tools block for admins only: one button, then the
same report the tool prints — what loaded, how many are in the registry,
per-file errors, and names in enabled.json nothing answers to.
It sits inside the existing MCP tab rather than a new one: the reload
rewrites the toolset of the whole server, not just this user's MCP keys,
and it belongs next to the thing it affects. Non-admins never see it.
dist rebuilt together with the source, as the server serves the bundle.
Eugene Sukhodolskiy
committed
2 days ago
|
webclient: keep MCP rows at content height on a phone
...
.mcp-keys-row stacks into a column below 768px, and .mcp-keys-info kept
its flex: 1 1 240px. That basis is a width in the desktop row and a
height once the row stacks, so every server reserved 240px and its text
sat at the top of the gap — 402px for a row whose content is 90px.
Back to content height on mobile only, and drop the kit's .form-group
bottom margin there: the row's own flex gap already separates the field
from the buttons.
Eugene Sukhodolskiy
committed
2 days ago
|
admin: POST /admin/tools/reload
...
reload_tools is granted to one profile only (tool_developer), so an admin
whose profile lacks it cannot reload at all — the tool answers "not
found" instead of reloading. The route calls the same reload_all() the
tool does, behind require_admin, for whoever is logged in as an admin.
An admin may read: ok, the tools loaded, the registry total, per-file
errors, names in enabled.json nothing answers to, and the MCP/providers
summary.
Eugene Sukhodolskiy
committed
2 days ago
|