| 2026-10-08 |

startup, llm, profiles, todo: five more the prod logs gave up
...
The same 5-day sweep of prod's journal as 5e51f87, continued. Every one of
these failed silently: no error in the UI, just a line in the log or a hole
in the database.
complete()/embed() retry a connection error before falling back.
stream_complete has always given its first chunk two attempts, but the
non-streaming path made exactly one per (server, model), and with a single
configured server (ollama.com) one ReadTimeout raised "All backends
exhausted: ReadTimeout" outright. That is the whole story of 2026-10-08,
when 27 calls died that way in a day — 23 of them memory summarisation, the
rest planning. Attempts now come from settings.llm_complete_retries (2) with
llm_retry_backoff_sec (2.0) between them; a missing model is still not
retried (a 404 is an answer, not a hiccup), and a lone server is still never
blacklisted — blacklisting it would block the next request for _TTL.
Startup waits for Postgres. create_container() opens the pool eagerly and
the retry loop for the DDL tables sits below it, so a host that comes up
before docker takes the whole lifespan with it. On 2026-10-07 15:19:40 the
agent ran `sudo reboot` on itself at the user's request ("Перезагрузи
себя"); the host booted 15:20:03, navi started at 15:20:18, hit
ConnectionRefusedError on the pool and logged "Application startup failed.
Exiting." Connection errors now retry for ~30s; a genuinely broken config
(a missing DATABASE_URL) still fails fast.
The profile loader reports its tally. A profile dropped by an unreadable
config.json is skipped by design so it cannot take the server down, but on
that same 15:20 boot four of them — modeler_3d, navi_code, secretary,
server_admin — were dropped with `Extra data: line 125 column 1`, and the
only trace was one error line each: the UI simply showed four profiles
fewer.
todo recovers a dropped `op`. Of the 3385 todo calls stored on prod, 73
arrived with the discriminator flattened away ({"index": 1, "status":
"in_progress"}, {"action": "view"}, {"": "add", "tasks": [...]}) and got a
bare "Unknown op: None" — a wasted round trip for an intent the remaining
arguments state plainly. The error message now names the five ops and what
each needs.
Message.created_at is stamped at creation. It defaulted to None and only
some construction sites filled it, so 85% of stored rows had no timestamp at
all (every tool message, 14 168 of 14 168) and no query could slice history
by time or measure a pause. Loading a row passes the column through
explicitly, so NULL rows stay NULL instead of being restamped with "now".
Every new test was checked to fail with its fix reverted. Full suite green
(1466 passed, 1 skipped). No frontend or dependency change, so the deploy is
a pull and a restart.
Eugene Sukhodolskiy
committed
6 hours ago
|

memory+llm: three failures the prod logs gave up
...
All three came out of a survey of the last days of journalctl on prod,
and each one silently destroyed something the user had already paid for.
memory_facts: the no-embedding INSERT bound $13 while its column list
had 12 entries, so every fact written while the embedding backend was
down died on PostgresSyntaxError and the extraction was lost — 47
embed failures in three days, most of them landing in this branch.
Embedding input now gets clipped instead of 400'd away. A 400 takes the
whole embedding with it and recall drops to ILIKE over everything; a
reaction-session prompt (a full event envelope inlined) did that 41
times. settings.embedding_max_chars (6000, 0 disables) caps the input
at the model's window, so recall still works on the head of the text.
Message strips NUL bytes at the model boundary. PostgreSQL text cannot
hold one, and a NUL arriving in a tool result (reading a binary file)
made the whole session_messages INSERT fail with
CharacterNotInRepertoireError — the turn died and the user lost it.
Three times on 2026-10-07. As a Message validator it covers every
writer downstream: session store, kv store, memory extraction.
Each new test was checked to fail with its fix reverted.
Eugene Sukhodolskiy
committed
6 hours ago
|

agent: a profile switch now survives the turn that made it
...
switch_profile repoints the session with one narrow UPDATE, and the run's own
save() — which lands right after every tool batch, the switch's own batch
included — wrote its stale in-memory profile back over it. The post-turn
reload then read the row it had just clobbered, saw no change, and left the
turn bound to the old profile's tools: the agent called exactly the tools
switch_profile had just advertised ("Newly available: ssh_exec, terminal, …")
and got "tool not found" for every one of them. Prod shows both halves — the
session that reported two switches kept the profile it started with, and
agent.profile_reloaded appears zero times in the whole boot.
The reload now compares against the profile the run is bound to
(bound_profile_id), which set_profile cannot move, instead of against
session.profile_id, which the store either rewrites or mutates to the new
value — making the old guard false in both branches by construction. And
save() leaves the profile_id column alone on conflict: set_profile owns it
after creation, the same way name/created_at are owned elsewhere.
A tool the live tool_map does not hold is now logged (tool.not_found, with
profile and live tool count). It previously left no trace at all in the
journal, which is why this took a database dig to find.
Three tests, each verified to fail without its fix: the next LLM call is
offered the new profile's tools, save() does not rewrite profile_id, and a
missing tool is logged.
Eugene Sukhodolskiy
committed
7 hours ago
|
mcp: point navi_ui at the port its server actually binds
...
navi_ui is navi's own MCP server (navi/mcp/ui_server), started in-process on
NAVI_UI_MCP_PORT — 8098. The config still carried the pre-8099/8098 default,
localhost:8001, where nothing has listened for a long time: on prod the connect
fails on every start and render_component only ever appeared as an unregistered
phantom in list_tools.
The URL now names the loopback literal rather than "localhost", which can
resolve to ::1 while the server binds IPv4 only.
Three tests assert the file against the live FastMCP object (host, port, path,
transport) instead of a second copy of the setting — with the old URL two of
them fail, so the drift that produced this cannot come back unnoticed.
Eugene Sukhodolskiy
committed
7 hours ago
|

tools: list_tools defaults to the profile the run executes as
...
Asking "what tools do I have?" required the agent to know its own profile id
— and to be right about it after a switch_profile had moved the session. The
run knows: run_stream now publishes the active profile id (re-published when
switch_profile reloads the run) alongside the model, SubAgentRunner publishes
its own rather than inheriting the parent's, and ToolContext carries it into
every execute() call.
list_tools reads that when profile_id is omitted, so the bare call answers the
question, and it says which profile it assumed ("current profile") so the
answer cannot be mistaken for another profile's.
scope='agent'|'subagent' selects which of the profile's two toolsets to list —
the set spawn_agent would hand a sub-agent running that profile, which is a
different list and until now had no way to be inspected.
Item D of the list_tools plan, kept separate from the accuracy/compactness work.
Eugene Sukhodolskiy
committed
7 hours ago
|

tools: list_tools tells the truth and costs a tenth of the context
...
The tool read the profile config, not the registry: a tool the config declares
but nothing registered — mcp__navi_ui__render_component, whose server never
connects — was advertised as callable, and the agent walked into "tool not
found" (28 such warnings on prod). It now resolves every name against the live
registry and reports the rest separately as "Not registered", so a phantom
entry reads as a broken server, not as a tool to try.
It also returned every description it could find. For server_admin that was
125 tools and 35 316 bytes per call — ~9k tokens to answer "do I have anything
for ssh". Names are now grouped by source (native, then one section per MCP
server) with descriptions behind verbose, and query filters by substring over
names and descriptions: the ssh question costs 88 bytes instead of 35 KB, and
a plain listing drops to 3.7 KB (9.5x smaller).
Tests cover the phantom-tool case, query matching by name and by description,
and the size property that makes names-only the default.
Eugene Sukhodolskiy
committed
7 hours ago
|
| 2026-10-07 |
config: let server_admin manage synapse hub keys
...
server_admin keeps the whole synapse server now — read, write and admin — so
it can issue and revoke hub API keys and change hub settings without a
detour into the admin panel. All 35 registered tools resolve.
secretary stays on read; the admin group is still withheld from every other
profile.
Eugene Sukhodolskiy
committed
1 day ago
|
config: give server_admin and secretary the synapse MCP tools
...
The synapse server was connecting and registering all 35 of its tools, but
no profile listed it under tools.*.mcp, so build_tool_list never handed a
single one to an agent — the hub was up and unreachable at the same time.
server_admin takes read + write (32 tools: events, deliveries, rules,
sources, targets, routing); secretary takes read only (15). The admin group
— key_issue, key_revoke, settings_put — is deliberately left out of both:
issuing and revoking hub keys stays a manual operation.
Eugene Sukhodolskiy
committed
1 day ago
|
config: run every profile on glm-5.3-flash:cloud first
...
Four chains still led with gemma4:31b-cloud, so those sessions were resolved
onto gemma4 even though glm-5.3-flash:cloud is the instance default. glm now
leads every profile.
The rest of each chain stays behind it as a fallback — that is what carried
today's 16:57 run through a ReadTimeout against ollama.com instead of failing
it. developer, navi_code, discuss and modeler_3d already led with glm and are
untouched.
Eugene Sukhodolskiy
committed
1 day ago
|
config: give tgclient the shared key it needs to handshake
...
The server declares a user_key slot and no shared credential, so the
startup handshake went out unauthenticated, got a 401, and the client
was marked disconnected — which kept its tools out of the registry
entirely, groups or no groups. The shared key now backs the handshake
and stays the fallback: a personal key from /mcp-keys still wins.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: hold the chat at the bottom when a stream ends
...
The list was pinned per streaming delta, but the last things to land arrive
after that final pin: the stats/rating footer, which renders only once
msg.done is set, and the copy buttons attached to code blocks after render.
Nothing re-clamped afterwards — the length watcher never fires (the message
stays in the array) and the landing loop only runs when a session opens — so
the view was left short of the bottom, looking like it had scrolled up.
Re-clamp through the same settle window the landing uses when streaming goes
true -> false, unless the user has scrolled up or a session is loading.
This addresses the late-layout half of the problem; the row is still remounted
when msg.id becomes h_<n>, which is the other source of a jump.
Eugene Sukhodolskiy
committed
1 day ago
|
config: tool groups for tgclient and synapse
...
Both servers declare no groups, so every profile asking for "tgclient":
["read", "write"] resolved to nothing: resolve_group reads the static
config, returned [], and no mcp__tgclient__* name ever reached the agent.
Worse, the server's own instructions did reach it — they are selected by
server name, not by group — so the model was told about tools it did not
have. synapse is not exposed by any profile, so its groups change nothing
today; without them, exposing it would repeat the same failure.
read — observe only.
write — changes to sources, types, targets and routing rules.
admin — hub-wide knobs, not routing: issuing and revoking a source's API
key, and settings overrides on top of .env. The same fence
gnexus-book puts around its service-operator tools.
Eugene Sukhodolskiy
committed
1 day ago
|
deps: declare html2text, which tools/gmail.py imports
...
The import worked on the server only because the package had been
installed into the venv by hand; a fresh sync would have pruned it and
tools/gmail.py would have stopped loading while tools/enabled.json kept
naming gmail. Reload now reports that drift, but the fix is to declare
the dependency. Lock gains html2text and nothing else.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: reload tools from a button in the MCP tab
...
Settings → MCP gains a Tools block for admins only: one button, then the
same report the tool prints — what loaded, how many are in the registry,
per-file errors, and names in enabled.json nothing answers to.
It sits inside the existing MCP tab rather than a new one: the reload
rewrites the toolset of the whole server, not just this user's MCP keys,
and it belongs next to the thing it affects. Non-admins never see it.
dist rebuilt together with the source, as the server serves the bundle.
Eugene Sukhodolskiy
committed
1 day ago
|
webclient: keep MCP rows at content height on a phone
...
.mcp-keys-row stacks into a column below 768px, and .mcp-keys-info kept
its flex: 1 1 240px. That basis is a width in the desktop row and a
height once the row stacks, so every server reserved 240px and its text
sat at the top of the gap — 402px for a row whose content is 90px.
Back to content height on mobile only, and drop the kit's .form-group
bottom margin there: the row's own flex gap already separates the field
from the buttons.
Eugene Sukhodolskiy
committed
1 day ago
|
admin: POST /admin/tools/reload
...
reload_tools is granted to one profile only (tool_developer), so an admin
whose profile lacks it cannot reload at all — the tool answers "not
found" instead of reloading. The route calls the same reload_all() the
tool does, behind require_admin, for whoever is logged in as an admin.
An admin may read: ok, the tools loaded, the registry total, per-file
errors, names in enabled.json nothing answers to, and the MCP/providers
summary.
Eugene Sukhodolskiy
committed
1 day ago
|

reload: pick up the new code, and drop MCP tools that are gone
...
reload_tools reported success while three separate things kept it from
doing what it says.
The bytecode cache is keyed on (mtime in whole seconds, file size), so a
tool edited to the same length inside the same second as its previous
load re-ran the OLD code — the reload was real, the new version was not
live. The loader now compiles the source itself instead of consulting
__pycache__ (importlib.invalidate_caches() does not help here).
MCP registrations only ever grew: register_mcp_tools called
register_external, and unregister_external was used in one place, so a
server removed from the config or a tool a server stopped exposing
stayed in the registry and failed only when the model called it. Reload
now clears external tools and rebuilds them.
enabled.json naming a tool that failed to load (the gmail/html2text case)
was visible only as a log line. It is now part of the report.
The reload itself moves to navi/core/reload.py, one implementation shared
by the tool and the admin route, so the two cannot leave different
toolsets behind. list_tools.py also read enabled.json through its own
cwd-relative path, which made it disagree with the real toolset, and the
class-based loader rejected execute(self, params, ctx=None) — the shape
every built-in uses.
Eugene Sukhodolskiy
committed
1 day ago
|
config: point the tgclient MCP server at the live host
...
The committed value was the local dev address (localhost:8710, where
~/Projects/tgclient-mcp serves it) and the server ran with the wrong public
host, so the tools never connected. Both clones now use
https://tgclientmcp.gnexus.space/mcp.
No default headers: the key is per-user by design (BYOD settings field), and
until one is entered the server answers 401.
Eugene Sukhodolskiy
committed
1 day ago
|
switch_profile: run the switch alone, then the batch on the new tools
...
A tool called in the same batch as switch_profile was still dispatched against
the old tool_map and died with "tool 'X' not found" — reload_tools did, twice
in one session. The batch is now split: the switch runs first, its
ProfileSwitched event is watched for the target profile, the tools are
re-resolved from it, and the remaining calls run against the new set.
The turn's own profile binding is deliberately left alone — the
end-of-iteration reload in run_stream() is what rebinds profile/llm/schemas,
and pre-setting session.profile_id here would make it skip that step.
Eugene Sukhodolskiy
committed
1 day ago
|
switch_profile: deliver profile_switched, report the new toolset
...
The tool took its sink as `ctx.event_sink if ctx else current_event_sink.get()`,
but the agent loop builds tool_ctx with event_sink=None — the ContextVar is the
real channel — so the branch always picked None, the event was silently
dropped, and the profile badge in the header never moved. plan.py already had
the fall-through (`if ctx and ctx.event_sink else current_event_sink.get()`);
this is the same fix.
The result also said nothing about what the switch changed, so the model kept
calling tools the target profile does not have (reload_tools after leaving
tool_developer). It now names the gained and lost tools, and the timing claim
is corrected: the new prompt and tools are in force from the next step of the
same turn, not "from the next message".
Eugene Sukhodolskiy
committed
1 day ago
|

Allocate message sequence numbers in the DB, not in memory
...
save() numbered new messages from Session.db_next_sequence, a counter read
when the session was loaded. Any second writer of the same session handed
out the numbers it still believed were free, and the two inserts collided on
UNIQUE(session_id, sequence_number) — in production, when switch_profile
loaded the session mid-run and saved it back. The turn died with
"Internal error: duplicate key value violates unique constraint".
The range is now claimed inside save()'s transaction with a single
UPDATE ... RETURNING, so concurrent writers serialize on the session row,
and GREATEST() seeds the pre-counter sessions whose next_sequence is still 0
instead of relying on a racy max()+1 fallback in memory.
switch_profile itself no longer saves a session at all: it repoints the row
through a narrow set_profile() UPDATE, which keeps it out of the running
turn's way. It runs mid-turn on a session the turn still holds, so saving a
second copy from there was the collision in the first place.
Eugene Sukhodolskiy
committed
1 day ago
|
config: add the gntodo MCP server
root
committed
1 day ago
|
config: capture the live server configuration
...
Profiles gain the gntodo / gnexus-book / gnexus-creds scopes for agent and
subagent runs, refreshed model lists, and write access where the live setup
has it. MCP server configs pick up their user_key slots, gnexus-book learns
delete_pending_change, and the hard-panel / synapse / tgclient servers join
the tree.
root
committed
1 day ago
|
Merge remote-tracking branch 'origin/master'
root
committed
1 day ago
|
Await the session pool in notify and three neighbours
...
PgSessionStore._get_pool() is async, and four call sites passed its coroutine
straight into a store constructor, so the first query inside died with
"'coroutine' object has no attribute 'fetchrow'": notify always failed, the
reaction runner never got past reading its settings, synapse_instructions was
unusable, and the BYOK resolver caught the AttributeError and silently fell
back to the default credential.
The tests missed it because each one mocked the store or the pool provider
away; the new ones run on a fake asyncpg pool with nothing faked below the tool.
Eugene Sukhodolskiy
committed
1 day ago
|
Merge origin/master (b7743f8): PWA icons, MCP settings tab, android push, runbook
root
committed
1 day ago
|
webclient: serve the PWA artwork past the images cache
...
/images/* is served cache-first from IMAGES_CACHE, which activate deliberately
keeps across builds. That is right for content whose URL is unique and wrong
for the app's own artwork: /images/icon-*, /images/logo-icon* and
/images/apple-splash/* keep their URLs while their contents change with the
logo, so a browser that had once loaded an icon would keep showing the old one
even after a deploy — and the apple-touch-icon is linked from index.html, so it
does travel through the page and the service worker.
Those paths now go network-first with a cached fallback (offline still works);
every other /images/ request is untouched.
Tests: frontend 148 passed; backend 1367 passed, 1 skipped.
Eugene Sukhodolskiy
committed
1 day ago
|

webclient: draw the PWA icons at full size again
...
The maskable icons and the apple-touch-icon carried the mark at 21% of the
canvas — the artwork scaled down and pasted in the centre — so on a home
screen the logo read as a fragment of itself. The mark takes 73.4% of the
canvas in logo.svg and in the launcher tile of the Android app icon; that is
the proportion all five files use now.
scripts/gen_pwa_icons.py redraws the mark from logo.svg's geometry with
Pillow (already a project dependency, no SVG rasterizer needed) and writes the
whole set, so the scale lives in one constant; --check measures what is on
disk and reports SUSPECT if it drifts again. Two runs produce identical bytes.
The regenerated icon-192/512 are geometrically identical to the rsvg-rendered
ones they replace: ink bbox 376x376 with 68px insets at 50% coverage, and the
same ink mass across the stroke. Only the antialiasing bytes differ.
dist/ rebuilt (it carries a copy of public/ and is served by the backend).
Tests: frontend 148 passed; backend 1367 passed, 1 skipped.
Eugene Sukhodolskiy
committed
1 day ago
|
deploy: runbook for updating a running instance
...
Adds deploy/UPDATE.md: preflight, the branch shapes (master vs the
tracked .env on deploy), what to rebuild before pushing (the frontend
dist is committed and the server never builds it), the verification
commands that prove the new bundle is the one being served, rollback,
and the traps — tracked secrets in mcp_servers.d/, gitignored vendor
kit dist, service-worker cache, Android app needing only a restart.
Linked from deploy/README.md.
Eugene Sukhodolskiy
committed
1 day ago
|

MCP settings tab: list every connected server, slot only where declared
...
The tab was empty on every install: it listed only servers whose config
declares a `user_key` slot, and no config declared one — which read as
"no MCP servers connected" even though five are wired to profiles.
- GET /mcp-keys now returns every server referenced by at least one
profile, keyed ones first, with `accepts_user_key`, the slot location
(null when there is none) and the profile ids that connect it. The
per-user key store is skipped entirely when nothing has a slot.
- gnexus-creds declares `user_key: {header: Authorization, prefix:
"Bearer "}` — it is the one server carrying a shared credential, so its
personal-key field is now real: users with a key run under their own,
users without one fall back to the shared default.
- The panel lists all servers (transport + profiles), dims the keyless
rows, and shows a key input only for slotted ones, spelling out the
shared-key fallback.
docs/api.md and docs/mcp.md updated; backend 1367 passed, webclient 148.
Eugene Sukhodolskiy
committed
1 day ago
|