| 2026-10-08 |

mcp: credentials leave the tree, and sub-agents lose the task board
...
Five tracked configs (gnexus-creds, gntodo, hard-panel, synapse, tgclient)
carried live bearer tokens — in git, in every clone, and handed to the admin
client by GET /admin/mcp/config. They now name a variable, and the value lives
in the service .env (chmod 600, untracked).
Substitution happens once, at transport open, next to the project-relative
path resolution: that is the single place a transport is built, so every
stored and serialised config stays placeholder-only and save_mcp_servers —
reached by create_mcp_server and PUT /admin/mcp/config — cannot write a secret
back into a tracked file. A variable that is not set refuses the connection
instead of dropping the header, which would fail open: a server may answer an
unauthenticated request as an anonymous user rather than returning 401.
Sub-agents also drop gntodo, gnexus-creds, tgclient and synapse from their
scopes, and the runner injects a scope anchor unconditionally — a sub-agent
asked to echo one word had picked a project off the task board and researched
it for forty iterations.
Docs: docs/mcp.md#secrets, a migration recipe in deploy/UPDATE.md, the variable
names in .env.example and deploy/env.template, and a regression guard over the
tracked configs.
The old values stay in git history; rotating all five at their sources is the
remediation, not a follow-up.
Eugene Sukhodolskiy
committed
6 hours ago
|

subagents, secrets: an honest result contract, and two leaks closed
...
A pass over the sub-agent system, plus the two findings from the previous one.
Leaks:
- `redact_args` / `is_sensitive_key` (navi/tools/_internal/redact.py) mask
credential-shaped keys at every log site that prints tool arguments — by key
name, for all tools, so a future MCP tool is covered without declaring
anything. ssh_exec passwords were sitting in the prod journal in plaintext.
- `mcp_servers.d/*.json` are tracked in git and carried the scaffolding
machine's absolute paths, so navi-3d/navi-web silently vanished from the
toolset off that machine. They now hold project-relative paths, resolved in
`McpClient.open_transport` — not at load time, because create_mcp_server
round-trips every config file through save_mcp_servers.
Sub-agent result contract:
- `run_ephemeral` returns a `SubAgentOutcome` (text + status: ok, timeout,
max_iterations, thinking_stall, user_stop, context_overflow) instead of
`(text, bool)`. Every exit routes through `SubAgentRunner._finish`, so token
accounting is never dropped and a partial run always carries a progress
report — the stop-after-stream path used to return a bare string.
- `spawn_agent` renders its header from the status. Before this, every short
run reached the parent as "hit iteration limit", and the parent repeats that
diagnosis to the user — a timeout was reported as an iteration limit.
- `ContextTooLargeError` no longer kills a sub-agent: the parent compresses its
own context, a sub-agent has no session to compress into, so it stops with
`context_overflow` and hands back what it did rather than surfacing as
"Sub-agent failed: …" with the whole run discarded.
- The result is capped at 8000 chars (head and tail kept) so a long run cannot
flood the parent's context, and tool arguments in the progress report go
through `redact_args`.
- The navi-3d calling convention left the runner for the server's own tool
schemas — a generic runner should not know one MCP server's path shapes.
- Validation moved inside spawn_agent's try: a bad argument is a clean
ToolResult the model can read, not an exception in the parent's loop.
Tests: 1636 passed, 1 skipped. New: SubAgentOutcome paths and headers, result
truncation, context overflow, clean argument failures.
Eugene Sukhodolskiy
committed
8 hours ago
|

profiles: tool_developer folds into developer
...
The profile was a duplicate on every axis we could measure. 22 of its 26
native tools were already developer's; the four it alone held —
reload_tools, create_mcp_server, test_mcp_tool, mcp_status — are 2.9 KB of
schema. Its model chain was the same seven models. Of its 14 KB prompt,
about 8 KB was copied verbatim from developer's (the whole `## Orchestration
model` block and everything from `## Editing policy` down), and most of the
remainder restated manuals/create_mcp_server.md, which already carried the
same ten-step workflow in more detail. It was not a specialisation, it was a
snapshot: `git log -S '"reload_tools"'` shows the tool lived in developer
until 61fa370 rewrote that profile around MCP and cut it off.
What kept it alive was a premise that no longer holds — that Navi's own
capabilities would be written as in-repo tools. They are MCP servers now,
and an MCP server is not a file in this repository with a life of its own:
it is an isolated process registered from mcp_servers.d/. So there is no
reason left for a profile whose only distinct feature is a toolset a general
developer profile can hold, and every reason to stop maintaining a second
prompt that drifts against the first.
- developer: + reload_tools, create_mcp_server, test_mcp_tool, mcp_status
(24 → 28 native). Its sub-agent gains tool_manual, which is what it
actually needed to reach the manual while writing a server — the previous
tool_developer sub-agent had it, developer's did not.
- server_admin: + reload_tools only (24 → 25). Adding a third-party MCP
server is something this profile does as often as developer does.
Deliberately not to its sub-agent: reconnecting the MCP manager is a
process-wide operation belonging to the main agent.
- The prompt and the manual took on what the deleted profile knew and
create_mcp_server.md did not: reload_tools before the first test_mcp_tool
(a freshly registered server is not connected, so the test fails and the
iteration is wasted), the smoke test read by exit code — 124 means timeout
killed a server still running, 0 means it exited on its own, usually a
main() without parentheses — absolute command/cwd, mcp_status as discovery
only, and the steps that stay inline instead of going to a sub-agent.
mcp_status and test_mcp_tool were built without an MCP manager, and their
fallback did `from navi.api.deps import _mcp_manager` — a name that does not
exist, so a live call raised ImportError rather than the intended "MCP
manager not available". The tools always passed a manager in tests, which is
why nothing caught it. Both now receive the manager at construction and fall
back to the live one lazily.
Sessions and profile_overrides are reassigned before the restart: agent.py
resolves the session's profile without a guard, so a deleted profile turns
every session that referenced it into an uncaught ProfileNotFound. Nothing
in the test suite pins the profile inventory, and profiles are read once at
import time — reload_tools does not re-read them — so this ships as a
restart, and the restart is also what makes it take effect.
Eugene Sukhodolskiy
committed
9 hours ago
|

tool_manual: a manual the agent cannot reach is not a manual
...
tool_manual asked the filesystem for the exact string the model sent, then
fell back to registry.get(). The executor already tolerated the ways models
mangle MCP names; the tool that documents them did not, so the agent calling
mcp__navi-3d__compile_scad got a schema dump while manuals/compile_scad.md sat
there unread. Resolving a name now happens in one place — resolve_tool moved
out of the executor into navi/core/tool_utils.py — so a name that works when
the tool is *called* also works when it is *asked about*. A miss suggests the
closest names instead of "not found", and a tool that exists but is not enabled
for this profile says exactly that: documenting it is still the useful answer,
but the agent must not go on to call it. params["tool_name"] was also a crash
waiting to happen — KeyError with the key omitted, AttributeError on null or a
number — and now anything that is not a usable string means "no name given",
which returns the index of manuals grouped by source.
The generated manual now renders the whole schema. An array of objects showed
as "(array, optional)" with its item shape invisible, so a model that could not
see the fields guessed them; objects, array items, oneOf/anyOf branches, enums
and defaults are all spelled out now, depth-capped. The header names the source
(native / mcp__<server>__) and says the text is a parameter contract, not a
curated manual — otherwise a thin schema reads as "this tool is simple".
Then the manuals themselves. A manual is filed under the tool's own name, and
five of the sixteen were not: write_tool.md and write_mcp_server.md documented
tools that do not exist anywhere (the tools are reload_tools and
create_mcp_server), model_3d.md documented compile_scad, render_3d.md documented
render_stl, and write_context_provider.md documented no tool at all — it is a
guide, and now lives in manuals/guides/. The prompts pointed at the same wrong
names (tool_developer's system prompt, persona_navi_code, docs/context_providers),
so fixing the files without the callers would have moved the breakage rather
than removed it. On HEAD the probe was blunt: tool_manual("mcp__navi-3d__compile_scad")
returned None and the agent got the schema; now it returns the manual.
Nothing was enforcing any of this, which is how it drifted. tests/unit/tools/
test_manual_drift.py reads the real repository tree: every manuals/*.md must be
named after a tool that exists (built-ins and tools/*.py by their name
assignments, MCP tools by @mcp.tool(name=…) in the server sources and by the
mcp_servers.d groups, whose file name is the server name), every manual must be
reachable by both its bare name and its full MCP spelling, a guide must never
shadow a tool's name, and every manuals/<x>.md or tool_manual("x") cited in a
doc or a prompt must resolve. The next rename fails the suite instead of
quietly wasting a manual.
Finally, manuals for the built-ins the agent uses most — ssh_exec, todo, memory,
plan, filesystem, notify, code_exec — written from the current schemas and the
code paths that produce the errors, so the gotchas are the real ones (todo's
`done` validation is enforced; memory's `list` returns categories, not facts;
filesystem's `delete` removes a tree with no prompt). reload_tools.md absorbs
the file format from the deleted write_tool.md, which is where a self-extension
recipe belongs now.
1556 passed, 1 skipped. ruff delta zero.
Eugene Sukhodolskiy
committed
10 hours ago
|

startup, llm, profiles, todo: five more the prod logs gave up
...
The same 5-day sweep of prod's journal as 5e51f87, continued. Every one of
these failed silently: no error in the UI, just a line in the log or a hole
in the database.
complete()/embed() retry a connection error before falling back.
stream_complete has always given its first chunk two attempts, but the
non-streaming path made exactly one per (server, model), and with a single
configured server (ollama.com) one ReadTimeout raised "All backends
exhausted: ReadTimeout" outright. That is the whole story of 2026-10-08,
when 27 calls died that way in a day — 23 of them memory summarisation, the
rest planning. Attempts now come from settings.llm_complete_retries (2) with
llm_retry_backoff_sec (2.0) between them; a missing model is still not
retried (a 404 is an answer, not a hiccup), and a lone server is still never
blacklisted — blacklisting it would block the next request for _TTL.
Startup waits for Postgres. create_container() opens the pool eagerly and
the retry loop for the DDL tables sits below it, so a host that comes up
before docker takes the whole lifespan with it. On 2026-10-07 15:19:40 the
agent ran `sudo reboot` on itself at the user's request ("Перезагрузи
себя"); the host booted 15:20:03, navi started at 15:20:18, hit
ConnectionRefusedError on the pool and logged "Application startup failed.
Exiting." Connection errors now retry for ~30s; a genuinely broken config
(a missing DATABASE_URL) still fails fast.
The profile loader reports its tally. A profile dropped by an unreadable
config.json is skipped by design so it cannot take the server down, but on
that same 15:20 boot four of them — modeler_3d, navi_code, secretary,
server_admin — were dropped with `Extra data: line 125 column 1`, and the
only trace was one error line each: the UI simply showed four profiles
fewer.
todo recovers a dropped `op`. Of the 3385 todo calls stored on prod, 73
arrived with the discriminator flattened away ({"index": 1, "status":
"in_progress"}, {"action": "view"}, {"": "add", "tasks": [...]}) and got a
bare "Unknown op: None" — a wasted round trip for an intent the remaining
arguments state plainly. The error message now names the five ops and what
each needs.
Message.created_at is stamped at creation. It defaulted to None and only
some construction sites filled it, so 85% of stored rows had no timestamp at
all (every tool message, 14 168 of 14 168) and no query could slice history
by time or measure a pause. Loading a row passes the column through
explicitly, so NULL rows stay NULL instead of being restamped with "now".
Every new test was checked to fail with its fix reverted. Full suite green
(1466 passed, 1 skipped). No frontend or dependency change, so the deploy is
a pull and a restart.
Eugene Sukhodolskiy
committed
11 hours ago
|
| 2026-10-07 |
config: let server_admin manage synapse hub keys
...
server_admin keeps the whole synapse server now — read, write and admin — so
it can issue and revoke hub API keys and change hub settings without a
detour into the admin panel. All 35 registered tools resolve.
secretary stays on read; the admin group is still withheld from every other
profile.
Eugene Sukhodolskiy
committed
1 day ago
|
config: give server_admin and secretary the synapse MCP tools
...
The synapse server was connecting and registering all 35 of its tools, but
no profile listed it under tools.*.mcp, so build_tool_list never handed a
single one to an agent — the hub was up and unreachable at the same time.
server_admin takes read + write (32 tools: events, deliveries, rules,
sources, targets, routing); secretary takes read only (15). The admin group
— key_issue, key_revoke, settings_put — is deliberately left out of both:
issuing and revoking hub keys stays a manual operation.
Eugene Sukhodolskiy
committed
1 day ago
|
config: run every profile on glm-5.3-flash:cloud first
...
Four chains still led with gemma4:31b-cloud, so those sessions were resolved
onto gemma4 even though glm-5.3-flash:cloud is the instance default. glm now
leads every profile.
The rest of each chain stays behind it as a fallback — that is what carried
today's 16:57 run through a ReadTimeout against ollama.com instead of failing
it. developer, navi_code, discuss and modeler_3d already led with glm and are
untouched.
Eugene Sukhodolskiy
committed
1 day ago
|
config: capture the live server configuration
...
Profiles gain the gntodo / gnexus-book / gnexus-creds scopes for agent and
subagent runs, refreshed model lists, and write access where the live setup
has it. MCP server configs pick up their user_key slots, gnexus-book learns
delete_pending_change, and the hard-panel / synapse / tgclient servers join
the tree.
root
committed
1 day ago
|
Merge remote-tracking branch 'origin/master'
...
# Conflicts:
# webclient/vendor/gnexus-ui-kit/package-lock.json
root
committed
1 day ago
|
| 2026-10-06 |
synapse reactions W3: reaction runner, notify tool, gateway hookup
...
- navi/synapse/reactions.py: fire-and-forget reaction run — per-user
reactions_enabled gate, one-shot dispatcher meta-pass (visible profiles
only, hidden dispatcher, secretary fallback), special=True session
headed by the resolved profile, headless agent run via the orchestrator
run registry, finalise: app push per completion_notify +
low-level reaction.finished/failed event to Synapse
- navi/synapse/outbound.py: source-side emit (syn_* key from env,
silent when not configured)
- navi/tools/notify.py: notify tool — message + level (info/warning/
intervention), both legs per push_target, honest delivered/skipped result
- push/service.py: notify_custom without cooldown
- gateway: verified + non-duplicate delivery schedules the reaction
- profiles: notify added to native tools (7 profiles)
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W2: hidden dispatcher profile + synapse_instructions tool
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-09-29 |
Merge remote-tracking branch 'origin/master'
root
committed
9 days ago
|
| 2026-09-26 |
e2e fixes: tasks tool in profiles, bg timeout lift, stepIcon fix, queue semantics docs
...
E2E findings addressed:
- tasks tool was registered but not in any profile's tools.agent.native —
added to all six profiles (agent could not check/wait/cancel bg tasks)
- detached terminal/code_exec/ssh_exec runs without explicit timeout are
lifted to 300s: foreground defaults (20/30/60s) marked long commands
'completed' with partial output while the process still ran
- ToolCard.vue: define stepIcon(status) — template referenced it but the
function was missing (render crash on task_update step)
- message_queued reachability documented: WS read loop is sequential, the
queue path is only reachable from a second socket/headless recall
Eugene Sukhodolskiy
committed
12 days ago
|
agent parallelism: background tools, bg spawn_agent, parallel tool batches, message queue
...
- backgroundable tools (terminal/ssh_exec/peer/spawn_agent/code_exec) detach
via TaskManager with per-session/global/spawn caps, rate limit and TTL
- completion delivery both ways: task_update event (out-of-band) + pending
notes injected into the next turn; tasks tool (list/check/wait/cancel)
- parallel tool-call batches (profile-level gate, off by default) with
ToolStarted up-front, tagged event mux, call-order results, one save,
plus dangling tool_call repair on session load
- user message queue instead of busy error: message_queued frame,
back-to-back drain on the same socket, headless fallback on disconnect
- pin mcp<2 (v2 renames FastMCP with breaking API changes)
- docs: tasks.md (new), websocket/api/agent/config updates, tasks manual,
spawn_agent background param, persona contract section
Eugene Sukhodolskiy
committed
12 days ago
|
| 2026-09-10 |
Merge branch 'master' of https://git.gnexus.space/git/root/navi-1
root
committed
29 days ago
|
.
root
committed
29 days ago
|
swarm stage 2: peer-to-peer channel — /peer endpoints + peer tool
...
Server side (navi/api/routes/peer.py, port 8099):
- GET /peer/hello — open discovery ping (name, uuid prefix, version)
- GET /peer/status — PSK; identity, uptime, machine facts, hive view
- POST /peer/ask — PSK; one-shot agent run under PEER_ASK_PROFILE
(default server_admin) answers, one concurrent ask at a time
- loop guard: answering agent runs without the peer tool (deterministic
recursion cut) + own-uuid asks refused with 409
- verify_swarm_key in navi/swarm.py accepts .swarm-key.previous
(rotation window, constant-time)
Client side (navi/tools/peer.py):
- peer tool: list (hive book, stale cache fallback marked), status,
ask by swarm name; asks go direct peer-to-peer, hive never on the
message path
- audit events on both sides (peer.ask_sent/received/answered/failed)
peer tool added to server_admin, navi_code, developer profiles.
Eugene Sukhodolskiy
committed
29 days ago
|
| 2026-09-09 |
.
root
committed
29 days ago
|
profiles: retire gate flags and retrain prompts for the plan tool
...
- drop planning_enabled / planning_mandatory / planning_phase2_enabled /
observe_skips_plan_enabled / adaptive_replan_enabled from AgentProfile,
the loader and admin serialization
- add `plan` to native tools of developer, navi_code, tool_developer,
secretary, server_admin, modeler_3d — without the gate they would
otherwise lose planning entirely
- system prompts: call plan for non-trivial multi-step work before
execution; complex plans get presented for confirmation; stuck or goal
changed -> plan with a reason
Eugene Sukhodolskiy
committed
29 days ago
|
Merge branch 'master' into feature/navi-code
...
# Conflicts:
# navi/api/websocket.py
# navi/config.py
# navi/core/agent.py
# navi/core/orchestrator.py
# navi/main.py
Eugene Sukhodolskiy
committed
29 days ago
|
navi_code: add gemma4 31b/26b qat models to the profile fallback chain
Eugene Sukhodolskiy
committed
29 days ago
|
| 2026-07-14 |

navi_code: add web lookup, ssh, and image_view — grow beyond local-only
...
navi_code was strictly local: no web, no remote hosts, no image viewing. That
made it lean on the developer profile whenever a coding task needed a doc
lookup or a remote command. Bring those capabilities in so navi_code covers
full coding work itself, while staying terminal-first and local-by-default.
config: agent gains image_view, ssh_exec, and the navi-web MCP (search/browse/
request); subagent gains image_view + navi-web (no ssh_exec — sub-agents don't
reach remote hosts, matching developer). Description/short_description/
full_description updated to "terminal-first with web and remote access". Also
includes the gemma4:12b-it-qat-128k model in the model list (local 128k-context
option alongside 31b-cloud).
prompt: soften "Local-only, no remote hosts" to "local by default, with web
lookup (docs/APIs) and ssh (remote ops) when the task needs it". Add compact
Web lookup and Remote access sections under Execution environment, and update
the sub-agent toolset briefing to list image_view + web and to exclude ssh_exec.
scope_boundary still applies; ssh destructive/system-wide actions need user
confirmation. Kept additions minimal to avoid bloating the prompt for the 12b
model.
Eugene Sukhodolskiy
committed
on 14 Jul
|

navi_code: stop delegating code work to the developer profile
...
navi_code was systematically spawning sub-agents as `developer` instead of
itself, even though it already covers local code work. Root cause was a
three-channel nudge: the navi_code prompt recommended `developer` for "general
code work" and framed omit (-> navi_code) as the exception; spawn_agent's
description listed `developer` as the coding example and never mentioned
navi_code; the injected Available-profiles list shows developer's broad
"General-purpose software development" blurb. The planner locks profile_id at
plan time on those same signals, so the bias is baked in before execution.
Flip the default in all three places: for code work OMIT profile_id so the
sub-agent runs as the current profile (navi_code); set profile_id only for a
different specialisation (secretary/server_admin/tool_developer). developer is
not removed as a profile — it is just no longer recommended from navi_code.
Eugene Sukhodolskiy
committed
on 14 Jul
|

compress: raise budget 6000 + richer prompt detail + input headroom
...
Increase the compression summary budget and push the prompt toward more
detail and completeness, so the agent keeps more durable facts when old turns
are folded into a summary.
Budget (output):
- context_summary_max_tokens 4000 → 6000 (global default).
- navi_code compression_max_tokens 4000 → 6000 (profile override).
Prompt (base, all profiles):
- New 'Intermediate Findings' section: durable facts from read/grep/log/terminal
that informed decisions (was forbidden by 'do not preserve intermediate
reasoning'). Existing sections get richer instructions (Active Files: key
symbol affected; Decisions: rationale + trigger; Completed: step + file +
verification outcome; Errors: snippet + what was tried + outcome).
- Loosen 'tight bullet points' → 'specific and complete; prefer listing over
generalizing — losing a fact costs more than a longer summary'.
- 'Do not write implementation code' → 'preserve short critical code verbatim
(final signatures, one-line fixes, key config lines); drop long blocks'.
Input headroom (so the summarizer can SEE what it must preserve):
- Non-critical tool result preview 300 → 800 chars.
- _MAX_SUMMARY_INPUT_CHARS 24000 → 32000 (fits 65k–128k windows with 6k output).
Profile prompt (navi_code/compression_prompt.txt): preserve short diagnostic
snippets verbatim (exact error line, failing assertion, one-line fix) instead
of blanket 'do not preserve long output'; keep code signatures + short key
snippets, not just names.
Tests: Intermediate Findings section in base prompt; 'short critical code'
phrasing; non-critical preview keeps 800 chars; meta-summary threshold raised
(2×6000 chars). Fixed flaky test_attach_session_seeds_context_fill (relied on
_startup worker timing after Etap 4 lengthened attach; now explicit attach
like its sibling + mock get_todos to skip a 30s×2 real-network timeout).
Full suite: 1008 passed, 1 skipped.
Eugene Sukhodolskiy
committed
on 14 Jul
|

todo: add op — append steps discovered mid-task (preserve statuses)
...
The todo tool only had set/view/update/clear. Adding a step that surfaced
mid-task meant either 'set' (which replaces the plan and resets every status
to pending) or 'replan' (1–3 LLM calls) — both heavy for a single new step, so
the agent usually skipped recording it. Add a cheap 'add' op:
- navi/tools/todo.py: op 'add' (tasks: list[str]) appends new _Task entries to
the end of the existing plan in pending status; existing steps and their
statuses are untouched. Requires an existing plan (points to 'set'
otherwise). Schema enum + op description + tool description updated.
- system_prompt.txt: tell the agent to record new subtasks with 'todo add'
right away (not held in head, not a replan); clarify the 'small todo edit'
line — add via 'todo add', drop/merge/reorder via 'set' (re-apply statuses).
- renderers/todo.py: TodoStartedRenderer renders the add card
('→ todo · add (N)') listing the new pending steps, not a JSON dump.
Tests: add appends + preserves statuses / requires existing plan / requires
tasks; renderer add card + empty-tasks. Full suite: 1005 passed, 1 skipped.
Eugene Sukhodolskiy
committed
on 14 Jul
|
| 2026-07-13 |

Автономность мелких моделей: OUTPUT DISCIPLINE, milestone-todo, перехват финала хода
...
Два последовательных захода над одной проблемой — мелкие модели (12–30B) в navi_code
теряют автономность на трудных шагах: «остановился поболтать» вместо действия и
слишком большое расстояние между пунктами плана.
Заход 1 (З1–З3):
- З1: OUTPUT DISCIPLINE в системном промпте navi_code — «act, don't announce»,
без few-shot антипримеров. Контракт хода: ответ без tool_calls = конец хода, поэтому
объявление намерения текстом убивает автономность; правила заставляют вызывать
инструмент в том же ходе.
- З2: плоский todo + метка группы milestone + декомпозиция. _parse_plan_steps
возвращает list[tuple[milestone, text]]; milestone — метка группировки (не сущность,
без статуса), «done» вычисляется при рендеринге; подшаги = больше плоских шагов
(без вложенности). TUI side-panel группирует по milestone (плоский фолбэк при пустом
milestone). Plan depth: max 15→20 + правило декомпозиции.
- З3: adaptive re-plan «длинный шаг» — nudge «разбей шаг» при in_progress ≥ порога
итераций без смены todo (порог 4, раньше общего anti-stall warning на 8).
Заход 2 (шаги 1–3, после cloud-теста 31b vs 12b):
Корневая структурная причина: весь спасательный механизм (anti-stall warning с явным
предложением reflect, adaptive re-plan) живёт только внутри tool-цикла — nudge
инжектируется в pre_turn *следующей* итерации, которой при «остановился поболтать»
нет (ход закрылся по return до post_turn). 31b застревала через «продолжаю
tool-итерации» → дожала до warning → спаслась; 12b — через «замолчала текстом» → мимо
всех nudge.
- Шаг 1: перехват финала хода. Если модель выдала bare-text, но в todo есть открытые
шаги (pending/in_progress) и лимит не исчерпан — НЕ эмитить StreamEnd, а сохранить
ассистентский текст в session.context, поставить системный nudge и continue
(без StreamEnd, без workers — консистентно с multi-iteration tool-турами). Счётчик
final_interceptions на AgentTurnContext, лимит final_intercept_limit (default 2),
эскалация жёсткости (мягкий → «second stop»). has_open_steps в todo.py: пустой
todo → False (защита casual-сообщений), failed/skipped терминальны. Профильные
флаги final_intercept_enabled/limit в base.py + loader + admin.
- Шаг 2: жёсткий reflect-триггер в промпте — «~3 tool attempts on the same step
without progress → call reflect IN THIS TURN (tool call, not reasoning aloud)».
- Шаг 3: открыть replan для застревания — «call replan when reflect showed the whole
approach is dead (not one failed step, but the approach itself)».
Тесты: 874 passed, 1 skipped. Новые — has_open_steps (5), final intercept (5),
milestone-группировка, adaptive long-step nudge, парсер шагов с milestone-маркером.
cloud: navi_code model → gemma4:31b-cloud для тестирования догадки (31b признала
застревание, 12b — нет); .env cloud-host уже gitignored.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 13 Jul
|
| 2026-07-12 |

permissions: remove broken client-side gate; add backend design doc
...
The terminal-TUI permission gate was a race: _execute_tools_with_sink
emits ToolStarted and starts the tool on the backend immediately, so by
the time the client dialog appeared the destructive action had already
run -- 'deny' could only stop the session, not un-execute the tool. It
also existed only in the terminal TUI (webclient/android had nothing),
and the system-prompt 'Strict Confirmation' nudge was an unreliable
duplicate.
Remove the broken machinery for a clean slate and plan the real
replacement from zero:
- delete clients/terminal/tui/permissions.py, screens/permission_dialog.py
and their tests
- strip tui_app.py (engine, _deny_tool, _show_permission_dialog,
_confirm_shell_command, _stop_session_worker, on_permission_request,
tool_started gate) and the PermissionRequest TUI event; !cmd now runs
directly (user-typed, no agent gate)
- drop the 'Strict Confirmation' prompt rule from navi_code/developer/
tool_developer
- add docs/permissions.md: authoritative backend gate design (engine +
registry + event + endpoint + agent gate + postgres policy + global
rules config, sub-agents gated) -- design only, not yet implemented
- update docs/index.md, navi_code_cli.md, profiles.md, testing.md
Until Phase 1 lands there is no destructive-action confirmation -- a
conscious period without the false security of the old race-gate.
Eugene Sukhodolskiy
committed
on 12 Jul
|
replan: integrated mid-task re-planning tool
...
Add a replan tool that re-runs the planner over the live session context
+ todo + scratchpad when the plan's structure is stale due to discoveries
(NOT failed steps -- that stays [Adaptive re-plan]). Integrated approach:
PlanningEngine.run gains is_replan/replan_context (suppresses DIRECT
shortcut and observe-skip, frames Phase 1 as a revision); ReplanRunner
packs reason/goal/todo/findings/errors and captures PlanReady; the tool
is exposed to navi_code/developer/tool_developer via a
current_replan_runner ContextVar set per-iteration in run_stream (correct
after switch_profile). New plan replaces the todo. Lazy events import
breaks the navi.tools -> navi.core -> navi.tools cycle.
Eugene Sukhodolskiy
committed
on 12 Jul
|

profiles: port context-org/planning instructions to developer & tool_developer
...
Port the context-organization and planning-machinery framing accumulated
in navi_code to the developer and tool_developer profiles, adapted per
profile:
- Working state & memory (todo/scratchpad/context_transfer/schedule_recall/
reflect/memory) replaces the stale "Context drift recovery" section.
- System signals subsection names the runtime-injected messages each profile
actually receives: [Goal anchor], [Anti-stall warning], [Adaptive re-plan]
(adaptive_replan is on for both), [Iteration N/M]. [Scope boundary] is
omitted — scope_boundary_enabled is off on these profiles.
- Reading & searching, Editing policy (tool_developer), Safety Rules, Git
discipline, and Project environments (isolated venv) added.
- developer: Workflow rewritten planner-aware with an observe carve-out
(observe_skips_plan now on); Project knowledge replaced by docs-first
Documentation; sub-agent briefing gains context_transfer + restricted-toolset
bullets.
- tool_developer: keeps its MCP-specific 10-step workflow and prerequisites;
sub-agent briefing gains context_transfer + restricted-toolset bullets
(sub-agent cannot reload_tools/test_mcp_tool/mcp_status — run those inline).
- config: observe_skips_plan_enabled=true on both, so observe requests skip
phase3 (no plan/todo for read/explain/inspect) — enables the Workflow
observe carve-out and saves an LLM call on info requests.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 12 Jul
|