| 2026-10-09 |

profiles: restricted profiles for ordinary users, and a role gate that holds
...
`is_admin_only` was checked in one place out of nine and was not read from
config.json at all, so all seven profiles were reachable by every account with
role `user`. The flag now lives in config.json — the file is the baseline, a
`profile_overrides` row still wins on top of it — and one predicate,
`admin_only_blocked`, is the single place the rule is expressed. The nine
surfaces that list, switch to, spawn or resolve a profile all consult it:
`POST /sessions`, the WebSocket, switch_profile, list_profiles, the system
prompt's "Available profiles" block, spawn_agent and the Synapse reaction
runner. The prompt cache is now keyed by (profile, role), so a user's prompt
can never be served an admin's profile list.
The seven existing profiles (developer, discuss, dispatcher, modeler_3d,
navi_code, secretary, server_admin) are marked admin-only. Three new ones take
their place for ordinary users: assistant, designer_3d and coder. They share one
native tool set — ssh_exec, peer, reload_tools, create_mcp_server, test_mcp_tool,
image_view and gmail are withheld — and differ only in system prompt, model and
MCP groups. navi-web's raw `request` group, and the whole of gnexus-creds and
tgclient, are withheld too.
MCP per-user keys gain the missing half of the rule: a server that declares a
`user_key` slot is refused to anyone but an admin who has no personal key, and
is left out of their tool list entirely, instead of quietly falling back to the
owner's credential and appearing as a tool that cannot work. The refusal names
the server and points at Settings.
Also closes `GET /agents/prompts`, which served every profile's system prompt to
anyone, with no user dependency at all.
The accepted residual risk is written down in docs/profiles.md: the working
directory is a convention, not a sandbox.
Eugene Sukhodolskiy
committed
4 hours ago
|
| 2026-10-08 |

subagents, secrets: an honest result contract, and two leaks closed
...
A pass over the sub-agent system, plus the two findings from the previous one.
Leaks:
- `redact_args` / `is_sensitive_key` (navi/tools/_internal/redact.py) mask
credential-shaped keys at every log site that prints tool arguments — by key
name, for all tools, so a future MCP tool is covered without declaring
anything. ssh_exec passwords were sitting in the prod journal in plaintext.
- `mcp_servers.d/*.json` are tracked in git and carried the scaffolding
machine's absolute paths, so navi-3d/navi-web silently vanished from the
toolset off that machine. They now hold project-relative paths, resolved in
`McpClient.open_transport` — not at load time, because create_mcp_server
round-trips every config file through save_mcp_servers.
Sub-agent result contract:
- `run_ephemeral` returns a `SubAgentOutcome` (text + status: ok, timeout,
max_iterations, thinking_stall, user_stop, context_overflow) instead of
`(text, bool)`. Every exit routes through `SubAgentRunner._finish`, so token
accounting is never dropped and a partial run always carries a progress
report — the stop-after-stream path used to return a bare string.
- `spawn_agent` renders its header from the status. Before this, every short
run reached the parent as "hit iteration limit", and the parent repeats that
diagnosis to the user — a timeout was reported as an iteration limit.
- `ContextTooLargeError` no longer kills a sub-agent: the parent compresses its
own context, a sub-agent has no session to compress into, so it stops with
`context_overflow` and hands back what it did rather than surfacing as
"Sub-agent failed: …" with the whole run discarded.
- The result is capped at 8000 chars (head and tail kept) so a long run cannot
flood the parent's context, and tool arguments in the progress report go
through `redact_args`.
- The navi-3d calling convention left the runner for the server's own tool
schemas — a generic runner should not know one MCP server's path shapes.
- Validation moved inside spawn_agent's try: a bad argument is a clean
ToolResult the model can read, not an exception in the parent's loop.
Tests: 1636 passed, 1 skipped. New: SubAgentOutcome paths and headers, result
truncation, context overflow, clean argument failures.
Eugene Sukhodolskiy
committed
8 hours ago
|
| 2026-10-07 |
admin: POST /admin/tools/reload
...
reload_tools is granted to one profile only (tool_developer), so an admin
whose profile lacks it cannot reload at all — the tool answers "not
found" instead of reloading. The route calls the same reload_all() the
tool does, behind require_admin, for whoever is logged in as an admin.
An admin may read: ok, the tools loaded, the registry total, per-file
errors, names in enabled.json nothing answers to, and the MCP/providers
summary.
Eugene Sukhodolskiy
committed
1 day ago
|

MCP settings tab: list every connected server, slot only where declared
...
The tab was empty on every install: it listed only servers whose config
declares a `user_key` slot, and no config declared one — which read as
"no MCP servers connected" even though five are wired to profiles.
- GET /mcp-keys now returns every server referenced by at least one
profile, keyed ones first, with `accepts_user_key`, the slot location
(null when there is none) and the profile ids that connect it. The
per-user key store is skipped entirely when nothing has a slot.
- gnexus-creds declares `user_key: {header: Authorization, prefix:
"Bearer "}` — it is the one server carrying a shared credential, so its
personal-key field is now real: users with a key run under their own,
users without one fall back to the shared default.
- The panel lists all servers (transport + profiles), dims the keyless
rows, and shows a key input only for slotted ones, spelling out the
shared-key fallback.
docs/api.md and docs/mcp.md updated; backend 1367 passed, webclient 148.
Eugene Sukhodolskiy
committed
1 day ago
|
mcp keys W1: user_key config slot, encrypted keystore, /mcp-keys REST
Eugene Sukhodolskiy
committed
1 day ago
|
| 2026-10-06 |
synapse settings: source_ready flag in GET/PUT response
...
Tells the UI whether Synapse-linked options are usable (source key
configured); not persisted, computed per request.
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W3: reaction runner, notify tool, gateway hookup
...
- navi/synapse/reactions.py: fire-and-forget reaction run — per-user
reactions_enabled gate, one-shot dispatcher meta-pass (visible profiles
only, hidden dispatcher, secretary fallback), special=True session
headed by the resolved profile, headless agent run via the orchestrator
run registry, finalise: app push per completion_notify +
low-level reaction.finished/failed event to Synapse
- navi/synapse/outbound.py: source-side emit (syn_* key from env,
silent when not configured)
- navi/tools/notify.py: notify tool — message + level (info/warning/
intervention), both legs per push_target, honest delivered/skipped result
- push/service.py: notify_custom without cooldown
- gateway: verified + non-duplicate delivery schedules the reaction
- profiles: notify added to native tools (7 profiles)
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W2: hidden dispatcher profile + synapse_instructions tool
Eugene Sukhodolskiy
committed
2 days ago
|
synapse reactions W1: special sessions flag + per-user reaction settings and instruction history
Eugene Sukhodolskiy
committed
2 days ago
|

handbook alignment: profile.updated mirror, health contract, Synapse gateway
...
Handbook gaps closed (10-platform auth/health/notifications):
- webhooks: handle user.profile_updated — mirror the gnexus-auth profile
into navi_users (name, contact fields, avatar_url — new column, boot
migration); role/permissions keep their dedicated events
- auth: avatar_url now flows through the login upsert and the API-token
resolution path
- /health: status is now the aggregate of the sub-checks (degraded embed
or hive no longer reports plain ok); embed probe cached for 10 s; body
grew the machine-readable checks{} map, legacy embed/hive payloads kept
- Synapse: s2s delivery gateway — POST /webhooks/synapse (both spellings),
per-user delivery secrets (synapse_targets, Fernet-encrypted, shown
once), verification via gnexus-synapse v0.1.2 client lib, idempotent
event ids; deliveries are verified and logged for now — navi generates
no outbound events by decision; web-push stays as the local browser
channel
- webclient settings: Synapse deliveries panel (add/revoke targets)
Eugene Sukhodolskiy
committed
2 days ago
|
webclient + navi: session title click-to-rename; empty rename regenerates via LLM
Eugene Sukhodolskiy
committed
2 days ago
|
navi: session naming folds agent's final replies into context
Eugene Sukhodolskiy
committed
2 days ago
|
navi: delete-timing log in session delete route
Eugene Sukhodolskiy
committed
2 days ago
|
| 2026-10-05 |

webclient: Backgrounds tab in artifacts panel + task toasts
...
- new Backgrounds tab: task list (running first, badges, per-tool icons)
with a detail view (status, timestamps, sub-agent tokens, result/progress)
- GET /sessions/{id}/tasks snapshot endpoint: task_update events are not
replayed on reconnect, so the client fetches the task list on session
load/reload; live task_update entries are merged on top (chat.fetchTasks)
- terminal task_update no longer removes the entry — it marks it finished
so the tab shows recent completions; _terminalTaskIds still blocks
resurrection by a late running update
- toasts for background task start / finish (info / success / error) in
the WS dispatch; silent for other sessions and unknown terminal ids
- persistent background-tasks chip removed (replaced by the tab); only
the queued-messages chip stays
- fix invisible status text in terminal detail rows: filled .status-*
backgrounds now scope to .terminal-status-badge only
Eugene Sukhodolskiy
committed
3 days ago
|

background tasks: review fixes B1-B11 batch
...
- stop mid-batch cancels in-flight tools (B1) and keeps real results
of already-finished ones, mixing them with synthetic stopped notes
in call order (B2)
- queued messages: headless drain publishes session_sync (B4), the
run's teardown broadcasts session_sync to other sockets but not the
owner socket (prevents double reload) (B5)
- task_update notifications are chained per task so late running
updates can't overtake the terminal one (B7; client drops late
re-flicker of a terminal task chip) (B8)
- client: ui_component results reference the owning message via
card.parentMsg instead of a stale msg reference (B6)
- tasks cancel reports the real outcome after a bounded wait instead
of an optimistic 'cancelled' (B9)
- session delete cancels its running background jobs and drops
pending result notes (B10)
- subagent tool loop routes background:true calls through
ToolExecutor._maybe_background like the main loop (B11)
- tests for all of the above; docs/tasks.md stop/cancel semantics
Eugene Sukhodolskiy
committed
3 days ago
|

local mode isolation: anonymous user is a NULL-owner user, not an alias
...
Local mode (NAVI_AUTH_ENABLED=false) used an anonymous admin whose admin
role widened every session listing to ALL users' chats — acceptable only
while assuming a private database, but a real leak on any shared one.
- session listings (list_all/list_page/count_all/search_list): new
scoping rule — non-admin with user_id=None sees only user_id IS NULL
rows; named owner unchanged; admin unchanged
- session create/list routes: owner id is None in local mode (sessions
are persisted ownerless), admin listing flag resolved only when auth
is enabled
- check_session_access: local mode allows only NULL-owner sessions —
a user-owned session by id is 403
- debug/admin endpoints (require_admin) stay reachable in local mode
- prod DB: 30 stray user_id='anonymous' rows reassigned to NULL owner
Tests: store scoping unit tests, local-mode integration tests (sidebar
filtering + access by id), updated check_session_access unit tests.
Full pytest 1239 passed, 1 skipped.
Eugene Sukhodolskiy
committed
3 days ago
|

chat history pagination: paged load instead of full-session fetch
...
Backend:
- PgSessionStore.get_meta — light session head (sessions row + COUNT),
no message payload
- PgSessionStore.get_history_page — page of display history (archive ∪
hot UNION), newest-first page, oldest-first result, limit+1 peek for
has_more; each message stamped metadata.display_index (ROW_NUMBER-1
over the full display history) so client ids stay stable across
paged loads; dangling-tool-call repair keeps real ranks by
sequence-number lookup
- GET /sessions/{id}/meta and GET /sessions/{id}/messages/history
(before_seq cursor, limit ≤200)
Webclient:
- loadSession/reloadSession fetch meta + newest 200-item page
(Promise.all), archive state from the page itself
- buildMessageList ids (h_*) and rawIndices use the global
display_index — feedback keys and search-jump stay stable when
history is paged
- scroll-up loads older pages through the same history endpoint
(archive table alone misses sessions whose threshold never moved)
- search-jump pulls pages until the target index is covered (bounded)
Tests: store unit tests (ranks, has_more peek, placeholder shift) +
vitest updates; full pytest 1231 passed, vitest 94 passed.
Eugene Sukhodolskiy
committed
3 days ago
|
| 2026-09-26 |
agent parallelism: background tools, bg spawn_agent, parallel tool batches, message queue
...
- backgroundable tools (terminal/ssh_exec/peer/spawn_agent/code_exec) detach
via TaskManager with per-session/global/spawn caps, rate limit and TTL
- completion delivery both ways: task_update event (out-of-band) + pending
notes injected into the next turn; tasks tool (list/check/wait/cancel)
- parallel tool-call batches (profile-level gate, off by default) with
ToolStarted up-front, tagged event mux, call-order results, one save,
plus dangling tool_call repair on session load
- user message queue instead of busy error: message_queued frame,
back-to-back drain on the same socket, headless fallback on disconnect
- pin mcp<2 (v2 renames FastMCP with breaking API changes)
- docs: tasks.md (new), websocket/api/agent/config updates, tasks manual,
spawn_agent background param, persona contract section
Eugene Sukhodolskiy
committed
12 days ago
|

PWA: installable webclient, offline shell, web push
...
Installability:
- public/manifest.webmanifest (standalone, theme #16161E) + PNG icons
generated from logo.svg (regular + maskable, served via /images mount)
- index.html: manifest link, theme-color, apple-touch-icon
Offline shell:
- hand-rolled sw.js (no workbox): navigation = network-first (3s race)
with cached-shell fallback + background refresh (a stale cached shell
would 404 on entry chunks after a deploy); /assets/* cache-first
(content-hashed); /images/* cache-first capped; /api,/ws,/auth,/push,
/content pass-through
- vite closeBundle plugin stamps __NAVI_BUILD_VERSION__ (digest of
index.html + asset names) into dist/sw.js; sw.js served no-store so
every deploy reactivates the SW and activation evicts old caches
- SW registration in main.js, PROD only (dev HMR untouched)
- OfflineBanner (useOnline composable) over the app shell
Web push (VAPID, pywebpush):
- navi/push/ package: push_subscriptions table (postgres, boot-time DDL),
PushSubscriptionStore, PushService (async fan-out, to_thread sends,
404/410 prunes dead endpoints, per-session cooldown)
- routes: GET /push/vapid-key, POST/DELETE /push/subscribe (auth-gated)
- trigger in orchestrator run_agent + run_recall: push on StreamEnd when
no WebSocket client watches the session; fire-and-forget, never
disturbs the run; anonymous fallback only when auth is off
- client: usePush composable + Notifications settings panel (enable/
disable via PushManager.subscribe with the server VAPID key)
- notification click focuses the app at /#<session_id> (hash routing
opens the right chat); payload body is a markdown-stripped <=140-char
preview
NAVIVAPID keys empty = push fully disabled (graceful, like other optional
integrations). dist/ artifacts committed per repo convention.
Tests: pytest push store/service/routes/trigger (+23), vitest usePush
(83 webclient tests green). Full suite 1143 passed.
Eugene Sukhodolskiy
committed
12 days ago
|
| 2026-09-10 |
swarm stage 2: peer-to-peer channel — /peer endpoints + peer tool
...
Server side (navi/api/routes/peer.py, port 8099):
- GET /peer/hello — open discovery ping (name, uuid prefix, version)
- GET /peer/status — PSK; identity, uptime, machine facts, hive view
- POST /peer/ask — PSK; one-shot agent run under PEER_ASK_PROFILE
(default server_admin) answers, one concurrent ask at a time
- loop guard: answering agent runs without the peer tool (deterministic
recursion cut) + own-uuid asks refused with 409
- verify_swarm_key in navi/swarm.py accepts .swarm-key.previous
(rotation window, constant-time)
Client side (navi/tools/peer.py):
- peer tool: list (hive book, stale cache fallback marked), status,
ask by swarm name; asks go direct peer-to-peer, hive never on the
message path
- audit events on both sides (peer.ask_sent/received/answered/failed)
peer tool added to server_admin, navi_code, developer profiles.
Eugene Sukhodolskiy
committed
29 days ago
|
| 2026-09-09 |
swarm stage 1: instance identity, PSK, hive registry, announce loop
...
- navi/identity.py: adjective-animal names + uuid in instance.json
(generated at install, renameable by hand)
- navi/swarm.py: HiveAnnouncer - periodic POST /announce to the hive
with non-blocking reachability tracking (transitional logs, /health
export, hive_status context provider reports outages to the agent)
- hive/: standalone FastAPI address book (SQLite, port 8087, PSK via
X-Swarm-Key with .swarm-key.previous rotation window, TTL online
status, host from client IP, port from payload). Not installed or
started by default - run manually on the main server.
- ports moved to 8099 (API) / 8098 (UI MCP)
- install.sh: PYTHONIOENCODING=utf-8 in the systemd unit, generates
instance.json and .swarm-key on fresh installs
Eugene Sukhodolskiy
committed
29 days ago
|
profiles: retire gate flags and retrain prompts for the plan tool
...
- drop planning_enabled / planning_mandatory / planning_phase2_enabled /
observe_skips_plan_enabled / adaptive_replan_enabled from AgentProfile,
the loader and admin serialization
- add `plan` to native tools of developer, navi_code, tool_developer,
secretary, server_admin, modeler_3d — without the gate they would
otherwise lose planning entirely
- system prompts: call plan for non-trivial multi-step work before
execution; complex plans get presented for confirmation; stuck or goal
changed -> plan with a reason
Eugene Sukhodolskiy
committed
29 days ago
|
Merge branch 'master' into feature/navi-code
...
# Conflicts:
# navi/api/websocket.py
# navi/config.py
# navi/core/agent.py
# navi/core/orchestrator.py
# navi/main.py
Eugene Sukhodolskiy
committed
29 days ago
|

security: critical batch 1 — RCE/XSS/CORS/webhook hardening
...
Backend:
- auth/deps: fix refresh-lock clock race (cleanup now monotonic like the cache)
- core/registry: add missing structlog logger (NameError on fallback path)
- config: GNAUTH_WEBHOOK_SECRET, NAVI_ALLOWED_ORIGINS (+list property)
- webhooks: HMAC-SHA256 signature verification (503 unconfigured in auth mode,
unsigned+warning in no-auth mode, 403 bad/stale signature, 400 bad JSON)
- main: CORS from NAVI_ALLOWED_ORIGINS in auth mode (fail-fast on empty),
eval router behind require_admin, /debug only registered in no-auth mode
- auth routes: mobile-done sid validation (32 hex) + CSP header + safe JS
escaping (reflected XSS); Secure cookie flag via helper (https base URL)
- websocket: anonymous WS rejected in auth mode (closes legacy-session RCE);
socket registered only after access checks; stop_session auth gate
- messages: REST agent start now takes session lock + busy flag (409 on
active run), contextvars reset in finally
Webclient:
- useMarkdown: DOMPurify.sanitize on all rendered markdown (stored XSS),
image-URL scheme whitelist, delegated error listener (no inline onerror)
- html.html artifact viewer: sandbox without allow-same-origin (opaque origin)
- tests: DOMPurify runs under jsdom (happy-dom Node.prototype.nodeName getter
breaks DOMPurify); useWebSocket tests get localStorage stub + ui-kit alias
Tests: pytest 1049 passed, 1 skipped; vitest 61 passed
Eugene Sukhodolskiy
committed
29 days ago
|
| 2026-07-13 |

terminal: REST list/close endpoints + async client helpers (Etap 2)
...
The /terminals modal and the status-bar count need a server source for the
session's persistent terminals (background dev servers, etc.) and a way to
force-close one. Add REST endpoints over the existing TerminalManager:
- GET /sessions/{id}/terminals → terminal_manager.list(session_id) →
{session_id, terminals: [summary,...]}. Access-checked (navi.sessions.read_all).
- POST /sessions/{id}/terminals/{name}/close → terminal_manager.close(...)
→ {session_id, terminal_name, closed: bool}. Name is URL-encoded by the
client (terminal names may contain spaces).
- clients/terminal/api.py (async): list_terminals(session_id),
close_terminal(session_id, name) — for the TUI modal + status-bar seed.
Tests: list empty / list returns active / close routes to manager / 404 on
unknown session (fake TerminalManager injected into the container, mirroring the
kv_store injection pattern). Full suite: 972 passed, 1 skipped.
Plan: docs/terminal_tool_plan.md — Etap 2 of 5.
Eugene Sukhodolskiy
committed
on 13 Jul
|

Автономность мелких моделей: OUTPUT DISCIPLINE, milestone-todo, перехват финала хода
...
Два последовательных захода над одной проблемой — мелкие модели (12–30B) в navi_code
теряют автономность на трудных шагах: «остановился поболтать» вместо действия и
слишком большое расстояние между пунктами плана.
Заход 1 (З1–З3):
- З1: OUTPUT DISCIPLINE в системном промпте navi_code — «act, don't announce»,
без few-shot антипримеров. Контракт хода: ответ без tool_calls = конец хода, поэтому
объявление намерения текстом убивает автономность; правила заставляют вызывать
инструмент в том же ходе.
- З2: плоский todo + метка группы milestone + декомпозиция. _parse_plan_steps
возвращает list[tuple[milestone, text]]; milestone — метка группировки (не сущность,
без статуса), «done» вычисляется при рендеринге; подшаги = больше плоских шагов
(без вложенности). TUI side-panel группирует по milestone (плоский фолбэк при пустом
milestone). Plan depth: max 15→20 + правило декомпозиции.
- З3: adaptive re-plan «длинный шаг» — nudge «разбей шаг» при in_progress ≥ порога
итераций без смены todo (порог 4, раньше общего anti-stall warning на 8).
Заход 2 (шаги 1–3, после cloud-теста 31b vs 12b):
Корневая структурная причина: весь спасательный механизм (anti-stall warning с явным
предложением reflect, adaptive re-plan) живёт только внутри tool-цикла — nudge
инжектируется в pre_turn *следующей* итерации, которой при «остановился поболтать»
нет (ход закрылся по return до post_turn). 31b застревала через «продолжаю
tool-итерации» → дожала до warning → спаслась; 12b — через «замолчала текстом» → мимо
всех nudge.
- Шаг 1: перехват финала хода. Если модель выдала bare-text, но в todo есть открытые
шаги (pending/in_progress) и лимит не исчерпан — НЕ эмитить StreamEnd, а сохранить
ассистентский текст в session.context, поставить системный nudge и continue
(без StreamEnd, без workers — консистентно с multi-iteration tool-турами). Счётчик
final_interceptions на AgentTurnContext, лимит final_intercept_limit (default 2),
эскалация жёсткости (мягкий → «second stop»). has_open_steps в todo.py: пустой
todo → False (защита casual-сообщений), failed/skipped терминальны. Профильные
флаги final_intercept_enabled/limit в base.py + loader + admin.
- Шаг 2: жёсткий reflect-триггер в промпте — «~3 tool attempts on the same step
without progress → call reflect IN THIS TURN (tool call, not reasoning aloud)».
- Шаг 3: открыть replan для застревания — «call replan when reflect showed the whole
approach is dead (not one failed step, but the approach itself)».
Тесты: 874 passed, 1 skipped. Новые — has_open_steps (5), final intercept (5),
milestone-группировка, adaptive long-step nudge, парсер шагов с milestone-маркером.
cloud: navi_code model → gemma4:31b-cloud для тестирования догадки (31b признала
застревание, 12b — нет); .env cloud-host уже gitignored.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 13 Jul
|
| 2026-07-11 |

Navi Code: force context compression via typed /compact control message
...
Forced /compact previously sent a chat message, which ran a full agent turn
that only produced summary text instead of running the real context compressor.
Worse, even when wired correctly, forced compact always reported "Nothing to
compact yet — the context is still small" regardless of context size: the typical
navi_code shape is a single long autonomous turn (1 user message + many tool
iterations = one turn), and partition_messages finds nothing to summarize when
turns <= keep_recent. The midturn auto-compress path already bypassed this via
keep_recent_messages (intra-turn split), but compact_stream passed
keep_recent_messages=None, so the fallback was disabled.
Changes:
- WS protocol: {"type":"compact"} control message (distinct from {"type":
"message"}); rejected while an agent turn is active to avoid racing the agent.
- Agent.compact_stream: forced compression that bypasses the token threshold
but still runs the real compressor; passes keep_recent_messages=max(12,
context_keep_recent*2) so a single long turn compresses via intra-turn split
(mirrors midturn auto-compress). Raises NothingToCompactError when context is
genuinely too small.
- Orchestrator.run_compact + clear_run: broadcast agent events to subscribers,
end with done marker, surface NothingToCompactError as an error event.
- Terminal client: ws_client.send accepts str|dict; CompactCommand enqueues
{"type":"compact"}; TUI distinguishes forced compact (no stream_start) from
in-turn auto-compress via the _streaming flag.
- Tests: compact_stream (incl. single-long-turn regression), WS handler
dispatch/rejection, run_compact event/error broadcasting, ws_client send,
compact command.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 11 Jul
|

Isolate sub-agent todo + render live todo in TUI side panel
...
Phase 1 — sub-agent todo isolation (backend):
- Add current_todo_session_id ContextVar; subagent_runner scopes it to the
sub-agent's ephemeral run id so its auto-populated plan and todo updates
land in an isolated KV row instead of clobbering the parent session's todo
(which the parent's goal-anchoring reads every iteration).
- todo._sid() and planning.set_tasks prefer current_todo_session_id; the
parent run leaves it unset, so all existing todo consumers (anti-stall,
goal anchor, get_progress_message) behave exactly as before.
Phase 2 — live todo in the TUI right column:
- TodoUpdated event + emit it from the agent loop after planning auto-populate
and after each tool-execution turn.
- GET /sessions/{id}/todos reads the parent session's todo KV row (explicit
user_id/session_id, optional injected kv).
- api.get_todos + TodoList/TodoPanel widgets: status-coloured markers
(pending dim, in_progress accent+bold, done success, failed error, skipped
dim), progress header, scrollable panel below the auto-height info block.
- Hybrid delivery: REST seeds the panel on attach/switch, todo_updated WS
events update it live.
Sub-agent todos as nested sub-lists is deferred to a later phase.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 11 Jul
|
| 2026-07-09 |

agent: bounded autonomy — scope boundary + observe-vs-act
...
navi_code had unwanted "free flight": an observe request ("look at a
directory") triggered the full Phase 3 plan with milestones + auto-todo,
and goal_anchoring then drove the agent to finish those steps, climbing
into sibling projects and executing milestone docs it found.
Two toggleable, default-off profile flags (on for navi_code):
- scope_boundary_enabled: injects a standing system message keeping the
agent within the literally requested scope; forbids acting on
discovered TODO/roadmap/milestone docs (report only).
- observe_skips_plan_enabled: Phase 1 classifies MODE: observe|act; an
observe request skips Phase 2/3 — no multi-step plan, no auto-todo, no
"execute step by step" prompt. The agent just gathers info and answers.
Independent of force_plan (observe on the first message still skips).
Free flight stays reproducible by flipping both flags off.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 9 Jul
|
fix(ws): include session_id/profile_id in session_sync, guard renderer
...
The server sent {"type": "session_sync"} without session_id/profile_id,
crashing the terminal client (render.py did None[:8]). Add the fields to
both session_sync sends and guard the renderer against a missing id.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 9 Jul
|