| 2026-10-09 |

profiles: restricted profiles for ordinary users, and a role gate that holds
...
`is_admin_only` was checked in one place out of nine and was not read from
config.json at all, so all seven profiles were reachable by every account with
role `user`. The flag now lives in config.json — the file is the baseline, a
`profile_overrides` row still wins on top of it — and one predicate,
`admin_only_blocked`, is the single place the rule is expressed. The nine
surfaces that list, switch to, spawn or resolve a profile all consult it:
`POST /sessions`, the WebSocket, switch_profile, list_profiles, the system
prompt's "Available profiles" block, spawn_agent and the Synapse reaction
runner. The prompt cache is now keyed by (profile, role), so a user's prompt
can never be served an admin's profile list.
The seven existing profiles (developer, discuss, dispatcher, modeler_3d,
navi_code, secretary, server_admin) are marked admin-only. Three new ones take
their place for ordinary users: assistant, designer_3d and coder. They share one
native tool set — ssh_exec, peer, reload_tools, create_mcp_server, test_mcp_tool,
image_view and gmail are withheld — and differ only in system prompt, model and
MCP groups. navi-web's raw `request` group, and the whole of gnexus-creds and
tgclient, are withheld too.
MCP per-user keys gain the missing half of the rule: a server that declares a
`user_key` slot is refused to anyone but an admin who has no personal key, and
is left out of their tool list entirely, instead of quietly falling back to the
owner's credential and appearing as a tool that cannot work. The refusal names
the server and points at Settings.
Also closes `GET /agents/prompts`, which served every profile's system prompt to
anyone, with no user dependency at all.
The accepted residual risk is written down in docs/profiles.md: the working
directory is a convention, not a sandbox.
Eugene Sukhodolskiy
committed
3 hours ago
|
| 2026-10-08 |

startup, llm, profiles, todo: five more the prod logs gave up
...
The same 5-day sweep of prod's journal as 5e51f87, continued. Every one of
these failed silently: no error in the UI, just a line in the log or a hole
in the database.
complete()/embed() retry a connection error before falling back.
stream_complete has always given its first chunk two attempts, but the
non-streaming path made exactly one per (server, model), and with a single
configured server (ollama.com) one ReadTimeout raised "All backends
exhausted: ReadTimeout" outright. That is the whole story of 2026-10-08,
when 27 calls died that way in a day — 23 of them memory summarisation, the
rest planning. Attempts now come from settings.llm_complete_retries (2) with
llm_retry_backoff_sec (2.0) between them; a missing model is still not
retried (a 404 is an answer, not a hiccup), and a lone server is still never
blacklisted — blacklisting it would block the next request for _TTL.
Startup waits for Postgres. create_container() opens the pool eagerly and
the retry loop for the DDL tables sits below it, so a host that comes up
before docker takes the whole lifespan with it. On 2026-10-07 15:19:40 the
agent ran `sudo reboot` on itself at the user's request ("Перезагрузи
себя"); the host booted 15:20:03, navi started at 15:20:18, hit
ConnectionRefusedError on the pool and logged "Application startup failed.
Exiting." Connection errors now retry for ~30s; a genuinely broken config
(a missing DATABASE_URL) still fails fast.
The profile loader reports its tally. A profile dropped by an unreadable
config.json is skipped by design so it cannot take the server down, but on
that same 15:20 boot four of them — modeler_3d, navi_code, secretary,
server_admin — were dropped with `Extra data: line 125 column 1`, and the
only trace was one error line each: the UI simply showed four profiles
fewer.
todo recovers a dropped `op`. Of the 3385 todo calls stored on prod, 73
arrived with the discriminator flattened away ({"index": 1, "status":
"in_progress"}, {"action": "view"}, {"": "add", "tasks": [...]}) and got a
bare "Unknown op: None" — a wasted round trip for an intent the remaining
arguments state plainly. The error message now names the five ops and what
each needs.
Message.created_at is stamped at creation. It defaulted to None and only
some construction sites filled it, so 85% of stored rows had no timestamp at
all (every tool message, 14 168 of 14 168) and no query could slice history
by time or measure a pause. Loading a row passes the column through
explicitly, so NULL rows stay NULL instead of being restamped with "now".
Every new test was checked to fail with its fix reverted. Full suite green
(1466 passed, 1 skipped). No frontend or dependency change, so the deploy is
a pull and a restart.
Eugene Sukhodolskiy
committed
9 hours ago
|
| 2026-09-09 |
profiles: retire gate flags and retrain prompts for the plan tool
...
- drop planning_enabled / planning_mandatory / planning_phase2_enabled /
observe_skips_plan_enabled / adaptive_replan_enabled from AgentProfile,
the loader and admin serialization
- add `plan` to native tools of developer, navi_code, tool_developer,
secretary, server_admin, modeler_3d — without the gate they would
otherwise lose planning entirely
- system prompts: call plan for non-trivial multi-step work before
execution; complex plans get presented for confirmation; stuck or goal
changed -> plan with a reason
Eugene Sukhodolskiy
committed
29 days ago
|
| 2026-07-13 |

Автономность мелких моделей: OUTPUT DISCIPLINE, milestone-todo, перехват финала хода
...
Два последовательных захода над одной проблемой — мелкие модели (12–30B) в navi_code
теряют автономность на трудных шагах: «остановился поболтать» вместо действия и
слишком большое расстояние между пунктами плана.
Заход 1 (З1–З3):
- З1: OUTPUT DISCIPLINE в системном промпте navi_code — «act, don't announce»,
без few-shot антипримеров. Контракт хода: ответ без tool_calls = конец хода, поэтому
объявление намерения текстом убивает автономность; правила заставляют вызывать
инструмент в том же ходе.
- З2: плоский todo + метка группы milestone + декомпозиция. _parse_plan_steps
возвращает list[tuple[milestone, text]]; milestone — метка группировки (не сущность,
без статуса), «done» вычисляется при рендеринге; подшаги = больше плоских шагов
(без вложенности). TUI side-panel группирует по milestone (плоский фолбэк при пустом
milestone). Plan depth: max 15→20 + правило декомпозиции.
- З3: adaptive re-plan «длинный шаг» — nudge «разбей шаг» при in_progress ≥ порога
итераций без смены todo (порог 4, раньше общего anti-stall warning на 8).
Заход 2 (шаги 1–3, после cloud-теста 31b vs 12b):
Корневая структурная причина: весь спасательный механизм (anti-stall warning с явным
предложением reflect, adaptive re-plan) живёт только внутри tool-цикла — nudge
инжектируется в pre_turn *следующей* итерации, которой при «остановился поболтать»
нет (ход закрылся по return до post_turn). 31b застревала через «продолжаю
tool-итерации» → дожала до warning → спаслась; 12b — через «замолчала текстом» → мимо
всех nudge.
- Шаг 1: перехват финала хода. Если модель выдала bare-text, но в todo есть открытые
шаги (pending/in_progress) и лимит не исчерпан — НЕ эмитить StreamEnd, а сохранить
ассистентский текст в session.context, поставить системный nudge и continue
(без StreamEnd, без workers — консистентно с multi-iteration tool-турами). Счётчик
final_interceptions на AgentTurnContext, лимит final_intercept_limit (default 2),
эскалация жёсткости (мягкий → «second stop»). has_open_steps в todo.py: пустой
todo → False (защита casual-сообщений), failed/skipped терминальны. Профильные
флаги final_intercept_enabled/limit в base.py + loader + admin.
- Шаг 2: жёсткий reflect-триггер в промпте — «~3 tool attempts on the same step
without progress → call reflect IN THIS TURN (tool call, not reasoning aloud)».
- Шаг 3: открыть replan для застревания — «call replan when reflect showed the whole
approach is dead (not one failed step, but the approach itself)».
Тесты: 874 passed, 1 skipped. Новые — has_open_steps (5), final intercept (5),
milestone-группировка, adaptive long-step nudge, парсер шагов с milestone-маркером.
cloud: navi_code model → gemma4:31b-cloud для тестирования догадки (31b признала
застревание, 12b — нет); .env cloud-host уже gitignored.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 13 Jul
|
| 2026-07-09 |

agent: bounded autonomy — scope boundary + observe-vs-act
...
navi_code had unwanted "free flight": an observe request ("look at a
directory") triggered the full Phase 3 plan with milestones + auto-todo,
and goal_anchoring then drove the agent to finish those steps, climbing
into sibling projects and executing milestone docs it found.
Two toggleable, default-off profile flags (on for navi_code):
- scope_boundary_enabled: injects a standing system message keeping the
agent within the literally requested scope; forbids acting on
discovered TODO/roadmap/milestone docs (report only).
- observe_skips_plan_enabled: Phase 1 classifies MODE: observe|act; an
observe request skips Phase 2/3 — no multi-step plan, no auto-todo, no
"execute step by step" prompt. The agent just gathers info and answers.
Independent of force_plan (observe on the first message still skips).
Free flight stays reproducible by flipping both flags off.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 9 Jul
|
| 2026-05-08 |
Add pagination, search, and sorting to admin sessions
...
Backend:
- Add count_all and search_list abstract methods to SessionStore
- Implement count_all and search_list in PgSessionStore (SQL with ILIKE)
- Implement count_all and search_list in InMemorySessionStore
- Update /admin/sessions to accept limit, offset, search, sort_by, sort_order
- Return {total, limit, offset, items} from /admin/sessions
Frontend:
- Add search input for sessions in admin panel
- Add clickable sortable column headers with asc/desc toggle
- Add pagination controls (prev/next, page size selector, item count)
- Debounce search input (300ms)
Tests:
- Add integration tests for pagination, offset, search, and sorting
- All 217 tests pass
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 8 May
|
| 2026-05-01 |
Disable thinking stalls for 3D subagents
Eugene Sukhodolskiy
committed
on 1 May
|
| 2026-04-29 |

Bootstrap test suite — Phase 1 unit tests
...
- docs/testing.md: testing strategy, mock strategy, phase breakdown
- tests/conftest.py: autouse fixture to reset navi.config.settings per test
- tests/conftest_factory.py: FakeLLMBackend, FakeTool, make_profile, make_registry helpers
- tests/unit/core/test_events.py: wire serialization for all 15 event dataclasses
- tests/unit/core/test_compressor.py: should_compress, partition_messages, format_for_summary, compress_context
- tests/unit/core/test_registry.py: ToolRegistry, ProfileRegistry, BackendRegistry
- tests/unit/core/test_context_builder.py: system prompt caching, persona injection, goal anchor, iteration budget
- tests/unit/profiles/test_base.py: Pydantic model coercion, defaults, extra fields
- navi/core/context_builder.py: use module-level `import navi.config` instead of `from navi.config import settings` so tests can swap the singleton
59 tests passing.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 29 Apr
|