| 2026-10-08 |

tools, mcp: a description is a hook, not a manual
...
With 24 tools in a profile the schemas are the largest fixed block of every
request, and they are paid for whether or not the tool is called. Eleven of
them carried a manual's worth of prose in their `description` — examples,
error tables, "common mistakes", the mechanics of the file areas. Reading
them against the manuals first, the manuals already held all of it, and
richer: this was duplication sitting in the one place the model cannot avoid
reading, not knowledge that had nowhere else to live. So the descriptions
became selection hooks — what the tool is for, when to pick it over its
neighbour, and the one rule that cannot wait — and the detail moved to
tool_manual("<tool>"), which costs nothing until the tool is actually in
play.
filesystem, spawn_agent, schedule_recall, manage_recall, scratchpad, plan,
share_file, content_publish, reflect, todo and peer. Descriptions drop from
18.4 KB to 11.5 KB, those tools' schemas from 25.5 KB to 16.8 KB, and a
profile's whole native toolset from 38.3 KB to 29.6 KB (server_admin,
developer; secretary 33.7 → 25.0, navi_code 32.9 → 25.9, tool_developer
37.9 → 29.2, modeler_3d 33.4 → 24.8). Stated as a budget rather than a byte
count: roughly 4k tokens of the model's window per request, back.
Two things were deliberately not shortened. The filesystem edit ladder
(`edit` → `edit_lines` → `smart_edit`, last resort, costs an LLM call) and
todo's mandatory `validation` on `done` both exist to prevent an extra
round trip; a shorter description there would trade tokens for turns. What
todo lost is the prose around the rule, not the rule. `enabled.json` was
left alone too: weather, gmail and get_current_datetime are opt-in and the
DB shows 12, 16 and 40 calls, so they are opt-in tools that are actually
used.
peer had no manual at all, which is how this surfaced: its description was
the only documentation the tool had. manuals/peer.md is now written from
the tool and its /peer route — the ask/status/list actions, the fact that
the hive is only a phone book and the question travels peer-to-peer, the
one-concurrent-answer semaphore, the two independent recursion guards, and
the six error codes.
The MCP server instructions leave the system prompt the same way. They were
8 KB of always-present prose per request, most of it workflow detail the
model only needs when it is about to use that server. Each server now
contributes one line — a new optional `summary` field in mcp_servers.d/*.json,
defaulting to the first sentence of its `instructions` — and
tool_manual("<server>") returns the full text plus the tool names the server
declares. server_admin's MCP block goes from 8.3 KB to 1.3 KB, secretary's
from 8.1 to 1.2. tool_manual learned to answer for a server (a tool name
still wins over a server of the same name, and dash/underscore spelling is
tolerated), and the index names the servers alongside the tools.
One judgement call worth recording: gnexus-book's and navi-web's
instructions end in an absolute "NEVER bypass these tools with filesystem,
terminal or code_exec" rule. Moving that on demand would have been a
behavioural regression dressed as a token saving, so those two `summary`
fields carry the prohibition verbatim alongside the hook, and it stays
always visible.
docs/tools.md gains a section on the split (a description is paid for every
request, a manual only when the tool is used), which is where the next
person will look before trimming a description back into a manual.
Full suite green (1577 passed, 1 skipped). No frontend change and no new
dependency, so the deploy is a pull and a restart.
Eugene Sukhodolskiy
committed
4 hours ago
|
| 2026-07-12 |

compression: fix auto-compress no-op on few-huge-messages + honest status
...
Root cause: the compression gate (should_compress) measures tokens, but the
partition measures message/turn count, and CompressionStarted was emitted
before the attempt. For navi_code's "few very large messages" shape (one big
file read = 1 user + assistant + 1 huge tool result, 66k tokens in 3 messages)
the gate fired, the UI showed "compression", but partition returned
to_summarize=[] -> compress_context None -> nothing shrank. The agent kept
going until the window overflowed. It wasn't running "during" compression —
there was no compression, just a no-op the user mistook for one.
A. Per-message head/tail truncation in context_builder.build(): oversized
tool/assistant messages (over context_message_token_budget, 0=num_ctx//6)
are capped head+marker+tail in the LLM view only (model_copy — stored
history and reloads are never affected). A single huge tool result can no
longer alone blow the window; user/system messages are never truncated.
B. Token-budget hard-truncate fallback in compress_session: when partition
no-ops but tokens exceed the threshold, drop oldest turns to num_ctx*0.5.
_hard_truncate is now token-aware (was a fixed message-count floor that
no-oped on <=6 messages even when huge). New would_compress() predicts
compress_session's real outcome with no LLM call.
C. Honest CompressionStarted: _compression_events_midturn/_preturn emit it
only after would_compress() confirms the partition (or token-budget
fallback) can actually shrink the stored context — no more "compression"
status with no ContextCompressed to follow.
Bonus: post-turn CompressionWorker now passes keep_recent_messages=
max(12, context_keep_recent*2), matching the midturn path, so a single long
autonomous turn compresses post-turn too (was always a no-op).
Tests (+14): would_compress agreement, token-budget fallback, token-aware
hard_truncate, build() truncation (preserves user, no mutation, head+tail),
agent no-CompressionStarted-when-nothing-to-compress, worker single-long-turn.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 12 Jul
|
| 2026-07-09 |
agent: cwd-aware memory fact filter for bounded autonomy
...
Long-term memory stored context-dependent path facts (e.g. "project_root
→ /home/.../navi-1") as global user facts. search_facts injected them into
any session, so when working in another project the agent was told "the
project root is navi-1" and drifted there.
When scope_boundary_enabled and a session cwd is set, _memory_facts_msg now
drops facts whose value is an absolute path outside the session cwd tree.
Facts are kept when working inside that path (then they are correct), and
non-path/relative facts always pass. Free flight stays reproducible by
toggling the flag off. No facts deleted, extractor untouched.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 9 Jul
|

agent: bounded autonomy — scope boundary + observe-vs-act
...
navi_code had unwanted "free flight": an observe request ("look at a
directory") triggered the full Phase 3 plan with milestones + auto-todo,
and goal_anchoring then drove the agent to finish those steps, climbing
into sibling projects and executing milestone docs it found.
Two toggleable, default-off profile flags (on for navi_code):
- scope_boundary_enabled: injects a standing system message keeping the
agent within the literally requested scope; forbids acting on
discovered TODO/roadmap/milestone docs (report only).
- observe_skips_plan_enabled: Phase 1 classifies MODE: observe|act; an
observe request skips Phase 2/3 — no multi-step plan, no auto-todo, no
"execute step by step" prompt. The agent just gathers info and answers.
Independent of force_plan (observe on the first message still skips).
Free flight stays reproducible by flipping both flags off.
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 9 Jul
|
| 2026-05-18 |
Make Settings immutable (frozen=True) and fix all test mutations
...
- Add frozen=True to SettingsConfigDict in navi/config.py
- Convert model_validator to mode="before" since mode="after" cannot mutate frozen instances
- Replace all field-level monkeypatches in tests with whole-Settings object replacement
- Ensure cross-module settings consistency (content_store, session_files, share_file, content_publish, filesystem)
392 passed, 1 skipped
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 18 May
|
| 2026-05-16 |
Enhance native toolset and add persistent KV store
...
- Add PostgreSQL-backed KvStore (navi/store/) for session-scoped data.
- Migrate todo and scratchpad from in-memory dicts to KvStore.
- Filesystem: add copy, grep, diff actions; compress description.
- CodeExec: remove language param, expose working_dir in schema.
- ImageView: resize to 1024px JPEG + Content-Type guard for URLs.
- Memory list: return distinct categories instead of all facts.
- SSH: add scp action with upload/download support.
- Update CLAUDE.md (Postgres-only), docs/tools.md, add docs/store.md.
- Fix agent/planning/context_builder async signatures for todo helpers.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 16 May
|
| 2026-05-12 |
Clarify knowledge persistence prompts
Eugene Sukhodolskiy
committed
on 12 May
|
| 2026-05-01 |
Improve 3D modeling validation prompts
Eugene Sukhodolskiy
committed
on 1 May
|
| 2026-04-29 |

Bootstrap test suite — Phase 1 unit tests
...
- docs/testing.md: testing strategy, mock strategy, phase breakdown
- tests/conftest.py: autouse fixture to reset navi.config.settings per test
- tests/conftest_factory.py: FakeLLMBackend, FakeTool, make_profile, make_registry helpers
- tests/unit/core/test_events.py: wire serialization for all 15 event dataclasses
- tests/unit/core/test_compressor.py: should_compress, partition_messages, format_for_summary, compress_context
- tests/unit/core/test_registry.py: ToolRegistry, ProfileRegistry, BackendRegistry
- tests/unit/core/test_context_builder.py: system prompt caching, persona injection, goal anchor, iteration budget
- tests/unit/profiles/test_base.py: Pydantic model coercion, defaults, extra fields
- navi/core/context_builder.py: use module-level `import navi.config` instead of `from navi.config import settings` so tests can swap the singleton
59 tests passing.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 29 Apr
|