| 2026-10-08 |

tools, mcp: a description is a hook, not a manual
...
With 24 tools in a profile the schemas are the largest fixed block of every
request, and they are paid for whether or not the tool is called. Eleven of
them carried a manual's worth of prose in their `description` — examples,
error tables, "common mistakes", the mechanics of the file areas. Reading
them against the manuals first, the manuals already held all of it, and
richer: this was duplication sitting in the one place the model cannot avoid
reading, not knowledge that had nowhere else to live. So the descriptions
became selection hooks — what the tool is for, when to pick it over its
neighbour, and the one rule that cannot wait — and the detail moved to
tool_manual("<tool>"), which costs nothing until the tool is actually in
play.
filesystem, spawn_agent, schedule_recall, manage_recall, scratchpad, plan,
share_file, content_publish, reflect, todo and peer. Descriptions drop from
18.4 KB to 11.5 KB, those tools' schemas from 25.5 KB to 16.8 KB, and a
profile's whole native toolset from 38.3 KB to 29.6 KB (server_admin,
developer; secretary 33.7 → 25.0, navi_code 32.9 → 25.9, tool_developer
37.9 → 29.2, modeler_3d 33.4 → 24.8). Stated as a budget rather than a byte
count: roughly 4k tokens of the model's window per request, back.
Two things were deliberately not shortened. The filesystem edit ladder
(`edit` → `edit_lines` → `smart_edit`, last resort, costs an LLM call) and
todo's mandatory `validation` on `done` both exist to prevent an extra
round trip; a shorter description there would trade tokens for turns. What
todo lost is the prose around the rule, not the rule. `enabled.json` was
left alone too: weather, gmail and get_current_datetime are opt-in and the
DB shows 12, 16 and 40 calls, so they are opt-in tools that are actually
used.
peer had no manual at all, which is how this surfaced: its description was
the only documentation the tool had. manuals/peer.md is now written from
the tool and its /peer route — the ask/status/list actions, the fact that
the hive is only a phone book and the question travels peer-to-peer, the
one-concurrent-answer semaphore, the two independent recursion guards, and
the six error codes.
The MCP server instructions leave the system prompt the same way. They were
8 KB of always-present prose per request, most of it workflow detail the
model only needs when it is about to use that server. Each server now
contributes one line — a new optional `summary` field in mcp_servers.d/*.json,
defaulting to the first sentence of its `instructions` — and
tool_manual("<server>") returns the full text plus the tool names the server
declares. server_admin's MCP block goes from 8.3 KB to 1.3 KB, secretary's
from 8.1 to 1.2. tool_manual learned to answer for a server (a tool name
still wins over a server of the same name, and dash/underscore spelling is
tolerated), and the index names the servers alongside the tools.
One judgement call worth recording: gnexus-book's and navi-web's
instructions end in an absolute "NEVER bypass these tools with filesystem,
terminal or code_exec" rule. Moving that on demand would have been a
behavioural regression dressed as a token saving, so those two `summary`
fields carry the prohibition verbatim alongside the hook, and it stays
always visible.
docs/tools.md gains a section on the split (a description is paid for every
request, a manual only when the tool is used), which is where the next
person will look before trimming a description back into a manual.
Full suite green (1577 passed, 1 skipped). No frontend change and no new
dependency, so the deploy is a pull and a restart.
Eugene Sukhodolskiy
committed
4 hours ago
|

tool_manual: a manual the agent cannot reach is not a manual
...
tool_manual asked the filesystem for the exact string the model sent, then
fell back to registry.get(). The executor already tolerated the ways models
mangle MCP names; the tool that documents them did not, so the agent calling
mcp__navi-3d__compile_scad got a schema dump while manuals/compile_scad.md sat
there unread. Resolving a name now happens in one place — resolve_tool moved
out of the executor into navi/core/tool_utils.py — so a name that works when
the tool is *called* also works when it is *asked about*. A miss suggests the
closest names instead of "not found", and a tool that exists but is not enabled
for this profile says exactly that: documenting it is still the useful answer,
but the agent must not go on to call it. params["tool_name"] was also a crash
waiting to happen — KeyError with the key omitted, AttributeError on null or a
number — and now anything that is not a usable string means "no name given",
which returns the index of manuals grouped by source.
The generated manual now renders the whole schema. An array of objects showed
as "(array, optional)" with its item shape invisible, so a model that could not
see the fields guessed them; objects, array items, oneOf/anyOf branches, enums
and defaults are all spelled out now, depth-capped. The header names the source
(native / mcp__<server>__) and says the text is a parameter contract, not a
curated manual — otherwise a thin schema reads as "this tool is simple".
Then the manuals themselves. A manual is filed under the tool's own name, and
five of the sixteen were not: write_tool.md and write_mcp_server.md documented
tools that do not exist anywhere (the tools are reload_tools and
create_mcp_server), model_3d.md documented compile_scad, render_3d.md documented
render_stl, and write_context_provider.md documented no tool at all — it is a
guide, and now lives in manuals/guides/. The prompts pointed at the same wrong
names (tool_developer's system prompt, persona_navi_code, docs/context_providers),
so fixing the files without the callers would have moved the breakage rather
than removed it. On HEAD the probe was blunt: tool_manual("mcp__navi-3d__compile_scad")
returned None and the agent got the schema; now it returns the manual.
Nothing was enforcing any of this, which is how it drifted. tests/unit/tools/
test_manual_drift.py reads the real repository tree: every manuals/*.md must be
named after a tool that exists (built-ins and tools/*.py by their name
assignments, MCP tools by @mcp.tool(name=…) in the server sources and by the
mcp_servers.d groups, whose file name is the server name), every manual must be
reachable by both its bare name and its full MCP spelling, a guide must never
shadow a tool's name, and every manuals/<x>.md or tool_manual("x") cited in a
doc or a prompt must resolve. The next rename fails the suite instead of
quietly wasting a manual.
Finally, manuals for the built-ins the agent uses most — ssh_exec, todo, memory,
plan, filesystem, notify, code_exec — written from the current schemas and the
code paths that produce the errors, so the gotchas are the real ones (todo's
`done` validation is enforced; memory's `list` returns categories, not facts;
filesystem's `delete` removes a tree with no prompt). reload_tools.md absorbs
the file format from the deleted write_tool.md, which is where a self-extension
recipe belongs now.
1556 passed, 1 skipped. ruff delta zero.
Eugene Sukhodolskiy
committed
5 hours ago
|
| 2026-05-23 |
Pass explicit ToolContext to tools instead of hidden ContextVars
...
Add ToolContext dataclass (session_id, event_sink, stop_event, model,
user_id, user_role, user_info) and thread it through the execution chain:
Agent._execute_tools_with_sink → ToolExecutor → tool.execute().
All ~25 tools updated to accept ctx parameter. Tools that previously
read ContextVar now prefer ctx when provided, falling back to
ContextVar for backward compatibility.
Tests updated to pass ToolContext explicitly — no more test fixtures
that set current_session_id / current_user_id ContextVars.
ContextVar setters remain as fallback for non-tool consumers
(ai_helper, context_builder, planning) and will be removed in a
follow-up refactor.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 23 May
|
| 2026-05-16 |
Enhance native toolset and add persistent KV store
...
- Add PostgreSQL-backed KvStore (navi/store/) for session-scoped data.
- Migrate todo and scratchpad from in-memory dicts to KvStore.
- Filesystem: add copy, grep, diff actions; compress description.
- CodeExec: remove language param, expose working_dir in schema.
- ImageView: resize to 1024px JPEG + Content-Type guard for URLs.
- Memory list: return distinct categories instead of all facts.
- SSH: add scp action with upload/download support.
- Update CLAUDE.md (Postgres-only), docs/tools.md, add docs/store.md.
- Fix agent/planning/context_builder async signatures for todo helpers.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 16 May
|
| 2026-05-15 |
cleanup: remove deprecated tools and orphaned memory tools
...
Removed (no profile used them, no cross-dependencies):
- write_tool.py, delete_tool.py, test_tool.py
- memory_save.py, memory_search.py, memory_forget.py
Updated:
- navi/core/registry.py — removed imports and registrations
- navi/tools/__init__.py — removed imports
- docs/tools.md — removed references, updated self-extension section to MCP
- navi/tools/tool_manual.py — updated example to create_mcp_server
Profile fixes:
- developer: +tool_manual, +ssh_exec
- discuss: +list_profiles
- tool_developer: +mcp_servers navi-web
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 15 May
|
Refactor: move tool helper modules into _internal subpackage
...
Moves non-tool infrastructure out of navi/tools/ root so that only
actual tool classes live there:
base.py → _internal/base.py
loader.py → _internal/loader.py
middleware.py → _internal/middleware.py
logging_middleware.py → _internal/logging_middleware.py
_time_parser.py → _internal/time_parser.py
All imports updated across core/, api/, mcp/, tools/, and tests/.
No proxy files remain in navi/tools/ root.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 15 May
|
| 2026-04-08 |
Add self-extension tool system: write_tool, list_tools, tool_manual
...
- loader.py: module-level format (name/description/parameters/execute)
preferred, class-based as fallback; isolated errors per file
- write_tool: validates + writes tools/name.py, reloads registry,
adds to tools/enabled.json in one call
- list_tools: live tool list from registry (prevents hallucination)
- tool_manual: serves manuals/*.md or auto-generates from schema
- reload_tools: hot-reload without server restart
- registry: registry injection pattern for tools that need it;
_builtin_names set to guard against reload overwriting builtins
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 8 Apr
|