diff --git a/docs/api.md b/docs/api.md index 2828e1e..c7b86c8 100644 --- a/docs/api.md +++ b/docs/api.md @@ -206,7 +206,7 @@ #### `GET /agents/prompts` -Return the fully resolved system prompt for each profile (persona + profile system_prompt + context provider injections + MCP server instructions). +Return the fully resolved system prompt for each profile (persona + profile system_prompt + context provider injections + one line per MCP server). **Response `200`** ```json diff --git a/docs/mechanics.md b/docs/mechanics.md index f8f7161..47000a0 100644 --- a/docs/mechanics.md +++ b/docs/mechanics.md @@ -91,7 +91,7 @@ | **Scope-boundary message** | Injects a standing `[Scope boundary]` system message keeping the agent within the user's literally requested scope — do not expand to sibling/parent dirs/projects, do not execute discovered TODO/roadmap/milestone docs unless explicitly asked. | `profile.scope_boundary_enabled` | `context_builder.py` | ✅ | | **Security policy message** | Injects `[Security policy]` based on `current_user_role`: admin = full access; user = sandbox + terminal allowlist. | `TERMINAL_ALLOWED_COMMANDS` | `context_builder.py` | ❌ | | **User context message** | Builds `[User context]` from `current_user_info` (display_name, email, locale, etc.). | None | `context_builder.py` | ❌ | -| **MCP context message** | Combines MCP server instructions from handshake with overlay instructions from `mcp_servers.d/*.json`. | `profile.tools.agent.mcp` | `context_builder.py` | ❌ | +| **MCP context message** | One line per reachable server: its configured `summary`, else the first sentence of its instructions, plus a pointer to `tool_manual("")` for the rest. The instructions themselves are *not* inlined — they are read on demand. | `profile.tools.agent.mcp` | `context_builder.py` | ❌ | | **Iteration budget message** | Appends `[Iteration N/M — K after this one]` with escalating urgency when ≤2 or ≤5 remaining. | `profile.iteration_budget_enabled` | `context_builder.py` | ✅ | | **Session context injection** | Appends session ID and exact `session_files_dir/{session_id}/` path. | `SESSION_FILES_DIR` | `context_builder.py` | ❌ | diff --git a/docs/tools.md b/docs/tools.md index 318a468..aa203d5 100644 --- a/docs/tools.md +++ b/docs/tools.md @@ -169,6 +169,18 @@ --- +## Schema descriptions vs manuals + +Every tool's `description` travels to the model on **every request**. With 25+ tools in a profile it is the largest fixed block of context the agent pays for, used or not. So a description is written as a **selection hook**, not as documentation: what the tool is for, when to pick it over its neighbour, and the one rule that cannot wait until later. The detail — the full action list, parameter semantics, examples, the mistakes that make a call fail — lives in `manuals/.md`, fetched on demand with `tool_manual("")`. + +The split is deliberate arithmetic: the description is paid for on every request whether or not the tool is called; the manual is paid for once, only when the tool is in play. + +When a hook points at a manual, the manual has to hold what the hook dropped — read `manuals/.md` before trimming a description. MCP servers work the same way: the system prompt keeps one line per server (`summary` in `mcp_servers.d/*.json`, defaulting to the first sentence of its `instructions`) and `tool_manual("")` returns the instructions in full; see [`mechanics.md`](mechanics.md). + +`tests/unit/tools/test_manual_drift.py` keeps the two ends honest — every manual is named after a real tool, and every `tool_manual` call cited in the docs resolves to something. + +--- + ## Scratchpad and Todo Both are per-session, backed by the PostgreSQL KV-store (`session_store` table) and survive server restarts. diff --git a/manuals/filesystem.md b/manuals/filesystem.md index 96f56e8..c568257 100644 --- a/manuals/filesystem.md +++ b/manuals/filesystem.md @@ -52,6 +52,28 @@ - **`find_up`** walks from `path` towards the root looking for an exact filename — the way to locate `pyproject.toml`/`.git` from a nested directory. - **`diff`** is a unified diff between two **files** — `path` against `destination` (both required; directories are rejected). +## `query` and `smart_edit` + +Both spend an LLM call, so both are for what the deterministic actions cannot do. + +`query` answers a question *instead of* returning the file — use it to avoid reading a large file whose content you only need one fact from: + +``` +{"action": "query", "path": "src/app.py", "question": "What does calculate() return?"} +{"action": "query", "path": "src/app.py", "question": "Where is class UserManager defined?"} +{"action": "query", "path": "deploy.sh", "question": "Which environment variables does this read?"} +``` + +`smart_edit` takes an instruction in natural language: + +``` +{"action": "smart_edit", "path": "src/app.py", "instruction": "Rename process to handle_request"} +{"action": "smart_edit", "path": "src/app.py", "instruction": "Add type hints to every function"} +{"action": "smart_edit", "path": "src/app.py", "instruction": "Replace the hardcoded URL with a constant BASE_URL"} +``` + +It reads the whole file and can touch more than you asked. When the change *is* expressible as exact text or line numbers, `edit`/`edit_lines` do it exactly, for free, and fail loudly instead of approximately. + ## Access In multi-user mode every path is resolved inside `user_data//`; a path that escapes it returns `Access denied: ... outside allowed paths`. Do not work around it by writing to `/tmp` or an absolute path — ask for the file to be placed in the workspace instead. In single-user/admin mode the allowed roots are the working directory and the configured paths. diff --git a/manuals/peer.md b/manuals/peer.md new file mode 100644 index 0000000..ba66d7b --- /dev/null +++ b/manuals/peer.md @@ -0,0 +1,61 @@ +# peer — Manual + +## What it does + +Talks to the other navi instances on the local network — machines that announced themselves to the same hive registry. Asking is real work on the other side: a full agent turn runs on the peer's machine, with the peer's own tools, files and services, and returns its answer as text. + +The **hive is only a phone book**. It answers "which machines exist"; the question itself travels directly to the peer's API port and the hive is never on the message path. That matters when the hive is down: `list` still works from cache, `ask` still reaches a peer you already know. + +## Actions + +| Action | What it does | +|---|---| +| `list` | Every machine in the swarm: name, address, online/offline, OS, core count. Our own machine is filtered out. | +| `status` | One peer's live state: uptime, version, port, machine facts, and whether the peer itself can reach the hive. | +| `ask` | A question for a peer. Answered by an agent run on that machine. | + +```json +{"action": "list"} +{"action": "status", "peer": "yuki"} +{"action": "ask", "peer": "yuki", "question": "Is the docker stack on this host healthy? Report container names and restart counts."} +``` + +## `ask` — how to write the question + +The peer sees **only your question**. It does not see this conversation, this machine, or anything you already know. So the question must be self-contained: + +- Name the machine implicitly ("the services on this host"), never "the server I mentioned". +- Say what you want back: a table, a yes/no with the evidence, the raw command output. +- Ask about **the peer's** machine. Anything you can check locally — your own files, your own services, your own process list — is cheaper and faster done here. Asking a peer about this machine's state is a normal way to waste a minute. +- Keep it under 4000 characters; that is the wire limit. + +An ask is not instant: the peer runs a real LLM turn, with a server-side ceiling of `PEER_ASK_TIMEOUT_SEC` (120 s) plus 30 s of client slack. Peers answer **one at a time** — a burst of asks queues behind a semaphore rather than fanning out onto the GPU. When the peer's turn hits its own iteration limit, the answer comes back with a note that it may be incomplete; that is a truncated answer, not a failure. + +For a question that may take a while, pass `background: true`: the call detaches immediately, returns a `task_id`, and the answer arrives later as a completion note (collect it with `tasks`). `peer` is one of the settings-configured backgroundable tools; only `ask` is worth detaching. + +## Errors, and what they mean + +| Error | Cause | +|---|---| +| `swarm_unconfigured` | `HIVE_URL` is empty, or `.swarm-key` is missing — this machine is not in a swarm. Nothing to retry. | +| `hive_unreachable` | The registry did not answer and no cached peer list exists yet. | +| `unknown_peer` | No peer by that name. The error lists the names that are known. | +| `peer_unreachable` | The peer is offline or its port is closed. Try `list` to see whether it is down. | +| `peer_refused` | The peer answered with an HTTP refusal — most often an ask that would loop back to itself. | +| `self_ask` | The name you asked is this very navi. Check locally instead. | + +A `list` result headed `(STALE — hive unreachable, last known list)` is the cached book: names and addresses are real, online flags may not be. + +## Loop safety + +`ask` cannot recurse. Two independent guards: + +1. The answering agent run is created with `peer` **excluded from its tools**, so the peer physically cannot ask onward. +2. An ask whose originating instance id matches the receiver's own uuid is refused at the boundary. + +That is also why a remote question can be answered but never relayed: a peer cannot forward your ask to a third machine. + +## Not to be confused with + +- **`spawn_agent`** — a sub-agent runs *here*, in this session's world, under the same navi. A peer is another machine with its own agent, reached over the network. +- **`ssh_exec`** — a raw shell on a remote host. A peer answers a *question* using its own agent and tools; you get judgement, not a command transcript. diff --git a/manuals/spawn_agent.md b/manuals/spawn_agent.md index ef9b957..16e22fb 100644 --- a/manuals/spawn_agent.md +++ b/manuals/spawn_agent.md @@ -7,6 +7,30 @@ **SYNCHRONOUS by default** — blocks until the sub-agent fully completes or times out (5 minutes hard limit). With `"background": true` the call detaches immediately and returns a `task_id` (`bt-...`); the sub-agent keeps running in isolation — see the `tasks` manual for `list`/`check`/`wait`/`cancel`, and PARALLELISM rules in your persona. Background sub-agents are capped separately (`tasks_max_spawn` per session) and their token usage is reported via `task_update`, not in your turn's token count. +## When to use it — and when not + +Use it when a step needs **3+ tool calls to complete as one logical unit**: a research question that takes several searches and reads, an ops task that chains SSH commands, an investigation with an unknown number of steps. + +Do **not** use it for a single tool call. If the step is "read this file", "run this test", "check whether the service is up" — call the tool yourself. A sub-agent costs a whole extra agent run and gives you back only its final text. + +The sub-agent also gets a **different, narrower tool set** than you have (see [Sub-agent tools](#sub-agent-tools)) and cannot see this conversation. If the step needs a tool the sub-agent will not have, or needs what you have already learned, do it yourself. + +## Choosing `profile_id` + +**Omit it by default.** The sub-agent then runs as the current session's profile — the right choice for most work, and the one that keeps its tools closest to yours. + +Set it to specialise: +- `server_admin` — remote ops over SSH, server state, infrastructure. +- `secretary` — research and writing, web-heavy work. +- `tool_developer` — implementing Navi tools and MCP servers. + +If your plan named a profile for this step, pass that exact id. The profile decides the sub-agent's model, system prompt and available tools — a wrong pick shows up as the sub-agent lacking what it needs. + +```json +{"task": "...", "briefing": "..."} // current profile +{"task": "...", "profile_id": "server_admin", "briefing": "..."} // specialised +``` + ## Parameters | Parameter | Required | Description | @@ -17,6 +41,7 @@ | `system_prompt` | no | Role specialisation injected into the sub-agent's system prompt between the executor persona and the briefing (e.g. "You are a security auditor. Report findings by severity."). | | `max_iterations` | no | Tool-call iteration limit (default: **40**). | | `background` | no | `true` → detach immediately, return `task_id`; result arrives later as a completion note (use `tasks` to check/wait/cancel). | +| `inherit_system_prompt` | no | `true` → the parent profile's full system prompt is the **base layer**, with the sub-agent's own specialisation overlaid on top. Use when the sub-agent must keep the parent's personality, rules and workflow. Default `false`: the sub-agent uses only its own `subagent_system_prompt` and ignores the parent's entirely. | ## Sub-agent system prompt structure diff --git a/mcp_servers.d/gnexus-book.json b/mcp_servers.d/gnexus-book.json index e4adc96..a34fd17 100644 --- a/mcp_servers.d/gnexus-book.json +++ b/mcp_servers.d/gnexus-book.json @@ -1,7 +1,10 @@ { "transport": "sse", "url": "http://192.168.1.170:8001/sse", - "user_key": { "header": "Authorization", "prefix": "Bearer " }, + "user_key": { + "header": "Authorization", + "prefix": "Bearer " + }, "groups": { "read": [ "search_docs", @@ -27,5 +30,6 @@ "delete_pending_change" ] }, - "instructions": "MANDATORY for profiles that expose gnexus-book tools: Before answering any question about infrastructure, servers, services, networks, documentation, or system inventory, call gnexus-book tools first.\n\nUse only gnexus-book tool names that are present in the current tool schema. In Navi they are exposed with the mcp__gnexus_book__ prefix (example: mcp__gnexus_book__search_docs), but each profile may expose only some groups. Do not invent or call gnexus-book tools that are not in the current tool list.\n\nQuery mapping by capability:\n- Status or facts about a server/service → search docs first, then read a specific doc or inventory item if those tools are available.\n- Service placement or topology → list inventory and relationships if available.\n- Documentation changes → read the target doc first, then propose a doc or inventory change if write tools are available.\n- Freshness questions → use freshness checks if available.\n- Repository validation/status → use repository tools only if they are available in the current tool schema; otherwise skip this step and continue with available read/write tools.\n\nDo not rely on memory for infrastructure facts. Memory is only for personal user facts and preferences. Always pull infrastructure state from gnexus-book when these tools are available to the active profile.\n\nDo not store raw secrets in documentation.\n\nABSOLUTE RULE — NEVER bypass MCP tools:\nYou MUST NOT use filesystem, terminal, code_exec, or any direct file access to read or write gnexus-book files. The MCP tools are the ONLY valid interface to this knowledge base. Violating this rule bypasses validation, corrupts repository state, and breaks consistency guarantees.\n- To read: use mcp__gnexus_book__search_docs, mcp__gnexus_book__read_doc, mcp__gnexus_book__list_inventory, mcp__gnexus_book__get_inventory_item.\n- To write: use mcp__gnexus_book__propose_doc_change, mcp__gnexus_book__propose_inventory_item_change, mcp__gnexus_book__apply_pending_change, mcp__gnexus_book__commit_changes.\n- NEVER call filesystem write, filesystem smart_edit, terminal, or code_exec on gnexus-book paths.\n\nBefore the final response, decide whether tool execution revealed stable reusable infrastructure facts, service configurations, or relationships. If yes and gnexus-book write tools are available, persist them before answering. If gnexus-book write tools are not available, report the facts that should be persisted. If the fact is user-specific rather than infrastructure documentation, use the memory tool instead. Choose the target based on scope, not habit." + "instructions": "MANDATORY for profiles that expose gnexus-book tools: Before answering any question about infrastructure, servers, services, networks, documentation, or system inventory, call gnexus-book tools first.\n\nUse only gnexus-book tool names that are present in the current tool schema. In Navi they are exposed with the mcp__gnexus_book__ prefix (example: mcp__gnexus_book__search_docs), but each profile may expose only some groups. Do not invent or call gnexus-book tools that are not in the current tool list.\n\nQuery mapping by capability:\n- Status or facts about a server/service → search docs first, then read a specific doc or inventory item if those tools are available.\n- Service placement or topology → list inventory and relationships if available.\n- Documentation changes → read the target doc first, then propose a doc or inventory change if write tools are available.\n- Freshness questions → use freshness checks if available.\n- Repository validation/status → use repository tools only if they are available in the current tool schema; otherwise skip this step and continue with available read/write tools.\n\nDo not rely on memory for infrastructure facts. Memory is only for personal user facts and preferences. Always pull infrastructure state from gnexus-book when these tools are available to the active profile.\n\nDo not store raw secrets in documentation.\n\nABSOLUTE RULE — NEVER bypass MCP tools:\nYou MUST NOT use filesystem, terminal, code_exec, or any direct file access to read or write gnexus-book files. The MCP tools are the ONLY valid interface to this knowledge base. Violating this rule bypasses validation, corrupts repository state, and breaks consistency guarantees.\n- To read: use mcp__gnexus_book__search_docs, mcp__gnexus_book__read_doc, mcp__gnexus_book__list_inventory, mcp__gnexus_book__get_inventory_item.\n- To write: use mcp__gnexus_book__propose_doc_change, mcp__gnexus_book__propose_inventory_item_change, mcp__gnexus_book__apply_pending_change, mcp__gnexus_book__commit_changes.\n- NEVER call filesystem write, filesystem smart_edit, terminal, or code_exec on gnexus-book paths.\n\nBefore the final response, decide whether tool execution revealed stable reusable infrastructure facts, service configurations, or relationships. If yes and gnexus-book write tools are available, persist them before answering. If gnexus-book write tools are not available, report the facts that should be persisted. If the fact is user-specific rather than infrastructure documentation, use the memory tool instead. Choose the target based on scope, not habit.", + "summary": "Infrastructure knowledge base. Answer questions about servers, services, networks or docs ONLY from its tools — never from memory — and never touch its files with filesystem, terminal or code_exec. Full rules: tool_manual(\"gnexus-book\")." } diff --git a/mcp_servers.d/navi-web.json b/mcp_servers.d/navi-web.json index 77b9dab..6ff1592 100644 --- a/mcp_servers.d/navi-web.json +++ b/mcp_servers.d/navi-web.json @@ -20,5 +20,6 @@ "http_request" ] }, - "instructions": "Navi Web MCP server provides web search, browsing, and raw HTTP tools.\n\nUse it when the task involves:\n- searching the web for current info, docs, or real-time data;\n- opening a URL in a browser to read human-readable content;\n- making REST API calls, webhooks, or raw HTTP requests.\n\nWorkflow:\n1. search — find relevant pages or facts.\n2. view — open promising URLs to read full content.\n3. request — call APIs or services requiring headers/auth.\n\nAll three tools are stateless and work with public URLs.\nNo session_id or filesystem paths are required.\n\nABSOLUTE RULE — NEVER bypass MCP tools:\nYou MUST NOT use filesystem, terminal, code_exec, or any direct file access to read or write web content. Use only the MCP tools listed above." + "instructions": "Navi Web MCP server provides web search, browsing, and raw HTTP tools.\n\nUse it when the task involves:\n- searching the web for current info, docs, or real-time data;\n- opening a URL in a browser to read human-readable content;\n- making REST API calls, webhooks, or raw HTTP requests.\n\nWorkflow:\n1. search — find relevant pages or facts.\n2. view — open promising URLs to read full content.\n3. request — call APIs or services requiring headers/auth.\n\nAll three tools are stateless and work with public URLs.\nNo session_id or filesystem paths are required.\n\nABSOLUTE RULE — NEVER bypass MCP tools:\nYou MUST NOT use filesystem, terminal, code_exec, or any direct file access to read or write web content. Use only the MCP tools listed above.", + "summary": "Web search, page browsing and raw HTTP. Never fetch web content with filesystem, terminal or code_exec — these are the only valid interface. Details: tool_manual(\"navi-web\")." } diff --git a/mcp_servers.d/navi_ui.json b/mcp_servers.d/navi_ui.json index d97d5f8..c96e84b 100644 --- a/mcp_servers.d/navi_ui.json +++ b/mcp_servers.d/navi_ui.json @@ -2,7 +2,10 @@ "transport": "streamable_http", "url": "http://127.0.0.1:8098/mcp", "groups": { - "ui": ["render_component"] + "ui": [ + "render_component" + ] }, - "instructions": "Tool: render_component. Arguments: {\"component_name\": \"card_grid\", \"payload\": {...}}. Only use component_name=\"card_grid\". session_id is injected automatically. For card_grid, payload must contain a non-empty \"cards\" array. Each card must have string \"id\" and string \"title\". Optional card fields: \"subtitle\", \"image\" (URL), \"meta\" (array of {\"label\", \"value\"}), \"description\" (short string), \"details\" (array of {\"label\", \"value\"} shown in modal), \"actions\" (array of {\"label\", \"url\"}). Limit card grid to 4 cards. After rendering the grid, still give a brief text summary and ask what the user wants next." + "instructions": "Tool: render_component. Arguments: {\"component_name\": \"card_grid\", \"payload\": {...}}. Only use component_name=\"card_grid\". session_id is injected automatically. For card_grid, payload must contain a non-empty \"cards\" array. Each card must have string \"id\" and string \"title\". Optional card fields: \"subtitle\", \"image\" (URL), \"meta\" (array of {\"label\", \"value\"}), \"description\" (short string), \"details\" (array of {\"label\", \"value\"} shown in modal), \"actions\" (array of {\"label\", \"url\"}). Limit card grid to 4 cards. After rendering the grid, still give a brief text summary and ask what the user wants next.", + "summary": "Renders a component in the user's chat window (card_grid). Use it when an answer is better shown as cards than as prose." } diff --git a/navi/core/context_builder.py b/navi/core/context_builder.py index d9212a6..300c161 100644 --- a/navi/core/context_builder.py +++ b/navi/core/context_builder.py @@ -358,24 +358,30 @@ return Message(role="system", content="\n".join(lines)) def _mcp_context_msg(self, profile: "AgentProfile | None" = None) -> "Message | None": - """Build a system message with MCP server instructions. + """Build a system message naming the MCP servers this run can reach. - Combines server-provided instructions (from MCP initialize handshake) - with overlay instructions from ``mcp_servers.d/*.json``. + One line per server — what it is for — and a pointer to the full text. The + instructions themselves used to be inlined here, which cost ~8 KB of prose + on *every* request of a heavy profile; they are now read on demand with + ``tool_manual("")``. """ if not self._mcp_manager: return None if profile is not None and not profile.get_agent_tools().mcp: return None server_names = set(profile.get_agent_tools().mcp.keys()) if profile is not None else None - instructions = self._mcp_manager.get_instructions(server_names) - if not instructions: + summaries = self._mcp_manager.get_summaries(server_names) + if not summaries: return None - lines = ["[MCP servers — external knowledge sources]"] - for name, text in instructions.items(): + lines = [ + "[MCP servers — external knowledge sources]", + "Their tools are named mcp____. Before using a server's tools, " + 'and whenever it is unclear which of them fits, read its full instructions ' + 'with tool_manual("").', + ] + for name in sorted(summaries): lines.append("") - lines.append(f"## {name}") - lines.append(text) + lines.append(f"## {name} — {summaries[name]}") return Message(role="system", content="\n".join(lines)) def _truncate_oversized(self, conv: list[Message]) -> list[Message]: diff --git a/navi/mcp/config.py b/navi/mcp/config.py index 931a3ab..031711f 100644 --- a/navi/mcp/config.py +++ b/navi/mcp/config.py @@ -55,10 +55,18 @@ # Profiles reference groups by name instead of listing individual tools. groups: dict[str, list[str]] = Field(default_factory=dict) - # Overlay instructions injected into Navi's system prompt alongside the - # instructions provided by the MCP server itself during the initialize handshake. + # Overlay instructions, merged with the instructions the MCP server itself + # provides during the initialize handshake. They are NOT injected into the + # system prompt any more — only one line per server is (see ``summary``), + # and the full text is fetched on demand with ``tool_manual("")``. instructions: str | None = None + # The one line that stays in the system prompt for this server: what it is + # for, in a few words. Defaults to the first sentence of ``instructions``. + # Set it explicitly when that first sentence is a poor hook (a long + # "MANDATORY ..." directive, or a server with no instructions at all). + summary: str | None = None + # Per-user credential slot (Bring Your Own Key). None = the server needs no # user key; all users share the default plaintext credential above. user_key: McpUserKey | None = None diff --git a/navi/mcp/manager.py b/navi/mcp/manager.py index 87e8b51..71c70b9 100644 --- a/navi/mcp/manager.py +++ b/navi/mcp/manager.py @@ -2,6 +2,7 @@ import asyncio import logging +import re import time from collections import OrderedDict from pathlib import Path @@ -12,6 +13,28 @@ logger = logging.getLogger(__name__) +# Longest one-line server summary kept in the system prompt. Servers whose first +# sentence is a poor hook should set ``summary`` in their config instead. +_SUMMARY_LIMIT = 200 + + +def summarize_instructions(text: str, limit: int = _SUMMARY_LIMIT) -> str: + """The first sentence of a server's instructions, capped at *limit* chars. + + This is the line the system prompt keeps for a server; the rest of the text is + one ``tool_manual("")`` call away. Everything a server says about itself + is aimed at a model deciding *whether* to reach for it, and that decision only + needs the opening claim — the query mappings and workflows underneath it are + read when the server is actually used. + """ + flat = " ".join((text or "").split()) + if not flat: + return "" + head = re.split(r"(?<=[.!?])\s+", flat, maxsplit=1)[0] + if len(head) > limit: + head = head[:limit].rsplit(" ", 1)[0].rstrip(" ,;:.—-") + "…" + return head + # Per-user client cache (BYOK): hard size cap to bound stdio subprocesses. _USER_CLIENT_LIMIT = 32 # Per-user clients unused for this long are dropped (health-check tick). @@ -196,6 +219,55 @@ out[name] = "\n".join(parts) return out + def configured_servers(self) -> list[str]: + """Every server named in ``mcp_servers.d/*.json``, connected or not.""" + return sorted(self._get_configs()) + + def get_summaries(self, server_names: list[str] | set[str] | None = None) -> dict[str, str]: + """One line per server, for the system prompt. + + The config's explicit ``summary`` wins; otherwise it is the first sentence + of the server's instructions. Servers with neither are omitted rather than + listed as a bare name — a name with no claim attached tells the agent nothing + it could not read off the tool prefix. + + Full instructions stay reachable: ``tool_manual("")``. + """ + configs = self._get_configs() + instructions = self.get_instructions(server_names) + if server_names is None: + names = set(instructions) + else: + names = set(server_names) + names |= {name for name, cfg in configs.items() if cfg.summary} + + out: dict[str, str] = {} + for name in sorted(names): + cfg = configs.get(name) + line = (cfg.summary if cfg and cfg.summary else "") or summarize_instructions( + instructions.get(name, "") + ) + if line: + out[name] = line + return out + + def server_tool_names(self, server_name: str) -> list[str]: + """Tool names this server exposes, from the static config groups. + + Read from the config rather than the live server so that a disconnected + server still documents what it would offer once it reconnects, and so a + profile that enables only part of it sees the whole catalogue. + """ + cfg = self._get_configs().get(server_name) + if cfg is None: + return [] + seen: list[str] = [] + for group_tools in cfg.groups.values(): + for tool_name in group_tools: + if tool_name not in seen: + seen.append(tool_name) + return sorted(seen) + async def call_tool( self, server_name: str, diff --git a/navi/tools/content_publish.py b/navi/tools/content_publish.py index 307b72e..560fa6d 100644 --- a/navi/tools/content_publish.py +++ b/navi/tools/content_publish.py @@ -18,27 +18,13 @@ class ContentPublishTool(Tool): name = "content_publish" description = ( - "Publish a file for inline viewing in the chat client. " - "Use this when you generate or produce content the user will want to see interactively " - "(3D models, HTML pages, SVG graphics, images, videos, PDFs, etc.).\n\n" - "There are two different file areas: workspace/ is for persistent private working files; " - "the session directory is for files visible to the user in this chat. " - "IMPORTANT — the file MUST already be inside the current session directory. " - "The default path is session_files/{session_id}/, but the root is configured by " - "SESSION_FILES_DIR. This tool does NOT copy from workspace/. " - "Before publishing, write, copy, or move the final file into the session folder. " - "If a file with the same name already exists in the session directory, choose a different name " - "or check the directory contents first with `filesystem list `.\n\n" - "Best practices:\n" - "- Use workspace/ for drafts and reusable work files\n" - "- Use the session directory for final artifacts the user should view now\n" - "- If a file already exists elsewhere and only needs a download link, use share_file instead\n" - "- If a file exists elsewhere but needs an inline viewer, copy it into the session directory first\n" - "- Use descriptive filenames (e.g., 'sales_chart.svg' not 'file.svg')\n" - "- After publishing, you can edit the file directly and the user will see changes immediately\n" - "- For images, use PNG or JPEG; for interactive content, use HTML or SVG\n" - "- For STL created from OpenSCAD, pass source_filename only if the .scad file exists in the session directory; " - "omit it for downloaded STL files or when no source exists" + "Publish a file from this session's directory for inline viewing in the chat — HTML, " + "SVG, PDF, images, video, STL models — as an interactive card.\n\n" + "The file MUST already be in the session directory: this tool registers it and does " + "NOT copy from workspace/. Write, copy or move it there first. Editing it afterwards " + "updates the card in place, so publish once.\n\n" + "A file kept elsewhere that only needs a download link is share_file's job.\n\n" + 'Filenames, re-publishing, STL sources: tool_manual("content_publish").' ) parameters = { "type": "object", diff --git a/navi/tools/filesystem.py b/navi/tools/filesystem.py index 87be385..404b377 100644 --- a/navi/tools/filesystem.py +++ b/navi/tools/filesystem.py @@ -342,24 +342,13 @@ class FilesystemTool(Tool): name = "filesystem" description = ( - "Operate the local filesystem. " - "Editing policy — prefer the cheapest deterministic method; reach for " - "smart_edit LAST:\n" - "1. Edit by exact text — 'edit' with 'old' (must occur exactly once) and " - "'new'. Read the file first and copy old text verbatim. This is the " - "default for almost all edits.\n" - "2. Edit by line numbers — 'edit_lines' with an operations array, when " - "you know the exact lines (e.g. 'change line 15'). Deterministic, no AI " - "call.\n" - "3. smart_edit (AI) — ONLY as a fallback when the change genuinely cannot " - "be expressed as exact text or line numbers (e.g. 'rename a symbol " - "everywhere', 'add type hints to every function'). It costs an LLM call " - "and reads the whole file, so try edit/edit_lines first.\n" - "4. Create or fully rewrite a file — 'write' (pass 'content').\n" - "5. Extract info from a file — 'query' (pass 'question').\n" - "6. Everything else — read, append, list, find, find_up, grep, diff, " - "info, copy, move, delete, exists, mkdir.\n" - "Tip: call 'info' before reading an unknown file to check size." + "Operate the local filesystem: read, write, edit, search, and manage files. " + "Editing policy — cheapest deterministic method first, smart_edit LAST: use " + "'edit' (exact text, the default) or 'edit_lines' (known line numbers); both " + "are deterministic. Reach for 'smart_edit' ONLY when the change cannot be " + "expressed as text or line numbers — it costs an LLM call and reads the whole " + "file. Call 'info' before reading an unknown file. Full action list, argument " + 'rules and access limits: tool_manual("filesystem").' ) parameters = { "type": "object", @@ -407,10 +396,8 @@ "operations": { "type": "array", "description": ( - "JSON array of line-based edit operations (required for edit_lines). " - 'Each op is {"op": "replace"|"delete"|"insert", "start": int, "end": int, "content": str}. ' - "Line numbers are 1-based and inclusive. Use edit_lines for fast deterministic edits " - "when you know the exact lines (e.g. 'change line 15 from X to Y')." + "Line-based edits (required for edit_lines). 1-based, inclusive. " + "'after' applies to insert." ), "items": { "type": "object", @@ -426,11 +413,11 @@ }, "numbered": { "type": "boolean", - "description": "Include 1-based line numbers in read output (default true). Set false to read raw file content without the number column.", + "description": "Number the lines in read output (default true). Set false to copy text into 'edit'.", }, "destination": { "type": "string", - "description": "Target path for move action.", + "description": "Second path: the target for move/copy, or the file to compare against for diff.", }, "pattern": { "type": "string", @@ -455,23 +442,13 @@ "question": { "type": "string", "description": ( - "Natural language question about the file's content (for query). " - "Examples: 'What does function calculate() return?', " - "'On which line is class UserManager defined?', " - "'What environment variables does this script read?', " - "'Are there any hardcoded passwords?'" + "Natural language question about the file (for query) — the answer comes " + "back instead of the file." ), }, "instruction": { "type": "string", - "description": ( - "Natural language edit instruction (for smart_edit). " - "Examples: 'Rename function process to handle_request', " - "'Add type hints to all function signatures', " - "'Replace the hardcoded URL with a constant BASE_URL', " - "'Delete the block comment on lines 10-20', " - "'Add logging to the save() method'" - ), + "description": "Natural language edit instruction (for smart_edit).", }, }, "required": ["action", "path"], diff --git a/navi/tools/manage_recall.py b/navi/tools/manage_recall.py index 5855bcf..608eafe 100644 --- a/navi/tools/manage_recall.py +++ b/navi/tools/manage_recall.py @@ -8,20 +8,10 @@ class ManageRecallTool(Tool): name = "manage_recall" description = ( - "Manage scheduled recalls for the current session. " - "Actions: cancel (remove pending), skip (defer recurring by one interval), list (show all). " - "Only one pending recall per session is allowed, so you often need cancel before creating a new one.\n\n" - "Actions explained:\n" - " cancel — deletes the pending recall. Use when the user no longer needs the callback, " - " or when you want to replace an existing recall with a new one.\n" - " skip — advances a recurring recall by interval_seconds. Use when the current scheduled check " - " is unnecessary (user already confirmed the build passed). Only works on recurring.\n" - " list — shows pending/fired/cancelled recalls for the session. Use before scheduling to verify " - " whether a recall already exists, or before telling the user they have no pending recalls.\n\n" - "Rules:\n" - " • Always call list before scheduling if you are unsure whether a recall already exists.\n" - " • The standard update pattern is: manage_recall(action=cancel) → schedule_recall(...).\n" - " • skip only works on recurring recalls; for one-time recalls use cancel to abort." + "Manage this session's scheduled recalls. Only one may be pending, so cancel " + "before scheduling a new one, and list first if you are unsure whether one " + "exists.\n\n" + 'Actions and the cancel → schedule pattern: tool_manual("manage_recall").' ) parameters = { "type": "object", diff --git a/navi/tools/peer.py b/navi/tools/peer.py index 96f2791..0b9f34f 100644 --- a/navi/tools/peer.py +++ b/navi/tools/peer.py @@ -52,17 +52,16 @@ class PeerTool(Tool): name = "peer" description = ( - "Communicate with other navi instances in the swarm (the local network of " - "navi agents this machine belongs to).\n\n" + "Communicate with the other navi instances in the swarm (the agents on this " + "local network).\n\n" "Actions:\n" - "· list — all machines in the swarm with name, address, online status\n" - "· status — one peer's live state: uptime, version, machine facts\n" - "· ask — ask a peer a question; its agent investigates " - "on ITS machine and returns a text answer. Use for anything this peer " - "can check better than you: its services, files, hardware, local state.\n\n" - "Address the peer by its swarm name (e.g. 'yuki'). " - "Answers can take up to a couple of minutes - the peer runs a real agent " - "turn. Do not ask peers about this machine's state; check it locally." + "· list — the machines in the swarm, with online status\n" + "· status — one peer's live state: uptime, version, machine, hive view\n" + "· ask — its agent investigates ON ITS OWN machine and " + "answers. Address it by swarm name (e.g. 'yuki'); a real agent turn takes up " + "to a minute, or pass background=true to detach. Never ask a peer about this " + "machine's state — check that locally.\n\n" + 'Errors, question-writing rules, swarm setup: tool_manual("peer").' ) parameters = { "type": "object", diff --git a/navi/tools/plan.py b/navi/tools/plan.py index 6d52994..50db7cf 100644 --- a/navi/tools/plan.py +++ b/navi/tools/plan.py @@ -142,21 +142,15 @@ class PlanTool(Tool): name = "plan" description = ( - "Run the planner: decompose a task into a structured execution plan (milestones + steps " - "with TOOL/AGENT/SELF executors) and auto-populate the todo.\n\n" - "Call this BEFORE starting execution when:\n" - "- The task is non-trivial: multiple steps, several files/systems, research, or real risk.\n" - "- The work needs decomposition, ordering, or sub-agent scoping decided up front.\n\n" - "Skip it for trivial work: single-file edits, one-off commands, questions, casual chat.\n\n" - "Re-plan mid-task by passing `reason` — what you discovered that invalidates the remaining " - "plan (a step turned out unnecessary, the real problem differs from the assumed one, new " - "constraints appeared) — plus optionally `updated_goal`. The new plan replaces the todo; " - "completed work is preserved in the scratchpad/conversation.\n\n" - "Costs 2 LLM calls (analysis + execution plan) — use selectively, like `reflect`.\n\n" - "Do NOT call plan for a single failed step: revise the todo inline instead. For a small " - "adjustment (drop/merge/reorder 1-2 steps) edit the `todo` directly — plan is for when the " - "remaining plan's overall structure is wrong (or missing). Use `reflect` when you're unsure " - "what's wrong (it surfaces assumptions, no plan change); use `plan` when you need a plan." + "Run the planner: decompose the task into milestones and steps (each tagged TOOL, " + "AGENT or SELF) and populate the todo with them. Costs 2 LLM calls.\n\n" + "Call it BEFORE execution when the work is non-trivial — several steps, several " + "files or systems, research, or real risk. Skip it for a single edit, a one-off " + "command, a question or chat.\n\n" + "Re-plan by passing `reason` when the remaining plan's overall shape is wrong. For " + "a one- or two-step adjustment edit the `todo` directly instead, and use `reflect` " + "when you are unsure what is wrong.\n\n" + 'What it returns, and the rules around re-planning: tool_manual("plan").' ) parameters = { "type": "object", diff --git a/navi/tools/reflect.py b/navi/tools/reflect.py index 1c45055..26c548c 100644 --- a/navi/tools/reflect.py +++ b/navi/tools/reflect.py @@ -77,18 +77,13 @@ class ReflectTool(Tool): name = "reflect" description = ( - "Get three independent expert perspectives on a situation before planning or when stuck.\n\n" - "Call this when:\n" - "- About to plan a complex or ambiguous task\n" - "- Stuck on a problem and need a fresh angle\n" - "- Unsure whether your approach is right\n\n" - "Three advisors analyse your situation in parallel:\n" - "· Critic — challenges assumptions, surfaces risks and flaws\n" - "· Pragmatist — finds the simplest path, cuts unnecessary complexity\n" - "· Detailer — spots missing requirements, edge cases, and gaps\n\n" - "IMPORTANT: The `assumptions` field is mandatory and is the most valuable input. " - "List every belief you are acting on without having verified it. " - "The act of listing assumptions often reveals the problem itself." + "Get three independent expert perspectives on a situation — the Critic on risk and " + "assumptions, the Pragmatist on the simplest path, the Detailer on gaps and edge " + "cases — before planning, or when you are stuck.\n\n" + "`assumptions` is mandatory and is the most valuable input: list every belief you " + "are acting on without having verified it. Listing them often reveals the problem " + "itself.\n\n" + 'When it pays off and when it does not: tool_manual("reflect").' ) parameters = { "type": "object", diff --git a/navi/tools/schedule_recall.py b/navi/tools/schedule_recall.py index 3034622..52aa914 100644 --- a/navi/tools/schedule_recall.py +++ b/navi/tools/schedule_recall.py @@ -23,26 +23,13 @@ class ScheduleRecallTool(Tool): name = "schedule_recall" description = ( - "Schedule a headless callback for the current session. " - "At the chosen time Navi wakes up, reads your self-instruction, and continues working using other tools. " - "Only one pending recall per session is allowed — cancel the old one first if you need a new timer.\n\n" - "CORE PRINCIPLE: this is a TOOL for CONTINUING WORK, not a chat reminder.\n" - " BAD message: 'Tell the user that 2 hours have passed.'\n" - " GOOD message: 'Read /tmp/build.log with filesystem. If errors, read last 50 lines and report. Otherwise confirm success.'\n\n" - "Call types:\n" - " once — single delayed action (check logs in 30m, continue after reboot).\n" - " recurring — periodic action with interval_seconds (poll API every 5 min, check inbox every 15 min).\n" - " immediate — fire ASAP; use to offload heavy multi-tool work without blocking the chat.\n\n" - "Popular scenarios:\n" - " 1. Hit iteration limit? Schedule immediate with context 'Continue Nginx config from step 3...'\n" - " 2. Waiting for a build/export? Schedule once for estimated finish time and tell yourself which files to check.\n" - " 3. Periodic monitoring? Schedule recurring 900s (15 min) with context 'Read /var/log/app/errors.log...'\n" - " 4. Heavy task with 20+ tool calls? Use immediate so user can keep chatting while you work headlessly.\n\n" - "Rules:\n" - " • additional_context_message must read like a todo item for yourself.\n" - " • Mention specific tools, files, or URLs — future-you has the same tools but not your short-term memory.\n" - " • Multi-phase tasks: chain recalls — phase 1 runs, then schedules recall for phase 2.\n" - " • Only one pending recall per session. Use manage_recall cancel before scheduling a new one." + "Schedule a headless callback for this session: at the chosen time Navi wakes, " + "reads the self-instruction you left, and continues working with other tools. " + "Only one pending recall per session — cancel the old one first.\n\n" + "This is for CONTINUING WORK, not a chat reminder. Write " + "additional_context_message as a todo for yourself, naming the tools and files " + "future-you will need.\n\n" + 'Scenarios and rules: tool_manual("schedule_recall").' ) parameters = { "type": "object", @@ -55,21 +42,17 @@ "when": { "type": "string", "description": ( - "When to trigger. Prefer RELATIVE formats — they work correctly regardless of timezone differences:\n" - " • '30m', '2h 15m', '1d 6h' — compact relative\n" - " • 'in 3 hours', 'in 2 days' — natural language relative\n" - " • 'tomorrow at 09:00' — tomorrow at specific local time\n" - "Use ABSOLUTE ISO datetime ONLY if you know the user's exact timezone offset (e.g., '2026-05-15T14:00:00+03:00'). " - "Never use UTC-only ISO like '2026-05-15T14:00:00+00:00' unless the user explicitly confirmed UTC.\n" - "Ignored for immediate calls." + "Relative formats are timezone-proof and preferred: '30m', '2h 15m', " + "'1d 6h', 'in 3 hours', 'tomorrow at 09:00'. Absolute ISO only if you " + "know the user's exact offset — never a bare UTC ISO. Ignored for " + "immediate." ), }, "timezone_offset": { "type": "string", "description": ( - "User's timezone offset in ±HH:MM format (e.g., '+03:00', '-05:00'). " - "Required when using absolute times or 'tomorrow at HH:MM' if the user's timezone is known. " - "If not provided, the server assumes UTC." + "Offset in ±HH:MM (e.g. '+03:00'). Required for absolute times and " + "'tomorrow at HH:MM'. Without it the server assumes UTC." ), }, "interval_seconds": { diff --git a/navi/tools/scratchpad.py b/navi/tools/scratchpad.py index b925ae9..6142a7d 100644 --- a/navi/tools/scratchpad.py +++ b/navi/tools/scratchpad.py @@ -35,23 +35,13 @@ class ScratchpadTool(Tool): name = "scratchpad" description = ( - "Working memory for the current session — for facts discovered mid-task, not for progress tracking. " - "Use this to save intermediate findings (file paths, URLs, error details) so they are not lost across tool calls. " - "Read before composing a final answer — saved findings may contain facts needed for the response.\n\n" - "JSON schema:\n" - " action: 'write' | 'append' | 'read' | 'clear'\n" - " section: string — which section to target (goal, findings, artifacts, errors, main). Defaults to 'main'.\n" - " content: string — text to write or append. REQUIRED for write and append.\n\n" - "Examples (copy this structure exactly):\n" - " {\"action\": \"write\", \"section\": \"findings\", \"content\": \"The config file is at /etc/app/config.yml\"}\n" - " {\"action\": \"append\", \"section\": \"findings\", \"content\": \"Port is 8080\"}\n" - " {\"action\": \"read\", \"section\": \"findings\"}\n" - " {\"action\": \"clear\", \"section\": \"errors\"}\n\n" - "Common mistakes to avoid:\n" - " - Do NOT omit 'content' on write/append — the call will fail.\n" - " - Do NOT pass the text as the action value (e.g. {\"action\": \"your text here\"}). " - " The action MUST be exactly one of: write, append, read, clear.\n" - " - Do NOT put the text inside 'section' — section is the category name, content is the text." + "Working memory for this session — save facts discovered mid-task (file paths, " + "URLs, error details) so they survive across tool calls, and read it before " + "composing a final answer: the facts the answer needs may be there. For progress " + "tracking use `todo`, not this.\n\n" + "Actions (write/append/read/clear), the standard sections, examples and the " + "argument mistakes that make the call fail: " + 'tool_manual("scratchpad").' ) parameters = { "type": "object", diff --git a/navi/tools/share_file.py b/navi/tools/share_file.py index 65e38ed..e62454a 100644 --- a/navi/tools/share_file.py +++ b/navi/tools/share_file.py @@ -22,21 +22,13 @@ class ShareFileTool(Tool): name = "share_file" description = ( - "Copy an existing local file into the current session directory and return a direct download link. " - "Use this when the user should receive a file to keep: archives, reports, exports, source bundles, " - "datasets, PDFs, CSV/JSON files, or other generated artifacts.\n\n" - "Mechanics: share_file takes an ABSOLUTE source path, copies that file into " - "SESSION_FILES_DIR/{session_id}/ under the optional clean filename, and returns a URL at " - "/api/sessions/{session_id}/files/{filename}. The source file remains where it was." - "If a file with the same name already exists in the session directory, share_file creates " - "a numbered filename instead of overwriting it. Max file size is SHARE_FILE_MAX_SIZE_MB " - "(default 1024 MB / 1 GB).\n\n" - "Do not confuse this with content_publish: share_file is for download links and may copy " - "from elsewhere; content_publish is for inline viewer cards and only registers a file that " - "already exists in the session directory.\n\n" - "IMPORTANT — path must be an ABSOLUTE path (e.g. /home/user/file.zip). Relative paths are rejected. " - "If you only know a relative path, resolve it first: use filesystem(action='info') or " - "terminal('realpath ') to get the absolute path, then call share_file." + "Copy a local file into this session's directory and return a download link — for a " + "file the user should keep: an archive, report, export, dataset, PDF, source bundle.\n\n" + "`path` must be ABSOLUTE; a relative one is rejected. Resolve it first with " + "filesystem(action='info') or terminal('realpath ...').\n\n" + "For an inline viewer card rather than a download, use content_publish, which " + "publishes a file that is already in the session directory.\n\n" + 'How to present the link: tool_manual("share_file").' ) parameters = { "type": "object", diff --git a/navi/tools/spawn_agent.py b/navi/tools/spawn_agent.py index b6b699a..31d26e5 100644 --- a/navi/tools/spawn_agent.py +++ b/navi/tools/spawn_agent.py @@ -18,26 +18,16 @@ class SpawnAgentTool(Tool): name = "spawn_agent" description = ( - "Delegate EXACTLY ONE step of your plan to an isolated sub-agent.\n\n" - "CRITICAL: one spawn_agent call = one plan step. " - "If your plan has three AGENT steps, you make three separate spawn_agent calls — " - "one per step. Never bundle multiple plan steps into a single sub-agent.\n\n" - "SYNCHRONOUS by default — blocks until the sub-agent fully completes. " - "For long research/ops sub-agents set background=true: the call returns " - "a task_id immediately, the sub-agent runs detached, and the result " - "arrives via the tasks tool and an automatic completion note.\n\n" - "USER CANNOT SEE sub-agent output — synthesise findings into your own response.\n\n" - "USE when a step requires 3+ tool calls to complete as a single logical unit. " - "DO NOT USE for a single tool call — call the tool directly.\n\n" - "PROFILE SELECTION: omit profile_id by default — the sub-agent then runs as the current " - "session's profile (e.g. navi_code for local coding), which is the right choice for most " - "code work. Set profile_id only to specialise: 'server_admin' for remote ops, 'secretary' " - "for research/writing, 'tool_developer' for Navi tool implementation. If your plan named a " - "specific profile, pass that exact profile_id.\n\n" - "Examples (copy this structure exactly):\n" - ' {\"task\": \"...\", \"profile_id\": \"server_admin\", \"briefing\": \"...\"}\n' - ' {\"task\": \"...\", \"profile_id\": \"tool_developer\"}\n' - ' {\"task\": \"...\"} ← omit profile_id to run as the current session\'s profile\n\n' + "Delegate EXACTLY ONE step of your plan to an isolated sub-agent with its own " + "context and tool loop.\n\n" + "Use it when a step needs 3+ tool calls as one logical unit — never for a single " + "call, and one call per plan step, never bundling several steps into one " + "sub-agent.\n\n" + "Synchronous by default (blocks until it finishes); background=true detaches and " + "returns a task_id. The user cannot see sub-agent output — synthesise findings " + "into your own response.\n\n" + "Omit profile_id unless the step needs another profile's specialisation.\n\n" + 'Details, examples and the profile list: tool_manual("spawn_agent").' ) parameters = { "type": "object", @@ -60,20 +50,17 @@ "profile_id": { "type": "string", "description": ( - "Profile to use for the sub-agent. Defaults to the current session's " - "profile (e.g. navi_code for local coding) — for plain code work, omit it. " - "Override to specialise: 'server_admin' for remote ops, 'secretary' for " - "research/writing, 'tool_developer' for Navi tool implementation. The " - "selected profile determines the sub-agent's model, prompt, and available tools." + "Defaults to the current session's profile — the right choice for most " + "work. Set it only to specialise: 'server_admin' (remote ops), " + "'secretary' (research/writing), 'tool_developer' (Navi tool " + "implementation)." ), }, "system_prompt": { "type": "string", "description": ( - "Optional role definition for this sub-agent, injected as a system-level " - "instruction on top of the profile default. Use to specialise the agent: " - "e.g. 'You are a security auditor. Report findings by severity.' " - "or 'You are a metrics collector. Return all values in a structured table.'" + "Optional role definition, injected on top of the profile default. " + "e.g. 'You are a security auditor. Report findings by severity.'" ), }, "max_iterations": { @@ -83,21 +70,16 @@ "background": { "type": "boolean", "description": ( - "Run detached: returns a task_id immediately and the sub-agent keeps " - "working while you continue. Collect via tasks check/wait; the result " - "also arrives as an automatic note next turn. Use for sub-agents expected " - "to run longer than ~45-60s. Default false (synchronous)." + "Detach: returns a task_id immediately; collect via tasks, or wait for the " + "completion note. Use above ~45-60s. Default false (synchronous)." ), }, "inherit_system_prompt": { "type": "boolean", "description": ( - "If true, the sub-agent starts with the parent profile's full system prompt " - "as a base layer, then overlays the sub-agent's own specialisation on top. " - "Use this when you want the sub-agent to keep the parent's personality, rules, " - "and workflow, plus any extra instructions from 'system_prompt' or 'briefing'. " - "If false (default), the sub-agent uses only its own subagent_system_prompt, " - "ignoring the parent's system prompt entirely." + "Start from the parent profile's full system prompt as a base, then overlay " + "this sub-agent's own specialisation. Default false — the sub-agent uses only " + "its own subagent prompt." ), }, }, diff --git a/navi/tools/todo.py b/navi/tools/todo.py index b71ee40..4e02d30 100644 --- a/navi/tools/todo.py +++ b/navi/tools/todo.py @@ -112,17 +112,18 @@ class TodoTool(Tool): name = "todo" description = ( - "Task plan tracker. Your todo list is automatically populated from the plan at the start of each task — " - "you do NOT need to call 'set'. " - "Indexes are 1-based. Call 'update' with status='in_progress' when you start a step. " - "Call 'update' immediately after completing or failing each step — before moving to the next. " - "When marking a step 'done', you MUST provide a 'validation' field describing how you verified the result. " - "When marking a step 'failed', provide 'validation' explaining what went wrong and what you tried. " - "Before final response, make sure every completed step, including the final step, is marked done with validation. " - "Call 'view' to re-orient yourself after sub-agent execution or long tool chains. " - "Use 'set' only when you need to replace the plan mid-task (rare). " - "Use 'add' to append new steps discovered mid-task — it preserves existing steps and their statuses (unlike 'set'). " - "Statuses: pending → in_progress → done / failed / skipped." + "Task plan tracker, already populated from the plan when a task starts — you do NOT " + "need to call 'set'. Indexes are 1-based.\n\n" + "Call 'update' with status='in_progress' when you start a step, and again " + "immediately after completing or failing it, before starting the next. Marking a " + "step 'done' REQUIRES a 'validation' field saying how you verified the result; " + "'failed' should carry one explaining what went wrong and what you tried. Before " + "your final response, every completed step — the last one included — must be 'done' " + "with validation.\n\n" + "'add' appends steps discovered mid-task and keeps existing steps and their " + "statuses; 'set' replaces the whole plan and discards them. 'view' re-orients you " + "after a sub-agent run or a long tool chain.\n\n" + 'Full rules, statuses and examples: tool_manual("todo").' ) parameters = { "type": "object", diff --git a/navi/tools/tool_manual.py b/navi/tools/tool_manual.py index efae9d3..59fac7b 100644 --- a/navi/tools/tool_manual.py +++ b/navi/tools/tool_manual.py @@ -30,9 +30,11 @@ name = "tool_manual" description = ( "Returns the detailed manual for a tool: full usage instructions, parameter " - "reference, and examples. Call this before using an unfamiliar tool, or when you " - "are unsure about the correct format or parameters. Name the tool bare " - "('compile_scad') or by its full MCP name ('mcp__navi-3d__compile_scad')." + "reference, and examples. Call this before using an unfamiliar tool, or when " + "you are unsure about the correct format or parameters. Name the tool bare " + "('compile_scad') or by its full MCP name ('mcp__navi-3d__compile_scad'); name " + "an MCP server ('gntodo') to get that server's instructions and tool list. " + "With no name it lists which docs exist." ) parameters = { "type": "object", @@ -40,12 +42,12 @@ "tool_name": { "type": "string", "description": ( - "Name of the tool to look up, e.g. 'filesystem', 'todo', " - "'mcp__navi-3d__compile_scad'." + "Tool to look up — 'filesystem', 'todo', 'mcp__navi-3d__compile_scad' " + "— or the name of an MCP server, e.g. 'gntodo'. Omit to list what " + "manuals exist." ), } }, - "required": ["tool_name"], } def __init__(self, registry=None, profile_registry=None, mcp_manager=None) -> None: @@ -89,6 +91,12 @@ ) return ToolResult(success=True, output=f"[{note}]\n\n{_auto_manual(elsewhere)}") + # Not a tool — but possibly a whole MCP server, which is where a server's + # instructions live now that the system prompt only keeps one line each. + server_manual = self._server_manual(tool_name) + if server_manual is not None: + return ToolResult(success=True, output=server_manual) + return self._not_found(tool_name, scoped) # ── helpers ────────────────────────────────────────────────────────── @@ -103,6 +111,24 @@ return {} return {tool.name: tool for tool in self._registry.all()} + def _server_manual(self, name: str) -> str | None: + """The manual for an MCP *server* — the server itself, not one of its tools. + + The system prompt names each reachable server in one line; the instructions + that used to follow that line are read here instead. A server is documented + whether or not it is connected: its config is on disk, and an offline server + is exactly when its tools' guidance is worth re-reading. + """ + if self._mcp_manager is None: + return None + server = _match_server(self._mcp_manager.configured_servers(), name) + if server is None: + return None + instructions = self._mcp_manager.get_instructions({server}).get(server, "") + return _render_server_manual( + server, instructions, self._mcp_manager.server_tool_names(server) + ) + def _profile_tool_map(self, profile_id: str) -> dict[str, Tool]: """The tools this profile can call, as {name: Tool} for resolve_tool(). @@ -163,12 +189,24 @@ guides = sorted(p.stem for p in _guide_paths()) if guides: lines.append(f"guides (not tools; readable by name): {', '.join(guides)}") + servers = self._documented_servers() + if servers: + lines.append( + f"MCP servers (call tool_manual with the server name for its full " + f"instructions and tool list): {', '.join(servers)}" + ) lines.append( "Any other tool still works: tool_manual returns a manual generated from " "its schema." ) return ToolResult(success=True, output="\n".join(lines)) + def _documented_servers(self) -> list[str]: + """Configured MCP servers that have something to say about themselves.""" + if self._mcp_manager is None: + return [] + return sorted(self._mcp_manager.get_summaries().keys()) + def _not_found(self, tool_name: str, scoped: dict[str, Tool]) -> ToolResult: """Nothing matched — suggest instead of leaving the agent to guess.""" known = { @@ -252,6 +290,48 @@ return sorted(GUIDES_DIR.glob("*.md")) +def _match_server(servers: list[str], name: str) -> str | None: + """Resolve a requested name against the configured server names. + + Servers are spelled both ways in practice — the config file is `navi-3d.json` + while the tool prefix the agent sees is `mcp__navi-3d__…`, and a model retyping + it may use either dash or underscore. Same tolerance ``resolve_tool`` gives + tool names, for the same reason. + """ + lowered = name.strip().lower() + normalized = lowered.replace("-", "_") + for server in servers: + candidate = server.lower() + if candidate == lowered or candidate.replace("-", "_") == normalized: + return server + return None + + +def _render_server_manual(server: str, instructions: str, tools: list[str]) -> str: + """The full text behind the one line the system prompt spends on a server.""" + lines = [ + f"# {server} — MCP server manual", + "", + f"An MCP server; its tools are enabled per profile and named " + f"`mcp__{server}__`.", + ] + if tools: + lines.append(f"Tools it exposes ({len(tools)}): " + ", ".join(f"`{t}`" for t in tools)) + lines.append("") + if instructions.strip(): + lines.append(f"> Written for this machine in `mcp_servers.d/{server}.json`, merged") + lines.append("> with whatever the server announces when it connects.") + lines.append("") + lines.append(instructions.strip()) + else: + lines.append( + "No instructions are configured for this server. Its tools' own descriptions " + "are all the guidance there is — `tool_manual(\"\")` returns any one of " + "them." + ) + return "\n".join(lines) + + def manual_names() -> set[str]: """Tool names with a hand-written manual (guides excluded — they are not tools).""" if not MANUALS_DIR.is_dir(): diff --git a/tests/unit/core/test_context_builder.py b/tests/unit/core/test_context_builder.py index 462a5f6..98c2654 100644 --- a/tests/unit/core/test_context_builder.py +++ b/tests/unit/core/test_context_builder.py @@ -10,6 +10,7 @@ class FakeMcpManager: def __init__(self): self.calls = [] + self.summary_calls = [] def get_instructions(self, server_names=None): self.calls.append(server_names) @@ -17,6 +18,12 @@ return {} return {name: f"{name} instructions" for name in server_names} + def get_summaries(self, server_names=None): + self.summary_calls.append(server_names) + if not server_names: + return {} + return {name: f"{name} summary" for name in server_names} + class TestBuildSystemPrompt: def test_includes_persona(self, monkeypatch): @@ -138,9 +145,10 @@ result = builder.build(context, profile, mem=None) assert not any("MCP servers" in (m.content or "") for m in result) - assert mcp.calls == [] + assert mcp.summary_calls == [] - def test_injects_only_profile_mcp_server_instructions(self): + def test_injects_only_profile_mcp_server_summaries(self): + """One line per server — the instructions behind it are fetched on demand.""" mcp = FakeMcpManager() builder = ContextBuilder(profile_registry=make_profile_registry(), mcp_manager=mcp) profile = make_profile("test", mcp_servers={"gnexus-book": ["read"]}) @@ -148,8 +156,23 @@ result = builder.build(context, profile, mem=None) - assert any("gnexus-book instructions" in (m.content or "") for m in result) - assert mcp.calls == [{"gnexus-book"}] + assert any("gnexus-book summary" in (m.content or "") for m in result) + assert mcp.summary_calls == [{"gnexus-book"}] + + def test_the_mcp_block_points_at_tool_manual_and_inlines_no_instructions(self): + mcp = FakeMcpManager() + builder = ContextBuilder(profile_registry=make_profile_registry(), mcp_manager=mcp) + profile = make_profile("test", mcp_servers={"gnexus-book": ["read"]}) + context = [Message(role="user", content="hi")] + + result = builder.build(context, profile, mem=None) + + block = next(m.content for m in result if "MCP servers" in (m.content or "")) + assert 'tool_manual("")' in block + assert "gnexus-book instructions" not in block + # Two prose lines, one heading line per server — nothing that scales with the + # size of the instructions themselves. + assert len(block.splitlines()) <= 4 class FakeMemoryStore: diff --git a/tests/unit/test_mcp.py b/tests/unit/test_mcp.py index 2f799b6..dab3fc3 100644 --- a/tests/unit/test_mcp.py +++ b/tests/unit/test_mcp.py @@ -6,7 +6,7 @@ from navi.mcp.client import McpClient from navi.mcp.config import McpServerConfig, load_mcp_servers -from navi.mcp.manager import McpManager +from navi.mcp.manager import McpManager, summarize_instructions from navi.mcp.tools import McpTool @@ -95,6 +95,85 @@ assert instructions == {"gnexus-book": "Use book."} +class TestServerSummaries: + """The one line the system prompt keeps per server. + + Instructions used to be inlined there in full — ~8 KB of prose on every request + of a heavy profile. Only the opening claim stays; the body is one + tool_manual("") call away. + """ + + def test_takes_the_first_sentence(self): + assert summarize_instructions("Use book. Then read it.") == "Use book." + + def test_keeps_a_single_sentence_whole(self): + assert summarize_instructions("MCP tools for gntodo.") == "MCP tools for gntodo." + + def test_flattens_newlines_so_it_stays_one_line(self): + assert summarize_instructions("Use book.\n\nThen\nread it.") == "Use book." + + def test_caps_a_runaway_first_sentence(self): + summary = summarize_instructions("word " * 100) + + assert len(summary) <= 201 + assert summary.endswith("…") + + def test_empty_text_has_no_summary(self): + assert summarize_instructions("") == "" + assert summarize_instructions(" \n ") == "" + + def test_summaries_fall_back_to_the_instructions(self, tmp_path): + path = tmp_path / "mcp_servers.json" + path.write_text('{"book": {"instructions": "Use book. Details follow."}}') + manager = McpManager(config_path=path) + + assert manager.get_summaries({"book"}) == {"book": "Use book."} + + def test_an_explicit_summary_wins(self, tmp_path): + path = tmp_path / "mcp_servers.json" + path.write_text( + '{"book": {"summary": "The book.", "instructions": "MANDATORY: use book now."}}' + ) + manager = McpManager(config_path=path) + + assert manager.get_summaries({"book"}) == {"book": "The book."} + + def test_a_server_with_nothing_to_say_is_not_listed(self, tmp_path): + path = tmp_path / "mcp_servers.json" + path.write_text('{"book": {"transport": "stdio", "command": "python"}}') + manager = McpManager(config_path=path) + + assert manager.get_summaries({"book"}) == {} + + def test_summaries_are_ordered(self, tmp_path): + path = tmp_path / "mcp_servers.json" + path.write_text( + '{"zeta": {"instructions": "Z."}, "alpha": {"instructions": "A."}}' + ) + manager = McpManager(config_path=path) + + assert list(manager.get_summaries({"zeta", "alpha"})) == ["alpha", "zeta"] + + def test_configured_servers_lists_the_config_not_the_pool(self, tmp_path): + """A disconnected server is exactly when its instructions are worth reading.""" + path = tmp_path / "mcp_servers.json" + path.write_text('{"book": {"instructions": "Use book."}, "todo": {}}') + manager = McpManager(config_path=path) + + assert manager.clients == {} + assert manager.configured_servers() == ["book", "todo"] + + def test_server_tool_names_dedupes_across_groups(self, tmp_path): + path = tmp_path / "mcp_servers.json" + path.write_text( + '{"book": {"groups": {"read": ["search", "get"], "write": ["get", "set"]}}}' + ) + manager = McpManager(config_path=path) + + assert manager.server_tool_names("book") == ["get", "search", "set"] + assert manager.server_tool_names("nope") == [] + + class TestMcpTool: def test_name_prefix(self): mock_manager = AsyncMock(spec=McpManager) diff --git a/tests/unit/tools/test_manual_drift.py b/tests/unit/tools/test_manual_drift.py index a20ddeb..339a1dc 100644 --- a/tests/unit/tools/test_manual_drift.py +++ b/tests/unit/tools/test_manual_drift.py @@ -272,12 +272,19 @@ def test_every_tool_manual_call_resolves(): - """`tool_manual("x")` must reach something: a hand-written manual, a guide, or — - at worst — a real tool whose schema the tool can render.""" + """`tool_manual("x")` must reach something: a hand-written manual, a guide, a real + tool whose schema the tool can render, or a configured MCP server — naming a server + is how its instructions are read now that the system prompt only keeps one line. + + A `` is documentation shorthand, not a name anyone will type. + """ + resolvable = KNOWN_TOOLS | set(MCP_BY_SERVER) unresolved = { name: where for name, where in _citations(_TOOL_MANUAL_CALL).items() - if tool_manual._read_manual(name) is None and name not in KNOWN_TOOLS + if tool_manual._read_manual(name) is None + and name not in resolvable + and not name.startswith("<") } assert not unresolved, ( f"tool_manual() calls that resolve to nothing: {unresolved} — name a real tool, or " diff --git a/tests/unit/tools/test_spawn_agent.py b/tests/unit/tools/test_spawn_agent.py index 9537be8..b6f38d4 100644 --- a/tests/unit/tools/test_spawn_agent.py +++ b/tests/unit/tools/test_spawn_agent.py @@ -83,4 +83,4 @@ assert background["type"] == "boolean" assert "task_id" in background["description"] # the default remains synchronous - assert "SYNCHRONOUS" in tool.description + assert "synchronous" in tool.description.lower() diff --git a/tests/unit/tools/test_tool_manual.py b/tests/unit/tools/test_tool_manual.py index dbede92..b40370e 100644 --- a/tests/unit/tools/test_tool_manual.py +++ b/tests/unit/tools/test_tool_manual.py @@ -39,12 +39,44 @@ class FakeMcpManager: """Stand-in for McpManager — group name to tool names, as a server config holds them.""" - def __init__(self, groups: dict[str, dict[str, list[str]]] | None = None) -> None: + def __init__( + self, + groups: dict[str, dict[str, list[str]]] | None = None, + instructions: dict[str, str] | None = None, + summaries: dict[str, str] | None = None, + ) -> None: self._groups = groups or {} + self._instructions = instructions or {} + self._summaries = summaries or {} def resolve_group(self, server_name: str, group_name: str) -> list[str]: return list(self._groups.get(server_name, {}).get(group_name, [])) + def configured_servers(self) -> list[str]: + return sorted(set(self._groups) | set(self._instructions) | set(self._summaries)) + + def get_instructions(self, server_names=None) -> dict[str, str]: + names = self._instructions if server_names is None else server_names + return {n: self._instructions[n] for n in names if n in self._instructions} + + def server_tool_names(self, server_name: str) -> list[str]: + names: list[str] = [] + for tools in self._groups.get(server_name, {}).values(): + names += [t for t in tools if t not in names] + return sorted(names) + + def get_summaries(self, server_names=None) -> dict[str, str]: + """One line per server — the explicit summary, else the first line of the docs.""" + names = set(self._instructions) | set(self._summaries) + if server_names is not None: + names &= set(server_names) + out: dict[str, str] = {} + for name in sorted(names): + line = self._summaries.get(name) or self._instructions.get(name, "").split("\n")[0] + if line: + out[name] = line + return out + def build(registry=None, profiles=None, mcp_manager=None) -> ToolManualTool: return ToolManualTool( @@ -140,6 +172,78 @@ assert "schema description" not in result.output +class TestServerManual: + """An MCP server is not a tool, but it is documented by the same call. + + The system prompt keeps one line per server; the instructions behind that line + are read here. Naming a server must therefore work even when the server is + disconnected — `configured_servers` reads the config on disk, not the live pool. + """ + + @staticmethod + def manager(**kwargs) -> FakeMcpManager: + return FakeMcpManager( + groups={"gntodo": {"read": ["list_projects", "get_project"]}}, + instructions={"gntodo": "MCP tools for gntodo.\n\nWorkflow: list_projects first."}, + **kwargs, + ) + + async def test_a_server_name_returns_its_instructions(self): + result = await build(mcp_manager=self.manager()).execute({"tool_name": "gntodo"}) + + assert result.success is True + assert "Workflow: list_projects first." in result.output + + async def test_it_lists_the_tools_the_server_exposes(self): + result = await build(mcp_manager=self.manager()).execute({"tool_name": "gntodo"}) + + assert "`list_projects`" in result.output + assert "`get_project`" in result.output + + async def test_a_dash_spelling_reaches_an_underscored_server(self): + """The config file is navi_ui.json; agents spell it both ways.""" + mcp = FakeMcpManager(instructions={"navi_ui": "Renders a UI component."}) + + result = await build(mcp_manager=mcp).execute({"tool_name": "navi-ui"}) + + assert result.success is True + assert "Renders a UI component." in result.output + + async def test_a_server_without_instructions_says_so(self): + mcp = FakeMcpManager(groups={"bare": {"all": ["do_thing"]}}) + + result = await build(mcp_manager=mcp).execute({"tool_name": "bare"}) + + assert result.success is True + assert "No instructions are configured" in result.output + + async def test_a_tool_name_is_not_mistaken_for_a_server(self, manuals): + """Server lookup runs after tool resolution — a tool always wins.""" + mcp = FakeMcpManager(instructions={"filesystem": "server instructions"}) + registry = registry_with(FakeTool("filesystem", description="the tool's schema")) + + result = await build(registry=registry, mcp_manager=mcp).execute( + {"tool_name": "filesystem"} + ) + + assert "the tool's schema" in result.output + assert "server instructions" not in result.output + + async def test_an_unknown_server_name_still_fails_usefully(self): + result = await build(mcp_manager=self.manager()).execute({"tool_name": "gntodoX"}) + + assert result.success is False + + async def test_the_index_names_the_servers(self, manuals): + (manuals / "todo.md").write_text("# todo\n\nhand-written\n") + + result = await build(mcp_manager=self.manager()).execute({}) + + assert result.success is True + assert "gntodo" in result.output + assert "full" in result.output + + class TestGeneratedManual: async def test_an_unwritten_manual_falls_back_to_the_schema(self, manuals): registry = registry_with(FakeTool("todo", description="Manage the todo list"))