| 2026-10-08 |

startup, llm, profiles, todo: five more the prod logs gave up
...
The same 5-day sweep of prod's journal as 5e51f87, continued. Every one of
these failed silently: no error in the UI, just a line in the log or a hole
in the database.
complete()/embed() retry a connection error before falling back.
stream_complete has always given its first chunk two attempts, but the
non-streaming path made exactly one per (server, model), and with a single
configured server (ollama.com) one ReadTimeout raised "All backends
exhausted: ReadTimeout" outright. That is the whole story of 2026-10-08,
when 27 calls died that way in a day — 23 of them memory summarisation, the
rest planning. Attempts now come from settings.llm_complete_retries (2) with
llm_retry_backoff_sec (2.0) between them; a missing model is still not
retried (a 404 is an answer, not a hiccup), and a lone server is still never
blacklisted — blacklisting it would block the next request for _TTL.
Startup waits for Postgres. create_container() opens the pool eagerly and
the retry loop for the DDL tables sits below it, so a host that comes up
before docker takes the whole lifespan with it. On 2026-10-07 15:19:40 the
agent ran `sudo reboot` on itself at the user's request ("Перезагрузи
себя"); the host booted 15:20:03, navi started at 15:20:18, hit
ConnectionRefusedError on the pool and logged "Application startup failed.
Exiting." Connection errors now retry for ~30s; a genuinely broken config
(a missing DATABASE_URL) still fails fast.
The profile loader reports its tally. A profile dropped by an unreadable
config.json is skipped by design so it cannot take the server down, but on
that same 15:20 boot four of them — modeler_3d, navi_code, secretary,
server_admin — were dropped with `Extra data: line 125 column 1`, and the
only trace was one error line each: the UI simply showed four profiles
fewer.
todo recovers a dropped `op`. Of the 3385 todo calls stored on prod, 73
arrived with the discriminator flattened away ({"index": 1, "status":
"in_progress"}, {"action": "view"}, {"": "add", "tasks": [...]}) and got a
bare "Unknown op: None" — a wasted round trip for an intent the remaining
arguments state plainly. The error message now names the five ops and what
each needs.
Message.created_at is stamped at creation. It defaulted to None and only
some construction sites filled it, so 85% of stored rows had no timestamp at
all (every tool message, 14 168 of 14 168) and no query could slice history
by time or measure a pause. Loading a row passes the column through
explicitly, so NULL rows stay NULL instead of being restamped with "now".
Every new test was checked to fail with its fix reverted. Full suite green
(1466 passed, 1 skipped). No frontend or dependency change, so the deploy is
a pull and a restart.
Eugene Sukhodolskiy
committed
9 hours ago
|

memory+llm: three failures the prod logs gave up
...
All three came out of a survey of the last days of journalctl on prod,
and each one silently destroyed something the user had already paid for.
memory_facts: the no-embedding INSERT bound $13 while its column list
had 12 entries, so every fact written while the embedding backend was
down died on PostgresSyntaxError and the extraction was lost — 47
embed failures in three days, most of them landing in this branch.
Embedding input now gets clipped instead of 400'd away. A 400 takes the
whole embedding with it and recall drops to ILIKE over everything; a
reaction-session prompt (a full event envelope inlined) did that 41
times. settings.embedding_max_chars (6000, 0 disables) caps the input
at the model's window, so recall still works on the head of the text.
Message strips NUL bytes at the model boundary. PostgreSQL text cannot
hold one, and a NUL arriving in a tool result (reading a binary file)
made the whole session_messages INSERT fail with
CharacterNotInRepertoireError — the turn died and the user lost it.
Three times on 2026-10-07. As a Message validator it covers every
writer downstream: session store, kv store, memory extraction.
Each new test was checked to fail with its fix reverted.
Eugene Sukhodolskiy
committed
9 hours ago
|
| 2026-07-10 |

tui: show the currently-served model in the status panel
...
The status panel's Model line was fed the global ollama_default_model, not the
session/profile model, and the server never told the client which model
actually served a call. Now:
- Backends stamp the resolved model onto LLMChunk (first chunk) / LLMResponse.
The fallback backend reports the model that survived its server+model
priority list (may differ from the profile's first choice).
- New ModelInfo event ({"type":"model_info","model":...}) emitted once per
turn from agent._consume_stream, re-emitted only when the model changes
across iterations. Additive WS event — old clients ignore it.
- TUI: attach_session/switch fetch the profile's configured model (first of
profile.model) via api.get_profile_model so the panel shows a value before
the first request; model_info then refines it to the actually-served model.
Not forwarded to the chat panel. raw CLI prints "[model] ...".
Co-Authored-By: Claude <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 10 Jul
|
| 2026-05-21 |
FallbackOllamaBackend: do not blacklist single server, empty file fallback
...
- When only one Ollama server is configured, LLMConnectionError no longer
adds it to the dead-server blacklist. This fixes the bug where a
transient failure permanently blocked all requests until server restart.
- LLMModelNotFoundError on a single server is also not blacklisted.
- _discover_backends now falls back to settings.ollama_host when the
ollama_backends_file is empty, missing, or returns no valid servers.
- Added unit tests covering single-server no-blacklist, multi-server
blacklist, model-not-found no-blacklist, and empty-file fallback.
400 passed, 1 skipped
Eugene Sukhodolskiy
committed
on 21 May
|
| 2026-05-18 |
Make Settings immutable (frozen=True) and fix all test mutations
...
- Add frozen=True to SettingsConfigDict in navi/config.py
- Convert model_validator to mode="before" since mode="after" cannot mutate frozen instances
- Replace all field-level monkeypatches in tests with whole-Settings object replacement
- Ensure cross-module settings consistency (content_store, session_files, share_file, content_publish, filesystem)
392 passed, 1 skipped
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 18 May
|
| 2026-05-11 |
Fix ollama_backends / FallbackOllamaBackend issues
...
- registry.py: always use FallbackOllamaBackend (unified backend).
Enables model priority lists in all deployments, not just multi-server.
- agent.py: add missing think=profile.think_enabled to run() (REST endpoint).
- compressor.py: fix model param type (str → list[str] | str | None).
- fallback.py: harden load_servers_from_file against missing/bad JSON files
and entries without host. Add clear_blacklists() for manual reset.
- admin.py: add POST /admin/ollama/clear-blacklists endpoint.
- tech_debt_review: document dead stream() methods.
- tests: add tests for single-server fallback, bad file handling,
missing host skipping, and blacklist clearing.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Eugene Sukhodolskiy
committed
on 11 May
|
| 2026-04-30 |
Improve content publishing UX
Eugene Sukhodolskiy
committed
on 30 Apr
|
| 2026-04-29 |
Add regression tests for content publishing and LLM timeouts
Eugene Sukhodolskiy
committed
on 29 Apr
|
Align Ollama HTTP timeout with LLM timeouts
Eugene Sukhodolskiy
committed
on 29 Apr
|