Files
echo/docs/IMPROVEMENT-PLANS.md
T
Jason Stedwell ed840eea73
Build and Push Docker Image / build (push) Successful in 9s
Plans: 2.3.0 index train shipped
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 00:10:32 -05:00

19 KiB
Raw Blame History

echo-memory — Improvement Plans (post-1.5.1 review)

Status: planned, not started. Source: full-code review 2026-07-28. Nine improvements, each scoped for an independent build session. The tenth item from that review — the MCP server — has its own detailed spec: docs/MCP-SERVER-SPEC.md.

Sequencing is at the bottom. Nothing here conflicts with TODO-1.6.md or ROADMAP-2.0.md; release targets assume 1.6 (skills migration) ships first.


1. Token-budgeted load (load --brief) + parallel orientation reads

Problem. The SessionStart hook injects the full text of six files into every session. Live measurement: 15.9KB of hook context on a single cold start, and two of the six files grow without bound (operator-preferences ## Observations, current-context ## Scope History). Every session pays this tax before any work happens. Separately, cmd_load (echo.py) issues its six GETs serially even though read_many() exists.

Design.

  • New echo.py load --brief (and make the SessionStart hook use it):
    • marker → one line (schema_version, bootstrapped yes/no).
    • operator-preferences → ## Fact / Pattern in full; ## Observations capped at the last 10 lines with a (+N older — /echo-load for full) note.
    • current-context → ## Scope + scope_updated + sessions-since (i.e. the scope show output); omit ## Scope History.
    • heartbeat → the pointer line + only the ## Summary/## Outcomes sections of the pointed-at session log (fetch it, extract those headings, cap ~30 lines).
    • today's daily note → ## Agent Log lines only (or "absent").
    • inbox → count + age of oldest item, not the contents (that's all the reconcile needs; /echo-triage --list has the details).
  • Budget guard: after assembly, if the brief output still exceeds ECHO_LOAD_BUDGET chars (default ~8000), truncate lowest-priority sections first (observations → agent log → session summary) with explicit (truncated) markers.
  • Full load unchanged; /echo-load keeps using it.
  • Fetch all reads via read_many() (the 6 targets + the heartbeat-pointed session log + the sessions listing fallback) instead of the serial loop at echo.py:640. Cache-put still applies per path.

Files. echo.py (cmd_load), echo_hook_session_start.py, SKILL.md (document brief-vs-full), eval/test_features.py.

Tests. Mock vault with oversized preferences/history files → assert brief output under budget, sections present, truncation markers correct; hook emits brief.

Size/target. Small-medium. 1.6.x point release — highest leverage per line.


2. Local-first recall index

Problem. echo_recall.update_note() does GET-whole-index → add → PUT-whole-index against the vault, under the global advisory lock, on every capture and every corpus put. The index carries full BM25 postings for the whole corpus, so this round-trip is O(vault) network per write and is the emerging bottleneck + lock hot-spot as the vault grows past a few hundred notes.

Design.

  • The live recall index moves to the local state dir (~/.echo-memory/recall-index.json, honoring ECHO_STATE_DIR) keyed by endpoint (hash of BASE) so multiple vaults don't collide.
  • update_note() becomes a local read-modify-write (file lock via os.O_EXCL sidecar or atomic os.replace), no vault round-trip, no advisory lock.
  • The vault copy (_agent/index/recall-index.json) becomes a snapshot, written by sweep.py and at session end (piggyback on the heartbeat write / future session-end verb). It exists so a fresh machine can seed its local index without a full rebuild.
  • Freshness rule on recall: local index used when present; if the vault snapshot's embedded built stamp is newer than the local one (another client swept), pull it. Staleness across clients is tolerable — recall degrades gracefully and improvement #4 (incremental sweep) trues it up cheaply.
  • Schema bump: recall-index schema 3 adds built (ISO timestamp) + endpoint_hash. Old schema-2 vault copies are still readable as seeds.
  • The entity index stays in the vault (it is small, and it is the shared routing authority for concurrent clients) — only the BM25 index moves.

Files. echo_recall.py (load/save/update_note/rebuild), echo_queue.py (state-dir helpers shared), sweep.py, echo.py (cmd_put upkeep path), routing.json (the vault snapshot path is unchanged), README performance section.

Tests. Capture with vault mock → assert zero recall-index PUTs; sweep → snapshot written; fresh state dir + existing snapshot → seeded without rebuild; two-vault (endpoint) separation.

Size/target. Medium. Pair with #4 in one index-focused minor (they touch the same maintenance loop). Fine to land before or after 2.0.


3. Offline durability for the high-level ops (capture path)

Problem. The low-level verbs queue on outage via safe_request, but _append_to_existing() and ensure_daily_log() (echo_ops.py) call echo.request directly — so during an outage a raw append survives, while a capture update and its Agent Log line silently vanish. Inverted from the documented contract ("capture is the default write").

Design. Queue semantic records, not byte-level requests, for the high-level path:

  • New outbox record type {"op": "capture", "args": {...}} alongside the existing raw records. On an unreachable vault, capture short-circuits: enqueue the full argument set (title/kind/body/tags/aliases/sources/date/domain/merge_into) and report queued (offline): capture ….
  • flush() replays op:capture records by calling echo_ops.capture(...) against the current index — so routing, the duplicate gate, aliasing, and linking all re-run with fresh state instead of replaying stale byte writes. A gate stop on replay (exit 76) keeps the record queued with a needs_attention flag and is surfaced by load/flush output ("1 queued capture needs --merge-into/--force").
  • Rationale for semantic-over-byte: a byte replay of _append_to_existing's PATCH could target a heading that moved, and the create-vs-update decision made offline may be wrong by replay time. Re-running the op is the idempotent, correct unit (the dated-bullet idempotency key already prevents double-append).
  • ensure_daily_log stays best-effort but routes its POST/PATCH through safe_request so a standalone agent-log line survives an outage too.
  • reflect/triage --apply inherit this for free (they call capture).

Files. echo_ops.py, echo_queue.py (record schema v2 + replay dispatch), eval/test_offline_queue.py.

Tests. Fault-injected mock: capture-update offline → queued; flush replays via capture; gate-on-replay keeps record + flags; agent-log line queued; no duplicate bullets on double flush.

Size/target. Small-medium. 1.6.x point release — it's a correctness gap.


4. Incremental sweep via content hashes

Problem. Human edits in Obsidian and writes by other clients never touch the entity/recall indexes; drift is only corrected by a full sweep someone must remember to run. Full sweep re-reads the whole vault (fast now, but O(vault)).

Design.

  • Entity index entries and the recall-index snapshot gain a per-note h — SHA-1 of the note body (entity index schema 2; tolerated-absent for old entries).
  • sweep.py --fast: walk the listing (cheap), read_many only paths that are (a) new, (b) missing a stored hash, or (c) whose hash mismatches after fetch — to avoid fetching everything just to hash it, fast mode fetches only paths whose listing is new/gone plus a rotating shard (e.g. hash(path) % 7 == weekday) of existing notes, so the whole vault is verified over a week of fast sweeps while each run stays tiny. --fast --all-shards forces full verification.
  • Deletions: listing walk catches removed files → drop index + recall entries.
  • Auto-run: load (brief or full) runs sweep --fast opportunistically when the last fast-sweep stamp (state dir) is older than ECHO_FAST_SWEEP_DAYS (default 7), in the background-tolerant sense: bounded by the same read_many concurrency, and skipped entirely when offline.
  • /echo-health reports last fast/full sweep ages.

Files. sweep.py, echo_index.py (schema 2 + hash field), echo_recall.py (snapshot hash reuse), echo.py (load hook-in), vault_lint.py (report), migrate note (index schema is machine-owned; no vault migration needed).

Tests. Mock: edit a note out-of-band → fast sweep in that shard re-indexes it; delete → entry dropped; stamp gating; full sweep unchanged.

Size/target. Medium. Ship with #2 (same schema-bump train).


5. Memory lifecycle — enforce forgetting

Problem. The vault only accretes. Working memory is "time-boxed" by convention only; Observations trimming is manual prose; stale-active detection only reports. Recall precision degrades as dead weight accumulates.

Design.

  • capture --kind working stamps expires: <today + ECHO_WORKING_TTL_DAYS> (default 14; --ttl <days> overrides). Template updated to carry the field.
  • vault_lint.py new checks: expired-working (past expires), stale-observations (Observations > 30 bullets) — advisory, like the rest.
  • sweep.py --decay (dry-run by default, --apply to write; always preview-first, honoring the operating contract):
    • expired working notes → status: archived + move body under a dated ## Archived marker (never delete);
    • projects/active/ untouched > 30 days → propose on-hold (folder move + status, the usual paired change) — apply only the ones confirmed;
    • Observations > 30 → trim oldest into a dated _agent/memory/episodic/ overflow note (observations-archive-YYYY.md append), keeping the last 30 in place. Nothing is lost, the hot file stays small (compounds with #1).
  • /echo-health output nudges: "run sweep --decay to review N aging items".

Files. echo_ops.py (--ttl/expires stamp), echo_index.py (KIND_REQUIRED_FM unchanged; working gains optional expires), vault_lint.py, sweep.py, scaffold working-memory template, SKILL.md (decay section), routing-map note.

Tests. Mock: expired note proposed+archived on apply; observations trim keeps last 30 and appends overflow; stale-active proposal list; dry-run writes nothing.

Size/target. Medium. Feature minor (1.7.x).


6. Contradiction / supersession handling

Problem. A new fact that contradicts a stored one lands beside it; both surface in recall with no signal about which is current. Biggest remaining trust gap in the write path.

Design.

  • Write side: capture --supersedes "<slug-or-quoted-line>":
    • if it names an entity slug → the old entity gets status: superseded + a superseded_by: frontmatter path; new note carries supersedes: back-link in ## Related. Recall's STATUS_FACTOR gains superseded: 0.5.
    • if it quotes a Fact/Pattern or Observations line in operator-preferences → the line is struck (~~old~~ (superseded YYYY-MM-DD)) and the new line appended, one PATCH each, in the same invocation.
  • Detect side (deterministic assist, model decides): echo_reflect.classify gains a conflict scan — for each proposal, run its distinctive tokens (_sig_tokens) against the Fact/Pattern + recent Observations lines; overlapping lines are attached as _conflicts: [...] and the preview row renders update-or-conflict? with the matched line(s). The model (which has the conversation) decides supersede-vs-append; the script only guarantees the question gets asked before the write.
  • Lint: superseded_by pointing at a missing note → violation.

Files. echo_ops.py, echo_reflect.py, echo_recall.py (STATUS_FACTOR), echo.py (flag plumbing), vault_lint.py, SKILL.md + operating-contract (supersession is additive: strike, never delete).

Tests. Mock: supersede entity → status/backlinks both sides; supersede line → strike + append idempotent; reflect preview flags a planted conflict; recall ranks superseded below active twin.

Size/target. Medium. Feature minor (1.7.x/1.8) — after #5 (shares the status-vocabulary touchpoints).


7. session-end bundle verb

Problem. Ending a session correctly is 4+ invocations (session-log PUT, Agent Log append, heartbeat PUT, optional scope set, optional reflect apply). Cost per session, and a partial failure leaves broken orientation state (log written, heartbeat stale).

Design.

  • echo.py session-end <bundle.json> [--apply], bundle:
    {
      "slug": "echo-mcp-spec",
      "log_body": "<full session-log markdown, frontmatter included>",
      "agent_log_line": "- 2026-07-28: wrote MCP spec [[...]]",
      "scope": "optional new scope text",
      "reflect": [ { ...PROPOSAL_SCHEMA... } ]
    }
    
  • Dry-run previews the whole plan (paths + reflect preview). --apply, in order, under one lock acquisition: PUT _agent/sessions/<date>-<HHMM>-<slug>.md (HHMM from ECHO_NOW env or wall clock, validated against the canonical regex) → agent-log line (ensure_daily_log) → reflect proposals via capture → scope set if given → heartbeat PUT last (it's the commit marker: a crash before it leaves the previous pointer intact, never a dangling one).
  • Reports per-step ok/queued/gated; offline → every step queues (rides on #3).
  • Stop hook reason text updated to name the one command; /echo-reflect docs point session logging at it. Prerequisite for the MCP echo_log_session tool — build this verb first so CLI and MCP share one implementation.

Files. new echo_session.py, echo.py (subcommand), echo_hook_stop.py (REASON text), SKILL.md (Session Logging section), commands/echo-reflect.md.

Tests. Mock: full bundle lands all five artifacts; heartbeat-last ordering (inject failure mid-way → heartbeat untouched); dry-run writes nothing; offline queues all steps.

Size/target. Small-medium. 1.6.x point release (and before the MCP build).


8. Recall robustness — stemming + alias query expansion

Problem. BM25 is purely lexical: "deploy"/"deployment"/"deployed" are distinct terms; a paraphrased query misses. The 8-query gold set can't see this class of miss.

Design.

  • Stemming: a light suffix-stripper (Porter-lite: s/es/ed/ing/ly/tion/ment/ness families, ~40 lines, pure stdlib) applied in tokenize() at both index and query time. Index schema bump (stemmed postings) → auto-rebuild on next recall/sweep, exactly like the 1→2 transition.
  • Alias query expansion: at recall time, run the query against fuzzy_candidates() (top 2, score ≥ 0.5); fold the matched entities' title + alias tokens into the BM25 query as down-weighted expansion terms (×0.5 idf contribution) so "the ECHO plugin" also scores echo-memory-alias vocabulary. Exact-match seeding stays as-is; expansion never creates hits ranked above genuine lexical matches without corpus support.
  • Eval: grow the gold set with ≥6 paraphrase queries (morphological variants + alias phrasings); publish before/after recall@5/MRR in the README table.

Files. echo_recall.py (tokenize, INDEX_SCHEMA, score, recall), new echo_stem.py (pure function, unit-testable), eval/run_eval.py + gold set.

Tests. Stemmer unit table (word → stem, incl. non-cases); paraphrase gold queries pass; schema-2 index discarded and rebuilt cleanly.

Size/target. Small-medium. Ride the same index-schema train as #2/#4.


9. Resolve-miss telemetry → alias suggestions

Problem. Every resolve that misses but ends in a confirmed candidate (--merge-into, or the operator picking a candidate) is a free mention→entity training pair, currently discarded.

Design.

  • echo_index.resolve() miss + fuzzy_candidates non-empty → append one NDJSON record to the local state dir (~/.echo-memory/resolve-misses.ndjson): {mention, candidates:[slugs], ts?} (no wall clock needed — session date via ECHO_TODAY). capture --merge-into <slug> appends the confirmation record {mention, resolved: slug} (this already learns the alias immediately — that path stays; the log covers the cases that never reach --merge-into).
  • sweep.py (and /echo-health) read the log: any {mention → slug} pair confirmed ≥2 times, or any mention that repeatedly fuzzy-matched a single candidate ≥3 times without confirmation, is proposed as an alias (sweep --apply writes it into the note's aliases: frontmatter + index, the existing fold-back path). Log entries consumed on apply; file capped (rotate at ~500 lines).
  • Privacy note: mentions can contain operator phrasing — the log lives in the state dir (never the vault), same trust level as the offline queue.

Files. echo_index.py (log hooks), echo_ops.py (confirmation hook), sweep.py (suggest/apply), vault_lint.py or sweep report (surface count), SKILL.md one-liner.

Tests. Miss → record written; confirm ×2 → sweep proposes; apply folds alias + consumes log; cap/rotation.

Size/target. Small. Feature minor, pairs naturally with #8.


Sequencing (unified with TODO-1.6 and ROADMAP-2.0)

TODO-1.6 folds in as the opening train — the skills migration restructures the files everything below touches, so it goes first. ROADMAP-2.0 stays a pure packaging major (its own stated thesis) and is not folded into any feature train; it slots between the quick wins and the MCP server so the MCP work lands on the final packaging baseline (single command format, Gitea releases, one manifest-verification pass).

Release Items Rationale
1.6.0 TODO-1.6: skills-format migration (+ per-skill allowed-tools, disable-model-invocation on write-heavy entries, $ECHO block dedupe) · lint blind-spot fix (retired-dir/leaf-README) · plugin.json homepage Restructures the command/SKILL files later work edits; allowed-tools removes prompt friction for every verb added below. Vault link repointing (TODO-1.6 §3) is vault data, not plugin code — any session.
1.6.0 — SHIPPED 2026-07-28 Skills migration + lint fix + homepage Verified on desktop + CoWork.
2.0.0 — SHIPPED 2026-07-28 ROADMAP-2.0: commands/ deleted · release-based artifacts · gitignore · Codex note Packaging-only major; skills-only install re-verified.
2.1.0 quick wins — SHIPPED 2026-07-28 #1 brief load · #3 offline capture · #7 session-end All three landed with +19 e2e checks; see CHANGELOG 2.1.0. The MCP spec's Phase 0 (return-not-print) did NOT ride along — it's the first phase of the MCP build session.
2.2 MCP server (containerized, echomcp.alwisp.com) — docs/MCP-SERVER-SPEC.md Needs Phase 0 (return-not-print refactor, first step of the build session). #7 session-end shipped, so echo_log_session wraps a real verb. Deployed via CI + PORT, not the plugin manifest.
2.3.0 index train — SHIPPED 2026-07-29 #2 local-first recall · #4 incremental sweep (sweep --fast, auto at load) · #8 stemming+alias expansion Recall schema 3 + entity schema 2 in one bump; eval 14-query gold set at recall@5/MRR 1.00/1.00. echo-mcp auto-redeployed with it.
2.4 memory train #5 lifecycle/decay · #6 supersession · #9 resolve-miss learning Capability tier; #6 after #5 (status vocabulary). Remaining alongside: the 2.2.x dividends (nightly vault backup FIRST, claude.ai OAuth shim, shadow history, background sweeps, alerting).