2026-07-21 13:35:51 -05:00
|
|
|
# chorus-memory eval — current-version metrics harness
|
2026-06-19 21:12:14 -05:00
|
|
|
|
2026-07-03 13:30:01 -05:00
|
|
|
A reproducible, credential-free eval of the shipped plugin against the deterministic
|
|
|
|
|
mock REST API. `run_eval.py` measures the four things the plugin claims (results print
|
|
|
|
|
as a table and land in `results/latest.json`):
|
|
|
|
|
|
|
|
|
|
1. **Retrieval** — hybrid recall (BM25 over entities **and** sessions/journal, fused
|
|
|
|
|
with freshness/status priors) vs a keyword-only baseline modeling pre-1.5 recall
|
|
|
|
|
(per-word substring ranking over entity notes only). Precision@5 / recall@5 / MRR,
|
|
|
|
|
session/journal answerability, freshness top-1.
|
|
|
|
|
2. **Dedup** — duplicate notes created replaying a capture stream with the pre-write
|
2026-07-21 13:35:51 -05:00
|
|
|
gate ON (default) vs OFF (`CHORUS_DUP_GATE=99` ≈ pre-1.5 warn-only), plus legitimate
|
2026-07-03 13:30:01 -05:00
|
|
|
captures wrongly blocked (must be 0).
|
|
|
|
|
3. **Write safety** — silent write failures under fault injection (transient 503,
|
|
|
|
|
phantom PUT, PATCH to a missing heading, replayed append). Must be 0; must be loud.
|
|
|
|
|
4. **Durability** — writes during an outage queue, replay exactly once, never
|
|
|
|
|
duplicate on re-flush.
|
|
|
|
|
|
|
|
|
|
> The original **0.6-vs-0.7 A/B harness** (token cost + silent-failure comparison that
|
|
|
|
|
> motivated the 0.7 client) lives in git history — `git log -- eval/run_eval.py`.
|
2026-06-19 21:12:14 -05:00
|
|
|
|
2026-07-03 12:57:52 -05:00
|
|
|
## Current test suites (run these for any change)
|
|
|
|
|
|
|
|
|
|
The A/B harness below is the historical 0.6-vs-0.7 comparison. The **current-version**
|
|
|
|
|
suites live alongside it and are what CI gates on:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
python3 test_features.py # end-to-end features vs the mock (capture/recall/gate/triage/hooks/…)
|
|
|
|
|
python3 test_reflect.py # reflection pipeline
|
|
|
|
|
python3 test_offline_queue.py # H2 outage queue + read cache
|
|
|
|
|
python3 test_patch_semantics.py # real PATCH semantics vs mock_olrapi_hifi.py (incl. fm create-or-replace)
|
2026-07-21 13:35:51 -05:00
|
|
|
# plus, in the plugin tree: scripts/test_chorus_client.py (offline unit + routing-sync guard)
|
2026-07-03 12:57:52 -05:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
## Run it (historical A/B)
|
2026-06-19 21:12:14 -05:00
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
cd eval
|
|
|
|
|
python3 run_eval.py # default params
|
|
|
|
|
python3 run_eval.py --recovery 2500 # sensitivity-test the recovery assumption
|
|
|
|
|
python3 run_eval.py --cpt 3.5 # different chars/token proxy
|
|
|
|
|
```
|
|
|
|
|
|
2026-06-21 11:46:54 -05:00
|
|
|
No network, no API key, no live vault. Pure stdlib Python 3 (the shipped client is now Python — no bash required).
|
2026-06-19 21:12:14 -05:00
|
|
|
Results table prints to stdout and a machine-readable copy lands in `results/latest.json`.
|
|
|
|
|
|
|
|
|
|
## How it works
|
|
|
|
|
|
|
|
|
|
- **`mock_olrapi.py`** — a deterministic mock of the Obsidian Local REST API surface the
|
|
|
|
|
plugin uses, reproducing its real behaviors and quirks (404 shape, the `/vault//`
|
|
|
|
|
double-slash 400, directory listings with `dir/` entries, `PATCH` heading targets that
|
|
|
|
|
return `400 invalid-target / 40080` when the heading is absent). Faults are triggered by
|
|
|
|
|
path markers so one server serves every scenario:
|
|
|
|
|
- `flaky` in the path → first write returns `503`, then succeeds (tests retry).
|
|
|
|
|
- `phantom` in the path → `PUT` returns `200` but does **not** persist (tests read-back verify).
|
|
|
|
|
- a `PATCH` to a missing heading → `400` (the silent-write-loss trigger).
|
|
|
|
|
- **`run_eval.py`** — for each scenario, runs both methods against a freshly reset +
|
|
|
|
|
re-seeded server (so faults are identical for both), then reads ground truth back
|
2026-07-21 13:35:51 -05:00
|
|
|
**independently** from the mock. The 0.7 side executes the *actual shipped `chorus.py`*;
|
2026-06-19 21:12:14 -05:00
|
|
|
the 0.6 side faithfully models the documented recipe (real HTTP, but no status check,
|
|
|
|
|
no retry, no verify, no dedupe).
|
|
|
|
|
|
|
|
|
|
## Metrics
|
|
|
|
|
|
|
|
|
|
| metric | meaning |
|
|
|
|
|
|---|---|
|
|
|
|
|
| `gen_tokens` | output tokens the model must generate for the op (`len(emitted)/cpt` proxy) |
|
|
|
|
|
| `silent_failure` | method **reported success** but ground truth is wrong (lost write or dup) — and nobody noticed |
|
|
|
|
|
| `detected` | the method surfaced the failure (nonzero exit) instead of hiding it |
|
|
|
|
|
| `effective_tokens` | `gen + silent_failures*recovery + detected*detect_cost` |
|
|
|
|
|
| `silent-error-free ops` | the headline accuracy number |
|
|
|
|
|
| `writes actually persisted` | did the single op land (separate from "was it silent") |
|
|
|
|
|
|
|
|
|
|
## Scenarios
|
|
|
|
|
|
|
|
|
|
1. **agent-log-missing-heading** — `PATCH` append to a note lacking the target heading (`400`).
|
|
|
|
|
2. **scope-switch** — clean `PATCH replace` (no fault; pure token comparison).
|
|
|
|
|
3. **inbox-capture-replayed** — same capture issued twice (retry/replay): dedup vs duplicate.
|
|
|
|
|
4. **session-log-flaky-network** — one-time `503`: retry vs single-shot.
|
|
|
|
|
5. **heartbeat-phantom-write** — accepted-but-not-persisted: read-back verify vs none.
|
|
|
|
|
6. **cold-start-load-6-reads** — 6 GETs (no fault; pure token comparison).
|
|
|
|
|
|
|
|
|
|
## Representative result (defaults)
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
generated tokens 723 -> 174 (+76% fewer)
|
|
|
|
|
silent failures 4 -> 0 (-4)
|
|
|
|
|
duplicate lines 1 -> 0 (-1)
|
|
|
|
|
silent-error-free ops 1/5 -> 5/5
|
|
|
|
|
effective tokens (assumed) 6723 -> 334
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
## Honest caveats
|
|
|
|
|
|
|
|
|
|
- **Mechanics, not reasoning.** This measures the deterministic plumbing differences. It
|
|
|
|
|
does **not** measure model judgment (routing choices, prose quality) — that needs a live model.
|
|
|
|
|
- **`recovery` and `detect-cost` are assumptions**, not measurements. The headline
|
|
|
|
|
"silent failures: 4 → 0" is a hard count from ground truth; the `effective_tokens` figure
|
|
|
|
|
is a model on top of it — tune `--recovery` to see the sensitivity.
|
|
|
|
|
- **`gen_tokens` is a `chars/cpt` proxy** for the I/O layer only, not a tokenizer count, and
|
|
|
|
|
excludes the one-time `+12%` SKILL.md context cost noted in the analysis (that's a
|
|
|
|
|
per-session context cost, not per-op).
|
|
|
|
|
|
|
|
|
|
## Extending to a live-model run (optional)
|
|
|
|
|
|
|
|
|
|
To measure real model behavior and true token counts:
|
|
|
|
|
1. Define the same six scenarios as natural-language tasks (e.g. "log a session note for X").
|
|
|
|
|
2. Run each twice — once with the 0.6 skill files, once with 0.7 — through the Agent SDK
|
2026-07-21 13:35:51 -05:00
|
|
|
against the **mock** server (point `CHORUS_BASE` at it) so faults stay deterministic.
|
2026-06-19 21:12:14 -05:00
|
|
|
3. Record `usage.output_tokens` per task from the API and whether the vault ended correct
|
|
|
|
|
(same ground-truth read used here).
|
|
|
|
|
|
|
|
|
|
The mock + ground-truth checks in this harness are reusable as-is for that; only the driver
|
|
|
|
|
changes from "scripted ops" to "model-driven ops".
|