feat(eval): block-yod effort calibration harness with read-only direction supply (#212)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s

בלוק י' (דיון והכרעה) הוא הבלוק החשוב ביותר אך ה-effort שלו (xhigh) נקבע
'by feel' ולא כויל אמפירית. ה-harness הכללי (#208) לא יכל לכייל אותו כי
write_block(block-yod) זורק "ללא כיוון מאושר" — בלוק-הדיון דורש
decisions.direction_doc.approved שתיקי-הכיול חסרים.

scripts/calibrate_block_yod.py מוסיף שכבת אספקת-כיוון read-only מעל מכונת
המדידה של #208:
- overlay על db.get_decision_by_case/get_decision לתהליך-הזה-בלבד מזרים
  direction_doc שמיש ל-write_block — אפס כתיבה ל-DB, מצב-תיק-הייצור לא
  משתנה (reverted ב-__exit__).
- אסטרטגיה per-תיק: stored (כיוון-מאושר שמור, נאמן) → reconstructed
  (מינימלי מ-decision.outcome + סקשן-הדיון של הסופי באותו תיק, INV-LRN5:
  אפס מהות חוצת-תיקים, מתויג reconstructed) → skip.
- משתמש-חוזר ב-style_distance.block_distance_to_final + _score_cell/
  recommend_effort מ-calibrate_effort (G2 — אין מסלול-מדידה מקביל).
- כתיבה אינקרמנטלית per-תיק + שער-per-תא לא-פטאלי (INV-G8).

ממצא-קורפוס: מ-8 סופיים 7 ניתנים-לכיול; רק 8126-03-25 stored+scorable
(8174-12-24 שמור-אך-חסר-סקשן-דיון) → הנסיגה reconstructed היא מה שמפעיל
את ה-harness היום.

Invariants: G8 (eval-harness), G2 (שימוש-חוזר ב-style_distance/block_writer,
אין fork), INV-LRN5 (כיוון משוחזר = סגנון/שיטה בלבד, אפס מהות חוצת-תיקים),
claude_session local-only (ייצור דרך claude -p, מקומי-בלבד).

אימות: --self-test ALL PASS (28 checks: מדידה + אספקת-כיוון + overlay
install/revert) · --dry-run מול DB חי (read-only): 7/8 scorable, 28
ייצורים מתוכננים · --allow-reconstructed false → 1 (נאמן בלבד) ·
--case מסנן. אפס mutation calls ב-harness (read-only מאומת).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-30 18:50:01 +00:00
parent ff07e0bc98
commit 858e05fb8b
3 changed files with 650 additions and 0 deletions

3
.gitignore vendored
View File

@@ -19,6 +19,9 @@ kiryat-yearim/
continuation-prompt.md continuation-prompt.md
node_modules/ node_modules/
data/eval/eval-report-* data/eval/eval-report-*
# effort-calibration run outputs (#208 / #212) — generated, not source
data/eval/effort-calibration-*
data/eval/block-yod-calibration-*
# revert snapshot for migrate_agent_adapter.py (runtime state) # revert snapshot for migrate_agent_adapter.py (runtime state)
data/adapter-migration-state.json data/adapter-migration-state.json
# frontmatter-stripped instruction copies for content_arg adapters (generated) # frontmatter-stripped instruction copies for content_arg adapters (generated)

View File

@@ -77,6 +77,7 @@
| `ab_halacha_opus48.py` | python | **A/B לא-הרסני לחילוץ הלכות (Claude)** — מריץ מחדש חילוץ הלכות על פסק-דין בודד דרך מודל/effort נבחרים (`AB_MODEL`/`AB_EFFORT`, ברירת-מחדל `claude-opus-4-8`/`xhigh`) ומשווה לסטטיסטיקות ההלכות הקיימות ב-DB **בלי למחוק/לכתוב כלום**. משכפל את `halacha_extractor.extract()` (אותם פרומפטים, בחירת-צ'אנקים, אימות-ציטוט) ומחליף רק את קריאת ה-LLM ב-`claude -p --model --effort`. מפיק `data/ab_halacha_<case>_<effort>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data .venv/bin/python scripts/ab_halacha_opus48.py <case_law_id>`. **ממצא 2026-05-31 (שטיין 1128-08-20):** Opus 4.8@xhigh חילץ 51 מול 124 בייצור (100% quote-verified מול 96%) אך ביטחון מכויל-נמוך יותר (חציון 0.75 מול 0.82) — ולכן **לא** מקטין את תור-האישור-הידני תחת sweep אוטו-אישור conf≥0.78 (26 מול 24). שיפור איכות, לא צמצום-תור. | ידני (החלטת מודל-חילוץ) | | `ab_halacha_opus48.py` | python | **A/B לא-הרסני לחילוץ הלכות (Claude)** — מריץ מחדש חילוץ הלכות על פסק-דין בודד דרך מודל/effort נבחרים (`AB_MODEL`/`AB_EFFORT`, ברירת-מחדל `claude-opus-4-8`/`xhigh`) ומשווה לסטטיסטיקות ההלכות הקיימות ב-DB **בלי למחוק/לכתוב כלום**. משכפל את `halacha_extractor.extract()` (אותם פרומפטים, בחירת-צ'אנקים, אימות-ציטוט) ומחליף רק את קריאת ה-LLM ב-`claude -p --model --effort`. מפיק `data/ab_halacha_<case>_<effort>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data .venv/bin/python scripts/ab_halacha_opus48.py <case_law_id>`. **ממצא 2026-05-31 (שטיין 1128-08-20):** Opus 4.8@xhigh חילץ 51 מול 124 בייצור (100% quote-verified מול 96%) אך ביטחון מכויל-נמוך יותר (חציון 0.75 מול 0.82) — ולכן **לא** מקטין את תור-האישור-הידני תחת sweep אוטו-אישור conf≥0.78 (26 מול 24). שיפור איכות, לא צמצום-תור. | ידני (החלטת מודל-חילוץ) |
| `ab_halacha_codex.py` | python | **A/B לא-הרסני לחילוץ הלכות (Codex/gpt-5.5)** — עמית ל-`ab_halacha_opus48` אך מחליף את `claude -p` ב-`codex exec --model gpt-5.5` (אימות ChatGPT, ללא OPENAI_API_KEY). אותם פרומפטים ואותו הסקת quote-verification. הפלט האחרון של הסוכן (`-o FILE`) נפענח כ-JSON. `AB_MODEL` (default `gpt-5.5`), `AB_REASONING` low/medium/high/xhigh (default `medium`), `AB_CONCURRENCY` (default 1), `CODEX_BIN`. מפיק `data/ab_halacha_codex_<case>_<model>_<reasoning>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data mcp-server/.venv/bin/python scripts/ab_halacha_codex.py <case_law_id>`. **ממצא 2026-06-17 (8181-21 האוניברסיטה העברית):** gpt-5.5@medium חילץ 27 מול 28 של Opus (quote-verified 100%/100%), ביטחון חציון 0.86 מול 0.78 — אך **0 פריטים מתחת ל-0.7** (לעומת 9/28 של Opus = 32%), דבר המצביע על over-confidence. holding↑ (12 מול 7), procedural↓ (4 מול 7). **מסקנה: ריאלי כ-fallback חירום; לא מוכן לייצור ללא כיול-ביטחון.** | ידני (בנצ'מרק מודל codex) | | `ab_halacha_codex.py` | python | **A/B לא-הרסני לחילוץ הלכות (Codex/gpt-5.5)** — עמית ל-`ab_halacha_opus48` אך מחליף את `claude -p` ב-`codex exec --model gpt-5.5` (אימות ChatGPT, ללא OPENAI_API_KEY). אותם פרומפטים ואותו הסקת quote-verification. הפלט האחרון של הסוכן (`-o FILE`) נפענח כ-JSON. `AB_MODEL` (default `gpt-5.5`), `AB_REASONING` low/medium/high/xhigh (default `medium`), `AB_CONCURRENCY` (default 1), `CODEX_BIN`. מפיק `data/ab_halacha_codex_<case>_<model>_<reasoning>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data mcp-server/.venv/bin/python scripts/ab_halacha_codex.py <case_law_id>`. **ממצא 2026-06-17 (8181-21 האוניברסיטה העברית):** gpt-5.5@medium חילץ 27 מול 28 של Opus (quote-verified 100%/100%), ביטחון חציון 0.86 מול 0.78 — אך **0 פריטים מתחת ל-0.7** (לעומת 9/28 של Opus = 32%), דבר המצביע על over-confidence. holding↑ (12 מול 7), procedural↓ (4 מול 7). **מסקנה: ריאלי כ-fallback חירום; לא מוכן לייצור ללא כיול-ביטחון.** | ידני (בנצ'מרק מודל codex) |
| `calibrate_effort.py` | python | **#208 (WS5/Q1, INV-G8 eval-harness) — כיול model×effort של הכותב מול הסופיים.** A/B per-(תיק,בלוק,effort) על `draft_final_pairs` בעלי `final_text`: מייצר מחדש כל בלוק דרך מסלול-הייצור (`block_writer.write_block(effort_override=…)``claude_session.query``claude -p`, Opus 4.8 נעוץ, **מקומי-בלבד**) ומודד מול ה**סקשן** המתאים בסופי דרך `services/style_distance.block_distance_to_final` (מקור-מדידה יחיד, G2): `change_percent` (compute_diff_stats) · `anti_pattern_total` (`lessons.ANTI_PATTERNS`, הסיגנל הנקי-לסגנון) · `golden_ratio_deviation_pp` · `distance` מרוכב. **ממליץ** per-בלוק על ה-effort בעל ה-distance-הממוצע-הנמוך (לצד ברירת-המחדל מ-#204). מפיק `data/eval/effort-calibration-<ts>.{json,md}` ב**כתיבה אינקרמנטלית** — הדוח נכתב מחדש אחרי כל בלוק שמסתיים, כך שקריסה באמצע הגריד לא מאבדת בלוקים שכבר יוצרו (INV-G8). כל תא-(תיק,בלוק,effort) שנכשל (rate-limit חולף, או חריגה) נרשם כ-warning ומדולג — לא מפיל את הריצה. **`block-yod` (דיון) הוצא מברירות-המחדל** (`CALIBRATABLE_BLOCKS`) — דורש כיוון-מאושר (brainstorm→approve_direction) שתיקי-הכיול חסרים ולכן `write_block` זורק; נשאר ב-`VALID_BLOCKS` לכפייה ידנית (`--blocks block-yod`) על-אחריות-המשתמש, כשהשער-per-תא שומר אותו לא-פטאלי. ⚠️ **גודל-מדגם מודפס בראש הדוח — עדות-כיוון, לא רגרסיה** (מעט סופיים-עלויים). `--self-test` (offline, אפס DB/CLI — מוכיח את לוגיקת-המדידה) · `--dry-run` (תכנון-גריד) · `--efforts`/`--blocks`/`--case`/`--repeats`. **חובה מקומי** (claude CLI; לא בקונטיינר/worktree-ללא-CLI). הרצה: `POSTGRES_PASSWORD=… mcp-server/.venv/bin/python scripts/calibrate_effort.py`. | ידני — לכיול ברירות-effort של הכותב | | `calibrate_effort.py` | python | **#208 (WS5/Q1, INV-G8 eval-harness) — כיול model×effort של הכותב מול הסופיים.** A/B per-(תיק,בלוק,effort) על `draft_final_pairs` בעלי `final_text`: מייצר מחדש כל בלוק דרך מסלול-הייצור (`block_writer.write_block(effort_override=…)``claude_session.query``claude -p`, Opus 4.8 נעוץ, **מקומי-בלבד**) ומודד מול ה**סקשן** המתאים בסופי דרך `services/style_distance.block_distance_to_final` (מקור-מדידה יחיד, G2): `change_percent` (compute_diff_stats) · `anti_pattern_total` (`lessons.ANTI_PATTERNS`, הסיגנל הנקי-לסגנון) · `golden_ratio_deviation_pp` · `distance` מרוכב. **ממליץ** per-בלוק על ה-effort בעל ה-distance-הממוצע-הנמוך (לצד ברירת-המחדל מ-#204). מפיק `data/eval/effort-calibration-<ts>.{json,md}` ב**כתיבה אינקרמנטלית** — הדוח נכתב מחדש אחרי כל בלוק שמסתיים, כך שקריסה באמצע הגריד לא מאבדת בלוקים שכבר יוצרו (INV-G8). כל תא-(תיק,בלוק,effort) שנכשל (rate-limit חולף, או חריגה) נרשם כ-warning ומדולג — לא מפיל את הריצה. **`block-yod` (דיון) הוצא מברירות-המחדל** (`CALIBRATABLE_BLOCKS`) — דורש כיוון-מאושר (brainstorm→approve_direction) שתיקי-הכיול חסרים ולכן `write_block` זורק; נשאר ב-`VALID_BLOCKS` לכפייה ידנית (`--blocks block-yod`) על-אחריות-המשתמש, כשהשער-per-תא שומר אותו לא-פטאלי. ⚠️ **גודל-מדגם מודפס בראש הדוח — עדות-כיוון, לא רגרסיה** (מעט סופיים-עלויים). `--self-test` (offline, אפס DB/CLI — מוכיח את לוגיקת-המדידה) · `--dry-run` (תכנון-גריד) · `--efforts`/`--blocks`/`--case`/`--repeats`. **חובה מקומי** (claude CLI; לא בקונטיינר/worktree-ללא-CLI). הרצה: `POSTGRES_PASSWORD=… mcp-server/.venv/bin/python scripts/calibrate_effort.py`. | ידני — לכיול ברירות-effort של הכותב |
| `calibrate_block_yod.py` | python | **#212 (WS5/Q1, INV-G8 eval-harness) — כיול effort ל-בלוק י' (דיון והכרעה) בלבד.** ה-harness ש-#208 לא יכל לכייל: `write_block(block-yod)` זורק "ללא כיוון מאושר" כי בלוק-הדיון דורש `decisions.direction_doc.approved` שתיקי-הכיול חסרים. **התוספת על #208 — שכבת אספקת-כיוון read-only:** לפני ייצור בלוק-הדיון, ה-harness מזרים `direction_doc` שמיש דרך **overlay על קריאת-ה-DB** (monkeypatch של `db.get_decision_by_case`/`db.get_decision` לתהליך-הזה-בלבד) — **אפס כתיבה ל-DB**, מצב-תיק-הייצור לא משתנה (השער reverted ב-finally). **אסטרטגיית-אספקה per-תיק:** (1) **stored** (מועדף, נאמן) — כיוון-מאושר שמור (`approved`+`selected_direction`, או סכמת-8126 עם `direction_id`/`structure`); (2) **reconstructed** (נסיגה, עדות-כיוון-בלבד) — `direction_doc` מינימלי שנגזר מ-`decision.outcome` + סקשן-הדיון של הסופי **באותו תיק** (INV-LRN5: outcome+מצביע-לסופי-עצמי, אפס מהות חוצת-תיקים, מתויג `reconstructed`); (3) **skip** אם אין דיון/כיוון. מודד דרך אותו `style_distance.block_distance_to_final` (מקור-מדידה יחיד, G2) ומשתמש-חוזר ב-`_score_cell`/`recommend_effort` מ-`calibrate_effort` (#208). **ממצא-קורפוס (2026-06-30):** מ-8 סופיים, 7 ניתנים-לכיול; רק 8126-03-25 הוא stored+scorable (8174-12-24 שמור-אך-חסר-סקשן-דיון) → `--allow-reconstructed false` נותן תא-יחיד; ברירת-המחדל (reconstructed פעיל) נותנת 7 → הנסיגה היא מה שמפעיל את ה-harness היום. כתיבה אינקרמנטלית per-תיק + שער-per-תא לא-פטאלי (INV-G8). ⚠️ **גודל-מדגם בראש הדוח — עדות-כיוון, לא רגרסיה.** `--self-test` (offline, אפס DB/CLI — מוכיח מדידה + לוגיקת-אספקת-כיוון + overlay) · `--dry-run` (תכנון-גריד + מצב-אספקה per-תיק, צריך DB) · `--efforts`/`--case`/`--repeats`/`--allow-reconstructed`. ברירת-מחדל נוכחית: `xhigh` (`block_writer.py:72`, 'by feel'). **חובה מקומי** (claude CLI; לא בקונטיינר/worktree-ללא-CLI). הרצת-כיול-חי (לבצע ע"י היו"ר): `POSTGRES_PASSWORD=… POSTGRES_HOST=127.0.0.1 POSTGRES_PORT=5433 mcp-server/.venv/bin/python scripts/calibrate_block_yod.py`. | ידני — לכיול ה-effort של בלוק-הדיון (היו"ר מריץ; long, צורך-מכסה) |
| `monitor_halacha_quality.py` | python | מנטר איכות חילוץ הלכות. בודק drift של `avg(confidence)` בין baseline היסטורי לחלון אחרון. מחזיר JSON מטריקות + alert ב-stderr אם drift > threshold (ברירת מחדל 5%). 2 סדרות: trusted (approved+published) ו-all_extracted. תומך `--window N` / `--threshold X` / `--min-sample N` / `--silent` / `--exit-on-alert`. רץ ב-container או מקומית עם `mcp-server/.venv` (אין תלות ב-LLM, רק SQL). **תזמון מומלץ**: `0 8 * * 1` (יום ראשון 08:00, שבועי) | `0 8 * * 1` (לתזמן) | | `monitor_halacha_quality.py` | python | מנטר איכות חילוץ הלכות. בודק drift של `avg(confidence)` בין baseline היסטורי לחלון אחרון. מחזיר JSON מטריקות + alert ב-stderr אם drift > threshold (ברירת מחדל 5%). 2 סדרות: trusted (approved+published) ו-all_extracted. תומך `--window N` / `--threshold X` / `--min-sample N` / `--silent` / `--exit-on-alert`. רץ ב-container או מקומית עם `mcp-server/.venv` (אין תלות ב-LLM, רק SQL). **תזמון מומלץ**: `0 8 * * 1` (יום ראשון 08:00, שבועי) | `0 8 * * 1` (לתזמן) |
| `audit_training_corpus.py` | python | audit של `style_corpus` — לכל החלטה: שדות מטא-דאטה מאוכלסים (`summary`/`outcome`/`key_principles`/`appeal_subtype`/`subject_categories`), קישור ל-`documents` (FK + chunks + embeddings). מפיק `data/audit/corpus-YYYY-MM-DD.json` + summary בקונסול. דרוש `POSTGRES_URL` או POSTGRES_*. אין תלויות חיצוניות מלבד asyncpg. **רץ מהמכונה המקומית** (לא קונטיינר) — חיבור ישיר ל-Postgres :5433 | ידני / קדם-עבודה לפני enrichment של מטא-דאטה | | `audit_training_corpus.py` | python | audit של `style_corpus` — לכל החלטה: שדות מטא-דאטה מאוכלסים (`summary`/`outcome`/`key_principles`/`appeal_subtype`/`subject_categories`), קישור ל-`documents` (FK + chunks + embeddings). מפיק `data/audit/corpus-YYYY-MM-DD.json` + summary בקונסול. דרוש `POSTGRES_URL` או POSTGRES_*. אין תלויות חיצוניות מלבד asyncpg. **רץ מהמכונה המקומית** (לא קונטיינר) — חיבור ישיר ל-Postgres :5433 | ידני / קדם-עבודה לפני enrichment של מטא-דאטה |
| `backfill_style_exemplars.py` | python | **T1 (style-acquisition)** — מאכלס `style_exemplars` מקורפוס דפנה (`style_corpus` + `internal_committee` chair=דפנה): מפצל לסעיפים (`chunker._split_into_sections`) → פסקאות (25-450 מילים) → embed (Voyage) → שמירה עם `section`/`outcome`/`practice_area`. מאפשר לכותב לאחזר פסקאות-בלוק אמיתיות של דפנה (T2/T3). מקור-סגנון בלבד (INV-LRN5). אידמפוטנטי (מנקה per-decision). `--dry-run` (default) / `--apply`. דורש POSTGRES_URL + Voyage. **רץ מקומית** (venv). | ידני (`python scripts/backfill_style_exemplars.py --apply`) | | `backfill_style_exemplars.py` | python | **T1 (style-acquisition)** — מאכלס `style_exemplars` מקורפוס דפנה (`style_corpus` + `internal_committee` chair=דפנה): מפצל לסעיפים (`chunker._split_into_sections`) → פסקאות (25-450 מילים) → embed (Voyage) → שמירה עם `section`/`outcome`/`practice_area`. מאפשר לכותב לאחזר פסקאות-בלוק אמיתיות של דפנה (T2/T3). מקור-סגנון בלבד (INV-LRN5). אידמפוטנטי (מנקה per-decision). `--dry-run` (default) / `--apply`. דורש POSTGRES_URL + Voyage. **רץ מקומית** (venv). | ידני (`python scripts/backfill_style_exemplars.py --apply`) |

View File

@@ -0,0 +1,646 @@
#!/usr/bin/env python3
"""#212 (WS5 / Q1, INV-G8 eval-harness) — DEDICATED effort calibration for בלוק י'
(block-yod, the discussion/decision block — דיון והכרעה).
WHY A SEPARATE HARNESS (the #208 harness could NOT calibrate block-yod):
block-yod is the most important block — its effort is currently `xhigh`, set
"by feel" (block_writer.py:72), never empirically calibrated. The general #208
harness (scripts/calibrate_effort.py) EXCLUDES it from CALIBRATABLE_BLOCKS
because block_writer.write_block(case_id, "block-yod", …) raises
ValueError("לא ניתן לכתוב בלוק דיון ללא כיוון מאושר…")
— the discussion block requires an APPROVED DIRECTION (brainstorm →
approve_direction → decisions.direction_doc{"approved": True, …}) that the
calibration cases (the draft_final_pairs finals) do not have. So block-yod is
not calibratable standalone without first SUPPLYING a direction.
WHAT THIS HARNESS ADDS over #208 — the DIRECTION-SUPPLY layer (read-only):
Before regenerating block-yod for a calibration case, it makes a usable
`direction_doc` available to write_block via a READ-ONLY DB OVERLAY (it
monkeypatches db.get_decision_by_case / db.get_decision FOR THIS PROCESS ONLY
to overlay the direction into the returned decision dict). NOTHING is written
to the DB — production case state is never mutated (no UPDATE), satisfying the
"do not pollute real case state" constraint. The overlay is reverted in a
finally block.
DIRECTION-SUPPLY STRATEGY (per-case, documented in the report):
1. STORED (preferred, most faithful): if the case's decision already carries
a usable approved direction (direction_doc.approved is truthy and it has a
selected_direction, OR the 8126-style schema with direction_id/structure),
use it verbatim. This is the production-faithful path.
2. RECONSTRUCTED (fallback, directional-only): otherwise synthesize a MINIMAL
direction_doc from the case's decision.outcome + the final's OWN discussion
section (the chair's signed reasoning is the ground truth for "what
direction was taken"). Clearly LABELLED reconstructed in the report → the
resulting numbers are directional, not faithful to a real brainstorm.
INV-LRN5: the reconstructed direction carries the case's OWN outcome +
a neutral pointer to its OWN final — no substance dragged across cases.
EMPIRICAL NOTE (measured on the current corpus, 2026-06-30): of the 8 finals,
only 8174-12-24 has a STORED approved direction — and its final has NO parsable
discussion section (0 words), so it is not scorable for block-yod anyway. The 7
finals that DO have a scorable discussion section all LACK a usable stored
direction. Hence the RECONSTRUCTED fallback is what makes this harness useful
today; --dry-run prints the per-case supply mode so the operator sees this.
WHAT IT MEASURES (identical to #208, REUSED — no parallel metric path, G2):
Regenerate block-yod at each effort via block_writer.write_block(effort_override)
(the PRODUCTION generation path → claude_session.query → `claude -p`, pinned
Opus 4.8, local-only), then score vs the final's DISCUSSION section via
services.style_distance.block_distance_to_final (change_percent /
anti_pattern_total / golden_ratio_deviation_pp / composite distance). The
recommend_effort, aggregate_cell, _score_cell and direction-free scaffolding are
imported from calibrate_effort (#208) — this file only supplies the direction
and restricts the grid to block-yod.
GENERATION PATH (do not violate — reference_claude_generation_path /
claude_session docstring): write_block → claude_session.query → `claude -p` uses
the local claude.ai session. It runs ONLY on the host where the `claude` CLI
exists (NOT the legal-ai container, NOT a symlinked worktree without CLI access).
Hence the live A/B is HOST-ONLY; --self-test proves the measurement +
direction-supply logic offline with zero model calls.
⚠️ SAMPLE-SIZE CAVEAT (honored, not hidden): very few cases have an uploaded
final WITH a parsable discussion section. The report prints n_finals PROMINENTLY
and labels the output DIRECTIONAL EVIDENCE, not a regression. With n<3 the
recommendation is advisory only; the chair/operator decides whether to adopt.
Usage (mcp-server venv; live needs POSTGRES + the `claude` CLI on the host):
PY=/home/chaim/legal-ai/mcp-server/.venv/bin/python
$PY scripts/calibrate_block_yod.py --self-test # offline proof, no DB/CLI
$PY scripts/calibrate_block_yod.py --dry-run # plan grid + supply modes (needs DB, no model)
POSTGRES_PASSWORD=… POSTGRES_HOST=127.0.0.1 POSTGRES_PORT=5433 \
$PY scripts/calibrate_block_yod.py # live A/B over finals (host-only)
… --efforts low,medium,high,xhigh # override the effort grid
… --case 1130-08-25 # a single case
… --repeats 2 # avg N gens/cell (noise)
… --allow-reconstructed false # STORED-direction cases only (faithful)
"""
from __future__ import annotations
import argparse
import asyncio
import json
import logging
import os
import sys
from datetime import datetime
from pathlib import Path
from statistics import mean
logger = logging.getLogger(__name__)
REPO_ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(REPO_ROOT / "mcp-server" / "src"))
sys.path.insert(0, str(REPO_ROOT / "scripts"))
if "POSTGRES_URL" not in os.environ:
os.environ["POSTGRES_URL"] = (
f"postgres://{os.environ.get('POSTGRES_USER','legal_ai')}:"
f"{os.environ.get('POSTGRES_PASSWORD','')}@"
f"{os.environ.get('POSTGRES_HOST','127.0.0.1')}:"
f"{os.environ.get('POSTGRES_PORT','5433')}/"
f"{os.environ.get('POSTGRES_DB','legal_ai')}"
)
# Reuse the #208 machinery wholesale (G2 — no parallel harness/metric path).
# These are pure/offline-safe imports; calibrate_effort does no work at import.
from calibrate_effort import ( # noqa: E402
IL_TZ,
OUT_DIR,
VALID_EFFORTS,
_current_default,
_finals_for_calibration,
_score_cell,
recommend_effort,
)
BLOCK_ID = "block-yod"
SECTION = "discussion" # block-yod's golden-ratio section (style_distance._BLOCK_TO_SECTION)
DEFAULT_EFFORTS = ["low", "medium", "high", "xhigh"]
# ── direction supply (the block-yod-specific piece) ──────────────────────────
def _stored_direction_is_usable(direction_doc: dict | None) -> bool:
"""True if the case's stored direction_doc can drive write_block(block-yod).
write_block gates on `direction_doc.get("approved")` being truthy. Two stored
schemas exist in the corpus:
• brainstorm schema (build_direction_doc): {"approved": True, "selected_direction": …}
• chair-approved schema (8126-style): {"direction_id", "direction_name",
"structure", "template", …} — has NO "approved" key, so write_block would
reject it as-is. We treat it as usable IFF it carries a concrete direction
(direction_id/direction_name), and _normalize_stored_direction() stamps
approved=True onto the OVERLAY copy (never the DB).
Empty / brainstorm-only ({"brainstorm": …, "approved": False}) → not usable.
"""
dd = direction_doc or {}
if dd.get("approved") and (dd.get("selected_direction") or dd.get("reasoning")):
return True
if dd.get("direction_id") or dd.get("direction_name"):
return True
return False
def _normalize_stored_direction(direction_doc: dict) -> dict:
"""Return an overlay copy of a usable stored direction with approved=True.
Pure dict transform on a COPY — the real decisions.direction_doc is untouched.
The 8126-style schema lacks `approved`; stamping it here lets write_block's gate
pass without us mutating the DB.
"""
dd = dict(direction_doc)
dd["approved"] = True
return dd
def reconstruct_minimal_direction(outcome: str, final_discussion: str) -> dict:
"""Build a MINIMAL, clearly-reconstructed direction_doc (fallback path).
Faithful enough to make write_block(block-yod) runnable, honest about being a
reconstruction: the chair's SIGNED final discussion is the ground truth for
"what direction was taken", and the case's OWN outcome fixes the structure.
INV-LRN5: case's own outcome + a neutral pointer to its own final only — no
cross-case substance. Pure (no DB/LLM) → unit-tested in --self-test.
"""
from legal_mcp.services.lessons import OUTCOME_LABELS_HE, canonical_outcome
oc = canonical_outcome(outcome or "rejection")
snippet = " ".join((final_discussion or "").split()[:120])
return {
"approved": True,
"reconstructed": True, # marker so the report can label it
"outcome": oc,
"outcome_hebrew": OUTCOME_LABELS_HE.get(oc, oc),
"reasoning": (
"כיוון משוחזר לצורך כיול-effort בלבד (לא סיעור-מוחות אמיתי): "
"התוצאה והכיוון נגזרים מההחלטה החתומה של היו\"ר בתיק זה. "
"תמצית פתיח-הדיון של הסופי: " + snippet
),
"selected_direction": {
"name": f"כיוון משוחזר — {OUTCOME_LABELS_HE.get(oc, oc)}",
"reasoning": [
"שחזר את שלד-הנימוק מהדיון החתום של היו\"ר בתיק זה.",
],
"precedents": [],
},
"additional_notes": (
"⚠️ כיוון משוחזר — תוצאת-הכיול היא עדות-כיוון, לא נאמנה למסמך-כיוון אמיתי."
),
}
def plan_direction_supply(
decision: dict | None,
outcome: str,
final_discussion: str,
allow_reconstructed: bool,
) -> tuple[dict | None, str]:
"""Decide how to supply block-yod's direction for one case.
Returns (overlay_direction_doc | None, mode) where mode ∈
{"stored", "reconstructed", "skip"}. Pure (no I/O) → unit-tested.
"""
stored = (decision or {}).get("direction_doc")
if _stored_direction_is_usable(stored):
return _normalize_stored_direction(stored), "stored"
if allow_reconstructed and (final_discussion or "").strip():
return reconstruct_minimal_direction(outcome, final_discussion), "reconstructed"
return None, "skip"
class _DirectionOverlay:
"""Read-only overlay: make db.get_decision_by_case / db.get_decision return the
case's decision WITH `direction_doc` replaced by the supplied overlay — for the
target case_id only, for the duration of one cell. Reverted in __exit__.
write_block re-fetches the decision internally (block_writer.py:383), so passing
a direction dict in isn't enough — the DB read must yield it. We patch the read,
not the write: ZERO DB mutation (no UPDATE ever issued).
"""
def __init__(self, case_id, direction_doc: dict):
self._case_id = str(case_id)
self._dir = direction_doc
self._orig_by_case = None
self._orig_by_id = None
def __enter__(self):
from legal_mcp.services import db
self._orig_by_case = db.get_decision_by_case
self._orig_by_id = db.get_decision
async def patched_by_case(case_id, _orig=self._orig_by_case):
dec = await _orig(case_id)
if dec and str(case_id) == self._case_id:
dec = dict(dec)
dec["direction_doc"] = self._dir
return dec
async def patched_by_id(decision_id, _orig=self._orig_by_id):
dec = await _orig(decision_id)
if dec and dec.get("case_id") == self._case_id:
dec = dict(dec)
dec["direction_doc"] = self._dir
return dec
db.get_decision_by_case = patched_by_case
db.get_decision = patched_by_id
return self
def __exit__(self, *exc):
from legal_mcp.services import db
if self._orig_by_case is not None:
db.get_decision_by_case = self._orig_by_case
if self._orig_by_id is not None:
db.get_decision = self._orig_by_id
return False
async def _score_block_yod_cell(case_id, effort, final_section, final_total_words,
outcome, repeats, direction_doc):
"""Score one (case, block-yod, effort) cell with the direction overlaid.
Wraps the #208 _score_cell (REUSED) inside the read-only direction overlay so
write_block's approved-direction gate passes without DB mutation.
"""
with _DirectionOverlay(case_id, direction_doc):
return await _score_cell(
case_id, BLOCK_ID, effort, final_section, final_total_words, outcome, repeats,
)
# ── self-test (no DB, no model) ──────────────────────────────────────────────
def _self_test() -> int:
ok = True
def chk(name, cond):
nonlocal ok
ok = ok and cond
print(f" {name:50} {'ok' if cond else 'FAIL'}")
# measurement reuse: block_distance_to_final on the discussion section.
from legal_mcp.services.style_distance import block_distance_to_final, split_final_by_section
final_disc = "אנו סבורים כי דין הערר להידחות. כידוע, הלכה פסוקה היא. " * 30
d_same = block_distance_to_final(BLOCK_ID, final_disc, final_disc, "rejection",
section_target_total_words=len(final_disc.split()))
chk("identical regen ⇒ change_percent==0", d_same["change_percent"] == 0.0)
chk("identical regen ⇒ section==discussion", d_same["section"] == SECTION)
clean = "אין בידנו לקבל את הטענה. יחד עם זאת, מצאנו טעם. " * 20
dirty = "## כותרת\n- נקודה\n- נקודה\n* עוד\n### תת\n" * 10
d_clean = block_distance_to_final(BLOCK_ID, clean, final_disc, "rejection")
d_dirty = block_distance_to_final(BLOCK_ID, dirty, final_disc, "rejection")
chk("dirty regen ⇒ more anti-patterns", d_dirty["anti_pattern_total"] > d_clean["anti_pattern_total"])
chk("dirty regen ⇒ larger distance", d_dirty["distance"] > d_clean["distance"])
# ── direction-supply logic ──
chk("empty direction ⇒ not usable", not _stored_direction_is_usable({}))
chk("brainstorm-only ⇒ not usable",
not _stored_direction_is_usable({"brainstorm": {"directions": [1]}, "approved": False}))
chk("approved+selected ⇒ usable",
_stored_direction_is_usable({"approved": True, "selected_direction": {"name": "x"}}))
chk("8126-style (direction_id) ⇒ usable",
_stored_direction_is_usable({"direction_id": "d1", "structure": "..."}))
src = {"direction_id": "d1", "structure": "..."}
norm = _normalize_stored_direction(src)
chk("normalize stamps approved=True", norm["approved"] is True)
chk("normalize does not mutate source", "approved" not in src)
rec = reconstruct_minimal_direction("rejection", final_disc)
chk("reconstruct ⇒ approved", rec["approved"] is True)
chk("reconstruct ⇒ marked reconstructed", rec["reconstructed"] is True)
chk("reconstruct ⇒ canonical outcome", rec["outcome"] == "rejection")
chk("reconstruct ⇒ has selected_direction", bool(rec.get("selected_direction")))
chk("reconstruct ⇒ no cross-case precedents (INV-LRN5)",
rec["selected_direction"]["precedents"] == [])
# _build_direction_context must accept BOTH supply shapes (write_block uses it).
from legal_mcp.services.block_writer import _build_direction_context
ctx_recon = _build_direction_context({"direction_doc": rec})
chk("recon direction renders (not 'לא אושר')", "כיוון לא אושר" not in ctx_recon)
ctx_stored = _build_direction_context({"direction_doc": norm})
chk("stored(8126) direction renders", "כיוון לא אושר" not in ctx_stored)
# plan_direction_supply: stored wins; else reconstructed; else skip.
dd_stored = {"direction_doc": {"approved": True, "selected_direction": {"name": "x"}}}
_, m1 = plan_direction_supply(dd_stored, "rejection", final_disc, allow_reconstructed=True)
chk("plan: stored-direction case ⇒ stored", m1 == "stored")
_, m2 = plan_direction_supply({"direction_doc": {}}, "rejection", final_disc, True)
chk("plan: no stored + discussion ⇒ reconstructed", m2 == "reconstructed")
_, m3 = plan_direction_supply({"direction_doc": {}}, "rejection", final_disc, False)
chk("plan: no stored + reconstructed disabled ⇒ skip", m3 == "skip")
_, m4 = plan_direction_supply({"direction_doc": {}}, "rejection", "", True)
chk("plan: no stored + no discussion ⇒ skip", m4 == "skip")
# overlay monkeypatch installs + reverts cleanly (no live DB needed).
import legal_mcp.services.db as real_db
saved_bc, saved_bi = real_db.get_decision_by_case, real_db.get_decision
async def fake_by_case(cid):
return {"id": "d", "case_id": "C1", "direction_doc": {"approved": False}}
async def fake_by_id(did):
return {"id": "d", "case_id": "C1", "direction_doc": {"approved": False}}
real_db.get_decision_by_case = fake_by_case
real_db.get_decision = fake_by_id
try:
ov_dir = {"approved": True, "reconstructed": True}
async def _probe():
with _DirectionOverlay("C1", ov_dir):
d1 = await real_db.get_decision_by_case("C1")
d2 = await real_db.get_decision_by_case("OTHER")
d3 = await real_db.get_decision("d")
return d1, d2, d3
# --self-test is dispatched at the __main__ entry, OUTSIDE asyncio.run,
# so a plain asyncio.run here is safe (no enclosing event loop).
d1, d2, d3 = asyncio.run(_probe())
chk("overlay: target case gets overlaid direction", d1["direction_doc"]["approved"] is True)
chk("overlay: non-target case untouched", d2["direction_doc"]["approved"] is False)
chk("overlay: get_decision path overlaid", d3["direction_doc"]["reconstructed"] is True)
chk("overlay: reverted after context", real_db.get_decision_by_case is fake_by_case)
finally:
real_db.get_decision_by_case = saved_bc
real_db.get_decision = saved_bi
# reused #208 primitives still behave.
r = recommend_effort([
{"effort": "high", "distance": 0.22, "anti_pattern_total": 2, "change_percent": 25, "n": 1},
{"effort": "xhigh", "distance": 0.22, "anti_pattern_total": 1, "change_percent": 22, "n": 1},
])
chk("recommend tie → fewer anti-patterns", r["effort"] == "xhigh")
chk("current default for block-yod == xhigh", _current_default(BLOCK_ID) == "xhigh")
secs = split_final_by_section(
"רקע עובדתי\nהמקרקעין. " * 8 + "\n\nדיון והכרעה\nאנו סבורים. " * 8)
chk("split returns golden sections only",
set(secs).issubset({"background", "claims", "discussion", "summary"}))
print("ALL PASS" if ok else "*** FAILURES ***")
return 0 if ok else 1
# ── live A/B (host-only — needs DB + `claude` CLI) ───────────────────────────
async def _gather_cases(case_filter, allow_reconstructed):
"""Build per-case metadata + direction-supply plan over the finals."""
from uuid import UUID
from legal_mcp.services import db
from legal_mcp.services.lessons import canonical_outcome
from legal_mcp.services.style_distance import split_final_by_section
finals = await _finals_for_calibration(case_filter)
cases = []
for f in finals:
case = await db.get_case_by_number(f["case_number"]) if f.get("case_number") else None
if not case:
continue
decision = await db.get_decision_by_case(UUID(case["id"]))
outcome = canonical_outcome((decision or {}).get("outcome", "rejection"))
sections = split_final_by_section(f.get("final_text", ""))
final_discussion = sections.get(SECTION, "")
final_total_words = len((f.get("final_text", "") or "").split())
overlay, mode = plan_direction_supply(
decision, outcome, final_discussion, allow_reconstructed)
cases.append({
"case_number": f["case_number"],
"case_id": case["id"],
"outcome": outcome,
"final_discussion": final_discussion,
"final_discussion_words": len(final_discussion.split()),
"final_total_words": final_total_words,
"direction": overlay,
"supply_mode": mode,
# scorable IFF it has a discussion section AND a supplied direction.
"scorable": bool(final_discussion.strip() and overlay is not None),
})
return cases
async def _run(args, ts):
from uuid import UUID
cases = await _gather_cases(args.case, args.allow_reconstructed)
scorable = [c for c in cases if c["scorable"]]
grid_summary = {
"block": BLOCK_ID,
"section": SECTION,
"n_finals": len(cases),
"n_scorable": len(scorable),
"current_default": _current_default(BLOCK_ID),
"efforts": args.efforts,
"repeats": args.repeats,
"total_generations": len(scorable) * len(args.efforts) * args.repeats,
"allow_reconstructed": args.allow_reconstructed,
"cases": [
{"case_number": c["case_number"], "supply_mode": c["supply_mode"],
"scorable": c["scorable"], "outcome": c["outcome"],
"discussion_words": c["final_discussion_words"]}
for c in cases
],
}
if args.dry_run:
return {"dry_run": True, "grid": grid_summary, "by_effort": []}
per_effort_runs = {e: [] for e in args.efforts}
per_case = []
for c in scorable:
case_cells = []
for effort in args.efforts:
# Per-cell guard (INV-G8): any single (case, effort) failure — a
# transient rate-limit, a too-large prompt, an unexpected raise — is
# logged and skipped, never allowed to abort the whole run.
try:
cell = await _score_block_yod_cell(
UUID(c["case_id"]), effort, c["final_discussion"],
c["final_total_words"], c["outcome"], args.repeats, c["direction"],
)
except Exception as exc: # noqa: BLE001 — harness must survive any cell failure
logger.warning(
"block-yod cell skipped: case=%s effort=%s mode=%s%s",
c["case_number"], effort, c["supply_mode"], exc,
)
continue
per_effort_runs[effort].append(cell)
case_cells.append({k: cell[k] for k in
("effort", "distance", "anti_pattern_total",
"change_percent", "golden_ratio_deviation_pp", "n")})
per_case.append({"case_number": c["case_number"],
"supply_mode": c["supply_mode"], "cells": case_cells})
# Incremental persistence (INV-G8): flush after EACH case so a crash later
# never loses cases already generated. _write_report tolerates partial rows.
try:
_write_report({"dry_run": False, "grid": grid_summary,
"by_effort": _effort_rows(per_effort_runs, args.efforts),
"per_case": per_case,
"recommended": (recommend_effort(_effort_rows(per_effort_runs, args.efforts)) or {}).get("effort")},
ts)
except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run
logger.warning("incremental report write failed after case=%s%s",
c["case_number"], exc)
effort_rows = _effort_rows(per_effort_runs, args.efforts)
rec = recommend_effort(effort_rows)
return {"dry_run": False, "grid": grid_summary, "by_effort": effort_rows,
"per_case": per_case, "recommended": rec["effort"] if rec else None}
def _effort_rows(per_effort_runs, efforts):
"""Mean each metric across cases for each effort → one comparable row/effort."""
rows = []
for effort in efforts:
cells = per_effort_runs[effort]
if not cells:
continue
ratios = [c["golden_ratio_deviation_pp"] for c in cells
if c.get("golden_ratio_deviation_pp") is not None]
rows.append({
"effort": effort,
"distance": round(mean(c["distance"] for c in cells), 4),
"anti_pattern_total": round(mean(c["anti_pattern_total"] for c in cells), 2),
"change_percent": round(mean(c["change_percent"] for c in cells), 2),
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
"n": len(cells),
})
return rows
# ── report ───────────────────────────────────────────────────────────────────
def _ts():
return datetime.now(IL_TZ).strftime("%Y%m%dT%H%M%S-IL")
def _write_report(result, ts):
OUT_DIR.mkdir(parents=True, exist_ok=True)
jp = OUT_DIR / f"block-yod-calibration-{ts}.json"
mp = OUT_DIR / f"block-yod-calibration-{ts}.md"
jp.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
g = result["grid"]
il_now = datetime.now(IL_TZ).strftime("%Y-%m-%d %H:%M")
n_recon = sum(1 for c in g["cases"] if c["supply_mode"] == "reconstructed")
n_stored = sum(1 for c in g["cases"] if c["supply_mode"] == "stored")
n_skip = sum(1 for c in g["cases"] if c["supply_mode"] == "skip")
lines = [
f"# #212 — כיול effort לבלוק י' (דיון והכרעה) מול הסופיים — {ts}\n",
f"> נוצר: {il_now} (שעון ישראל · Asia/Jerusalem)\n",
f"> ⚠️ **גודל-מדגם: {g['n_scorable']} תיקים ניתנים-לכיול** "
f"(מתוך {g['n_finals']} סופיים). זוהי **עדות-כיוון, לא רגרסיה**. "
"ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n",
f"> אספקת-כיוון: {n_stored} מאוחסן · {n_recon} משוחזר · {n_skip} מדולג. "
"כיוון **משוחזר** = נגזר מתוצאת-ההחלטה + פתיח-הדיון של הסופי באותו תיק "
"(לא סיעור-מוחות אמיתי) → התוצאה נאמנה-לכיוון פחות מתיק עם כיוון-מאושר שמור.\n",
f"- בלוק: {g['block']} (section={g['section']}) · ברירת-מחדל נוכחית: "
f"**{g['current_default'] or ''}** (נקבע 'by feel', block_writer.py:72)",
f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}",
f"- סך ייצורי-מודל: {g['total_generations']}",
"",
"## אספקת-כיוון per-תיק\n",
"| case | supply_mode | discussion_words | outcome | scorable |",
"|---|---|---|---|---|",
]
for c in g["cases"]:
lines.append(
f"| {c['case_number']} | {c['supply_mode']} | {c['discussion_words']} | "
f"{c['outcome']} | {'' if c['scorable'] else ''} |")
lines.append("")
if result.get("dry_run"):
lines += ["## DRY-RUN — תכנון הגריד בלבד (ללא ייצור)\n",
f"- תיקים ניתנים-לכיול: **{g['n_scorable']}** · "
f"ייצורים מתוכננים: **{g['total_generations']}**",
""]
mp.write_text("\n".join(lines) + "\n", encoding="utf-8")
return jp, mp
rec = result.get("recommended") or ""
mark = "" if rec == g["current_default"] else " ⬅︎ שינוי מומלץ"
lines += [
f"## המלצה: effort = **{rec}**{mark} (current = {g['current_default'] or ''})\n",
"## טבלת effort (distance נמוך = קרוב יותר לדיון של דפנה)\n",
"| effort | distance | anti_total | change% | ratioΔpp | n |",
"|---|---|---|---|---|---|",
]
for r in result.get("by_effort", []):
star = "" if r["effort"] == result.get("recommended") else ""
ratio = r["golden_ratio_deviation_pp"]
lines.append(
f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | "
f"{r['change_percent']} | {ratio if ratio is not None else ''} | {r['n']} |")
lines.append("")
lines += ["## פירוט per-תיק\n", "| case | mode | effort | distance | anti | change% |",
"|---|---|---|---|---|---|"]
for pc in result.get("per_case", []):
for cell in pc["cells"]:
lines.append(
f"| {pc['case_number']} | {pc['supply_mode']} | {cell['effort']} | "
f"{cell['distance']:.4f} | {cell['anti_pattern_total']} | {cell['change_percent']} |")
lines.append("")
lines.append("> change% מערבב סגנון עם שלמות-תוכן (07-learning §0.7); "
"anti_total הוא הסיגנל הנקי-יותר לסגנון. תיקים עם supply_mode=reconstructed "
"תורמים עדות-כיוון בלבד.\n")
mp.write_text("\n".join(lines) + "\n", encoding="utf-8")
return jp, mp
async def main() -> int:
ap = argparse.ArgumentParser(description="#212 block-yod (discussion) effort calibration harness")
ap.add_argument("--self-test", action="store_true",
help="offline measurement + direction-supply proof (no DB/CLI)")
ap.add_argument("--dry-run", action="store_true",
help="plan the grid + per-case direction-supply mode, no model calls (needs DB)")
ap.add_argument("--efforts", default=",".join(DEFAULT_EFFORTS),
help=f"comma effort grid (default {','.join(DEFAULT_EFFORTS)})")
ap.add_argument("--case", default=None, help="restrict to a single case_number")
ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)")
ap.add_argument("--allow-reconstructed", default="true",
help="true|false — allow the reconstructed-direction fallback (default true). "
"false = STORED-direction cases only (most faithful, may yield 0 cells).")
args = ap.parse_args()
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
args.efforts = [e.strip() for e in args.efforts.split(",") if e.strip()]
bad = [e for e in args.efforts if e not in VALID_EFFORTS]
if bad:
print(f"invalid effort(s): {bad}. valid: {sorted(VALID_EFFORTS)}", file=sys.stderr)
return 2
args.allow_reconstructed = str(args.allow_reconstructed).strip().lower() in ("1", "true", "yes", "y")
ts = _ts()
result = await _run(args, ts)
jp, mp = _write_report(result, ts)
g = result["grid"]
print(f"BLOCK-YOD CALIBRATION: {g['n_scorable']}/{g['n_finals']} scorable finals — "
"DIRECTIONAL EVIDENCE, not a regression")
n_recon = sum(1 for c in g["cases"] if c["supply_mode"] == "reconstructed")
n_stored = sum(1 for c in g["cases"] if c["supply_mode"] == "stored")
print(f" direction supply: {n_stored} stored · {n_recon} reconstructed · "
f"{sum(1 for c in g['cases'] if c['supply_mode']=='skip')} skipped")
if g["n_scorable"] == 0:
print(" no scorable finals (need a parsable discussion section + a suppliable direction).")
elif result.get("dry_run"):
print(f" dry-run: {g['total_generations']} generations planned "
f"({len(g['efforts'])} efforts × {g['n_scorable']} cases × {g['repeats']} repeats)")
else:
print(f" current default={g['current_default']} → recommended={result.get('recommended')}")
print(f" report: {mp}")
return 0
if __name__ == "__main__":
# --self-test is OFFLINE and synchronous (it spins its own short-lived loop for
# the overlay probe), so dispatch it BEFORE asyncio.run(main()) — otherwise the
# probe's asyncio.run would nest inside main's running loop.
if "--self-test" in sys.argv:
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
sys.exit(_self_test())
sys.exit(asyncio.run(main()))