feat(eval): block-yod effort calibration harness with read-only direction supply (#212)
בלוק י' (דיון והכרעה) הוא הבלוק החשוב ביותר אך ה-effort שלו (xhigh) נקבע 'by feel' ולא כויל אמפירית. ה-harness הכללי (#208) לא יכל לכייל אותו כי write_block(block-yod) זורק "ללא כיוון מאושר" — בלוק-הדיון דורש decisions.direction_doc.approved שתיקי-הכיול חסרים. scripts/calibrate_block_yod.py מוסיף שכבת אספקת-כיוון read-only מעל מכונת המדידה של #208: - overlay על db.get_decision_by_case/get_decision לתהליך-הזה-בלבד מזרים direction_doc שמיש ל-write_block — אפס כתיבה ל-DB, מצב-תיק-הייצור לא משתנה (reverted ב-__exit__). - אסטרטגיה per-תיק: stored (כיוון-מאושר שמור, נאמן) → reconstructed (מינימלי מ-decision.outcome + סקשן-הדיון של הסופי באותו תיק, INV-LRN5: אפס מהות חוצת-תיקים, מתויג reconstructed) → skip. - משתמש-חוזר ב-style_distance.block_distance_to_final + _score_cell/ recommend_effort מ-calibrate_effort (G2 — אין מסלול-מדידה מקביל). - כתיבה אינקרמנטלית per-תיק + שער-per-תא לא-פטאלי (INV-G8). ממצא-קורפוס: מ-8 סופיים 7 ניתנים-לכיול; רק 8126-03-25 stored+scorable (8174-12-24 שמור-אך-חסר-סקשן-דיון) → הנסיגה reconstructed היא מה שמפעיל את ה-harness היום. Invariants: G8 (eval-harness), G2 (שימוש-חוזר ב-style_distance/block_writer, אין fork), INV-LRN5 (כיוון משוחזר = סגנון/שיטה בלבד, אפס מהות חוצת-תיקים), claude_session local-only (ייצור דרך claude -p, מקומי-בלבד). אימות: --self-test ALL PASS (28 checks: מדידה + אספקת-כיוון + overlay install/revert) · --dry-run מול DB חי (read-only): 7/8 scorable, 28 ייצורים מתוכננים · --allow-reconstructed false → 1 (נאמן בלבד) · --case מסנן. אפס mutation calls ב-harness (read-only מאומת). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
3
.gitignore
vendored
3
.gitignore
vendored
@@ -19,6 +19,9 @@ kiryat-yearim/
|
|||||||
continuation-prompt.md
|
continuation-prompt.md
|
||||||
node_modules/
|
node_modules/
|
||||||
data/eval/eval-report-*
|
data/eval/eval-report-*
|
||||||
|
# effort-calibration run outputs (#208 / #212) — generated, not source
|
||||||
|
data/eval/effort-calibration-*
|
||||||
|
data/eval/block-yod-calibration-*
|
||||||
# revert snapshot for migrate_agent_adapter.py (runtime state)
|
# revert snapshot for migrate_agent_adapter.py (runtime state)
|
||||||
data/adapter-migration-state.json
|
data/adapter-migration-state.json
|
||||||
# frontmatter-stripped instruction copies for content_arg adapters (generated)
|
# frontmatter-stripped instruction copies for content_arg adapters (generated)
|
||||||
|
|||||||
@@ -77,6 +77,7 @@
|
|||||||
| `ab_halacha_opus48.py` | python | **A/B לא-הרסני לחילוץ הלכות (Claude)** — מריץ מחדש חילוץ הלכות על פסק-דין בודד דרך מודל/effort נבחרים (`AB_MODEL`/`AB_EFFORT`, ברירת-מחדל `claude-opus-4-8`/`xhigh`) ומשווה לסטטיסטיקות ההלכות הקיימות ב-DB **בלי למחוק/לכתוב כלום**. משכפל את `halacha_extractor.extract()` (אותם פרומפטים, בחירת-צ'אנקים, אימות-ציטוט) ומחליף רק את קריאת ה-LLM ב-`claude -p --model --effort`. מפיק `data/ab_halacha_<case>_<effort>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data .venv/bin/python scripts/ab_halacha_opus48.py <case_law_id>`. **ממצא 2026-05-31 (שטיין 1128-08-20):** Opus 4.8@xhigh חילץ 51 מול 124 בייצור (100% quote-verified מול 96%) אך ביטחון מכויל-נמוך יותר (חציון 0.75 מול 0.82) — ולכן **לא** מקטין את תור-האישור-הידני תחת sweep אוטו-אישור conf≥0.78 (26 מול 24). שיפור איכות, לא צמצום-תור. | ידני (החלטת מודל-חילוץ) |
|
| `ab_halacha_opus48.py` | python | **A/B לא-הרסני לחילוץ הלכות (Claude)** — מריץ מחדש חילוץ הלכות על פסק-דין בודד דרך מודל/effort נבחרים (`AB_MODEL`/`AB_EFFORT`, ברירת-מחדל `claude-opus-4-8`/`xhigh`) ומשווה לסטטיסטיקות ההלכות הקיימות ב-DB **בלי למחוק/לכתוב כלום**. משכפל את `halacha_extractor.extract()` (אותם פרומפטים, בחירת-צ'אנקים, אימות-ציטוט) ומחליף רק את קריאת ה-LLM ב-`claude -p --model --effort`. מפיק `data/ab_halacha_<case>_<effort>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data .venv/bin/python scripts/ab_halacha_opus48.py <case_law_id>`. **ממצא 2026-05-31 (שטיין 1128-08-20):** Opus 4.8@xhigh חילץ 51 מול 124 בייצור (100% quote-verified מול 96%) אך ביטחון מכויל-נמוך יותר (חציון 0.75 מול 0.82) — ולכן **לא** מקטין את תור-האישור-הידני תחת sweep אוטו-אישור conf≥0.78 (26 מול 24). שיפור איכות, לא צמצום-תור. | ידני (החלטת מודל-חילוץ) |
|
||||||
| `ab_halacha_codex.py` | python | **A/B לא-הרסני לחילוץ הלכות (Codex/gpt-5.5)** — עמית ל-`ab_halacha_opus48` אך מחליף את `claude -p` ב-`codex exec --model gpt-5.5` (אימות ChatGPT, ללא OPENAI_API_KEY). אותם פרומפטים ואותו הסקת quote-verification. הפלט האחרון של הסוכן (`-o FILE`) נפענח כ-JSON. `AB_MODEL` (default `gpt-5.5`), `AB_REASONING` low/medium/high/xhigh (default `medium`), `AB_CONCURRENCY` (default 1), `CODEX_BIN`. מפיק `data/ab_halacha_codex_<case>_<model>_<reasoning>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data mcp-server/.venv/bin/python scripts/ab_halacha_codex.py <case_law_id>`. **ממצא 2026-06-17 (8181-21 האוניברסיטה העברית):** gpt-5.5@medium חילץ 27 מול 28 של Opus (quote-verified 100%/100%), ביטחון חציון 0.86 מול 0.78 — אך **0 פריטים מתחת ל-0.7** (לעומת 9/28 של Opus = 32%), דבר המצביע על over-confidence. holding↑ (12 מול 7), procedural↓ (4 מול 7). **מסקנה: ריאלי כ-fallback חירום; לא מוכן לייצור ללא כיול-ביטחון.** | ידני (בנצ'מרק מודל codex) |
|
| `ab_halacha_codex.py` | python | **A/B לא-הרסני לחילוץ הלכות (Codex/gpt-5.5)** — עמית ל-`ab_halacha_opus48` אך מחליף את `claude -p` ב-`codex exec --model gpt-5.5` (אימות ChatGPT, ללא OPENAI_API_KEY). אותם פרומפטים ואותו הסקת quote-verification. הפלט האחרון של הסוכן (`-o FILE`) נפענח כ-JSON. `AB_MODEL` (default `gpt-5.5`), `AB_REASONING` low/medium/high/xhigh (default `medium`), `AB_CONCURRENCY` (default 1), `CODEX_BIN`. מפיק `data/ab_halacha_codex_<case>_<model>_<reasoning>.json`. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data mcp-server/.venv/bin/python scripts/ab_halacha_codex.py <case_law_id>`. **ממצא 2026-06-17 (8181-21 האוניברסיטה העברית):** gpt-5.5@medium חילץ 27 מול 28 של Opus (quote-verified 100%/100%), ביטחון חציון 0.86 מול 0.78 — אך **0 פריטים מתחת ל-0.7** (לעומת 9/28 של Opus = 32%), דבר המצביע על over-confidence. holding↑ (12 מול 7), procedural↓ (4 מול 7). **מסקנה: ריאלי כ-fallback חירום; לא מוכן לייצור ללא כיול-ביטחון.** | ידני (בנצ'מרק מודל codex) |
|
||||||
| `calibrate_effort.py` | python | **#208 (WS5/Q1, INV-G8 eval-harness) — כיול model×effort של הכותב מול הסופיים.** A/B per-(תיק,בלוק,effort) על `draft_final_pairs` בעלי `final_text`: מייצר מחדש כל בלוק דרך מסלול-הייצור (`block_writer.write_block(effort_override=…)` → `claude_session.query` → `claude -p`, Opus 4.8 נעוץ, **מקומי-בלבד**) ומודד מול ה**סקשן** המתאים בסופי דרך `services/style_distance.block_distance_to_final` (מקור-מדידה יחיד, G2): `change_percent` (compute_diff_stats) · `anti_pattern_total` (`lessons.ANTI_PATTERNS`, הסיגנל הנקי-לסגנון) · `golden_ratio_deviation_pp` · `distance` מרוכב. **ממליץ** per-בלוק על ה-effort בעל ה-distance-הממוצע-הנמוך (לצד ברירת-המחדל מ-#204). מפיק `data/eval/effort-calibration-<ts>.{json,md}` ב**כתיבה אינקרמנטלית** — הדוח נכתב מחדש אחרי כל בלוק שמסתיים, כך שקריסה באמצע הגריד לא מאבדת בלוקים שכבר יוצרו (INV-G8). כל תא-(תיק,בלוק,effort) שנכשל (rate-limit חולף, או חריגה) נרשם כ-warning ומדולג — לא מפיל את הריצה. **`block-yod` (דיון) הוצא מברירות-המחדל** (`CALIBRATABLE_BLOCKS`) — דורש כיוון-מאושר (brainstorm→approve_direction) שתיקי-הכיול חסרים ולכן `write_block` זורק; נשאר ב-`VALID_BLOCKS` לכפייה ידנית (`--blocks block-yod`) על-אחריות-המשתמש, כשהשער-per-תא שומר אותו לא-פטאלי. ⚠️ **גודל-מדגם מודפס בראש הדוח — עדות-כיוון, לא רגרסיה** (מעט סופיים-עלויים). `--self-test` (offline, אפס DB/CLI — מוכיח את לוגיקת-המדידה) · `--dry-run` (תכנון-גריד) · `--efforts`/`--blocks`/`--case`/`--repeats`. **חובה מקומי** (claude CLI; לא בקונטיינר/worktree-ללא-CLI). הרצה: `POSTGRES_PASSWORD=… mcp-server/.venv/bin/python scripts/calibrate_effort.py`. | ידני — לכיול ברירות-effort של הכותב |
|
| `calibrate_effort.py` | python | **#208 (WS5/Q1, INV-G8 eval-harness) — כיול model×effort של הכותב מול הסופיים.** A/B per-(תיק,בלוק,effort) על `draft_final_pairs` בעלי `final_text`: מייצר מחדש כל בלוק דרך מסלול-הייצור (`block_writer.write_block(effort_override=…)` → `claude_session.query` → `claude -p`, Opus 4.8 נעוץ, **מקומי-בלבד**) ומודד מול ה**סקשן** המתאים בסופי דרך `services/style_distance.block_distance_to_final` (מקור-מדידה יחיד, G2): `change_percent` (compute_diff_stats) · `anti_pattern_total` (`lessons.ANTI_PATTERNS`, הסיגנל הנקי-לסגנון) · `golden_ratio_deviation_pp` · `distance` מרוכב. **ממליץ** per-בלוק על ה-effort בעל ה-distance-הממוצע-הנמוך (לצד ברירת-המחדל מ-#204). מפיק `data/eval/effort-calibration-<ts>.{json,md}` ב**כתיבה אינקרמנטלית** — הדוח נכתב מחדש אחרי כל בלוק שמסתיים, כך שקריסה באמצע הגריד לא מאבדת בלוקים שכבר יוצרו (INV-G8). כל תא-(תיק,בלוק,effort) שנכשל (rate-limit חולף, או חריגה) נרשם כ-warning ומדולג — לא מפיל את הריצה. **`block-yod` (דיון) הוצא מברירות-המחדל** (`CALIBRATABLE_BLOCKS`) — דורש כיוון-מאושר (brainstorm→approve_direction) שתיקי-הכיול חסרים ולכן `write_block` זורק; נשאר ב-`VALID_BLOCKS` לכפייה ידנית (`--blocks block-yod`) על-אחריות-המשתמש, כשהשער-per-תא שומר אותו לא-פטאלי. ⚠️ **גודל-מדגם מודפס בראש הדוח — עדות-כיוון, לא רגרסיה** (מעט סופיים-עלויים). `--self-test` (offline, אפס DB/CLI — מוכיח את לוגיקת-המדידה) · `--dry-run` (תכנון-גריד) · `--efforts`/`--blocks`/`--case`/`--repeats`. **חובה מקומי** (claude CLI; לא בקונטיינר/worktree-ללא-CLI). הרצה: `POSTGRES_PASSWORD=… mcp-server/.venv/bin/python scripts/calibrate_effort.py`. | ידני — לכיול ברירות-effort של הכותב |
|
||||||
|
| `calibrate_block_yod.py` | python | **#212 (WS5/Q1, INV-G8 eval-harness) — כיול effort ל-בלוק י' (דיון והכרעה) בלבד.** ה-harness ש-#208 לא יכל לכייל: `write_block(block-yod)` זורק "ללא כיוון מאושר" כי בלוק-הדיון דורש `decisions.direction_doc.approved` שתיקי-הכיול חסרים. **התוספת על #208 — שכבת אספקת-כיוון read-only:** לפני ייצור בלוק-הדיון, ה-harness מזרים `direction_doc` שמיש דרך **overlay על קריאת-ה-DB** (monkeypatch של `db.get_decision_by_case`/`db.get_decision` לתהליך-הזה-בלבד) — **אפס כתיבה ל-DB**, מצב-תיק-הייצור לא משתנה (השער reverted ב-finally). **אסטרטגיית-אספקה per-תיק:** (1) **stored** (מועדף, נאמן) — כיוון-מאושר שמור (`approved`+`selected_direction`, או סכמת-8126 עם `direction_id`/`structure`); (2) **reconstructed** (נסיגה, עדות-כיוון-בלבד) — `direction_doc` מינימלי שנגזר מ-`decision.outcome` + סקשן-הדיון של הסופי **באותו תיק** (INV-LRN5: outcome+מצביע-לסופי-עצמי, אפס מהות חוצת-תיקים, מתויג `reconstructed`); (3) **skip** אם אין דיון/כיוון. מודד דרך אותו `style_distance.block_distance_to_final` (מקור-מדידה יחיד, G2) ומשתמש-חוזר ב-`_score_cell`/`recommend_effort` מ-`calibrate_effort` (#208). **ממצא-קורפוס (2026-06-30):** מ-8 סופיים, 7 ניתנים-לכיול; רק 8126-03-25 הוא stored+scorable (8174-12-24 שמור-אך-חסר-סקשן-דיון) → `--allow-reconstructed false` נותן תא-יחיד; ברירת-המחדל (reconstructed פעיל) נותנת 7 → הנסיגה היא מה שמפעיל את ה-harness היום. כתיבה אינקרמנטלית per-תיק + שער-per-תא לא-פטאלי (INV-G8). ⚠️ **גודל-מדגם בראש הדוח — עדות-כיוון, לא רגרסיה.** `--self-test` (offline, אפס DB/CLI — מוכיח מדידה + לוגיקת-אספקת-כיוון + overlay) · `--dry-run` (תכנון-גריד + מצב-אספקה per-תיק, צריך DB) · `--efforts`/`--case`/`--repeats`/`--allow-reconstructed`. ברירת-מחדל נוכחית: `xhigh` (`block_writer.py:72`, 'by feel'). **חובה מקומי** (claude CLI; לא בקונטיינר/worktree-ללא-CLI). הרצת-כיול-חי (לבצע ע"י היו"ר): `POSTGRES_PASSWORD=… POSTGRES_HOST=127.0.0.1 POSTGRES_PORT=5433 mcp-server/.venv/bin/python scripts/calibrate_block_yod.py`. | ידני — לכיול ה-effort של בלוק-הדיון (היו"ר מריץ; long, צורך-מכסה) |
|
||||||
| `monitor_halacha_quality.py` | python | מנטר איכות חילוץ הלכות. בודק drift של `avg(confidence)` בין baseline היסטורי לחלון אחרון. מחזיר JSON מטריקות + alert ב-stderr אם drift > threshold (ברירת מחדל 5%). 2 סדרות: trusted (approved+published) ו-all_extracted. תומך `--window N` / `--threshold X` / `--min-sample N` / `--silent` / `--exit-on-alert`. רץ ב-container או מקומית עם `mcp-server/.venv` (אין תלות ב-LLM, רק SQL). **תזמון מומלץ**: `0 8 * * 1` (יום ראשון 08:00, שבועי) | `0 8 * * 1` (לתזמן) |
|
| `monitor_halacha_quality.py` | python | מנטר איכות חילוץ הלכות. בודק drift של `avg(confidence)` בין baseline היסטורי לחלון אחרון. מחזיר JSON מטריקות + alert ב-stderr אם drift > threshold (ברירת מחדל 5%). 2 סדרות: trusted (approved+published) ו-all_extracted. תומך `--window N` / `--threshold X` / `--min-sample N` / `--silent` / `--exit-on-alert`. רץ ב-container או מקומית עם `mcp-server/.venv` (אין תלות ב-LLM, רק SQL). **תזמון מומלץ**: `0 8 * * 1` (יום ראשון 08:00, שבועי) | `0 8 * * 1` (לתזמן) |
|
||||||
| `audit_training_corpus.py` | python | audit של `style_corpus` — לכל החלטה: שדות מטא-דאטה מאוכלסים (`summary`/`outcome`/`key_principles`/`appeal_subtype`/`subject_categories`), קישור ל-`documents` (FK + chunks + embeddings). מפיק `data/audit/corpus-YYYY-MM-DD.json` + summary בקונסול. דרוש `POSTGRES_URL` או POSTGRES_*. אין תלויות חיצוניות מלבד asyncpg. **רץ מהמכונה המקומית** (לא קונטיינר) — חיבור ישיר ל-Postgres :5433 | ידני / קדם-עבודה לפני enrichment של מטא-דאטה |
|
| `audit_training_corpus.py` | python | audit של `style_corpus` — לכל החלטה: שדות מטא-דאטה מאוכלסים (`summary`/`outcome`/`key_principles`/`appeal_subtype`/`subject_categories`), קישור ל-`documents` (FK + chunks + embeddings). מפיק `data/audit/corpus-YYYY-MM-DD.json` + summary בקונסול. דרוש `POSTGRES_URL` או POSTGRES_*. אין תלויות חיצוניות מלבד asyncpg. **רץ מהמכונה המקומית** (לא קונטיינר) — חיבור ישיר ל-Postgres :5433 | ידני / קדם-עבודה לפני enrichment של מטא-דאטה |
|
||||||
| `backfill_style_exemplars.py` | python | **T1 (style-acquisition)** — מאכלס `style_exemplars` מקורפוס דפנה (`style_corpus` + `internal_committee` chair=דפנה): מפצל לסעיפים (`chunker._split_into_sections`) → פסקאות (25-450 מילים) → embed (Voyage) → שמירה עם `section`/`outcome`/`practice_area`. מאפשר לכותב לאחזר פסקאות-בלוק אמיתיות של דפנה (T2/T3). מקור-סגנון בלבד (INV-LRN5). אידמפוטנטי (מנקה per-decision). `--dry-run` (default) / `--apply`. דורש POSTGRES_URL + Voyage. **רץ מקומית** (venv). | ידני (`python scripts/backfill_style_exemplars.py --apply`) |
|
| `backfill_style_exemplars.py` | python | **T1 (style-acquisition)** — מאכלס `style_exemplars` מקורפוס דפנה (`style_corpus` + `internal_committee` chair=דפנה): מפצל לסעיפים (`chunker._split_into_sections`) → פסקאות (25-450 מילים) → embed (Voyage) → שמירה עם `section`/`outcome`/`practice_area`. מאפשר לכותב לאחזר פסקאות-בלוק אמיתיות של דפנה (T2/T3). מקור-סגנון בלבד (INV-LRN5). אידמפוטנטי (מנקה per-decision). `--dry-run` (default) / `--apply`. דורש POSTGRES_URL + Voyage. **רץ מקומית** (venv). | ידני (`python scripts/backfill_style_exemplars.py --apply`) |
|
||||||
|
|||||||
646
scripts/calibrate_block_yod.py
Normal file
646
scripts/calibrate_block_yod.py
Normal file
@@ -0,0 +1,646 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""#212 (WS5 / Q1, INV-G8 eval-harness) — DEDICATED effort calibration for בלוק י'
|
||||||
|
(block-yod, the discussion/decision block — דיון והכרעה).
|
||||||
|
|
||||||
|
WHY A SEPARATE HARNESS (the #208 harness could NOT calibrate block-yod):
|
||||||
|
block-yod is the most important block — its effort is currently `xhigh`, set
|
||||||
|
"by feel" (block_writer.py:72), never empirically calibrated. The general #208
|
||||||
|
harness (scripts/calibrate_effort.py) EXCLUDES it from CALIBRATABLE_BLOCKS
|
||||||
|
because block_writer.write_block(case_id, "block-yod", …) raises
|
||||||
|
ValueError("לא ניתן לכתוב בלוק דיון ללא כיוון מאושר…")
|
||||||
|
— the discussion block requires an APPROVED DIRECTION (brainstorm →
|
||||||
|
approve_direction → decisions.direction_doc{"approved": True, …}) that the
|
||||||
|
calibration cases (the draft_final_pairs finals) do not have. So block-yod is
|
||||||
|
not calibratable standalone without first SUPPLYING a direction.
|
||||||
|
|
||||||
|
WHAT THIS HARNESS ADDS over #208 — the DIRECTION-SUPPLY layer (read-only):
|
||||||
|
Before regenerating block-yod for a calibration case, it makes a usable
|
||||||
|
`direction_doc` available to write_block via a READ-ONLY DB OVERLAY (it
|
||||||
|
monkeypatches db.get_decision_by_case / db.get_decision FOR THIS PROCESS ONLY
|
||||||
|
to overlay the direction into the returned decision dict). NOTHING is written
|
||||||
|
to the DB — production case state is never mutated (no UPDATE), satisfying the
|
||||||
|
"do not pollute real case state" constraint. The overlay is reverted in a
|
||||||
|
finally block.
|
||||||
|
|
||||||
|
DIRECTION-SUPPLY STRATEGY (per-case, documented in the report):
|
||||||
|
1. STORED (preferred, most faithful): if the case's decision already carries
|
||||||
|
a usable approved direction (direction_doc.approved is truthy and it has a
|
||||||
|
selected_direction, OR the 8126-style schema with direction_id/structure),
|
||||||
|
use it verbatim. This is the production-faithful path.
|
||||||
|
2. RECONSTRUCTED (fallback, directional-only): otherwise synthesize a MINIMAL
|
||||||
|
direction_doc from the case's decision.outcome + the final's OWN discussion
|
||||||
|
section (the chair's signed reasoning is the ground truth for "what
|
||||||
|
direction was taken"). Clearly LABELLED reconstructed in the report → the
|
||||||
|
resulting numbers are directional, not faithful to a real brainstorm.
|
||||||
|
INV-LRN5: the reconstructed direction carries the case's OWN outcome +
|
||||||
|
a neutral pointer to its OWN final — no substance dragged across cases.
|
||||||
|
|
||||||
|
EMPIRICAL NOTE (measured on the current corpus, 2026-06-30): of the 8 finals,
|
||||||
|
only 8174-12-24 has a STORED approved direction — and its final has NO parsable
|
||||||
|
discussion section (0 words), so it is not scorable for block-yod anyway. The 7
|
||||||
|
finals that DO have a scorable discussion section all LACK a usable stored
|
||||||
|
direction. Hence the RECONSTRUCTED fallback is what makes this harness useful
|
||||||
|
today; --dry-run prints the per-case supply mode so the operator sees this.
|
||||||
|
|
||||||
|
WHAT IT MEASURES (identical to #208, REUSED — no parallel metric path, G2):
|
||||||
|
Regenerate block-yod at each effort via block_writer.write_block(effort_override)
|
||||||
|
(the PRODUCTION generation path → claude_session.query → `claude -p`, pinned
|
||||||
|
Opus 4.8, local-only), then score vs the final's DISCUSSION section via
|
||||||
|
services.style_distance.block_distance_to_final (change_percent /
|
||||||
|
anti_pattern_total / golden_ratio_deviation_pp / composite distance). The
|
||||||
|
recommend_effort, aggregate_cell, _score_cell and direction-free scaffolding are
|
||||||
|
imported from calibrate_effort (#208) — this file only supplies the direction
|
||||||
|
and restricts the grid to block-yod.
|
||||||
|
|
||||||
|
GENERATION PATH (do not violate — reference_claude_generation_path /
|
||||||
|
claude_session docstring): write_block → claude_session.query → `claude -p` uses
|
||||||
|
the local claude.ai session. It runs ONLY on the host where the `claude` CLI
|
||||||
|
exists (NOT the legal-ai container, NOT a symlinked worktree without CLI access).
|
||||||
|
Hence the live A/B is HOST-ONLY; --self-test proves the measurement +
|
||||||
|
direction-supply logic offline with zero model calls.
|
||||||
|
|
||||||
|
⚠️ SAMPLE-SIZE CAVEAT (honored, not hidden): very few cases have an uploaded
|
||||||
|
final WITH a parsable discussion section. The report prints n_finals PROMINENTLY
|
||||||
|
and labels the output DIRECTIONAL EVIDENCE, not a regression. With n<3 the
|
||||||
|
recommendation is advisory only; the chair/operator decides whether to adopt.
|
||||||
|
|
||||||
|
Usage (mcp-server venv; live needs POSTGRES + the `claude` CLI on the host):
|
||||||
|
PY=/home/chaim/legal-ai/mcp-server/.venv/bin/python
|
||||||
|
$PY scripts/calibrate_block_yod.py --self-test # offline proof, no DB/CLI
|
||||||
|
$PY scripts/calibrate_block_yod.py --dry-run # plan grid + supply modes (needs DB, no model)
|
||||||
|
POSTGRES_PASSWORD=… POSTGRES_HOST=127.0.0.1 POSTGRES_PORT=5433 \
|
||||||
|
$PY scripts/calibrate_block_yod.py # live A/B over finals (host-only)
|
||||||
|
… --efforts low,medium,high,xhigh # override the effort grid
|
||||||
|
… --case 1130-08-25 # a single case
|
||||||
|
… --repeats 2 # avg N gens/cell (noise)
|
||||||
|
… --allow-reconstructed false # STORED-direction cases only (faithful)
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import asyncio
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from statistics import mean
|
||||||
|
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||||
|
sys.path.insert(0, str(REPO_ROOT / "mcp-server" / "src"))
|
||||||
|
sys.path.insert(0, str(REPO_ROOT / "scripts"))
|
||||||
|
|
||||||
|
if "POSTGRES_URL" not in os.environ:
|
||||||
|
os.environ["POSTGRES_URL"] = (
|
||||||
|
f"postgres://{os.environ.get('POSTGRES_USER','legal_ai')}:"
|
||||||
|
f"{os.environ.get('POSTGRES_PASSWORD','')}@"
|
||||||
|
f"{os.environ.get('POSTGRES_HOST','127.0.0.1')}:"
|
||||||
|
f"{os.environ.get('POSTGRES_PORT','5433')}/"
|
||||||
|
f"{os.environ.get('POSTGRES_DB','legal_ai')}"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Reuse the #208 machinery wholesale (G2 — no parallel harness/metric path).
|
||||||
|
# These are pure/offline-safe imports; calibrate_effort does no work at import.
|
||||||
|
from calibrate_effort import ( # noqa: E402
|
||||||
|
IL_TZ,
|
||||||
|
OUT_DIR,
|
||||||
|
VALID_EFFORTS,
|
||||||
|
_current_default,
|
||||||
|
_finals_for_calibration,
|
||||||
|
_score_cell,
|
||||||
|
recommend_effort,
|
||||||
|
)
|
||||||
|
|
||||||
|
BLOCK_ID = "block-yod"
|
||||||
|
SECTION = "discussion" # block-yod's golden-ratio section (style_distance._BLOCK_TO_SECTION)
|
||||||
|
DEFAULT_EFFORTS = ["low", "medium", "high", "xhigh"]
|
||||||
|
|
||||||
|
|
||||||
|
# ── direction supply (the block-yod-specific piece) ──────────────────────────
|
||||||
|
def _stored_direction_is_usable(direction_doc: dict | None) -> bool:
|
||||||
|
"""True if the case's stored direction_doc can drive write_block(block-yod).
|
||||||
|
|
||||||
|
write_block gates on `direction_doc.get("approved")` being truthy. Two stored
|
||||||
|
schemas exist in the corpus:
|
||||||
|
• brainstorm schema (build_direction_doc): {"approved": True, "selected_direction": …}
|
||||||
|
• chair-approved schema (8126-style): {"direction_id", "direction_name",
|
||||||
|
"structure", "template", …} — has NO "approved" key, so write_block would
|
||||||
|
reject it as-is. We treat it as usable IFF it carries a concrete direction
|
||||||
|
(direction_id/direction_name), and _normalize_stored_direction() stamps
|
||||||
|
approved=True onto the OVERLAY copy (never the DB).
|
||||||
|
Empty / brainstorm-only ({"brainstorm": …, "approved": False}) → not usable.
|
||||||
|
"""
|
||||||
|
dd = direction_doc or {}
|
||||||
|
if dd.get("approved") and (dd.get("selected_direction") or dd.get("reasoning")):
|
||||||
|
return True
|
||||||
|
if dd.get("direction_id") or dd.get("direction_name"):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_stored_direction(direction_doc: dict) -> dict:
|
||||||
|
"""Return an overlay copy of a usable stored direction with approved=True.
|
||||||
|
|
||||||
|
Pure dict transform on a COPY — the real decisions.direction_doc is untouched.
|
||||||
|
The 8126-style schema lacks `approved`; stamping it here lets write_block's gate
|
||||||
|
pass without us mutating the DB.
|
||||||
|
"""
|
||||||
|
dd = dict(direction_doc)
|
||||||
|
dd["approved"] = True
|
||||||
|
return dd
|
||||||
|
|
||||||
|
|
||||||
|
def reconstruct_minimal_direction(outcome: str, final_discussion: str) -> dict:
|
||||||
|
"""Build a MINIMAL, clearly-reconstructed direction_doc (fallback path).
|
||||||
|
|
||||||
|
Faithful enough to make write_block(block-yod) runnable, honest about being a
|
||||||
|
reconstruction: the chair's SIGNED final discussion is the ground truth for
|
||||||
|
"what direction was taken", and the case's OWN outcome fixes the structure.
|
||||||
|
INV-LRN5: case's own outcome + a neutral pointer to its own final only — no
|
||||||
|
cross-case substance. Pure (no DB/LLM) → unit-tested in --self-test.
|
||||||
|
"""
|
||||||
|
from legal_mcp.services.lessons import OUTCOME_LABELS_HE, canonical_outcome
|
||||||
|
|
||||||
|
oc = canonical_outcome(outcome or "rejection")
|
||||||
|
snippet = " ".join((final_discussion or "").split()[:120])
|
||||||
|
return {
|
||||||
|
"approved": True,
|
||||||
|
"reconstructed": True, # marker so the report can label it
|
||||||
|
"outcome": oc,
|
||||||
|
"outcome_hebrew": OUTCOME_LABELS_HE.get(oc, oc),
|
||||||
|
"reasoning": (
|
||||||
|
"כיוון משוחזר לצורך כיול-effort בלבד (לא סיעור-מוחות אמיתי): "
|
||||||
|
"התוצאה והכיוון נגזרים מההחלטה החתומה של היו\"ר בתיק זה. "
|
||||||
|
"תמצית פתיח-הדיון של הסופי: " + snippet
|
||||||
|
),
|
||||||
|
"selected_direction": {
|
||||||
|
"name": f"כיוון משוחזר — {OUTCOME_LABELS_HE.get(oc, oc)}",
|
||||||
|
"reasoning": [
|
||||||
|
"שחזר את שלד-הנימוק מהדיון החתום של היו\"ר בתיק זה.",
|
||||||
|
],
|
||||||
|
"precedents": [],
|
||||||
|
},
|
||||||
|
"additional_notes": (
|
||||||
|
"⚠️ כיוון משוחזר — תוצאת-הכיול היא עדות-כיוון, לא נאמנה למסמך-כיוון אמיתי."
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def plan_direction_supply(
|
||||||
|
decision: dict | None,
|
||||||
|
outcome: str,
|
||||||
|
final_discussion: str,
|
||||||
|
allow_reconstructed: bool,
|
||||||
|
) -> tuple[dict | None, str]:
|
||||||
|
"""Decide how to supply block-yod's direction for one case.
|
||||||
|
|
||||||
|
Returns (overlay_direction_doc | None, mode) where mode ∈
|
||||||
|
{"stored", "reconstructed", "skip"}. Pure (no I/O) → unit-tested.
|
||||||
|
"""
|
||||||
|
stored = (decision or {}).get("direction_doc")
|
||||||
|
if _stored_direction_is_usable(stored):
|
||||||
|
return _normalize_stored_direction(stored), "stored"
|
||||||
|
if allow_reconstructed and (final_discussion or "").strip():
|
||||||
|
return reconstruct_minimal_direction(outcome, final_discussion), "reconstructed"
|
||||||
|
return None, "skip"
|
||||||
|
|
||||||
|
|
||||||
|
class _DirectionOverlay:
|
||||||
|
"""Read-only overlay: make db.get_decision_by_case / db.get_decision return the
|
||||||
|
case's decision WITH `direction_doc` replaced by the supplied overlay — for the
|
||||||
|
target case_id only, for the duration of one cell. Reverted in __exit__.
|
||||||
|
|
||||||
|
write_block re-fetches the decision internally (block_writer.py:383), so passing
|
||||||
|
a direction dict in isn't enough — the DB read must yield it. We patch the read,
|
||||||
|
not the write: ZERO DB mutation (no UPDATE ever issued).
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, case_id, direction_doc: dict):
|
||||||
|
self._case_id = str(case_id)
|
||||||
|
self._dir = direction_doc
|
||||||
|
self._orig_by_case = None
|
||||||
|
self._orig_by_id = None
|
||||||
|
|
||||||
|
def __enter__(self):
|
||||||
|
from legal_mcp.services import db
|
||||||
|
|
||||||
|
self._orig_by_case = db.get_decision_by_case
|
||||||
|
self._orig_by_id = db.get_decision
|
||||||
|
|
||||||
|
async def patched_by_case(case_id, _orig=self._orig_by_case):
|
||||||
|
dec = await _orig(case_id)
|
||||||
|
if dec and str(case_id) == self._case_id:
|
||||||
|
dec = dict(dec)
|
||||||
|
dec["direction_doc"] = self._dir
|
||||||
|
return dec
|
||||||
|
|
||||||
|
async def patched_by_id(decision_id, _orig=self._orig_by_id):
|
||||||
|
dec = await _orig(decision_id)
|
||||||
|
if dec and dec.get("case_id") == self._case_id:
|
||||||
|
dec = dict(dec)
|
||||||
|
dec["direction_doc"] = self._dir
|
||||||
|
return dec
|
||||||
|
|
||||||
|
db.get_decision_by_case = patched_by_case
|
||||||
|
db.get_decision = patched_by_id
|
||||||
|
return self
|
||||||
|
|
||||||
|
def __exit__(self, *exc):
|
||||||
|
from legal_mcp.services import db
|
||||||
|
|
||||||
|
if self._orig_by_case is not None:
|
||||||
|
db.get_decision_by_case = self._orig_by_case
|
||||||
|
if self._orig_by_id is not None:
|
||||||
|
db.get_decision = self._orig_by_id
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
async def _score_block_yod_cell(case_id, effort, final_section, final_total_words,
|
||||||
|
outcome, repeats, direction_doc):
|
||||||
|
"""Score one (case, block-yod, effort) cell with the direction overlaid.
|
||||||
|
|
||||||
|
Wraps the #208 _score_cell (REUSED) inside the read-only direction overlay so
|
||||||
|
write_block's approved-direction gate passes without DB mutation.
|
||||||
|
"""
|
||||||
|
with _DirectionOverlay(case_id, direction_doc):
|
||||||
|
return await _score_cell(
|
||||||
|
case_id, BLOCK_ID, effort, final_section, final_total_words, outcome, repeats,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ── self-test (no DB, no model) ──────────────────────────────────────────────
|
||||||
|
def _self_test() -> int:
|
||||||
|
ok = True
|
||||||
|
|
||||||
|
def chk(name, cond):
|
||||||
|
nonlocal ok
|
||||||
|
ok = ok and cond
|
||||||
|
print(f" {name:50} {'ok' if cond else 'FAIL'}")
|
||||||
|
|
||||||
|
# measurement reuse: block_distance_to_final on the discussion section.
|
||||||
|
from legal_mcp.services.style_distance import block_distance_to_final, split_final_by_section
|
||||||
|
|
||||||
|
final_disc = "אנו סבורים כי דין הערר להידחות. כידוע, הלכה פסוקה היא. " * 30
|
||||||
|
d_same = block_distance_to_final(BLOCK_ID, final_disc, final_disc, "rejection",
|
||||||
|
section_target_total_words=len(final_disc.split()))
|
||||||
|
chk("identical regen ⇒ change_percent==0", d_same["change_percent"] == 0.0)
|
||||||
|
chk("identical regen ⇒ section==discussion", d_same["section"] == SECTION)
|
||||||
|
|
||||||
|
clean = "אין בידנו לקבל את הטענה. יחד עם זאת, מצאנו טעם. " * 20
|
||||||
|
dirty = "## כותרת\n- נקודה\n- נקודה\n* עוד\n### תת\n" * 10
|
||||||
|
d_clean = block_distance_to_final(BLOCK_ID, clean, final_disc, "rejection")
|
||||||
|
d_dirty = block_distance_to_final(BLOCK_ID, dirty, final_disc, "rejection")
|
||||||
|
chk("dirty regen ⇒ more anti-patterns", d_dirty["anti_pattern_total"] > d_clean["anti_pattern_total"])
|
||||||
|
chk("dirty regen ⇒ larger distance", d_dirty["distance"] > d_clean["distance"])
|
||||||
|
|
||||||
|
# ── direction-supply logic ──
|
||||||
|
chk("empty direction ⇒ not usable", not _stored_direction_is_usable({}))
|
||||||
|
chk("brainstorm-only ⇒ not usable",
|
||||||
|
not _stored_direction_is_usable({"brainstorm": {"directions": [1]}, "approved": False}))
|
||||||
|
chk("approved+selected ⇒ usable",
|
||||||
|
_stored_direction_is_usable({"approved": True, "selected_direction": {"name": "x"}}))
|
||||||
|
chk("8126-style (direction_id) ⇒ usable",
|
||||||
|
_stored_direction_is_usable({"direction_id": "d1", "structure": "..."}))
|
||||||
|
|
||||||
|
src = {"direction_id": "d1", "structure": "..."}
|
||||||
|
norm = _normalize_stored_direction(src)
|
||||||
|
chk("normalize stamps approved=True", norm["approved"] is True)
|
||||||
|
chk("normalize does not mutate source", "approved" not in src)
|
||||||
|
|
||||||
|
rec = reconstruct_minimal_direction("rejection", final_disc)
|
||||||
|
chk("reconstruct ⇒ approved", rec["approved"] is True)
|
||||||
|
chk("reconstruct ⇒ marked reconstructed", rec["reconstructed"] is True)
|
||||||
|
chk("reconstruct ⇒ canonical outcome", rec["outcome"] == "rejection")
|
||||||
|
chk("reconstruct ⇒ has selected_direction", bool(rec.get("selected_direction")))
|
||||||
|
chk("reconstruct ⇒ no cross-case precedents (INV-LRN5)",
|
||||||
|
rec["selected_direction"]["precedents"] == [])
|
||||||
|
|
||||||
|
# _build_direction_context must accept BOTH supply shapes (write_block uses it).
|
||||||
|
from legal_mcp.services.block_writer import _build_direction_context
|
||||||
|
ctx_recon = _build_direction_context({"direction_doc": rec})
|
||||||
|
chk("recon direction renders (not 'לא אושר')", "כיוון לא אושר" not in ctx_recon)
|
||||||
|
ctx_stored = _build_direction_context({"direction_doc": norm})
|
||||||
|
chk("stored(8126) direction renders", "כיוון לא אושר" not in ctx_stored)
|
||||||
|
|
||||||
|
# plan_direction_supply: stored wins; else reconstructed; else skip.
|
||||||
|
dd_stored = {"direction_doc": {"approved": True, "selected_direction": {"name": "x"}}}
|
||||||
|
_, m1 = plan_direction_supply(dd_stored, "rejection", final_disc, allow_reconstructed=True)
|
||||||
|
chk("plan: stored-direction case ⇒ stored", m1 == "stored")
|
||||||
|
_, m2 = plan_direction_supply({"direction_doc": {}}, "rejection", final_disc, True)
|
||||||
|
chk("plan: no stored + discussion ⇒ reconstructed", m2 == "reconstructed")
|
||||||
|
_, m3 = plan_direction_supply({"direction_doc": {}}, "rejection", final_disc, False)
|
||||||
|
chk("plan: no stored + reconstructed disabled ⇒ skip", m3 == "skip")
|
||||||
|
_, m4 = plan_direction_supply({"direction_doc": {}}, "rejection", "", True)
|
||||||
|
chk("plan: no stored + no discussion ⇒ skip", m4 == "skip")
|
||||||
|
|
||||||
|
# overlay monkeypatch installs + reverts cleanly (no live DB needed).
|
||||||
|
import legal_mcp.services.db as real_db
|
||||||
|
saved_bc, saved_bi = real_db.get_decision_by_case, real_db.get_decision
|
||||||
|
|
||||||
|
async def fake_by_case(cid):
|
||||||
|
return {"id": "d", "case_id": "C1", "direction_doc": {"approved": False}}
|
||||||
|
|
||||||
|
async def fake_by_id(did):
|
||||||
|
return {"id": "d", "case_id": "C1", "direction_doc": {"approved": False}}
|
||||||
|
|
||||||
|
real_db.get_decision_by_case = fake_by_case
|
||||||
|
real_db.get_decision = fake_by_id
|
||||||
|
try:
|
||||||
|
ov_dir = {"approved": True, "reconstructed": True}
|
||||||
|
|
||||||
|
async def _probe():
|
||||||
|
with _DirectionOverlay("C1", ov_dir):
|
||||||
|
d1 = await real_db.get_decision_by_case("C1")
|
||||||
|
d2 = await real_db.get_decision_by_case("OTHER")
|
||||||
|
d3 = await real_db.get_decision("d")
|
||||||
|
return d1, d2, d3
|
||||||
|
|
||||||
|
# --self-test is dispatched at the __main__ entry, OUTSIDE asyncio.run,
|
||||||
|
# so a plain asyncio.run here is safe (no enclosing event loop).
|
||||||
|
d1, d2, d3 = asyncio.run(_probe())
|
||||||
|
chk("overlay: target case gets overlaid direction", d1["direction_doc"]["approved"] is True)
|
||||||
|
chk("overlay: non-target case untouched", d2["direction_doc"]["approved"] is False)
|
||||||
|
chk("overlay: get_decision path overlaid", d3["direction_doc"]["reconstructed"] is True)
|
||||||
|
chk("overlay: reverted after context", real_db.get_decision_by_case is fake_by_case)
|
||||||
|
finally:
|
||||||
|
real_db.get_decision_by_case = saved_bc
|
||||||
|
real_db.get_decision = saved_bi
|
||||||
|
|
||||||
|
# reused #208 primitives still behave.
|
||||||
|
r = recommend_effort([
|
||||||
|
{"effort": "high", "distance": 0.22, "anti_pattern_total": 2, "change_percent": 25, "n": 1},
|
||||||
|
{"effort": "xhigh", "distance": 0.22, "anti_pattern_total": 1, "change_percent": 22, "n": 1},
|
||||||
|
])
|
||||||
|
chk("recommend tie → fewer anti-patterns", r["effort"] == "xhigh")
|
||||||
|
chk("current default for block-yod == xhigh", _current_default(BLOCK_ID) == "xhigh")
|
||||||
|
secs = split_final_by_section(
|
||||||
|
"רקע עובדתי\nהמקרקעין. " * 8 + "\n\nדיון והכרעה\nאנו סבורים. " * 8)
|
||||||
|
chk("split returns golden sections only",
|
||||||
|
set(secs).issubset({"background", "claims", "discussion", "summary"}))
|
||||||
|
|
||||||
|
print("ALL PASS" if ok else "*** FAILURES ***")
|
||||||
|
return 0 if ok else 1
|
||||||
|
|
||||||
|
|
||||||
|
# ── live A/B (host-only — needs DB + `claude` CLI) ───────────────────────────
|
||||||
|
async def _gather_cases(case_filter, allow_reconstructed):
|
||||||
|
"""Build per-case metadata + direction-supply plan over the finals."""
|
||||||
|
from uuid import UUID
|
||||||
|
from legal_mcp.services import db
|
||||||
|
from legal_mcp.services.lessons import canonical_outcome
|
||||||
|
from legal_mcp.services.style_distance import split_final_by_section
|
||||||
|
|
||||||
|
finals = await _finals_for_calibration(case_filter)
|
||||||
|
cases = []
|
||||||
|
for f in finals:
|
||||||
|
case = await db.get_case_by_number(f["case_number"]) if f.get("case_number") else None
|
||||||
|
if not case:
|
||||||
|
continue
|
||||||
|
decision = await db.get_decision_by_case(UUID(case["id"]))
|
||||||
|
outcome = canonical_outcome((decision or {}).get("outcome", "rejection"))
|
||||||
|
sections = split_final_by_section(f.get("final_text", ""))
|
||||||
|
final_discussion = sections.get(SECTION, "")
|
||||||
|
final_total_words = len((f.get("final_text", "") or "").split())
|
||||||
|
overlay, mode = plan_direction_supply(
|
||||||
|
decision, outcome, final_discussion, allow_reconstructed)
|
||||||
|
cases.append({
|
||||||
|
"case_number": f["case_number"],
|
||||||
|
"case_id": case["id"],
|
||||||
|
"outcome": outcome,
|
||||||
|
"final_discussion": final_discussion,
|
||||||
|
"final_discussion_words": len(final_discussion.split()),
|
||||||
|
"final_total_words": final_total_words,
|
||||||
|
"direction": overlay,
|
||||||
|
"supply_mode": mode,
|
||||||
|
# scorable IFF it has a discussion section AND a supplied direction.
|
||||||
|
"scorable": bool(final_discussion.strip() and overlay is not None),
|
||||||
|
})
|
||||||
|
return cases
|
||||||
|
|
||||||
|
|
||||||
|
async def _run(args, ts):
|
||||||
|
from uuid import UUID
|
||||||
|
|
||||||
|
cases = await _gather_cases(args.case, args.allow_reconstructed)
|
||||||
|
scorable = [c for c in cases if c["scorable"]]
|
||||||
|
|
||||||
|
grid_summary = {
|
||||||
|
"block": BLOCK_ID,
|
||||||
|
"section": SECTION,
|
||||||
|
"n_finals": len(cases),
|
||||||
|
"n_scorable": len(scorable),
|
||||||
|
"current_default": _current_default(BLOCK_ID),
|
||||||
|
"efforts": args.efforts,
|
||||||
|
"repeats": args.repeats,
|
||||||
|
"total_generations": len(scorable) * len(args.efforts) * args.repeats,
|
||||||
|
"allow_reconstructed": args.allow_reconstructed,
|
||||||
|
"cases": [
|
||||||
|
{"case_number": c["case_number"], "supply_mode": c["supply_mode"],
|
||||||
|
"scorable": c["scorable"], "outcome": c["outcome"],
|
||||||
|
"discussion_words": c["final_discussion_words"]}
|
||||||
|
for c in cases
|
||||||
|
],
|
||||||
|
}
|
||||||
|
|
||||||
|
if args.dry_run:
|
||||||
|
return {"dry_run": True, "grid": grid_summary, "by_effort": []}
|
||||||
|
|
||||||
|
per_effort_runs = {e: [] for e in args.efforts}
|
||||||
|
per_case = []
|
||||||
|
for c in scorable:
|
||||||
|
case_cells = []
|
||||||
|
for effort in args.efforts:
|
||||||
|
# Per-cell guard (INV-G8): any single (case, effort) failure — a
|
||||||
|
# transient rate-limit, a too-large prompt, an unexpected raise — is
|
||||||
|
# logged and skipped, never allowed to abort the whole run.
|
||||||
|
try:
|
||||||
|
cell = await _score_block_yod_cell(
|
||||||
|
UUID(c["case_id"]), effort, c["final_discussion"],
|
||||||
|
c["final_total_words"], c["outcome"], args.repeats, c["direction"],
|
||||||
|
)
|
||||||
|
except Exception as exc: # noqa: BLE001 — harness must survive any cell failure
|
||||||
|
logger.warning(
|
||||||
|
"block-yod cell skipped: case=%s effort=%s mode=%s — %s",
|
||||||
|
c["case_number"], effort, c["supply_mode"], exc,
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
per_effort_runs[effort].append(cell)
|
||||||
|
case_cells.append({k: cell[k] for k in
|
||||||
|
("effort", "distance", "anti_pattern_total",
|
||||||
|
"change_percent", "golden_ratio_deviation_pp", "n")})
|
||||||
|
per_case.append({"case_number": c["case_number"],
|
||||||
|
"supply_mode": c["supply_mode"], "cells": case_cells})
|
||||||
|
|
||||||
|
# Incremental persistence (INV-G8): flush after EACH case so a crash later
|
||||||
|
# never loses cases already generated. _write_report tolerates partial rows.
|
||||||
|
try:
|
||||||
|
_write_report({"dry_run": False, "grid": grid_summary,
|
||||||
|
"by_effort": _effort_rows(per_effort_runs, args.efforts),
|
||||||
|
"per_case": per_case,
|
||||||
|
"recommended": (recommend_effort(_effort_rows(per_effort_runs, args.efforts)) or {}).get("effort")},
|
||||||
|
ts)
|
||||||
|
except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run
|
||||||
|
logger.warning("incremental report write failed after case=%s — %s",
|
||||||
|
c["case_number"], exc)
|
||||||
|
|
||||||
|
effort_rows = _effort_rows(per_effort_runs, args.efforts)
|
||||||
|
rec = recommend_effort(effort_rows)
|
||||||
|
return {"dry_run": False, "grid": grid_summary, "by_effort": effort_rows,
|
||||||
|
"per_case": per_case, "recommended": rec["effort"] if rec else None}
|
||||||
|
|
||||||
|
|
||||||
|
def _effort_rows(per_effort_runs, efforts):
|
||||||
|
"""Mean each metric across cases for each effort → one comparable row/effort."""
|
||||||
|
rows = []
|
||||||
|
for effort in efforts:
|
||||||
|
cells = per_effort_runs[effort]
|
||||||
|
if not cells:
|
||||||
|
continue
|
||||||
|
ratios = [c["golden_ratio_deviation_pp"] for c in cells
|
||||||
|
if c.get("golden_ratio_deviation_pp") is not None]
|
||||||
|
rows.append({
|
||||||
|
"effort": effort,
|
||||||
|
"distance": round(mean(c["distance"] for c in cells), 4),
|
||||||
|
"anti_pattern_total": round(mean(c["anti_pattern_total"] for c in cells), 2),
|
||||||
|
"change_percent": round(mean(c["change_percent"] for c in cells), 2),
|
||||||
|
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
|
||||||
|
"n": len(cells),
|
||||||
|
})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
# ── report ───────────────────────────────────────────────────────────────────
|
||||||
|
def _ts():
|
||||||
|
return datetime.now(IL_TZ).strftime("%Y%m%dT%H%M%S-IL")
|
||||||
|
|
||||||
|
|
||||||
|
def _write_report(result, ts):
|
||||||
|
OUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||||
|
jp = OUT_DIR / f"block-yod-calibration-{ts}.json"
|
||||||
|
mp = OUT_DIR / f"block-yod-calibration-{ts}.md"
|
||||||
|
jp.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
g = result["grid"]
|
||||||
|
il_now = datetime.now(IL_TZ).strftime("%Y-%m-%d %H:%M")
|
||||||
|
n_recon = sum(1 for c in g["cases"] if c["supply_mode"] == "reconstructed")
|
||||||
|
n_stored = sum(1 for c in g["cases"] if c["supply_mode"] == "stored")
|
||||||
|
n_skip = sum(1 for c in g["cases"] if c["supply_mode"] == "skip")
|
||||||
|
lines = [
|
||||||
|
f"# #212 — כיול effort לבלוק י' (דיון והכרעה) מול הסופיים — {ts}\n",
|
||||||
|
f"> נוצר: {il_now} (שעון ישראל · Asia/Jerusalem)\n",
|
||||||
|
f"> ⚠️ **גודל-מדגם: {g['n_scorable']} תיקים ניתנים-לכיול** "
|
||||||
|
f"(מתוך {g['n_finals']} סופיים). זוהי **עדות-כיוון, לא רגרסיה**. "
|
||||||
|
"ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n",
|
||||||
|
f"> אספקת-כיוון: {n_stored} מאוחסן · {n_recon} משוחזר · {n_skip} מדולג. "
|
||||||
|
"כיוון **משוחזר** = נגזר מתוצאת-ההחלטה + פתיח-הדיון של הסופי באותו תיק "
|
||||||
|
"(לא סיעור-מוחות אמיתי) → התוצאה נאמנה-לכיוון פחות מתיק עם כיוון-מאושר שמור.\n",
|
||||||
|
f"- בלוק: {g['block']} (section={g['section']}) · ברירת-מחדל נוכחית: "
|
||||||
|
f"**{g['current_default'] or '—'}** (נקבע 'by feel', block_writer.py:72)",
|
||||||
|
f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}",
|
||||||
|
f"- סך ייצורי-מודל: {g['total_generations']}",
|
||||||
|
"",
|
||||||
|
"## אספקת-כיוון per-תיק\n",
|
||||||
|
"| case | supply_mode | discussion_words | outcome | scorable |",
|
||||||
|
"|---|---|---|---|---|",
|
||||||
|
]
|
||||||
|
for c in g["cases"]:
|
||||||
|
lines.append(
|
||||||
|
f"| {c['case_number']} | {c['supply_mode']} | {c['discussion_words']} | "
|
||||||
|
f"{c['outcome']} | {'✓' if c['scorable'] else '—'} |")
|
||||||
|
lines.append("")
|
||||||
|
|
||||||
|
if result.get("dry_run"):
|
||||||
|
lines += ["## DRY-RUN — תכנון הגריד בלבד (ללא ייצור)\n",
|
||||||
|
f"- תיקים ניתנים-לכיול: **{g['n_scorable']}** · "
|
||||||
|
f"ייצורים מתוכננים: **{g['total_generations']}**",
|
||||||
|
""]
|
||||||
|
mp.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||||
|
return jp, mp
|
||||||
|
|
||||||
|
rec = result.get("recommended") or "—"
|
||||||
|
mark = "" if rec == g["current_default"] else " ⬅︎ שינוי מומלץ"
|
||||||
|
lines += [
|
||||||
|
f"## המלצה: effort = **{rec}**{mark} (current = {g['current_default'] or '—'})\n",
|
||||||
|
"## טבלת effort (distance נמוך = קרוב יותר לדיון של דפנה)\n",
|
||||||
|
"| effort | distance | anti_total | change% | ratioΔpp | n |",
|
||||||
|
"|---|---|---|---|---|---|",
|
||||||
|
]
|
||||||
|
for r in result.get("by_effort", []):
|
||||||
|
star = " ⭐" if r["effort"] == result.get("recommended") else ""
|
||||||
|
ratio = r["golden_ratio_deviation_pp"]
|
||||||
|
lines.append(
|
||||||
|
f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | "
|
||||||
|
f"{r['change_percent']} | {ratio if ratio is not None else '—'} | {r['n']} |")
|
||||||
|
lines.append("")
|
||||||
|
lines += ["## פירוט per-תיק\n", "| case | mode | effort | distance | anti | change% |",
|
||||||
|
"|---|---|---|---|---|---|"]
|
||||||
|
for pc in result.get("per_case", []):
|
||||||
|
for cell in pc["cells"]:
|
||||||
|
lines.append(
|
||||||
|
f"| {pc['case_number']} | {pc['supply_mode']} | {cell['effort']} | "
|
||||||
|
f"{cell['distance']:.4f} | {cell['anti_pattern_total']} | {cell['change_percent']} |")
|
||||||
|
lines.append("")
|
||||||
|
lines.append("> change% מערבב סגנון עם שלמות-תוכן (07-learning §0.7); "
|
||||||
|
"anti_total הוא הסיגנל הנקי-יותר לסגנון. תיקים עם supply_mode=reconstructed "
|
||||||
|
"תורמים עדות-כיוון בלבד.\n")
|
||||||
|
mp.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||||
|
return jp, mp
|
||||||
|
|
||||||
|
|
||||||
|
async def main() -> int:
|
||||||
|
ap = argparse.ArgumentParser(description="#212 block-yod (discussion) effort calibration harness")
|
||||||
|
ap.add_argument("--self-test", action="store_true",
|
||||||
|
help="offline measurement + direction-supply proof (no DB/CLI)")
|
||||||
|
ap.add_argument("--dry-run", action="store_true",
|
||||||
|
help="plan the grid + per-case direction-supply mode, no model calls (needs DB)")
|
||||||
|
ap.add_argument("--efforts", default=",".join(DEFAULT_EFFORTS),
|
||||||
|
help=f"comma effort grid (default {','.join(DEFAULT_EFFORTS)})")
|
||||||
|
ap.add_argument("--case", default=None, help="restrict to a single case_number")
|
||||||
|
ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)")
|
||||||
|
ap.add_argument("--allow-reconstructed", default="true",
|
||||||
|
help="true|false — allow the reconstructed-direction fallback (default true). "
|
||||||
|
"false = STORED-direction cases only (most faithful, may yield 0 cells).")
|
||||||
|
args = ap.parse_args()
|
||||||
|
|
||||||
|
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
||||||
|
|
||||||
|
args.efforts = [e.strip() for e in args.efforts.split(",") if e.strip()]
|
||||||
|
bad = [e for e in args.efforts if e not in VALID_EFFORTS]
|
||||||
|
if bad:
|
||||||
|
print(f"invalid effort(s): {bad}. valid: {sorted(VALID_EFFORTS)}", file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
args.allow_reconstructed = str(args.allow_reconstructed).strip().lower() in ("1", "true", "yes", "y")
|
||||||
|
|
||||||
|
ts = _ts()
|
||||||
|
result = await _run(args, ts)
|
||||||
|
jp, mp = _write_report(result, ts)
|
||||||
|
|
||||||
|
g = result["grid"]
|
||||||
|
print(f"BLOCK-YOD CALIBRATION: {g['n_scorable']}/{g['n_finals']} scorable finals — "
|
||||||
|
"DIRECTIONAL EVIDENCE, not a regression")
|
||||||
|
n_recon = sum(1 for c in g["cases"] if c["supply_mode"] == "reconstructed")
|
||||||
|
n_stored = sum(1 for c in g["cases"] if c["supply_mode"] == "stored")
|
||||||
|
print(f" direction supply: {n_stored} stored · {n_recon} reconstructed · "
|
||||||
|
f"{sum(1 for c in g['cases'] if c['supply_mode']=='skip')} skipped")
|
||||||
|
if g["n_scorable"] == 0:
|
||||||
|
print(" no scorable finals (need a parsable discussion section + a suppliable direction).")
|
||||||
|
elif result.get("dry_run"):
|
||||||
|
print(f" dry-run: {g['total_generations']} generations planned "
|
||||||
|
f"({len(g['efforts'])} efforts × {g['n_scorable']} cases × {g['repeats']} repeats)")
|
||||||
|
else:
|
||||||
|
print(f" current default={g['current_default']} → recommended={result.get('recommended')}")
|
||||||
|
print(f" report: {mp}")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
# --self-test is OFFLINE and synchronous (it spins its own short-lived loop for
|
||||||
|
# the overlay probe), so dispatch it BEFORE asyncio.run(main()) — otherwise the
|
||||||
|
# probe's asyncio.run would nest inside main's running loop.
|
||||||
|
if "--self-test" in sys.argv:
|
||||||
|
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
||||||
|
sys.exit(_self_test())
|
||||||
|
sys.exit(asyncio.run(main()))
|
||||||
Reference in New Issue
Block a user