Files
legal-ai/scripts/calibrate_block_yod.py
Chaim 858e05fb8b
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
feat(eval): block-yod effort calibration harness with read-only direction supply (#212)
בלוק י' (דיון והכרעה) הוא הבלוק החשוב ביותר אך ה-effort שלו (xhigh) נקבע
'by feel' ולא כויל אמפירית. ה-harness הכללי (#208) לא יכל לכייל אותו כי
write_block(block-yod) זורק "ללא כיוון מאושר" — בלוק-הדיון דורש
decisions.direction_doc.approved שתיקי-הכיול חסרים.

scripts/calibrate_block_yod.py מוסיף שכבת אספקת-כיוון read-only מעל מכונת
המדידה של #208:
- overlay על db.get_decision_by_case/get_decision לתהליך-הזה-בלבד מזרים
  direction_doc שמיש ל-write_block — אפס כתיבה ל-DB, מצב-תיק-הייצור לא
  משתנה (reverted ב-__exit__).
- אסטרטגיה per-תיק: stored (כיוון-מאושר שמור, נאמן) → reconstructed
  (מינימלי מ-decision.outcome + סקשן-הדיון של הסופי באותו תיק, INV-LRN5:
  אפס מהות חוצת-תיקים, מתויג reconstructed) → skip.
- משתמש-חוזר ב-style_distance.block_distance_to_final + _score_cell/
  recommend_effort מ-calibrate_effort (G2 — אין מסלול-מדידה מקביל).
- כתיבה אינקרמנטלית per-תיק + שער-per-תא לא-פטאלי (INV-G8).

ממצא-קורפוס: מ-8 סופיים 7 ניתנים-לכיול; רק 8126-03-25 stored+scorable
(8174-12-24 שמור-אך-חסר-סקשן-דיון) → הנסיגה reconstructed היא מה שמפעיל
את ה-harness היום.

Invariants: G8 (eval-harness), G2 (שימוש-חוזר ב-style_distance/block_writer,
אין fork), INV-LRN5 (כיוון משוחזר = סגנון/שיטה בלבד, אפס מהות חוצת-תיקים),
claude_session local-only (ייצור דרך claude -p, מקומי-בלבד).

אימות: --self-test ALL PASS (28 checks: מדידה + אספקת-כיוון + overlay
install/revert) · --dry-run מול DB חי (read-only): 7/8 scorable, 28
ייצורים מתוכננים · --allow-reconstructed false → 1 (נאמן בלבד) ·
--case מסנן. אפס mutation calls ב-harness (read-only מאומת).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 18:50:01 +00:00

647 lines
32 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env python3
"""#212 (WS5 / Q1, INV-G8 eval-harness) — DEDICATED effort calibration for בלוק י'
(block-yod, the discussion/decision block — דיון והכרעה).
WHY A SEPARATE HARNESS (the #208 harness could NOT calibrate block-yod):
block-yod is the most important block — its effort is currently `xhigh`, set
"by feel" (block_writer.py:72), never empirically calibrated. The general #208
harness (scripts/calibrate_effort.py) EXCLUDES it from CALIBRATABLE_BLOCKS
because block_writer.write_block(case_id, "block-yod", …) raises
ValueError("לא ניתן לכתוב בלוק דיון ללא כיוון מאושר…")
— the discussion block requires an APPROVED DIRECTION (brainstorm →
approve_direction → decisions.direction_doc{"approved": True, …}) that the
calibration cases (the draft_final_pairs finals) do not have. So block-yod is
not calibratable standalone without first SUPPLYING a direction.
WHAT THIS HARNESS ADDS over #208 — the DIRECTION-SUPPLY layer (read-only):
Before regenerating block-yod for a calibration case, it makes a usable
`direction_doc` available to write_block via a READ-ONLY DB OVERLAY (it
monkeypatches db.get_decision_by_case / db.get_decision FOR THIS PROCESS ONLY
to overlay the direction into the returned decision dict). NOTHING is written
to the DB — production case state is never mutated (no UPDATE), satisfying the
"do not pollute real case state" constraint. The overlay is reverted in a
finally block.
DIRECTION-SUPPLY STRATEGY (per-case, documented in the report):
1. STORED (preferred, most faithful): if the case's decision already carries
a usable approved direction (direction_doc.approved is truthy and it has a
selected_direction, OR the 8126-style schema with direction_id/structure),
use it verbatim. This is the production-faithful path.
2. RECONSTRUCTED (fallback, directional-only): otherwise synthesize a MINIMAL
direction_doc from the case's decision.outcome + the final's OWN discussion
section (the chair's signed reasoning is the ground truth for "what
direction was taken"). Clearly LABELLED reconstructed in the report → the
resulting numbers are directional, not faithful to a real brainstorm.
INV-LRN5: the reconstructed direction carries the case's OWN outcome +
a neutral pointer to its OWN final — no substance dragged across cases.
EMPIRICAL NOTE (measured on the current corpus, 2026-06-30): of the 8 finals,
only 8174-12-24 has a STORED approved direction — and its final has NO parsable
discussion section (0 words), so it is not scorable for block-yod anyway. The 7
finals that DO have a scorable discussion section all LACK a usable stored
direction. Hence the RECONSTRUCTED fallback is what makes this harness useful
today; --dry-run prints the per-case supply mode so the operator sees this.
WHAT IT MEASURES (identical to #208, REUSED — no parallel metric path, G2):
Regenerate block-yod at each effort via block_writer.write_block(effort_override)
(the PRODUCTION generation path → claude_session.query → `claude -p`, pinned
Opus 4.8, local-only), then score vs the final's DISCUSSION section via
services.style_distance.block_distance_to_final (change_percent /
anti_pattern_total / golden_ratio_deviation_pp / composite distance). The
recommend_effort, aggregate_cell, _score_cell and direction-free scaffolding are
imported from calibrate_effort (#208) — this file only supplies the direction
and restricts the grid to block-yod.
GENERATION PATH (do not violate — reference_claude_generation_path /
claude_session docstring): write_block → claude_session.query → `claude -p` uses
the local claude.ai session. It runs ONLY on the host where the `claude` CLI
exists (NOT the legal-ai container, NOT a symlinked worktree without CLI access).
Hence the live A/B is HOST-ONLY; --self-test proves the measurement +
direction-supply logic offline with zero model calls.
⚠️ SAMPLE-SIZE CAVEAT (honored, not hidden): very few cases have an uploaded
final WITH a parsable discussion section. The report prints n_finals PROMINENTLY
and labels the output DIRECTIONAL EVIDENCE, not a regression. With n<3 the
recommendation is advisory only; the chair/operator decides whether to adopt.
Usage (mcp-server venv; live needs POSTGRES + the `claude` CLI on the host):
PY=/home/chaim/legal-ai/mcp-server/.venv/bin/python
$PY scripts/calibrate_block_yod.py --self-test # offline proof, no DB/CLI
$PY scripts/calibrate_block_yod.py --dry-run # plan grid + supply modes (needs DB, no model)
POSTGRES_PASSWORD=… POSTGRES_HOST=127.0.0.1 POSTGRES_PORT=5433 \
$PY scripts/calibrate_block_yod.py # live A/B over finals (host-only)
… --efforts low,medium,high,xhigh # override the effort grid
… --case 1130-08-25 # a single case
… --repeats 2 # avg N gens/cell (noise)
… --allow-reconstructed false # STORED-direction cases only (faithful)
"""
from __future__ import annotations
import argparse
import asyncio
import json
import logging
import os
import sys
from datetime import datetime
from pathlib import Path
from statistics import mean
logger = logging.getLogger(__name__)
REPO_ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(REPO_ROOT / "mcp-server" / "src"))
sys.path.insert(0, str(REPO_ROOT / "scripts"))
if "POSTGRES_URL" not in os.environ:
os.environ["POSTGRES_URL"] = (
f"postgres://{os.environ.get('POSTGRES_USER','legal_ai')}:"
f"{os.environ.get('POSTGRES_PASSWORD','')}@"
f"{os.environ.get('POSTGRES_HOST','127.0.0.1')}:"
f"{os.environ.get('POSTGRES_PORT','5433')}/"
f"{os.environ.get('POSTGRES_DB','legal_ai')}"
)
# Reuse the #208 machinery wholesale (G2 — no parallel harness/metric path).
# These are pure/offline-safe imports; calibrate_effort does no work at import.
from calibrate_effort import ( # noqa: E402
IL_TZ,
OUT_DIR,
VALID_EFFORTS,
_current_default,
_finals_for_calibration,
_score_cell,
recommend_effort,
)
BLOCK_ID = "block-yod"
SECTION = "discussion" # block-yod's golden-ratio section (style_distance._BLOCK_TO_SECTION)
DEFAULT_EFFORTS = ["low", "medium", "high", "xhigh"]
# ── direction supply (the block-yod-specific piece) ──────────────────────────
def _stored_direction_is_usable(direction_doc: dict | None) -> bool:
"""True if the case's stored direction_doc can drive write_block(block-yod).
write_block gates on `direction_doc.get("approved")` being truthy. Two stored
schemas exist in the corpus:
• brainstorm schema (build_direction_doc): {"approved": True, "selected_direction": …}
• chair-approved schema (8126-style): {"direction_id", "direction_name",
"structure", "template", …} — has NO "approved" key, so write_block would
reject it as-is. We treat it as usable IFF it carries a concrete direction
(direction_id/direction_name), and _normalize_stored_direction() stamps
approved=True onto the OVERLAY copy (never the DB).
Empty / brainstorm-only ({"brainstorm": …, "approved": False}) → not usable.
"""
dd = direction_doc or {}
if dd.get("approved") and (dd.get("selected_direction") or dd.get("reasoning")):
return True
if dd.get("direction_id") or dd.get("direction_name"):
return True
return False
def _normalize_stored_direction(direction_doc: dict) -> dict:
"""Return an overlay copy of a usable stored direction with approved=True.
Pure dict transform on a COPY — the real decisions.direction_doc is untouched.
The 8126-style schema lacks `approved`; stamping it here lets write_block's gate
pass without us mutating the DB.
"""
dd = dict(direction_doc)
dd["approved"] = True
return dd
def reconstruct_minimal_direction(outcome: str, final_discussion: str) -> dict:
"""Build a MINIMAL, clearly-reconstructed direction_doc (fallback path).
Faithful enough to make write_block(block-yod) runnable, honest about being a
reconstruction: the chair's SIGNED final discussion is the ground truth for
"what direction was taken", and the case's OWN outcome fixes the structure.
INV-LRN5: case's own outcome + a neutral pointer to its own final only — no
cross-case substance. Pure (no DB/LLM) → unit-tested in --self-test.
"""
from legal_mcp.services.lessons import OUTCOME_LABELS_HE, canonical_outcome
oc = canonical_outcome(outcome or "rejection")
snippet = " ".join((final_discussion or "").split()[:120])
return {
"approved": True,
"reconstructed": True, # marker so the report can label it
"outcome": oc,
"outcome_hebrew": OUTCOME_LABELS_HE.get(oc, oc),
"reasoning": (
"כיוון משוחזר לצורך כיול-effort בלבד (לא סיעור-מוחות אמיתי): "
"התוצאה והכיוון נגזרים מההחלטה החתומה של היו\"ר בתיק זה. "
"תמצית פתיח-הדיון של הסופי: " + snippet
),
"selected_direction": {
"name": f"כיוון משוחזר — {OUTCOME_LABELS_HE.get(oc, oc)}",
"reasoning": [
"שחזר את שלד-הנימוק מהדיון החתום של היו\"ר בתיק זה.",
],
"precedents": [],
},
"additional_notes": (
"⚠️ כיוון משוחזר — תוצאת-הכיול היא עדות-כיוון, לא נאמנה למסמך-כיוון אמיתי."
),
}
def plan_direction_supply(
decision: dict | None,
outcome: str,
final_discussion: str,
allow_reconstructed: bool,
) -> tuple[dict | None, str]:
"""Decide how to supply block-yod's direction for one case.
Returns (overlay_direction_doc | None, mode) where mode ∈
{"stored", "reconstructed", "skip"}. Pure (no I/O) → unit-tested.
"""
stored = (decision or {}).get("direction_doc")
if _stored_direction_is_usable(stored):
return _normalize_stored_direction(stored), "stored"
if allow_reconstructed and (final_discussion or "").strip():
return reconstruct_minimal_direction(outcome, final_discussion), "reconstructed"
return None, "skip"
class _DirectionOverlay:
"""Read-only overlay: make db.get_decision_by_case / db.get_decision return the
case's decision WITH `direction_doc` replaced by the supplied overlay — for the
target case_id only, for the duration of one cell. Reverted in __exit__.
write_block re-fetches the decision internally (block_writer.py:383), so passing
a direction dict in isn't enough — the DB read must yield it. We patch the read,
not the write: ZERO DB mutation (no UPDATE ever issued).
"""
def __init__(self, case_id, direction_doc: dict):
self._case_id = str(case_id)
self._dir = direction_doc
self._orig_by_case = None
self._orig_by_id = None
def __enter__(self):
from legal_mcp.services import db
self._orig_by_case = db.get_decision_by_case
self._orig_by_id = db.get_decision
async def patched_by_case(case_id, _orig=self._orig_by_case):
dec = await _orig(case_id)
if dec and str(case_id) == self._case_id:
dec = dict(dec)
dec["direction_doc"] = self._dir
return dec
async def patched_by_id(decision_id, _orig=self._orig_by_id):
dec = await _orig(decision_id)
if dec and dec.get("case_id") == self._case_id:
dec = dict(dec)
dec["direction_doc"] = self._dir
return dec
db.get_decision_by_case = patched_by_case
db.get_decision = patched_by_id
return self
def __exit__(self, *exc):
from legal_mcp.services import db
if self._orig_by_case is not None:
db.get_decision_by_case = self._orig_by_case
if self._orig_by_id is not None:
db.get_decision = self._orig_by_id
return False
async def _score_block_yod_cell(case_id, effort, final_section, final_total_words,
outcome, repeats, direction_doc):
"""Score one (case, block-yod, effort) cell with the direction overlaid.
Wraps the #208 _score_cell (REUSED) inside the read-only direction overlay so
write_block's approved-direction gate passes without DB mutation.
"""
with _DirectionOverlay(case_id, direction_doc):
return await _score_cell(
case_id, BLOCK_ID, effort, final_section, final_total_words, outcome, repeats,
)
# ── self-test (no DB, no model) ──────────────────────────────────────────────
def _self_test() -> int:
ok = True
def chk(name, cond):
nonlocal ok
ok = ok and cond
print(f" {name:50} {'ok' if cond else 'FAIL'}")
# measurement reuse: block_distance_to_final on the discussion section.
from legal_mcp.services.style_distance import block_distance_to_final, split_final_by_section
final_disc = "אנו סבורים כי דין הערר להידחות. כידוע, הלכה פסוקה היא. " * 30
d_same = block_distance_to_final(BLOCK_ID, final_disc, final_disc, "rejection",
section_target_total_words=len(final_disc.split()))
chk("identical regen ⇒ change_percent==0", d_same["change_percent"] == 0.0)
chk("identical regen ⇒ section==discussion", d_same["section"] == SECTION)
clean = "אין בידנו לקבל את הטענה. יחד עם זאת, מצאנו טעם. " * 20
dirty = "## כותרת\n- נקודה\n- נקודה\n* עוד\n### תת\n" * 10
d_clean = block_distance_to_final(BLOCK_ID, clean, final_disc, "rejection")
d_dirty = block_distance_to_final(BLOCK_ID, dirty, final_disc, "rejection")
chk("dirty regen ⇒ more anti-patterns", d_dirty["anti_pattern_total"] > d_clean["anti_pattern_total"])
chk("dirty regen ⇒ larger distance", d_dirty["distance"] > d_clean["distance"])
# ── direction-supply logic ──
chk("empty direction ⇒ not usable", not _stored_direction_is_usable({}))
chk("brainstorm-only ⇒ not usable",
not _stored_direction_is_usable({"brainstorm": {"directions": [1]}, "approved": False}))
chk("approved+selected ⇒ usable",
_stored_direction_is_usable({"approved": True, "selected_direction": {"name": "x"}}))
chk("8126-style (direction_id) ⇒ usable",
_stored_direction_is_usable({"direction_id": "d1", "structure": "..."}))
src = {"direction_id": "d1", "structure": "..."}
norm = _normalize_stored_direction(src)
chk("normalize stamps approved=True", norm["approved"] is True)
chk("normalize does not mutate source", "approved" not in src)
rec = reconstruct_minimal_direction("rejection", final_disc)
chk("reconstruct ⇒ approved", rec["approved"] is True)
chk("reconstruct ⇒ marked reconstructed", rec["reconstructed"] is True)
chk("reconstruct ⇒ canonical outcome", rec["outcome"] == "rejection")
chk("reconstruct ⇒ has selected_direction", bool(rec.get("selected_direction")))
chk("reconstruct ⇒ no cross-case precedents (INV-LRN5)",
rec["selected_direction"]["precedents"] == [])
# _build_direction_context must accept BOTH supply shapes (write_block uses it).
from legal_mcp.services.block_writer import _build_direction_context
ctx_recon = _build_direction_context({"direction_doc": rec})
chk("recon direction renders (not 'לא אושר')", "כיוון לא אושר" not in ctx_recon)
ctx_stored = _build_direction_context({"direction_doc": norm})
chk("stored(8126) direction renders", "כיוון לא אושר" not in ctx_stored)
# plan_direction_supply: stored wins; else reconstructed; else skip.
dd_stored = {"direction_doc": {"approved": True, "selected_direction": {"name": "x"}}}
_, m1 = plan_direction_supply(dd_stored, "rejection", final_disc, allow_reconstructed=True)
chk("plan: stored-direction case ⇒ stored", m1 == "stored")
_, m2 = plan_direction_supply({"direction_doc": {}}, "rejection", final_disc, True)
chk("plan: no stored + discussion ⇒ reconstructed", m2 == "reconstructed")
_, m3 = plan_direction_supply({"direction_doc": {}}, "rejection", final_disc, False)
chk("plan: no stored + reconstructed disabled ⇒ skip", m3 == "skip")
_, m4 = plan_direction_supply({"direction_doc": {}}, "rejection", "", True)
chk("plan: no stored + no discussion ⇒ skip", m4 == "skip")
# overlay monkeypatch installs + reverts cleanly (no live DB needed).
import legal_mcp.services.db as real_db
saved_bc, saved_bi = real_db.get_decision_by_case, real_db.get_decision
async def fake_by_case(cid):
return {"id": "d", "case_id": "C1", "direction_doc": {"approved": False}}
async def fake_by_id(did):
return {"id": "d", "case_id": "C1", "direction_doc": {"approved": False}}
real_db.get_decision_by_case = fake_by_case
real_db.get_decision = fake_by_id
try:
ov_dir = {"approved": True, "reconstructed": True}
async def _probe():
with _DirectionOverlay("C1", ov_dir):
d1 = await real_db.get_decision_by_case("C1")
d2 = await real_db.get_decision_by_case("OTHER")
d3 = await real_db.get_decision("d")
return d1, d2, d3
# --self-test is dispatched at the __main__ entry, OUTSIDE asyncio.run,
# so a plain asyncio.run here is safe (no enclosing event loop).
d1, d2, d3 = asyncio.run(_probe())
chk("overlay: target case gets overlaid direction", d1["direction_doc"]["approved"] is True)
chk("overlay: non-target case untouched", d2["direction_doc"]["approved"] is False)
chk("overlay: get_decision path overlaid", d3["direction_doc"]["reconstructed"] is True)
chk("overlay: reverted after context", real_db.get_decision_by_case is fake_by_case)
finally:
real_db.get_decision_by_case = saved_bc
real_db.get_decision = saved_bi
# reused #208 primitives still behave.
r = recommend_effort([
{"effort": "high", "distance": 0.22, "anti_pattern_total": 2, "change_percent": 25, "n": 1},
{"effort": "xhigh", "distance": 0.22, "anti_pattern_total": 1, "change_percent": 22, "n": 1},
])
chk("recommend tie → fewer anti-patterns", r["effort"] == "xhigh")
chk("current default for block-yod == xhigh", _current_default(BLOCK_ID) == "xhigh")
secs = split_final_by_section(
"רקע עובדתי\nהמקרקעין. " * 8 + "\n\nדיון והכרעה\nאנו סבורים. " * 8)
chk("split returns golden sections only",
set(secs).issubset({"background", "claims", "discussion", "summary"}))
print("ALL PASS" if ok else "*** FAILURES ***")
return 0 if ok else 1
# ── live A/B (host-only — needs DB + `claude` CLI) ───────────────────────────
async def _gather_cases(case_filter, allow_reconstructed):
"""Build per-case metadata + direction-supply plan over the finals."""
from uuid import UUID
from legal_mcp.services import db
from legal_mcp.services.lessons import canonical_outcome
from legal_mcp.services.style_distance import split_final_by_section
finals = await _finals_for_calibration(case_filter)
cases = []
for f in finals:
case = await db.get_case_by_number(f["case_number"]) if f.get("case_number") else None
if not case:
continue
decision = await db.get_decision_by_case(UUID(case["id"]))
outcome = canonical_outcome((decision or {}).get("outcome", "rejection"))
sections = split_final_by_section(f.get("final_text", ""))
final_discussion = sections.get(SECTION, "")
final_total_words = len((f.get("final_text", "") or "").split())
overlay, mode = plan_direction_supply(
decision, outcome, final_discussion, allow_reconstructed)
cases.append({
"case_number": f["case_number"],
"case_id": case["id"],
"outcome": outcome,
"final_discussion": final_discussion,
"final_discussion_words": len(final_discussion.split()),
"final_total_words": final_total_words,
"direction": overlay,
"supply_mode": mode,
# scorable IFF it has a discussion section AND a supplied direction.
"scorable": bool(final_discussion.strip() and overlay is not None),
})
return cases
async def _run(args, ts):
from uuid import UUID
cases = await _gather_cases(args.case, args.allow_reconstructed)
scorable = [c for c in cases if c["scorable"]]
grid_summary = {
"block": BLOCK_ID,
"section": SECTION,
"n_finals": len(cases),
"n_scorable": len(scorable),
"current_default": _current_default(BLOCK_ID),
"efforts": args.efforts,
"repeats": args.repeats,
"total_generations": len(scorable) * len(args.efforts) * args.repeats,
"allow_reconstructed": args.allow_reconstructed,
"cases": [
{"case_number": c["case_number"], "supply_mode": c["supply_mode"],
"scorable": c["scorable"], "outcome": c["outcome"],
"discussion_words": c["final_discussion_words"]}
for c in cases
],
}
if args.dry_run:
return {"dry_run": True, "grid": grid_summary, "by_effort": []}
per_effort_runs = {e: [] for e in args.efforts}
per_case = []
for c in scorable:
case_cells = []
for effort in args.efforts:
# Per-cell guard (INV-G8): any single (case, effort) failure — a
# transient rate-limit, a too-large prompt, an unexpected raise — is
# logged and skipped, never allowed to abort the whole run.
try:
cell = await _score_block_yod_cell(
UUID(c["case_id"]), effort, c["final_discussion"],
c["final_total_words"], c["outcome"], args.repeats, c["direction"],
)
except Exception as exc: # noqa: BLE001 — harness must survive any cell failure
logger.warning(
"block-yod cell skipped: case=%s effort=%s mode=%s%s",
c["case_number"], effort, c["supply_mode"], exc,
)
continue
per_effort_runs[effort].append(cell)
case_cells.append({k: cell[k] for k in
("effort", "distance", "anti_pattern_total",
"change_percent", "golden_ratio_deviation_pp", "n")})
per_case.append({"case_number": c["case_number"],
"supply_mode": c["supply_mode"], "cells": case_cells})
# Incremental persistence (INV-G8): flush after EACH case so a crash later
# never loses cases already generated. _write_report tolerates partial rows.
try:
_write_report({"dry_run": False, "grid": grid_summary,
"by_effort": _effort_rows(per_effort_runs, args.efforts),
"per_case": per_case,
"recommended": (recommend_effort(_effort_rows(per_effort_runs, args.efforts)) or {}).get("effort")},
ts)
except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run
logger.warning("incremental report write failed after case=%s%s",
c["case_number"], exc)
effort_rows = _effort_rows(per_effort_runs, args.efforts)
rec = recommend_effort(effort_rows)
return {"dry_run": False, "grid": grid_summary, "by_effort": effort_rows,
"per_case": per_case, "recommended": rec["effort"] if rec else None}
def _effort_rows(per_effort_runs, efforts):
"""Mean each metric across cases for each effort → one comparable row/effort."""
rows = []
for effort in efforts:
cells = per_effort_runs[effort]
if not cells:
continue
ratios = [c["golden_ratio_deviation_pp"] for c in cells
if c.get("golden_ratio_deviation_pp") is not None]
rows.append({
"effort": effort,
"distance": round(mean(c["distance"] for c in cells), 4),
"anti_pattern_total": round(mean(c["anti_pattern_total"] for c in cells), 2),
"change_percent": round(mean(c["change_percent"] for c in cells), 2),
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
"n": len(cells),
})
return rows
# ── report ───────────────────────────────────────────────────────────────────
def _ts():
return datetime.now(IL_TZ).strftime("%Y%m%dT%H%M%S-IL")
def _write_report(result, ts):
OUT_DIR.mkdir(parents=True, exist_ok=True)
jp = OUT_DIR / f"block-yod-calibration-{ts}.json"
mp = OUT_DIR / f"block-yod-calibration-{ts}.md"
jp.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
g = result["grid"]
il_now = datetime.now(IL_TZ).strftime("%Y-%m-%d %H:%M")
n_recon = sum(1 for c in g["cases"] if c["supply_mode"] == "reconstructed")
n_stored = sum(1 for c in g["cases"] if c["supply_mode"] == "stored")
n_skip = sum(1 for c in g["cases"] if c["supply_mode"] == "skip")
lines = [
f"# #212 — כיול effort לבלוק י' (דיון והכרעה) מול הסופיים — {ts}\n",
f"> נוצר: {il_now} (שעון ישראל · Asia/Jerusalem)\n",
f"> ⚠️ **גודל-מדגם: {g['n_scorable']} תיקים ניתנים-לכיול** "
f"(מתוך {g['n_finals']} סופיים). זוהי **עדות-כיוון, לא רגרסיה**. "
"ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n",
f"> אספקת-כיוון: {n_stored} מאוחסן · {n_recon} משוחזר · {n_skip} מדולג. "
"כיוון **משוחזר** = נגזר מתוצאת-ההחלטה + פתיח-הדיון של הסופי באותו תיק "
"(לא סיעור-מוחות אמיתי) → התוצאה נאמנה-לכיוון פחות מתיק עם כיוון-מאושר שמור.\n",
f"- בלוק: {g['block']} (section={g['section']}) · ברירת-מחדל נוכחית: "
f"**{g['current_default'] or ''}** (נקבע 'by feel', block_writer.py:72)",
f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}",
f"- סך ייצורי-מודל: {g['total_generations']}",
"",
"## אספקת-כיוון per-תיק\n",
"| case | supply_mode | discussion_words | outcome | scorable |",
"|---|---|---|---|---|",
]
for c in g["cases"]:
lines.append(
f"| {c['case_number']} | {c['supply_mode']} | {c['discussion_words']} | "
f"{c['outcome']} | {'' if c['scorable'] else ''} |")
lines.append("")
if result.get("dry_run"):
lines += ["## DRY-RUN — תכנון הגריד בלבד (ללא ייצור)\n",
f"- תיקים ניתנים-לכיול: **{g['n_scorable']}** · "
f"ייצורים מתוכננים: **{g['total_generations']}**",
""]
mp.write_text("\n".join(lines) + "\n", encoding="utf-8")
return jp, mp
rec = result.get("recommended") or ""
mark = "" if rec == g["current_default"] else " ⬅︎ שינוי מומלץ"
lines += [
f"## המלצה: effort = **{rec}**{mark} (current = {g['current_default'] or ''})\n",
"## טבלת effort (distance נמוך = קרוב יותר לדיון של דפנה)\n",
"| effort | distance | anti_total | change% | ratioΔpp | n |",
"|---|---|---|---|---|---|",
]
for r in result.get("by_effort", []):
star = "" if r["effort"] == result.get("recommended") else ""
ratio = r["golden_ratio_deviation_pp"]
lines.append(
f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | "
f"{r['change_percent']} | {ratio if ratio is not None else ''} | {r['n']} |")
lines.append("")
lines += ["## פירוט per-תיק\n", "| case | mode | effort | distance | anti | change% |",
"|---|---|---|---|---|---|"]
for pc in result.get("per_case", []):
for cell in pc["cells"]:
lines.append(
f"| {pc['case_number']} | {pc['supply_mode']} | {cell['effort']} | "
f"{cell['distance']:.4f} | {cell['anti_pattern_total']} | {cell['change_percent']} |")
lines.append("")
lines.append("> change% מערבב סגנון עם שלמות-תוכן (07-learning §0.7); "
"anti_total הוא הסיגנל הנקי-יותר לסגנון. תיקים עם supply_mode=reconstructed "
"תורמים עדות-כיוון בלבד.\n")
mp.write_text("\n".join(lines) + "\n", encoding="utf-8")
return jp, mp
async def main() -> int:
ap = argparse.ArgumentParser(description="#212 block-yod (discussion) effort calibration harness")
ap.add_argument("--self-test", action="store_true",
help="offline measurement + direction-supply proof (no DB/CLI)")
ap.add_argument("--dry-run", action="store_true",
help="plan the grid + per-case direction-supply mode, no model calls (needs DB)")
ap.add_argument("--efforts", default=",".join(DEFAULT_EFFORTS),
help=f"comma effort grid (default {','.join(DEFAULT_EFFORTS)})")
ap.add_argument("--case", default=None, help="restrict to a single case_number")
ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)")
ap.add_argument("--allow-reconstructed", default="true",
help="true|false — allow the reconstructed-direction fallback (default true). "
"false = STORED-direction cases only (most faithful, may yield 0 cells).")
args = ap.parse_args()
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
args.efforts = [e.strip() for e in args.efforts.split(",") if e.strip()]
bad = [e for e in args.efforts if e not in VALID_EFFORTS]
if bad:
print(f"invalid effort(s): {bad}. valid: {sorted(VALID_EFFORTS)}", file=sys.stderr)
return 2
args.allow_reconstructed = str(args.allow_reconstructed).strip().lower() in ("1", "true", "yes", "y")
ts = _ts()
result = await _run(args, ts)
jp, mp = _write_report(result, ts)
g = result["grid"]
print(f"BLOCK-YOD CALIBRATION: {g['n_scorable']}/{g['n_finals']} scorable finals — "
"DIRECTIONAL EVIDENCE, not a regression")
n_recon = sum(1 for c in g["cases"] if c["supply_mode"] == "reconstructed")
n_stored = sum(1 for c in g["cases"] if c["supply_mode"] == "stored")
print(f" direction supply: {n_stored} stored · {n_recon} reconstructed · "
f"{sum(1 for c in g['cases'] if c['supply_mode']=='skip')} skipped")
if g["n_scorable"] == 0:
print(" no scorable finals (need a parsable discussion section + a suppliable direction).")
elif result.get("dry_run"):
print(f" dry-run: {g['total_generations']} generations planned "
f"({len(g['efforts'])} efforts × {g['n_scorable']} cases × {g['repeats']} repeats)")
else:
print(f" current default={g['current_default']} → recommended={result.get('recommended')}")
print(f" report: {mp}")
return 0
if __name__ == "__main__":
# --self-test is OFFLINE and synchronous (it spins its own short-lived loop for
# the overlay probe), so dispatch it BEFORE asyncio.run(main()) — otherwise the
# probe's asyncio.run would nest inside main's running loop.
if "--self-test" in sys.argv:
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
sys.exit(_self_test())
sys.exit(asyncio.run(main()))