feat(calibration): כיול-אמפירי model×effort מול הסופיים (#208)
A/B harness שמכייל את ה-effort הנעוץ per-בלוק (#204) מול הסופיים של דפנה: מייצר מחדש כל בלוק דרך מסלול-הייצור (write_block(effort_override=…) → claude_session.query → claude -p, Opus 4.8, מקומי-בלבד) ומודד מול הסקשן המתאים בסופי דרך style_distance.block_distance_to_final (change_percent, anti_pattern_total, golden-ratio deviation, composite distance). ממליץ per-בלוק על ה-effort הקרוב-ביותר לסופי. - scripts/calibrate_effort.py — ההארנס (מודל eval_retrieval.py): --self-test (offline, מוכיח מדידה+המלצה, אפס DB/CLI) · --dry-run · --efforts/--blocks/ --case/--repeats. דוח data/eval/effort-calibration-<ts>.{json,md} עם גודל-מדגם בולט — עדות-כיוון, לא רגרסיה (מעט סופיים-עלויים). - style_distance.py — block_distance_to_final + split_final_by_section (מקור-מדידה יחיד, G2; reuse compute_diff_stats/count_anti_patterns/chunker). - block_writer.py — write_block(effort_override=) להזרקת effort per-קריאה בלי לדרוס את ברירות-המחדל הנעוצות; רושם את ה-effort האפקטיבי. - scripts/SCRIPTS.md — ערך חדש. Invariants: G8 (eval-harness — מדידה אמפירית, לא הנחה) · G2 (reuse של style_distance/learning_loop — אין מסלול-מדד מקביל) · claude_session local-only (reference_claude_generation_path) · INV-LRN4/5 (השוואה מול הסופי, מדידת-סגנון; אין מהות-תיק נגררת). אומת: --self-test 14/14 PASS. הריצה החיה host-only (claude CLI) — לא ניתנת-להרצה ב-worktree/קונטיינר. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -355,6 +355,7 @@ async def write_block(
|
||||
case_id: UUID,
|
||||
block_id: str,
|
||||
instructions: str = "",
|
||||
effort_override: str | None = None,
|
||||
) -> dict:
|
||||
"""כתיבת בלוק יחיד בהחלטה.
|
||||
|
||||
@@ -362,6 +363,11 @@ async def write_block(
|
||||
case_id: מזהה התיק
|
||||
block_id: מזהה הבלוק (block-alef, block-he, block-yod, ...)
|
||||
instructions: הנחיות נוספות
|
||||
effort_override: optional per-call reasoning effort (low/medium/high/
|
||||
xhigh/max). When set, overrides BLOCK_CONFIG[block_id].effort for
|
||||
THIS call only — used by the #208 model/effort calibration harness
|
||||
to A/B efforts without mutating the pinned defaults. Production
|
||||
callers leave it None and get the deterministic per-block effort.
|
||||
|
||||
Returns:
|
||||
dict עם content, word_count, block_id, generation_type
|
||||
@@ -472,7 +478,7 @@ async def write_block(
|
||||
# reasoning effort so generation is structurally deterministic — these were
|
||||
# previously NOT forwarded (the source of inconsistency). model/effort flow
|
||||
# through claude_session.query → `claude -p --model … --effort …`.
|
||||
effort = block_cfg.get("effort", DEFAULT_EFFORT)
|
||||
effort = effort_override or block_cfg.get("effort", DEFAULT_EFFORT)
|
||||
timeout = claude_session.LONG_TIMEOUT if effort in _LONG_EFFORTS else claude_session.DEFAULT_TIMEOUT
|
||||
content = await claude_session.query(
|
||||
prompt,
|
||||
@@ -485,6 +491,10 @@ async def write_block(
|
||||
sources = await _collect_block_sources(case_id, block_id)
|
||||
sources["case_law_ids"] = _precedent_case_law_ids
|
||||
result = _build_result(block_id, content, block_cfg)
|
||||
# Record the EFFECTIVE effort (override wins) so the harness can attribute
|
||||
# the measured distance to the effort that actually produced the text.
|
||||
if result.get("effort") is not None:
|
||||
result["effort"] = effort
|
||||
result["sources"] = sources
|
||||
return result
|
||||
|
||||
|
||||
Reference in New Issue
Block a user