Compare commits
6 Commits
worktree-a
...
worktree-o
| Author | SHA1 | Date | |
|---|---|---|---|
| 86e66cc5bd | |||
| 10e05700cc | |||
| 11acdac337 | |||
| 4eb3312e9b | |||
| 574998021e | |||
| cc4d757fce |
@@ -820,45 +820,38 @@ ls data/cases/$CASE_NUMBER/documents/research/analysis-and-research.md
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
**תבנית issue לכותב ההחלטה — חובה בכל issue שמוקצה לכותב:**
|
**מסמך-ההכוונה לכותב — הפק את התדריך שהיית רוצה לקבל, לא טופס למילוי:**
|
||||||
|
|
||||||
כל issue לכותב חייב לכלול את **כל** הסעיפים הבאים. אסור לשלוח issue עם משפט כמו "הועבר לכתיבה" — זה חסר תועלת. הכותב צריך הכל מוכן מראש.
|
כשאתה מעביר תיק לכותב אתה מבצע את **פעולת-ההיסק המרכזית שלך**: להמיר את ניתוח-המנתח + הכרעות-היו"ר למסמך שמאפשר לכותב לנסח החלטה חדה בסגנון דפנה **בלי לחזור אליך**. אל תמלא טופס — הפעל שיפוט משפטי. תדריך טוב:
|
||||||
|
|
||||||
|
- **מוביל בהכרעה ובסוגיה המכריעה** — קבע איזו סוגיה נושאת את התוצאה ומה מייתר את מה, והצב אותה ראשונה.
|
||||||
|
- **בונה כל סוגיה כסילוגיזם** (כלל → עובדות → מסקנה) עם התקדים והמסמך הספציפיים.
|
||||||
|
- **מזהה אדנים עצמאיים** — אם יותר מנימוק אחד מספיק לבדו לתוצאה, אמור זאת מפורשות, כך שנפילת אדן בערעור לא תפיל את ההחלטה.
|
||||||
|
- **בודק עקביות פנימית** — אם שתי הכרעות עלולות להיראות סותרות (למשל דחיית טענה פרשנית אחת וקבלת אחרת), סמן את המתח והסבר את האבחנה לפני שעורך-דין יטען לו.
|
||||||
|
- **עונה לנקודה החזקה של הצד המפסיד** — לא מתעלם ממנה.
|
||||||
|
- **משקלל את הכרעות-היו"ר** ומעביר אותן מילולית.
|
||||||
|
|
||||||
|
**מה התדריך חייב להכיל** (החוזה מול הכותב — אל תשמיט אף רכיב; אל תשלח issue עם "הועבר לכתיבה"):
|
||||||
|
|
||||||
```markdown
|
```markdown
|
||||||
## הנחיות כתיבה — ערר {case_number}
|
## הנחיות כתיבה — ערר {case_number}
|
||||||
|
|
||||||
### 1. תוצאה ומצב
|
### 1. תוצאה ומצב
|
||||||
- **תוצאה:** {דחייה / קבלה חלקית / קבלה מלאה}
|
- **תוצאה:** {דחייה / קבלה חלקית / קבלה מלאה} — עם נימוק קצר ומהי הראיה הניצחת.
|
||||||
- **טיוטה קיימת:** {כן/לא}. אם כן: נתיב מלא לקובץ + הנחיה "קרא את הטיוטה, השתמש בה כבסיס, אל תכתוב מאפס"
|
- **טיוטה קיימת:** {כן/לא}. אם כן: נתיב מלא + "קרא, השתמש כבסיס, אל תכתוב מאפס".
|
||||||
- **הוראות עריכה מתוך הטיוטה:** {רשימה מדויקת של מה חיים ביקש לשנות — פסקאות, תוכן, placeholders}
|
- **הוראות עריכה מהטיוטה:** {מה חיים ביקש לשנות — פסקאות, תוכן, placeholders}.
|
||||||
|
|
||||||
### 2. סדר סוגיות + מבנה סילוגיסטי
|
### 2. סוגיות — סדר סילוגיסטי, המכריעה מובילה
|
||||||
לכל סוגיה שצריך לכתוב/לערוך — מבנה סילוגיסטי מלא:
|
לכל סוגיה: סוג-ניתוח (כלל ברור / איזון / מידתיות / שיקול-דעת) · כלל (ציטוט מדויק של הוראת-תכנית/חוק/הלכה) · עובדות (בהפניה למסמך-מקור ספציפי) · מסקנה · תקדימים (שם + מה קובע + רלוונטיות) · מסמכי-מקור (ב-data/cases/{case_number}/documents/originals/). סמן אדנים עצמאיים, מוקשי-עקביות ומענה לצד המפסיד היכן שהם קיימים.
|
||||||
|
|
||||||
**סוגיה N: {כותרת}**
|
|
||||||
- סוג ניתוח: {כלל ברור / איזון אינטרסים / מידתיות / שיקול דעת}
|
|
||||||
- כלל (הנחה עליונה): {הוראת תכנית / סעיף חוק / הלכה — ציטוט מדויק}
|
|
||||||
- עובדות (הנחה תחתונה): {העובדות הספציפיות שצריך להחיל — הפנייה למסמך מקור ספציפי}
|
|
||||||
- מסקנה: {מה נובע מהחלת הכלל על העובדות}
|
|
||||||
- תקדימים: {שם פסק דין + מה הוא קובע + למה רלוונטי}
|
|
||||||
- מסמכי מקור: {שמות קבצים ספציפיים ב-data/cases/{case_number}/documents/originals/}
|
|
||||||
|
|
||||||
### 3. טיפול בטענות
|
### 3. טיפול בטענות
|
||||||
| # | טענה | טיפול | סוגיה |
|
טבלה: # | טענה | טיפול (דיון מלא / קיבוץ / דילוג) | סוגיה.
|
||||||
|---|------|-------|-------|
|
|
||||||
| 1 | {טענה} | דיון מלא / קיבוץ / דילוג | {באיזו סוגיה} |
|
|
||||||
...
|
|
||||||
|
|
||||||
### 4. chair directions
|
### 4. הנחיות-היו"ר
|
||||||
- העתק מלא של עמדות הוועדה מ-analysis-and-research.md (או הפנייה: "קרא get_chair_directions").
|
העתק מילולי של עמדות-הוועדה מ-analysis-and-research.md (או "קרא get_chair_directions"), **עטוף ב-`<chair_directions>…</chair_directions>`** — טקסט מילולי בלבד בלי פרפרזה, כדי שהכותב לא ידרוס אותן.
|
||||||
- **עטוף את ההעתק המילולי בתגית `<chair_directions>…</chair_directions>`** — כך הכותב מבחין בין הוראות-היו"ר המילוליות לבין הערות ה-CEO, ואינו דורס אותן. בתוך התגית: טקסט מילולי בלבד, בלי פרפרזה.
|
|
||||||
|
|
||||||
### 5. הנחיות סגנון
|
### 5. הנחיות סגנון
|
||||||
- ניטרליות: בלוק ו = עובדות בלבד, בלי ציטוטים מצדדים
|
ניטרליות (בלוק ו = עובדות בלבד, בלי ציטוטי-צדדים) · ללא כפילות (בלוק י מפנה לקודמים) · טענות מקוריות (בלוק ז) · דיון ≥ 1,500 מילים · ≥ 3 תקדימים בדיון.
|
||||||
- ללא כפילות: בלוק י מפנה לבלוקים קודמים
|
|
||||||
- טענות מקוריות: בלוק ז = כתבי טענות מקוריים
|
|
||||||
- אורך מינימלי לדיון: 1,500 מילים לבלוק י
|
|
||||||
- פסיקה: חובה לצטט לפחות 3 תקדימים בדיון
|
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -369,6 +369,7 @@ async def write_block(
|
|||||||
block_id: str,
|
block_id: str,
|
||||||
instructions: str = "",
|
instructions: str = "",
|
||||||
effort_override: str | None = None,
|
effort_override: str | None = None,
|
||||||
|
model_override: str | None = None,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
"""כתיבת בלוק יחיד בהחלטה.
|
"""כתיבת בלוק יחיד בהחלטה.
|
||||||
|
|
||||||
@@ -381,6 +382,12 @@ async def write_block(
|
|||||||
THIS call only — used by the #208 model/effort calibration harness
|
THIS call only — used by the #208 model/effort calibration harness
|
||||||
to A/B efforts without mutating the pinned defaults. Production
|
to A/B efforts without mutating the pinned defaults. Production
|
||||||
callers leave it None and get the deterministic per-block effort.
|
callers leave it None and get the deterministic per-block effort.
|
||||||
|
model_override: optional per-call generation model id (e.g.
|
||||||
|
"claude-opus-5"). Same contract as effort_override — the #208
|
||||||
|
harness A/Bs MODELS without mutating the pinned GENERATION_MODEL.
|
||||||
|
Pass the BASE id only: the 1M-context escalation (#216) is applied
|
||||||
|
on top automatically for large prompts, so an override never
|
||||||
|
silently loses the 1M window. Production callers leave it None.
|
||||||
|
|
||||||
Returns:
|
Returns:
|
||||||
dict עם content, word_count, block_id, generation_type
|
dict עם content, word_count, block_id, generation_type
|
||||||
@@ -478,7 +485,12 @@ async def write_block(
|
|||||||
# escalate to the 1M-context build (`[1m]`) instead of failing the block —
|
# escalate to the 1M-context build (`[1m]`) instead of failing the block —
|
||||||
# block-yod legitimately carries the whole case as source-context. The 400K
|
# block-yod legitimately carries the whole case as source-context. The 400K
|
||||||
# ceiling was an artifact of the old 200K-only build, NOT a model limit.
|
# ceiling was an artifact of the old 200K-only build, NOT a model limit.
|
||||||
gen_model = GENERATION_MODEL_1M if len(prompt) > _CTX_1M_THRESHOLD_CHARS else GENERATION_MODEL
|
# model_override (#208 harness) swaps the BASE id only — the 1M decision below
|
||||||
|
# still applies, so an A/B'd model keeps the same context-window behaviour as
|
||||||
|
# the pinned default instead of silently falling back to the 200K build.
|
||||||
|
_base_model = model_override or GENERATION_MODEL
|
||||||
|
_model_1m = GENERATION_MODEL_1M if _base_model == GENERATION_MODEL else f"{_base_model}[1m]"
|
||||||
|
gen_model = _model_1m if len(prompt) > _CTX_1M_THRESHOLD_CHARS else _base_model
|
||||||
|
|
||||||
# Final guard: even the 1M build is finite (~2M Hebrew chars of input). Cap at
|
# Final guard: even the 1M build is finite (~2M Hebrew chars of input). Cap at
|
||||||
# 1.5M chars (~750K tokens) to leave room for output + a safety margin under 1M.
|
# 1.5M chars (~750K tokens) to leave room for output + a safety margin under 1M.
|
||||||
|
|||||||
@@ -176,7 +176,13 @@ def block_distance_to_final(
|
|||||||
outcome = canonical_outcome(outcome)
|
outcome = canonical_outcome(outcome)
|
||||||
diff = compute_diff_stats(regenerated_text or "", final_section_text or "")
|
diff = compute_diff_stats(regenerated_text or "", final_section_text or "")
|
||||||
change_percent = diff["change_percent"]
|
change_percent = diff["change_percent"]
|
||||||
anti_total = count_anti_patterns(regenerated_text or "")["total"]
|
anti = count_anti_patterns(regenerated_text or "")
|
||||||
|
anti_total = anti["total"]
|
||||||
|
# Per-pattern breakdown, not just the total: a calibration run that only
|
||||||
|
# reports "anti=4" cannot tell you WHICH rule was broken, so it cannot say
|
||||||
|
# what to fix. (Diagnosing the 2026-07-28 model A/B needed exactly this and
|
||||||
|
# had to fall back on inference.)
|
||||||
|
anti_by_pattern = {name: h["count"] for name, h in anti["by_pattern"].items()}
|
||||||
|
|
||||||
section = _BLOCK_TO_SECTION.get(block_id)
|
section = _BLOCK_TO_SECTION.get(block_id)
|
||||||
regen_words = len((regenerated_text or "").split())
|
regen_words = len((regenerated_text or "").split())
|
||||||
@@ -205,6 +211,7 @@ def block_distance_to_final(
|
|||||||
"final_words": final_words,
|
"final_words": final_words,
|
||||||
"change_percent": change_percent,
|
"change_percent": change_percent,
|
||||||
"anti_pattern_total": anti_total,
|
"anti_pattern_total": anti_total,
|
||||||
|
"anti_by_pattern": anti_by_pattern,
|
||||||
"golden_ratio_deviation_pp": ratio_dev,
|
"golden_ratio_deviation_pp": ratio_dev,
|
||||||
"distance": distance,
|
"distance": distance,
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -166,17 +166,33 @@ def aggregate_cell(per_run: list[dict]) -> dict:
|
|||||||
"""Mean each metric across repeated generations of the same (case, block, effort)."""
|
"""Mean each metric across repeated generations of the same (case, block, effort)."""
|
||||||
if not per_run:
|
if not per_run:
|
||||||
return {"distance": 1.0, "anti_pattern_total": 0.0, "change_percent": 100.0,
|
return {"distance": 1.0, "anti_pattern_total": 0.0, "change_percent": 100.0,
|
||||||
"golden_ratio_deviation_pp": None, "n": 0}
|
"golden_ratio_deviation_pp": None, "anti_by_pattern": {}, "n": 0}
|
||||||
ratios = [r["golden_ratio_deviation_pp"] for r in per_run if r.get("golden_ratio_deviation_pp") is not None]
|
ratios = [r["golden_ratio_deviation_pp"] for r in per_run if r.get("golden_ratio_deviation_pp") is not None]
|
||||||
return {
|
return {
|
||||||
"distance": round(mean(r["distance"] for r in per_run), 4),
|
"distance": round(mean(r["distance"] for r in per_run), 4),
|
||||||
"anti_pattern_total": round(mean(r["anti_pattern_total"] for r in per_run), 2),
|
"anti_pattern_total": round(mean(r["anti_pattern_total"] for r in per_run), 2),
|
||||||
"change_percent": round(mean(r["change_percent"] for r in per_run), 2),
|
"change_percent": round(mean(r["change_percent"] for r in per_run), 2),
|
||||||
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
|
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
|
||||||
|
"anti_by_pattern": _mean_by_pattern(per_run),
|
||||||
"n": len(per_run),
|
"n": len(per_run),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _mean_by_pattern(per_run: list[dict]) -> dict:
|
||||||
|
"""Mean hits PER anti-pattern name across runs — the 'which rule broke' view.
|
||||||
|
|
||||||
|
A pattern absent from a run counts as 0 (count_anti_patterns omits zero-hit
|
||||||
|
keys), so the mean is over ALL runs, not only the ones that tripped it.
|
||||||
|
"""
|
||||||
|
names: set[str] = set()
|
||||||
|
for r in per_run:
|
||||||
|
names |= set((r.get("anti_by_pattern") or {}).keys())
|
||||||
|
return {
|
||||||
|
name: round(mean((r.get("anti_by_pattern") or {}).get(name, 0) for r in per_run), 2)
|
||||||
|
for name in sorted(names)
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def _current_default(block_id: str) -> str | None:
|
def _current_default(block_id: str) -> str | None:
|
||||||
from legal_mcp.services.block_writer import BLOCK_CONFIG, DEFAULT_EFFORT
|
from legal_mcp.services.block_writer import BLOCK_CONFIG, DEFAULT_EFFORT
|
||||||
cfg = BLOCK_CONFIG.get(block_id, {})
|
cfg = BLOCK_CONFIG.get(block_id, {})
|
||||||
@@ -420,13 +436,30 @@ async def _finals_for_calibration(case_filter: str | None) -> list[dict]:
|
|||||||
|
|
||||||
|
|
||||||
async def _score_cell(case_id, block_id: str, effort: str, final_section: str,
|
async def _score_cell(case_id, block_id: str, effort: str, final_section: str,
|
||||||
final_total_words: int, outcome: str, repeats: int) -> dict:
|
final_total_words: int, outcome: str, repeats: int,
|
||||||
"""Generate `block_id` at `effort` `repeats` times; score each vs the final section."""
|
model: str | None = None, instructions: str = "") -> dict:
|
||||||
|
"""Generate `block_id` at `effort` `repeats` times; score each vs the final section.
|
||||||
|
|
||||||
|
`model` (optional) A/Bs the generation model via write_block(model_override=…).
|
||||||
|
None ⇒ the pinned GENERATION_MODEL, i.e. the production path unchanged.
|
||||||
|
|
||||||
|
`instructions` (optional) is appended to the block prompt for EVERY cell in
|
||||||
|
the run — a prompt-variant A/B (e.g. an explicit formatting rule). It is
|
||||||
|
applied to all models so the comparison stays a model comparison rather
|
||||||
|
than silently becoming a prompt comparison.
|
||||||
|
"""
|
||||||
from legal_mcp.services import block_writer
|
from legal_mcp.services import block_writer
|
||||||
from legal_mcp.services.style_distance import block_distance_to_final
|
from legal_mcp.services.style_distance import block_distance_to_final
|
||||||
runs: list[dict] = []
|
runs: list[dict] = []
|
||||||
|
models_used: list[str] = []
|
||||||
for _ in range(repeats):
|
for _ in range(repeats):
|
||||||
res = await block_writer.write_block(case_id, block_id, effort_override=effort)
|
res = await block_writer.write_block(
|
||||||
|
case_id, block_id, instructions=instructions,
|
||||||
|
effort_override=effort, model_override=model,
|
||||||
|
)
|
||||||
|
# Record what the CLI was actually asked to run, so a silent fallback to
|
||||||
|
# a different build is visible in the report rather than mis-attributed.
|
||||||
|
models_used.append(res.get("model_used") or "?")
|
||||||
scored = block_distance_to_final(
|
scored = block_distance_to_final(
|
||||||
block_id, res.get("content", ""), final_section, outcome,
|
block_id, res.get("content", ""), final_section, outcome,
|
||||||
section_target_total_words=final_total_words,
|
section_target_total_words=final_total_words,
|
||||||
@@ -434,6 +467,8 @@ async def _score_cell(case_id, block_id: str, effort: str, final_section: str,
|
|||||||
runs.append(scored)
|
runs.append(scored)
|
||||||
agg = aggregate_cell(runs)
|
agg = aggregate_cell(runs)
|
||||||
agg["effort"] = effort
|
agg["effort"] = effort
|
||||||
|
agg["model"] = model
|
||||||
|
agg["models_used"] = sorted(set(models_used))
|
||||||
agg["runs"] = runs
|
agg["runs"] = runs
|
||||||
return agg
|
return agg
|
||||||
|
|
||||||
@@ -446,6 +481,7 @@ async def _run(args, ts: str) -> dict:
|
|||||||
|
|
||||||
efforts = args.efforts
|
efforts = args.efforts
|
||||||
blocks = args.blocks
|
blocks = args.blocks
|
||||||
|
models = args.models
|
||||||
finals = await _finals_for_calibration(args.case)
|
finals = await _finals_for_calibration(args.case)
|
||||||
|
|
||||||
cases_meta = []
|
cases_meta = []
|
||||||
@@ -468,11 +504,15 @@ async def _run(args, ts: str) -> dict:
|
|||||||
section = _BLOCK_TO_SECTION.get(block_id)
|
section = _BLOCK_TO_SECTION.get(block_id)
|
||||||
plan[block_id] = [c for c in cases_meta if section and c["sections"].get(section)]
|
plan[block_id] = [c for c in cases_meta if section and c["sections"].get(section)]
|
||||||
|
|
||||||
total_cells = sum(len(plan[b]) for b in blocks) * len(efforts) * args.repeats
|
total_cells = sum(len(plan[b]) for b in blocks) * len(efforts) * args.repeats * len(models)
|
||||||
grid_summary = {
|
grid_summary = {
|
||||||
"n_finals": len(cases_meta),
|
"n_finals": len(cases_meta),
|
||||||
"finals": [c["case_number"] for c in cases_meta],
|
"finals": [c["case_number"] for c in cases_meta],
|
||||||
"blocks": blocks, "efforts": efforts, "repeats": args.repeats,
|
"blocks": blocks, "efforts": efforts, "repeats": args.repeats,
|
||||||
|
"models": models,
|
||||||
|
# Provenance: a prompt-variant run is NOT comparable to a baseline run,
|
||||||
|
# so the instruction text is recorded in the report, not just the shell.
|
||||||
|
"instructions": getattr(args, "instructions", "") or "",
|
||||||
"total_generations": total_cells,
|
"total_generations": total_cells,
|
||||||
"per_block_n": {b: len(plan[b]) for b in blocks},
|
"per_block_n": {b: len(plan[b]) for b in blocks},
|
||||||
}
|
}
|
||||||
@@ -480,6 +520,27 @@ async def _run(args, ts: str) -> dict:
|
|||||||
if args.dry_run:
|
if args.dry_run:
|
||||||
return {"dry_run": True, "grid": grid_summary, "by_block": {}}
|
return {"dry_run": True, "grid": grid_summary, "by_block": {}}
|
||||||
|
|
||||||
|
by_model: dict[str, dict] = {}
|
||||||
|
for model in models:
|
||||||
|
by_block = await _run_blocks_for_model(
|
||||||
|
model, blocks, efforts, plan, args, ts, grid_summary, by_model, _BLOCK_TO_SECTION,
|
||||||
|
)
|
||||||
|
by_model[model] = by_block
|
||||||
|
|
||||||
|
# `by_block` stays the single-model shape (first model) so --rerank and the
|
||||||
|
# existing per-block report path keep working unchanged (G2 — no second
|
||||||
|
# result schema); multi-model runs additionally carry by_model.
|
||||||
|
out = {"dry_run": False, "grid": grid_summary, "by_block": by_model[models[0]]}
|
||||||
|
if len(models) > 1:
|
||||||
|
out["by_model"] = by_model
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
async def _run_blocks_for_model(model, blocks, efforts, plan, args, ts, grid_summary,
|
||||||
|
by_model_so_far, _BLOCK_TO_SECTION) -> dict:
|
||||||
|
"""The per-block × per-effort grid for ONE generation model."""
|
||||||
|
from uuid import UUID
|
||||||
|
|
||||||
by_block: dict[str, dict] = {}
|
by_block: dict[str, dict] = {}
|
||||||
for block_id in blocks:
|
for block_id in blocks:
|
||||||
section = _BLOCK_TO_SECTION.get(block_id)
|
section = _BLOCK_TO_SECTION.get(block_id)
|
||||||
@@ -497,11 +558,12 @@ async def _run(args, ts: str) -> dict:
|
|||||||
cell = await _score_cell(
|
cell = await _score_cell(
|
||||||
UUID(c["case_id"]), block_id, effort, final_section,
|
UUID(c["case_id"]), block_id, effort, final_section,
|
||||||
c["final_total_words"], c["outcome"], args.repeats,
|
c["final_total_words"], c["outcome"], args.repeats,
|
||||||
|
model=model, instructions=getattr(args, "instructions", "") or "",
|
||||||
)
|
)
|
||||||
except Exception as exc: # noqa: BLE001 — harness must survive any cell failure
|
except Exception as exc: # noqa: BLE001 — harness must survive any cell failure
|
||||||
logger.warning(
|
logger.warning(
|
||||||
"calibration cell skipped: case=%s block=%s effort=%s — %s",
|
"calibration cell skipped: case=%s block=%s effort=%s model=%s — %s",
|
||||||
c["case_number"], block_id, effort, exc,
|
c["case_number"], block_id, effort, model, exc,
|
||||||
)
|
)
|
||||||
continue
|
continue
|
||||||
per_effort_runs[effort].append(cell)
|
per_effort_runs[effort].append(cell)
|
||||||
@@ -523,6 +585,7 @@ async def _run(args, ts: str) -> dict:
|
|||||||
"anti_pattern_total": round(mean(r["anti_pattern_total"] for r in rows), 2),
|
"anti_pattern_total": round(mean(r["anti_pattern_total"] for r in rows), 2),
|
||||||
"change_percent": round(mean(r["change_percent"] for r in rows), 2),
|
"change_percent": round(mean(r["change_percent"] for r in rows), 2),
|
||||||
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
|
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
|
||||||
|
"anti_by_pattern": _mean_by_pattern(rows),
|
||||||
"n": len(rows),
|
"n": len(rows),
|
||||||
})
|
})
|
||||||
rec = recommend_effort(effort_rows)
|
rec = recommend_effort(effort_rows)
|
||||||
@@ -532,6 +595,11 @@ async def _run(args, ts: str) -> dict:
|
|||||||
"recommended": rec["effort"] if rec else None,
|
"recommended": rec["effort"] if rec else None,
|
||||||
"confidence": rec["confidence"] if rec else None,
|
"confidence": rec["confidence"] if rec else None,
|
||||||
"confidence_margin": rec.get("confidence_margin") if rec else None,
|
"confidence_margin": rec.get("confidence_margin") if rec else None,
|
||||||
|
"model": model,
|
||||||
|
# Model builds the CLI actually reported across this block's cells —
|
||||||
|
# a mismatch vs `model` means a silent fallback, not a real A/B.
|
||||||
|
"models_used": sorted({m for e in per_effort_runs.values()
|
||||||
|
for cell in e for m in cell.get("models_used", [])}),
|
||||||
"efforts": effort_rows,
|
"efforts": effort_rows,
|
||||||
"per_case": per_case,
|
"per_case": per_case,
|
||||||
}
|
}
|
||||||
@@ -541,11 +609,15 @@ async def _run(args, ts: str) -> dict:
|
|||||||
# Blocks not yet done are simply absent from by_block; _write_report tolerates
|
# Blocks not yet done are simply absent from by_block; _write_report tolerates
|
||||||
# partial results. main() does the final flush once the loop finishes.
|
# partial results. main() does the final flush once the loop finishes.
|
||||||
try:
|
try:
|
||||||
_write_report({"dry_run": False, "grid": grid_summary, "by_block": by_block}, ts)
|
snap = {"dry_run": False, "grid": grid_summary, "by_block": by_block}
|
||||||
|
if by_model_so_far or len(grid_summary.get("models", [])) > 1:
|
||||||
|
snap["by_model"] = {**by_model_so_far, model: by_block}
|
||||||
|
_write_report(snap, ts)
|
||||||
except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run
|
except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run
|
||||||
logger.warning("incremental report write failed after block=%s — %s", block_id, exc)
|
logger.warning("incremental report write failed after block=%s model=%s — %s",
|
||||||
|
block_id, model, exc)
|
||||||
|
|
||||||
return {"dry_run": False, "grid": grid_summary, "by_block": by_block}
|
return by_block
|
||||||
|
|
||||||
|
|
||||||
IL_TZ = ZoneInfo("Asia/Jerusalem")
|
IL_TZ = ZoneInfo("Asia/Jerusalem")
|
||||||
@@ -578,7 +650,10 @@ def _write_report(result: dict, ts: str) -> tuple[Path, Path]:
|
|||||||
"ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n",
|
"ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n",
|
||||||
f"- בלוקים: {', '.join(g['blocks'])}",
|
f"- בלוקים: {', '.join(g['blocks'])}",
|
||||||
f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}",
|
f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}",
|
||||||
|
f"- models: {', '.join(m or 'pinned-default' for m in g.get('models', [None]))}",
|
||||||
f"- סך ייצורי-מודל: {g['total_generations']}",
|
f"- סך ייצורי-מודל: {g['total_generations']}",
|
||||||
|
(f"- ⚠️ **וריאנט-פרומפט** (לא בר-השוואה לריצת-בסיס): `{g['instructions']}`"
|
||||||
|
if g.get("instructions") else "- וריאנט-פרומפט: — (פרומפט ייצור כפי-שהוא)"),
|
||||||
"",
|
"",
|
||||||
]
|
]
|
||||||
if result.get("dry_run"):
|
if result.get("dry_run"):
|
||||||
@@ -613,6 +688,54 @@ def _write_report(result: dict, ts: str) -> tuple[Path, Path]:
|
|||||||
f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | "
|
f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | "
|
||||||
f"{r['change_percent']} | {ratio if ratio is not None else '—'} | {r['n']} |")
|
f"{r['change_percent']} | {ratio if ratio is not None else '—'} | {r['n']} |")
|
||||||
lines.append("")
|
lines.append("")
|
||||||
|
by_model = result.get("by_model") or {}
|
||||||
|
if len(by_model) > 1:
|
||||||
|
lines += ["## השוואת-מודלים (אותו block, אותו effort, אותם סופיים)\n",
|
||||||
|
"| block | effort | model | anti_total | change% | ratioΔpp | distance | n |",
|
||||||
|
"|---|---|---|---|---|---|---|---|"]
|
||||||
|
for b in g["blocks"]:
|
||||||
|
for eff in g["efforts"]:
|
||||||
|
rows = []
|
||||||
|
for m, bb in by_model.items():
|
||||||
|
for r in (bb.get(b) or {}).get("efforts", []):
|
||||||
|
if r["effort"] == eff:
|
||||||
|
rows.append((m, r))
|
||||||
|
if len(rows) < 2:
|
||||||
|
continue # nothing to compare for this cell — don't fake a row
|
||||||
|
best = min(rows, key=lambda mr: (mr[1]["anti_pattern_total"],
|
||||||
|
mr[1]["golden_ratio_deviation_pp"] or 0,
|
||||||
|
mr[1]["distance"]))[0]
|
||||||
|
for m, r in rows:
|
||||||
|
ratio = r["golden_ratio_deviation_pp"]
|
||||||
|
star = " ⭐" if m == best else ""
|
||||||
|
lines.append(
|
||||||
|
f"| {b} | {eff} | {m}{star} | {r['anti_pattern_total']} | "
|
||||||
|
f"{r['change_percent']} | {ratio if ratio is not None else '—'} | "
|
||||||
|
f"{r['distance']:.4f} | {r['n']} |")
|
||||||
|
lines.append("")
|
||||||
|
|
||||||
|
# WHICH rule broke — a total alone can't tell you what to fix.
|
||||||
|
bd_rows = [(b, eff, m, r) for b in g["blocks"] for eff in g["efforts"]
|
||||||
|
for m, bb in by_model.items()
|
||||||
|
for r in (bb.get(b) or {}).get("efforts", []) if r["effort"] == eff]
|
||||||
|
if any(r.get("anti_by_pattern") for *_, r in bd_rows):
|
||||||
|
names = sorted({n for *_, r in bd_rows for n in (r.get("anti_by_pattern") or {})})
|
||||||
|
lines += ["### פילוח אנטי-דפוסים (איזה כלל הופר)\n",
|
||||||
|
"| block | effort | model | " + " | ".join(names) + " |",
|
||||||
|
"|---|---|---|" + "---|" * len(names)]
|
||||||
|
for b, eff, m, r in bd_rows:
|
||||||
|
cells = " | ".join(str((r.get("anti_by_pattern") or {}).get(n, 0)) for n in names)
|
||||||
|
lines.append(f"| {b} | {eff} | {m} | {cells} |")
|
||||||
|
lines.append("")
|
||||||
|
# A silent CLI fallback would make the whole comparison meaningless — surface it.
|
||||||
|
for m, bb in by_model.items():
|
||||||
|
for b, bd in bb.items():
|
||||||
|
used = bd.get("models_used") or []
|
||||||
|
if used and any(not u.startswith(str(m)) for u in used):
|
||||||
|
lines.append(f"> ⚠️ **{b} / {m}**: ה-CLI דיווח `{', '.join(used)}` — "
|
||||||
|
"ייתכן fallback שקט; ההשוואה לתא זה אינה תקפה.\n")
|
||||||
|
lines.append("")
|
||||||
|
|
||||||
lines.append("> דירוג-ההמלצה **style-clean** (#213): anti_total ראשי → ratioΔ → distance (tiebreak). "
|
lines.append("> דירוג-ההמלצה **style-clean** (#213): anti_total ראשי → ratioΔ → distance (tiebreak). "
|
||||||
"**change% מדווח-לא-מדורג** — מערבב סגנון עם שלמות-תוכן (07-learning §0.7), "
|
"**change% מדווח-לא-מדורג** — מערבב סגנון עם שלמות-תוכן (07-learning §0.7), "
|
||||||
"anti_total הוא הסיגנל הנקי-לסגנון. confidence=⚠️weak ⇒ הבחירה בתוך-הרעש "
|
"anti_total הוא הסיגנל הנקי-לסגנון. confidence=⚠️weak ⇒ הבחירה בתוך-הרעש "
|
||||||
@@ -633,6 +756,12 @@ async def main() -> int:
|
|||||||
help="comma block ids to calibrate")
|
help="comma block ids to calibrate")
|
||||||
ap.add_argument("--case", default=None, help="restrict to a single case_number")
|
ap.add_argument("--case", default=None, help="restrict to a single case_number")
|
||||||
ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)")
|
ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)")
|
||||||
|
ap.add_argument("--models", default="",
|
||||||
|
help="comma generation-model ids to A/B (e.g. claude-opus-4-8,claude-opus-5). "
|
||||||
|
"Empty (default) = the pinned GENERATION_MODEL, i.e. production unchanged.")
|
||||||
|
ap.add_argument("--instructions", default="",
|
||||||
|
help="extra prompt instruction appended to EVERY cell (prompt-variant A/B). "
|
||||||
|
"Applied to all models — the run stays a model comparison. Recorded in the report.")
|
||||||
args = ap.parse_args()
|
args = ap.parse_args()
|
||||||
|
|
||||||
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
||||||
@@ -653,6 +782,9 @@ async def main() -> int:
|
|||||||
if bad_b:
|
if bad_b:
|
||||||
print(f"non-calibratable block(s): {bad_b}. valid: {VALID_BLOCKS}", file=sys.stderr)
|
print(f"non-calibratable block(s): {bad_b}. valid: {VALID_BLOCKS}", file=sys.stderr)
|
||||||
return 2
|
return 2
|
||||||
|
# [None] = "use the pinned GENERATION_MODEL" — keeps the default run byte-identical
|
||||||
|
# to the pre-#models behaviour instead of hard-coding the id in a second place (G2).
|
||||||
|
args.models = [m.strip() for m in args.models.split(",") if m.strip()] or [None]
|
||||||
|
|
||||||
ts = _ts()
|
ts = _ts()
|
||||||
result = await _run(args, ts)
|
result = await _run(args, ts)
|
||||||
|
|||||||
Reference in New Issue
Block a user