Compare commits

..

9 Commits

Author SHA1 Message Date
2ebaa82f85 Merge remote-tracking branch 'origin/main' into worktree-anti-pattern-directive-position
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
2026-07-28 11:53:23 +00:00
79d9ea55d8 Merge pull request 'feat(eval): ממד model×prompt ל-harness הכיול (#208) — A/B מודל-ייצור מול הסופיים' (#420) from worktree-opus5-model-calibration into main
Some checks failed
Build & Deploy / build-and-deploy (push) Has been cancelled
G12 Leak-Guard / leak-guard (push) Has been cancelled
Lint — undefined names / undefined-names (push) Has been cancelled
2026-07-28 11:52:57 +00:00
86e66cc5bd feat(eval): פילוח אנטי-דפוסים per-ריצה — "איזה כלל הופר", לא רק כמה
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
`block_distance_to_final` החזיר `anti_pattern_total` בלבד. ריצת-כיול שמדווחת
"anti=4" אינה יכולה לומר מה לתקן — באבחון ה-A/B של 2026-07-28 נאלצנו להסיק
את הדפוס האשם מהקשר במקום למדוד אותו.

- `anti_by_pattern` (שם-דפוס → מספר-פגיעות) נוסף לתא-המדידה, מ-
  `count_anti_patterns` הקיים — אין ספירה מקבילה.
- `_mean_by_pattern` ממצע על **כל** הריצות: דפוס שלא נורה בריצה נספר כ-0
  ולא מושמט, אחרת הממוצע היה מוטה כלפי מעלה.
- הדוח מקבל טבלת "פילוח אנטי-דפוסים (איזה כלל הופר)" per block×effort×model.

invariants: INV-G8 (eval-harness) · G2 (מרונדר מ-count_anti_patterns/
ANTI_PATTERNS הקנוניים — מקור אחד).

אימות: self-test ALL PASS · 470 passed · בדיקת-שפיות ישירה —
טקסט עם 1 כותרת + 2 תבליטים + 1 פיצול-מיני מפולח נכון ל-
{markdown_headers:1, bullet_lists:2, inline_numbered_fragments:1}.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:17:35 +00:00
42ea1a7c58 fix(writer): כלל-הסגנון בסוף הפרומפט — הוא היה שם, במקום שבו הוא לא תופס
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 35s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
הרשימה הקנונית של lessons.ANTI_PATTERNS כבר הוזרקה לכותב, אבל בתו ~46,781
מתוך 46,950 של style_context — שהוא עצמו מקטע אחד מתוך ~12 בפרומפט. הטיוטות
המשיכו לפלוט בדיוק את מה שהיא אוסרת.

A/B מדוד מול הסופיים החתומים (9 תיקים, 60 ייצורים, 2026-07-28) הראה שאותו
כלל, בסוף הפרומפט, חותך anti_pattern_total ב-72–93%:

  block-vav   opus-4-8 1.75→0.12 · opus-5 2.25→0.62
  block-zayin opus-4-8 4.57→0.43 · opus-5 4.43→0.43

וב-distance: −12%/−29% (4-8), −9%/−27% (5). זה שיפור גדול פי-3 מכל הבדל
שנמדד בין המודלים עצמם.

- `lessons.anti_pattern_directive()` — רינדור שני של אותה רשימה קנונית
  (מקור אחד, שתי תצוגות — לא שני כללים).
- מתווסף **אחרון** בשני מסלולי-הכתיבה: `write_block` (בתהליך) ו-
  `get_block_context` (סוכן legal-writer). אילו הוחל רק באחד, שני הכותבים
  היו נפרדים בסגנון (G2).
- **תיקון בליעה-שקטה (§6):** הרשימה הקנונית רונדרה בתוך לולאת-ה-overrides,
  כך שכשל-DB בקטגוריה מוקדמת (golden_ratios) הפיל את הלולאה והשמיט את
  אינווריאנטי-הסגנון כליל — עם אזהרה גנרית בלבד. עכשיו היא מרונדרת ללא
  תנאי, לפני כל קריאת-DB; הערות-היו"ר מתווספות מעליה.

invariants: G11 (תוכן משפטי — סגנון דפנה) · G2 (מקור-אמת יחיד לכלל, ושני
מסלולי-הכתיבה מיושרים) · §6 (אין בליעה שקטה).

בדיקות: 473 passed (3 חדשות — רינדור מלא, שני המסלולים, שרידות לכשל-DB).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:14:47 +00:00
10e05700cc feat(eval): --instructions ל-harness — A/B של וריאנט-פרומפט
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 34s
Lint — undefined names / undefined-names (pull_request) Successful in 12s
מוחל על כל המודלים בריצה (אחרת השוואת-מודלים הופכת בשקט להשוואת-פרומפטים),
ונרשם ב-grid_summary + בכותרת הדוח כדי שריצת-וריאנט לא תושווה בטעות לבסיס.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 07:47:05 +00:00
11acdac337 feat(eval): ממד-מודל ל-harness הכיול (#208) — A/B של מודל-ייצור מול הסופיים
`calibrate_effort.py` כייל `effort` בלבד, על מודל נעוץ (GENERATION_MODEL).
כדי להשוות מודל-ייצור (opus-4-8 מול opus-5) מול הסופיים החתומים של דפנה
נדרש ממד שני — ללא מסלול-מדידה מקביל.

- `block_writer.write_block(model_override=…)` — אותו חוזה כמו
  `effort_override` הקיים (נוצר בדיוק ל-#208). מקבל את המזהה הבסיסי בלבד;
  אסקלציית ההקשר-1M (#216) מוחלת מעליו, כך ש-override לא מאבד בשקט את
  חלון ה-1M.
- `--models` ל-harness (ריק = המודל הנעוץ ⇒ ריצת ברירת-המחדל זהה לקודם).
- הדוח מקבל טבלת השוואת-מודלים (block × effort × model) ומסמן  לפי
  אותו דירוג style-clean (#213): anti_total → ratioΔ → distance.
- כל תא מתעד את `model_used` שה-CLI דיווח בפועל; אי-התאמה מסומנת כאזהרת
  fallback-שקט במקום להיזקף בטעות למודל המבוקש.

invariants: INV-G8 (eval-harness — מדידה, לא הרגשה) · G2 (אין מסלול מקביל:
משתמש ב-style_distance/learning_loop הקיימים ובשדה result הקיים
`model_used`) · §6 (אין בליעה שקטה — כשל-תא מדווח ומדולג).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 06:39:33 +00:00
4eb3312e9b Merge pull request 'refactor(legal-ceo): גיזום מסמך-ההכוונה לכותב — שיפוט במקום טופס' (#419) from ceo-prompt-trim into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 10s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-07-25 21:29:53 +00:00
574998021e refactor(legal-ceo): גיזום מסמך-ההכוונה לכותב — שיפוט במקום טופס
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 35s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
מיישם את ממצא ה-A/B (fable-5/opus-4-8 על תיק 1043-02-26): פרומפט
מגוזם-ומוכוון-שיפוט מפיק היסק משפטי חד יותר מתבנית נוקשה. משכתב את
תבנית מסמך-ההכוונה של ה-CEO לכותב בלבד — משמר את החוזה המלא (5 הרכיבים,
chair_directions בתגית, אילוצי-הסגנון) ומזריק את רמזי-השיפוט שהוכחו:
אדנים עצמאיים, מוקשי-עקביות-פנימית, מענה לצד המפסיד, הובלה בסוגיה
המכריעה. שאר מכונת-התזמור התפעולית (שערים, סטטוסים, API, MCP-race)
לא נגעה. net -27/+20.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-25 21:28:16 +00:00
cc4d757fce Merge pull request 'refactor(agents): גיזום כפילויות + rubric-קבלה משותף (context-engineering לדור-5)' (#418) from worktree-agent-prompts-trim into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 11s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-07-25 17:33:52 +00:00
6 changed files with 291 additions and 43 deletions

View File

@@ -820,45 +820,38 @@ ls data/cases/$CASE_NUMBER/documents/research/analysis-and-research.md
--- ---
**תבנית issue לכותב ההחלטה — חובה בכל issue שמוקצה לכותב:** **מסמך-ההכוונה לכותב — הפק את התדריך שהיית רוצה לקבל, לא טופס למילוי:**
כל issue לכותב חייב לכלול את **כל** הסעיפים הבאים. אסור לשלוח issue עם משפט כמו "הועבר לכתיבה" — זה חסר תועלת. הכותב צריך הכל מוכן מראש. כשאתה מעביר תיק לכותב אתה מבצע את **פעולת-ההיסק המרכזית שלך**: להמיר את ניתוח-המנתח + הכרעות-היו"ר למסמך שמאפשר לכותב לנסח החלטה חדה בסגנון דפנה **בלי לחזור אליך**. אל תמלא טופס — הפעל שיפוט משפטי. תדריך טוב:
- **מוביל בהכרעה ובסוגיה המכריעה** — קבע איזו סוגיה נושאת את התוצאה ומה מייתר את מה, והצב אותה ראשונה.
- **בונה כל סוגיה כסילוגיזם** (כלל → עובדות → מסקנה) עם התקדים והמסמך הספציפיים.
- **מזהה אדנים עצמאיים** — אם יותר מנימוק אחד מספיק לבדו לתוצאה, אמור זאת מפורשות, כך שנפילת אדן בערעור לא תפיל את ההחלטה.
- **בודק עקביות פנימית** — אם שתי הכרעות עלולות להיראות סותרות (למשל דחיית טענה פרשנית אחת וקבלת אחרת), סמן את המתח והסבר את האבחנה לפני שעורך-דין יטען לו.
- **עונה לנקודה החזקה של הצד המפסיד** — לא מתעלם ממנה.
- **משקלל את הכרעות-היו"ר** ומעביר אותן מילולית.
**מה התדריך חייב להכיל** (החוזה מול הכותב — אל תשמיט אף רכיב; אל תשלח issue עם "הועבר לכתיבה"):
```markdown ```markdown
## הנחיות כתיבה — ערר {case_number} ## הנחיות כתיבה — ערר {case_number}
### 1. תוצאה ומצב ### 1. תוצאה ומצב
- **תוצאה:** {דחייה / קבלה חלקית / קבלה מלאה} - **תוצאה:** {דחייה / קבלה חלקית / קבלה מלאה} — עם נימוק קצר ומהי הראיה הניצחת.
- **טיוטה קיימת:** {כן/לא}. אם כן: נתיב מלא לקובץ + הנחיה "קרא את הטיוטה, השתמש בה כבסיס, אל תכתוב מאפס" - **טיוטה קיימת:** {כן/לא}. אם כן: נתיב מלא + "קרא, השתמש כבסיס, אל תכתוב מאפס".
- **הוראות עריכה מתוך הטיוטה:** {רשימה מדויקת של מה חיים ביקש לשנות — פסקאות, תוכן, placeholders} - **הוראות עריכה מהטיוטה:** {מה חיים ביקש לשנות — פסקאות, תוכן, placeholders}.
### 2. סדר סוגיות + מבנה סילוגיסטי ### 2. סוגיות — סדר סילוגיסטי, המכריעה מובילה
לכל סוגיה שצריך לכתוב/לערוך — מבנה סילוגיסטי מלא: לכל סוגיה: סוג-ניתוח (כלל ברור / איזון / מידתיות / שיקול-דעת) · כלל (ציטוט מדויק של הוראת-תכנית/חוק/הלכה) · עובדות (בהפניה למסמך-מקור ספציפי) · מסקנה · תקדימים (שם + מה קובע + רלוונטיות) · מסמכי-מקור (ב-data/cases/{case_number}/documents/originals/). סמן אדנים עצמאיים, מוקשי-עקביות ומענה לצד המפסיד היכן שהם קיימים.
**סוגיה N: {כותרת}**
- סוג ניתוח: {כלל ברור / איזון אינטרסים / מידתיות / שיקול דעת}
- כלל (הנחה עליונה): {הוראת תכנית / סעיף חוק / הלכה — ציטוט מדויק}
- עובדות (הנחה תחתונה): {העובדות הספציפיות שצריך להחיל — הפנייה למסמך מקור ספציפי}
- מסקנה: {מה נובע מהחלת הכלל על העובדות}
- תקדימים: {שם פסק דין + מה הוא קובע + למה רלוונטי}
- מסמכי מקור: {שמות קבצים ספציפיים ב-data/cases/{case_number}/documents/originals/}
### 3. טיפול בטענות ### 3. טיפול בטענות
| # | טענה | טיפול | סוגיה | טבלה: # | טענה | טיפול (דיון מלא / קיבוץ / דילוג) | סוגיה.
|---|------|-------|-------|
| 1 | {טענה} | דיון מלא / קיבוץ / דילוג | {באיזו סוגיה} |
...
### 4. chair directions ### 4. הנחיות-היו"ר
- העתק מלא של עמדות הוועדה מ-analysis-and-research.md (או הפנייה: "קרא get_chair_directions"). העתק מילולי של עמדות-הוועדה מ-analysis-and-research.md (או "קרא get_chair_directions"), **עטוף ב-`<chair_directions>…</chair_directions>`** — טקסט מילולי בלבד בלי פרפרזה, כדי שהכותב לא ידרוס אותן.
- **עטוף את ההעתק המילולי בתגית `<chair_directions>…</chair_directions>`** — כך הכותב מבחין בין הוראות-היו"ר המילוליות לבין הערות ה-CEO, ואינו דורס אותן. בתוך התגית: טקסט מילולי בלבד, בלי פרפרזה.
### 5. הנחיות סגנון ### 5. הנחיות סגנון
- ניטרליות: בלוק ו = עובדות בלבד, בלי ציטוטים מצדדים ניטרליות (בלוק ו = עובדות בלבד, בלי ציטוטי-צדדים) · ללא כפילות (בלוק י מפנה לקודמים) · טענות מקוריות (בלוק ז) · דיון ≥ 1,500 מילים · ≥ 3 תקדימים בדיון.
- ללא כפילות: בלוק י מפנה לבלוקים קודמים
- טענות מקוריות: בלוק ז = כתבי טענות מקוריים
- אורך מינימלי לדיון: 1,500 מילים לבלוק י
- פסיקה: חובה לצטט לפחות 3 תקדימים בדיון
``` ```
--- ---

View File

@@ -23,8 +23,10 @@ from pathlib import Path
from legal_mcp import config from legal_mcp import config
from legal_mcp.services import db, embeddings, claude_session, audit, storage from legal_mcp.services import db, embeddings, claude_session, audit, storage
from legal_mcp.services.lessons import ( from legal_mcp.services.lessons import (
ANTI_PATTERNS as _ANTI_PATTERNS,
OUTCOME_LABELS_HE, OUTCOME_LABELS_HE,
PRACTICE_AREA_OVERRIDES, PRACTICE_AREA_OVERRIDES,
anti_pattern_directive,
canonical_outcome, canonical_outcome,
get_content_checklist, get_content_checklist,
get_methodology_summary, get_methodology_summary,
@@ -369,6 +371,7 @@ async def write_block(
block_id: str, block_id: str,
instructions: str = "", instructions: str = "",
effort_override: str | None = None, effort_override: str | None = None,
model_override: str | None = None,
) -> dict: ) -> dict:
"""כתיבת בלוק יחיד בהחלטה. """כתיבת בלוק יחיד בהחלטה.
@@ -381,6 +384,12 @@ async def write_block(
THIS call only — used by the #208 model/effort calibration harness THIS call only — used by the #208 model/effort calibration harness
to A/B efforts without mutating the pinned defaults. Production to A/B efforts without mutating the pinned defaults. Production
callers leave it None and get the deterministic per-block effort. callers leave it None and get the deterministic per-block effort.
model_override: optional per-call generation model id (e.g.
"claude-opus-5"). Same contract as effort_override — the #208
harness A/Bs MODELS without mutating the pinned GENERATION_MODEL.
Pass the BASE id only: the 1M-context escalation (#216) is applied
on top automatically for large prompts, so an override never
silently loses the 1M window. Production callers leave it None.
Returns: Returns:
dict עם content, word_count, block_id, generation_type dict עם content, word_count, block_id, generation_type
@@ -468,6 +477,12 @@ async def write_block(
if instructions: if instructions:
prompt += f"\n\n## הנחיות נוספות:\n{instructions}" prompt += f"\n\n## הנחיות נוספות:\n{instructions}"
# LAST in the prompt, deliberately (see lessons.anti_pattern_directive): the
# same canonical rule already appears inside style_context, but ~47K chars
# deep, where it measurably fails to bind. Restating it here is the only
# change the A/B isolated as effective — so nothing may be appended after it.
prompt += "\n\n" + anti_pattern_directive()
# Block י requires approved direction # Block י requires approved direction
if block_id == "block-yod": if block_id == "block-yod":
dir_doc = (decision or {}).get("direction_doc") or {} dir_doc = (decision or {}).get("direction_doc") or {}
@@ -478,7 +493,12 @@ async def write_block(
# escalate to the 1M-context build (`[1m]`) instead of failing the block — # escalate to the 1M-context build (`[1m]`) instead of failing the block —
# block-yod legitimately carries the whole case as source-context. The 400K # block-yod legitimately carries the whole case as source-context. The 400K
# ceiling was an artifact of the old 200K-only build, NOT a model limit. # ceiling was an artifact of the old 200K-only build, NOT a model limit.
gen_model = GENERATION_MODEL_1M if len(prompt) > _CTX_1M_THRESHOLD_CHARS else GENERATION_MODEL # model_override (#208 harness) swaps the BASE id only — the 1M decision below
# still applies, so an A/B'd model keeps the same context-window behaviour as
# the pinned default instead of silently falling back to the 200K build.
_base_model = model_override or GENERATION_MODEL
_model_1m = GENERATION_MODEL_1M if _base_model == GENERATION_MODEL else f"{_base_model}[1m]"
gen_model = _model_1m if len(prompt) > _CTX_1M_THRESHOLD_CHARS else _base_model
# Final guard: even the 1M build is finite (~2M Hebrew chars of input). Cap at # Final guard: even the 1M build is finite (~2M Hebrew chars of input). Cap at
# 1.5M chars (~750K tokens) to leave room for output + a safety margin under 1M. # 1.5M chars (~750K tokens) to leave room for output + a safety margin under 1M.
@@ -1107,6 +1127,16 @@ async def _build_style_context(practice_area: str = "") -> str:
# ── למידה מצטברת (T15) — עריכות היו"ר ב-/methodology + לקחי /training ── # ── למידה מצטברת (T15) — עריכות היו"ר ב-/methodology + לקחי /training ──
# גובר על ברירות-המחדל לעיל. כך כל מה שלמדנו עד היום מגיע לכותב. # גובר על ברירות-המחדל לעיל. כך כל מה שלמדנו עד היום מגיע לכותב.
learned: list[str] = [] learned: list[str] = []
# The canonical anti-patterns are rendered UNCONDITIONALLY, before any DB
# call. They used to be produced inside the overrides loop below — so a
# failure on an EARLIER category (e.g. golden_ratios) aborted the loop and
# dropped the style invariants from the prompt silently, with only a generic
# "overrides not loaded" warning to show for it (§6). A chair-override
# outage must not be able to un-teach Dafna's structural style.
learned.append("\n**אנטי-דפוסים (להימנע) — כתוב נרטיב משפטי רציף; הימנע מ:**")
for ap in _ANTI_PATTERNS:
learned.append(f"- {ap['note']}")
try: try:
for cat, label in ( for cat, label in (
("golden_ratios", "יחסי-זהב (אחוזי-סעיפים)"), ("golden_ratios", "יחסי-זהב (אחוזי-סעיפים)"),
@@ -1125,10 +1155,8 @@ async def _build_style_context(practice_area: str = "") -> str:
# corrects them, and drafts keep emitting them (the gap that left # corrects them, and drafts keep emitting them (the gap that left
# 8137 with 28 hits). Chair additions layer on top; they never # 8137 with 28 hits). Chair additions layer on top; they never
# remove the canonical ones. # remove the canonical ones.
from legal_mcp.services.lessons import ANTI_PATTERNS as _ANTI # The canonical list is already rendered above, outside this try —
learned.append(f"\n**{label} — כתוב נרטיב משפטי רציף; הימנע מ:**") # here we only layer the chair's ADDITIONS on top of it.
for ap in _ANTI:
learned.append(f"- {ap['note']}")
for k, v in (ov or {}).items(): for k, v in (ov or {}).items():
learned.append(f"- (יו\"ר) {k}: {json.dumps(v, ensure_ascii=False)}") learned.append(f"- (יו\"ר) {k}: {json.dumps(v, ensure_ascii=False)}")
continue continue
@@ -1264,6 +1292,12 @@ async def get_block_context(case_id: UUID, block_id: str, instructions: str = ""
if instructions: if instructions:
formatted_prompt += f"\n\n## הנחיות נוספות:\n{instructions}" formatted_prompt += f"\n\n## הנחיות נוספות:\n{instructions}"
# Same closing directive, same position, same canonical source as write_block.
# This is the EXTERNAL-writer path (legal-writer agent) — if the rule were
# applied only in write_block, agent-written blocks would keep emitting the
# anti-patterns and the two writers would drift apart (G2).
formatted_prompt += "\n\n" + anti_pattern_directive()
# Block י requires approved direction # Block י requires approved direction
if block_id == "block-yod": if block_id == "block-yod":
dir_doc = (decision or {}).get("direction_doc") or {} dir_doc = (decision or {}).get("direction_doc") or {}

View File

@@ -59,6 +59,28 @@ ANTI_PATTERNS: list[dict] = [
"note": "רשימות תבליטים באנליזה — דפנה כותבת נרטיב רציף"}, "note": "רשימות תבליטים באנליזה — דפנה כותבת נרטיב רציף"},
] ]
def anti_pattern_directive() -> str:
"""The closing style directive, rendered from ANTI_PATTERNS (the same list
style_distance scores against — one source, two renderings, not two rules).
WHY THIS EXISTS SEPARATELY FROM the style-context rendering: the rule was
already reaching the writer, buried ~47K chars deep inside style_context,
and drafts kept emitting the very patterns it forbids. A measured A/B over
the signed finals (9 cases, 60 generations, 2026-07-28) showed that the SAME
rule restated at the END of the assembled prompt cuts anti-pattern hits by
7293% on both blocks and both models:
block-vav opus-4-8 1.75 → 0.12 | opus-5 2.25 → 0.62
block-zayin opus-4-8 4.57 → 0.43 | opus-5 4.43 → 0.43
So this is a POSITION fix, not a new instruction. Keep it last in the prompt.
"""
lines = ["## כלל-סגנון מחייב (גובר על כל דוגמה בהקשר שלמעלה)",
"כתוב נרטיב משפטי רציף בלבד — פסקאות שלמות. אסור:"]
lines += [f"- {ap['note']}" for ap in ANTI_PATTERNS]
return "\n".join(lines)
# ── Paragraph length guidance (word counts) ──────────────────────── # ── Paragraph length guidance (word counts) ────────────────────────
PARAGRAPH_LENGTHS = { PARAGRAPH_LENGTHS = {

View File

@@ -176,7 +176,13 @@ def block_distance_to_final(
outcome = canonical_outcome(outcome) outcome = canonical_outcome(outcome)
diff = compute_diff_stats(regenerated_text or "", final_section_text or "") diff = compute_diff_stats(regenerated_text or "", final_section_text or "")
change_percent = diff["change_percent"] change_percent = diff["change_percent"]
anti_total = count_anti_patterns(regenerated_text or "")["total"] anti = count_anti_patterns(regenerated_text or "")
anti_total = anti["total"]
# Per-pattern breakdown, not just the total: a calibration run that only
# reports "anti=4" cannot tell you WHICH rule was broken, so it cannot say
# what to fix. (Diagnosing the 2026-07-28 model A/B needed exactly this and
# had to fall back on inference.)
anti_by_pattern = {name: h["count"] for name, h in anti["by_pattern"].items()}
section = _BLOCK_TO_SECTION.get(block_id) section = _BLOCK_TO_SECTION.get(block_id)
regen_words = len((regenerated_text or "").split()) regen_words = len((regenerated_text or "").split())
@@ -205,6 +211,7 @@ def block_distance_to_final(
"final_words": final_words, "final_words": final_words,
"change_percent": change_percent, "change_percent": change_percent,
"anti_pattern_total": anti_total, "anti_pattern_total": anti_total,
"anti_by_pattern": anti_by_pattern,
"golden_ratio_deviation_pp": ratio_dev, "golden_ratio_deviation_pp": ratio_dev,
"distance": distance, "distance": distance,
} }

View File

@@ -0,0 +1,60 @@
"""The style invariants must actually REACH the writer.
Both tests here cover defects found by the 2026-07-28 model×prompt A/B over the
signed finals: the canonical anti-patterns were present in the prompt but buried
~47K chars into style_context (where they measurably failed to bind), and they
were rendered inside a try/except that an unrelated DB failure could abort.
"""
import pytest
from legal_mcp.services import block_writer
from legal_mcp.services.lessons import ANTI_PATTERNS, anti_pattern_directive
def test_directive_renders_every_canonical_anti_pattern():
"""One source, two renderings — the directive may not drift from the list
style_distance scores against."""
text = anti_pattern_directive()
for ap in ANTI_PATTERNS:
assert ap["note"] in text, f"missing anti-pattern in directive: {ap['name']}"
def test_both_writer_paths_append_the_directive_last():
"""write_block (in-process) and get_block_context (legal-writer agent) must
both close with the directive — otherwise the two writers drift (G2)."""
import inspect
src = inspect.getsource(block_writer)
for fn in ("async def write_block(", "async def get_block_context("):
start = src.index(fn)
# bound the search to this function: up to the next top-level def
rest = src[start + len(fn):]
nxt = rest.find("\nasync def ")
body = rest[: nxt if nxt != -1 else len(rest)]
assert "anti_pattern_directive()" in body, f"{fn} does not append the style directive"
@pytest.mark.asyncio
async def test_style_context_keeps_anti_patterns_when_overrides_fail(monkeypatch):
"""A chair-override outage must not silently un-teach the structural style.
Regression: the canonical list used to be emitted inside the overrides loop,
so a throw on an EARLIER category (golden_ratios) dropped it entirely.
"""
async def _boom(*a, **k):
raise RuntimeError("methodology table unavailable")
async def _empty(*a, **k):
return []
# Every DB accessor this function touches is stubbed — the test must not open
# a real connection (a live pool here leaks across the shared event loop and
# breaks unrelated tests later in the run).
monkeypatch.setattr(block_writer.db, "get_style_patterns", _empty)
monkeypatch.setattr(block_writer.db, "get_methodology_overrides", _boom)
monkeypatch.setattr(block_writer.db, "get_recent_decision_lessons", _empty)
ctx = await block_writer._build_style_context("היטל השבחה")
assert "נרטיב משפטי רציף" in ctx
for ap in ANTI_PATTERNS:
assert ap["note"] in ctx, f"anti-pattern dropped on override failure: {ap['name']}"

View File

@@ -166,17 +166,33 @@ def aggregate_cell(per_run: list[dict]) -> dict:
"""Mean each metric across repeated generations of the same (case, block, effort).""" """Mean each metric across repeated generations of the same (case, block, effort)."""
if not per_run: if not per_run:
return {"distance": 1.0, "anti_pattern_total": 0.0, "change_percent": 100.0, return {"distance": 1.0, "anti_pattern_total": 0.0, "change_percent": 100.0,
"golden_ratio_deviation_pp": None, "n": 0} "golden_ratio_deviation_pp": None, "anti_by_pattern": {}, "n": 0}
ratios = [r["golden_ratio_deviation_pp"] for r in per_run if r.get("golden_ratio_deviation_pp") is not None] ratios = [r["golden_ratio_deviation_pp"] for r in per_run if r.get("golden_ratio_deviation_pp") is not None]
return { return {
"distance": round(mean(r["distance"] for r in per_run), 4), "distance": round(mean(r["distance"] for r in per_run), 4),
"anti_pattern_total": round(mean(r["anti_pattern_total"] for r in per_run), 2), "anti_pattern_total": round(mean(r["anti_pattern_total"] for r in per_run), 2),
"change_percent": round(mean(r["change_percent"] for r in per_run), 2), "change_percent": round(mean(r["change_percent"] for r in per_run), 2),
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None, "golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
"anti_by_pattern": _mean_by_pattern(per_run),
"n": len(per_run), "n": len(per_run),
} }
def _mean_by_pattern(per_run: list[dict]) -> dict:
"""Mean hits PER anti-pattern name across runs — the 'which rule broke' view.
A pattern absent from a run counts as 0 (count_anti_patterns omits zero-hit
keys), so the mean is over ALL runs, not only the ones that tripped it.
"""
names: set[str] = set()
for r in per_run:
names |= set((r.get("anti_by_pattern") or {}).keys())
return {
name: round(mean((r.get("anti_by_pattern") or {}).get(name, 0) for r in per_run), 2)
for name in sorted(names)
}
def _current_default(block_id: str) -> str | None: def _current_default(block_id: str) -> str | None:
from legal_mcp.services.block_writer import BLOCK_CONFIG, DEFAULT_EFFORT from legal_mcp.services.block_writer import BLOCK_CONFIG, DEFAULT_EFFORT
cfg = BLOCK_CONFIG.get(block_id, {}) cfg = BLOCK_CONFIG.get(block_id, {})
@@ -420,13 +436,30 @@ async def _finals_for_calibration(case_filter: str | None) -> list[dict]:
async def _score_cell(case_id, block_id: str, effort: str, final_section: str, async def _score_cell(case_id, block_id: str, effort: str, final_section: str,
final_total_words: int, outcome: str, repeats: int) -> dict: final_total_words: int, outcome: str, repeats: int,
"""Generate `block_id` at `effort` `repeats` times; score each vs the final section.""" model: str | None = None, instructions: str = "") -> dict:
"""Generate `block_id` at `effort` `repeats` times; score each vs the final section.
`model` (optional) A/Bs the generation model via write_block(model_override=…).
None ⇒ the pinned GENERATION_MODEL, i.e. the production path unchanged.
`instructions` (optional) is appended to the block prompt for EVERY cell in
the run — a prompt-variant A/B (e.g. an explicit formatting rule). It is
applied to all models so the comparison stays a model comparison rather
than silently becoming a prompt comparison.
"""
from legal_mcp.services import block_writer from legal_mcp.services import block_writer
from legal_mcp.services.style_distance import block_distance_to_final from legal_mcp.services.style_distance import block_distance_to_final
runs: list[dict] = [] runs: list[dict] = []
models_used: list[str] = []
for _ in range(repeats): for _ in range(repeats):
res = await block_writer.write_block(case_id, block_id, effort_override=effort) res = await block_writer.write_block(
case_id, block_id, instructions=instructions,
effort_override=effort, model_override=model,
)
# Record what the CLI was actually asked to run, so a silent fallback to
# a different build is visible in the report rather than mis-attributed.
models_used.append(res.get("model_used") or "?")
scored = block_distance_to_final( scored = block_distance_to_final(
block_id, res.get("content", ""), final_section, outcome, block_id, res.get("content", ""), final_section, outcome,
section_target_total_words=final_total_words, section_target_total_words=final_total_words,
@@ -434,6 +467,8 @@ async def _score_cell(case_id, block_id: str, effort: str, final_section: str,
runs.append(scored) runs.append(scored)
agg = aggregate_cell(runs) agg = aggregate_cell(runs)
agg["effort"] = effort agg["effort"] = effort
agg["model"] = model
agg["models_used"] = sorted(set(models_used))
agg["runs"] = runs agg["runs"] = runs
return agg return agg
@@ -446,6 +481,7 @@ async def _run(args, ts: str) -> dict:
efforts = args.efforts efforts = args.efforts
blocks = args.blocks blocks = args.blocks
models = args.models
finals = await _finals_for_calibration(args.case) finals = await _finals_for_calibration(args.case)
cases_meta = [] cases_meta = []
@@ -468,11 +504,15 @@ async def _run(args, ts: str) -> dict:
section = _BLOCK_TO_SECTION.get(block_id) section = _BLOCK_TO_SECTION.get(block_id)
plan[block_id] = [c for c in cases_meta if section and c["sections"].get(section)] plan[block_id] = [c for c in cases_meta if section and c["sections"].get(section)]
total_cells = sum(len(plan[b]) for b in blocks) * len(efforts) * args.repeats total_cells = sum(len(plan[b]) for b in blocks) * len(efforts) * args.repeats * len(models)
grid_summary = { grid_summary = {
"n_finals": len(cases_meta), "n_finals": len(cases_meta),
"finals": [c["case_number"] for c in cases_meta], "finals": [c["case_number"] for c in cases_meta],
"blocks": blocks, "efforts": efforts, "repeats": args.repeats, "blocks": blocks, "efforts": efforts, "repeats": args.repeats,
"models": models,
# Provenance: a prompt-variant run is NOT comparable to a baseline run,
# so the instruction text is recorded in the report, not just the shell.
"instructions": getattr(args, "instructions", "") or "",
"total_generations": total_cells, "total_generations": total_cells,
"per_block_n": {b: len(plan[b]) for b in blocks}, "per_block_n": {b: len(plan[b]) for b in blocks},
} }
@@ -480,6 +520,27 @@ async def _run(args, ts: str) -> dict:
if args.dry_run: if args.dry_run:
return {"dry_run": True, "grid": grid_summary, "by_block": {}} return {"dry_run": True, "grid": grid_summary, "by_block": {}}
by_model: dict[str, dict] = {}
for model in models:
by_block = await _run_blocks_for_model(
model, blocks, efforts, plan, args, ts, grid_summary, by_model, _BLOCK_TO_SECTION,
)
by_model[model] = by_block
# `by_block` stays the single-model shape (first model) so --rerank and the
# existing per-block report path keep working unchanged (G2 — no second
# result schema); multi-model runs additionally carry by_model.
out = {"dry_run": False, "grid": grid_summary, "by_block": by_model[models[0]]}
if len(models) > 1:
out["by_model"] = by_model
return out
async def _run_blocks_for_model(model, blocks, efforts, plan, args, ts, grid_summary,
by_model_so_far, _BLOCK_TO_SECTION) -> dict:
"""The per-block × per-effort grid for ONE generation model."""
from uuid import UUID
by_block: dict[str, dict] = {} by_block: dict[str, dict] = {}
for block_id in blocks: for block_id in blocks:
section = _BLOCK_TO_SECTION.get(block_id) section = _BLOCK_TO_SECTION.get(block_id)
@@ -497,11 +558,12 @@ async def _run(args, ts: str) -> dict:
cell = await _score_cell( cell = await _score_cell(
UUID(c["case_id"]), block_id, effort, final_section, UUID(c["case_id"]), block_id, effort, final_section,
c["final_total_words"], c["outcome"], args.repeats, c["final_total_words"], c["outcome"], args.repeats,
model=model, instructions=getattr(args, "instructions", "") or "",
) )
except Exception as exc: # noqa: BLE001 — harness must survive any cell failure except Exception as exc: # noqa: BLE001 — harness must survive any cell failure
logger.warning( logger.warning(
"calibration cell skipped: case=%s block=%s effort=%s%s", "calibration cell skipped: case=%s block=%s effort=%s model=%s%s",
c["case_number"], block_id, effort, exc, c["case_number"], block_id, effort, model, exc,
) )
continue continue
per_effort_runs[effort].append(cell) per_effort_runs[effort].append(cell)
@@ -523,6 +585,7 @@ async def _run(args, ts: str) -> dict:
"anti_pattern_total": round(mean(r["anti_pattern_total"] for r in rows), 2), "anti_pattern_total": round(mean(r["anti_pattern_total"] for r in rows), 2),
"change_percent": round(mean(r["change_percent"] for r in rows), 2), "change_percent": round(mean(r["change_percent"] for r in rows), 2),
"golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None, "golden_ratio_deviation_pp": round(mean(ratios), 2) if ratios else None,
"anti_by_pattern": _mean_by_pattern(rows),
"n": len(rows), "n": len(rows),
}) })
rec = recommend_effort(effort_rows) rec = recommend_effort(effort_rows)
@@ -532,6 +595,11 @@ async def _run(args, ts: str) -> dict:
"recommended": rec["effort"] if rec else None, "recommended": rec["effort"] if rec else None,
"confidence": rec["confidence"] if rec else None, "confidence": rec["confidence"] if rec else None,
"confidence_margin": rec.get("confidence_margin") if rec else None, "confidence_margin": rec.get("confidence_margin") if rec else None,
"model": model,
# Model builds the CLI actually reported across this block's cells —
# a mismatch vs `model` means a silent fallback, not a real A/B.
"models_used": sorted({m for e in per_effort_runs.values()
for cell in e for m in cell.get("models_used", [])}),
"efforts": effort_rows, "efforts": effort_rows,
"per_case": per_case, "per_case": per_case,
} }
@@ -541,11 +609,15 @@ async def _run(args, ts: str) -> dict:
# Blocks not yet done are simply absent from by_block; _write_report tolerates # Blocks not yet done are simply absent from by_block; _write_report tolerates
# partial results. main() does the final flush once the loop finishes. # partial results. main() does the final flush once the loop finishes.
try: try:
_write_report({"dry_run": False, "grid": grid_summary, "by_block": by_block}, ts) snap = {"dry_run": False, "grid": grid_summary, "by_block": by_block}
if by_model_so_far or len(grid_summary.get("models", [])) > 1:
snap["by_model"] = {**by_model_so_far, model: by_block}
_write_report(snap, ts)
except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run except Exception as exc: # noqa: BLE001 — a write hiccup must not abort the run
logger.warning("incremental report write failed after block=%s%s", block_id, exc) logger.warning("incremental report write failed after block=%s model=%s %s",
block_id, model, exc)
return {"dry_run": False, "grid": grid_summary, "by_block": by_block} return by_block
IL_TZ = ZoneInfo("Asia/Jerusalem") IL_TZ = ZoneInfo("Asia/Jerusalem")
@@ -578,7 +650,10 @@ def _write_report(result: dict, ts: str) -> tuple[Path, Path]:
"ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n", "ההמלצה אדוויזורית; ההכרעה בידי היו\"ר/המפעיל.\n",
f"- בלוקים: {', '.join(g['blocks'])}", f"- בלוקים: {', '.join(g['blocks'])}",
f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}", f"- efforts: {', '.join(g['efforts'])} · repeats/cell: {g['repeats']}",
f"- models: {', '.join(m or 'pinned-default' for m in g.get('models', [None]))}",
f"- סך ייצורי-מודל: {g['total_generations']}", f"- סך ייצורי-מודל: {g['total_generations']}",
(f"- ⚠️ **וריאנט-פרומפט** (לא בר-השוואה לריצת-בסיס): `{g['instructions']}`"
if g.get("instructions") else "- וריאנט-פרומפט: — (פרומפט ייצור כפי-שהוא)"),
"", "",
] ]
if result.get("dry_run"): if result.get("dry_run"):
@@ -613,6 +688,54 @@ def _write_report(result: dict, ts: str) -> tuple[Path, Path]:
f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | " f"| {r['effort']}{star} | {r['distance']:.4f} | {r['anti_pattern_total']} | "
f"{r['change_percent']} | {ratio if ratio is not None else ''} | {r['n']} |") f"{r['change_percent']} | {ratio if ratio is not None else ''} | {r['n']} |")
lines.append("") lines.append("")
by_model = result.get("by_model") or {}
if len(by_model) > 1:
lines += ["## השוואת-מודלים (אותו block, אותו effort, אותם סופיים)\n",
"| block | effort | model | anti_total | change% | ratioΔpp | distance | n |",
"|---|---|---|---|---|---|---|---|"]
for b in g["blocks"]:
for eff in g["efforts"]:
rows = []
for m, bb in by_model.items():
for r in (bb.get(b) or {}).get("efforts", []):
if r["effort"] == eff:
rows.append((m, r))
if len(rows) < 2:
continue # nothing to compare for this cell — don't fake a row
best = min(rows, key=lambda mr: (mr[1]["anti_pattern_total"],
mr[1]["golden_ratio_deviation_pp"] or 0,
mr[1]["distance"]))[0]
for m, r in rows:
ratio = r["golden_ratio_deviation_pp"]
star = "" if m == best else ""
lines.append(
f"| {b} | {eff} | {m}{star} | {r['anti_pattern_total']} | "
f"{r['change_percent']} | {ratio if ratio is not None else ''} | "
f"{r['distance']:.4f} | {r['n']} |")
lines.append("")
# WHICH rule broke — a total alone can't tell you what to fix.
bd_rows = [(b, eff, m, r) for b in g["blocks"] for eff in g["efforts"]
for m, bb in by_model.items()
for r in (bb.get(b) or {}).get("efforts", []) if r["effort"] == eff]
if any(r.get("anti_by_pattern") for *_, r in bd_rows):
names = sorted({n for *_, r in bd_rows for n in (r.get("anti_by_pattern") or {})})
lines += ["### פילוח אנטי-דפוסים (איזה כלל הופר)\n",
"| block | effort | model | " + " | ".join(names) + " |",
"|---|---|---|" + "---|" * len(names)]
for b, eff, m, r in bd_rows:
cells = " | ".join(str((r.get("anti_by_pattern") or {}).get(n, 0)) for n in names)
lines.append(f"| {b} | {eff} | {m} | {cells} |")
lines.append("")
# A silent CLI fallback would make the whole comparison meaningless — surface it.
for m, bb in by_model.items():
for b, bd in bb.items():
used = bd.get("models_used") or []
if used and any(not u.startswith(str(m)) for u in used):
lines.append(f"> ⚠️ **{b} / {m}**: ה-CLI דיווח `{', '.join(used)}` — "
"ייתכן fallback שקט; ההשוואה לתא זה אינה תקפה.\n")
lines.append("")
lines.append("> דירוג-ההמלצה **style-clean** (#213): anti_total ראשי → ratioΔ → distance (tiebreak). " lines.append("> דירוג-ההמלצה **style-clean** (#213): anti_total ראשי → ratioΔ → distance (tiebreak). "
"**change% מדווח-לא-מדורג** — מערבב סגנון עם שלמות-תוכן (07-learning §0.7), " "**change% מדווח-לא-מדורג** — מערבב סגנון עם שלמות-תוכן (07-learning §0.7), "
"anti_total הוא הסיגנל הנקי-לסגנון. confidence=⚠weak ⇒ הבחירה בתוך-הרעש " "anti_total הוא הסיגנל הנקי-לסגנון. confidence=⚠weak ⇒ הבחירה בתוך-הרעש "
@@ -633,6 +756,12 @@ async def main() -> int:
help="comma block ids to calibrate") help="comma block ids to calibrate")
ap.add_argument("--case", default=None, help="restrict to a single case_number") ap.add_argument("--case", default=None, help="restrict to a single case_number")
ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)") ap.add_argument("--repeats", type=int, default=1, help="generations per cell (avg out gen noise)")
ap.add_argument("--models", default="",
help="comma generation-model ids to A/B (e.g. claude-opus-4-8,claude-opus-5). "
"Empty (default) = the pinned GENERATION_MODEL, i.e. production unchanged.")
ap.add_argument("--instructions", default="",
help="extra prompt instruction appended to EVERY cell (prompt-variant A/B). "
"Applied to all models — the run stays a model comparison. Recorded in the report.")
args = ap.parse_args() args = ap.parse_args()
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s") logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
@@ -653,6 +782,9 @@ async def main() -> int:
if bad_b: if bad_b:
print(f"non-calibratable block(s): {bad_b}. valid: {VALID_BLOCKS}", file=sys.stderr) print(f"non-calibratable block(s): {bad_b}. valid: {VALID_BLOCKS}", file=sys.stderr)
return 2 return 2
# [None] = "use the pinned GENERATION_MODEL" — keeps the default run byte-identical
# to the pre-#models behaviour instead of hard-coding the id in a second place (G2).
args.models = [m.strip() for m in args.models.split(",") if m.strip()] or [None]
ts = _ts() ts = _ts()
result = await _run(args, ts) result = await _run(args, ts)