Files
legal-ai/scripts/regenerate_voice_fingerprint.py
Chaim 9a19db1ec1
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
feat(learning): voice-fingerprint corpus-refresh proposal generator (#161, prose half)
The voice-fingerprint (docs/daphna-voice-fingerprint.md) is the PRIMARY abstract
style channel the writer consumes (07-learning §0.2 channel A), but it was
hand-authored and never regenerated from the growing corpus — the "voice" the
writer adapts stayed frozen while finals accumulated. PR #345 unfroze the
structural half (corpus-measured ratios → writer); this is the prose half.

scripts/regenerate_voice_fingerprint.py reads the corpus (style_corpus stats +
a bounded block-level style_exemplars sample + measured section ratios) and asks
Opus (local) to PROPOSE refinements/additions to the fingerprint — STYLE / METHOD
/ VOICE / LEXICON only, grounded in corpus evidence, never case substance
(INV-LRN5), preserving curated insights rather than replacing them.

CHAIR-GATED, HARD (INV-LRN1): the prose profile is voice-knowledge, not a
low-risk style rule, so it does NOT auto-flow. The script writes a PROPOSAL to
data/curator-proposals/ (the existing curator-proposal surface) and NEVER
overwrites the live fingerprint; the chair reviews and hand-commits accepted parts
— the same manual gate the fingerprint already uses. Runs on the host (claude CLI).

Reuses claude_session + measure_corpus_ratios + the curator-proposals dir (G2 —
no parallel synthesis or gate). SCRIPTS.md + spec §0.2 updated.

Invariants: INV-LRN1 (hard chair gate for voice-knowledge prose), INV-LRN5
(style-only synthesis), G2 (reuses existing synthesis inputs + proposal surface).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 22:26:04 +00:00

114 lines
6.5 KiB
Python
Executable File
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env python3
"""Propose a corpus-refreshed update to Dafna's voice-fingerprint (TaskMaster #161).
The voice-fingerprint (docs/daphna-voice-fingerprint.md) is the PRIMARY abstract
style channel the writer consumes (07-learning §0.2, channel A). It was hand-authored
and never regenerated from the growing corpus — so the "voice" the writer adapts has
been frozen while the corpus grew. This script reads the corpus (style_corpus stats +
block-level style_exemplars sample + measured section ratios) and asks Opus (local)
to PROPOSE refinements/additions to the fingerprint — STYLE / METHOD / VOICE / LEXICON
ONLY, never case substance (INV-LRN5).
CHAIR-GATED (INV-LRN1, HARD gate — the prose profile is voice-knowledge, not a
low-risk style rule, so it does NOT auto-flow): the proposal is written to
data/curator-proposals/ and NEVER overwrites the live fingerprint. The chair reviews
it (in /training, with the other curator proposals) and hand-commits the parts she
accepts to docs/daphna-voice-fingerprint.md — the same manual gate the fingerprint
already uses. This is the prose half of the voice-profile refresh; the structural
half (corpus-measured ratios → writer) shipped in PR #345.
Runs on the HOST (claude_session needs the local claude CLI). Usage:
DOTENV_PATH=/home/chaim/.env DATA_DIR=/home/chaim/legal-ai/data \\
/home/chaim/legal-ai/mcp-server/.venv/bin/python scripts/regenerate_voice_fingerprint.py
"""
from __future__ import annotations
import asyncio
import logging
from datetime import datetime, timezone
from legal_mcp import config
from legal_mcp.services import claude_session, db
from legal_mcp.services.style_distance import measure_corpus_ratios
logging.basicConfig(level=logging.INFO, format="%(message)s")
log = logging.getLogger("voice_fingerprint_refresh")
FINGERPRINT = config.DATA_DIR.parent / "docs" / "daphna-voice-fingerprint.md"
PROPOSALS = config.DATA_DIR / "curator-proposals"
_PER_SECTION = 5 # exemplars per section in the evidence pack
_MAX_PARA_CHARS = 700 # cap each exemplar so the prompt stays bounded
_SYSTEM = """אתה מזקק את טביעת-אצבע-הקול של עו"ד דפנה תמיר — פרופיל-סגנון מופשט (איך היא כותבת: קול, שיטה, מבנה, לקסיקון, אנטי-דפוסים), לא תוכן.
חוקים מחייבים:
- **סגנון/שיטה בלבד (INV-LRN5):** אסור לכלול מהות משפטית ספציפית — הלכה, עובדה, תקדים, או ניסוח מתיק קונקרטי. רק הכללות-סגנון.
- **לחדד ולהרחיב, לא לזרוק:** שמר את התובנות האצורות בפרופיל הנוכחי; הצע תוספות/חידודים מעוגנים בראיות-הקורפוס שלהלן.
- **לעגן:** כל קביעה חדשה נשענת על דפוס שחוזר בראיות. אם אין עיגון — אל תמציא ואל תוסיף.
- שמר על מבנה-המסמך והסעיפים הקיימים.
הפלט: גרסת-fingerprint מוצעת מלאה (Markdown), מוכנה לסקירת-יו"ר — בלי הקדמות או הסברים מסביב."""
async def _gather_evidence() -> tuple[str, str]:
pool = await db.get_pool()
async with pool.acquire() as conn:
n_corpus = await conn.fetchval("SELECT count(*) FROM style_corpus WHERE coalesce(full_text,'') <> ''")
n_ex = await conn.fetchval("SELECT count(*) FROM style_exemplars")
rows = await conn.fetch(
"""SELECT section, decision_number, outcome, paragraph_text FROM (
SELECT section, decision_number, outcome, paragraph_text,
row_number() OVER (PARTITION BY section
ORDER BY decision_number DESC, word_count DESC) AS rn
FROM style_exemplars
) t WHERE rn <= $1 ORDER BY section, rn""",
_PER_SECTION,
)
ratios = await measure_corpus_ratios()
parts = [f"קורפוס: {n_corpus} החלטות, {n_ex} דוגמאות-בלוק.", "", "יחסי-מבנה מדודים מהקורפוס (אחוז-מהסך, לפי תוצאה):"]
for outcome, entry in (ratios or {}).items():
secs = ", ".join(f"{s} {round(p)}%" for s, p in (entry.get("sections") or {}).items())
parts.append(f" - {outcome} (n={entry.get('n', 0)}): {secs}")
parts += ["", "דוגמאות-סגנון מייצגות (פסקאות אמיתיות של דפנה — מקור ללמידת-קול, לא להעתקת-מהות):"]
for r in rows:
para = (r["paragraph_text"] or "").strip()[:_MAX_PARA_CHARS]
parts.append(f"\n[{r['section']} · {r['decision_number']} · {r['outcome'] or ''}]\n{para}")
return "\n".join(parts), f"{n_corpus} החלטות / {n_ex} דוגמאות"
async def main() -> int:
if not FINGERPRINT.exists():
log.error("fingerprint not found: %s", FINGERPRINT)
return 1
current = FINGERPRINT.read_text(encoding="utf-8")
evidence, summary = await _gather_evidence()
user = (
f"## הפרופיל הנוכחי (לחדד ולהרחיב, לא להחליף):\n{current}\n\n"
f"## ראיות-קורפוס:\n{evidence}\n\n"
"## המשימה:\nהצע גרסת-fingerprint מעודכנת המשלבת חידודים/תוספות-סגנון מעוגנים בראיות לעיל, "
"תוך שימור כל התובנות האצורות. סגנון/שיטה בלבד — בלי מהות."
)
log.info("synthesizing voice-fingerprint proposal from %s (Opus, local)…", summary)
proposed = await claude_session.query(
user, timeout=claude_session.LONG_TIMEOUT, system=_SYSTEM,
model="claude-opus-4-8", tools="",
)
if not proposed or not proposed.strip():
log.error("empty proposal from model — aborting (nothing written)")
return 1
PROPOSALS.mkdir(parents=True, exist_ok=True)
ts = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
out = PROPOSALS / f"voice-fingerprint-{ts}.md"
header = (
f"<!-- הצעת-חידוש לטביעת-אצבע-הקול · {ts} · מקור: {summary} -->\n"
"<!-- שער-יו\"ר (INV-LRN1): הצעה בלבד. לסקור ולקבע ידנית ל-docs/daphna-voice-fingerprint.md. -->\n\n"
)
out.write_text(header + proposed, encoding="utf-8")
log.info("✓ proposal written (chair-gated, NOT applied): %s", out)
log.info(" review in /training → commit accepted parts to %s", FINGERPRINT)
return 0
if __name__ == "__main__":
raise SystemExit(asyncio.run(main()))