78 Commits

Author SHA1 Message Date
97f3f50d0b Merge pull request 'feat(learning): chair-feedback style corrections auto-flow to writer (P1 #6)' (#344) from worktree-chair-feedback-flow into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m28s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 22:07:30 +00:00
8b23542dec feat(learning): chair-feedback style corrections auto-flow to the writer (P1 #6)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The chair-feedback chain was DEAD: 27 feedback rows captured, 0 ever became a
lesson the writer reads (no weekly-analysis job ever existed). The chair's own
corrections — the most authoritative signal of all, and exactly "learn from every
returned draft" — went nowhere.

decision_lessons can't carry them: it's FK-coupled to a style_corpus row (a signed
final), which a case-in-progress doesn't have. So chair STYLE feedback instead
rides the discussion_rules channel that already reaches the writer for all blocks —
the same path /training promote uses.

- db.append_global_rule: the locked read-modify-write append, extracted from web
  `_append_methodology_override` into one shared impl (G2). The web function is now
  a thin wrapper that seeds defaults; chair-feedback calls it directly.
- record_chair_feedback (MCP tool): a STYLE-category correction (style/wrong_tone/
  wrong_structure) with a lesson_extracted flows immediately to discussion_rules.
  The chair IS the gate — no separate approval (INV-LRN1 graduated gate). SUBSTANCE
  feedback (missing_content/factual_error/other) is case-specific → recorded only.

Invariants: INV-LRN1 (chair-authored style = highest authority, flows; substance
not auto-flowed), G2 (single append impl shared by promote + feedback), INV-LRN5
(style channel only).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 22:06:58 +00:00
b8eb0ee123 Merge pull request 'feat(learning): ניתוח-§A של האוצֵר רץ על mark-final — לכידת source='curator' על כל סופי (#159)' (#342) from worktree-curator-autofire into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 8s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 22:03:15 +00:00
b0a6c2fe01 Merge pull request 'feat(learning): grow style-exemplars from every final (P1 #5)' (#343) from worktree-grow-exemplars into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m28s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 21:58:12 +00:00
93a9404663 feat(learning): grow style-exemplars from every final (P1 #5)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The style-exemplar corpus (channel B — the writer's block-level retrieval of
Dafna's real prose) was FROZEN at the one-time seed backfill: new finals were
enrolled into style_corpus but never broken into exemplars, so the richest style
channel never grew (8137/8126/8174 had 0 exemplars). The writer kept retrieving
only March–April seed paragraphs no matter how many finals were signed.

Extract the per-decision exemplar logic (section→paragraph→Voyage-embed→replace)
into a shared service `legal_mcp.services.style_exemplars.extract_and_store` —
the SINGLE implementation now used by BOTH the one-time backfill and the live
enrollment path (G2; no parallel extractor). `_enroll_final_in_library` calls it
on every final upload (source='internal_committee', the same source the writer's
search_style_exemplars reads). Voyage embeds over REST → container-safe;
best-effort, surfaced in the upload response, never fails the upload.

Effect: every signed final now grows the exemplar corpus, so the writer's
block-level style retrieval improves with each decision — the core "learn from
every decision" fix for channel B. Path A (style_distance_history) will track
whether the larger exemplar pool reduces style-distance over time.

Invariants: G2 (one extraction path shared by backfill + enroll), INV-LRN5
(style/structure prose only — substance routes elsewhere), INV-LRN4 (the
draft↔final loop now feeds the exemplar channel, not just the lesson channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:57:38 +00:00
4d6268820c feat(learning): curator §A fires on mark-final — capture source='curator' findings on every final (#159)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
חלופה ז (agent-centric) — סוגר את פער-המימוש מ-#157:

1. final_learning_pipeline.py — סדר-מחדש: enroll_style_corpus רץ ראשון (לפני הדיסטילציה של עד
   30 דק'), כך שרשומת style_corpus קיימת תוך שניות — והאוצֵר יכול לרשום ממצאים בלי מירוץ.
   enroll/ingest בלתי-תלויים; panel נשאר אחרון. תוויות [1/3]/[2/3] עודכנו.

2. hermes-curator.md PIPELINE-WAKE BRANCH — מבדיל final_learning_* (ממשיך ל-§A, מצב AUTO) מ-
   final_halacha_* (exit כקודם). ב-AUTO: §A.1–§A.5b → record_curator_findings, דילוג על §A.6
   (interaction) כדי לא להעיר את דפנה (הממצאים proposed ונסקרים ב-/training). §A.5b קיבל retry
   כרשת-ביטחון למירוץ-שארית.

3. docs/spec/07-learning.md §0.6 + §1.1 — עודכנו לסדר-הצינור החדש; הערת "פער-מימוש פתוח" → "נסגר".

Invariants: G2 (מחולל-יחיד לממצאי-curator — הסוכן; אין כפילות-תבונה), INV-LRN1/G10 (proposed,
שער-יו"ר), INV-LRN3, INV-DUR1 (סדר-צעדים בשמות יציבים; checkpoint resume — edge-case של ריצה
מקבילה-לפריסה, סיכון נמוך). depends-on #157.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:56:36 +00:00
07bd8bee48 Merge pull request 'feat(learning): graduated gate — panel-consensus style lessons auto-flow to writer (P0)' (#341) from worktree-graduated-gate into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 7s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 21:46:29 +00:00
044de2c0f0 feat(learning): graduated gate — panel-consensus style lessons auto-flow to writer (P0)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The voice-learning panel produced 81 vetted style lessons — all stuck at
review_status='proposed' with 0 ever approved — so the writer (which reads only
'approved') received NONE of them. The system captured learning but never let it
flow. Chair decision (2026-06-28): a GRADUATED gate by content risk.

- STYLE lessons (categories style/structure/lexicon/tabular) the 2/2 panel kept →
  created as review_status='approved' → flow to the writer immediately, reversibly
  (chair vetoes in /training). The deepseek+gemini panel only emits style_method
  and only on 2/2 consensus, so this is exactly the gate's criterion; substance is
  already filtered out and skipped, and routes through the strict halacha gate.
- SUBSTANCE (halacha/precedent/fact) stays a HARD chair gate — unchanged.

style_lesson_panel.py: _review_status_for(category) sets the gate explicitly
(approved for _STYLE_CATEGORIES, else proposed); the apply loop passes it to
db.add_decision_lesson (which already accepts review_status — no DB change).

Spec: INV-LRN1 rewritten as the graduated gate (hard for substance; reversible
auto-flow for consensus style — still "under user control" via veto, per
NCSC/CEPEJ); §0.6 updated to match.

A one-time backfill of the 81 existing panel style lessons to 'approved' runs
separately (DB, post-merge).

Invariants: INV-LRN1 (amended — graduated), INV-LRN5 (style-only, substance never
auto-flows), G10 (human control preserved as reversible veto).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:45:56 +00:00
cfcfe4df48 Merge pull request 'fix(learning): כרטיס-אוצֵר כן-3-ערוצים + לכידת ממצאי-אוצֵר (source='curator')' (#339) from worktree-curator-learning-surface into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m29s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 21:28:53 +00:00
6427911baf Merge pull request 'fix(writer): always feed canonical anti-patterns to the writer' (#340) from worktree-writer-anti-patterns into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m25s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 21:24:57 +00:00
d7201736f2 fix(writer): always feed canonical anti-patterns to the writer
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The learning loop measured anti-patterns but never corrected them. style_distance
detects markdown headers / bullet lists / mid-paragraph mini-lists (from the
canonical lessons.ANTI_PATTERNS), but the writer only received anti-pattern
guidance if a chair `anti_patterns` override existed in appeal_type_rules — and
none does. So _build_style_context's `if ov:` branch injected nothing, the writer
was never told to avoid them, and drafts kept emitting them (8137: 28 hits, the
worst, newest — anti-patterns were trending UP, not down).

Anti-patterns are structural invariants of Dafna's voice (continuous legal
narrative — no markdown, no bullets), not overridable preferences. So inject the
canonical ANTI_PATTERNS notes ALWAYS, from the same list style_distance measures
against (single source of truth), with any chair additions layered on top. This
closes the measure-but-don't-correct gap: the next draft should show the markdown/
bullet anti-patterns drop, and Path A (style_distance_history) will confirm it.

Invariants: G1 (correct at source — the writer, not a post-hoc stripper), G2
(one canonical anti-pattern list shared by detection and instruction — no parallel
list), INV-LRN4 (closes the feedback half of the draft↔final loop).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:24:24 +00:00
be774ab87e fix(learning): honest 3-channel curator card + capture curator findings (source='curator')
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 9s
תיקוני-מבנה ללולאת-הלמידה (TaskMaster #157):

1. כרטיס ה-curator (/training טאב "אוצֵר") ספר decision_lessons WHERE source='curator'
   — ערך שאף מסלול-קוד לא כתב → תמיד 0. עוצב-מחדש (דרך שער-העיצוב Claude Design, אושר)
   להציג ביושר את שלושת ערוצי-ההזנה לכותב: דיסטילציה→appeal_type_rules (180, זורם),
   פאנל→decision_lessons (81, ממתין), אוצֵר→source='curator'. get_curator_stats שוכתב.

2. drift ספ↔סוכן: 07-learning.md §1.1 + INV-LRN3 קבעו שהאוצֵר רושם ממצאים כ-decision_lesson
   source='curator', אך הסוכן כתב comments בלבד — הממצאים האיכותיים אבדו. נוסף כלי-MCP
   record_curator_findings + §A.5b ב-hermes-curator.md (read-only נשמר; הצעה מגודרת-שער).

3. get_recent_decision_lessons(limit=15) חתך בשקט — נוסף WARN על מה שנחתך (חוקה §6);
   הפתרון האמיתי = סינתזת-לקחים (TaskMaster #158).

Invariants: מקיים INV-LRN1/G10 (שער-יו"ר), INV-LRN3 (לכידה מובנית), INV-IA2/IA5 (מקור-אמת
יחיד), G12 (leak-guard עובר), G2 (אין מסלול מקביל — אותם מאגרים). פער-מימוש פתוח מתועד:
§A לא רץ אוטומטית על mark-final (pipeline-wake exits) → נדחה לתכנון נפרד.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:20:05 +00:00
c00e7b4f49 Merge pull request 'docs(spec): Path A — prospective held-out style-distance trend (07-learning §0.7)' (#338) from worktree-spec-path-a into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 8s
G12 Leak-Guard / leak-guard (push) Successful in 3s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 21:16:35 +00:00
927be5c6bb docs(spec): add Path A — prospective held-out style-distance trend (07-learning §0.7)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
Documents the prospective held-out methodology shipped in PR #337: why a clean
retrospective held-out is impossible (lessons stored universal/untagged → no
leave-one-out), and how the system instead captures a clean generalization
datapoint at final-upload — before this case's lessons are folded — via a
style_distance snapshot + lesson-pool size into style_distance_history,
surfaced by GET /api/learning/style-distance-history.

Adds §0.7, wires step [7] MEASUREMENT and INV-LRN4 to it. Notes the
change_percent style/content confound and the one-time clean window.

Invariants documented: INV-LRN4 (this is its trend surface), G2 (reuses
style_distance + appeal_type_rules — no parallel metric path).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:16:11 +00:00
84afdcb36c Merge pull request 'feat(learning): prospective held-out style-distance trend (Path A)' (#337) from worktree-style-distance-history into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m28s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 21:03:50 +00:00
ab510328af feat(learning): prospective held-out style-distance trend (Path A)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 9s
A clean held-out test of voice learning was no longer runnable: every
final-uploaded case already has its lessons folded, and lessons are stored
universal/untagged so leave-one-out is impossible. Path A makes the test
prospective instead — capture the generalization datapoint at the one moment
it's clean.

On final upload, after the draft↔final pair is created but BEFORE this
case's lessons are folded (folding is a separate manual /training step), we
snapshot style_distance (anti_pattern_total, golden-ratio max-deviation,
change_percent) alongside the current voice-lesson pool size. Because the
draft was written with only the PRIOR pool, each row is a clean "with N
accumulated lessons, our draft on this unseen case scored X" datapoint. As
the pool grows over cases, a downward trend = learning generalizes.

- db: SCHEMA_V45 style_distance_history (append-only) + helpers
  voice_lesson_pool_sizes / record_style_distance_snapshot /
  get_style_distance_history.
- app: best-effort capture in api_upload_final_decision (never fails the
  upload); GET /api/learning/style-distance-history for the trend.

Reuses the existing style_distance service + appeal_type_rules pool — no
parallel metric path. The 8 existing cases are already folded, so the table
starts empty and fills from the next final (their clean window is past).

Invariants: G2 (reuse style_distance/appeal_type_rules — one path),
INV-LRN4 (measure the draft↔final gap; this is its trend surface). LLM-free
(style_distance is deterministic) so it runs in the container.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 21:03:10 +00:00
8d0e44937d Merge pull request 'fix(compose): wire staged-pipeline indicators to live learning-status' (#336) from worktree-compose-learning-indicators into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 44s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 17:55:02 +00:00
baf478d9dc fix(compose): wire staged-pipeline indicators to live learning-status
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The two stage indicators in the compose page's "השלמה והעברה" rail
("הרץ למידת-קול" / "הרץ אימות-הלכות") were static placeholder text
translated from mockup 03 and never wired to data — they always read
"ממתין להעלאת הסופי", even after the final was uploaded and both
pipelines completed. (The real status was already shown correctly by
LearningStatusBadges on the case page.)

Wire them to useCaseLearningStatus (/api/cases/{n}/learning-status) — the
same source the drafts-panel badges use (G2). The trailing text is now
derived: "ממתין להעלאת הסופי" until the final is uploaded, then the live
state ("✓ הושלם · 12 לקחים הופקו · 12 הוצעו לאישור" / "✓ הושלם · חולצו 48
· 33 אושרו · 15 נדחו", or running/queued/failed).

Data-binding bugfix only — identical markup/classes (mockup-03 layout
preserved), so no visual redesign; within the design-gate bugfix exception
(chair-approved approach). tsc + eslint clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 17:54:33 +00:00
6c487a2e85 Merge pull request 'fix(learning): extract precedent metadata inline on final-decision enroll' (#335) from worktree-enroll-inline-metadata into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m27s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 9s
2026-06-28 10:49:23 +00:00
8e2cfd602c fix(learning): extract precedent metadata inline on final-decision enroll
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 9s
When a chair's signed final decision is uploaded it is enrolled into the
precedent library, but _enroll_final_in_library only set deterministic
fields (citation, proceeding_type, date) and copied subject_tags from the
case's subject_categories — which is usually empty. The Gemini metadata
pass (subject_tags / summary / headnote / key_quote) was never triggered,
so the row sat at metadata_extraction_status='pending' with no subject
tags until a drain happened to pick it up. In practice a freshly-enrolled
final showed an empty "תגיות נושא" in the precedent edit UI.

Fix: after enrollment + citation, call the existing reextract_metadata
path inline (G2 — one path, full status lifecycle). It runs on Gemini
Flash over REST (GOOGLE_GEMINI_API_KEY, already in Coolify), so it is
container-safe — unlike the halacha path (claude CLI, host-only).
apply_to_record fills only empty fields, so the deterministic seeds and
any chair-curated subject_categories are preserved. Result surfaced in the
upload response (out["metadata"]); failures logged, upload still succeeds.

Also corrects two now-stale comments: the "Gemini returns no_metadata for
internal decisions" note in the enroll loop, and the "MCP-tool-only path"
docstring on reextract_metadata (true for halacha, not for Gemini REST).

Invariants: G1 (fill at source on enroll, not a read-time workaround),
G2 (reuse reextract_metadata — no parallel extraction path). LLM call is
Gemini REST, not claude_session, so the container LLM-call constraint holds.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 10:48:46 +00:00
0bf787929c Merge pull request 'fix(renumber): update Paperclip issue linkage on case renumber' (#334) from worktree-renumber-paperclip-linkage into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 11s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-28 10:20:00 +00:00
8932bd1f54 fix(renumber): update Paperclip issue linkage on case renumber
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
renumber_cases.py step 6 only renamed the Paperclip project, leaving the
case↔issue linkage on the OLD case number. get_case_issues() keys on two
text surfaces the script never touched:
  • plugin_state.legal-case-number (value_json) — the authoritative linkage
  • issues.title — the '[ערר {cn}] …' tag the title-path lookup matches

Result: after a renumber the issues kept the old number, get_case_issues
returned [], and post-final actions (run-learning / run-halacha) silently
skipped with reason "no_issue" — surfaced in the UI as "לא הופעלה למידה".
Discovered on case 8137-11-24 (renamed from 8137-24).

Now the Paperclip step rewrites all three surfaces (projects.name +
plugin_state + issue titles); inspect_paperclip + the dry-run report show
the linkage-row and issue-title counts that will be rewritten.

The 11-case migration already ran; the matching DB rows were fixed
manually (incl. this commit's logic) for all stale cases. This change is
for any future renumber.

Invariants: G1 (normalize at source — the linkage, not a read-time
workaround). G12 note: this is a one-time host-side migration script that
already connects directly to the Paperclip DB by design (below the platform
port, like the existing projects/MinIO/Gitea surgery) — it does not add a
parallel runtime path through agent_platform_port.py.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 10:19:24 +00:00
4de555367d chore: fix .gitignore inline comments + add untracked code files
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m32s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 11s
Inline # comments in gitignore are not supported — they were silently
breaking three patterns (data/checkpoints/, data/adapter-migration-state.json,
.claude/agents/.generated/). Moved comments to their own lines and added
missing entries for runtime dirs (data/audit/, data/logs/, etc.) and
temp files (.interaction_tmp.json, .design-build/, .taskmaster bak files).

Also tracks previously untracked legitimate files: scripts, tests, docs,
skills references, .env.example, taskmaster templates.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 12:49:32 +00:00
b4a68cf5da feat(ocr): replace Google Vision with Mistral OCR as PDF fallback
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m27s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-27 10:15:13 +00:00
9ae7304d44 feat(ocr): replace Google Vision with Mistral OCR as PDF fallback
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
Switches the scanned-PDF fallback from Google Cloud Vision to
Mistral OCR (mistral-ocr-latest) for better Hebrew accuracy and
robustness against broken embedded OCR layers (e.g. case 1044-03-26
which returned English garbage through Vision).

Routing strategy (document-level, not per-page):
- PyMuPDF extracts all pages; pages that pass _text_quality_ok()
  use PyMuPDF output directly (free, ~50ms).
- If ANY page fails quality → Mistral OCR called once for the whole
  PDF, returning per-page Markdown for all pages (consistent format,
  no plain-text/Markdown mix within a document).

Markdown output preserved: Mistral returns ## headers and |tables|;
chunker updated to recognise ATX Markdown headers (##/###) as section
boundaries in _split_into_sections().

Config: GOOGLE_CLOUD_VISION_API_KEY → MISTRAL_API_KEY; allowlist
updated vision.googleapis.com → api.mistral.ai.
MISTRAL_API_KEY added to Coolify container env.

Invariants: G1 (single OCR fallback path, not parallel), G2 (no
duplicate extractor route), INV-AH (Mistral handles gershayim
natively; quote-fix only on PyMuPDF path).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 10:11:45 +00:00
f2a264a7da Merge pull request 'fix(status): add qa_passed to legacy display labels' (#332) from worktree-qa-passed-status into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 45s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-24 15:04:16 +00:00
e051fda0cb fix(status): add qa_passed to legacy labels + migrate 8174-12-24 to drafted
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
qa_passed was set by the old agent pipeline but not included in the
trimmed 10-status CASE_STATUSES list or LEGACY_STATUS_LABELS, causing
the status badge and workflow timeline to render nothing ("לא ידוע").

Added qa_passed → "טיוטה" to LEGACY_STATUS_LABELS as a display-only
fallback so any case still carrying this value renders correctly until
migrated. Case 8174-12-24 status updated to drafted via PUT /api/cases.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 15:04:00 +00:00
331b8c4249 Merge pull request 'feat(agents): reset button + fix comment routing fallback' (#331) from worktree-agent-reset into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 47s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-24 14:54:25 +00:00
82844a63c2 feat(agents): reset button + endpoint to clear stuck agent error state
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 5s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
Adds a one-click 'reset agents' action for cases where writer/QA agents
are stuck in Paperclip's error state (triggered by recovery loop or
failed run). Addresses the root cause documented in
reference_paperclip_recovery_loops: reassigns open issues from agents
back to the chair user, and calls reset_agent_session for each error
agent to clear wedged runtime sessions.

Changes:
- paperclip_client.py: reset_case_agents() — reassigns stuck issues +
  clears error status in Paperclip DB + calls reset_agent_session
- agent_platform_port.py: exports pc_reset_case_agents (G12 gate)
- app.py: POST /api/cases/{case_number}/agents/reset endpoint
- agents.ts: useResetCaseAgents mutation hook + AgentResetResult type
- agent-status-widget.tsx: 'אפס' button (shown only when error agents
  exist) with Dialog confirmation + loading state + toast feedback

Invariants: G12 (Paperclip only via port), G2 (no parallel path —
uses existing reset_agent_session + pc_get_case_issues).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 14:53:50 +00:00
a9d04c1e9f Merge pull request 'fix(arguments): validate claim_ids + per-row savepoint in argument_aggregator (#156)' (#330) from worktree-fix-argument-aggregator-savepoint into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 3m12s
G12 Leak-Guard / leak-guard (push) Successful in 3s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-24 12:59:33 +00:00
c9d83431e0 fix(arguments): validate claim_ids + per-row savepoint in argument_aggregator
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
aggregate_claims_to_arguments crashed with "current transaction is aborted,
commands ignored until end of transaction block" on large cases (confirmed on
1027-04-26, 195 claims; reported via CMP-186).

Root cause: the proposition INSERT (legal_argument_propositions) was wrapped in
a broad except Exception with no savepoint. When the LLM echoes a
syntactically-valid-but-nonexistent claim_id, the FK violation
(legal_argument_propositions_claim_id_fkey) puts the asyncpg transaction into
aborted state. The except caught only that INSERT's error but never issued
ROLLBACK TO SAVEPOINT, so the next statement (the following argument's INSERT
into legal_arguments) raised InFailedSQLTransactionError uncaught and crashed
the whole call. With many claims the bad-UUID probability is high -> consistent
failure.

Fix:
- Validate each claim_id against the known set of claim ids fetched for the
  case before INSERT, so a hallucinated id never reaches the DB (G1: fix at
  source). Malformed UUIDs are already dropped in _normalize_argument; this
  catches the valid-but-nonexistent ones.
- Wrap the INSERT in a per-row savepoint (async with conn.transaction()) as
  defense-in-depth, so any unexpected constraint failure rolls back to the
  savepoint instead of poisoning the surrounding transaction.

Invariants: G1 (normalize at source, not symptom-catch on read). No parallel
path (G2). No silent swallow — skips are logged.

TaskMaster: legal-ai #156

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 12:58:03 +00:00
b57cd17408 Merge pull request 'fix(case-page): resolve full-page review findings (bugs + RTL + resilience)' (#329) from worktree-case-page-fixes into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 43s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 18:58:21 +00:00
2b591f5018 fix(case-page): resolve full-page review findings (bugs + RTL + resilience)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
Full code review of the case-detail page (14 components) surfaced these,
all fixed here:

Logic bugs
- case-edit-dialog: form.reset ran on the 5s useCase refetch while the dialog
  was open, clobbering in-progress edits. Now resets only on open→true.
- status-changer: `selected` never synced to async/external `currentStatus`
  (stale dropdown). Reworked to track currentStatus until an explicit pick;
  resets to tracking after save.
- decision-blocks-panel: `block.content`/`word_count` accessed without null
  guards (endpoint has no response model) → potential render crash. Coerced
  with `?? ""` / `?? 0`. `STATUS_LABELS[status]` now falls back to the raw
  status instead of rendering literal "undefined".
- document-type-editor: `await mutateAsync()` in async click handlers without
  try/catch → unhandled promise rejection. Wrapped (errors still surface via
  isError).

Resilience / hygiene
- page.tsx: a transient 5xx on the background poll flipped the WHOLE page to
  the error card and discarded loaded data. Now gated on `!data`, plus a
  "נסה שוב" retry.
- cases.ts useUpdateCase: invalidated casesKeys.all, which re-invalidated the
  detail it had just optimistically patched. Scoped to the list prefix.

RTL correctness (logical properties)
- agent-status-widget `mr-auto`→`ms-auto`; agent-activity-feed `mr-auto`→
  `me-auto`, icon `ml-1/ml-2`→`me-1/me-2`, required `*` `mr-1`→`ms-1`;
  document-type-editor list `pr-4`→`ps-4`.

Minors
- drafts-panel: `<a href>`→`next/link` (operations + citation links) for SPA
  nav. agent-activity-feed: issueMap memoized; comment Textarea aria-label.
  upload-sheet: `unknown` status no longer shown as green success (neutral
  icon + "רענן לאישור"). citations.ts: case_name typed `string | null`.

Design gate: visual-touching items (RTL gap side, retry button, neutral
upload icon) were chair-authorized via the reviewed-findings approval ("fix
all"); none alter an approved page layout — they are correctness fixes.
tsc clean; eslint clean (1 pre-existing form.watch warning, untouched line).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:57:43 +00:00
148b4b9bf6 Merge pull request 'feat(arguments): inline status banners + fix double-card nesting' (#328) from worktree-arguments-banners into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 43s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 18:36:21 +00:00
31029b2d43 feat(arguments): inline status banners + fix double-card nesting
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
Implements the Claude-Design-approved inline status banners for the
"חשב טיעונים" panel (mockup 25-legal-arguments-panel): the four endpoint
states (queued / exists / no_claims / skipped) now render as tone-coded
banners below the header instead of a transient toast. Toast is kept only
for hard transport errors.

Also fixes a pre-existing double-card bug found while reviewing the page:
the page's "arguments" tab already wraps the panel in <Card><CardContent>
(page.tsx:151-155), yet LegalArgumentsPanel rendered its OWN <Card> too —
unlike its sibling tab panels (DecisionBlocksPanel/DraftsPanel/
AgentActivityFeed) which render a plain <div>. The panel now matches that
convention (single card, no double border/padding), consistent with the
approved single-card mockup.

Design gate: card 25 approved by חיים. tsc + eslint clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:35:55 +00:00
c9970a5955 Merge pull request 'fix(operations): contain agent run-log text inside the "פלט" popup' (#327) from worktree-fix-runlog-overflow into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 43s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 18:32:17 +00:00
38234d9b4f fix(operations): contain agent run-log text inside the "פלט" popup
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
The RunLogDialog rendered the live agent log inside a Radix ScrollArea.
Radix wraps Viewport children in an inner `display:table` div that
shrink-wraps to the widest unbreakable token — the long `toolu_…` IDs and
`/home/chaim/…` file paths in the raw JSON log. That table cell expands
instead of constraining width, so `whitespace-pre-wrap break-words` never
had a width to wrap against and the text spilled out past the dialog
borders.

Replace the ScrollArea with a plain width-constrained scrollable div and
switch break-words → break-all so the long unbreakable IDs/paths wrap too.
Pure presentational fix (no invariant surface).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:31:17 +00:00
a0b158b2c8 Merge pull request 'fix(arguments): route "חשב טיעונים" through the legal-analyst agent' (#326) from worktree-fix-aggregate-btn into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 41s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 18:20:48 +00:00
a3df05e067 fix(arguments): route "חשב טיעונים" through the legal-analyst agent
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The /aggregate-arguments endpoint ran an in-container BackgroundTask that
called claude_session (the local `claude` CLI) — which does not exist in the
FastAPI container. The button silently produced nothing, and on `force` it
destructively DELETEd existing arguments *before* the doomed LLM call.

Replace the inline task with the established delegation pattern used by
"חלץ עובדות שמאיות" (extract-appraiser-facts): a cheap in-container DB
pre-check (no_claims / exists), then a Paperclip wakeup of the company's
legal-analyst, which runs mcp__legal-ai__aggregate_claims_to_arguments
locally (where the CLI lives) and reports back. `force` now runs locally too,
so delete+recompute are atomic on the host — no more destructive failure.

Frontend: AggregateArgumentsResult becomes a discriminated union
(queued | no_claims | exists | skipped) and the toast is status-accurate
instead of the misleading fixed "refresh in a minute".

Invariants: G12 (Paperclip touch confined to paperclip_client behind the
agent_platform_port), G2 (replaces the broken path, no parallel capability),
engineering §6 (explicit statuses, no silent swallow).

UI change is logic/toast only (no visual-layout change) — within the
Claude-Design-gate bug-fix exemption. Richer inline status panels deferred
to the gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:19:57 +00:00
a3e612f22b Merge pull request 'feat(citation-verify): frontend "אימות פסיקה" tab (#154)' (#325) from worktree-citation-verify-frontend into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 42s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 18:19:28 +00:00
5108c854cf feat(citation-verify): frontend "אימות פסיקה" tab in the decision editor (#154)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
Frontend half of the citation-verification panel (backend in #323), as the agreed
third tab in /compose — matching the approved X17 mockup 24-citation-verification:

- lib/api/citation-verification.ts: hand-written types + useCitationVerification
  query + useVerifyCitation mutation (POST verify, invalidates the view).
- components/compose/citation-verification-panel.tsx: per legal argument →
  supporting corpus precedents with the cited_by authority chips (אומץ ×N /
  אובחן ×N), the exact supporting_quote, ✓ מאמת / ✗ לא רלוונטי verify actions, a
  per-citation chair-note field, and the per-issue 📡 radar (unlinked digests,
  pointer-only). Summary strip + INV-AH/INV-DIG1 reminders.
- compose page: third tab "אימות פסיקה" alongside עורך הבלוקים / עמדות וטענות.

No api:types regen needed (endpoints return assembled dicts; types hand-written
per the lib/api convention). tsc + eslint clean.

Invariants: INV-AH (writer cites only verified), INV-DIG1 (digest never cited),
G2 (consumes the one backend view). Visual matches the gate-approved mockup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:19:03 +00:00
d7855f6284 Merge pull request 'fix(ui): case header — metadata parallel to title, parties clamped to 2 lines' (#324) from worktree-case-header-layout into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 43s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 18:11:55 +00:00
80809ca406 fix(ui): case header — keep metadata parallel to title, clamp parties to 2 lines
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
The hearing-date/updated/sync metadata dl dropped below the title (taking rows and
pushing the band down) when the appellants/respondents line was long, because the
outer row used flex-wrap. Fix:

- Drop flex-wrap on the title↔metadata row (and flex-1 on the title block) so the
  metadata dl stays parallel to the H1; the title block shrinks instead.
- Clamp the parties line (and the subject fallback) to 2 lines with line-clamp-2 +
  title tooltip, so a long party list no longer grows the band.

Chair-directed layout fix; existing components, no new design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:11:32 +00:00
abe4c53df1 Merge pull request 'feat(citation-verify): backend for "אימות פסיקה" — per-argument support + verify gate (#154)' (#323) from worktree-citation-verify-backend into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m29s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 18:08:46 +00:00
5ede8a9653 feat(citation-verify): backend for the "אימות פסיקה" panel — per-argument support + verify gate (#154)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 12s
Backend half of the citation-verification tab (frontend follows in a separate PR):

- Schema V44: case_precedents gains argument_id (the legal argument it supports),
  case_law_id (the corpus ruling — for cited_by + dedup), verified (the INV-AH gate
  the writer respects) + verified_at. All nullable; chair_note already existed.
- db: create_case_precedent(+argument_id/case_law_id/verified),
  set_case_precedent_verified(id, verified, chair_note), list_case_precedents(+cols).
- services/case_citation_verification.build_view(case): per legal_argument →
  in-corpus suggestions (search_library per issue) each with the cited_by authority
  breakdown (db.citation_authority, X11) + merged attach/verify state + per-issue
  radar (case_digest_radar grouped by matched issue). Pure read/assembly; reuses the
  one corpus search + one authority query + one radar (G2).
- endpoints: GET /api/cases/{n}/citation-verification (the view) and
  POST .../verify (upsert verify/un-verify + chair note).

Validated on 1044-03-26: 14 arguments, all with support; 317/10 surfaces with
אומץ×11/אובחן×1; radar leads attributed per issue. Auto-approval untouched (chair
gate, INV-G10). Verify defaults to false — nothing is authoritative until the chair
marks it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 18:08:21 +00:00
33c10e4147 Merge pull request 'fix(ui): nav & layout — decision-editor in band, compose back-link, wider cases table' (#322) from worktree-case-nav-fixes into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 53s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 17:56:44 +00:00
81050181d7 fix(ui): nav & layout — decision-editor in band, compose back-link, wider cases table
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
Chair-directed navigation/layout fixes (existing components, no new design):

1. "פתח עורך החלטה" moved into the case-page band actions (right after "העלאת
   מסמכים") so it's reachable from EVERY tab, not only סקירה. Removed the now-
   redundant full-width CTA from the overview tab.
2. Prominent "→ חזרה לדף התיק" back-link added to the /compose header (the
   breadcrumb link was too subtle).
3. Home cases table: rail trimmed 360→280px and the title cell made min-w-0 so the
   table gets the width it needs and no longer shows a horizontal scrollbar.

No backend / API change. The larger NEW citation-verification panel stays gated
behind Claude Design (preview 24-citation-verification already pushed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 17:56:21 +00:00
00c8083cc9 Merge pull request 'feat(digests): calibrate case radar to distilled issues + per-issue attribution' (#321) from worktree-digest-radar-calibrate into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m27s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 17:07:58 +00:00
21ff52aff9 feat(digests): calibrate case radar to the analyst's distilled issues + per-issue attribution
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 12s
The radar's query was built from dozens of raw claims (procedural-heavy noise), so
matches were thematic but imprecise and gave no reason WHY a lead is relevant. Now:

- Prefer the analyst's distilled legal_arguments (argument_title + legal_topic — one
  crisp CREAC issue per row) over raw claims.
- Search EACH issue separately and MERGE, so every lead is attributed to the case
  issue(s) it answers (`matched_issues`) — the chair sees "this ruling is for your
  'זכות עמידה' issue", not just a blended score.
- Fall back to the raw-claims blended query pre-aggregation; `source` reports the path.
- Shared `_radar_enrich` helper (gap status + action + matched_issues), bounded to 25
  issues to cap the per-issue fan-out.

Validated: 8124-09-24 (32 args → per-issue) surfaces betterment rulings each tagged to
its issue (היעדר השבחה / זהות הנישום / סעיף 7(ב)); 1044-03-26 (0 args) falls back to
claims unchanged. No tool/endpoint signature change (new fields pass through the dict).

Invariants: G2 (reuses the one digest search + arg accessor), INV-DIG1 (radar only).
No schema change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 17:07:15 +00:00
cc8fd0b853 Merge pull request 'feat(digests): case-contextual digest radar — surface unlinked digests as chair leads (X12)' (#320) from worktree-digest-radar into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m53s
G12 Leak-Guard / leak-guard (push) Successful in 7s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 16:45:48 +00:00
70f93c3bd4 feat(digests): case-contextual digest radar — surface unlinked digests as chair leads (X12)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 5s
Lint — undefined names / undefined-names (pull_request) Successful in 14s
A digest pointing at a ruling we don't hold yet ("unlinked") was captured globally
(missing_precedents inbox) but never surfaced IN THE CONTEXT of the case being
decided — so a relevant ruling known only via a digest could fall through the cracks
at the moment it matters. Adds the case-contextual radar:

- db.search_digests_semantic: new `linked_only` filter (False = unlinked-only target set).
- digest_library.case_digest_radar(case_number): builds the case topic from title +
  appeal_subtype + the analyst's claims, embeds once, matches against UNLINKED digests,
  and returns leads enriched with the underlying ruling's gap status + suggested action
  (new_lead / gap_open / fetched / available_link).
- MCP tool `digest_radar` + endpoint GET /api/cases/{n}/digest-radar (for the agent and
  the future case-page lead).

INV-DIG1 preserved: radar only — every lead points at the underlying RULING (fetch /
upload / link), never cites the digest. Read-only.

Validated on 8124-09-24 (היטל השבחה): 5 on-topic unlinked digests, scores 0.64–0.72.
The visible case-page panel is a separate UI change → goes through the design gate.

Invariants: G2 (reuses the one digest semantic-search + gap/citation resolvers),
INV-DIG1 (no digest citation), INV-DIG3 (gap surfacing). No schema change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 16:45:14 +00:00
4f45fa416b Merge pull request 'feat(x11): treatment-aware citation authority wired into research agents (#154)' (#319) from worktree-x11-treatment-wiring into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m26s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 16:04:36 +00:00
ccc5a73bc8 feat(x11): treatment-aware citation authority wired into research agents (#154)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The internal citation graph fed only RANKING (raw in-degree), and the per-citation
TREATMENT was never classified — so a precedent distinguished N times got the same
authority boost as one followed N times (INV-COR2 violation), and the signal never
reached the agents' reasoning. Wires the full path:

Phase 1 — scripts/classify_citation_treatments.py: classify each linked edge's
  treatment (followed/distinguished/…) from its match_context via
  corroboration.classify_treatment (Opus 4.8 @ xhigh, local), filling
  precedent_internal_citations.treatment. Idempotent.
Phase 2 — db.refresh_verified_layer: count only NON-negative treatments toward
  verified/cite_count (INV-COR2/COR4). Unclassified counts as neutral-positive so
  the signal degrades gracefully before classification runs.
Phase 3 — db.citation_authority(ids): per-precedent {total, positive, negative,
  unclassified, by_treatment}. Surfaced as `cited_by` in search_precedent_library
  hits and precedent_library_get, and `treatment` per incoming citation.
Phase 4 — legal-researcher/analyst/writer prompts: weigh & ARGUE authority
  ("הלכה שאומצה ב-N החלטות ועדת-ערר"), flag distinguished/overruled, never invent
  the count (INV-AH; writer is read-only of the analyst).

Auto-approval stays kill-switched off (chair gate preserved, INV-G10). No schema
change (treatment column already existed). Operational: run the classifier +
refresh_verified_layer over the 379 edges, then sync agents across companies.

Invariants: G2 (one classifier + one authority query, reused), INV-COR2/COR3/COR4
(negative never corroborates; point-specific; ≥N), INV-G10 (no auto-approval),
INV-AH (no invented numbers).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 16:04:04 +00:00
b8c49a1269 Merge pull request 'fix(precedents): מראה-מקום never blank — seed at upload + inline Gemini enrichment' (#318) from worktree-citation-autofill-root into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m28s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 15:30:40 +00:00
b0bcdbeeef fix(precedents): מראה-מקום never blank — seed at upload + inline Gemini enrichment (root cause)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
Root cause: the upload form "מראה המקום" field is stored only as case_number;
citation_formatted was left empty and filled ONLY by the async metadata
extraction, which runs in the local drainer (not the container) because it was
lumped with halacha extraction. Result: the chair saw an empty מראה-מקום during
the window and re-typed it by hand.

Fix (hybrid — Option C):
- Seed: _create_external_record / create_external_case_law store the chair's typed
  citation as citation_formatted at INSERT, so the field is never blank. A
  cited_only→external promotion preserves an existing non-empty value (prior edit).
- Enrich: ingest_precedent runs metadata extraction INLINE in-container
  (gemini_session is direct REST, no local CLI) with force_citation=True, upgrading
  the seed to the canonical derived citation (parties + reporter + date) before the
  chair opens the page. Best-effort: on no key / API failure the seed remains and
  the queued metadata drain stays as the fallback. Halacha extraction is untouched
  (stays local — claude_session).
- apply_to_record / extract_and_apply gain force_citation: re-assemble even when
  citation_formatted is non-empty, writing only when assembly SUCCEEDS (seed
  preserved on abstention). The drainer keeps the default (False) so a chair's
  manual edit in /precedents/[id] is never clobbered.

Ops: GOOGLE_GEMINI_API_KEY added to the legal-ai Coolify app (runtime) — the SoT
value from Infisical nautilus:/external-apis/gemini, validated against the API.

Invariants: G2 (one citation-resolution + one metadata-extraction path, reused —
no parallel logic), INV-ID2/X1§3 (citation_formatted stays a derived field; the
seed is provisional and upgraded), INV-G10 (chair edits preserved). No schema change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 15:30:00 +00:00
2c515966c5 Merge pull request 'feat(precedents): surface auto-detected incoming citations in the "ציטוטים מקושרים" panel' (#317) from worktree-incoming-citations-panel into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m30s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 14:55:01 +00:00
91c521922f feat(precedents): surface auto-detected incoming citations in the "ציטוטים מקושרים" panel
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The precedent-detail "ציטוטים מקושרים" panel rendered only MANUAL appeal-chain
links (case_law_relations), so it stayed empty even when decisions cite the
ruling — the automatic citation graph (precedent_internal_citations) was never
surfaced. e.g. עע"מ 317/10 (שפר) has 12 incoming citations in the DB, none shown.

Fills the existing panel from the citation graph (data/logic only, no new
visual design — gate-exempt per web-ui/AGENTS.md):
- citation_extractor.list_citations_to_case_law: return source precedent_level/
  court/date too (additive; reuses the one canonical incoming query — G2).
- precedent_library.get_precedent: add incoming_citations[], shaped like the
  RelatedCase row. No parallel resolution path.
- web-ui: render incoming citations in the same row style, read-only (no unlink),
  de-duped against manual links; manual "קשר" linking preserved alongside.

Invariants: G2 (single citation-resolution path, reused), G1 (graph self-heals
via existing relink_orphan_citations on upload). No schema change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 14:54:12 +00:00
ad29f6033f Merge pull request 'fix(graph): רצפת מסנן-השנה ל-1980 + תקרה דינמית' (#316) from worktree-graph-year-floor into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 45s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 14:36:29 +00:00
9d4960f28f fix(graph): lower year-filter floor to 1980, dynamic ceiling
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The /graph "משנה / עד שנה" dropdown was hardcoded to 1994–2026, so dated
precedents below 1994 were unreachable by the year filter. After backfilling
decision dates, the corpus now has precedents from 1982 (ע"א 725/81), 1988
(910/86) and 1990 (בג"ץ 1578/90).

Floor → 1980 (covers the oldest); ceiling → current year via getFullYear()
so the range never ages out. Pure UI-logic change — same year_from/year_to
params, no new data path (G2), no backend touch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 14:35:49 +00:00
38b3ffc587 Merge pull request 'feat(corpus): עיצוב-מחדש קורפוס-הפסיקה — ביטול תור-ההלכות, שכבת-מאומת-מאזכורים, דירוג-בזמן-אחזור (#153)' (#315) from worktree-canonical-synthesis into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m32s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
Merge PR #315: corpus redesign — no queue, verified-by-citation, rank-at-retrieval (#153)
2026-06-20 13:55:49 +00:00
2ae68c5896 Merge pull request 'fix(missing-precedents): load full set so accordion counts match the header' (#314) from worktree-mp-loadall into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 42s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 13:35:05 +00:00
dd0312e457 fix(missing-precedents): load full set so accordion counts match the header
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
The table fetched limit:200 while the header shows the true open count (~497,
via a separate COUNT). With accordion grouping this became visible: sections
summed to exactly 200 (e.g. 16 chair / 173 digest / 11 other) — and 90 of the
106 committee rows (the high-value "cited by Dafna" ones) were hidden past row
200. True buckets: chair 106 / digest 328 / other 63.

- web-ui: list limit 200 → 1000 so all open rows load; accordion section counts
  (computed from the loaded set) now equal the header total.
- backend: list cap 500 → 2000 to allow it (response stays a few hundred KB).

Logic/data only — no visual change (bug-fix exception to the design gate).
tsc --noEmit clean; py_compile clean.

Follow-up: when the backlog (digests grow daily) approaches the cap, move to
server-side per-bucket counts + pagination/virtualization (design gate).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 13:34:36 +00:00
eaf80dbe9a Merge pull request 'fix(citations): extract internal citations from discussion sections only (G1/G2)' (#313) from worktree-citation-sections into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m30s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 13:29:42 +00:00
f63bd4df0f fix(citations): extract internal citations from discussion sections only (G1/G2)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
extract_and_store read the WHOLE full_text, so a ruling a party cited in its
own arguments (block ז / appellant_claims / parties_claims) was recorded as the
deciding committee citing it — and surfaced as "צוטט ע״י דפנה". Example: ערר
130/10 appears only in the appellant_claims chunk of decision 1200-12-25, yet
was attributed to the chair.

Fix: build the extraction text from the discussion/ruling chunks via the
halacha extractor's _select_extractable_chunks (G2 — one definition of
"reasoning sections": legal_analysis/ruling/conclusion + discussion-anchor,
excluding facts/claims/intro). Falls back to full_text only for un-chunked rows.
A citation now reflects the deciding body's reliance, not a party's argument.

Existing data reconciled via a one-off (backup
data/audit/citation-edges-backup-*.csv): per source decision, dropped edges
whose cited number is absent from the discussion-only set — 82 of 398 removed
(incl. ערר 130/10 ↔ 1200-12-25); 316 genuine edges kept.

Invariants: G1 (correct at the source) · G2 (single reasoning-section selector,
shared with halacha extraction) · INV-LRN2 (quality-at-source).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 13:29:14 +00:00
25b41af6a3 Merge pull request 'feat(missing-precedents): accordion grouping by discovery source' (#312) from worktree-mp-accordion into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 43s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 12:53:41 +00:00
979ec17a45 feat(missing-precedents): accordion grouping by discovery source
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
Chair-approved (Claude Design card 09): split the flat table into collapsible
sections by where the gap surfaced, so the high-value "cited by Dafna" rulings
aren't buried among the 330 digest rows.

- Three <details> accordion sections within the active status tab:
  · "צוטט ע״י דפנה" — rows with a resolved chair from the citation-graph bridge
    (cited_by_chairs non-empty); open by default.
  · "יומון" — discovery_source='digest'; collapsed.
  · "אחר" — party-cited / manual remainder; collapsed.
- Row JSX extracted to renderRow() and reused across sections; shared
  TableHeaderRow(); grouping via useMemo over the current result set.
- Empty sections are hidden; section header shows source chip + count.

No backend change — groups on existing cited_by_chairs / discovery_source.
tsc --noEmit clean.

Invariants: INV-IA* (task-oriented surface, single gate) — UI only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 12:53:08 +00:00
eb0182ecf8 Merge pull request 'fix(graph): normalize case_law subject_tags to underscore convention at write chokepoint' (#311) from worktree-subject-tag-normalize into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m30s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 12:39:50 +00:00
a401197204 fix(graph): normalize case_law subject_tags to underscore convention at write chokepoint
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
Topic hubs in /graph are built from case_law.subject_tags. The documented
extraction contract (precedent_library tool examples: קווי_בניין,
מועד_קביעת_שומה) and ~99% of the corpus store plain multi-word Hebrew tags
with underscores between words ("היטל_השבחה"). A single case (8126-03-25,
יעקב עמיאל) was tagged with spaces ("היטל השבחה"), which the graph renders as
a SECOND, distinct topic hub — a duplicate of the underscore form. The data
was normalized separately; this enforces the convention at the source so no
write path can re-introduce the split.

_normalize_subject_tags() is applied at the three (and only) case_law write
chokepoints in db.py — create_external_case_law, create_internal_committee_decision,
update_case_law — so the rule cannot be bypassed (G1: normalize at source, not
in the read/graph path). Tags carrying punctuation/digits/dashes
(e.g. "פטור מותנה — סעיף 19(ג)") are left untouched; only plain Hebrew word
phrases ([א-ת]+ separated by spaces) are converted. Also dedups post-normalize.

digests.subject_tags (TEXT[], a different graph layer) and canonical_halachot
are intentionally out of scope.

Invariants: maintains G1 (fix at source, not in the read projection),
G2 (graph_api stays a pure read projection — no parallel normalization there).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 12:39:06 +00:00
688de7f842 Merge pull request 'feat(missing-precedents): cited-by-chair bridge + open/closed tabs + single-line citation' (#310) from worktree-missing-prec-page into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m28s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 12:29:10 +00:00
a18ed8ffb7 feat(missing-precedents): bridge cited-by-chair + decision number; open/closed tabs; single-line citation
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
Implements the chair-approved redesign of /missing-precedents (Claude Design
card 09). The "צוטט ע״י" and "תיק" columns were empty for ~98% of rows because
the corpus-decision reliance lives in precedent_internal_citations (keyed to
case_law), a different system than missing_precedents (keyed to cases).

Backend (G2 — read-time bridge, no stored duplication):
- list_missing_precedents: new cited_by_chairs / cited_by_decisions columns,
  resolved from the citation graph by normalized docket number (same
  normalization the relinker uses). For an open gap cited by committee
  decisions, surfaces the chair(s) (e.g. דפנה תמיר) + the deciding case numbers.

Frontend (matches approved mockup):
- "צוטט ע״י" shows the chair chip (bridge) with priority over the generic
  discovery-source chips; "תיק" shows the deciding decision number(s) + "+N".
- Status filter reduced to פתוח / נסגר (removed הועלה, לא-רלוונטי, הכל).
- "פסיקה" column collapses to a single line when case_name is empty — fixes the
  duplicated-citation two-row rendering.
- Lifecycle note simplified to open → closed.

Verified: bridge SQL returns chair+decisions on live data; tsc --noEmit clean;
py_compile clean.

Follow-up (data, separate): LLM backfill of legal_topic (נושא); create
missing_precedents rows for the 28 graph-cited rulings not yet tracked.

Invariants: G2 (single source of truth — bridge derived at read, no duplicated
column) · INV-AH context (closing source gaps).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 12:28:38 +00:00
086913ce8d Merge pull request 'fix(corpus): re-link orphan citations when a cited ruling is uploaded (G1)' (#309) from worktree-citation-relink into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m26s
G12 Leak-Guard / leak-guard (push) Successful in 5s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 11:47:17 +00:00
f7bf437f67 fix(corpus): re-link orphan citations when a cited ruling is uploaded (G1)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
precedent_internal_citations resolves cited_case_law_id only at extraction
time. When a decision cites a ruling that isn't in the corpus yet, the edge is
stored with cited_case_law_id=NULL. The ruling is often uploaded days later —
but the orphan edge was never re-resolved, so the citation graph permanently
under-counted real connectivity (same stale-derived-value pattern as the
searchable flag). Worse, external rulings store case_number WITHOUT the court
prefix ("3213/97") while citations carry it ("ע\"א 3213/97"), so a large share
of "missing" citations were actually present-but-unlinked.

Fix:
- new citation_extractor.relink_orphan_citations(case_law_id=None): re-resolves
  orphan edges via the canonical _resolve_case_law_id (no parallel matching, G2);
  scoped to one row, or full sweep when None. Idempotent.
- ingest_document calls it on every upload (non-fatal) so the graph self-heals.

Existing backlog already re-linked via a one-off sweep against the live DB:
linked 128 edges → in-corpus distinct rulings 67→178, unlinked distinct
262→143. So real "cited but not uploaded" ≈ 143, not the 262 the stale graph
implied.

Invariants: G1 (re-resolve the derived link at the write) · G2 (single resolver).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 11:46:47 +00:00
473c54fc43 Merge pull request 'fix(corpus): refresh searchable flag on metadata patch (INV-DM1, G1)' (#308) from worktree-searchable-freshness into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m31s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 11s
2026-06-20 11:36:26 +00:00
92f0c7d208 fix(corpus): refresh searchable flag on metadata patch (INV-DM1, G1)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
The `searchable` completeness flag is recomputed only at ingest end and after
the Gemini metadata extractor. Internal-committee decisions skip Gemini, so
when their summary/subject_tags/metadata are filled after ingest the flag goes
stale and the row stays invisible to RAG despite being complete. Found 4 such
decisions (1049-06-21, 8126-03-25 [יעקב עמיאל], 8181-21 [האוניברסיטה העברית],
85074-04-25) — all complete (chunks+embeddings, extraction+halacha completed,
metadata present) yet searchable=false.

Fix: recompute_searchable() at the end of update_case_law — the single
chokepoint for metadata patches — so the derived flag never drifts from the
content (G1). Existing stale rows already corrected via a one-off canonical
recompute_searchable(None) run (corpus: 328 -> 332 searchable).

Invariants: INV-DM1 (completeness contract) · G1 (normalize derived value at
the write, no read-time patch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 11:35:59 +00:00
422d79f9e8 Merge pull request 'fix(corpus): חיפה is its own district, not folded into צפון (G2)' (#307) from worktree-haifa-district into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m32s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 11:25:46 +00:00
fb32c278e6 fix(corpus): חיפה is its own district, not folded into צפון (G2)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 3s
Lint — undefined names / undefined-names (pull_request) Successful in 10s
Haifa (מחוז חיפה) and the North (מחוז הצפון) are separate Israeli planning
districts, each with its own appeals committee. Two divergent bugs collapsed
Haifa into צפון:
- services/internal_decisions.py _VALID_DISTRICTS omitted "חיפה" (while the
  tools-layer VALID_DISTRICTS and db.py citation-abbrev map both include it —
  a G2 single-source-of-truth divergence)
- _COURT_TO_DISTRICT mapped "חיפה" -> "צפון", so Haifa courts derived צפון

Result: 2 Haifa committee decisions (1074-08-23, 8508-03-24) were mis-filed
under צפון, and "חיפה" never appeared as a district.

Changes:
- add "חיפה" to service _VALID_DISTRICTS
- _COURT_TO_DISTRICT: ("חיפה","צפון") -> ("חיפה","חיפה")
- SCHEMA_V43: reclassify legacy internal rows whose court mentions חיפה from
  צפון -> חיפה (normalize at source, G1; idempotent)

Verified against live DB (rollback txn): 2 rows move צפון(6)->צפון(4)+חיפה(2).

Invariants: G1 (normalize at source) · G2 (single source of truth for the
valid-district set — follow-up: consolidate the 3 divergent lists).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 11:25:20 +00:00
9e0f483e83 Merge pull request 'fix(corpus): committee decisions are persuasive, never binding (INV-DM7)' (#306) from worktree-committee-binding into main
All checks were successful
Build & Deploy / build-and-deploy (push) Successful in 1m35s
G12 Leak-Guard / leak-guard (push) Successful in 4s
Lint — undefined names / undefined-names (push) Successful in 10s
2026-06-20 10:53:27 +00:00
5a69980adf fix(corpus): committee decisions are persuasive, never binding (INV-DM7)
All checks were successful
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 9s
ועדת-ערר decisions (source_kind=internal_committee / source_type=
appeals_committee) are persuasive authority — they do not bind another
committee, nor the committee itself. The stored is_binding column wrongly
defaulted to True across the FastAPI form + service/db layers, so 46 of 92
committee rows were marked binding and got the BINDING halacha-extraction
prompt. Authority is structural for this source (INV-DM7) — normalize at the
source (G1), not trust the input.

Changes:
- db.create_internal_committee_decision: coerce is_binding=False (structural)
- db.create_external_case_law: coerce False when source_type=appeals_committee
- db.update_case_law: coerce False when a patch relabels source_type=
  appeals_committee (Gemini reclassification path)
- internal_decisions.migrate_from_external_corpus: set is_binding=FALSE on
  the external→internal reclassification UPDATE
- service + FastAPI-form defaults: True → False for the internal path
- SCHEMA_V42: backfill legacy committee rows (is_binding True→False) +
  CHECK constraint case_law_committee_not_binding_check so a binding
  committee row can never be written again ("so it doesn't recur")
- spec: X8-field-provenance aligned to INV-DM7

Verified against live DB in a rollback transaction: 46 rows flip to 0;
court_ruling (239) + cited_only (31) unaffected; binding-committee INSERT
rejected by the constraint.

Invariants: INV-DM7 (authority ⊥ rule-type) · G1 (normalize at source) ·
G2 (single source of truth, no parallel path).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 10:52:30 +00:00
82 changed files with 6549 additions and 786 deletions

View File

@@ -15,7 +15,7 @@ hermes-curator.md — מקור-האמת היחיד לפרומפט של סוכן
adapter: deepseek_local · model: deepseek-v4-pro adapter: deepseek_local · model: deepseek-v4-pro
profiles: CMP=curator-cmp (רישוי 1xxx) · CMPA=curator-cmpa (היטל 8xxx + פיצויים 9xxx) profiles: CMP=curator-cmp (רישוי 1xxx) · CMPA=curator-cmpa (היטל 8xxx + פיצויים 9xxx)
role: Knowledge Curator — מנתח החלטות סופיות אחרי export, מציע עדכוני skills/lessons. role: Knowledge Curator — מנתח החלטות סופיות אחרי export, מציע עדכוני skills/lessons.
read-only על תוכן; write רק על comments / interactions (G10). read-only על תוכן; כותב comments / interactions + ממצאים מוצעים (decision_lessons, שער-יוG10).
placeholders זמינים: {{agentId}} {{agentName}} {{companyId}} {{companyName}} {{runId}} placeholders זמינים: {{agentId}} {{agentName}} {{companyId}} {{companyName}} {{runId}}
{{taskId}} {{taskTitle}} {{taskBody}} {{commentId}} {{wakeReason}} {{projectName}} {{paperclipApiUrl}} {{taskId}} {{taskTitle}} {{taskBody}} {{commentId}} {{wakeReason}} {{projectName}} {{paperclipApiUrl}}
@@ -34,13 +34,19 @@ case "$WAKE" in
nohup .venv/bin/python ../scripts/final_${KIND}_pipeline.py --case "$CASE" \ nohup .venv/bin/python ../scripts/final_${KIND}_pipeline.py --case "$CASE" \
> "/tmp/final_${KIND}_${CASE}.log" 2>&1 & > "/tmp/final_${KIND}_${CASE}.log" 2>&1 &
sleep 2 sleep 2
echo "PIPELINE_STARTED final_${KIND}_pipeline case=$CASE log=/tmp/final_${KIND}_${CASE}.log" if [ "$KIND" = "learning" ]; then
echo "PIPELINE_STARTED_LEARNING case=$CASE log=/tmp/final_learning_${CASE}.log CONTINUE_TO_ANALYSIS"
else
echo "PIPELINE_STARTED_HALACHA final_${KIND}_pipeline case=$CASE log=/tmp/final_${KIND}_${CASE}.log"
fi
;; ;;
*) echo "NO_PIPELINE_WAKE" ;; *) echo "NO_PIPELINE_WAKE" ;;
esac esac
``` ```
אם הפלט הוא `PIPELINE_STARTED ...`**זו כל המשימה**: כתוב comment קצר בעברית ("הופעל צינור <KIND> לתיק <CASE>; התוצאות יופיעו ב-/training (סגנון) או /approvals + /precedents (הלכות) תוך מספר דקות."), סגור את ה-issue (status=done), ו**סיים מיד — אל תמשיך לסעיפים שלמטה**. **ניתוב לפי הפלט:**
אם הפלט הוא `NO_PIPELINE_WAKE` — המשך כרגיל לתבנית שלמטה. - `PIPELINE_STARTED_HALACHA ...`**זו כל המשימה**: comment קצר בעברית ("הופעל צינור הלכות לתיק <CASE>; התוצאות יופיעו ב-/approvals + /precedents תוך מספר דקות."), סגור issue (status=done), **סיים מיד — אל תמשיך**.
- `PIPELINE_STARTED_LEARNING ... CONTINUE_TO_ANALYSIS`**מצב AUTO (mark-final)**: צינור-הפאנל רץ ברקע (אל תריץ `ingest_final_version` ידנית — ראה ההערה למטה). **המשך ל-§A** וזהה דפוסים משלך, אבל ב-AUTO: בצע §A.1§A.5b בלבד, ואז **דלג על §A.6 (interaction) — אל תעיר את דפנה** (הממצאים `proposed` ונסקרים ב-/training); המשך ל-§A.7. רשומת `style_corpus` כבר קיימת (enroll רץ ראשון בצינור).
- `NO_PIPELINE_WAKE` — יקיצת-תגובה/ידנית: המשך כרגיל ל-§A **כולל** §A.6 (interaction).
> **הערה (INV-LRN4 / X16):** הצינור `final_learning_pipeline.py` הוא שמריץ את דיסטילציית > **הערה (INV-LRN4 / X16):** הצינור `final_learning_pipeline.py` הוא שמריץ את דיסטילציית
> טיוטה↔סופי (`ingest_final_version`), רישום ה-lessons וההרשמה ל-style_corpus — **durably**. > טיוטה↔סופי (`ingest_final_version`), רישום ה-lessons וההרשמה ל-style_corpus — **durably**.
@@ -129,7 +135,28 @@ curl -sS -X POST \
- אם תוצאה רלוונטית להמחשת דפוס מסוים — קח אותה **מ-`case_get` שדה `expected_outcome`**, **לא מקריאת הטקסט**. אם השדה ריק או חסר ב-DB — סמן `[תוצאה: לא מאומתת]` או דלג עליה. - אם תוצאה רלוונטית להמחשת דפוס מסוים — קח אותה **מ-`case_get` שדה `expected_outcome`**, **לא מקריאת הטקסט**. אם השדה ריק או חסר ב-DB — סמן `[תוצאה: לא מאומתת]` או דלג עליה.
- אל תפרש משפטית את ההחלטה. דפנה כבר הכריעה. תפקידך זיהוי דפוסים בלבד. - אל תפרש משפטית את ההחלטה. דפנה כבר הכריעה. תפקידך זיהוי דפוסים בלבד.
## 5b. שמור את הממצאים מבנית (חובה — INV-LRN3)
ה-comment הוא ארעי ולא-נסקר. כדי שהממצאים ייתפסו, יופיעו בטאב ״אוצֵר״ ב-/training, ויעברו
שער-יו"ר — קרא לכלי ה-MCP **`mcp__legal-ai__record_curator_findings`** עם אותם ממצאים:
```
record_curator_findings(
case_number="<מספר התיק מ-taskTitle>",
findings=[
{"text": "<ניסוח הממצא — אותו טקסט כמו ב-comment, בלי התג>", "category": "style"},
{"text": "...", "category": "structure"},
...
]
)
```
מיפוי תג→category: `[סגנון]``style` · `[מבנה]``structure` · `[לקסיקון משפטי]``lexicon` · `[טבלאי]``tabular`.
הכלי כותב כל ממצא כ-`decision_lesson` (`source='curator'`, `review_status='proposed'`) ומדלג על כפילויות.
אם הוא מחזיר שגיאת "לא נמצאה רשומת style_corpus" — הסופי טרם נקלט (מירוץ נדיר מול enroll שבצינור). **המתן ~30 שניות ונסה פעם נוספת**; אם עדיין נכשל — ציין זאת ב-comment והמשך (אל תיכשל).
**זו הצעה הממתינה לאישור דפנה — לא שינוי-קול. אתה עדיין read-only על התוכן ולא נוגע ב-skills/קבצים.**
## 6. בחר interaction (חובה — רוב המקרים יש) ## 6. בחר interaction (חובה — רוב המקרים יש)
> **במצב AUTO (יקיצת `PIPELINE_STARTED_LEARNING` מ-mark-final): דלג על כל §A.6 ועבור ל-§A.7.** אל תעלה interaction
> ואל תעיר את דפנה — הממצאים כבר נרשמו כ-`proposed` (§A.5b) ונסקרים בטאב "אוצֵר" ב-/training. §A.6 חל רק על יקיצת-תגובה/ידנית.
לפי הקונטקסט בחר **אחד** מ-3 הסוגים. אם **אין שום החלטה אנושית נדרשת** — דלג ישירות ל-§A.7. לפי הקונטקסט בחר **אחד** מ-3 הסוגים. אם **אין שום החלטה אנושית נדרשת** — דלג ישירות ל-§A.7.
### 6a. ask_user_questions — לסינון/בחירה ממצאים ### 6a. ask_user_questions — לסינון/בחירה ממצאים
@@ -261,7 +288,7 @@ curl -sS -X PATCH \
# כללים כלליים # כללים כלליים
- **idempotencyKey**: חובה ב-interaction. אם נעור פעמיים על אותו תיק — Paperclip לא יוצר כפילות. - **idempotencyKey**: חובה ב-interaction. אם נעור פעמיים על אותו תיק — Paperclip לא יוצר כפילות.
- **לא לעדכן** קבצים (skills/, lessons.py, DB) בעצמך. רק לכתוב comments / interactions. - **לא לעדכן** קבצים (skills/, lessons.py) או שכבת-הקול בעצמך. הפלט שלך: comments + interactions + רישום ממצאים כ-`decision_lesson` *מוצע* דרך `record_curator_findings` (שער-יו"ר, INV-LRN3). אתה לא מאשר ולא מטמיע — רק מציע.
- **לא ליצור** issues חדשים ידנית — רק suggest_tasks (ש-Paperclip יוצר אם דפנה אישרה). - **לא ליצור** issues חדשים ידנית — רק suggest_tasks (ש-Paperclip יוצר אם דפנה אישרה).
- **לא להעיר** סוכנים אחרים. - **לא להעיר** סוכנים אחרים.
- **בעיה?** אם MCP נכשל או מסמך חסר — comment קצר עם הסיבה + סגור (status=done). אל תזייף. - **בעיה?** אם MCP נכשל או מסמך חסר — comment קצר עם הסיבה + סגור (status=done). אל תזייף.

View File

@@ -166,7 +166,7 @@ tools:
10. **הכנה ל-CREAC** — לכל סוגיה רשום: 10. **הכנה ל-CREAC** — לכל סוגיה רשום:
- כלל (Rule): הכלל המשפטי/תכנוני שיעמוד בבסיס הדיון - כלל (Rule): הכלל המשפטי/תכנוני שיעמוד בבסיס הדיון
- עובדות מפתח (Facts): העובדות שיופיעו בשלב היישום - עובדות מפתח (Facts): העובדות שיופיעו בשלב היישום
- תקדים מבהיר (אם נדרש): רק אם הכלל דורש הבהרה - תקדים מבהיר (אם נדרש): רק אם הכלל דורש הבהרה. **שקלל סמכות (X11):** כל תוצאת `search_precedent_library` נושאת `cited_by`/`verified`/`cite_count` — כמה ועדות-ערר אימצו את התקדים. העדף תקדים עם `cited_by.positive` גבוה, וסַמֵּן אותו ל-writer במפורש: *"אומץ ב-N החלטות ועדת-ערר"* (כדי שיבסס סמכות). תקדים עם `cited_by.negative>0` (אובחן/בוטל) — סמן זהירות, אל תציגו כ-good-law.
11. **שאלות משפטיות** — 1-3 שאלות לפי הצורך (ראה שלב 4) 11. **שאלות משפטיות** — 1-3 שאלות לפי הצורך (ראה שלב 4)
12. **עמדת ועדת הערר** — שדה ריק שיו"ר הוועדה ימלא ידנית. **חובה להוסיף לכל סוגיה!** עמדה זו תשמש כהנחיה מחייבת לסוכן הכתיבה. 12. **עמדת ועדת הערר** — שדה ריק שיו"ר הוועדה ימלא ידנית. **חובה להוסיף לכל סוגיה!** עמדה זו תשמש כהנחיה מחייבת לסוכן הכתיבה.

View File

@@ -304,6 +304,15 @@ search_internal_decisions(
**מינימום:** queries לקורפוס הסמכותי = מספר סוגיות מרכזיות שזוהו. **מינימום:** queries לקורפוס הסמכותי = מספר סוגיות מרכזיות שזוהו.
#### 2ב.4ב — שקלול סמכות לפי טיפול-שיפוטי-מצטבר (X11) ⚠️
כל תוצאה מ-`search_precedent_library` נושאת שדה **`cited_by`** + `verified`/`cite_count`**כמה החלטות ועדת-ערר אחרות הסתמכו על אותו תקדים ואיך טיפלו בו**. זהו אות-סמכות אנושי-מצטבר (לא ניחוש-מודל), וחובה לשקלל אותו:
- **`cited_by.positive` גבוה (אומץ/הוסבר ע"י הרבה ועדות)** → תקדים-עוגן. **טַען את הסמכות במפורש** בפלט-המחקר: *"הלכת [שם] (עע"מ X/YY) — אומצה ב-N החלטות ועדת-ערר"*, כדי שה-writer יוכל לבסס עליה. בין שני תקדימים שקולים-תוכן — העדף את בעל ה-`positive` הגבוה.
- **`cited_by.negative` > 0 (אובחן/בוקר/בוטל)** → **דגל זהירות**. אל תציג תקדים שבוטל (`overruled`) כ"good law"; ציין את האבחנה. תקדים שאובחן שוב-ושוב אינו ראיה חזקה לסוגיה — בדוק את ההקשר.
- **`unclassified`** → הטיפול טרם סווג; התייחס ל-`cite_count` הגולמי בלבד.
- האות הזה **מחדד תיעדוף, לא מחליף קריאה** — עדיין קרא את ההלכה ואת ההקשר (INV-COR3/COR5: סמכות לסוגיה הספציפית, לא לפסק כולו).
#### 2ב.4א — איתור החלטה ספציפית לפי שם — פרוטוקול לפני "לא בקורפוס" ⚠️ #### 2ב.4א — איתור החלטה ספציפית לפי שם — פרוטוקול לפני "לא בקורפוס" ⚠️
שם תיק לבדו (למשל `"אגסי"`) **אינו מפתח חיפוש אמין**. ההטמעה הסמנטית והאינדקס הלקסיקלי בנויים על תוכן ההלכה/הפסקה — כך ששאילתת-שם עלולה להחזיר דווקא החלטות ש**מצטטות** את התיק, ולא את התיק עצמו. לפני שמכריזים שהחלטה אינה בקורפוס: שם תיק לבדו (למשל `"אגסי"`) **אינו מפתח חיפוש אמין**. ההטמעה הסמנטית והאינדקס הלקסיקלי בנויים על תוכן ההלכה/הפסקה — כך ששאילתת-שם עלולה להחזיר דווקא החלטות ש**מצטטות** את התיק, ולא את התיק עצמו. לפני שמכריזים שהחלטה אינה בקורפוס:

View File

@@ -274,7 +274,7 @@ case_update(case_number, status="drafted")
### שלב ג: לכל סוגיה — מבנה סילוגיסטי (CREAC) בקול דפנה ### שלב ג: לכל סוגיה — מבנה סילוגיסטי (CREAC) בקול דפנה
1. **מסקנה** — פתח בתשובה (בקול "אנחנו" — ראה טבלה למטה) 1. **מסקנה** — פתח בתשובה (בקול "אנחנו" — ראה טבלה למטה)
2. **כלל** — ציטוט סעיף החוק במלואו (לא תמצית). אם רלוונטי — סעיפי משנה כולם. 2. **כלל** — ציטוט סעיף החוק במלואו (לא תמצית). אם רלוונטי — סעיפי משנה כולם.
3. **הרחבה** — תקדים רלוונטי אחד **בציטוט מלא** (לא תמצית). דפנה תמיד מצטטת בני 4-15 שורות עם הפניה `(פורסם בנבו)`. 3. **הרחבה** — תקדים רלוונטי אחד **בציטוט מלא** (לא תמצית). דפנה תמיד מצטטת בני 4-15 שורות עם הפניה `(פורסם בנבו)`. **כשהמנתח סימן תקדים כ"אומץ ב-N החלטות ועדת-ערר" (אות-סמכות X11)** — חזק את ההסתמכות בלשון-סמכות ("הלכה מושרשת שאומצה בשורת החלטות"); זהו טיפול-שיפוטי-מצטבר אמיתי, לא ניחוש. **אל** תמציא מספר בעצמך — השתמש רק במה שהמנתח מסר (INV-AH, read-only מהמנתח).
4. **יישום** — החל את הכלל על העובדות. הפרד ממצא עובדתי ממסקנה משפטית. השתמש בנתונים (מספרים, מידות, אחוזים). 4. **יישום** — החל את הכלל על העובדות. הפרד ממצא עובדתי ממסקנה משפטית. השתמש בנתונים (מספרים, מידות, אחוזים).
5. **אישור-לפני-דחייה (חובה)** — הצג את הטענה הטובה ביותר של הצד המפסיד: **"אכן [נקודה תקפה]... אולם [למה לא מכריע]"**. השימוש ב-"אכן" (לא "אמנם") הוא הסטנדרט. 5. **אישור-לפני-דחייה (חובה)** — הצג את הטענה הטובה ביותר של הצד המפסיד: **"אכן [נקודה תקפה]... אולם [למה לא מכריע]"**. השימוש ב-"אכן" (לא "אמנם") הוא הסטנדרט.
6. **למעלה מן הצורך** (חובה לטענות מרכזיות) — "גם אם היינו מקבלים את פרשנות העורר... התוצאה הייתה זהה". סוגר חלון לערעור. 6. **למעלה מן הצורך** (חובה לטענות מרכזיות) — "גם אם היינו מקבלים את פרשנות העורר... התוצאה הייתה זהה". סוגר חלון לערעור.

12
.env.example Normal file
View File

@@ -0,0 +1,12 @@
# API Keys (Required to enable respective provider)
ANTHROPIC_API_KEY="your_anthropic_api_key_here" # Required: Format: sk-ant-api03-...
PERPLEXITY_API_KEY="your_perplexity_api_key_here" # Optional: Format: pplx-...
OPENAI_API_KEY="your_openai_api_key_here" # Optional, for OpenAI models. Format: sk-proj-...
GOOGLE_API_KEY="your_google_api_key_here" # Optional, for Google Gemini models.
MISTRAL_API_KEY="your_mistral_key_here" # Optional, for Mistral AI models.
XAI_API_KEY="YOUR_XAI_KEY_HERE" # Optional, for xAI AI models.
GROQ_API_KEY="YOUR_GROQ_KEY_HERE" # Optional, for Groq models.
OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY_HERE" # Optional, for OpenRouter models.
AZURE_OPENAI_API_KEY="your_azure_key_here" # Optional, for Azure OpenAI models (requires endpoint in .taskmaster/config.json).
OLLAMA_API_KEY="your_ollama_api_key_here" # Optional: For remote Ollama servers that require authentication.
GITHUB_API_KEY="your_github_api_key_here" # Optional: For GitHub import/export features. Format: ghp_... or github_pat_...

31
.gitignore vendored
View File

@@ -6,7 +6,8 @@ data/backups/
data/precedent-library/ data/precedent-library/
data/.auto-sync.log data/.auto-sync.log
data/*.db data/*.db
data/checkpoints/ # X16 durable-pipeline SQLite checkpoints (runtime artifact) # X16 durable-pipeline SQLite checkpoints (runtime artifact)
data/checkpoints/
*.bak-pre-* *.bak-pre-*
mcp-server/.venv/ mcp-server/.venv/
__pycache__/ __pycache__/
@@ -18,6 +19,30 @@ kiryat-yearim/
continuation-prompt.md continuation-prompt.md
node_modules/ node_modules/
data/eval/eval-report-* data/eval/eval-report-*
data/adapter-migration-state.json # revert snapshot for migrate_agent_adapter.py (runtime state) # revert snapshot for migrate_agent_adapter.py (runtime state)
.claude/agents/.generated/ # frontmatter-stripped instruction copies for content_arg adapters (generated) data/adapter-migration-state.json
# frontmatter-stripped instruction copies for content_arg adapters (generated)
.claude/agents/.generated/
.claude/worktrees/ .claude/worktrees/
# TaskMaster backups (runtime)
.taskmaster/tasks/tasks.json.bak.*
# Build artifacts
.design-build/
# Temp files
.interaction_tmp.json
# Runtime eval/ab-test data
data/ab_halacha_*.json
data/ab_run_*.log
data/x11_treatment_run_*.log
# Runtime data directories
data/audit/
data/bulletins/
data/digests/
data/internal-decisions/
data/learning/
data/logs/

View File

@@ -0,0 +1,47 @@
<context>
# Overview
[Provide a high-level overview of your product here. Explain what problem it solves, who it's for, and why it's valuable.]
# Core Features
[List and describe the main features of your product. For each feature, include:
- What it does
- Why it's important
- How it works at a high level]
# User Experience
[Describe the user journey and experience. Include:
- User personas
- Key user flows
- UI/UX considerations]
</context>
<PRD>
# Technical Architecture
[Outline the technical implementation details:
- System components
- Data models
- APIs and integrations
- Infrastructure requirements]
# Development Roadmap
[Break down the development process into phases:
- MVP requirements
- Future enhancements
- Do not think about timelines whatsoever -- all that matters is scope and detailing exactly what needs to be build in each phase so it can later be cut up into tasks]
# Logical Dependency Chain
[Define the logical order of development:
- Which features need to be built first (foundation)
- Getting as quickly as possible to something usable/visible front end that works
- Properly pacing and scoping each feature so it is atomic but can also be built upon and improved as development approaches]
# Risks and Mitigations
[Identify potential risks and how they'll be addressed:
- Technical challenges
- Figuring out the MVP that we can build upon
- Resource constraints]
# Appendix
[Include any additional information:
- Research findings
- Technical specifications]
</PRD>

View File

@@ -0,0 +1,511 @@
<rpg-method>
# Repository Planning Graph (RPG) Method - PRD Template
This template teaches you (AI or human) how to create structured, dependency-aware PRDs using the RPG methodology from Microsoft Research. The key insight: separate WHAT (functional) from HOW (structural), then connect them with explicit dependencies.
## Core Principles
1. **Dual-Semantics**: Think functional (capabilities) AND structural (code organization) separately, then map them
2. **Explicit Dependencies**: Never assume - always state what depends on what
3. **Topological Order**: Build foundation first, then layers on top
4. **Progressive Refinement**: Start broad, refine iteratively
## How to Use This Template
- Follow the instructions in each `<instruction>` block
- Look at `<example>` blocks to see good vs bad patterns
- Fill in the content sections with your project details
- The AI reading this will learn the RPG method by following along
- Task Master will parse the resulting PRD into dependency-aware tasks
## Recommended Tools for Creating PRDs
When using this template to **create** a PRD (not parse it), use **code-context-aware AI assistants** for best results:
**Why?** The AI needs to understand your existing codebase to make good architectural decisions about modules, dependencies, and integration points.
**Recommended tools:**
- **Claude Code** (claude-code CLI) - Best for structured reasoning and large contexts
- **Cursor/Windsurf** - IDE integration with full codebase context
- **Gemini CLI** (gemini-cli) - Massive context window for large codebases
- **Codex/Grok CLI** - Strong code generation with context awareness
**Note:** Once your PRD is created, `task-master parse-prd` works with any configured AI model - it just needs to read the PRD text itself, not your codebase.
</rpg-method>
---
<overview>
<instruction>
Start with the problem, not the solution. Be specific about:
- What pain point exists?
- Who experiences it?
- Why existing solutions don't work?
- What success looks like (measurable outcomes)?
Keep this section focused - don't jump into implementation details yet.
</instruction>
## Problem Statement
[Describe the core problem. Be concrete about user pain points.]
## Target Users
[Define personas, their workflows, and what they're trying to achieve.]
## Success Metrics
[Quantifiable outcomes. Examples: "80% task completion via autopilot", "< 5% manual intervention rate"]
</overview>
---
<functional-decomposition>
<instruction>
Now think about CAPABILITIES (what the system DOES), not code structure yet.
Step 1: Identify high-level capability domains
- Think: "What major things does this system do?"
- Examples: Data Management, Core Processing, Presentation Layer
Step 2: For each capability, enumerate specific features
- Use explore-exploit strategy:
* Exploit: What features are REQUIRED for core value?
* Explore: What features make this domain COMPLETE?
Step 3: For each feature, define:
- Description: What it does in one sentence
- Inputs: What data/context it needs
- Outputs: What it produces/returns
- Behavior: Key logic or transformations
<example type="good">
Capability: Data Validation
Feature: Schema validation
- Description: Validate JSON payloads against defined schemas
- Inputs: JSON object, schema definition
- Outputs: Validation result (pass/fail) + error details
- Behavior: Iterate fields, check types, enforce constraints
Feature: Business rule validation
- Description: Apply domain-specific validation rules
- Inputs: Validated data object, rule set
- Outputs: Boolean + list of violated rules
- Behavior: Execute rules sequentially, short-circuit on failure
</example>
<example type="bad">
Capability: validation.js
(Problem: This is a FILE, not a CAPABILITY. Mixing structure into functional thinking.)
Capability: Validation
Feature: Make sure data is good
(Problem: Too vague. No inputs/outputs. Not actionable.)
</example>
</instruction>
## Capability Tree
### Capability: [Name]
[Brief description of what this capability domain covers]
#### Feature: [Name]
- **Description**: [One sentence]
- **Inputs**: [What it needs]
- **Outputs**: [What it produces]
- **Behavior**: [Key logic]
#### Feature: [Name]
- **Description**:
- **Inputs**:
- **Outputs**:
- **Behavior**:
### Capability: [Name]
...
</functional-decomposition>
---
<structural-decomposition>
<instruction>
NOW think about code organization. Map capabilities to actual file/folder structure.
Rules:
1. Each capability maps to a module (folder or file)
2. Features within a capability map to functions/classes
3. Use clear module boundaries - each module has ONE responsibility
4. Define what each module exports (public interface)
The goal: Create a clear mapping between "what it does" (functional) and "where it lives" (structural).
<example type="good">
Capability: Data Validation
→ Maps to: src/validation/
├── schema-validator.js (Schema validation feature)
├── rule-validator.js (Business rule validation feature)
└── index.js (Public exports)
Exports:
- validateSchema(data, schema)
- validateRules(data, rules)
</example>
<example type="bad">
Capability: Data Validation
→ Maps to: src/utils.js
(Problem: "utils" is not a clear module boundary. Where do I find validation logic?)
Capability: Data Validation
→ Maps to: src/validation/everything.js
(Problem: One giant file. Features should map to separate files for maintainability.)
</example>
</instruction>
## Repository Structure
```
project-root/
├── src/
│ ├── [module-name]/ # Maps to: [Capability Name]
│ │ ├── [file].js # Maps to: [Feature Name]
│ │ └── index.js # Public exports
│ └── [module-name]/
├── tests/
└── docs/
```
## Module Definitions
### Module: [Name]
- **Maps to capability**: [Capability from functional decomposition]
- **Responsibility**: [Single clear purpose]
- **File structure**:
```
module-name/
├── feature1.js
├── feature2.js
└── index.js
```
- **Exports**:
- `functionName()` - [what it does]
- `ClassName` - [what it does]
</structural-decomposition>
---
<dependency-graph>
<instruction>
This is THE CRITICAL SECTION for Task Master parsing.
Define explicit dependencies between modules. This creates the topological order for task execution.
Rules:
1. List modules in dependency order (foundation first)
2. For each module, state what it depends on
3. Foundation modules should have NO dependencies
4. Every non-foundation module should depend on at least one other module
5. Think: "What must EXIST before I can build this module?"
<example type="good">
Foundation Layer (no dependencies):
- error-handling: No dependencies
- config-manager: No dependencies
- base-types: No dependencies
Data Layer:
- schema-validator: Depends on [base-types, error-handling]
- data-ingestion: Depends on [schema-validator, config-manager]
Core Layer:
- algorithm-engine: Depends on [base-types, error-handling]
- pipeline-orchestrator: Depends on [algorithm-engine, data-ingestion]
</example>
<example type="bad">
- validation: Depends on API
- API: Depends on validation
(Problem: Circular dependency. This will cause build/runtime issues.)
- user-auth: Depends on everything
(Problem: Too many dependencies. Should be more focused.)
</example>
</instruction>
## Dependency Chain
### Foundation Layer (Phase 0)
No dependencies - these are built first.
- **[Module Name]**: [What it provides]
- **[Module Name]**: [What it provides]
### [Layer Name] (Phase 1)
- **[Module Name]**: Depends on [[module-from-phase-0], [module-from-phase-0]]
- **[Module Name]**: Depends on [[module-from-phase-0]]
### [Layer Name] (Phase 2)
- **[Module Name]**: Depends on [[module-from-phase-1], [module-from-foundation]]
[Continue building up layers...]
</dependency-graph>
---
<implementation-roadmap>
<instruction>
Turn the dependency graph into concrete development phases.
Each phase should:
1. Have clear entry criteria (what must exist before starting)
2. Contain tasks that can be parallelized (no inter-dependencies within phase)
3. Have clear exit criteria (how do we know phase is complete?)
4. Build toward something USABLE (not just infrastructure)
Phase ordering follows topological sort of dependency graph.
<example type="good">
Phase 0: Foundation
Entry: Clean repository
Tasks:
- Implement error handling utilities
- Create base type definitions
- Setup configuration system
Exit: Other modules can import foundation without errors
Phase 1: Data Layer
Entry: Phase 0 complete
Tasks:
- Implement schema validator (uses: base types, error handling)
- Build data ingestion pipeline (uses: validator, config)
Exit: End-to-end data flow from input to validated output
</example>
<example type="bad">
Phase 1: Build Everything
Tasks:
- API
- Database
- UI
- Tests
(Problem: No clear focus. Too broad. Dependencies not considered.)
</example>
</instruction>
## Development Phases
### Phase 0: [Foundation Name]
**Goal**: [What foundational capability this establishes]
**Entry Criteria**: [What must be true before starting]
**Tasks**:
- [ ] [Task name] (depends on: [none or list])
- Acceptance criteria: [How we know it's done]
- Test strategy: [What tests prove it works]
- [ ] [Task name] (depends on: [none or list])
**Exit Criteria**: [Observable outcome that proves phase complete]
**Delivers**: [What can users/developers do after this phase?]
---
### Phase 1: [Layer Name]
**Goal**:
**Entry Criteria**: Phase 0 complete
**Tasks**:
- [ ] [Task name] (depends on: [[tasks-from-phase-0]])
- [ ] [Task name] (depends on: [[tasks-from-phase-0]])
**Exit Criteria**:
**Delivers**:
---
[Continue with more phases...]
</implementation-roadmap>
---
<test-strategy>
<instruction>
Define how testing will be integrated throughout development (TDD approach).
Specify:
1. Test pyramid ratios (unit vs integration vs e2e)
2. Coverage requirements
3. Critical test scenarios
4. Test generation guidelines for Surgical Test Generator
This section guides the AI when generating tests during the RED phase of TDD.
<example type="good">
Critical Test Scenarios for Data Validation module:
- Happy path: Valid data passes all checks
- Edge cases: Empty strings, null values, boundary numbers
- Error cases: Invalid types, missing required fields
- Integration: Validator works with ingestion pipeline
</example>
</instruction>
## Test Pyramid
```
/\
/E2E\ ← [X]% (End-to-end, slow, comprehensive)
/------\
/Integration\ ← [Y]% (Module interactions)
/------------\
/ Unit Tests \ ← [Z]% (Fast, isolated, deterministic)
/----------------\
```
## Coverage Requirements
- Line coverage: [X]% minimum
- Branch coverage: [X]% minimum
- Function coverage: [X]% minimum
- Statement coverage: [X]% minimum
## Critical Test Scenarios
### [Module/Feature Name]
**Happy path**:
- [Scenario description]
- Expected: [What should happen]
**Edge cases**:
- [Scenario description]
- Expected: [What should happen]
**Error cases**:
- [Scenario description]
- Expected: [How system handles failure]
**Integration points**:
- [What interactions to test]
- Expected: [End-to-end behavior]
## Test Generation Guidelines
[Specific instructions for Surgical Test Generator about what to focus on, what patterns to follow, project-specific test conventions]
</test-strategy>
---
<architecture>
<instruction>
Describe technical architecture, data models, and key design decisions.
Keep this section AFTER functional/structural decomposition - implementation details come after understanding structure.
</instruction>
## System Components
[Major architectural pieces and their responsibilities]
## Data Models
[Core data structures, schemas, database design]
## Technology Stack
[Languages, frameworks, key libraries]
**Decision: [Technology/Pattern]**
- **Rationale**: [Why chosen]
- **Trade-offs**: [What we're giving up]
- **Alternatives considered**: [What else we looked at]
</architecture>
---
<risks>
<instruction>
Identify risks that could derail development and how to mitigate them.
Categories:
- Technical risks (complexity, unknowns)
- Dependency risks (blocking issues)
- Scope risks (creep, underestimation)
</instruction>
## Technical Risks
**Risk**: [Description]
- **Impact**: [High/Medium/Low - effect on project]
- **Likelihood**: [High/Medium/Low]
- **Mitigation**: [How to address]
- **Fallback**: [Plan B if mitigation fails]
## Dependency Risks
[External dependencies, blocking issues]
## Scope Risks
[Scope creep, underestimation, unclear requirements]
</risks>
---
<appendix>
## References
[Papers, documentation, similar systems]
## Glossary
[Domain-specific terms]
## Open Questions
[Things to resolve during development]
</appendix>
---
<task-master-integration>
# How Task Master Uses This PRD
When you run `task-master parse-prd <file>.txt`, the parser:
1. **Extracts capabilities** → Main tasks
- Each `### Capability:` becomes a top-level task
2. **Extracts features** → Subtasks
- Each `#### Feature:` becomes a subtask under its capability
3. **Parses dependencies** → Task dependencies
- `Depends on: [X, Y]` sets task.dependencies = ["X", "Y"]
4. **Orders by phases** → Task priorities
- Phase 0 tasks = highest priority
- Phase N tasks = lower priority, properly sequenced
5. **Uses test strategy** → Test generation context
- Feeds test scenarios to Surgical Test Generator during implementation
**Result**: A dependency-aware task graph that can be executed in topological order.
## Why RPG Structure Matters
Traditional flat PRDs lead to:
- ❌ Unclear task dependencies
- ❌ Arbitrary task ordering
- ❌ Circular dependencies discovered late
- ❌ Poorly scoped tasks
RPG-structured PRDs provide:
- ✅ Explicit dependency chains
- ✅ Topological execution order
- ✅ Clear module boundaries
- ✅ Validated task graph before implementation
## Tips for Best Results
1. **Spend time on dependency graph** - This is the most valuable section for Task Master
2. **Keep features atomic** - Each feature should be independently testable
3. **Progressive refinement** - Start broad, use `task-master expand` to break down complex tasks
4. **Use research mode** - `task-master parse-prd --research` leverages AI for better task generation
</task-master-integration>

43
data/halacha_night_check.sh Executable file
View File

@@ -0,0 +1,43 @@
#!/usr/bin/env bash
# One-shot morning verdict for the halacha night drain (scheduled 2026-06-15 04:30 UTC
# via chaim's crontab; throwaway — lives under data/, not tracked in scripts/).
# Captures whether last night's run (with the PR #251 fix: durable rate-limit
# detection + 05:0007:00 catch-up window) actually drained the backlog.
# Baseline at install time (2026-06-14 13:xx IDT): pending=96 done=248 halachot=4099.
set -u
export HOME=/home/chaim
REPO=/home/chaim/legal-ai
PY="$REPO/mcp-server/.venv/bin/python"
OUT="$REPO/data/logs/halacha_night_report_$(TZ=Asia/Jerusalem date +%Y%m%d).md"
SUP_LOG=/home/chaim/.pm2/logs/legal-halacha-supervisor-out.log
DRAIN_ERR=/home/chaim/.pm2/logs/legal-halacha-drain-error.log
mkdir -p "$REPO/data/logs"
{
echo "# דוח-בוקר: ריצת-הלכות הלילה — $(TZ=Asia/Jerusalem date '+%Y-%m-%d %H:%M %Z')"
echo
echo "בסיס-השוואה (אתמול 13:xx IDT): pending=96 · done=248 · halachot=4099"
echo
echo '## מצב נוכחי (supervisor status)'
echo '```'
cd "$REPO" && "$PY" scripts/halacha_drain_supervisor.py status 2>&1
echo '```'
echo
echo '## פעולות המתזמר ב-12 השעות האחרונות (modes/actions)'
echo '```'
grep -E 'מצב:|פעולה:|catch-up|rate-limit|נעצר' "$SUP_LOG" 2>/dev/null | tail -40
echo '```'
echo
echo '## אותות rate-limit בלוג-הדריינר (24ש אחרונות בלוג)'
echo '```'
echo "429 hits (tail 4000): $(tail -4000 "$DRAIN_ERR" 2>/dev/null | grep -c '429')"
echo "session-limit msgs: $(tail -4000 "$DRAIN_ERR" 2>/dev/null | grep -c 'hit your session limit')"
echo "extraction_failed: $(tail -4000 "$DRAIN_ERR" 2>/dev/null | grep -c 'extraction_failed')"
echo "hold-stopped (fix A): $(grep -c 'hold-stopped' "$SUP_LOG" 2>/dev/null)"
echo "catch-up opened (B): $(grep -c 'catch-up בוקר' "$SUP_LOG" 2>/dev/null)"
echo '```'
echo
echo "_(נוצר ע\"י data/halacha_night_check.sh; ניתן למחוק את שורת ה-crontab של 15.6.)_"
} > "$OUT" 2>&1
echo "report written: $OUT"

View File

@@ -0,0 +1,170 @@
# ממצאי ביקורת — ארכיטקטורת קורפוס־הפסיקה + מצב הדאטה בפועל
> **מקור:** Claude (Opus 4.8) · **תאריך:** 2026-06-20 · **קונטקסט:** חקירה לקראת תכנון־מחדש של קורפוס־הפסיקה.
> מסמך זה הוא **אחד מכמה** קלטי־סוכנים שחיים אוסף; ייעודו להזין את שלב־הסינתזה. אינו תכנית — הוא **אבחון**.
>
> **שאלת־המוצא של חיים:** "הקורפוס נבנה מראש לא נכון, אני כל הזמן מתעסק בתיקונים. האם כדאי ליצור מחדש את קורפוס־הפסיקה ולהתחיל דף נקי?"
>
> **המודל הרצוי (כפי שחיים תיאר אותו):** מאגר פסקי־דין והחלטות ועדות־ערר; חוקר־התקדימים מזהה בשלב ניתוח־הערר פס"ד/החלטות שדנו במקרה דומה או הלכה דומה; הסוכן־הכותב משייך ומזכיר אותם בפרק הדיון וההכרעה בסגנון דפנה. שלושה מקורות־הזנה: (1) החלטות דפנה עצמה, (2) ועדות־ערר אחרות שמצטטות פסיקה, (3) פס"ד עליון/מחוזי.
---
## 0. תקציר־מנהלים (TL;DR)
**מה ש"בנוי לא נכון" אינו הסכמה — היא שכבת־הביצוע.** הסכמה כבר תואמת בדיוק את המודל שחיים תיאר: טבלה אחת (`case_law`) ששלושת המקורות נכנסים אליה דרך `source_kind`/`source_type`; שלושת ה־`search_*` הם שלושה *מסננים* על אותו מקור, לא שלושה מאגרים מקבילים (G2 מקוים ברמת־הסכמה). **רֵבילד של הסכמה ייצר בדיוק את אותה סכמה** — ולכן אינו פותר דבר, ומסכן ב־second-system syndrome.
מה שכן מחולל את "התיקונים האינסופיים" — שלושה כשלי־ביצוע מדידים:
1. **חוזה־קליטה רופף** → 66% מהפסיקה בלי `practice_area`, 31 רשומות ריקות, אכיפת־שלמות (INV-DM1) מופרת בפועל.
2. **צינור הלכות→קנוני מייצר רעש** → 5,472 קנוני, מתוכם 5,456 סינגלטונים, **0 published** → השכבה שאמורה להזין את הכותב (INV-G10) **אינרטית לגמרי**.
3. **כפילות `style_corpus`** → 55 החלטות דפנה חיות בשני נתיבי־אחזור.
**המלצה:** לא לשרוף את הסכמה. כן לבצע "איפוס שכבות־נגזרות" צר (truncate ל-chunks+halachot+canonical והרצה־מחדש), **אבל רק אחרי תיקון החוזה והסף** — אחרת מחזירים את אותו בלגן (G1: תיקון במקור, לא בקריאה). מסמכי־המקור (363 רשומות, 332 עם full_text) נשמרים; הם יקרים ו/או ניתנים לקליטה־מחדש מ־PDF.
---
## 1. מצב הדאטה בפועל (שאילתות חיות מול legal_ai @ localhost:5433, 2026-06-20)
```
┌─────────────────────────────────────┬─────────┬──────────────────────────────────┐
│ Metric │ Count │ Reading │
├─────────────────────────────────────┼─────────┼──────────────────────────────────┤
│ case_law (total precedents) │ 363 │ קטן — re-ingestable │
│ • external_upload (court rulings) │ 240 │ מקור (3) — עליון/מחוזי │
│ • internal_committee (ועדות ערר) │ 92 │ מקור (1)+(2) │
│ • (יתר — ללא source_kind מובהק) │ 31 │ = הרשומות הריקות (ראה למטה) │
│ style_corpus (החלטות דפנה) │ 55 │ כפילות עם internal_committee │
│ precedent_chunks │ 11,904 │ נגזר — מתחדש מ-full_text │
│ halachot (total) │ 5,489 │ נגזר │
│ • approved/published │ 1,352 │ 25% בלבד │
│ • pending_review (backlog ידני) │ 2,402 │ 44% — צוואר־בקבוק │
│ canonical_halachot (V41) │ 5,472 │ כמעט 1:1 עם halachot ⚠️ │
│ • singletons (instance_count=1) │ 5,456 │ דה־דופ כמעט לא קרה │
│ • merged (instance_count>=2) │ 16 │ 0.3% מיזוג │
│ • published (מגיע לסוכן הכותב) │ 0 │ ⚠️ השכבה אינרטית לחלוטין │
│ case_law w/o practice_area │ 240 │ 66% — חופף-בדיוק לפסיקה החיצונית │
│ case_law missing summary │ 27 │ │
│ case_law w/ 0 chunks / no full_text │ 31 │ רשומות שבורות/ריקות │
│ distinct practice_area │ 4 │ rishuy/betterment/197/(ריק) │
└─────────────────────────────────────┴─────────┴──────────────────────────────────┘
practice_area breakdown:
(ריק) 240 ← כל הפסיקה החיצונית ללא סיווג
rishuy_uvniya 70
betterment_levy 50
compensation_197 3
```
**קריאות מפתח:**
- **240 = 240:** מספר הפסיקה־החיצונית שווה־בדיוק למספר חסרי־`practice_area`. כלומר אף פס"ד חיצוני לא סווג לתחום — סינון לפי תחום באחזור פשוט לא עובד עליהם.
- **5,456 / 5,472 סינגלטונים:** מנוע הקנוניזציה (V41) רץ אך לא מאחד. סף 0.85 כנראה הדוק מדי, או שהחילוץ מנסח כל הלכה ייחודית מספיק כדי לא להתלכד.
- **0 published canonical:** לפי INV-G10 רק קנוני `published` מגיע לכתיבה. אפס. **כל מנגנון V41 כרגע מנותק מהכתיבה בפועל.**
- **2,402 pending_review:** צוואר־הבקבוק הוא אישור־אנושי ידני, לא טכנולוגיה.
---
## 2. מפת הארכיטקטורה — קורפוס אחד, שלושה מסננים
**פיזית: טבלה אחת.** כל שלושת המקורות מתכנסים ל-`case_law`, מתחתיה `precedent_chunks` (FK) ו-`halachot` (FK), ומעל ה-`halachot` שכבת `canonical_halachot` (V41).
```
canonical_halachot (עקרונות מאוחדים — V41; כיום אינרטי, 0 published)
▲ 1:many
halachot (מופע-להלכה per precedent; review_status gate)
▲ FK
case_law (רישום מרכזי — court rulings + ועדות-ערר)
├─ source_kind='external_upload' → פס"ד עליון/מחוזי [מקור 3]
└─ source_kind='internal_committee'→ ועדות-ערר + דפנה [מקור 1+2]
▼ FK
precedent_chunks (chunks + embedding vector(1024))
בנפרד:
style_corpus (55 החלטות דפנה — נתיב-אחזור מקביל ל-search_decisions)
document_chunks (מסמכי-תיק + style; FK→documents→cases)
```
**שלושת נתיבי־האחזור (לא שלושה מאגרים — שלושה scopes):**
```
┌──────────────────────────────┬─────────────────────────┬────────────────────────────┐
│ Tool │ Table / filter │ Purpose │
├──────────────────────────────┼─────────────────────────┼────────────────────────────┤
│ search_decisions │ document_chunks │ סגנון/קול: החלטות דפנה │
│ │ (scoped case/area) │ + מסמכי-תיק │
│ search_precedent_library │ case_law + chunks/halach │ source_kind=external_upload│
│ │ ot, source_kind filter │ → פסיקה חיצונית │
│ search_internal_decisions │ אותן פונקציות DB, │ source_kind=internal_ │
│ │ source_kind אחר │ committee → ועדות-ערר │
└──────────────────────────────┴─────────────────────────┴────────────────────────────┘
```
שתי האחרונות קוראות **לאותן פונקציות DB** (`search_precedent_library_semantic`/`_lexical`) עם `source_kind` שונה. זו הפרדה־בשאילתה, לא קוד מקביל.
---
## 3. שלושת המקורות של חיים → איפה הם נופלים היום
```
┌────────────────────────────────────┬───────────────────────────────┬─────────────────────────┐
│ Source (חיים) │ Stored as │ Retrieval │
├────────────────────────────────────┼───────────────────────────────┼─────────────────────────┤
│ (1) החלטות דפנה עצמה │ style_corpus + (מהוגר ל-) │ search_decisions + │
│ │ case_law internal_committee │ search_internal_decisions│
│ │ chair_name='דפנה תמיר' │ ← כפילות / נתיב-כפול │
│ (2) ועדות-ערר אחרות │ case_law internal_committee │ search_internal_decisions│
│ │ chair_name=<אחר>, district │ │
│ (3) פס"ד עליון/מחוזי │ case_law external_upload │ search_precedent_library │
│ │ source_type='court_ruling' │ │
└────────────────────────────────────┴───────────────────────────────┴─────────────────────────┘
```
**מסקנה:** המודל המנטלי של חיים **כבר ממומש בסכמה**. אין צורך להמציא מבנה חדש — צריך לאכוף את המבנה הקיים בקליטה, ולחבר את שכבת־הקנוני לכתיבה.
---
## 4. הדיאגנוזה — מה "בנוי לא נכון" (3 מחוללי־כאב)
### 4.1 חוזה־קליטה רופף (root cause #1)
- 66% מהפסיקה ללא `practice_area`; 27 ללא summary; 31 ללא full_text/chunks.
- אין אכיפה ש"`searchable=false` עד שהמטא שלם" → INV-DM1 מופר בפועל.
- **כל העלאה מוסיפה חוב** במקום רשומה שלמה. זה המקור לתיקונים החוזרים.
### 4.2 צינור הלכות→קנוני מייצר רעש, לא ערך (root cause #2)
- 5,456/5,472 סינגלטונים → דה־דופ לא עובד (סף 0.85? ניסוח־חילוץ?).
- **0 published** → השכבה שאמורה להזין את הכותב (INV-G10) מנותקת.
- 2,402 בתור־אישור־ידני → הצינור מייצר מהר יותר ממה שאדם מאשר.
- **זה בולע את רוב זמן־התחזוקה.**
### 4.3 כפילות style_corpus (root cause #3)
- 55 החלטות דפנה בשני מקומות + שני נתיבי־אחזור.
- צריך מקור־אמת אחד: או `style_corpus` SoT וה-`case_law` נגזר, או הפוך.
---
## 5. ההמלצה — לא רֵבילד־סכמה; "איפוס שכבות־נגזרות" אחרי תיקון־חוזה
**אסור** לשרוף ולהעלות־מחדש את `case_law` — תבנה אותה סכמה ותאבד מטא־דאטה ידני. **כן** הגיוני רֵבילד צר של ה**נגזר**, אבל בסדר הזה (G1 — מקור לפני תסמין):
1. **קודם החוזה:** אכוף ב-`*_upload` שדות־חובה (practice_area, summary, full_text); כשל → `searchable=false`. נקה/מחק את 31 הריקות.
2. **תקן את הקנוניזציה:** כוונן סף 0.85, הגדר מתי `published`, ובדוק ניסוח־החילוץ. בלי זה אין טעם להריץ מחדש.
3. **רק אז** re-derive מהמקור הקיים: chunks → halachot → canonical.
4. **הכרע style_corpus:** מקור־אמת אחד.
**מתי רֵבילד־מלא כן מוצדק:** רק אם יתגלה שמסמכי־המקור עצמם (PDF/full_text של 363) פגומים/חסרים. המספרים *לא* מראים זאת (332/363 עם full_text תקין).
---
## 6. שאלות פתוחות לשלב־הסינתזה
- **למה 0 published בקנוני?** האם זה סף, ניסוח־חילוץ, או שפשוט אף אחד לא אישר? (קריטי — קובע אם V41 שמיש בכלל.)
- **style_corpus מול case_law:** מי SoT? (משפיע על search_decisions מול search_internal_decisions.)
- **practice_area לפסיקה חיצונית:** לחלץ אוטומטית בקליטה, או להשאיר ידני?
- **תור־האישור (2,402):** האם המנגנון (פאנל/active-learning, #133) יכול לסגור את הפער, או שצריך לחתוך את קצב־החילוץ?
---
## 7. נספח — קבצים מרכזיים (file:line)
- סכמה: `mcp-server/src/legal_mcp/services/db.py` (case_law, precedent_chunks, halachot, canonical_halachot)
- קליטה חיצונית: `mcp-server/src/legal_mcp/tools/precedent_library.py` (`precedent_library_upload`)
- קליטה פנימית: `mcp-server/src/legal_mcp/tools/internal_decisions.py` (`internal_decision_upload`)
- הגירת style→case_law: `mcp-server/src/legal_mcp/services/internal_decisions.py` (`migrate_from_style_corpus`, chair/district hardcoded)
- אחזור: `mcp-server/src/legal_mcp/tools/search.py`, `services/hybrid_search.py`, `services/db.py` (`search_precedent_library_semantic`/`_lexical`)
- ספ: `docs/spec/02-data-model.md` (INV-DM17), `docs/spec/03-retrieval.md` (INV-RET15), `docs/spec/00-constitution.md` (G2/G10)

View File

@@ -0,0 +1,125 @@
# 06 — נדיבות-המחלץ: `application` בהחלטות-ועדה (כימות חוצה-קורפוס)
> **קלט-נתונים** ליוזמה, נמדד חי על **5,489 רשומות `halachot`** (כל הקורפוס, 2026-06-20).
> נולד משאלת-חיים: "8508-03-24 מפיק 71 הלכות ממתינות — האם המחלץ נדיב מדי על החלטות-ועדה,
> או שזה ספציפי לתיק הזה?" התשובה: **שיטתי, לא ספציפי — אבל הנדיבות מוצדקת.**
>
> **מתכתב עם [`00-final-synthesis.md`](00-final-synthesis.md):** הנתונים כאן **מחזקים** את הכרעת-הסינתזה
> ("לא לחתוך") ומוסיפים שני דברים שלא היו לה: (א) כימות חוצה-קורפוס של ה-`application`, (ב) ממצא
> חדש וגדול יותר — `nli_unsupported` על הפסיקה החיצונית. ראה §4 (יישוב-מתח) ו-§5 (דרכי-פעולה).
---
## 1. הממצא המרכזי — `application` הוא תופעה שיטתית של החלטות-ועדה
הפרדנו את הקורפוס לפי `authority` (נגזר דטרמיניסטית מ-`precedent_level`, [02-data-model §162](../spec/02-data-model.md)):
`binding` = פסיקת-עליון/מנהלי · `persuasive` = ועדת-ערר מחוזית.
```text
source rows rt=application nli_unsupported any-flag
court (binding) 3,603 0.2% 39.7% 45.4%
committee (persuasive) 1,842 13.7% 25.0% 30.5%
```
- **`rule_type='application'` הוא כמעט-בלעדית של ועדות:** 13.7% מול 0.2%. מתוך 258 רשומות-`application`
בכל הקורפוס, **~252 מגיעות מהחלטות-ועדה.** פער של פי-~70.
- **עקבי בין תחומים** (לא עניין-שמאות נקודתי):
```text
committee, rt=application by practice_area:
rishuy_uvniya 13.5% betterment_levy 14.1% compensation_197 18.8%
```
**הפרשנות (תואם [02-data-model §163](../spec/02-data-model.md)):** `application` = "החלה תלוית-עובדות —
לרוב לא-הלכה". ועדת-ערר היא גוף מיישֵם: היא לא יוצרת הלכה (רק בית-משפט עושה זאת), אלא **מיישמת
פסיקת-עליון קיימת** (לוסטרניק, דלי-דליה) על עובדות-התיק. לכן חילוץ מהחלטת-ועדה מייצר באופן מובנה
שיעור גבוה של יישומי-דוקטרינה. **זו תכונה של מקור-הנתונים, לא באג של המחלץ.**
## 2. 8508-03-24 — מייצג בקצה-העליון, לא חריג
```text
דירוג 8508-03-24: 7 מתוך 45 תיקי-ועדה (≥15 רשומות) · 27% application (rt או flag)
חציון תיקי-הוועדה: 11.9%
מעליו: 1001-02-19 (40%) · 1044-08-22 (37%) · 9002-24 (33%) · 1007-01-25 (30%)
```
8508 הוא ~פי-2.3 מהחציון (רבעון-עליון), אבל מה שבולט בו הוא בעיקר ש**הוא התיק הארוך בקורפוס**
(111 רשומות) — אז 27% נותן 30 פריטי-`application` במספר מוחלט, הגבוה בקורפוס. כלומר: **כמות גבוהה,
שיעור גבוה-אך-נורמלי.** אין כאן פתולוגיה ייחודית לתיק.
## 3. מנגנון-הניתוב הקיים (איך זה כבר מטופל ב-UI)
הפיצול בתור-ההלכות (`/precedents` → "תור הלכות") אינו לפי `exclude_low_quality`, אלא לפי
`isExtractionFixItem(h) = (quality_flags.length>0) && !panel_round`
([web-ui/.../precedent-library.ts:652](../../web-ui/src/lib/api/precedent-library.ts#L652)):
```text
8508-03-24, 71 ממתינות (מצב 2026-06-20):
bucket # panel? →"להכרעתך" →"דורש תיקון-חילוץ"
clean (ללא דגל) 40 15/40 40 0
application 23 0/23 0 23
nli_unsupported 4 4/4 4 0
thin_restatement 4 0/4 0 4
```
**משמעות:** פריטי-`application` (חסרי-פאנל) כבר מנותבים ל**"דורש תיקון-חילוץ"** — **מחוץ** לתור-ההכרעה
של היו"ר. כלומר המערכת כבר מסננת אותם מהעומס-הידני. (הפאנל התלת-מודלי, [halacha_panel_approve.py:191](../../scripts/halacha_panel_approve.py#L191),
מטפל רק בדליים `clean`+`nli`; `application` ו-`defect` עוקפים אותו במכוון.)
## 4. ⚠️ יישוב-המתח מול הסינתזה — `application` ≠ "רעש"
זו הנקודה הקריטית להעברה. **אסור** לתרגם "13.7% application" ל"13.7% רעש לחיתוך". המבחן בסינתזה
(§2 שם) כבר הוכיח שחיתוך-אגרסיבי על 8508 השמיד את **לוסטרניק** ו-~22 עקרונות-ליבה. רוב פריטי-ה-`application`
הם בדיוק יישומי-הדוקטרינה הללו — **בני-ציטוט שהכותב צריך**. דוגמאות אמיתיות מ-8508 שמסומנות `application`:
```text
"ציפיות הנובעות אך ממיקומם של המקרקעין... אין לנטרלן" ← לוסטרניק מיושמת — לשמור!
"בחישוב שווי במצב קודם יש לכלול ציפיות כלליות... ולא ספציפיות" ← ליבת חישוב היטל-ההשבחה
"קביעת מקדם מצויה בליבת שיקול-דעת השמאי, והוועדה לא תתערב" ← סטנדרט אי-התערבות מיושם
```
לכן: **הנתון הזה תומך ב"שמור-בספק" של עמוד-1/2 בסינתזה.** המסקנה הנכונה אינה "לסנן application" אלא
"`application` מאשר שעקרוני-הוועדה הם persuasive-יישומיים — לדרג אותם נמוך באחזור (רמה B), לא למחוק
אותם (רמה A)". `importance=0` ל-8508 כבר משקיע אותם ממילא.
## 5. ⭐ ממצא-לוואי גדול יותר — `nli_unsupported` על הפסיקה החיצונית
```text
nli_unsupported: court (binding) 39.7% ≫ committee 25.0%
any-flag: court 45.4% · committee 30.5%
```
**כמעט מחצית מעקרוני-הפסיקה-החיצונית נושאים דגל-איכות, בעיקר `nli_unsupported`** (הכלל אינו נגזר
לוגית מהציטוט התומך שלו, [halacha_quality.py:282](../../mcp-server/src/legal_mcp/services/halacha_quality.py#L282)).
זה **מגמד מספרית** את סוגיית-ה-`application`, ונוגע ישירות ל"רמה A = ניקוי-רעש" של הסינתזה:
**מהו ה"רעש" שמנקים?** הדגל הדומיננטי הוא `nli`, והוא מרוכז בפסיקה, לא בוועדות.
שתי השערות מתחרות, שצריך להכריע ביניהן **לפני** שמשתמשים ב-`nli` כמסנן-רעש ברמה A:
- **(א) הבודק מחמיר מדי** — סף-ה-NLI חותך יישורים לגיטימיים → 40% הם false-positives, וה"רעש" מדומה.
- **(ב) החילוץ-מהפסיקה לקוי** — ציטוטים תומכים שלא מיישרים לכלל → בעיית-חילוץ אמיתית בקנה-מידה.
ההכרעה משנה את כל אסטרטגיית רמה-A. **אנו ממליצים לאמת זאת על מדגם-זהב לפני כל שימוש ב-`nli` כסיגנל.**
---
## 6. דרכי-פעולה מוצעות (לסוכן-הקורפוס)
ממוין מהשמרני לאגרסיבי. ההמלצה שלנו: **B כברירת-מחדל + C כעבודה-מקבילה**; להימנע מ-A ומ-D.
| # | פעולה | טיעון | סיכון | המלצה |
|---|-------|-------|-------|-------|
| **A** | לכוונן את ה-prompt לדכא `application` מוועדות במקור | מטפל-בשורש (G1); חוסך 14% רשומות | **גבוה** — סותר את מבחן-8508; משמיד יישומי-לוסטרניק בני-ציטוט | ✗ לא |
| **B** | להשאיר את החילוץ; לסמוך על הניתוב הקיים (`application`→"תיקון", מחוץ לתור-היו"ר) + לדרג נמוך ב-RRF | תואם-סינתזה (שמור-בספק + דרג-בזמן-אחזור); אפס סיכון-אובדן | הרעש נשאר ב-DB (אחסון בלבד) | ✓ **כן — ברירת-מחדל** |
| **C** | לחקור קודם את `nli_unsupported` (40% פסיקה): מדגם-זהב, להכריע (א) מחמיר-מדי מול (ב) חילוץ-לקוי | זה הסיגנל הגדול; הכרחי לפני שמגדירים "רעש" ברמה A | דורש מדגם מתויג-ידנית | ✓ **כן — במקביל** |
| **D** | להפסיק חילוץ-עקרונות מוועדות לגמרי | רוב עקרוני-הוועדה הם שכתוב-persuasive של עליון | **קיצוני** — נוגד 07-learning §61 (ועדות ברות-ציטוט במכוון); נוגע INV-LRN | ✗ לא (אלא בהכרעת-יו"ר מפורשת) |
### ההמלצה המזוקקת
1. **לא לגעת בחילוץ-מוועדות** — הנתון מאשר שהנדיבות מוצדקת; `application`=יישום-בר-ציטוט, לא זבל. עקבי עם
הכרעת-הסינתזה "לא לחתוך".
2. **רמה B עושה את העבודה**`importance` boost ב-RRF מטביע את עקרוני-הוועדה ה-persuasive מתחת
לפסיקה-המחייבת, בלי למחוק דבר. 8508 (`importance≈0`) שוקע ממילא.
3. **להעביר את ה-`nli` לראש תור-המחקר** — לפני שמשתמשים בו כמסנן-רעש ברמה A, לאמת אם 40% אמיתי.
---
> **מקור-הנתונים:** `GET /api/halachot?limit=100000` (5,489 שורות) + per-case `?case_law_id=…`.
> ניתן לשחזר את כל המספרים מהשאילתות האלו. הקאנון הוא live — שיעורים ינועו ככל שהדריינר/פאנל רצים.

View File

@@ -40,7 +40,7 @@
3. **C — deep-read (נקודתי):** voice-XXXX.md — worked example לתיק-מופת. 3. **C — deep-read (נקודתי):** voice-XXXX.md — worked example לתיק-מופת.
### 0.3 הצינור החוזר per-final (7 שלבים) ### 0.3 הצינור החוזר per-final (7 שלבים)
`mark-final` → [1] INTAKE (snapshot של הטיוטה) → [2] PAIRING (בלוק↔בלוק) → [3] ALIGNMENT (diff פר-בלוק) → [4] DISTILLATION (מפריד סגנון↔מהות) → [5] CURATION (Hermes + שער-יו"ר) → [6] FEEDBACK (ניתוב לערוץ A/B/C) → [7] MEASUREMENT (מדד-מרחק-סגנון). `mark-final` → [1] INTAKE (snapshot של הטיוטה) → [2] PAIRING (בלוק↔בלוק) → [3] ALIGNMENT (diff פר-בלוק) → [4] DISTILLATION (מפריד סגנון↔מהות) → [5] CURATION (Hermes + שער-יו"ר) → [6] FEEDBACK (ניתוב לערוץ A/B/C) → [7] MEASUREMENT (מדד-מרחק-סגנון + snapshot-held-out פרוספקטיבי, §0.7).
### 0.4 ניהול ב-UI ### 0.4 ניהול ב-UI
`/methodology` = **עורך-הפרופיל היחיד** (declarative: יחסי-זהב, כללי-דיון, צ׳קליסטים, ביטויי-מעבר, אנטי-דפוסים, voice-invariants). `/training` = **שולחן-הלמידה** (קורפוס, פורטרט-סגנון, השוואת draft↔final, curator, מדד-מרחק, פנקס-התאמה). `/methodology` = **עורך-הפרופיל היחיד** (declarative: יחסי-זהב, כללי-דיון, צ׳קליסטים, ביטויי-מעבר, אנטי-דפוסים, voice-invariants). `/training` = **שולחן-הלמידה** (קורפוס, פורטרט-סגנון, השוואת draft↔final, curator, מדד-מרחק, פנקס-התאמה).
@@ -48,7 +48,7 @@
**שער-אישור אחד · טרנזקציית-כותב אחת (INV-IA3 → [X17](X17-information-architecture.md)):** ל-`decision_lesson` יש **סטטוס-יחיד** שקובע "זורם-לכותב" — `review_status='approved'` (INV-LRN1/G10). הדגל `applied_to_skill` **הוסר** (היה אינפורמטיבי-בלבד, נכתב-לשומקום → בלבל את היו"ר ב"שני שערים"; גל-2 #131). לקח שהיו"ר מחבר ידנית נוצר כבר כ-`approved`; לקח-פאנל נוצר כ-`proposed` וממתין לשער. promote של זוג draft↔final מטמיע את הלקחים/הביטויים שהיו"ר בחר **דרך appeal_type_rules בטרנזקציה אחת נעולה (FOR UPDATE)** — מסלול-כתיבה-יחיד, ללא read-modify-write מתפצל מול עורך-המתודולוגיה (MET-2/3, להלן G2 הפרות-ידועות). **שער-אישור אחד · טרנזקציית-כותב אחת (INV-IA3 → [X17](X17-information-architecture.md)):** ל-`decision_lesson` יש **סטטוס-יחיד** שקובע "זורם-לכותב" — `review_status='approved'` (INV-LRN1/G10). הדגל `applied_to_skill` **הוסר** (היה אינפורמטיבי-בלבד, נכתב-לשומקום → בלבל את היו"ר ב"שני שערים"; גל-2 #131). לקח שהיו"ר מחבר ידנית נוצר כבר כ-`approved`; לקח-פאנל נוצר כ-`proposed` וממתין לשער. promote של זוג draft↔final מטמיע את הלקחים/הביטויים שהיו"ר בחר **דרך appeal_type_rules בטרנזקציה אחת נעולה (FOR UPDATE)** — מסלול-כתיבה-יחיד, ללא read-modify-write מתפצל מול עורך-המתודולוגיה (MET-2/3, להלן G2 הפרות-ידועות).
### 0.5 Invariants חדשים ### 0.5 Invariants חדשים
**INV-LRN4 (ניגוד-אמת → G10/G9):** למידת-קול מבוססת **pairing draft↔final ברמת-בלוק**, לא קריאת-final בלבד. כל החלטה אינה "סגורה" עד שהושוותה מול הסופי; כל סופי מנותח מול הטיוטה. נשמר פנקס-התאמה (`draft_final_pairs`) עם מצב-חיים `draft_done → final_received → analyzed → lessons_folded`. **INV-LRN4 (ניגוד-אמת → G10/G9):** למידת-קול מבוססת **pairing draft↔final ברמת-בלוק**, לא קריאת-final בלבד. כל החלטה אינה "סגורה" עד שהושוותה מול הסופי; כל סופי מנותח מול הטיוטה. נשמר פנקס-התאמה (`draft_final_pairs`) עם מצב-חיים `draft_done → final_received → analyzed → lessons_folded`. ההכללה (האם הלמידה משפרת תיקים שלא-נלמדו) נמדדת פרוספקטיבית — §0.7.
*מקורות:* imitation-learning-from-expert-edits · contrastive personalization (arxiv 2504.08745) · author-profiling. *סטטוס: verified.* *מקורות:* imitation-learning-from-expert-edits · contrastive personalization (arxiv 2504.08745) · author-profiling. *סטטוס: verified.*
**INV-LRN5 (טוהר-הקול → G4/G11):** שכבת-ידע-הקול (voice-fingerprint, style_patterns, exemplars) **לא תכיל הלכות/עובדות ספציפיות** — רק סגנון ושיטה. מהות מנותבת ל-precedent_library/halacha. ה-distillation מפריד במקור. **INV-LRN5 (טוהר-הקול → G4/G11):** שכבת-ידע-הקול (voice-fingerprint, style_patterns, exemplars) **לא תכיל הלכות/עובדות ספציפיות** — רק סגנון ושיטה. מהות מנותבת ל-precedent_library/halacha. ה-distillation מפריד במקור.
@@ -62,9 +62,18 @@
3. **בדיקת-ציטוטים**`extract_internal_citations` מקשר את הפסיקה שההחלטה מצטטת לספרייה; כל ציטוט שאינו בספרייה **מסומן אוטומטית** כ-`missing_precedent` (open) להעלאה ע"י היו"ר. 3. **בדיקת-ציטוטים**`extract_internal_citations` מקשר את הפסיקה שההחלטה מצטטת לספרייה; כל ציטוט שאינו בספרייה **מסומן אוטומטית** כ-`missing_precedent` (open) להעלאה ע"י היו"ר.
4. הציטוטים-המקושרים מזינים את **לולאת-ה-corroboration** (X11): ציטוט-נכנס מההחלטה שלנו מחזק את ההלכות של התקדים המצוטט (`corroboration_rebuild`). 4. הציטוטים-המקושרים מזינים את **לולאת-ה-corroboration** (X11): ציטוט-נכנס מההחלטה שלנו מחזק את ההלכות של התקדים המצוטט (`corroboration_rebuild`).
ואז שני שלבים אוטומטיים נפרדים (`run-learning` / `run-halacha`) המעירים worker מקומי (claude/DeepSeek/Gemini מקומיים בלבד): ואז שני שלבים אוטומטיים נפרדים (`run-learning` / `run-halacha`) המעירים worker מקומי (claude/DeepSeek/Gemini מקומיים בלבד):
- **למידה:** `ingest_final_version` (Opus distillation) → **פאנל-סגנון דו-סוכני** (DeepSeek+Gemini, "למידה כפולה") שמצביע על כל לקח-style_method; הסכמה 2/2 → `decision_lesson` (`source=panel:deepseek+gemini`); פיצול → ליו"ר. - **למידה:** הצינור (`final_learning_pipeline.py`) רץ בסדר `enroll_style_corpus` (יצירת רשומת-הקורפוס, מהיר — **ראשון** מאז #159 כדי שהקורפוס יהיה זמין לפני הדיסטילציה הארוכה) → `ingest_final_version` (Opus distillation) → **פאנל-סגנון דו-סוכני** (DeepSeek+Gemini, "למידה כפולה") שמצביע על כל לקח-style_method; הסכמה 2/2 → `decision_lesson` (`source=panel:deepseek+gemini`) **שזורם אוטומטית לכותב** כ-`approved` (שער-מדורג, INV-LRN1) — הפיך (veto-יו"ר ב-/training); פיצול → ליו"ר. **במקביל** (יקיצת `final_learning_*`) האוצֵר ממשיך ל-§A ורושם ממצאי-`source='curator'` (ערוץ ג׳, §1.1).
- **הלכות:** `extract_internal_citations``precedent_extract_halachot``corroboration_rebuild`**פאנל-הלכות תלת-סוכני** (`halacha_panel_approve.py --apply`). - **הלכות:** `extract_internal_citations``precedent_extract_halachot``corroboration_rebuild`**פאנל-הלכות תלת-סוכני** (`halacha_panel_approve.py --apply`).
שני הפאנלים **הפיכים** (גיבוי-CSV ל-`data/audit/`) ומסלימים מחלוקות. ההטמעה הסופית ל-`SKILL.md`/`legal-decision-lessons.md` נשארת **אישור-יו"ר ידני** (INV-LRN1/G10) — הפאנל יוצר *הצעות* בלבד. שני הפאנלים **הפיכים** (גיבוי-CSV ל-`data/audit/`) ומסלימים מחלוקות. ההטמעה ל-`SKILL.md`/`legal-decision-lessons.md` ולכל **מהות** נשארת **אישור-יו"ר ידני קשיח** (INV-LRN1/G10); לקחי-**סגנון** בקונצנזוס זורמים אוטומטית-והפיך לכותב.
### 0.7 מדד-ההכללה הפרוספקטיבי (held-out trend — "מסלול A")
**השאלה שזה עונה:** האם הלמידה באמת *מכלילה* — כלומר משפרת טיוטות של תיקים **שלא נלמדו** — ולא רק "משננת" תיקים שכבר ראינו. מדד-מרחק-הסגנון (שלב [7]) על תיק שלקחיו כבר הוטמעו אינו held-out; מדידה רטרואקטיבית בלתי-אפשרית כי הלקחים נשמרים `appeal_type_rules` (`universal`) **ללא תיוג-מקור** → אין leave-one-out. לכן המדידה **פרוספקטיבית**: נלכדת ברגע היחיד שבו היא נקייה.
**המנגנון (אוטומטי, ב-`final/upload`, ללא LLM):** מיד אחרי פתיחת הזוג `draft_final_pair` אבל **לפני** הטמעת-לקחי-התיק (ה-fold הוא שלב-`promote` ידני נפרד ב-`/training`, §0.4), נלכד snapshot של `style_distance` (`anti_pattern_total`, סטיית-יחסי-זהב, `change_percent`) יחד עם **גודל-בריכת-הלקחים** באותו רגע (מספר `discussion_rules`+`transition_phrases` ב-`universal`). מכיוון שהטיוטה נכתבה עם הבריכה ה**קודמת** בלבד, כל שורה היא נקודת-נתון held-out נקייה: *"עם N לקחים מצטברים, הטיוטה שלנו על תיק שלא-נראה קיבלה ציון X"*.
**הקריאה:** טבלת `style_distance_history` (append-only) + `GET /api/learning/style-distance-history`. ירידה ב-`anti_pattern_total`/`change_percent` ככל ש-N גדל = **הוכחה מתגלגלת שהלמידה מכלילה** (INV-LRN4 — זהו משטח-המגמה של "ניגוד-האמת"). **אזהרת-פרשנות:** `change_percent` מערבב סגנון עם שלמות-תוכן (לפעמים היו"ר מכפילה אורך כי חסרה מהות, לפעמים חותכת) → `anti_pattern_total` הוא הסיגנל הנקי-יותר לסגנון. שימוש-חוזר בשירות `style_distance` ובבריכת `appeal_type_rules` — אין מסלול-מדד מקביל (G2).
> **חלון נקי חד-פעמי:** תיק שכבר `lessons_folded` פספס את חלון ה-held-out שלו — אין backfill. הטבלה מתמלאת קדימה מהסופי הבא.
--- ---
@@ -87,10 +96,24 @@
(`hermes-curator.md:60-70`). (`hermes-curator.md:60-70`).
- מזהה **35 דפוסים/פערים** חדשים, כל ממצא מתויג `[סגנון]` / `[מבנה]` / - מזהה **35 דפוסים/פערים** חדשים, כל ממצא מתויג `[סגנון]` / `[מבנה]` /
`[לקסיקון משפטי]` / `[טבלאי]` (`hermes-curator.md:99-108`). `[לקסיקון משפטי]` / `[טבלאי]` (`hermes-curator.md:99-108`).
- **מציע** — comment ב-Paperclip + רישום כל ממצא כ-`decision_lesson` דרך - **מציע** — comment + interaction בערוץ-הפלטפורמה, וגם **רושם כל ממצא מבנית** כ-`decision_lesson`
`POST /api/training/corpus/{corpus_id}/lessons` (`source:"curator"`) שמופיע ב-UI (`source='curator'`, `review_status='proposed'`) דרך כלי-ה-MCP `record_curator_findings`
תחת הטאב "מה למדנו" (`hermes-curator.md:73-96`). (`hermes-curator.md` §A.5b → `tools/workflow.py::record_curator_findings``db.add_decision_lesson`).
- **אינו מעדכן** קבצים בעצמו (skills/, lessons.py, DB) — רק מציע (`hermes-curator.md:125-130`). הממצאים מופיעים ב-`/training` טאב "אוצֵר" (ערוץ ג׳) ועוברים את שער-היו"ר (INV-LRN1).
> **דריפט-מימוש שנסגר (2026-06-28):** עד גל זה הסוכן כתב comment בלבד והספ תיאר endpoint שלא נקרא —
> הממצאים האיכותיים אבדו, וכרטיס-האוצֵר (`get_curator_stats`) ספר `source='curator'` שאיש לא כתב → 0.
> `record_curator_findings` מממש את INV-LRN3 בפועל.
- **אינו מאשר ואינו מטמיע** — read-only על התוכן; רישום-הממצא הוא *הצעה* מגודרת-שער, לא שינוי-קול.
אינו עורך `skills/` / `lessons.py` / שכבת-הקול בעצמו (`hermes-curator.md` §"כללים כלליים").
- **שני ערוצי-למידה נפרדים על אותו סופי:** ערוץ-האוצֵר (כאן, qualitative, `source='curator'`) ו**פאנל-הסגנון**
הדו-סוכני שהצינור מפעיל (§0.6 — DeepSeek+Gemini, `source='panel:deepseek+gemini'`). שניהם
`decision_lessons` ממתיני-שער; ערוץ הדיסטילציה (`appeal_type_rules`, §0.4) הוא השלישי.
> **נסגר (TaskMaster #159, 2026-06-28):** ה-`PIPELINE-WAKE BRANCH` (`hermes-curator.md`) מבדיל כעת
> `final_learning_*` מ-`final_halacha_*`: על **learning** הסוכן מפעיל את הצינור ברקע ו**ממשיך ל-§A**
> (מצב AUTO — רושם ממצאי-`curator`, מדלג על interaction), על **halacha** יוצא מיד כקודם. המירוץ מול
> `enroll_style_corpus` נסגר בכך ש-enroll רץ **ראשון** בצינור (§0.6) — הקורפוס קיים תוך שניות, הרבה לפני
> שהסוכן/ה-§A מסיים את ניתוח-ה-LLM; `record_curator_findings` כולל גם retry כרשת-ביטחון.
### 1.2 לולאת-פידבק-היו"ר (capture → ניתוח שבועי → לקחים) ### 1.2 לולאת-פידבק-היו"ר (capture → ניתוח שבועי → לקחים)
@@ -162,13 +185,12 @@
## 3. Invariants של התחום ## 3. Invariants של התחום
### INV-LRN1: עדכון-ידע דורש אישור-יוידני — אין auto-commit (governance →G10) ### INV-LRN1: עדכון-ידע דורש שער-יו— **שער מדורג** לפי סיכון (governance →G10)
**כלל:** מנגנוני-הלמידה (Hermes, ניתוח-פידבק שבועי) **מציעים בלבד**. כל שינוי ב- **כלל (מעודכן 2026-06-28, הכרעת-יו"ר):** השער **מדורג לפי סיכון-התוכן**, לא אחיד:
[SKILL.md](../../skills/decision/SKILL.md) או ב-[legal-decision-lessons.md](../legal-decision-lessons.md) - **מהות** (הלכה / תקדים / עובדה / כל שינוי ב-[SKILL.md](../../skills/decision/SKILL.md) או ב-[legal-decision-lessons.md](../legal-decision-lessons.md)) → **שער-קשיח**: בחינה ואישור ידניים של היו"ר/חיים ואז commit ידני — **לעולם לא auto-committed**. פאנל-ההלכות התלת-סוכני מציע בלבד.
מחייב **בחינה ואישור ידניים של היו"ר/חיים** ואז commit ידני — **לעולם לא auto-committed**. - **סגנון** (`decision_lessons` בקטגוריות style/structure/lexicon/tabular, וכן discussion_rules / transition_phrases / anti_patterns) שעבר **קונצנזוס-פאנל 2/2****זורם אוטומטית לכותב** (`review_status='approved'`), **הפיך**: היו"ר רואה ויכול **לבטל בדיעבד** ב-/training. זהו שינוי-קול נמוך-סיכון (טהור-מהות לפי INV-LRN5), בקרת-איכות מובנית (2 מודלים), והפיכוּת — ולכן עדיין "תחת בקרת-המשתמש" במובן NCSC/CEPEJ (veto, לא אישור-מראש על כל פריט).
Hermes כותב comment + `decision_lesson`, לא קבצים; ה-CEO השבועי כותב לקובץ אך הצעותיו
מאומתות ידנית לפני קיבוע. זהו פֶּאֶט של [INV-G10](00-constitution.md#inv-g10-המערכת-מסייעת--שערים-אנושיים-הם-invariant) זהו פֶּאֶט של [INV-G10](00-constitution.md#inv-g10-המערכת-מסייעת--שערים-אנושיים-הם-invariant) על שכבת-הידע: הלמידה כפופה לשיקול-הדעת האנושי — קשיח למהות, הפיך-veto לסגנון. **מימוש:** `scripts/style_lesson_panel.py:_review_status_for` קובע `approved` לקטגוריות-סגנון בקונצנזוס, `proposed` אחרת; הכותב צורך רק `approved` ([db.get_recent_decision_lessons]).
על שכבת-הידע: גם הלמידה כפופה לשיקול-הדעת האנושי.
**מקורות:** NCSC/JTC — *Principles & Practices for AI Use in Courts* (human-in-the-loop; **מקורות:** NCSC/JTC — *Principles & Practices for AI Use in Courts* (human-in-the-loop;
never replace human judgment) · Council of Europe / CEPEJ (2018, under user control) · never replace human judgment) · Council of Europe / CEPEJ (2018, under user control) ·
Federal Judicial Center — *Judicial Writing Manual* (2d ed.) | סטטוס: verified Federal Judicial Center — *Judicial Writing Manual* (2d ed.) | סטטוס: verified

View File

@@ -35,14 +35,14 @@
| `case_number` (citation) | CHAIR (חובה) | מפתח idempotency | | `case_number` (citation) | CHAIR (חובה) | מפתח idempotency |
| `full_text`, `extraction_status`, `source_kind` | DETERMINISTIC | — | | `full_text`, `extraction_status`, `source_kind` | DETERMINISTIC | — |
| `case_name`, `court`, `date`, `headnote`, `summary`, `key_quote`, `subject_tags`, `appeal_subtype`, `precedent_level`, `source_type`, `citation_formatted` | CHAIR או OPUS | Opus ממלא רק אם ריק | | `case_name`, `court`, `date`, `headnote`, `summary`, `key_quote`, `subject_tags`, `appeal_subtype`, `precedent_level`, `source_type`, `citation_formatted` | CHAIR או OPUS | Opus ממלא רק אם ריק |
| `is_binding` | CHAIR (default true) | קובע prompt-הלכה | | `is_binding` | CHAIR (default true) **אך** DETERMINISTIC-false ל-`source_type=appeals_committee` | קובע prompt-הלכה. סמכות מבנית ([INV-DM7](02-data-model.md#inv-dm7-סיווג-הלכה--סמכות-נגזרת--תפקיד-כלל-מסווג-שני-צירים-לא-enum-אחד)): מקור-ועדה הוא תמיד persuasive — נכפה ל-false ב-`create_external_case_law` ונאכף ע"י `case_law_committee_not_binding_check` |
| chunks (`content`/`section_type`/`page_number`) | DETERMINISTIC | — | | chunks (`content`/`section_type`/`page_number`) | DETERMINISTIC | — |
| `embedding` (chunks) | Voyage (לא-LLM-reasoning) | ⚠ לא-GENERATED ([gap-audit GAP-09](gap-audit.md)) | | `embedding` (chunks) | Voyage (לא-LLM-reasoning) | ⚠ לא-GENERATED ([gap-audit GAP-09](gap-audit.md)) |
| כל `halachot` | OPUS | נכנס pending_review | | כל `halachot` | OPUS | נכנס pending_review |
### 2ב. החלטה פנימית (`case_law`, source_kind=`internal_committee`) ### 2ב. החלטה פנימית (`case_law`, source_kind=`internal_committee`)
כמו 2א, ובנוסף: `case_number` **חובה**; `chair_name`/`district`/`proceeding_type` — CHAIR או OPUS או DERIVED; כמו 2א, ובנוסף: `case_number` **חובה**; `chair_name`/`district`/`proceeding_type` — CHAIR או OPUS או DERIVED;
`source_type` = `appeals_committee` (DETERMINISTIC קבוע). placeholder `"(טרם חולץ)"` מסומן ל-chair_name/district `source_type` = `appeals_committee` (DETERMINISTIC קבוע); `is_binding` = `false` (DETERMINISTIC קבוע — נכפה ב-`create_internal_committee_decision`, [INV-DM7](02-data-model.md#inv-dm7-סיווג-הלכה--סמכות-נגזרת--תפקיד-כלל-מסווג-שני-צירים-לא-enum-אחד): החלטת-ועדה אינה מחייבת ועדה אחרת ולא את עצמה — תמיד persuasive). placeholder `"(טרם חולץ)"` מסומן ל-chair_name/district
ריקים ומטופל כריק ע"י ה-extractor. ריקים ומטופל כריק ע"י ה-extractor.
### 2ג. מסמך-תיק (`documents`) ### 2ג. מסמך-תיק (`documents`)

View File

@@ -0,0 +1,348 @@
# Hooks: Case Status Webhooks Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** כאשר status של תיק משתנה ב-legal-ai (e.g. `qa_failed`, `exported`), ה-plugin מקבל webhook, מעדכן את ה-issue ב-Paperclip, ומעיר את ה-CEO במקרה הצורך.
**Architecture:** Legal-ai REST API קורא `pc_request("POST", "/api/plugins/marcusgroup.legal-ai/webhooks/case-status", ...)` אחרי כל שינוי status. ה-plugin מטפל ב-`onWebhook()` ומגיב: מוסיף תגובה לissue, מעיר CEO אם QA נכשל.
**Tech Stack:** TypeScript (plugin-legal-ai), Python/FastAPI (legal-ai web), `@paperclipai/plugin-sdk@2026.325.0`, `httpx` (Python), `pc_request` helper.
---
## File Map
| Action | File |
|--------|------|
| Modify | `plugin-legal-ai/src/worker.ts` — add `onWebhook()` to `definePlugin({})` |
| Modify | `plugin-legal-ai/plugin.json` — add `"webhooks.receive"` capability |
| Modify | `plugin-legal-ai/src/manifest.ts` — add webhook capability |
| Modify | `legal-ai/web/app.py` — emit webhook after `PUT /api/cases/{case_number}` |
| Modify | `legal-ai/web/paperclip_api.py` — add `emit_webhook()` helper |
---
## Task 1: Add `emit_webhook` helper ב-Python
**Files:**
- Modify: `legal-ai/web/paperclip_api.py`
- [ ] **Step 1: קרא את הקובץ הקיים**
```bash
head -90 /home/chaim/legal-ai/web/paperclip_api.py
```
- [ ] **Step 2: הוסף את ה-helper בסוף הקובץ**
פתח `/home/chaim/legal-ai/web/paperclip_api.py` והוסף אחרי הפונקציה `pc_request`:
```python
async def emit_case_status_webhook(
case_number: str,
old_status: str,
new_status: str,
company_id: str | None = None,
run_id: str | None = None,
) -> None:
"""Notify the Paperclip plugin that a case status changed.
Fire-and-forget: logs errors but never raises, so callers aren't blocked.
"""
try:
await pc_request(
"POST",
"/api/plugins/marcusgroup.legal-ai/webhooks/case-status",
json={
"caseNumber": case_number,
"oldStatus": old_status,
"newStatus": new_status,
"companyId": company_id,
"timestamp": datetime.utcnow().isoformat() + "Z",
},
run_id=run_id,
timeout=5.0,
)
except Exception as exc:
logger.warning("emit_case_status_webhook failed: %s", exc)
```
> **הערה:** `datetime` ו-`logger` כבר מיובאים ב-`app.py`. בדוק שהם מיובאים גם ב-`paperclip_api.py` — אם לא, הוסף `from datetime import datetime` ו-`import logging; logger = logging.getLogger(__name__)` בראש הקובץ.
- [ ] **Step 3: Commit**
```bash
cd /home/chaim/legal-ai
git add web/paperclip_api.py
git commit -m "feat: add emit_case_status_webhook helper"
```
---
## Task 2: צרף webhook לendpoint `PUT /api/cases/{case_number}`
**Files:**
- Modify: `legal-ai/web/app.py`
- [ ] **Step 1: מצא את ה-endpoint**
```bash
grep -n "PUT\|case_number\|update_case" /home/chaim/legal-ai/web/app.py | head -30
```
- [ ] **Step 2: קרא את הenable endpoint המלא**
זהה את הסקציה המלאה של ה-endpoint ואת הייבוא הקיים.
- [ ] **Step 3: הוסף import ל-emit_webhook**
בראש `app.py`, בסקציית ה-imports מ-`paperclip_api`:
```python
from .paperclip_api import pc_request, emit_case_status_webhook
```
- [ ] **Step 4: הוסף webhook emit בתוך הendpoint**
אחרי שהקוד מעדכן את ה-case (לפני ה-`return`), הוסף:
```python
# Notify plugin about status change (fire-and-forget)
if updates.get("status") and old_status != updates["status"]:
background_tasks.add_task(
emit_case_status_webhook,
case_number=case_number,
old_status=old_status,
new_status=updates["status"],
company_id=str(case.get("company_id")),
)
```
> אם ה-endpoint כבר מקבל `background_tasks: BackgroundTasks` — השתמש בו. אם לא, הוסף `background_tasks: BackgroundTasks` לחתימת הפונקציה. הוסף `from fastapi import BackgroundTasks` ל-imports.
- [ ] **Step 5: שמור את ה-`old_status` לפני ה-update**
בתחילת ה-endpoint handler, לפני קריאת ה-DB update:
```python
old_status = (await db.get_case(case_number) or {}).get("status", "")
```
- [ ] **Step 6: בדיקה בסיסית — שלח PUT ידני**
```bash
curl -s -X PUT https://legal-ai.nautilus.marcusgroup.org/api/cases/1130-25 \
-H "Content-Type: application/json" \
-d '{"status": "in_progress"}' | jq .status
```
בדוק ב-Paperclip logs שהwebhook נשלח (עדיין לא מטופל בצד ה-plugin):
```bash
pm2 logs paperclip --lines 20
```
- [ ] **Step 7: Commit**
```bash
cd /home/chaim/legal-ai
git add web/app.py
git commit -m "feat: emit case-status webhook on PUT /api/cases/:case"
```
---
## Task 3: הוסף `onWebhook()` ל-plugin
**Files:**
- Modify: `plugin-legal-ai/src/worker.ts`
- [ ] **Step 1: קרא את `definePlugin({})` הקיים**
```bash
grep -n "definePlugin\|onWebhook\|onHealth\|onShutdown" /home/chaim/plugin-legal-ai/src/worker.ts
```
- [ ] **Step 2: הוסף את `onWebhook` handler**
בתוך הobject שמועבר ל-`definePlugin({})`, אחרי `setup(ctx)`:
```typescript
onWebhook: async (input) => {
const { endpointKey, payload, companyId } = input as {
endpointKey: string;
payload: {
caseNumber: string;
oldStatus: string;
newStatus: string;
companyId: string;
timestamp: string;
};
companyId: string;
};
if (endpointKey !== "case-status") return;
const { caseNumber, oldStatus, newStatus } = payload;
ctx.logger.info(`Webhook: ${caseNumber} ${oldStatus}${newStatus}`);
// Find the Paperclip issue linked to this case
const stateKey = `case:${caseNumber}`;
const issueId = await ctx.state.get({ companyId }, stateKey);
if (!issueId) {
ctx.logger.warn(`No issue found for case ${caseNumber}`);
return;
}
const statusLabels: Record<string, string> = {
in_progress: "🔄 בעבודה",
drafted: "✍️ טיוטה מוכנה",
qa_failed: "❌ QA נכשל",
exported: "📄 יוצא ל-DOCX",
reviewed: "✅ נבדק",
final: "🎯 סופי",
};
const label = statusLabels[newStatus] ?? newStatus;
// Post a status comment on the issue
await ctx.issues.createComment({
issueId: issueId as string,
body: `**עדכון סטטוס:** ${label} (היה: ${oldStatus})`,
});
// Wake CEO if QA failed
if (newStatus === "qa_failed") {
const companies = await ctx.companies.list();
const company = companies.find((c) => c.id === companyId);
if (!company) return;
const CEO_IDS: Record<string, string> = {
"42a7acd0-30c5-4cbd-ac97-7424f65df294": "752cebdd-6748-4a04-aacd-c7ab0294ef33",
"8639e837-4c9d-47fa-a76b-95788d651896": "cdbfa8bc-3d61-41a4-a2e7-677ec7d34562",
};
const ceoId = CEO_IDS[companyId];
if (ceoId) {
await ctx.agents.invoke(ceoId, companyId, {
prompt: `תיק ${caseNumber} נכשל בבדיקת QA. עיין בתוצאות QA ותקן את הבעיות.`,
reason: "qa_failed webhook",
});
}
}
},
```
> **הערה:** `ctx` חייב להיות נגיש ב-`onWebhook`. אם `ctx` מוגדר בתוך `setup()` בלבד — הוצא אותו ל-closure חיצוני של ה-plugin object (ראה את הדפוס הקיים ב-`worker.ts`).
- [ ] **Step 3: בדק TypeScript**
```bash
cd /home/chaim/plugin-legal-ai && npx tsc --noEmit
```
Expected: 0 errors.
- [ ] **Step 4: Build**
```bash
cd /home/chaim/plugin-legal-ai && npm run build
```
Expected: `dist/worker.js` נוצר ללא שגיאות.
- [ ] **Step 5: Commit**
```bash
cd /home/chaim/plugin-legal-ai
git add src/worker.ts
git commit -m "feat: add onWebhook handler for case-status events"
```
---
## Task 4: הוסף capability ל-`plugin.json` ול-`manifest.ts`
**Files:**
- Modify: `plugin-legal-ai/plugin.json`
- Modify: `plugin-legal-ai/src/manifest.ts`
- [ ] **Step 1: הוסף `"webhooks.receive"` ל-capabilities**
ב-`plugin-legal-ai/plugin.json`, בarray `"capabilities"`, הוסף:
```json
"webhooks.receive"
```
- [ ] **Step 2: הוסף גם ב-`manifest.ts`**
```bash
grep -n "capabilities\|webhooks" /home/chaim/plugin-legal-ai/src/manifest.ts
```
הוסף `"webhooks.receive"` לarray שם.
- [ ] **Step 3: Re-install plugin**
```bash
cd /home/chaim/plugin-legal-ai && npm run build
npx paperclipai plugin uninstall marcusgroup.legal-ai \
--api-base http://localhost:3100 --api-key pcapi_legal_install_key_2026
npx paperclipai plugin install /home/chaim/plugin-legal-ai \
--api-base http://localhost:3100 --api-key pcapi_legal_install_key_2026
pm2 restart paperclip
```
- [ ] **Step 4: Commit**
```bash
cd /home/chaim/plugin-legal-ai
git add plugin.json src/manifest.ts
git commit -m "feat: add webhooks.receive capability to plugin manifest"
```
---
## Task 5: בדיקה end-to-end
- [ ] **Step 1: Deploy legal-ai**
```bash
cd /home/chaim/legal-ai
git push origin main
# המתן לבנייה (~2-4 דקות)
```
- [ ] **Step 2: שנה סטטוס תיק**
```bash
curl -s -X PUT https://legal-ai.nautilus.marcusgroup.org/api/cases/1130-25 \
-H "Content-Type: application/json" \
-d '{"status": "qa_failed"}' | jq .status
```
- [ ] **Step 3: בדוק שהתגובה הוספה ל-Paperclip issue**
```bash
# מצא את ה-issue הקשור לתיק 1130-25
pm2 logs paperclip --lines 30 | grep "1130-25\|webhook\|qa_failed"
```
Expected: תגובה "❌ QA נכשל" הוספה לissue. CEO הועיר.
- [ ] **Step 4: בדוק שינוי סטטוס שגרתי (לא QA)**
```bash
curl -s -X PUT https://legal-ai.nautilus.marcusgroup.org/api/cases/1130-25 \
-H "Content-Type: application/json" \
-d '{"status": "drafted"}' | jq .status
```
Expected: תגובה "✍️ טיוטה מוכנה" בissue. CEO **לא** הועיר.
---
## אימות סופי
| בדיקה | פקודה | תוצאה מצופה |
|-------|-------|-------------|
| QA נכשל → CEO מועיר | `PUT status=qa_failed` | תגובה + agent invocation |
| exported → תגובה בלבד | `PUT status=exported` | תגובה בלבד |
| שינוי ללא status | `PUT title=...` | שום webhook |
| תיק ללא issue | webhook לתיק חדש | לוג warning, ללא crash |

View File

@@ -0,0 +1,306 @@
# Per-Agent CLAUDE.md Versioning & Validation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** הוסף validation ו-version tracking לקבצי הוראות הסוכנים. כעת, לפני שה-sync script מחיל שינויים, הוא מאמת שכל `instructionsFilePath` קיים. בנוסף, metadata של הסוכן ב-DB יכיל `claude_md_mtime` — השינוי האחרון בקובץ — כדי לזהות drift.
**Architecture:** `sync_agents_across_companies.py` מקבל `--check-instructions` flag שסורק את כל הסוכנים ומדווח על קבצים חסרים/ישנים. ב-`--apply`, מתווספת בדיקת pre-flight שמבטלת את הsync אם קובץ חסר. `agents.metadata` מקבל `claude_md_mtime` עם ה-mtime בפועל של הקובץ.
**Tech Stack:** Python 3.10+, asyncpg, httpx, `os.path.getmtime()`, Paperclip REST API (`PATCH /api/agents/{id}`).
---
## File Map
| Action | File |
|--------|------|
| Modify | `legal-ai/scripts/sync_agents_across_companies.py``--check-instructions` flag, pre-flight, metadata update |
זה הקובץ היחיד שצריך לגעת בו. כל שאר הלוגיקה קיימת כבר.
---
## Task 1: הוסף `--check-instructions` flag
**Files:**
- Modify: `legal-ai/scripts/sync_agents_across_companies.py`
- [ ] **Step 1: קרא את החלק של `argparse` בסקריפט**
```bash
grep -n "argparse\|add_argument\|--verify\|--dry-run\|--apply" \
/home/chaim/legal-ai/scripts/sync_agents_across_companies.py | head -20
```
- [ ] **Step 2: הוסף את הargument**
מצא את הסקציה שמגדירה `args` והוסף:
```python
parser.add_argument(
"--check-instructions",
action="store_true",
help="Scan all agents' instructionsFilePath and report missing/outdated files",
)
```
- [ ] **Step 3: הוסף את הפונקציה `check_instructions()`**
הוסף לפני `async def main()`:
```python
async def check_instructions(agents: list[dict]) -> bool:
"""Print a report of all agents' instruction files. Returns True if all OK."""
all_ok = True
print(f"\n{'Agent':<30} {'File':<60} {'Status':<15} {'Size':>8} {'Modified'}")
print("-" * 120)
for agent in agents:
adapter_cfg = agent.get("adapter_config") or {}
if isinstance(adapter_cfg, str):
import json as _json
adapter_cfg = _json.loads(adapter_cfg)
file_path = adapter_cfg.get("instructionsFilePath", "")
name = agent.get("name", agent.get("id", "?"))[:29]
if not file_path:
print(f"{name:<30} {'(none)':<60} {'⚠️ NOT SET':<15}")
continue
if not os.path.exists(file_path):
print(f"{name:<30} {file_path[-59:]:<60} {'❌ MISSING':<15}")
all_ok = False
continue
stat = os.stat(file_path)
size_kb = stat.st_size // 1024
mtime = datetime.fromtimestamp(stat.st_mtime).strftime("%Y-%m-%d %H:%M")
# Compare with DB metadata
metadata = agent.get("metadata") or {}
if isinstance(metadata, str):
import json as _json
metadata = _json.loads(metadata)
db_mtime = metadata.get("claude_md_mtime", "")
actual_mtime = str(int(stat.st_mtime))
drift = " ⚠️ DRIFT" if db_mtime and db_mtime != actual_mtime else ""
print(f"{name:<30} {file_path[-59:]:<60} {'✅ OK':<15} {size_kb:>6}KB {mtime}{drift}")
print()
return all_ok
```
> `from datetime import datetime` ו-`import os` — בדוק שמיובאים בראש הסקריפט. אם לא, הוסף.
- [ ] **Step 4: הוסף קריאה ל-`check_instructions()` ב-`main()`**
בתוך `async def main()`, אחרי שloading הagents מה-DB:
```python
if args.check_instructions:
all_ok = await check_instructions(master_agents + mirror_agents)
sys.exit(0 if all_ok else 1)
```
- [ ] **Step 5: בדיקה**
```bash
cd /home/chaim/legal-ai
python scripts/sync_agents_across_companies.py --check-instructions
```
Expected: טבלה עם כל הסוכנים, paths, סטטוס ✅/❌.
- [ ] **Step 6: Commit**
```bash
cd /home/chaim/legal-ai
git add scripts/sync_agents_across_companies.py
git commit -m "feat: add --check-instructions flag to sync script"
```
---
## Task 2: הוסף pre-flight validation לפני `--apply`
**Files:**
- Modify: `legal-ai/scripts/sync_agents_across_companies.py`
- [ ] **Step 1: מצא את נקודת הכניסה של `--apply`**
```bash
grep -n "args.apply\|if.*apply\|apply.*mode" \
/home/chaim/legal-ai/scripts/sync_agents_across_companies.py | head -10
```
- [ ] **Step 2: הוסף pre-flight לפני apply**
בתחילת בלוק `--apply`, לפני כל שינוי:
```python
if args.apply:
# Pre-flight: abort if any agent is missing its instructions file
print("🔍 Pre-flight: checking instruction files...")
all_ok = await check_instructions(master_agents + mirror_agents)
if not all_ok:
print("❌ Abort: one or more instruction files are missing. Fix paths before --apply.")
sys.exit(1)
print("✅ Pre-flight passed.\n")
# ... rest of apply logic ...
```
- [ ] **Step 3: בדיקה — הפעל עם קובץ חסר (סימולציה)**
```bash
# שנה זמנית path לקובץ שלא קיים
cd /home/chaim/legal-ai
python scripts/sync_agents_across_companies.py --dry-run 2>&1 | head -5
```
Expected: dry-run עובר. אם תנסה `--apply` עם agent שhis file חסר — הsync יבוטל.
- [ ] **Step 4: Commit**
```bash
cd /home/chaim/legal-ai
git add scripts/sync_agents_across_companies.py
git commit -m "feat: add pre-flight instruction file validation before --apply"
```
---
## Task 3: עדכן `agents.metadata` עם `claude_md_mtime`
**Files:**
- Modify: `legal-ai/scripts/sync_agents_across_companies.py`
- [ ] **Step 1: מצא את `compute_diff()` או `apply_diff()`**
```bash
grep -n "def compute_diff\|def apply_diff\|def build_patch\|metadata" \
/home/chaim/legal-ai/scripts/sync_agents_across_companies.py | head -20
```
- [ ] **Step 2: הוסף פונקציה `get_claude_md_mtime()`**
```python
def get_claude_md_mtime(adapter_config: dict) -> str | None:
"""Return the Unix mtime of the agent's instructionsFilePath, or None if missing."""
path = adapter_config.get("instructionsFilePath", "")
if not path or not os.path.exists(path):
return None
return str(int(os.path.getmtime(path)))
```
- [ ] **Step 3: שלב mtime בbuild של metadata patch**
מצא את המקום שמכין את ה-`metadata` לsync. הוסף:
```python
# Build metadata patch with claude_md_mtime
current_metadata = master_agent.get("metadata") or {}
if isinstance(current_metadata, str):
import json as _json
current_metadata = _json.loads(current_metadata)
adapter_cfg = master_agent.get("adapter_config") or {}
if isinstance(adapter_cfg, str):
import json as _json
adapter_cfg = _json.loads(adapter_cfg)
mtime = get_claude_md_mtime(adapter_cfg)
if mtime:
current_metadata["claude_md_mtime"] = mtime
current_metadata["claude_md_last_synced"] = datetime.utcnow().isoformat() + "Z"
```
כלול את ה-`current_metadata` המעודכן ב-PATCH לAPI.
- [ ] **Step 4: בדוק שה-metadata מתעדכן**
הפעל `--dry-run` וחפש `claude_md_mtime` בoutput:
```bash
python scripts/sync_agents_across_companies.py --dry-run 2>&1 | grep -i "mtime\|metadata" | head -10
```
לאחר `--apply`, בדוק ב-DB:
```bash
psql -h localhost -p 54329 -U paperclip -c \
"SELECT name, metadata->>'claude_md_mtime' AS mtime FROM agents WHERE metadata->>'claude_md_mtime' IS NOT NULL LIMIT 5" \
paperclip
```
- [ ] **Step 5: Commit**
```bash
cd /home/chaim/legal-ai
git add scripts/sync_agents_across_companies.py
git commit -m "feat: track claude_md_mtime in agents.metadata during sync"
```
---
## Task 4: הוסף `make check-agents` shortcut
**Files:**
- Modify: `legal-ai/Makefile` (אם קיים) אחרת הוסף alias
- [ ] **Step 1: בדוק אם Makefile קיים**
```bash
ls /home/chaim/legal-ai/Makefile 2>/dev/null && echo "EXISTS" || echo "MISSING"
```
- [ ] **Step 2א: אם Makefile קיים — הוסף target**
```makefile
check-agents:
python scripts/sync_agents_across_companies.py --check-instructions
sync-agents-dry:
python scripts/sync_agents_across_companies.py --dry-run
sync-agents:
python scripts/sync_agents_across_companies.py --apply
```
- [ ] **Step 2ב: אם Makefile לא קיים — הוסף alias ל-`~/.bashrc`**
```bash
echo "alias check-agents='cd /home/chaim/legal-ai && python scripts/sync_agents_across_companies.py --check-instructions'" >> ~/.bashrc
source ~/.bashrc
```
- [ ] **Step 3: בדוק**
```bash
# אם Makefile:
make -C /home/chaim/legal-ai check-agents
# אם alias:
check-agents
```
Expected: טבלת סוכנים מוצגת.
- [ ] **Step 4: Commit (אם Makefile)**
```bash
cd /home/chaim/legal-ai
git add Makefile
git commit -m "feat: add check-agents and sync-agents make targets"
```
---
## אימות סופי
| בדיקה | פקודה | תוצאה מצופה |
|-------|-------|-------------|
| `--check-instructions` | `python sync_agents... --check-instructions` | טבלה עם ✅ לכל agent |
| Pre-flight בולם apply | מחק קובץ זמנית + `--apply` | Abort עם הודעה ברורה |
| mtime ב-DB | `SELECT metadata->>'claude_md_mtime' FROM agents` | timestamp לכל agent |
| DRIFT זוהה | שנה קובץ + `--check-instructions` | ⚠️ DRIFT מוצג |
| shortcut | `check-agents` או `make check-agents` | עובד |

View File

@@ -0,0 +1,412 @@
# Scheduled Background Agents Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** הוסף 2 cron jobs לplugin: (1) תזכורת על תיקים תקועים 3+ ימים, (2) ניתוח פידבק יו"ר שבועי עם עדכון `decision-lessons.md` אוטומטי.
**Architecture:** שני jobs חדשים נרשמים ב-`ctx.jobs.register()`. Job 1 קורא `/api/cases/stale?days=3` (endpoint חדש ב-legal-ai) ומוסיף תגובה לissues תקועים. Job 2 קורא `/api/chair-feedback/weekly-summary` (endpoint חדש), שולח לCEO agent שמעדכן את `decision-lessons.md`.
**Tech Stack:** TypeScript (plugin-legal-ai jobs), Python/FastAPI (legal-ai web), asyncpg, Paperclip SDK `ctx.jobs.register()`, `ctx.agents.invoke()`.
---
## File Map
| Action | File |
|--------|------|
| Modify | `plugin-legal-ai/src/worker.ts` — register 2 new jobs |
| Modify | `plugin-legal-ai/plugin.json` — declare 2 new job entries |
| Modify | `plugin-legal-ai/src/manifest.ts` — add to jobs array |
| Modify | `legal-ai/web/app.py` — add `GET /api/cases/stale` + `GET /api/chair-feedback/weekly-summary` |
| Modify | `legal-ai/web/database.py` (or equivalent DB module) — add `get_stale_cases()` + `get_weekly_chair_feedback()` |
---
## Task 1: הוסף endpoint `GET /api/cases/stale`
**Files:**
- Modify: `legal-ai/web/app.py`
- Modify: `legal-ai/web/database.py` (שם DB queries מנוהלים — בדוק עם `grep -n "async def get_cases\|from.*database\|import.*db" /home/chaim/legal-ai/web/app.py | head -10`)
- [ ] **Step 1: מצא את module ה-DB**
```bash
grep -n "^from\|^import\|db\." /home/chaim/legal-ai/web/app.py | head -20
```
זהה את שם הmodule שמכיל את DB queries (בד"כ `database.py` או `db.py`).
- [ ] **Step 2: הוסף `get_stale_cases()` לmodule ה-DB**
```python
async def get_stale_cases(days: int = 3) -> list[dict]:
"""Return cases whose status is not 'final' and haven't been updated in `days` days."""
async with get_db_connection() as conn:
rows = await conn.fetch(
"""
SELECT case_number, title, status, company_id,
updated_at,
now() - updated_at AS age
FROM cases
WHERE status NOT IN ('final', 'new')
AND updated_at < now() - ($1 || ' days')::interval
ORDER BY updated_at ASC
""",
str(days),
)
return [dict(r) for r in rows]
```
> `get_db_connection()` — השתמש בדפוס הקיים בקובץ. אם זה `asyncpg.connect()` ישיר, `asyncpg.create_pool()`, או context manager — העתק את הדפוס.
- [ ] **Step 3: הוסף endpoint ב-`app.py`**
```python
@app.get("/api/cases/stale")
async def list_stale_cases(days: int = 3):
"""Cases stuck in non-final status for more than `days` days."""
cases = await db.get_stale_cases(days=days)
return {
"cases": [
{
"case_number": c["case_number"],
"title": c["title"],
"status": c["status"],
"company_id": str(c["company_id"]),
"days_stale": c["age"].days,
}
for c in cases
],
"total": len(cases),
}
```
- [ ] **Step 4: בדיקה**
```bash
curl -s "https://legal-ai.nautilus.marcusgroup.org/api/cases/stale?days=1" | jq .total
```
Expected: JSON עם רשימת תיקים.
- [ ] **Step 5: Commit**
```bash
cd /home/chaim/legal-ai
git add web/app.py web/database.py # או השם הנכון
git commit -m "feat: add GET /api/cases/stale endpoint"
```
---
## Task 2: הוסף endpoint `GET /api/chair-feedback/weekly-summary`
**Files:**
- Modify: `legal-ai/web/app.py`
- Modify: DB module
- [ ] **Step 1: בדוק את מבנה טבלת `chair_feedback`**
```bash
sqlite3 /home/chaim/.paperclip/instances/default/data/app.db \
".schema chair_feedback" 2>/dev/null || \
psql -h localhost -p 5433 -U legal_ai -c "\d chair_feedback" legal_ai 2>/dev/null || \
grep -rn "chair_feedback" /home/chaim/legal-ai/mcp-server/src/ | head -10
```
- [ ] **Step 2: הוסף `get_weekly_chair_feedback()` לDB module**
```python
async def get_weekly_chair_feedback(days: int = 7) -> list[dict]:
"""Return chair feedback entries from the last `days` days."""
async with get_db_connection() as conn:
rows = await conn.fetch(
"""
SELECT cf.case_number, cf.feedback_text, cf.created_at,
cf.feedback_type, c.title
FROM chair_feedback cf
JOIN cases c ON c.case_number = cf.case_number
WHERE cf.created_at > now() - ($1 || ' days')::interval
ORDER BY cf.created_at DESC
""",
str(days),
)
return [dict(r) for r in rows]
```
> אם שמות השדות שונים (בדוק ב-Step 1) — התאם.
- [ ] **Step 3: הוסף endpoint**
```python
@app.get("/api/chair-feedback/weekly-summary")
async def get_chair_feedback_weekly(days: int = 7):
"""Feedback entries from the past week, formatted for the learning agent."""
entries = await db.get_weekly_chair_feedback(days=days)
if not entries:
return {"summary": "", "entry_count": 0}
lines = [
f"- תיק {e['case_number']} ({e['title']}): {e['feedback_text']}"
for e in entries
]
summary = "\n".join(lines)
return {"summary": summary, "entry_count": len(entries), "entries": entries}
```
- [ ] **Step 4: בדיקה**
```bash
curl -s "https://legal-ai.nautilus.marcusgroup.org/api/chair-feedback/weekly-summary" | jq .entry_count
```
- [ ] **Step 5: Commit**
```bash
cd /home/chaim/legal-ai
git add web/app.py web/database.py
git commit -m "feat: add GET /api/chair-feedback/weekly-summary endpoint"
```
---
## Task 3: הוסף jobs לplugin
**Files:**
- Modify: `plugin-legal-ai/plugin.json`
- Modify: `plugin-legal-ai/src/manifest.ts`
- [ ] **Step 1: קרא את ה-`jobs` הקיים ב-`plugin.json`**
```bash
cat /home/chaim/plugin-legal-ai/plugin.json | python3 -c "import json,sys; d=json.load(sys.stdin); print(json.dumps(d['jobs'], indent=2))"
```
- [ ] **Step 2: הוסף 2 jobs חדשים לarray `"jobs"` ב-`plugin.json`**
```json
{
"jobKey": "stale-case-reminder",
"displayName": "תזכורת תיקים תקועים",
"description": "מזהה תיקים שלא עודכנו 3+ ימים ומוסיף תגובה לissue",
"schedule": "0 8 * * *"
},
{
"jobKey": "weekly-feedback-analysis",
"displayName": "ניתוח פידבק שבועי",
"description": "מסכם פידבק יו\"ר מהשבוע האחרון ומעדכן את decision-lessons.md",
"schedule": "0 19 * * 0"
}
```
> `"0 8 * * *"` = כל יום בשעה 08:00. `"0 19 * * 0"` = כל ראשון ב-19:00.
- [ ] **Step 3: עדכן `manifest.ts`**
```bash
grep -n "jobs\|jobKey\|schedule" /home/chaim/plugin-legal-ai/src/manifest.ts
```
הוסף את אותם 2 objects לarray `jobs` ב-`manifest.ts`.
- [ ] **Step 4: Commit**
```bash
cd /home/chaim/plugin-legal-ai
git add plugin.json src/manifest.ts
git commit -m "feat: declare stale-case-reminder and weekly-feedback-analysis jobs"
```
---
## Task 4: Implement job handlers ב-`worker.ts`
**Files:**
- Modify: `plugin-legal-ai/src/worker.ts`
- [ ] **Step 1: קרא את handler של `sync-case-status` הקיים**
```bash
grep -n "sync-case-status\|jobs.register\|jobKey" /home/chaim/plugin-legal-ai/src/worker.ts
```
העתק את הדפוס.
- [ ] **Step 2: הוסף את `stale-case-reminder` handler**
בתוך `setup(ctx)`, אחרי רישום ה-job הקיים:
```typescript
ctx.jobs.register("stale-case-reminder", async (job) => {
ctx.logger.info("stale-case-reminder: starting");
const config = await ctx.config.get();
const apiBase = (config.legalApiBaseUrl as string) ?? "http://localhost:8085";
const resp = await ctx.http.fetch(`${apiBase}/api/cases/stale?days=3`);
if (!resp.ok) {
ctx.logger.error(`stale-case-reminder: API error ${resp.status}`);
return;
}
const data = (await resp.json()) as {
cases: Array<{
case_number: string;
title: string;
status: string;
company_id: string;
days_stale: number;
}>;
};
for (const staleCase of data.cases) {
const issueId = await ctx.state.get(
{ companyId: staleCase.company_id },
`case:${staleCase.case_number}`
);
if (!issueId) continue;
await ctx.issues.createComment({
issueId: issueId as string,
body: `⚠️ **תיק תקוע** — ${staleCase.days_stale} ימים ללא עדכון (סטטוס: ${staleCase.status}). האם יש צורך בפעולה?`,
});
ctx.logger.info(
`stale-case-reminder: posted reminder for ${staleCase.case_number} (${staleCase.days_stale}d stale)`
);
}
ctx.logger.info(`stale-case-reminder: done. ${data.cases.length} cases reminded`);
});
```
- [ ] **Step 3: הוסף את `weekly-feedback-analysis` handler**
```typescript
ctx.jobs.register("weekly-feedback-analysis", async (job) => {
ctx.logger.info("weekly-feedback-analysis: starting");
const config = await ctx.config.get();
const apiBase = (config.legalApiBaseUrl as string) ?? "http://localhost:8085";
const resp = await ctx.http.fetch(`${apiBase}/api/chair-feedback/weekly-summary`);
if (!resp.ok) {
ctx.logger.error(`weekly-feedback-analysis: API error ${resp.status}`);
return;
}
const data = (await resp.json()) as {
summary: string;
entry_count: number;
};
if (data.entry_count === 0) {
ctx.logger.info("weekly-feedback-analysis: no feedback this week, skipping");
return;
}
// Invoke the CEO agent to process feedback and update decision-lessons.md
const companies = await ctx.companies.list();
for (const company of companies) {
// CEO IDs per company
const CEO_IDS: Record<string, string> = {
"42a7acd0-30c5-4cbd-ac97-7424f65df294": "752cebdd-6748-4a04-aacd-c7ab0294ef33",
"8639e837-4c9d-47fa-a76b-95788d651896": "cdbfa8bc-3d61-41a4-a2e7-677ec7d34562",
};
const ceoId = CEO_IDS[company.id];
if (!ceoId) continue;
await ctx.agents.invoke(ceoId, company.id, {
prompt: `ניתוח פידבק שבועי יו"ר (${data.entry_count} פריטים):
${data.summary}
המשימה: עדכן את /home/chaim/legal-ai/docs/legal-decision-lessons.md עם הלקחים החדשים שעולים מהפידבק. הוסף רק לקחים חדשים שלא קיימים כבר. קבץ לפי נושא.`,
reason: "weekly-feedback-analysis scheduled job",
});
ctx.logger.info(
`weekly-feedback-analysis: invoked CEO for company ${company.id} (${data.entry_count} feedback entries)`
);
break; // One CEO is enough — lessons file is shared
}
});
```
- [ ] **Step 4: TypeScript check**
```bash
cd /home/chaim/plugin-legal-ai && npx tsc --noEmit
```
Expected: 0 errors.
- [ ] **Step 5: Build**
```bash
cd /home/chaim/plugin-legal-ai && npm run build
```
- [ ] **Step 6: Commit**
```bash
cd /home/chaim/plugin-legal-ai
git add src/worker.ts
git commit -m "feat: implement stale-case-reminder and weekly-feedback-analysis jobs"
```
---
## Task 5: Re-install plugin + בדיקה
- [ ] **Step 1: Deploy legal-ai**
```bash
cd /home/chaim/legal-ai && git push origin main
# המתן ~3 דקות
curl -s https://legal-ai.nautilus.marcusgroup.org/api/health | jq .status
```
- [ ] **Step 2: Re-install plugin**
```bash
cd /home/chaim/plugin-legal-ai && npm run build
npx paperclipai plugin uninstall marcusgroup.legal-ai \
--api-base http://localhost:3100 --api-key pcapi_legal_install_key_2026
npx paperclipai plugin install /home/chaim/plugin-legal-ai \
--api-base http://localhost:3100 --api-key pcapi_legal_install_key_2026
pm2 restart paperclip
```
- [ ] **Step 3: בדוק שה-jobs רשומים**
```bash
curl -s -H "Authorization: Bearer pcapi_legal_install_key_2026" \
http://localhost:3100/api/plugins/marcusgroup.legal-ai/jobs | jq .[].jobKey
```
Expected: `"sync-case-status"`, `"stale-case-reminder"`, `"weekly-feedback-analysis"`.
- [ ] **Step 4: הפעל job ידנית לבדיקה**
```bash
curl -s -X POST -H "Authorization: Bearer pcapi_legal_install_key_2026" \
http://localhost:3100/api/plugins/marcusgroup.legal-ai/jobs/stale-case-reminder/run | jq .
```
Expected: Job הופעל. בדוק logs:
```bash
pm2 logs paperclip --lines 30 | grep "stale-case-reminder"
```
---
## אימות סופי
| בדיקה | פקודה | תוצאה מצופה |
|-------|-------|-------------|
| API stale endpoint | `curl .../api/cases/stale?days=1` | JSON עם cases |
| API feedback endpoint | `curl .../api/chair-feedback/weekly-summary` | JSON עם summary |
| Jobs רשומים | `GET .../api/plugins/.../jobs` | 3 jobs רשומים |
| Stale reminder ידני | `POST .../jobs/stale-case-reminder/run` | תגובות בissues |
| Feedback analysis ידני | `POST .../jobs/weekly-feedback-analysis/run` | CEO מועיר |

View File

@@ -275,8 +275,8 @@ HALACHA_CANONICAL_SYNTH_MODEL = os.environ.get("HALACHA_CANONICAL_SYNTH_MODEL",
HALACHA_CANONICAL_SYNTH_EFFORT = os.environ.get("HALACHA_CANONICAL_SYNTH_EFFORT", "high") HALACHA_CANONICAL_SYNTH_EFFORT = os.environ.get("HALACHA_CANONICAL_SYNTH_EFFORT", "high")
HALACHA_CANONICAL_SYNTH_DRIFT_FLOOR = float(os.environ.get("HALACHA_CANONICAL_SYNTH_DRIFT_FLOOR", "0.80")) HALACHA_CANONICAL_SYNTH_DRIFT_FLOOR = float(os.environ.get("HALACHA_CANONICAL_SYNTH_DRIFT_FLOOR", "0.80"))
# Google Cloud Vision (OCR for scanned PDFs) # Mistral OCR (fallback for scanned PDFs — replaces Google Cloud Vision)
GOOGLE_CLOUD_VISION_API_KEY = os.environ.get("GOOGLE_CLOUD_VISION_API_KEY", "") MISTRAL_API_KEY = os.environ.get("MISTRAL_API_KEY", "")
# Data directory # Data directory
DATA_DIR = Path(os.environ.get("DATA_DIR", str(Path.home() / "legal-ai" / "data"))) DATA_DIR = Path(os.environ.get("DATA_DIR", str(Path.home() / "legal-ai" / "data")))
@@ -361,7 +361,7 @@ PARENT_DOC_CHILD_OVERLAP_TOKENS = int(
# External service allowlist — case materials may ONLY be sent to these domains # External service allowlist — case materials may ONLY be sent to these domains
ALLOWED_EXTERNAL_SERVICES = { ALLOWED_EXTERNAL_SERVICES = {
"api.voyageai.com", # Voyage AI (embeddings) "api.voyageai.com", # Voyage AI (embeddings)
"vision.googleapis.com", # Google Cloud Vision (OCR) "api.mistral.ai", # Mistral OCR (scanned PDFs)
} }
# Audit # Audit

View File

@@ -418,6 +418,12 @@ async def digest_process_pending(limit: int = 20) -> str:
return await digest_tools.digest_process_pending(_clamp_limit(limit)) return await digest_tools.digest_process_pending(_clamp_limit(limit))
@mcp.tool()
async def digest_radar(case_number: str, limit: int = 5, min_score: float = 0.45) -> str:
"""רדאר-יומונים הקשרי-לתיק (X12) — מחזיר יומונים לא-מקושרים (פס"ד שעוד אין לנו) שהנושא שלהם קרוב סמנטית לתיק, כדי שלא ייפול פס"ד רלוונטי בזמן ההכרעה. כל ליד מצביע על מראה-המקום של הפס"ד המקורי + סטטוס-הפער + פעולה מוצעת. ⚠️ radar בלבד — היומון אינו מצוטט (INV-DIG1)."""
return await digest_tools.digest_radar(case_number, _clamp_limit(limit), min_score)
@mcp.tool() @mcp.tool()
async def halacha_review( async def halacha_review(
halacha_id: str, halacha_id: str,
@@ -1170,6 +1176,15 @@ async def list_chair_feedback(
return await workflow.list_chair_feedback(case_number, category, unresolved_only, _clamp_limit(limit)) return await workflow.list_chair_feedback(case_number, category, unresolved_only, _clamp_limit(limit))
@mcp.tool()
async def record_curator_findings(case_number: str, findings: list[dict]) -> str:
"""לכידת ממצאי-האוצֵר על החלטה סופית כ-decision_lessons מובְנים (source='curator',
ממתינים לשער-יו"ר) — INV-LRN3. כל ממצא: {"text": "...", "category": style/structure/
lexicon/tabular} (או "tag" עברי). מחזיר כמה נרשמו וכמה כפילויות דולגו. read-only על
התוכן — הרישום הצעה הממתינה לאישור דפנה ב-/training, לא שינוי-קול ישיר (G10)."""
return await workflow.record_curator_findings(case_number, findings)
@mcp.tool() @mcp.tool()
async def halacha_corroboration(halacha_id: str) -> dict: async def halacha_corroboration(halacha_id: str) -> dict:
"""החזר את ה-corroboration של הלכה: הציטוטים שמתקפים אותה, הטיפול, וסיכום (X11, read-only).""" """החזר את ה-corroboration של הלכה: הציטוטים שמתקפים אותה, הטיפול, וסיכום (X11, read-only)."""

View File

@@ -230,6 +230,15 @@ async def aggregate_claims_to_arguments(
party = "unknown" party = "unknown"
by_party.setdefault(party, []).append(dict(r)) by_party.setdefault(party, []).append(dict(r))
# Valid claim_ids for this case == the ids of the claims we just fetched.
# The LLM is asked to echo back supporting claim_ids, but it may hallucinate
# a syntactically-valid-but-nonexistent UUID (malformed ones are already
# dropped in ``_normalize_argument``). Validating against this known set at
# source keeps a doomed INSERT — which would poison the surrounding asyncpg
# transaction (FK violation -> "current transaction is aborted") — out of
# the transaction entirely (G1: fix at source, not symptom).
valid_claim_ids: set[UUID] = {r["id"] for r in rows}
party_counts: dict[str, int] = {} party_counts: dict[str, int] = {}
inserted = 0 inserted = 0
errors: list[str] = [] errors: list[str] = []
@@ -275,17 +284,30 @@ async def aggregate_claims_to_arguments(
arg["priority"], arg["priority"],
) )
for cid in arg["claim_ids"]: for cid in arg["claim_ids"]:
try: if cid not in valid_claim_ids:
await conn.execute( # Hallucinated claim_id that doesn't belong to this
"""INSERT INTO legal_argument_propositions # case. Skip it rather than letting the FK violation
(argument_id, claim_id) # abort the whole transaction.
VALUES ($1, $2) logger.warning(
ON CONFLICT DO NOTHING""", "argument_aggregator: skipped unknown claim_id %s for arg %s",
arg_id, cid, cid, arg_id,
) )
continue
try:
# Per-row savepoint: even after the validation above,
# wrap the INSERT so any unexpected constraint failure
# rolls back to the savepoint instead of poisoning the
# surrounding transaction (asyncpg nests transaction()
# as SAVEPOINT when already inside one).
async with conn.transaction():
await conn.execute(
"""INSERT INTO legal_argument_propositions
(argument_id, claim_id)
VALUES ($1, $2)
ON CONFLICT DO NOTHING""",
arg_id, cid,
)
except Exception as e: # noqa: BLE001 except Exception as e: # noqa: BLE001
# Likely FK violation if the LLM hallucinated
# a claim_id. Log and continue.
logger.warning( logger.warning(
"argument_aggregator: skipped bad claim_id %s for arg %s: %s", "argument_aggregator: skipped bad claim_id %s for arg %s: %s",
cid, arg_id, e, cid, arg_id, e,

View File

@@ -983,6 +983,22 @@ async def _build_style_context(practice_area: str = "") -> str:
("anti_patterns", "אנטי-דפוסים (להימנע)"), ("anti_patterns", "אנטי-דפוסים (להימנע)"),
): ):
ov = await db.get_methodology_overrides(cat) ov = await db.get_methodology_overrides(cat)
if cat == "anti_patterns":
# Anti-patterns are STRUCTURAL INVARIANTS of Dafna's style (no
# markdown headers, no bullet lists, no mid-paragraph mini-lists —
# she writes continuous legal narrative). They must reach the writer
# ALWAYS, from the SAME canonical list style_distance measures against
# (lessons.ANTI_PATTERNS) — otherwise the loop detects them but never
# corrects them, and drafts keep emitting them (the gap that left
# 8137 with 28 hits). Chair additions layer on top; they never
# remove the canonical ones.
from legal_mcp.services.lessons import ANTI_PATTERNS as _ANTI
learned.append(f"\n**{label} — כתוב נרטיב משפטי רציף; הימנע מ:**")
for ap in _ANTI:
learned.append(f"- {ap['note']}")
for k, v in (ov or {}).items():
learned.append(f"- (יו\"ר) {k}: {json.dumps(v, ensure_ascii=False)}")
continue
if ov: if ov:
learned.append(f"\n**{label} — ערכי היו\"ר (גוברים על ברירת-המחדל):**") learned.append(f"\n**{label} — ערכי היו\"ר (גוברים על ברירת-המחדל):**")
for k, v in ov.items(): for k, v in ov.items():

View File

@@ -0,0 +1,134 @@
"""Citation-verification view (X11 Phase 2 / #154) — the chair's "אימות פסיקה" tab.
Assembles, per legal ARGUMENT of a case, the supporting precedents the chair should
verify before the writer cites them:
• in-corpus suggestions — per-issue semantic retrieval over the authoritative
precedent library (``search_library``), each carrying the cumulative authority
signal (``cited_by``: followed/distinguished — db.citation_authority, X11).
• attached/verified state — any ``case_precedents`` row already attached to the
argument (verified flag + chair_note), merged onto the matching suggestion.
• radar — UNLINKED digests relevant to the same issue (rulings we don't hold yet),
from ``case_digest_radar`` grouped by matched issue.
Pure read/assembly — never writes, never cites (INV-DIG1/INV-AH). The chair verifies
through ``db.set_case_precedent_verified`` / attach; the writer consumes only verified
rows. Reuses the one corpus search + the one authority query + the one radar (G2).
"""
from __future__ import annotations
import logging
from uuid import UUID
from legal_mcp.services import (
argument_aggregator,
db,
digest_library,
precedent_library,
)
logger = logging.getLogger(__name__)
_SUGGEST_PER_ISSUE = 4
_SUGGEST_FLOOR = 0.45
async def build_view(case_number: str) -> dict:
case = await db.get_case_by_number(case_number)
if not case:
return {"status": "case_not_found", "case_number": case_number, "arguments": []}
case_id = case["id"]
if isinstance(case_id, str):
case_id = UUID(case_id)
ctx = " ".join(x for x in [case.get("title") or "", case.get("appeal_subtype") or ""] if x).strip()
args = await argument_aggregator.get_legal_arguments(case_id)
# Attached precedents already on the case → grouped by argument_id, keyed by
# the resolved corpus ruling so we can merge verify-state onto a suggestion.
attached = await db.list_case_precedents(case_id)
attached_by_arg: dict[str, dict[str, dict]] = {}
for p in attached:
aid = str(p.get("argument_id") or "")
clid = str(p.get("case_law_id") or "")
if aid and clid:
attached_by_arg.setdefault(aid, {})[clid] = p
# Radar (unlinked digests) once, grouped by the issue label it matched.
radar_by_issue: dict[str, list[dict]] = {}
try:
radar = await digest_library.case_digest_radar(case_number, limit=12, min_score=0.42)
for lead in radar.get("leads", []):
for label in (lead.get("matched_issues") or [""]):
radar_by_issue.setdefault(label, []).append(lead)
except Exception as e: # noqa: BLE001 — radar is best-effort
logger.warning("citation_verification radar failed for %s: %s", case_number, e)
out_args: list[dict] = []
n_verified = 0
for a in args:
aid = str(a["id"])
title = (a.get("argument_title") or "").strip()
topic = (a.get("legal_topic") or "").strip()
query = f"{ctx} {title}. {topic}".strip()
hits = []
try:
hits = await precedent_library.search_library(
query=query, limit=_SUGGEST_PER_ISSUE, include_halachot=True)
except Exception as e: # noqa: BLE001
logger.warning("citation_verification search failed (%s): %s", title[:30], e)
# Resolve the authority breakdown for the hit set in one batched query.
clids = [UUID(str(h["case_law_id"])) for h in hits
if h.get("case_law_id") and float(h.get("score", 0) or 0) >= _SUGGEST_FLOOR]
authority = await db.citation_authority(clids) if clids else {}
seen: set[str] = set()
supporting: list[dict] = []
for h in hits:
clid = str(h.get("case_law_id") or "")
if not clid or clid in seen:
continue
if float(h.get("score", 0) or 0) < _SUGGEST_FLOOR:
continue
seen.add(clid)
att = attached_by_arg.get(aid, {}).get(clid)
if att and att.get("verified"):
n_verified += 1
supporting.append({
"case_law_id": clid,
"case_number": h.get("case_number") or "",
"case_name": h.get("case_name") or "",
"quote": h.get("supporting_quote") or h.get("rule_statement") or "",
"score": round(float(h.get("score", 0) or 0), 3),
"cited_by": authority.get(clid, {"total": 0, "positive": 0,
"negative": 0, "unclassified": 0,
"by_treatment": {}}),
"attached_id": str(att["id"]) if att else None,
"verified": bool(att.get("verified")) if att else False,
"chair_note": (att.get("chair_note") or "") if att else "",
})
out_args.append({
"argument_id": aid,
"title": title,
"legal_topic": topic,
"priority": a.get("priority") or "",
"party": a.get("party") or "",
"supporting": supporting,
"radar": radar_by_issue.get(title, []),
})
return {
"status": "ok",
"case_number": case_number,
"arguments": out_args,
"summary": {
"arguments_total": len(out_args),
"arguments_with_support": sum(1 for x in out_args if x["supporting"]),
"verified": n_verified,
"radar_leads": sum(len(x["radar"]) for x in out_args),
},
}

View File

@@ -144,9 +144,12 @@ def _split_into_sections(text: str) -> list[tuple[str, str]]:
markers: list[tuple[int, str]] = [] markers: list[tuple[int, str]] = []
for pattern, section_type in SECTION_PATTERNS: for pattern, section_type in SECTION_PATTERNS:
# ^ + MULTILINE: line start only. Optional leading spaces/tabs and an # ^ + MULTILINE: line start only. Optional leading spaces/tabs, an
# optional Markdown ATX header prefix (``## ``/``### ``), and an
# optional ordinal prefix ("5.", "5)", "ג.") before the keyword. # optional ordinal prefix ("5.", "5)", "ג.") before the keyword.
anchored = rf"^[ \t]*(?:\d+[.)]\s*|[א-ת][.)]\s*)?(?:{pattern})" # The Markdown prefix handles Mistral OCR output where section
# titles are rendered as ``## נימוקי הערר`` etc.
anchored = rf"^[ \t]*(?:#{1,3}\s+)?(?:\d+[.)]\s*|[א-ת][.)]\s*)?(?:{pattern})"
for match in re.finditer(anchored, text, re.MULTILINE): for match in re.finditer(anchored, text, re.MULTILINE):
markers.append((match.start(), section_type)) markers.append((match.start(), section_type))

View File

@@ -236,7 +236,18 @@ async def extract_and_store(case_law_id: UUID) -> dict:
if not row: if not row:
return {"extracted": 0, "linked": 0, "new": 0, "skipped": 0, "error": "not_found"} return {"extracted": 0, "linked": 0, "new": 0, "skipped": 0, "error": "not_found"}
text = row["full_text"] or "" # A citation counts as the DECIDING body's reliance ONLY when it appears in
# the discussion/ruling sections — never where a party cites a ruling in its
# own arguments (block ז). Reuse the halacha extractor's section selection
# (G2: a single definition of "reasoning sections" — legal_analysis/ruling/
# conclusion + the discussion-anchor, excluding facts/claims/intro) so a
# ruling cited only in appellant_claims/parties_claims is never attributed to
# the chair. Fall back to full_text only for un-chunked rows.
from legal_mcp.services import halacha_extractor as _hx
disc_chunks, _used_fallback = await _hx._select_extractable_chunks(case_law_id)
text = "\n".join((c.get("content") or "") for c in disc_chunks).strip()
if not text:
text = row["full_text"] or ""
own_norm = _normalize_case_number(row["case_number"] or "") own_norm = _normalize_case_number(row["case_number"] or "")
extracted = 0 extracted = 0
@@ -388,12 +399,16 @@ async def list_citations_to_case_law(case_law_id: UUID) -> list[dict]:
pic.cited_case_number, pic.cited_case_number,
pic.match_context, pic.match_context,
pic.match_pattern, pic.match_pattern,
pic.treatment,
pic.confidence::float AS confidence, pic.confidence::float AS confidence,
pic.created_at, pic.created_at,
cl.case_number AS source_case_number, cl.case_number AS source_case_number,
cl.case_name AS source_case_name, cl.case_name AS source_case_name,
cl.chair_name AS source_chair_name, cl.chair_name AS source_chair_name,
cl.district AS source_district cl.district AS source_district,
cl.precedent_level AS source_precedent_level,
cl.court AS source_court,
cl.date AS source_date
FROM precedent_internal_citations pic FROM precedent_internal_citations pic
JOIN case_law cl ON cl.id = pic.source_case_law_id JOIN case_law cl ON cl.id = pic.source_case_law_id
WHERE pic.cited_case_law_id = $1 WHERE pic.cited_case_law_id = $1
@@ -432,3 +447,42 @@ async def get_cited_case_law_ids(source_case_law_ids: list[UUID]) -> dict[str, l
for r in rows: for r in rows:
out.setdefault(r["source_id"], []).append(r["cited_id"]) out.setdefault(r["source_id"], []).append(r["cited_id"])
return out return out
async def relink_orphan_citations(case_law_id: "UUID | None" = None) -> int:
"""Back-fill ``cited_case_law_id`` on orphan citation edges that now resolve.
Citations are resolved to a corpus row only at extraction time (see
``extract_and_store``). When a cited ruling is uploaded LATER, its existing
orphan edges (``cited_case_law_id IS NULL``) are never re-resolved, so the
citation graph under-counts real connectivity. Call on every precedent
upload (``case_law_id`` set → link only edges that resolve to that row) or as
a one-off sweep (``case_law_id=None`` → re-resolve every orphan). Reuses the
canonical resolver ``_resolve_case_law_id`` (no parallel matching logic, G2);
idempotent — a second run links nothing new.
"""
pool = await db.get_pool()
async with pool.acquire() as conn:
orphans = await conn.fetch(
"SELECT DISTINCT cited_case_number FROM precedent_internal_citations "
"WHERE cited_case_law_id IS NULL AND coalesce(cited_case_number, '') <> ''"
)
linked = 0
for r in orphans:
cited = await _resolve_case_law_id(_normalize_case_number(r["cited_case_number"]))
if cited is None:
continue
if case_law_id is not None and cited != case_law_id:
continue
async with pool.acquire() as conn:
res = await conn.execute(
"UPDATE precedent_internal_citations "
"SET cited_case_law_id = $1, confidence = GREATEST(confidence, 0.90) "
"WHERE cited_case_law_id IS NULL AND cited_case_number = $2",
cited, r["cited_case_number"],
)
try:
linked += int(res.split()[-1])
except (ValueError, IndexError):
pass
return linked

View File

@@ -1697,6 +1697,80 @@ CREATE INDEX IF NOT EXISTS idx_hcc_canonical
""" """
SCHEMA_V42_SQL = """
-- INV-DM7: authority (binding/persuasive) is STRUCTURAL for committee sources.
-- An appeals-committee decision (source_kind='internal_committee', or any row
-- with source_type='appeals_committee') is persuasive, NEVER binding — a
-- committee's ruling does not bind another committee, nor itself. is_binding
-- stays a stored column (it selects the halacha-extraction prompt), but for
-- committee sources it is a faithful cache of the derived value, not a chair
-- guess. Mirrors 02-data-model.md §INV-DM7 / X8-field-provenance.md.
-- Step 1 — normalize legacy rows at the source (G1), idempotent.
UPDATE case_law SET is_binding = FALSE
WHERE is_binding IS TRUE
AND (source_kind = 'internal_committee' OR source_type = 'appeals_committee');
-- Step 2 — constrain so a binding committee row can never be written again.
DO $$ BEGIN
ALTER TABLE case_law ADD CONSTRAINT case_law_committee_not_binding_check
CHECK (NOT ((source_kind = 'internal_committee'
OR source_type = 'appeals_committee') AND is_binding));
EXCEPTION WHEN duplicate_object THEN NULL; END $$;
"""
SCHEMA_V43_SQL = """
-- חיפה (Haifa) is a distinct planning district, separate from הצפון (North).
-- A bug in _district_from_court mapped "חיפה" -> "צפון", and the service-layer
-- _VALID_DISTRICTS omitted "חיפה" entirely, so Haifa committee decisions were
-- mis-filed under צפון. Both fixed in internal_decisions.py; this reclassifies
-- the legacy rows at the source (G1) by their court text. Idempotent.
UPDATE case_law SET district = 'חיפה'
WHERE source_kind = 'internal_committee'
AND district = 'צפון'
AND court ILIKE '%חיפה%';
"""
SCHEMA_V44_SQL = """
-- Citation-verification panel (#154): the chair verifies each corpus precedent
-- against the specific legal ARGUMENT before the writer may cite it (INV-AH gate).
-- case_precedents gains: the argument it supports, the resolved case_law row
-- (for the cited_by authority signal + dedup), a verified flag (the gate), and a
-- verification timestamp. chair_note already exists. All nullable so legacy
-- section-scoped rows stay valid.
ALTER TABLE case_precedents ADD COLUMN IF NOT EXISTS argument_id UUID
REFERENCES legal_arguments(id) ON DELETE SET NULL;
ALTER TABLE case_precedents ADD COLUMN IF NOT EXISTS case_law_id UUID
REFERENCES case_law(id) ON DELETE SET NULL;
ALTER TABLE case_precedents ADD COLUMN IF NOT EXISTS verified BOOLEAN DEFAULT false;
ALTER TABLE case_precedents ADD COLUMN IF NOT EXISTS verified_at TIMESTAMPTZ;
CREATE INDEX IF NOT EXISTS idx_case_precedents_argument ON case_precedents(argument_id);
"""
# V45 (Path A — prospective held-out): a style-distance snapshot captured at
# final-upload time, BEFORE this case's lessons are folded. Because the draft was
# written with only the PRIOR lesson pool, each row is a clean generalization
# datapoint: "with N accumulated lessons, our draft on this unseen case scored X".
# pool_* record how much voice-learning had accumulated when the draft was made,
# so the trend (style-distance vs pool size / time) shows whether learning
# generalizes. Append-only; one row per final upload. (07-learning §0, INV-LRN4.)
SCHEMA_V45_SQL = """
CREATE TABLE IF NOT EXISTS style_distance_history (
id UUID PRIMARY KEY DEFAULT uuid_generate_v4(),
case_number TEXT NOT NULL,
pair_id UUID REFERENCES draft_final_pairs(id) ON DELETE SET NULL,
measured_at TIMESTAMPTZ DEFAULT now(),
pool_discussion_rules INT DEFAULT 0,
pool_transition_phrases INT DEFAULT 0,
anti_pattern_total INT,
ratio_max_deviation REAL,
change_percent REAL
);
CREATE INDEX IF NOT EXISTS idx_style_distance_history_measured ON style_distance_history(measured_at);
"""
# Stable, arbitrary key for the session-level advisory lock that serialises # Stable, arbitrary key for the session-level advisory lock that serialises
# schema DDL across processes. Every short-lived process (cron drains, services) # schema DDL across processes. Every short-lived process (cron drains, services)
# re-runs the idempotent migrations on startup; without this lock two processes # re-runs the idempotent migrations on startup; without this lock two processes
@@ -1714,7 +1788,7 @@ async def _run_schema_migrations(pool: asyncpg.Pool) -> None:
await _apply_schema_ddl(conn) await _apply_schema_ddl(conn)
finally: finally:
await conn.execute("SELECT pg_advisory_unlock($1)", _MIGRATION_LOCK_KEY) await conn.execute("SELECT pg_advisory_unlock($1)", _MIGRATION_LOCK_KEY)
logger.info("Database schema initialized (v1-v41)") logger.info("Database schema initialized (v1-v43)")
async def _apply_schema_ddl(conn: asyncpg.Connection) -> None: async def _apply_schema_ddl(conn: asyncpg.Connection) -> None:
@@ -1760,6 +1834,10 @@ async def _apply_schema_ddl(conn: asyncpg.Connection) -> None:
await conn.execute(SCHEMA_V39_SQL) await conn.execute(SCHEMA_V39_SQL)
await conn.execute(SCHEMA_V40_SQL) await conn.execute(SCHEMA_V40_SQL)
await conn.execute(SCHEMA_V41_SQL) await conn.execute(SCHEMA_V41_SQL)
await conn.execute(SCHEMA_V42_SQL)
await conn.execute(SCHEMA_V43_SQL)
await conn.execute(SCHEMA_V44_SQL)
await conn.execute(SCHEMA_V45_SQL)
async def init_schema() -> None: async def init_schema() -> None:
@@ -2597,6 +2675,22 @@ async def add_decision_lesson(
return dict(row) if row else {} return dict(row) if row else {}
async def get_style_corpus_id_by_decision(decision_number: str) -> UUID | None:
"""Resolve the style_corpus row id for a final decision by its number.
The curator knows the case_number; lessons attach to the corpus row the
learning pipeline enrolled (enroll_style_corpus). Returns None if the final
has not been enrolled yet (caller surfaces it — no silent attach).
"""
pool = await get_pool()
async with pool.acquire() as conn:
row = await conn.fetchrow(
"SELECT id FROM style_corpus WHERE decision_number = $1 LIMIT 1",
decision_number,
)
return row["id"] if row else None
async def update_decision_lesson( async def update_decision_lesson(
lesson_id: UUID, lesson_id: UUID,
*, *,
@@ -2781,6 +2875,48 @@ async def get_style_patterns(pattern_type: str | None = None) -> list[dict]:
return [dict(r) for r in rows] return [dict(r) for r in rows]
async def append_global_rule(
category: str, key: str, items: list[str], seed_if_missing: list | None = None,
) -> int:
"""Append items to a _global appeal_type_rules list value (the writer-consumed
methodology channel), idempotently and under a row lock. Returns how many NEW
items were added (existing duplicates are skipped). Single locked read-modify-
write (MET-2/3) so concurrent appends can't drop items. Shared by the /training
promote gate (web `_append_methodology_override`) and chair-feedback auto-flow
(G2 — one append implementation)."""
pool = await get_pool()
async with pool.acquire() as conn:
async with conn.transaction():
row = await conn.fetchrow(
"SELECT rule_value FROM appeal_type_rules "
"WHERE appeal_type = '_global' AND rule_category = $1 AND rule_key = $2 "
"FOR UPDATE",
category, key,
)
if row:
current = row["rule_value"]
if isinstance(current, str):
try:
current = json.loads(current)
except (json.JSONDecodeError, TypeError):
current = []
else:
current = list(seed_if_missing or [])
if not isinstance(current, list):
current = []
added = [s for s in items if s and s not in current]
if not added and row:
return 0
merged = current + added
await conn.execute(
"INSERT INTO appeal_type_rules (id, appeal_type, rule_category, rule_key, rule_value) "
"VALUES (gen_random_uuid(), '_global', $1, $2, $3::text::jsonb) "
"ON CONFLICT (appeal_type, rule_category, rule_key) DO UPDATE SET rule_value = $3::text::jsonb",
category, key, json.dumps(merged, ensure_ascii=False),
)
return len(added)
async def get_methodology_overrides(category: str) -> dict: async def get_methodology_overrides(category: str) -> dict:
"""Chair's /methodology edits for one category (golden_ratios / discussion_rules / """Chair's /methodology edits for one category (golden_ratios / discussion_rules /
content_checklists). Returns {rule_key: parsed_value}. These OVERRIDE the hardcoded content_checklists). Returns {rule_key: parsed_value}. These OVERRIDE the hardcoded
@@ -2805,6 +2941,66 @@ async def get_methodology_overrides(category: str) -> dict:
return out return out
async def voice_lesson_pool_sizes() -> dict:
"""How many folded voice-learning items are in the writer-consumed pool right now
(universal discussion_rules + transition_phrases). Used by the prospective held-out
snapshot to record how much learning had accumulated when a draft was produced."""
pool = await get_pool()
async with pool.acquire() as conn:
rows = await conn.fetch(
"SELECT rule_category, COALESCE(jsonb_array_length(rule_value), 0) AS n "
"FROM appeal_type_rules "
"WHERE appeal_type = '_global' AND rule_key = 'universal' "
"AND rule_category IN ('discussion_rules', 'transition_phrases')",
)
out = {"discussion_rules": 0, "transition_phrases": 0}
for r in rows:
out[r["rule_category"]] = r["n"]
return out
async def record_style_distance_snapshot(
case_number: str, pair_id: str | None, pool_rules: int, pool_phrases: int,
anti_pattern_total: int | None, ratio_max_deviation: float | None,
change_percent: float | None,
) -> None:
"""Append one prospective held-out datapoint (Path A). Append-only; never updates."""
pool = await get_pool()
async with pool.acquire() as conn:
await conn.execute(
"INSERT INTO style_distance_history (case_number, pair_id, "
"pool_discussion_rules, pool_transition_phrases, anti_pattern_total, "
"ratio_max_deviation, change_percent) VALUES ($1, $2, $3, $4, $5, $6, $7)",
case_number, UUID(pair_id) if pair_id else None,
pool_rules, pool_phrases, anti_pattern_total,
ratio_max_deviation, change_percent,
)
async def get_style_distance_history() -> list[dict]:
"""The prospective held-out trend, oldest-first (Path A). Each row = one final
upload's style-distance measured before that case's lessons were folded."""
pool = await get_pool()
async with pool.acquire() as conn:
rows = await conn.fetch(
"SELECT case_number, measured_at, pool_discussion_rules, "
"pool_transition_phrases, anti_pattern_total, ratio_max_deviation, "
"change_percent FROM style_distance_history ORDER BY measured_at",
)
return [
{
"case_number": r["case_number"],
"measured_at": r["measured_at"].isoformat() if r["measured_at"] else None,
"pool_discussion_rules": r["pool_discussion_rules"],
"pool_transition_phrases": r["pool_transition_phrases"],
"anti_pattern_total": r["anti_pattern_total"],
"ratio_max_deviation": r["ratio_max_deviation"],
"change_percent": r["change_percent"],
}
for r in rows
]
async def get_recent_decision_lessons(limit: int = 15, practice_area: str = "") -> list[dict]: async def get_recent_decision_lessons(limit: int = 15, practice_area: str = "") -> list[dict]:
"""Per-decision learnings the chair/curator attached in /training (decision_lessons), """Per-decision learnings the chair/curator attached in /training (decision_lessons),
so the writer consumes them too (T15). Prefers style/structure/lexicon, recent first. so the writer consumes them too (T15). Prefers style/structure/lexicon, recent first.
@@ -2812,6 +3008,12 @@ async def get_recent_decision_lessons(limit: int = 15, practice_area: str = "")
Gate (INV-LRN1/G10): returns only CHAIR-APPROVED lessons (review_status='approved'). Gate (INV-LRN1/G10): returns only CHAIR-APPROVED lessons (review_status='approved').
Panel-written proposals stay out of the writer's context until the chair approves Panel-written proposals stay out of the writer's context until the chair approves
them in /training — the chair's review is the gate, not the panel's 2/2 vote. them in /training — the chair's review is the gate, not the panel's 2/2 vote.
No silent cap (חוקה §6): when more approved lessons match than `limit`, the
overflow is dropped by ORDER BY created_at DESC — log it so coverage loss is
observable, not invisible. The real fix is lesson synthesis (TaskMaster #158):
consolidate overlapping lessons into fewer, richer ones so the set fits without
truncation. Until then this WARN is the audit trail.
""" """
pool = await get_pool() pool = await get_pool()
async with pool.acquire() as conn: async with pool.acquire() as conn:
@@ -2826,6 +3028,20 @@ async def get_recent_decision_lessons(limit: int = 15, practice_area: str = "")
LIMIT $1""", LIMIT $1""",
limit, practice_area, limit, practice_area,
) )
if len(rows) >= limit:
total = await conn.fetchval(
"""SELECT count(*) FROM decision_lessons dl
JOIN style_corpus sc ON sc.id = dl.style_corpus_id
WHERE dl.review_status = 'approved'
AND ($1 = '' OR sc.practice_area = $1)""",
practice_area,
) or 0
if total > limit:
logger.warning(
"get_recent_decision_lessons: capped %d%d approved lessons "
"(practice_area=%r) — %d not reaching the writer. See TaskMaster #158 (synthesis).",
total, limit, practice_area or "*", total - limit,
)
return [dict(r) for r in rows] return [dict(r) for r in rows]
@@ -3295,28 +3511,58 @@ async def create_case_precedent(
chair_note: str = "", chair_note: str = "",
pdf_document_id: UUID | None = None, pdf_document_id: UUID | None = None,
practice_area: str | None = None, practice_area: str | None = None,
argument_id: UUID | None = None,
case_law_id: UUID | None = None,
verified: bool = False,
) -> dict: ) -> dict:
"""Insert a new precedent attached to a case.""" """Insert a new precedent attached to a case.
``argument_id``/``case_law_id``/``verified`` (X11 #154) link the attachment to
the specific legal argument it supports and the corpus ruling, and mark whether
the chair verified it (the INV-AH gate the writer respects)."""
pool = await get_pool() pool = await get_pool()
row = await pool.fetchrow( row = await pool.fetchrow(
""" """
INSERT INTO case_precedents INSERT INTO case_precedents
(case_id, section_id, quote, citation, chair_note, pdf_document_id, practice_area) (case_id, section_id, quote, citation, chair_note, pdf_document_id,
VALUES ($1, $2, $3, $4, $5, $6, $7) practice_area, argument_id, case_law_id, verified, verified_at)
VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10,
CASE WHEN $10 THEN now() ELSE NULL END)
RETURNING * RETURNING *
""", """,
case_id, section_id, quote, citation, chair_note, pdf_document_id, practice_area, case_id, section_id, quote, citation, chair_note, pdf_document_id,
practice_area, argument_id, case_law_id, verified,
) )
return dict(row) return dict(row)
async def set_case_precedent_verified(
precedent_id: UUID, verified: bool, chair_note: str | None = None,
) -> dict | None:
"""Toggle the chair-verification gate on an attached precedent (#154), optionally
updating the chair note. Sets verified_at when verifying, clears it when un-verifying."""
pool = await get_pool()
sets = ["verified = $2", "verified_at = CASE WHEN $2 THEN now() ELSE NULL END",
"updated_at = now()"]
params: list = [precedent_id, verified]
if chair_note is not None:
sets.append(f"chair_note = ${len(params) + 1}")
params.append(chair_note)
row = await pool.fetchrow(
f"UPDATE case_precedents SET {', '.join(sets)} WHERE id = $1 RETURNING *",
*params,
)
return dict(row) if row else None
async def list_case_precedents(case_id: UUID) -> list[dict]: async def list_case_precedents(case_id: UUID) -> list[dict]:
"""List all precedents attached to a case, ordered by section then creation time.""" """List all precedents attached to a case, ordered by section then creation time."""
pool = await get_pool() pool = await get_pool()
rows = await pool.fetch( rows = await pool.fetch(
""" """
SELECT id, case_id, section_id, quote, citation, chair_note, SELECT id, case_id, section_id, quote, citation, chair_note,
pdf_document_id, practice_area, created_at, updated_at pdf_document_id, practice_area, argument_id, case_law_id,
verified, verified_at, created_at, updated_at
FROM case_precedents FROM case_precedents
WHERE case_id = $1 WHERE case_id = $1
ORDER BY section_id NULLS LAST, created_at ORDER BY section_id NULLS LAST, created_at
@@ -4253,6 +4499,34 @@ async def case_number_collides(case_number: str, exclude_id: UUID) -> bool:
)) ))
# Subject-tag convention for case_law topic hubs (powers /graph topic nodes).
# A simple multi-word Hebrew phrase is stored with underscores between words
# ("היטל_השבחה") — the documented extraction contract (precedent_library tool
# examples: קווי_בניין, מועד_קביעת_שומה) that ~99% of the corpus already follows.
# The same phrase arriving with spaces ("היטל השבחה") is the same topic in
# disguise and spawns a duplicate graph hub. Normalize at the single case_law
# write chokepoint (G1) so no path can re-introduce the split. Tags carrying
# punctuation/digits/dashes (e.g. "פטור מותנה — סעיף 19(ג)") are left untouched —
# the convention covers only plain Hebrew word phrases.
_PURE_HEBREW_WORDS = re.compile(r"^[א-ת]+(?: [א-ת]+)+$")
def _normalize_subject_tags(tags) -> list[str]:
"""Strip, apply the underscore convention to plain Hebrew phrases, dedup."""
out: list[str] = []
seen: set[str] = set()
for raw in tags or []:
t = str(raw).strip()
if not t:
continue
if _PURE_HEBREW_WORDS.match(t):
t = t.replace(" ", "_")
if t not in seen:
seen.add(t)
out.append(t)
return out
async def create_external_case_law( async def create_external_case_law(
case_number: str, case_number: str,
case_name: str, case_name: str,
@@ -4270,15 +4544,27 @@ async def create_external_case_law(
precedent_level: str = "", precedent_level: str = "",
is_binding: bool = True, is_binding: bool = True,
document_id: UUID | None = None, document_id: UUID | None = None,
citation_formatted: str = "",
) -> dict: ) -> dict:
"""Insert a chair-uploaded external precedent into case_law. """Insert a chair-uploaded external precedent into case_law.
If a row with this ``case_number`` already exists with If a row with this ``case_number`` already exists with
source_kind='cited_only' (auto-discovered), promote it to source_kind='cited_only' (auto-discovered), promote it to
source_kind='external_upload' and fill in the missing fields. source_kind='external_upload' and fill in the missing fields.
``citation_formatted`` seeds the מראה-מקום from the chair's typed upload-form
value so the field is NEVER blank between upload and metadata extraction (the
inline enrichment later UPGRADES it to the canonical derived form via
``apply_to_record(force_citation=True)``). On a cited_only→external promotion
an existing non-empty value (a prior chair edit) is preserved.
""" """
# INV-DM7: an appeals-committee source is persuasive even when uploaded via
# the external path — coerce to non-binding so it matches the committee
# invariant and case_law_committee_not_binding_check (G1).
if source_type == "appeals_committee":
is_binding = False
pool = await get_pool() pool = await get_pool()
tags_json = json.dumps(subject_tags or [], ensure_ascii=False) tags_json = json.dumps(_normalize_subject_tags(subject_tags), ensure_ascii=False)
async with pool.acquire() as conn: async with pool.acquire() as conn:
# Atomic upsert on the V15 partial unique index # Atomic upsert on the V15 partial unique index
# uq_case_law_external_number (case_number) WHERE source_kind <> 'internal_committee'. # uq_case_law_external_number (case_number) WHERE source_kind <> 'internal_committee'.
@@ -4295,11 +4581,12 @@ async def create_external_case_law(
summary, key_quote, full_text, source_url, summary, key_quote, full_text, source_url,
source_kind, document_id, extraction_status, source_kind, document_id, extraction_status,
halacha_extraction_status, practice_area, appeal_subtype, halacha_extraction_status, practice_area, appeal_subtype,
headnote, source_type, precedent_level, is_binding, content_hash headnote, source_type, precedent_level, is_binding, content_hash,
citation_formatted
) VALUES ( ) VALUES (
$1, $2, $3, $4, $5, $6, $7, $8, $9, $1, $2, $3, $4, $5, $6, $7, $8, $9,
'external_upload', $10, 'processing', 'pending', 'external_upload', $10, 'processing', 'pending',
$11, $12, $13, $14, $15, $16, $17 $11, $12, $13, $14, $15, $16, $17, $18
) )
ON CONFLICT (case_number) WHERE source_kind <> 'internal_committee' ON CONFLICT (case_number) WHERE source_kind <> 'internal_committee'
DO UPDATE SET DO UPDATE SET
@@ -4321,14 +4608,16 @@ async def create_external_case_law(
source_kind = 'external_upload', source_kind = 'external_upload',
extraction_status = 'processing', extraction_status = 'processing',
halacha_extraction_status = 'pending', halacha_extraction_status = 'pending',
content_hash = EXCLUDED.content_hash content_hash = EXCLUDED.content_hash,
citation_formatted = COALESCE(
NULLIF(case_law.citation_formatted, ''), EXCLUDED.citation_formatted)
RETURNING * RETURNING *
""", """,
case_number, case_name, court, decision_date, tags_json, case_number, case_name, court, decision_date, tags_json,
summary, key_quote, full_text, source_url, summary, key_quote, full_text, source_url,
document_id, practice_area, appeal_subtype, headnote, document_id, practice_area, appeal_subtype, headnote,
source_type, precedent_level, is_binding, source_type, precedent_level, is_binding,
_content_hash(full_text), _content_hash(full_text), (citation_formatted or "").strip(),
) )
return _row_to_case_law(row) return _row_to_case_law(row)
@@ -4345,7 +4634,7 @@ async def create_internal_committee_decision(
appeal_subtype: str = "", appeal_subtype: str = "",
subject_tags: list[str] | None = None, subject_tags: list[str] | None = None,
summary: str = "", summary: str = "",
is_binding: bool = True, is_binding: bool = False,
document_id: UUID | None = None, document_id: UUID | None = None,
proceeding_type: str = "ערר", proceeding_type: str = "ערר",
) -> dict: ) -> dict:
@@ -4355,9 +4644,13 @@ async def create_internal_committee_decision(
exist as both 'ערר' and 'בל"מ' (an extension-of-time request can be exist as both 'ערר' and 'בל"מ' (an extension-of-time request can be
filed against an existing appeal with the same number). filed against an existing appeal with the same number).
""" """
# INV-DM7: authority is structural for this source — an appeals-committee
# decision is persuasive, never binding. Coerce regardless of caller so the
# value can never violate case_law_committee_not_binding_check (G1).
is_binding = False
pool = await get_pool() pool = await get_pool()
case_number = _canonical_case_number(case_number) case_number = _canonical_case_number(case_number)
tags_json = json.dumps(subject_tags or [], ensure_ascii=False) tags_json = json.dumps(_normalize_subject_tags(subject_tags), ensure_ascii=False)
async with pool.acquire() as conn: async with pool.acquire() as conn:
# Atomic upsert on V15 partial unique index # Atomic upsert on V15 partial unique index
# uq_case_law_internal_number_proc (case_number, proceeding_type) # uq_case_law_internal_number_proc (case_number, proceeding_type)
@@ -4498,6 +4791,12 @@ async def update_case_law(case_law_id: UUID, **fields) -> dict | None:
"proceeding_type", "citation_formatted", "parties", "proceeding_type", "citation_formatted", "parties",
} }
updates = {k: v for k, v in fields.items() if k in allowed} updates = {k: v for k, v in fields.items() if k in allowed}
# INV-DM7: reclassifying a row to an appeals-committee source makes it
# persuasive — coerce is_binding in the same patch so the UPDATE can't
# violate case_law_committee_not_binding_check (e.g. the Gemini metadata
# extractor relabels an external court_ruling as appeals_committee).
if updates.get("source_type") == "appeals_committee":
updates["is_binding"] = False
if not updates: if not updates:
return await get_case_law(case_law_id) return await get_case_law(case_law_id)
@@ -4506,12 +4805,21 @@ async def update_case_law(case_law_id: UUID, **fields) -> dict | None:
params: list = [case_law_id] params: list = [case_law_id]
for i, (k, v) in enumerate(updates.items(), start=2): for i, (k, v) in enumerate(updates.items(), start=2):
if k == "subject_tags": if k == "subject_tags":
v = json.dumps(v or [], ensure_ascii=False) v = json.dumps(_normalize_subject_tags(v), ensure_ascii=False)
set_parts.append(f"{k} = ${i}") set_parts.append(f"{k} = ${i}")
params.append(v) params.append(v)
sql = f"UPDATE case_law SET {', '.join(set_parts)} WHERE id = $1 RETURNING *" sql = f"UPDATE case_law SET {', '.join(set_parts)} WHERE id = $1 RETURNING *"
row = await pool.fetchrow(sql, *params) row = await pool.fetchrow(sql, *params)
return _row_to_case_law(row) if row else None if row is None:
return None
# `searchable` is a DERIVED completeness flag (INV-DM1). It is computed at
# ingest end and after the Gemini metadata extractor — but internal-committee
# decisions skip Gemini, so when their summary/tags/metadata are filled later
# (manual edit, deterministic enrichment) the flag would otherwise go stale
# and the row stays invisible to RAG. Recompute at this write so the derived
# value never drifts from the content (G1: normalize at the source).
await recompute_searchable(case_law_id)
return await get_case_law(case_law_id)
async def set_case_law_extraction_status(case_law_id: UUID, status: str) -> None: async def set_case_law_extraction_status(case_law_id: UUID, status: str) -> None:
@@ -4920,14 +5228,24 @@ async def search_digests_semantic(
subject_tag: str = "", subject_tag: str = "",
concept_tag: str = "", concept_tag: str = "",
limit: int = 10, limit: int = 10,
linked_only: bool | None = None,
) -> list[dict]: ) -> list[dict]:
"""Pure-semantic search over the digests radar (X12). Single vector per row """Pure-semantic search over the digests radar (X12). Single vector per row
(no chunks/halachot), so no RRF here — see X12 §6. Joins the linked ruling's (no chunks/halachot), so no RRF here — see X12 §6. Joins the linked ruling's
citation when present so the researcher sees the pointer target directly.""" citation when present so the researcher sees the pointer target directly.
``linked_only``: None = all digests (default); False = only UNLINKED digests
(``linked_case_law_id IS NULL`` — rulings we don't hold yet, the case radar's
target set); True = only linked digests.
"""
pool = await get_pool() pool = await get_pool()
conditions = ["d.embedding IS NOT NULL"] conditions = ["d.embedding IS NOT NULL"]
params: list = [query_embedding, limit] params: list = [query_embedding, limit]
idx = 3 idx = 3
if linked_only is True:
conditions.append("d.linked_case_law_id IS NOT NULL")
elif linked_only is False:
conditions.append("d.linked_case_law_id IS NULL")
if practice_area: if practice_area:
conditions.append(f"d.practice_area = ${idx}") conditions.append(f"d.practice_area = ${idx}")
params.append(practice_area) params.append(practice_area)
@@ -6438,8 +6756,13 @@ async def refresh_verified_layer() -> dict:
"""Recompute the verified/cite_count layer from chair citations (#153). """Recompute the verified/cite_count layer from chair citations (#153).
'verified' = the principle's SOURCE precedent was cited by a chair (any 'verified' = the principle's SOURCE precedent was cited by a chair (any
committee decision). 'cite_count' = # distinct chair decisions citing it. This committee decision) WITHOUT a negative treatment. 'cite_count' = # distinct
is the ONLY trust signal — never human review. Idempotent (full recompute). chair decisions citing it whose treatment is NOT negative (X11 §4 / INV-COR2:
a precedent *distinguished*/*criticized*/*questioned*/*overruled* must never
gain authority from those citations). Unclassified edges (treatment='') count
as neutral-positive until ``classify_citation_treatments.py`` labels them, so
the signal degrades gracefully before classification has run. This is the ONLY
trust signal — never human review. Idempotent (full recompute).
Returns {verified_principles, verified_precedents}. Returns {verified_principles, verified_precedents}.
""" """
pool = await get_pool() pool = await get_pool()
@@ -6456,6 +6779,8 @@ async def refresh_verified_layer() -> dict:
" JOIN case_law src ON src.id = pic.source_case_law_id " " JOIN case_law src ON src.id = pic.source_case_law_id "
" WHERE src.source_kind='internal_committee' " " WHERE src.source_kind='internal_committee' "
" AND pic.cited_case_law_id IS NOT NULL " " AND pic.cited_case_law_id IS NOT NULL "
" AND coalesce(pic.treatment,'') NOT IN "
" ('distinguished','criticized','questioned','overruled') "
" GROUP BY pic.cited_case_law_id) " " GROUP BY pic.cited_case_law_id) "
"UPDATE halachot h SET verified=true, cite_count=cc.n, updated_at=now() " "UPDATE halachot h SET verified=true, cite_count=cc.n, updated_at=now() "
"FROM cc WHERE h.case_law_id = cc.id") "FROM cc WHERE h.case_law_id = cc.id")
@@ -6466,6 +6791,55 @@ async def refresh_verified_layer() -> dict:
return {"verified_principles": row["vp"], "verified_precedents": row["vc"]} return {"verified_principles": row["vp"], "verified_precedents": row["vc"]}
# X11 §4 treatment buckets (mirrors corroboration.TREATMENT_POSITIVE/NEGATIVE) —
# kept here so the SQL layer can label a breakdown without importing the service.
_TREATMENT_POSITIVE = ("followed", "explained")
_TREATMENT_NEGATIVE = ("distinguished", "criticized", "questioned", "overruled")
async def citation_authority(case_law_ids: list["UUID"]) -> dict[str, dict]:
"""Per-precedent incoming-citation breakdown by treatment (X11 Phase 2, #154).
For each precedent id → how many DISTINCT committee decisions cite it, split
into positive (followed/explained), negative (distinguished/criticized/
questioned/overruled) and unclassified (treatment not yet labelled). This is the
'cited_by N (X אומצו, Y אובחנו)' authority signal surfaced to research agents so
they can argue authority — and avoid leaning on a precedent that was repeatedly
distinguished/overruled. Counts distinct sources; a source with no treatment yet
falls in 'unclassified'. Returns {} for ids with no incoming committee citations.
"""
if not case_law_ids:
return {}
pool = await get_pool()
rows = await pool.fetch(
"SELECT pic.cited_case_law_id::text AS id, "
" coalesce(NULLIF(pic.treatment, ''), 'unclassified') AS t, "
" count(DISTINCT pic.source_case_law_id) AS n "
"FROM precedent_internal_citations pic "
"JOIN case_law src ON src.id = pic.source_case_law_id "
"WHERE src.source_kind = 'internal_committee' "
" AND pic.cited_case_law_id = ANY($1::uuid[]) "
"GROUP BY 1, 2",
case_law_ids,
)
out: dict[str, dict] = {}
for r in rows:
d = out.setdefault(r["id"], {
"total": 0, "positive": 0, "negative": 0, "unclassified": 0,
"by_treatment": {},
})
t, n = r["t"], int(r["n"])
d["by_treatment"][t] = d["by_treatment"].get(t, 0) + n
d["total"] += n
if t in _TREATMENT_POSITIVE:
d["positive"] += n
elif t in _TREATMENT_NEGATIVE:
d["negative"] += n
else:
d["unclassified"] += n
return out
async def list_canonical_instances(canonical_id: "UUID") -> list[dict]: async def list_canonical_instances(canonical_id: "UUID") -> list[dict]:
"""List all halachot (instances) sharing a canonical_id — used by the UI accordion.""" """List all halachot (instances) sharing a canonical_id — used by the UI accordion."""
pool = await get_pool() pool = await get_pool()
@@ -8030,7 +8404,33 @@ _MP_PROVENANCE_COLS = """,
WHERE pic.cited_case_law_id = mp.linked_case_law_id WHERE pic.cited_case_law_id = mp.linked_case_law_id
AND COALESCE(src.case_number, '') <> '' AND COALESCE(src.case_number, '') <> ''
) AS cited_by_precedents, ) AS cited_by_precedents,
substring(mp.notes from 'מס''?\\s*([0-9]+)') AS yomon_number""" substring(mp.notes from 'מס''?\\s*([0-9]+)') AS yomon_number,
-- Bridge to the corpus citation graph (G2: read-time, no stored
-- duplication). For an OPEN gap (no linked_case_law_id yet) find
-- which committee DECISIONS cite this ruling, matched on the
-- normalized docket number (same normalization the relinker uses:
-- strip to digits/dashes, slash->dash). Surfaces "צוטט ע\"י <יו\"ר>"
-- + the deciding case number in the UI.
(SELECT array_agg(DISTINCT src.chair_name ORDER BY src.chair_name)
FROM precedent_internal_citations picc
JOIN case_law src ON src.id = picc.source_case_law_id
AND src.source_kind = 'internal_committee'
WHERE split_part(mp.citation_norm, '|', 2) <> ''
AND regexp_replace(replace(picc.cited_case_number, '/', '-'),
'[^0-9-]', '', 'g')
= split_part(mp.citation_norm, '|', 2)
AND COALESCE(src.chair_name, '') <> ''
) AS cited_by_chairs,
(SELECT array_agg(DISTINCT src.case_number ORDER BY src.case_number)
FROM precedent_internal_citations picd
JOIN case_law src ON src.id = picd.source_case_law_id
AND src.source_kind = 'internal_committee'
WHERE split_part(mp.citation_norm, '|', 2) <> ''
AND regexp_replace(replace(picd.cited_case_number, '/', '-'),
'[^0-9-]', '', 'g')
= split_part(mp.citation_norm, '|', 2)
AND COALESCE(src.case_number, '') <> ''
) AS cited_by_decisions"""
async def list_missing_precedents( async def list_missing_precedents(

View File

@@ -381,6 +381,119 @@ async def link_digest(digest_id: UUID | str, case_law_id: UUID | str) -> dict:
} }
async def _radar_enrich(h: dict, score: float, matched_issues: list[str]) -> dict:
"""Shape one radar hit into a chair lead: gap status + suggested action +
which case ISSUE(s) it answers. The action points at the underlying RULING,
never the digest (INV-DIG1)."""
cit = (h.get("underlying_citation") or "").strip()
gap = await db.find_missing_precedent_by_citation(cit) if cit else None
in_corpus = await db.find_case_law_by_citation_fuzzy(cit) if cit else None
if in_corpus:
action = "available_link" # ruling actually IS in the corpus → just link the digest
elif gap and (gap.get("status") in ("uploaded", "closed")):
action = "fetched" # already obtained
elif gap:
action = "gap_open" # flagged as missing — can request a fetch
else:
action = "new_lead" # not even flagged yet — the highest-value alert
return {
"digest_id": str(h["id"]),
"yomon_number": h.get("yomon_number"),
"headline": h.get("headline_holding") or h.get("summary") or "",
"underlying_citation": cit,
"underlying_court": h.get("underlying_court") or "",
"score": round(score, 3),
"matched_issues": matched_issues,
"missing_precedent_id": str(gap["id"]) if gap else None,
"missing_precedent_status": gap.get("status") if gap else None,
"action": action,
}
async def case_digest_radar(
case_number: str,
limit: int = 5,
min_score: float = 0.45,
) -> dict:
"""Case-contextual digest radar (X12) — the chair-facing "שים לב" lead.
Surfaces UNLINKED digests (``linked_case_law_id IS NULL`` — rulings we don't hold
yet) whose topic is semantically close to THIS case, so a relevant ruling we only
know about via a digest doesn't fall through the cracks while the case is decided.
Query calibration: prefer the analyst's DISTILLED legal arguments (one crisp issue
per row — ``argument_title`` + ``legal_topic``) over the dozens of raw claims. Each
issue is searched separately and the leads are MERGED, so every lead is attributed
to the case issue(s) it answers (``matched_issues``) — far higher precision than
one blended query of noisy claims. Falls back to raw claims pre-aggregation
(``source`` reports which path ran). Each lead carries the underlying ruling's gap
status + a suggested action.
INV-DIG1: this is RADAR — the digest is never cited; the lead points at the
underlying *ruling* (fetch / upload / link), never the digest itself. Read-only.
"""
case = await db.get_case_by_number(case_number)
if not case:
return {"status": "case_not_found", "case_number": case_number, "leads": [], "count": 0}
case_id = case["id"]
if isinstance(case_id, str):
case_id = UUID(case_id)
ctx = " ".join(x for x in [case.get("title") or "", case.get("appeal_subtype") or ""] if x).strip()
# Preferred source: the analyst's distilled CREAC issues (one per legal_argument).
issues: list[tuple[str, str]] = [] # (label, embed_text)
try:
from legal_mcp.services import argument_aggregator
for a in await argument_aggregator.get_legal_arguments(case_id):
label = (a.get("argument_title") or a.get("legal_topic") or "").strip()
topic = (a.get("legal_topic") or "").strip()
body = f"{label}. {topic}".strip(". ").strip()
if body:
issues.append((label, f"{ctx} {body}".strip()))
except Exception as e: # noqa: BLE001 — arguments are optional; fall back to claims
logger.warning("case_digest_radar: get_legal_arguments failed for %s: %s", case_number, e)
merged: dict[str, dict] = {} # digest_id -> {"hit", "best", "issues": set}
if issues:
source = "legal_arguments"
issues = issues[:25] # already distilled — bound the per-issue search fan-out
vecs = await embeddings.embed_texts([t for _, t in issues], input_type="query")
for (label, _), vec in zip(issues, vecs):
for h in await db.search_digests_semantic(vec, limit=6, linked_only=False):
s = float(h.get("score", 0) or 0)
if s < min_score:
continue
m = merged.setdefault(str(h["id"]), {"hit": h, "best": s, "issues": set()})
m["best"] = max(m["best"], s)
if label:
m["issues"].add(label)
else:
# Fallback: blended query from raw claims (pre-aggregation), one search.
source = "claims"
parts = [ctx]
try:
claims = await db.get_claims(case_id)
parts += [(c.get("claim_text") or "") for c in claims[:20]]
except Exception as e: # noqa: BLE001
logger.warning("case_digest_radar: get_claims failed for %s: %s", case_number, e)
query = " ".join(p for p in parts if p).strip()
if not query:
return {"status": "no_topic", "case_number": case_number,
"source": source, "leads": [], "count": 0}
vec = (await embeddings.embed_texts([query], input_type="query"))[0]
for h in await db.search_digests_semantic(vec, limit=max(limit * 4, 20), linked_only=False):
s = float(h.get("score", 0) or 0)
if s < min_score:
continue
merged.setdefault(str(h["id"]), {"hit": h, "best": s, "issues": set()})
ordered = sorted(merged.values(), key=lambda m: -m["best"])[:max(1, limit)]
leads = [await _radar_enrich(m["hit"], m["best"], sorted(m["issues"])) for m in ordered]
return {"status": "ok", "case_number": case_number, "source": source,
"issues_used": [lbl for lbl, _ in issues] if source == "legal_arguments" else None,
"leads": leads, "count": len(leads)}
async def relink_digest(digest_id: UUID | str) -> dict: async def relink_digest(digest_id: UUID | str) -> dict:
"""Re-run autolink for an unlinked digest. No-op if already linked / no match.""" """Re-run autolink for an unlinked digest. No-op if already linked / no match."""
digest = await db.get_digest(digest_id) digest = await db.get_digest(digest_id)

View File

@@ -1,23 +1,33 @@
"""Text extraction from PDF, DOCX, DOC, and RTF files. """Text extraction from PDF, DOCX, DOC, and RTF files.
Primary PDF extraction: PyMuPDF direct text (for born-digital PDFs). Primary PDF extraction: PyMuPDF direct text (for born-digital PDFs).
Fallback: Google Cloud Vision OCR (for scanned documents). Fallback: Mistral OCR (for scanned documents or broken OCR layers).
Routing logic (document-level, not per-page):
1. PyMuPDF extracts text from every page.
2. Pages are quality-checked via _text_quality_ok().
3. If ALL pages pass → use PyMuPDF output (free, ~50ms, no API call).
4. If ANY page fails → call Mistral OCR once for the entire PDF.
Mistral returns per-page Markdown; page_offsets are computed from it.
DOC files: converted to DOCX via LibreOffice before extraction. DOC files: converted to DOCX via LibreOffice before extraction.
Post-processing: Hebrew abbreviation quote fixer. Post-processing: Hebrew abbreviation quote fixer (PyMuPDF path only;
Mistral handles gershayim natively).
""" """
from __future__ import annotations from __future__ import annotations
import asyncio import asyncio
import base64
import io import io
import logging import logging
import re import re
import subprocess import subprocess
import tempfile import tempfile
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING
import fitz # PyMuPDF import fitz # PyMuPDF
import httpx
from PIL import Image from PIL import Image
from docx import Document as DocxDocument from docx import Document as DocxDocument
from striprtf.striprtf import rtf_to_text from striprtf.striprtf import rtf_to_text
@@ -25,36 +35,65 @@ from striprtf.striprtf import rtf_to_text
from legal_mcp import config from legal_mcp import config
from legal_mcp.services import storage from legal_mcp.services import storage
if TYPE_CHECKING:
from google.cloud import vision
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# ── Google Cloud Vision client (imported lazily — saves ~550ms at MCP startup) ── # ── Mistral OCR ───────────────────────────────────────────────────
_vision_client: "vision.ImageAnnotatorClient | None" = None _MISTRAL_OCR_URL = "https://api.mistral.ai/v1/ocr"
_MISTRAL_OCR_MODEL = "mistral-ocr-latest"
def _get_vision_client() -> "vision.ImageAnnotatorClient": async def _call_mistral_ocr(path: Path) -> list[str]:
global _vision_client """Call Mistral OCR API on a PDF. Returns per-page Markdown text list.
if _vision_client is None:
from google.cloud import vision The Mistral response contains ``pages[i].markdown`` for each page.
_vision_client = vision.ImageAnnotatorClient( If the response has fewer pages than the PDF, trailing pages are
client_options={"api_key": config.GOOGLE_CLOUD_VISION_API_KEY} padded with empty strings by the caller.
"""
if not config.MISTRAL_API_KEY:
raise RuntimeError(
"MISTRAL_API_KEY not configured — cannot OCR scanned PDF. "
"Set the env var in Coolify."
) )
return _vision_client
pdf_b64 = base64.b64encode(path.read_bytes()).decode()
async with httpx.AsyncClient(timeout=300.0) as client:
resp = await client.post(
_MISTRAL_OCR_URL,
headers={
"Authorization": f"Bearer {config.MISTRAL_API_KEY}",
"Content-Type": "application/json",
},
json={
"model": _MISTRAL_OCR_MODEL,
"document": {
"type": "document_url",
"document_url": f"data:application/pdf;base64,{pdf_b64}",
},
"include_image_base64": False,
},
)
if resp.status_code != 200:
raise RuntimeError(
f"Mistral OCR returned {resp.status_code}: {resp.text[:400]}"
)
pages = resp.json().get("pages", [])
return [p.get("markdown", "") for p in pages]
# ── Hebrew text quality detection ──────────────────────────────── # ── Hebrew text quality detection ────────────────────────────────
_HEBREW_RE = re.compile(r'[\u0590-\u05FF]') _HEBREW_RE = re.compile(r'[֐-׿]')
_WORD_RE = re.compile(r'\S+') _WORD_RE = re.compile(r'\S+')
def _text_quality_ok(text: str) -> bool: def _text_quality_ok(text: str) -> bool:
"""Check if extracted text is real content vs broken OCR layer. """Check if PyMuPDF-extracted text is genuine Hebrew legal content.
Returns True if text appears to be genuine Hebrew legal content. Returns True if text appears to be real content.
Broken OCR layers from scanned PDFs often have: Broken OCR layers from scanned PDFs often have:
- Very short words / single-character fragments - Very short words / single-character fragments
- Each word on its own line (high words-per-line ratio) - Each word on its own line (high words-per-line ratio)
@@ -64,26 +103,19 @@ def _text_quality_ok(text: str) -> bool:
if len(words) < 10: if len(words) < 10:
return False return False
# Average word length — real Hebrew words avg 4-6 chars.
avg_len = sum(len(w) for w in words) / len(words) avg_len = sum(len(w) for w in words) / len(words)
if avg_len < 2.5: if avg_len < 2.5:
return False return False
# Percentage of single-character "words"
single_char_pct = sum(1 for w in words if len(w) == 1) / len(words) single_char_pct = sum(1 for w in words if len(w) == 1) / len(words)
if single_char_pct > 0.4: if single_char_pct > 0.4:
return False return False
# Words per line — broken OCR puts each word on its own line.
# Real text has 5-15 words per line; broken OCR has ~1-2.
lines = [l for l in text.split("\n") if l.strip()] lines = [l for l in text.split("\n") if l.strip()]
if lines: if lines and len(words) / len(lines) < 3.0:
words_per_line = len(words) / len(lines) return False
if words_per_line < 3.0:
return False
# Hebrew character ratio among letter characters letters = re.findall(r'[a-zA-Z֐-׿]', text)
letters = re.findall(r'[a-zA-Z\u0590-\u05FF]', text)
if letters: if letters:
hebrew_pct = sum(1 for c in letters if _HEBREW_RE.match(c)) / len(letters) hebrew_pct = sum(1 for c in letters if _HEBREW_RE.match(c)) / len(letters)
if hebrew_pct < 0.5: if hebrew_pct < 0.5:
@@ -92,7 +124,7 @@ def _text_quality_ok(text: str) -> bool:
return True return True
# ── Hebrew abbreviation quote fixer ────────────────────────────── # ── Hebrew abbreviation quote fixer (PyMuPDF path only) ──────────
_HEBREW_ABBREV_FIXES: dict[str, str] = { _HEBREW_ABBREV_FIXES: dict[str, str] = {
'עוהייד': 'עוה"ד', 'עוהייד': 'עוה"ד',
@@ -111,50 +143,133 @@ _HEBREW_ABBREV_FIXES: dict[str, str] = {
'יחייד': 'יח"ד', 'יחייד': 'יח"ד',
'בייכ': 'ב"כ', 'בייכ': 'ב"כ',
# Patterns where double-yod (יי) substitutes for gershayim (״) in born-digital PDFs # Patterns where double-yod (יי) substitutes for gershayim (״) in born-digital PDFs
'בליימ': 'בל"מ', # בקשה להארכת מועד — appears in RTL legal docs 'בליימ': 'בל"מ',
'תמייא': 'תמ"א', # תכנית מתאר ארצית 'תמייא': 'תמ"א',
} }
_ABBREV_PATTERN = re.compile( _ABBREV_PATTERN = re.compile(
'|'.join(re.escape(k) for k in sorted(_HEBREW_ABBREV_FIXES, key=len, reverse=True)) '|'.join(re.escape(k) for k in sorted(_HEBREW_ABBREV_FIXES, key=len, reverse=True))
) )
# Matches Hebrew law year abbreviations where gershayim was encoded as double-yod.
# e.g. תשכייה → תשכ"ה, תשנייב → תשנ"ב
_HEBREW_YEAR_RE = re.compile(r'(תש[א-ת]+)יי([א-ת])') _HEBREW_YEAR_RE = re.compile(r'(תש[א-ת]+)יי([א-ת])')
def _fix_hebrew_quotes(text: str) -> str: def _fix_hebrew_quotes(text: str) -> str:
"""Fix known Hebrew abbreviation quote replacements. """Fix gershayim encoded as double-yod in born-digital PDFs."""
Applied to both Google Vision OCR output and direct PyMuPDF extraction —
some born-digital PDFs encode gershayim (״) as double-yod (יי), producing
the same corruption patterns as OCR.
"""
text = _ABBREV_PATTERN.sub(lambda m: _HEBREW_ABBREV_FIXES[m.group()], text) text = _ABBREV_PATTERN.sub(lambda m: _HEBREW_ABBREV_FIXES[m.group()], text)
text = _HEBREW_YEAR_RE.sub(r'\1"\2', text) text = _HEBREW_YEAR_RE.sub(r'\1"\2', text)
return text return text
# ── Extraction ─────────────────────────────────────────────────── # ── Page joining ──────────────────────────────────────────────────
# Separator used when joining per-page text. Constant so chunker / # Separator used when joining per-page text. Constant so chunker /
# retrofit can reproduce the join when computing page offsets. # retrofit can reproduce the join when computing page offsets.
PAGE_SEPARATOR = "\n\n" PAGE_SEPARATOR = "\n\n"
def _join_pages(pages_text: list[str]) -> tuple[str, list[int]]:
"""Join per-page text with PAGE_SEPARATOR while recording start offsets."""
offsets: list[int] = []
parts: list[str] = []
cursor = 0
for i, pg in enumerate(pages_text):
offsets.append(cursor)
parts.append(pg)
cursor += len(pg)
if i < len(pages_text) - 1:
parts.append(PAGE_SEPARATOR)
cursor += len(PAGE_SEPARATOR)
return "".join(parts), offsets
# ── PDF extraction ────────────────────────────────────────────────
async def _extract_pdf(path: Path) -> tuple[str, int, list[int]]:
"""Extract text from PDF using document-level routing.
Stage 1 — PyMuPDF pre-screen (free, ~50ms, no API call):
Run on every page. Collect per-page text; flag pages where
PyMuPDF returns < 50 chars or _text_quality_ok() fails
(scanned, blank, or broken embedded OCR layer).
Stage 2 — Mistral OCR (triggered when any page fails Stage 1):
Send the entire PDF to Mistral once. Use its per-page Markdown
for ALL pages — consistent source, no mixed plain/Markdown formats.
Mistral handles gershayim natively; no quote-fix applied.
Page offsets are always computed so the chunker can attribute each
chunk to its source page number (multimodal hybrid retrieval).
"""
doc = fitz.open(str(path))
page_count = len(doc)
# Stage 1: PyMuPDF pre-screen
pymupdf_pages: list[str] = []
failed: list[int] = []
for i in range(page_count):
text = doc[i].get_text().strip()
if len(text) > 50 and _text_quality_ok(text):
pymupdf_pages.append(_fix_hebrew_quotes(text))
else:
pymupdf_pages.append("")
failed.append(i)
doc.close()
if not failed:
logger.debug(
"PDF %s: all %d pages digital — PyMuPDF only", path.name, page_count
)
joined, offsets = _join_pages(pymupdf_pages)
return joined, page_count, offsets
# Stage 2: Mistral OCR for entire document
logger.info(
"PDF %s: %d/%d pages failed quality check → Mistral OCR",
path.name, len(failed), page_count,
)
mistral_pages = await _call_mistral_ocr(path)
# Pad if Mistral returns fewer pages than PyMuPDF counted
while len(mistral_pages) < page_count:
mistral_pages.append("")
joined, offsets = _join_pages(mistral_pages[:page_count])
return joined, page_count, offsets
def page_at_offset(offset: int, page_offsets: list[int]) -> int:
"""Return the 1-based page number containing a given char offset.
page_offsets[i] is the start of page (i+1) in the joined text.
"""
if not page_offsets:
return 1
page = 1
for i, start in enumerate(page_offsets):
if start <= offset:
page = i + 1
else:
break
return page
# ── Public entry point ────────────────────────────────────────────
async def extract_text(file_path: str) -> tuple[str, int, list[int] | None]: async def extract_text(file_path: str) -> tuple[str, int, list[int] | None]:
"""Extract text from a document file. """Extract text from a document file.
Returns: Returns:
``(text, page_count, page_offsets)`` where: ``(text, page_count, page_offsets)`` where:
- ``text``: concatenated extracted text - ``text``: extracted text. Plain text for PyMuPDF path;
Markdown for Mistral path (tables, ``##`` headers preserved).
- ``page_count``: number of pages (0 for non-PDF) - ``page_count``: number of pages (0 for non-PDF)
- ``page_offsets``: ``page_offsets[i]`` = char start offset of - ``page_offsets``: char start of each page inside ``text``,
page (i+1) inside ``text``. ``None`` for non-PDFs (where the or ``None`` for non-PDF formats
notion of pages doesn't apply). Used by the chunker to assign
a ``page_number`` to each chunk.
""" """
path = Path(file_path) path = Path(file_path)
suffix = path.suffix.lower() suffix = path.suffix.lower()
@@ -173,95 +288,11 @@ async def extract_text(file_path: str) -> tuple[str, int, list[int] | None]:
raise ValueError(f"Unsupported file type: {suffix}") raise ValueError(f"Unsupported file type: {suffix}")
def _join_pages(pages_text: list[str]) -> tuple[str, list[int]]: # ── Non-PDF formats ───────────────────────────────────────────────
"""Join per-page text with PAGE_SEPARATOR while recording the start
offset of each page in the joined output."""
offsets: list[int] = []
parts: list[str] = []
cursor = 0
for i, pg in enumerate(pages_text):
offsets.append(cursor)
parts.append(pg)
cursor += len(pg)
if i < len(pages_text) - 1:
parts.append(PAGE_SEPARATOR)
cursor += len(PAGE_SEPARATOR)
return "".join(parts), offsets
async def _extract_pdf(path: Path) -> tuple[str, int, list[int]]:
"""Extract text from PDF.
Try direct text first, fall back to Google Cloud Vision for scanned
or broken-OCR pages.
"""
doc = fitz.open(str(path))
page_count = len(doc)
pages_text: list[str] = []
for page_num in range(page_count):
page = doc[page_num]
text = page.get_text().strip()
if len(text) > 50 and _text_quality_ok(text):
pages_text.append(_fix_hebrew_quotes(text))
logger.debug("Page %d: direct extraction (%d chars, quality OK)", page_num + 1, len(text))
else:
reason = "insufficient text" if len(text) <= 50 else "low quality OCR layer"
logger.info("Page %d: Google Vision OCR (%s)", page_num + 1, reason)
pix = page.get_pixmap(dpi=300)
img_bytes = pix.tobytes("png")
ocr_text = await asyncio.to_thread(
_ocr_with_google_vision, img_bytes, page_num + 1
)
pages_text.append(ocr_text)
doc.close()
joined, offsets = _join_pages(pages_text)
return joined, page_count, offsets
def page_at_offset(offset: int, page_offsets: list[int]) -> int:
"""Look up the page number containing a given char offset.
page_offsets[i] is the start of page (i+1) in the joined text;
a chunk starting at ``offset`` belongs to the highest-indexed page
whose start is ``<= offset``. Returns 1-based page number.
"""
if not page_offsets:
return 1
# Linear scan is fine — page_offsets is short (≤ ~200 for our PDFs).
page = 1
for i, start in enumerate(page_offsets):
if start <= offset:
page = i + 1
else:
break
return page
def _ocr_with_google_vision(image_bytes: bytes, page_num: int) -> str:
"""OCR a single page image using Google Cloud Vision API."""
from google.cloud import vision # lazy: keeps MCP startup fast
client = _get_vision_client()
image = vision.Image(content=image_bytes)
response = client.document_text_detection(
image=image,
image_context=vision.ImageContext(language_hints=["he"]),
)
if response.error.message:
raise RuntimeError(
f"Google Vision error on page {page_num}: {response.error.message}"
)
text = response.full_text_annotation.text if response.full_text_annotation else ""
return _fix_hebrew_quotes(text)
def _extract_doc(path: Path) -> str: def _extract_doc(path: Path) -> str:
"""Extract text from legacy .doc file by converting to .docx via LibreOffice.""" """Extract text from legacy .doc via LibreOffice → DOCX conversion."""
with tempfile.TemporaryDirectory() as tmp_dir: with tempfile.TemporaryDirectory() as tmp_dir:
# Isolate the LibreOffice user profile per call: headless soffice # Isolate the LibreOffice user profile per call: headless soffice
# locks a single shared profile, so concurrent .doc conversions would # locks a single shared profile, so concurrent .doc conversions would
@@ -296,13 +327,13 @@ def _extract_rtf(path: Path) -> str:
# ── Multimodal page rendering (V9) ─────────────────────────────── # ── Multimodal page rendering (V9) ───────────────────────────────
# Unchanged — multimodal embedding always uses PyMuPDF-rendered images
# regardless of whether text extraction used PyMuPDF or Mistral.
def _pixmap_to_pil(pix: fitz.Pixmap) -> Image.Image: def _pixmap_to_pil(pix: fitz.Pixmap) -> Image.Image:
"""Convert a PyMuPDF pixmap to PIL.Image (RGB) without going through """Convert a PyMuPDF pixmap to PIL.Image (RGB)."""
PNG bytes. Faster than tobytes('png') → Image.open()."""
if pix.alpha: if pix.alpha:
# Drop alpha channel — voyage multimodal expects RGB.
pix = fitz.Pixmap(pix, 0) pix = fitz.Pixmap(pix, 0)
return Image.frombytes("RGB", (pix.width, pix.height), pix.samples) return Image.frombytes("RGB", (pix.width, pix.height), pix.samples)
@@ -314,12 +345,9 @@ def render_pages_for_multimodal(
thumbnail_dir: Path | None = None, thumbnail_dir: Path | None = None,
) -> list[tuple[Image.Image, Path | None]]: ) -> list[tuple[Image.Image, Path | None]]:
"""Render each PDF page as PIL.Image at ``embed_dpi`` for the """Render each PDF page as PIL.Image at ``embed_dpi`` for the
multimodal embedder, and optionally save a smaller JPEG thumbnail multimodal embedder, and optionally save JPEG thumbnails.
at ``thumb_dpi`` to ``thumbnail_dir`` for UI preview.
Returns ``[(pil_image, thumb_path_or_None), ...]`` in page order. Returns ``[(pil_image, thumb_path_or_None), ...]`` in page order.
The full-DPI image stays in memory only — only the thumbnail is
persisted to disk.
""" """
src = Path(pdf_path) src = Path(pdf_path)
if not src.is_file(): if not src.is_file():
@@ -338,17 +366,12 @@ def render_pages_for_multimodal(
thumb_path: Path | None = None thumb_path: Path | None = None
if thumbnail_dir is not None and thumb_dpi: if thumbnail_dir is not None and thumb_dpi:
thumb_path = thumbnail_dir / f"p{page_num:03d}.jpg" thumb_path = thumbnail_dir / f"p{page_num:03d}.jpg"
# Downsample the same render rather than re-rendering
# with PyMuPDF — far faster.
ratio = thumb_dpi / embed_dpi ratio = thumb_dpi / embed_dpi
thumb_size = ( thumb_size = (
max(1, int(img.width * ratio)), max(1, int(img.width * ratio)),
max(1, int(img.height * ratio)), max(1, int(img.height * ratio)),
) )
thumb = img.resize(thumb_size, Image.Resampling.LANCZOS) thumb = img.resize(thumb_size, Image.Resampling.LANCZOS)
# Persist the thumbnail (a DERIVED, regenerable artifact)
# through the storage layer (INV-STG1). Under the filesystem
# backend it lands at thumb_path exactly as before.
_tbuf = io.BytesIO() _tbuf = io.BytesIO()
thumb.save(_tbuf, "JPEG", quality=75, optimize=True) thumb.save(_tbuf, "JPEG", quality=75, optimize=True)
try: try:
@@ -366,44 +389,28 @@ def render_pages_for_multimodal(
return out return out
# ── Nevo preamble stripping ────────────────────────────────────── # ── Nevo preamble stripping ──────────────────────────────────────
_NEVO_MARKERS = ("ספרות:", "חקיקה שאוזכרה:", "מיני-רציו:", "פסקי דין שאוזכרו:", _NEVO_MARKERS = ("ספרות:", "חקיקה שאוזכרה:", "מיני-רציו:", "פסקי דין שאוזכרו:",
"כתבי עת:", "הועתק מנבו") "כתבי עת:", "הועתק מנבו")
# Markers for where the actual decision body begins (everything before is Nevo
# preamble: bibliography + מיני-רציו). Two families:
# - ועדת ערר / district openings (בפנינו / הערר שבנדון / ...)
# - COURT-RULING openings (#86.1): a פסק-דין header or the authoring judge's
# line. Without these, Nevo court judgments — exactly the ones carrying a
# מיני-רציו — slipped through unstripped (e.g. בג"ץ 1764/05).
#
# #86.2 hardening — two over-strip bugs found while backfilling:
# 1. ``פסק-דין`` headers are often markdown-wrapped (``**פסק דין**``); the old
# ``^פסק[- ]דין`` required the keyword to be the very first char of the line
# and allowed only one separator, so it missed the header and fell through
# to a citation 32K deep (עמ"נ 50567-07-21). We now tolerate leading
# markdown/whitespace and 0-3 separators.
# 2. Bare ``השופט``/``הנשיא`` matched *citations* ("השופט מ' חשין, פסקה 23"),
# stripping real decision body. The authoring-judge line ends with a COLON
# ("השופט י' עמית:"); citations use a comma. We now require the colon.
_DECISION_START = re.compile( _DECISION_START = re.compile(
r"^[ \t>*_#]{0,6}(?:" r"^[ \t>*_#]{0,6}(?:"
r"בפנינו|לפנינו|לפניי|הערר שבנדון|ועדת הערר לתכנון|רקע עובדתי|עסקינן|" r"בפנינו|לפנינו|לפניי|הערר שבנדון|ועדת הערר לתכנון|רקע עובדתי|עסקינן|"
r"פסק[ \t\-]{0,3}די(?:ן|נו)|" # פסק-דין / פסק דין / **פסק דין** header (final-nun ן vs דינו) r"פסק[ \t\-]{0,3}די(?:ן|נו)|"
r"(?:כב(?:וד)?['׳\"]?\s*)?(?:ה?שופט[ת]?|ה?נשיא[ה]?|המשנה לנשיא)\s+[^\n,]{1,40}:" # author line → colon r"(?:כב(?:וד)?['׳\"]?\s*)?(?:ה?שופט[ת]?|ה?נשיא[ה]?|המשנה לנשיא)\s+[^\n,]{1,40}:"
r")", r")",
re.MULTILINE, re.MULTILINE,
) )
def strip_nevo_preamble(text: str) -> str: def strip_nevo_preamble(text: str) -> str:
"""Remove Nevo database preamble (bibliography, legislation, mini-ratio) from decision text. """Remove Nevo database preamble (bibliography, legislation, mini-ratio).
Returns the original text unchanged if no preamble is detected. Returns the original text unchanged if no preamble is detected.
Works on both plain text (PyMuPDF) and Markdown (Mistral) since
_DECISION_START already tolerates leading ``[ \t>*_#]{0,6}``.
""" """
# Window wide enough to catch the Nevo markers even when a long court/parties
# header precedes them (court rulings push חקיקה שאוזכרה:/מיני-רציו: down).
head = text[:1500] head = text[:1500]
if not any(marker in head for marker in _NEVO_MARKERS): if not any(marker in head for marker in _NEVO_MARKERS):
return text return text
@@ -419,17 +426,7 @@ _RATIO_MARKER = "מיני-רציו:"
def extract_nevo_ratio(text: str) -> str: def extract_nevo_ratio(text: str) -> str:
"""Return the Nevo מיני-רציו block (editorial holdings summary), or ''. """Return the Nevo מיני-רציו block (editorial holdings summary), or ''."""
The mini-ratio is Nevo's own headnote — a concise, professionally-written
list of the holdings. We capture it *before* :func:`strip_nevo_preamble`
discards it, to serve as a free gold-set for benchmarking how well our
halacha extractor covers the real holdings (#86.3).
The block runs from the ``מיני-רציו:`` marker to whichever comes first:
the decision body (``_DECISION_START``) or the next preamble marker
(bibliography / legislation). Returns '' when there is no mini-ratio.
"""
if not text: if not text:
return "" return ""
start = text.find(_RATIO_MARKER) start = text.find(_RATIO_MARKER)
@@ -437,9 +434,6 @@ def extract_nevo_ratio(text: str) -> str:
return "" return ""
body = text[start + len(_RATIO_MARKER):] body = text[start + len(_RATIO_MARKER):]
# End at the earliest of: decision body start, or a following preamble
# marker (ספרות: / חקיקה שאוזכרה: / ...). Both are measured relative to
# the ratio body so we never run past it into the judgment itself.
end = len(body) end = len(body)
dm = _DECISION_START.search(body) dm = _DECISION_START.search(body)
if dm: if dm:

View File

@@ -233,6 +233,17 @@ async def ingest_document(
await db.request_metadata_extraction(case_law_id) await db.request_metadata_extraction(case_law_id)
await db.request_halacha_extraction(case_law_id) await db.request_halacha_extraction(case_law_id)
await db.recompute_searchable(case_law_id) await db.recompute_searchable(case_law_id)
# Citations resolve to a corpus row only at extraction time. A ruling
# uploaded AFTER it was first cited leaves orphan edges (NULL link),
# so re-link them to this freshly-ingested row now — the citation
# graph self-heals instead of permanently under-counting (G1). Non-fatal.
try:
from legal_mcp.services import citation_extractor as _ce
relinked = await _ce.relink_orphan_citations(case_law_id)
if relinked:
logger.info("relinked %d orphan citation(s) -> %s", relinked, case_law_id)
except Exception as e: # noqa: BLE001 — relink is best-effort
logger.warning("citation relink failed (non-fatal): %s", e)
await progress("completed", 100, await progress("completed", 100,
f"נקלט: {stored_chunks} chunks. חילוץ הלכות ומטא-דאטה ממתינים בתור.") f"נקלט: {stored_chunks} chunks. חילוץ הלכות ומטא-דאטה ממתינים בתור.")

View File

@@ -28,14 +28,14 @@ logger = logging.getLogger(__name__)
INTERNAL_DECISIONS_DIR = Path(config.DATA_DIR) / "internal-decisions" INTERNAL_DECISIONS_DIR = Path(config.DATA_DIR) / "internal-decisions"
_VALID_PRACTICE_AREAS = frozenset({"", "rishuy_uvniya", "betterment_levy", "compensation_197"}) _VALID_PRACTICE_AREAS = frozenset({"", "rishuy_uvniya", "betterment_levy", "compensation_197"})
_VALID_DISTRICTS = frozenset({"", "ירושלים", "מרכז", "תל אביב", "צפון", "דרום", "ארצי"}) _VALID_DISTRICTS = frozenset({"", "ירושלים", "מרכז", "תל אביב", "חיפה", "צפון", "דרום", "ארצי"})
_COURT_TO_DISTRICT = [ _COURT_TO_DISTRICT = [
("ירושלים", "ירושלים"), ("ירושלים", "ירושלים"),
("תל אביב", "תל אביב"), ("תל אביב", "תל אביב"),
('ת"א', "תל אביב"), ('ת"א', "תל אביב"),
("מרכז", "מרכז"), ("מרכז", "מרכז"),
("חיפה", "צפון"), ("חיפה", "חיפה"),
("צפון", "צפון"), ("צפון", "צפון"),
("דרום", "דרום"), ("דרום", "דרום"),
("ארצי", "ארצי"), ("ארצי", "ארצי"),
@@ -77,7 +77,7 @@ async def _create_internal_record(**kw) -> dict:
appeal_subtype=(kw.get("appeal_subtype") or "").strip(), appeal_subtype=(kw.get("appeal_subtype") or "").strip(),
subject_tags=list(kw.get("subject_tags") or []), subject_tags=list(kw.get("subject_tags") or []),
summary=(kw.get("summary") or "").strip(), summary=(kw.get("summary") or "").strip(),
is_binding=kw.get("is_binding", True), is_binding=kw.get("is_binding", False), # INV-DM7: committee = persuasive
document_id=kw.get("document_id"), document_id=kw.get("document_id"),
proceeding_type=kw.get("proceeding_type") or "ערר", proceeding_type=kw.get("proceeding_type") or "ערר",
) )
@@ -108,7 +108,7 @@ async def ingest_internal_decision(
appeal_subtype: str = "", appeal_subtype: str = "",
subject_tags: list[str] | None = None, subject_tags: list[str] | None = None,
summary: str = "", summary: str = "",
is_binding: bool = True, is_binding: bool = False, # INV-DM7: committee sources are persuasive
file_path: str | Path | None = None, file_path: str | Path | None = None,
text: str | None = None, text: str | None = None,
document_id: UUID | None = None, document_id: UUID | None = None,
@@ -229,6 +229,7 @@ async def migrate_from_external_corpus(dry_run: bool = False) -> dict:
await conn.execute( await conn.execute(
"""UPDATE case_law """UPDATE case_law
SET source_kind = 'internal_committee', SET source_kind = 'internal_committee',
is_binding = FALSE, -- INV-DM7: committee = persuasive
district = CASE WHEN $2 <> '' THEN $2 ELSE district END district = CASE WHEN $2 <> '' THEN $2 ELSE district END
WHERE id = $1""", WHERE id = $1""",
row["id"], district, row["id"], district,

View File

@@ -68,7 +68,11 @@ def _external_staging_subdir(inputs: dict) -> str:
async def _create_external_record(**kw) -> dict: async def _create_external_record(**kw) -> dict:
"""Adapter: maps canonical inputs (citation) to create_external_case_law(case_number).""" """Adapter: maps canonical inputs (citation) to create_external_case_law(case_number).
The chair's typed citation seeds ``citation_formatted`` so the מראה-מקום is never
blank before metadata extraction; the inline enrichment upgrades it to the
canonical form (see ``ingest_precedent``)."""
return await db.create_external_case_law( return await db.create_external_case_law(
case_number=kw["citation"].strip(), case_number=kw["citation"].strip(),
case_name=kw["case_name"], case_name=kw["case_name"],
@@ -84,6 +88,7 @@ async def _create_external_record(**kw) -> dict:
precedent_level=kw.get("precedent_level", ""), precedent_level=kw.get("precedent_level", ""),
is_binding=kw.get("is_binding", True), is_binding=kw.get("is_binding", True),
document_id=kw.get("document_id"), document_id=kw.get("document_id"),
citation_formatted=kw["citation"].strip(),
) )
@@ -126,10 +131,28 @@ async def ingest_precedent(
"appeal_subtype": appeal_subtype, "subject_tags": subject_tags, "appeal_subtype": appeal_subtype, "subject_tags": subject_tags,
"is_binding": is_binding, "headnote": headnote, "summary": summary, "is_binding": is_binding, "headnote": headnote, "summary": summary,
} }
return await ingest.ingest_document( result = await ingest.ingest_document(
_EXTERNAL_SPEC, inputs=inputs, file_path=file_path, _EXTERNAL_SPEC, inputs=inputs, file_path=file_path,
document_id=document_id, progress=progress, document_id=document_id, progress=progress,
) )
# Inline metadata enrichment (Gemini, in-container — gemini_session is direct
# REST, no local CLI). UPGRADES the seeded provisional citation_formatted to the
# canonical derived form (parties + reporter + date) so the מראה-מקום is correct
# the moment the chair opens /precedents/[id] — not deferred to the local drainer.
# force_citation=True overwrites the seed only; chair edits aren't reachable yet
# (row just created). Best-effort: on no key / API failure the seed remains and
# the queued metadata drain (request_metadata_extraction, already set) is the
# fallback. Halacha extraction is NOT touched here — it stays local (claude_session).
cid = result.get("case_law_id") if isinstance(result, dict) else None
if cid:
try:
from legal_mcp.services import precedent_metadata_extractor as _pme
r = await _pme.extract_and_apply(UUID(str(cid)), force_citation=True)
if r.get("status") == "completed":
await db.set_case_law_metadata_status(UUID(str(cid)), "completed")
except Exception as e: # noqa: BLE001 — enrichment is best-effort; drainer is fallback
logger.warning("inline metadata enrichment failed for %s (drainer will retry): %s", cid, e)
return result
async def reextract_halachot( async def reextract_halachot(
@@ -360,7 +383,10 @@ async def reextract_metadata(
appeal_subtype, and case_name when it equals the citation). User appeal_subtype, and case_name when it equals the citation). User
values are preserved. values are preserved.
**MCP-tool-only path** — same constraint as :func:`reextract_halachot`. **Container-safe** — unlike :func:`reextract_halachot` (claude CLI, host-only),
metadata extraction runs on Gemini Flash over REST (GOOGLE_GEMINI_API_KEY), so
this path is callable from the FastAPI container too. The final-decision
enrollment loop (``_enroll_final_in_library``) calls it inline on upload.
""" """
from legal_mcp.services import precedent_metadata_extractor from legal_mcp.services import precedent_metadata_extractor
@@ -411,6 +437,37 @@ async def get_precedent(case_law_id: UUID | str) -> dict | None:
return None return None
record["halachot"] = await db.list_halachot(case_law_id=case_law_id, limit=500) record["halachot"] = await db.list_halachot(case_law_id=case_law_id, limit=500)
record["related_cases"] = await db.get_case_law_relations(case_law_id) record["related_cases"] = await db.get_case_law_relations(case_law_id)
# Auto-detected citation graph: which decisions cite THIS ruling (incoming
# edges in precedent_internal_citations). Reuses the canonical query (G2) —
# no parallel resolution path. Shaped to match the front-end RelatedCase row
# so the existing "ציטוטים מקושרים" panel renders them with no visual change.
from legal_mcp.services import citation_extractor as _ce
raw = await _ce.list_citations_to_case_law(case_law_id)
record["incoming_citations"] = [
{
"id": r["source_case_law_id"],
"case_number": r.get("source_case_number") or r.get("cited_case_number") or "",
"case_name": r.get("source_case_name") or "",
"court": r.get("source_court") or "",
"precedent_level": r.get("source_precedent_level") or "",
"chair_name": r.get("source_chair_name") or "",
"date": (
r["source_date"].isoformat()
if r.get("source_date") is not None else None
),
"treatment": r.get("treatment") or "",
"confidence": r.get("confidence"),
}
for r in raw
]
# Authority signal (X11 Phase 2, #154): how the citing committee decisions
# TREATED this ruling (followed/distinguished/…) — surfaced so the chair (and
# research agents via the tool output) can argue authority and avoid leaning on
# a repeatedly-distinguished precedent.
record["cited_by"] = (await db.citation_authority([case_law_id])).get(
str(case_law_id),
{"total": 0, "positive": 0, "negative": 0, "unclassified": 0, "by_treatment": {}},
)
return record return record

View File

@@ -260,6 +260,7 @@ async def apply_to_record(
case_law_id: UUID | str, case_law_id: UUID | str,
suggested: dict, suggested: dict,
overwrite_case_number: bool = False, overwrite_case_number: bool = False,
force_citation: bool = False,
) -> dict: ) -> dict:
"""Merge suggested metadata into the case_law row, filling ONLY empty fields. """Merge suggested metadata into the case_law row, filling ONLY empty fields.
@@ -274,6 +275,13 @@ async def apply_to_record(
overwrite_case_number: when True, update case_number from case_number_clean overwrite_case_number: when True, update case_number from case_number_clean
even if the field already has a value (used for one-time migration enrichment). even if the field already has a value (used for one-time migration enrichment).
force_citation: when True, (re)assemble citation_formatted even if the field
is non-empty — used by the at-upload inline enrichment to UPGRADE the seeded
provisional citation (the raw chair input) to the canonical derived form. The
write still happens only when the deterministic assembly SUCCEEDS (a missing
component → no write → the seed is preserved). The drainer keeps the default
(False) so a chair's manual edit in /precedents/[id] is never clobbered.
""" """
if isinstance(case_law_id, str): if isinstance(case_law_id, str):
case_law_id = UUID(case_law_id) case_law_id = UUID(case_law_id)
@@ -482,7 +490,7 @@ async def apply_to_record(
# source_type/district/proceeding_type/parties). Only fill when empty so chair # source_type/district/proceeding_type/parties). Only fill when empty so chair
# edits in /precedents/[id] are preserved; abstains (no write) when a component # edits in /precedents/[id] are preserved; abstains (no write) when a component
# is missing. # is missing.
if not (record.get("citation_formatted") or "").strip(): if force_citation or not (record.get("citation_formatted") or "").strip():
eff = {**record, **fields_to_update} eff = {**record, **fields_to_update}
eff_parties = ( eff_parties = (
fields_to_update.get("parties") or record.get("parties") or "" fields_to_update.get("parties") or record.get("parties") or ""
@@ -505,6 +513,7 @@ async def apply_to_record(
async def extract_and_apply( async def extract_and_apply(
case_law_id: UUID | str, case_law_id: UUID | str,
overwrite_case_number: bool = False, overwrite_case_number: bool = False,
force_citation: bool = False,
) -> dict: ) -> dict:
"""Convenience wrapper: extract → merge into row → return summary.""" """Convenience wrapper: extract → merge into row → return summary."""
suggested = await extract_metadata(case_law_id) suggested = await extract_metadata(case_law_id)
@@ -523,7 +532,11 @@ async def extract_and_apply(
"status": "extraction_failed" if has_text else "no_metadata", "status": "extraction_failed" if has_text else "no_metadata",
"fields": [], "fields": [],
} }
result = await apply_to_record(case_law_id, suggested, overwrite_case_number=overwrite_case_number) result = await apply_to_record(
case_law_id, suggested,
overwrite_case_number=overwrite_case_number,
force_citation=force_citation,
)
if result["updated"]: if result["updated"]:
await db.recompute_searchable(case_law_id) await db.recompute_searchable(case_law_id)
return { return {

View File

@@ -0,0 +1,83 @@
"""Block-level style-exemplar extraction (channel B of Style Acquisition).
Splits one of Dafna's decisions into section→paragraph units, embeds them
(Voyage), and stores them in `style_exemplars` so the writer can retrieve real
block-level prose by section/outcome/practice_area (07-learning §0.2 channel B).
This is the SINGLE source of truth for exemplar extraction (G2): both the
one-time backfill (`scripts/backfill_style_exemplars.py`) and the live
final-enrollment path (`_enroll_final_in_library`) call `extract_and_store`.
Before this, exemplars were frozen at the seed backfill — new finals enrolled
into style_corpus but were never broken into exemplars, so the richest style
channel never grew. Embedding is Voyage-over-REST → container-safe.
"""
from __future__ import annotations
import logging
from legal_mcp.services import db, embeddings
from legal_mcp.services.chunker import _split_into_sections
logger = logging.getLogger(__name__)
# chunker section_type → style_exemplars.section
_SECTION_MAP = {
"facts": "background",
"appellant_claims": "claims",
"respondent_claims": "claims",
"legal_analysis": "discussion",
"conclusion": "summary",
"ruling": "summary",
"intro": "other",
"other": "other",
}
MIN_WORDS = 25 # skip tiny fragments
MAX_WORDS = 450 # skip over-long blobs (likely un-split)
MAX_PER_SECTION = 15
def _paragraphs(section_text: str) -> list[str]:
"""Split a section into paragraph units (blank-line separated; fall back to lines)."""
raw = [p.strip() for p in section_text.split("\n\n")]
if len(raw) <= 1:
raw = [p.strip() for p in section_text.split("\n")]
out = []
for p in raw:
wc = len(p.split())
if MIN_WORDS <= wc <= MAX_WORDS:
out.append(p)
return out[:MAX_PER_SECTION]
def units_for(full_text: str) -> list[tuple[str, str]]:
"""(section, paragraph) units for a decision — pure, no I/O."""
units: list[tuple[str, str]] = []
for section_type, section_text in _split_into_sections(full_text or ""):
section = _SECTION_MAP.get(section_type, "other")
for para in _paragraphs(section_text):
units.append((section, para))
return units
async def extract_and_store(
decision_number: str, source: str, full_text: str,
practice_area: str = "", outcome: str = "",
) -> int:
"""Idempotently (re)build a decision's block-level style exemplars: split →
embed → replace. Returns the number of exemplars stored. Raises on failure;
callers in best-effort paths (enroll) should wrap in try/except."""
units = units_for(full_text)
if not units:
return 0
texts = [u[1] for u in units]
vecs = await embeddings.embed_texts(texts, input_type="document")
await db.delete_style_exemplars(decision_number, source)
for (section, para), vec in zip(units, vecs):
await db.insert_style_exemplar(
decision_number=decision_number, source=source,
practice_area=practice_area, outcome=outcome,
section=section, paragraph_text=para, word_count=len(para.split()),
embedding=vec,
)
return len(units)

View File

@@ -170,3 +170,30 @@ async def digest_process_pending(limit: int = 20) -> str:
except Exception as e: except Exception as e:
return _err(str(e)) return _err(str(e))
return _ok(result) return _ok(result)
async def digest_radar(case_number: str, limit: int = 5, min_score: float = 0.45) -> str:
"""רדאר-יומונים הקשרי-לתיק (X12) — "שים לב" ליו"ר.
מחזיר יומונים **לא-מקושרים** (פס"ד שעוד אין לנו בקורפוס) שהנושא שלהם קרוב
סמנטית לתיק הזה — כדי שפס"ד רלוונטי שמוכר רק דרך יומון לא ייפול בין הכיסאות
בזמן הכרעת התיק. נושא-התיק נבנה מ-title + appeal_subtype + טענות-הצדדים, ומותאם
מול יומונים לא-מקושרים. כל ליד נושא את מראה-המקום של הפס"ד המקורי, סטטוס-הפער
(new_lead / gap_open / fetched / available_link) ופעולה מוצעת.
INV-DIG1: זהו radar — היומון לעולם אינו מצוטט; הליד מצביע על **הפס"ד** (להזמין
משיכה / להעלות / לקשר), לא על היומון. read-only.
Args:
case_number: מספר התיק (למשל "8124-09-24").
limit: מספר לידים מקסימלי.
min_score: סף-דמיון סמנטי (0-1) לסינון רעש.
"""
if not case_number.strip():
return _err("case_number חובה")
try:
result = await digest_library.case_digest_radar(
case_number.strip(), limit=max(1, int(limit)), min_score=float(min_score))
except Exception as e:
return _err(str(e))
return _ok(result)

View File

@@ -300,6 +300,20 @@ async def search_precedent_library(
limit=limit, limit=limit,
include_halachot=include_halachot, include_halachot=include_halachot,
) )
# X11 Phase 2 (#154): attach the incoming-citation authority breakdown so the
# research agent can WEIGH and ARGUE authority ("הלכה שאומצה ב-N החלטות ועדת-ערר")
# — and steer clear of a precedent that committees repeatedly distinguished /
# overruled. Batched: one query for the whole result page.
try:
ids = {str(r.get("case_law_id")) for r in results if r.get("case_law_id")}
if ids:
auth = await db.citation_authority([UUID(i) for i in ids])
for r in results:
cb = auth.get(str(r.get("case_law_id")))
if cb:
r["cited_by"] = cb
except Exception: # noqa: BLE001 — authority is an additive signal; never break search
pass
elapsed_ms = int((time.perf_counter() - t0) * 1000) elapsed_ms = int((time.perf_counter() - t0) * 1000)
telemetry.log_search_bg( telemetry.log_search_bg(
search_type="precedent_library", search_type="precedent_library",

View File

@@ -394,13 +394,107 @@ async def record_chair_feedback(
lesson_extracted=lesson_extracted, lesson_extracted=lesson_extracted,
) )
# Auto-flow chair-authored STYLE feedback to the writer (closes the dead
# chair_feedback→lesson chain — 27 feedback rows had produced 0 lessons). The
# chair is the highest authority, so a style correction she writes flows
# immediately — the chair IS the gate. It rides the SAME discussion_rules channel
# promote uses (db.append_global_rule, G2), reaching every block, without the
# style_corpus coupling decision_lessons require. SUBSTANCE feedback
# (missing_content/factual_error/other) is case-specific → recorded only.
# (INV-LRN1 graduated gate; 07-learning §1.2.)
_STYLE_FB = {"style", "wrong_tone", "wrong_structure"}
flowed = 0
if lesson_extracted.strip() and category in _STYLE_FB:
try:
flowed = await db.append_global_rule(
"discussion_rules", "universal", [lesson_extracted.strip()],
)
except Exception as e:
logger.warning("chair-feedback auto-flow failed for %s: %s", case_number, e)
msg = f"הערה נרשמה בהצלחה. קטגוריה: {category}."
if flowed:
msg += " הלקח (סגנון) זרם אוטומטית לכותב."
return ok({ return ok({
"feedback_id": str(feedback_id), "feedback_id": str(feedback_id),
"flowed_to_writer": bool(flowed),
"next_steps": [ "next_steps": [
"כדי להפיק לקח מההערה, הפעל: analyze_chair_feedback", "כדי להפיק לקח מההערה, הפעל: analyze_chair_feedback",
"כדי לסמן כמטופל: resolve_chair_feedback", "כדי לסמן כמטופל: resolve_chair_feedback",
], ],
}, message=f"הערה נרשמה בהצלחה. קטגוריה: {category}.") }, message=msg)
_CURATOR_FINDING_CATEGORIES = {"style", "structure", "lexicon", "tabular", "general"}
# Agent tags ([סגנון]/[מבנה]/[לקסיקון משפטי]/[טבלאי]) → decision_lessons.category
_CURATOR_TAG_TO_CATEGORY = {
"סגנון": "style", "מבנה": "structure",
"לקסיקון משפטי": "lexicon", "לקסיקון": "lexicon", "טבלאי": "tabular",
}
async def record_curator_findings(case_number: str, findings: list[dict]) -> str:
"""לכידת ממצאי-האוצֵר כ-decision_lessons מובְנים (source='curator', proposed) — INV-LRN3.
האוצֵר מזהה דפוסי-סגנון בקריאת הסופי; עד כה הם חיו רק כהערה ארעית בערוץ-התגובות
(אובד, לא נסקר). כאן הם נתפסים מבנית כך שיופיעו בטאב ״אוצֵר״ ויעברו שער-יו"ר (INV-LRN1/G10).
האוצֵר נשאר read-only על התוכן — הרישום הוא הצעה הממתינה לאישור, לא שינוי-קול.
Args:
case_number: מספר התיק הסופי (= decision_number בקורפוס-הסגנון).
findings: רשימת ממצאים, כל אחד {"text": "...", "category"/"tag": "..."}.
category ∈ style/structure/lexicon/tabular/general (או tag עברי).
"""
if not findings:
return err("findings ריק — אין ממצאים לרשום.")
corpus_id = await db.get_style_corpus_id_by_decision(case_number)
if not corpus_id:
return err(
f"לא נמצאה רשומת style_corpus ל-{case_number} — ודא שהסופי נקלט לקורפוס-הסגנון "
"(enroll_style_corpus) לפני רישום ממצאים."
)
# Dedup against lessons already on this corpus (any source) — re-running §A
# must not pile duplicates (INV-LRN3 reliability).
existing = {(_norm(r["lesson_text"])) for r in await db.list_decision_lessons(corpus_id)}
written, skipped_dup, skipped_empty = [], 0, 0
for f in findings:
text = (f.get("text") or "").strip()
if not text:
skipped_empty += 1
continue
if _norm(text) in existing:
skipped_dup += 1
continue
raw_cat = (f.get("category") or f.get("tag") or "general").strip()
category = _CURATOR_TAG_TO_CATEGORY.get(raw_cat, raw_cat)
if category not in _CURATOR_FINDING_CATEGORIES:
category = "general"
row = await db.add_decision_lesson(
corpus_id,
lesson_text=text,
category=category,
source="curator",
created_by="curator",
review_status="proposed",
)
if row:
written.append(str(row["id"]))
existing.add(_norm(text))
return ok({
"corpus_id": str(corpus_id),
"written": len(written),
"skipped_duplicate": skipped_dup,
"skipped_empty": skipped_empty,
"lesson_ids": written,
}, message=(
f"נרשמו {len(written)} ממצאי-אוצֵר (source=curator, ממתינים לשער-יו\"ר ב-/training). "
f"{skipped_dup} כפילויות דולגו."
))
def _norm(s: str) -> str:
"""Normalize lesson text for dedup — collapse whitespace, strip."""
return " ".join((s or "").split())
async def list_chair_feedback( async def list_chair_feedback(

View File

@@ -0,0 +1,21 @@
"""Adapter-profile compatibility gates for Paperclip migration."""
from __future__ import annotations
import importlib.util
from pathlib import Path
_SCRIPT = Path(__file__).resolve().parents[2] / "scripts" / "adapter_profiles.py"
_spec = importlib.util.spec_from_file_location("adapter_profiles", _SCRIPT)
profiles = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(profiles)
def test_codex_local_profile_accepts_openai_model_ids():
assert profiles.model_matches_provider("gpt-5.3-codex", "codex_local")
assert profiles.model_matches_provider("o4-mini", "codex_local")
assert profiles.model_matches_provider("codex-mini-latest", "codex_local")
def test_codex_local_profile_rejects_foreign_model_ids():
assert not profiles.model_matches_provider("claude-opus-4-8", "codex_local")
assert not profiles.model_matches_provider("gemini-3.1-pro-preview", "codex_local")

View File

@@ -15,6 +15,7 @@
| `pc.sh` | bash | **wrapper לכל קריאות Paperclip API מסוכנים** — מוסיף Authorization, X-Paperclip-Run-Id (audit trail), Content-Type, base URL. תחביר: `pc.sh <METHOD> <PATH> [BODY_JSON]`. אסור `curl` ישיר ל-`$PAPERCLIP_API_URL`. ראה `HEARTBEAT.md §0`. counterpart ב-Python: `web/paperclip_api.py`. | נקרא ע"י סוכנים | | `pc.sh` | bash | **wrapper לכל קריאות Paperclip API מסוכנים** — מוסיף Authorization, X-Paperclip-Run-Id (audit trail), Content-Type, base URL. תחביר: `pc.sh <METHOD> <PATH> [BODY_JSON]`. אסור `curl` ישיר ל-`$PAPERCLIP_API_URL`. ראה `HEARTBEAT.md §0`. counterpart ב-Python: `web/paperclip_api.py`. | נקרא ע"י סוכנים |
| `sync_agents_across_companies.py` | python | **סנכרון סוכנים מ-CMP (1xxx, master) ל-CMPA (8xxx, mirror)** — Gap #25. משווה adapter_config (model/timeout/instructions/skills/etc), runtime_config (heartbeat), ושדות top-level (budget/metadata/icon/title/role). מסנן אוטומטית local skills שלא קיימים ב-mirror. לוגיקת subset (mirror יכול להחזיק יותר skills כי ה-API מוסיף required runtime skills). תומך `--verify`/`--dry-run`/`--apply [--only NAME]`. גיבוי אוטומטי. דורש `PAPERCLIP_BOARD_API_KEY`. **להריץ אחרי כל שינוי הגדרות ב-CMP.** **⚠ אם `adapter_type` שונה בין CMP ל-CMPA — `--apply` מדלג על הסוכן; `--verify` מדווח אותו רם כ-DRIFT.** בעת מעבר adapter (למשל ל-`deepseek_local`) חובה לעדכן ידנית בשתי החברות. **`--verify` יוצא exit≠0 על כל drift** (needs-sync / adapter-mismatch / missing-in-mirror) — שמיש כ-gate ל-cron/CI (GAP-21/FU-8a). | ידני אחרי כל שינוי | | `sync_agents_across_companies.py` | python | **סנכרון סוכנים מ-CMP (1xxx, master) ל-CMPA (8xxx, mirror)** — Gap #25. משווה adapter_config (model/timeout/instructions/skills/etc), runtime_config (heartbeat), ושדות top-level (budget/metadata/icon/title/role). מסנן אוטומטית local skills שלא קיימים ב-mirror. לוגיקת subset (mirror יכול להחזיק יותר skills כי ה-API מוסיף required runtime skills). תומך `--verify`/`--dry-run`/`--apply [--only NAME]`. גיבוי אוטומטי. דורש `PAPERCLIP_BOARD_API_KEY`. **להריץ אחרי כל שינוי הגדרות ב-CMP.** **⚠ אם `adapter_type` שונה בין CMP ל-CMPA — `--apply` מדלג על הסוכן; `--verify` מדווח אותו רם כ-DRIFT.** בעת מעבר adapter (למשל ל-`deepseek_local`) חובה לעדכן ידנית בשתי החברות. **`--verify` יוצא exit≠0 על כל drift** (needs-sync / adapter-mismatch / missing-in-mirror) — שמיש כ-gate ל-cron/CI (GAP-21/FU-8a). | ידני אחרי כל שינוי |
| `fix_paperclipai_skills_drift.py` | python | סקריפט חד-פעמי (בוצע 2026-05-04) שניקה drift על `paperclipai/*` skills בין CMP ל-CMPA. הסיר `paperclip-dev` מכל 14 הסוכנים, ודאג ש-`paperclip-converting-plans-to-tasks` קיים רק על CEO ו-analyst. תומך `--apply` (ברירת מחדל: dry-run). דורש `PAPERCLIP_BOARD_API_KEY`. נשמר לרפרנס למקרה שhdrift חוזר. | חד-פעמי (בוצע) | | `fix_paperclipai_skills_drift.py` | python | סקריפט חד-פעמי (בוצע 2026-05-04) שניקה drift על `paperclipai/*` skills בין CMP ל-CMPA. הסיר `paperclip-dev` מכל 14 הסוכנים, ודאג ש-`paperclip-converting-plans-to-tasks` קיים רק על CEO ו-analyst. תומך `--apply` (ברירת מחדל: dry-run). דורש `PAPERCLIP_BOARD_API_KEY`. נשמר לרפרנס למקרה שhdrift חוזר. | חד-פעמי (בוצע) |
| `classify_citation_treatments.py` | python | **סיווג-טיפול לקצוות-ציטוט (X11 Phase 2, #154)** — לכל קצה ב-`precedent_internal_citations` (החלטת-ועדה מצטטת תקדים) מסווג את ה-`treatment` מתוך ה-`match_context` דרך `corroboration.classify_treatment` (Opus 4.8 @ xhigh, claude_session **מקומי** — לא בקונטיינר): followed/explained=חיובי, distinguished/criticized/questioned/overruled=שלילי. ממלא `precedent_internal_citations.treatment` כך ש-`refresh_verified_layer` לא יספור ציטוט שלילי כסמכות (INV-COR2) ו-`db.citation_authority` יציג פירוק לסוכנים. אידמפוטנטי (מדלג על מסווגים). `--apply`/`--limit N`/`--case-law-id UUID`. **אחרי `--apply` הרץ `build_verified_layer.py`.** דורש `HOME=/home/chaim`. | ידני / אחרי גלי-ציטוט חדשים |
| `adapter_profiles.py` | python (module) | **רישום-פרופילי-אדפטר** — מקור-אמת יחיד ל-3 צירי-הכשל של מעבר-אדפטר: provider/default_model, instructions_mode (`file_path` בטוח-frontmatter מול `content_arg` ששובר `---`), ו-tool_config (`gemini_global` excludeTools / `frontmatter` / `hermes` / `codex_home`). כולל `codex_local` עם משפחת מודלי OpenAI/Codex (`gpt-*`, `o3*`, `o4*`, `codex-*`). מיובא ע"י `migrate_agent_adapter.py`. הוספת אדפטר עתידי = רשומה אחת. לא מורץ ישירות. | תשתית | | `adapter_profiles.py` | python (module) | **רישום-פרופילי-אדפטר** — מקור-אמת יחיד ל-3 צירי-הכשל של מעבר-אדפטר: provider/default_model, instructions_mode (`file_path` בטוח-frontmatter מול `content_arg` ששובר `---`), ו-tool_config (`gemini_global` excludeTools / `frontmatter` / `hermes` / `codex_home`). כולל `codex_local` עם משפחת מודלי OpenAI/Codex (`gpt-*`, `o3*`, `o4*`, `codex-*`). מיובא ע"י `migrate_agent_adapter.py`. הוספת אדפטר עתידי = רשומה אחת. לא מורץ ישירות. | תשתית |
| `migrate_agent_adapter.py` | python | **מעבר-אדפטר בטוח לכל סוכן ← כל אדפטר, בשתי החברות יחד (INV-MC1)**. מיישב model↔provider, גורס frontmatter לעותק `.generated/<name>.nofm.md` ל-content_arg adapters (אחרת קריסת `gemini --prompt`/`hermes -q` על `---`), ומשחרר excludeTools גלובלי של gemini (`--relax-tools`). `--check` (preflight בלבד, exit≠0 על שגיאה — שער FU-8a) / `--apply` / `--revert` (שחזור מדויק מ-sidecar `data/adapter-migration-state.json`) / `--verify` (מסמן מצב לא-תואם/א-סימטרי, exit≠0). `--agent "<שם>"\|all --to <adapter> [--model X] [--relax-tools]`. PATCH דרך `/api/agents/{id}` (לא DB). דורש `PAPERCLIP_BOARD_API_KEY`. הרץ עם `mcp-server/.venv/bin/python`. **fallback-חירום כשנגמרים טוקני-Claude; החזר ל-claude_local כשחוזרים.** | ידני לפי צורך | | `migrate_agent_adapter.py` | python | **מעבר-אדפטר בטוח לכל סוכן ← כל אדפטר, בשתי החברות יחד (INV-MC1)**. מיישב model↔provider, גורס frontmatter לעותק `.generated/<name>.nofm.md` ל-content_arg adapters (אחרת קריסת `gemini --prompt`/`hermes -q` על `---`), ומשחרר excludeTools גלובלי של gemini (`--relax-tools`). `--check` (preflight בלבד, exit≠0 על שגיאה — שער FU-8a) / `--apply` / `--revert` (שחזור מדויק מ-sidecar `data/adapter-migration-state.json`) / `--verify` (מסמן מצב לא-תואם/א-סימטרי, exit≠0). `--agent "<שם>"\|all --to <adapter> [--model X] [--relax-tools]`. PATCH דרך `/api/agents/{id}` (לא DB). דורש `PAPERCLIP_BOARD_API_KEY`. הרץ עם `mcp-server/.venv/bin/python`. **fallback-חירום כשנגמרים טוקני-Claude; החזר ל-claude_local כשחוזרים.** | ידני לפי צורך |
@@ -152,6 +153,7 @@
| Script | Type | Purpose | Scheduled | | Script | Type | Purpose | Scheduled |
|--------|------|---------|-----------| |--------|------|---------|-----------|
| `ocr_benchmark_mistral.py` | python | **בנצ'מרק OCR — Mistral מול Google Vision** (מחקר חד-פעמי, הוביל למעבר ל-Mistral). מוריד מסמכים מ-MinIO, קורא טקסט קיים מה-DB, שולח ל-Mistral OCR, מחשב מטריקות (כיסוי/ניקיון/עברית%), שומר דוח ל-`data/audit/ocr-benchmark-mistral.md` + טקסטים גולמיים ל-`data/audit/ocr-benchmark-raw/`. הרצה: `mcp-server/.venv/bin/python scripts/ocr_benchmark_mistral.py`. דורש `MISTRAL_API_KEY` ו-mcli alias `legalminio`. **ממצא:** Mistral מנצח ב-4/5 תיקים בדוגמה ומטפל נכון ב-OCR שבור (1044-03-26 שהחזיר "English garbage" ב-Vision). | חד-פעמי (בוצע 2026-06-27) |
| `backfill_missing_precedents.py` | python | **הזנת `missing_precedents` פתוחים לתור-האחזור (X13)** — מסווג כל פער-פתוח; עליון-סדרתי→Tier-0(supremedecisions), נט-format→Tier-1; ועדת-ערר/לא-מזוהה→דילוג. יוצר `court_fetch_jobs` (idempotent). `--apply` (ברירת-מחדל dry-run). אחרי הרצה: drain-court-fetch קולט. | ידני (חד-פעמי/לפי-צורך) | | `backfill_missing_precedents.py` | python | **הזנת `missing_precedents` פתוחים לתור-האחזור (X13)** — מסווג כל פער-פתוח; עליון-סדרתי→Tier-0(supremedecisions), נט-format→Tier-1; ועדת-ערר/לא-מזוהה→דילוג. יוצר `court_fetch_jobs` (idempotent). `--apply` (ברירת-מחדל dry-run). אחרי הרצה: drain-court-fetch קולט. | ידני (חד-פעמי/לפי-צורך) |
| `derive_missing_from_cited_only.py` | python | **#143 — איחוד cited_only↔missing_precedents (G2)**: גוזר רשומת `missing_precedents` 'open' לכל stub `cited_only` (פסיקה מצוטטת ללא טקסט), כך ש-31 ה-stubs מופיעים בדף "פסיקה חסרה" (היו היו חפיפה≈0). (1) backfill `citation_norm` (מפתח-dedup designator-aware — `court_citation.citation_dedup_key`) ל-291 הקיימים; (2) לכל stub → `create_missing_precedent(discovery_source='cited_only', linked_case_law_id=stub, notes=מצטטים)` עם dedup. `linked_case_law_id`=זהות-קנונית-ידועה, `status='open'` עד העלאת-טקסט (→ promote-in-place דרך ON CONFLICT). אידמפוטנטי, dry-run / `--apply`. הרצה: `HOME=/home/chaim mcp-server/.venv/bin/python scripts/derive_missing_from_cited_only.py --apply`. | חד-פעמי / re-runnable | | `derive_missing_from_cited_only.py` | python | **#143 — איחוד cited_only↔missing_precedents (G2)**: גוזר רשומת `missing_precedents` 'open' לכל stub `cited_only` (פסיקה מצוטטת ללא טקסט), כך ש-31 ה-stubs מופיעים בדף "פסיקה חסרה" (היו היו חפיפה≈0). (1) backfill `citation_norm` (מפתח-dedup designator-aware — `court_citation.citation_dedup_key`) ל-291 הקיימים; (2) לכל stub → `create_missing_precedent(discovery_source='cited_only', linked_case_law_id=stub, notes=מצטטים)` עם dedup. `linked_case_law_id`=זהות-קנונית-ידועה, `status='open'` עד העלאת-טקסט (→ promote-in-place דרך ON CONFLICT). אידמפוטנטי, dry-run / `--apply`. הרצה: `HOME=/home/chaim mcp-server/.venv/bin/python scripts/derive_missing_from_cited_only.py --apply`. | חד-פעמי / re-runnable |
| `backfill_digest_missing_precedents.py` | python | **#136 — חיבור יומונים-לא-מקושרים ל"פסיקה חסרה"**: לכל digest עם `underlying_citation` ו-`linked_case_law_id IS NULL` (461) מריץ את `digest_library.try_autolink` הקנוני (G2) — מקשר אם אפשר, אחרת פותח gap: ערר/בל"מ/unknown → `missing_precedent` (discovery_source='digest', dedup designator-aware), פס"ד בתי-משפט → `court_fetch_job` (X13). dry-run מציג פילוח-tier (369 ערר + 21 unknown → MP; 71 fetchable → court_fetch). אידמפוטנטי. הרצה: `HOME=/home/chaim mcp-server/.venv/bin/python scripts/backfill_digest_missing_precedents.py --apply`. | חד-פעמי / re-runnable | | `backfill_digest_missing_precedents.py` | python | **#136 — חיבור יומונים-לא-מקושרים ל"פסיקה חסרה"**: לכל digest עם `underlying_citation` ו-`linked_case_law_id IS NULL` (461) מריץ את `digest_library.try_autolink` הקנוני (G2) — מקשר אם אפשר, אחרת פותח gap: ערר/בל"מ/unknown → `missing_precedent` (discovery_source='digest', dedup designator-aware), פס"ד בתי-משפט → `court_fetch_job` (X13). dry-run מציג פילוח-tier (369 ערר + 21 unknown → MP; 71 fetchable → court_fetch). אידמפוטנטי. הרצה: `HOME=/home/chaim mcp-server/.venv/bin/python scripts/backfill_digest_missing_precedents.py --apply`. | חד-פעמי / re-runnable |
@@ -198,7 +200,7 @@
| `drain_digests.py` | python | ריקון תור ההעשרה של יומונים (X12): מעבד כל digest בסטטוס `pending` דרך `digest_library.enrich_digest` (חילוץ-LLM Sonnet + embedding + autolink). מקבילי (CONCURRENCY=3, env-tunable), idempotent. מוסיף `~/.local/bin` ל-PATH כדי שה-claude CLI יימצא תחת cron. בודק דגל `drain_controls('legal-digest-drain')` ב-startup → no-op כשכבוי מ-/operations. | דרך `legal-digest-drain.config.cjs` (pm2 cron) + ידני אחרי backfill. חלופת-MCP: `digest_process_pending` | | `drain_digests.py` | python | ריקון תור ההעשרה של יומונים (X12): מעבד כל digest בסטטוס `pending` דרך `digest_library.enrich_digest` (חילוץ-LLM Sonnet + embedding + autolink). מקבילי (CONCURRENCY=3, env-tunable), idempotent. מוסיף `~/.local/bin` ל-PATH כדי שה-claude CLI יימצא תחת cron. בודק דגל `drain_controls('legal-digest-drain')` ב-startup → no-op כשכבוי מ-/operations. | דרך `legal-digest-drain.config.cjs` (pm2 cron) + ידני אחרי backfill. חלופת-MCP: `digest_process_pending` |
| `legal-digest-drain.config.cjs` | pm2/js | **תזמון כל שעתיים של `drain_digests.py`** (cron `12 */2 * * *`, `DIGEST_DRAIN_CRON` לעקיפה; דקת-הצתה `:12` כדי לא לחלוק דקה עם metadata-drain `:00` — מונע deadlock של DDL-המיגרציה) — הועבר מ-crontab של המערכת ל-pm2 כדי שיופיע ויהיה שליט בדף `/operations` (הרץ-עכשיו/הפעל/כבה). `autorestart:false` (one-shot per tick). דורש claude CLI + `VOYAGE_API_KEY`. התקנה: `pm2 start scripts/legal-digest-drain.config.cjs && pm2 save`. | pm2 cron (host-side) | | `legal-digest-drain.config.cjs` | pm2/js | **תזמון כל שעתיים של `drain_digests.py`** (cron `12 */2 * * *`, `DIGEST_DRAIN_CRON` לעקיפה; דקת-הצתה `:12` כדי לא לחלוק דקה עם metadata-drain `:00` — מונע deadlock של DDL-המיגרציה) — הועבר מ-crontab של המערכת ל-pm2 כדי שיופיע ויהיה שליט בדף `/operations` (הרץ-עכשיו/הפעל/כבה). `autorestart:false` (one-shot per tick). דורש claude CLI + `VOYAGE_API_KEY`. התקנה: `pm2 start scripts/legal-digest-drain.config.cjs && pm2 save`. | pm2 cron (host-side) |
| `renumber_cases.py` | python | **מיגרציה חד-פעמית (בוצעה 2026-06-12)** — תיקון 11 מספרי-תיקים לפורמט קנוני `NNNN-MM-YY` (הוספת ספרות-חודש; 1046-26→1024-02-26 תיקון-סידורי). רץ על ה-host (לא בקונטיינר): DB pool של האפליקציה + `mcli` (MinIO) + Gitea API + Paperclip DB. אטומי per-case עם גיבוי ל-`data/audit/` ואימות-אחרי. FK-ים על `cases.id` (UUID) לא נגעו; משכתב כל עמודה עם `cases/{old}/` (file_path **וגם** image_thumbnail_path שהוא storage-key בלי `/data`), מנרמל זהות חוצת-קורפוס (case_law/style_corpus/style_exemplars/citations — לא תוכן/full_text), מעביר מפתחי-MinIO ב-3 buckets (legal-immutable=WORM copy-only), משנה-שם repo ב-Gitea, ומעדכן שם-פרויקט ב-Paperclip. dry-run כברירת-מחדל; `--apply --tier clean\|archive`. **מיצוי — לא להריץ שוב** (ה-MAPPING היסטורי). | חד-פעמי — בוצע | | `renumber_cases.py` | python | **מיגרציה חד-פעמית (בוצעה 2026-06-12)** — תיקון 11 מספרי-תיקים לפורמט קנוני `NNNN-MM-YY` (הוספת ספרות-חודש; 1046-26→1024-02-26 תיקון-סידורי). רץ על ה-host (לא בקונטיינר): DB pool של האפליקציה + `mcli` (MinIO) + Gitea API + Paperclip DB. אטומי per-case עם גיבוי ל-`data/audit/` ואימות-אחרי. FK-ים על `cases.id` (UUID) לא נגעו; משכתב כל עמודה עם `cases/{old}/` (file_path **וגם** image_thumbnail_path שהוא storage-key בלי `/data`), מנרמל זהות חוצת-קורפוס (case_law/style_corpus/style_exemplars/citations — לא תוכן/full_text), מעביר מפתחי-MinIO ב-3 buckets (legal-immutable=WORM copy-only), משנה-שם repo ב-Gitea, ומעדכן ב-Paperclip את **שם-הפרויקט + `plugin_state.legal-case-number` + כותרות-issues** (3 המשטחים ש-`get_case_issues` נשען עליהם — בלי (b)+(c) ה-issues נשארים על המספר הישן ו-run-learning/run-halacha מדלגים בשקט עם "no_issue"; תוקן 2026-06-28 אחרי שהפער התגלה ב-8137-11-24). dry-run כברירת-מחדל; `--apply --tier clean\|archive`. **מיצוי — לא להריץ שוב** (ה-MAPPING היסטורי); לתבנית-עתידית: ה-Paperclip-step המעודכן הוא הרפרנס. | חד-פעמי — בוצע |
## סקריפטים שנמחקו (git history בלבד) ## סקריפטים שנמחקו (git history בלבד)

View File

@@ -19,40 +19,15 @@ import argparse
import asyncio import asyncio
import logging import logging
from legal_mcp.services import db, embeddings from legal_mcp.services import db
from legal_mcp.services.chunker import _split_into_sections from legal_mcp.services import style_exemplars as sx
logging.basicConfig(level=logging.INFO, format="%(message)s") logging.basicConfig(level=logging.INFO, format="%(message)s")
log = logging.getLogger("backfill_exemplars") log = logging.getLogger("backfill_exemplars")
# chunker section_type → style_exemplars.section # Section mapping + paragraph splitting now live in the shared service
_SECTION_MAP = { # (legal_mcp.services.style_exemplars) so the backfill and the live
"facts": "background", # final-enrollment path use ONE extraction implementation (G2).
"appellant_claims": "claims",
"respondent_claims": "claims",
"legal_analysis": "discussion",
"conclusion": "summary",
"ruling": "summary",
"intro": "other",
"other": "other",
}
MIN_WORDS = 25 # skip tiny fragments
MAX_WORDS = 450 # skip over-long blobs (likely un-split)
MAX_PER_SECTION = 15
def _paragraphs(section_text: str) -> list[str]:
"""Split a section into paragraph units (blank-line separated; fall back to lines)."""
raw = [p.strip() for p in section_text.split("\n\n")]
if len(raw) <= 1:
raw = [p.strip() for p in section_text.split("\n")]
out = []
for p in raw:
wc = len(p.split())
if MIN_WORDS <= wc <= MAX_WORDS:
out.append(p)
return out[:MAX_PER_SECTION]
async def _gather_sources() -> list[dict]: async def _gather_sources() -> list[dict]:
@@ -94,28 +69,18 @@ async def main(apply: bool) -> None:
total_paras = 0 total_paras = 0
for src in sources: for src in sources:
units: list[tuple[str, str]] = [] # (section, paragraph) n = len(sx.units_for(src["full_text"]))
for section_type, section_text in _split_into_sections(src["full_text"]): if not n:
section = _SECTION_MAP.get(section_type, "other")
for para in _paragraphs(section_text):
units.append((section, para))
if not units:
continue continue
total_paras += len(units) total_paras += n
log.info(" %-14s %-16s%d פסקאות", src["source"], src["decision_number"], len(units)) log.info(" %-14s %-16s%d פסקאות", src["source"], src["decision_number"], n)
if not apply: if not apply:
continue continue
await sx.extract_and_store(
await db.delete_style_exemplars(src["decision_number"], src["source"]) decision_number=src["decision_number"], source=src["source"],
texts = [u[1] for u in units] full_text=src["full_text"], practice_area=src["practice_area"],
vecs = await embeddings.embed_texts(texts, input_type="document") outcome=src["outcome"],
for (section, para), vec in zip(units, vecs): )
await db.insert_style_exemplar(
decision_number=src["decision_number"], source=src["source"],
practice_area=src["practice_area"], outcome=src["outcome"],
section=section, paragraph_text=para, word_count=len(para.split()),
embedding=vec,
)
if apply: if apply:
cov = await db.count_style_exemplars() cov = await db.count_style_exemplars()

View File

@@ -0,0 +1,99 @@
"""Classify the TREATMENT of each internal citation edge (X11 Phase 2, #154).
Each row in ``precedent_internal_citations`` records that a committee decision
cited a precedent, with the surrounding ``match_context``. Until now the edge's
``treatment`` column was empty, so the verified/authority layer counted every
citation as if it were positive — a precedent *distinguished* N times got the
same authority boost as one *followed* N times (an INV-COR2 violation).
This script fills ``treatment`` per edge by classifying the context with
``corroboration.classify_treatment`` (Opus 4.8 @ xhigh via the local
claude_session bridge — LOCAL ONLY, the claude CLI is not in the container):
followed | explained → positive (counts toward authority)
distinguished | criticized |
questioned | overruled → negative (never counts; overruled = demote)
Scope: only LINKED edges (``cited_case_law_id IS NOT NULL``) with an empty
``treatment`` and a non-empty ``match_context``. Idempotent — a second run skips
rows already classified. After applying, run ``scripts/build_verified_layer.py``
(or ``db.refresh_verified_layer``) so the treatment-aware count takes effect.
Run (dry-run, default — classifies and PRINTS, writes nothing):
HOME=/home/chaim mcp-server/.venv/bin/python scripts/classify_citation_treatments.py
Apply:
HOME=/home/chaim mcp-server/.venv/bin/python scripts/classify_citation_treatments.py --apply
Options:
--limit N classify at most N edges (smoke test)
--case-law-id UUID restrict to citations TO this one precedent
"""
from __future__ import annotations
import argparse
import asyncio
import os
import sys
from uuid import UUID
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "mcp-server", "src"))
from legal_mcp.services import corroboration, db # noqa: E402
async def _pending(limit: int | None, case_law_id: str | None) -> list[dict]:
pool = await db.get_pool()
where = ["cited_case_law_id IS NOT NULL", "coalesce(treatment,'') = ''",
"coalesce(match_context,'') <> ''"]
params: list = []
if case_law_id:
params.append(UUID(case_law_id))
where.append(f"cited_case_law_id = ${len(params)}")
sql = (f"SELECT id, cited_case_number, match_context "
f"FROM precedent_internal_citations WHERE {' AND '.join(where)} "
f"ORDER BY created_at")
if limit:
sql += f" LIMIT {int(limit)}"
rows = await pool.fetch(sql, *params)
return [dict(r) for r in rows]
async def main() -> None:
ap = argparse.ArgumentParser()
ap.add_argument("--apply", action="store_true", help="write changes (default: dry-run)")
ap.add_argument("--limit", type=int, default=None)
ap.add_argument("--case-law-id", type=str, default=None)
args = ap.parse_args()
rows = await _pending(args.limit, args.case_law_id)
print(f"קצוות לא-מסווגים לעיבוד: {len(rows)}\n")
pool = await db.get_pool()
counts: dict[str, int] = {}
errors = 0
for r in rows:
try:
t = await corroboration.classify_treatment(
r["cited_case_number"] or "", r["match_context"] or "")
except Exception as e: # noqa: BLE001 — one bad row must not abort the batch
errors += 1
print(f" ✗ [error] {r['cited_case_number']}: {type(e).__name__}: {e}")
continue
counts[t] = counts.get(t, 0) + 1
sign = "" if corroboration.is_positive(t) else ("" if corroboration.is_negative(t) else "·")
print(f" {sign} {t:<14} {r['cited_case_number']}")
if args.apply:
await pool.execute(
"UPDATE precedent_internal_citations SET treatment = $2 WHERE id = $1",
r["id"], t,
)
print(f"\nסיכום טיפול: {counts} שגיאות={errors}"
+ ("" if args.apply else " (dry-run — לא נכתב)"))
if args.apply:
print("הרץ עכשיו: scripts/build_verified_layer.py (או db.refresh_verified_layer) "
"כדי שהספירה מודעת-הטיפול תיכנס לתוקף.")
if __name__ == "__main__":
asyncio.run(main())

View File

@@ -0,0 +1,145 @@
"""Ad-hoc: executive summary (סיכום מנהלים) DOCX for case 1043-02-26.
Reuses the dafna decision template styles (David font + RTL) via the
analysis_docx_exporter helpers. One-off prep document for chaim's meeting
with the chair — NOT a decision draft.
"""
from pathlib import Path
import sys
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "mcp-server" / "src"))
from docx import Document
from legal_mcp.services.analysis_docx_exporter import (
TEMPLATE_PATH,
_clear_body,
_add_paragraph,
_add_runs_with_inline_bold,
_mark_paragraph_rtl,
_mark_run_rtl,
)
CASE = "1043-02-26"
OUT = Path(f"/home/chaim/legal-ai/data/cases/{CASE}/exports/סיכום-מנהלים-v1.docx")
def H1(doc, t):
_add_paragraph(doc, t, "Heading 1")
def H2(doc, t):
_add_paragraph(doc, t, "Heading 2")
def P(doc, t):
p = doc.add_paragraph(style="Normal")
_add_runs_with_inline_bold(p, t)
_mark_paragraph_rtl(p)
return p
def BULLET(doc, t):
p = doc.add_paragraph(style="List Paragraph")
_add_runs_with_inline_bold(p, t)
_mark_paragraph_rtl(p)
return p
def LABEL(doc, label, value):
p = doc.add_paragraph(style="Normal")
r = p.add_run(label + ": ")
r.bold = True
_mark_run_rtl(r)
r2 = p.add_run(value)
_mark_run_rtl(r2)
_mark_paragraph_rtl(p)
return p
def main():
doc = Document(str(TEMPLATE_PATH))
_clear_body(doc)
H1(doc, "סיכום מנהלים — הכנה לדיון")
P(doc, "**ערר 1043-02-26 — הקמת מתקן למיון פסולת (אתר \"קומפוסט דלילה\", תכנית מי/1030)**")
P(doc, "מסמך הכנה פנימי לקראת דיון עם יו\"ר הוועדה. אינו החלטה ואינו טיוטת החלטה.")
H2(doc, "פרטי התיק")
LABEL(doc, "סוג הערר", "רישוי ובנייה (1xxx) — ערר על סירוב בקשה להיתר")
LABEL(doc, "עוררים (מבקשי ההיתר)", "קיבוץ נחשון; חברת אלקטרה אקו גרין פארק")
LABEL(doc, "משיבים", "הוועדה המקומית מטה יהודה; מושבי גפן/תירוש/כפר הריף; קיבוצי כפר מנחם/רבדים; מועצה אזורית יואב (מתנגדים)")
LABEL(doc, "מושא הערר", "החלטת הוועדה המקומית מטה יהודה מיום 9.2.26 לסרב לבקשה (מס' 20240972)")
LABEL(doc, "המקרקעין", "גוש 5093 ח\"ל 4 מגרש 1, אתר \"דלילה\" (~137 דונם), מצפון לכפר מנחם, ליד כביש 383")
LABEL(doc, "תקן ביקורת", "ועדת הערר כמוסד תכנון בעל סמכות מקורית — שיקול דעת תכנוני עצמאי; ביקורת רחבה יותר בחלק המשפטי-פרשני")
H2(doc, "מהות המחלוקת")
P(doc, "מבקשי ההיתר מבקשים להקים מתקן מיון פסולת/קומפוסטציה הכולל מבנה קומפוסטציה אחוד (~19,118 מ\"ר עיקרי), משטחי הבשלה פתוחים ובריכה תפעולית, בצירוף שתי הקלות: הגבהת גובה מ-8 ל-20 מ', והגדלת תכסית מ-3.5% ל-22%. הוועדה המקומית סירבה, בקובעה כי השינוי מהותי ומקומו בעדכון התכנית ולא בהליך רישוי. השאלה המרכזית: האם הבקשה תואמת את תכנית מי/1030, או חורגת ממנה באופן המחייב תיקון תכנית.")
H2(doc, "החלטת הוועדה המקומית (מושא הערר)")
P(doc, "**דחתה** את טענת המתנגדים לפקיעת התכנית (בהסתמך על ע\"א 3213/97 נקר) ואת טענת השימושים האסורים.")
P(doc, "**קיבלה** והפכה לבסיס הסירוב: (1) היקף הבינוי חורג משלב א' לפי הוראת השלביות; (2) הבינוי המונוליטי שונה דרמטית מנספח הבינוי; (3) משטחי ההבשלה הפתוחים מנוגדים לסעיף 6.9 (טיפול במבנים סגורים) — שהפרתו מוגדרת בתכנית כסטייה ניכרת.")
H2(doc, "טענות סף (לדיון תמציתי)")
BULLET(doc, "**פקיעת תכנית** — סעיף הפקיעה (5 שנים) הושמט מהנוסח המאושר (ראו ממצאי ניתוח-העומק).")
BULLET(doc, "**מיצוי התנגדות / \"מועד ב'\"** — האם טענות שהוכרעו בשלב התכנית מועלות מחדש בשלב הרישוי.")
BULLET(doc, "**זכות עמידה** — מועד חתימת חוזה החכירה מול מועד פתיחת הבקשה (הוכרע: קיבוץ נחשון חוכר רשום).")
BULLET(doc, "**פגמי פרסום / היעדר יידוע** (ס' 149(א)(2א)) — היעדר תשריט, אי-יידוע ועדה מקומית יואב.")
BULLET(doc, "**איחור בהגשת ההתנגדויות** — מול אינטרס ההסתמכות (רכישת קרקע ב-35 מיליון ₪).")
H2(doc, "חמש הסוגיות המהותיות")
P(doc, "**סוגיה 1 (מכריעה) — שלביות הביצוע: בינוי מול היקף פעילות**")
BULLET(doc, "השאלה: האם הוראת השלביות (ס' 7.1) חלה על הבינוי הפיזי או רק על היקף ההפעלה?")
BULLET(doc, "עמדות: עוררים — נוגעת להיקף הפעילות; ועדה/משיבים — חלה על הבינוי (4 מבנים ≈ 8,000 מ\"ר מול ~19,118 מבוקשים).")
BULLET(doc, "לאן נוטה: **לטובת הוועדה** — לשון \"הקמת 4 מבני קומפוסט\" מתייחסת לבינוי; נספחים 06, 11 מצביעים על מנגנון פיילוט מדורג מהותי.")
BULLET(doc, "תקדים: עע\"מ 10089/07 אירוס הגלבוע (אין לעקוף תכנון נדרש דרך היתר).")
P(doc, "**סוגיה 2 (מכריעה) — משטחי ההבשלה וסעיף 6.9: \"טיפול\" מול \"אחסון\"**")
BULLET(doc, "השאלה: האם ההבשלה הפתוחה היא \"טיפול\" החוסה תחת 6.9 (→ סטייה ניכרת החוסמת הקלה, ס' 151), או \"אחסון\"?")
BULLET(doc, "עמדות: עוררים (ד\"ר ענבר, נספח 22) — אחסון מוצר סופי נטול ריח; ועדה — חלק בלתי נפרד מהטיפול.")
BULLET(doc, "לאן נוטה: **נחלשה לוועדה** — נספח 07 §122-127: ועדת המשנה לעררים אישרה הבשלה פתוחה כ\"נכון וסביר\" (אך כינתה זאת \"שלב אחרון של טיפול\"). מתח פרשני אמיתי.")
BULLET(doc, "תקדים: עע\"מ 402/03 עמותת העצמאים (ס' 151 — סטייה ניכרת).")
P(doc, "**סוגיה 3 — ההקלות בגובה ובתכסית + טענת \"טעות סופר\"**")
BULLET(doc, "השאלה: האם ההקלות בגדרי הקלה או סטייה ניכרת? והאם התכסית 3.5% היא \"טעות סופר\"?")
BULLET(doc, "לאן נוטה: טענת טעות הסופר **התחזקה מאוד** — נספח 12 (מופקד) מראה ש-3.5% נגזרה מתפיסה שהחממות \"אינן שטח לבניה\"; נספחים 06 §14 ו-07 §130 הורו לתקן; הסכם רמ\"י נוקב ב-30,280 מ\"ר.")
BULLET(doc, "תקדים: עמ\"נ 25955-11-22 ברק-רחביה (פרשנות הרמונית); בג\"ץ 2667/17 מטה בנימין (\"פרשנות אפשרית\").")
P(doc, "**סוגיה 4 — תצורת הבינוי מול נספח בינוי מנחה ונספח נופי מחייב**")
BULLET(doc, "השאלה: האם המבנה המונוליטי + הבריכה חורגים ממרחב הגמישות של נספח מנחה, לנוכח הנספח הנופי המחייב?")
BULLET(doc, "תקדים: **ערר 1033-25 אבו גוש** (תקדים דפנה ישיר — נספח בינוי מנחה אינו המלצה בלבד); בג\"ץ 6525/15 עמק שווה.")
P(doc, "**סוגיה 5 — שיקולים זרים / NIMBY בהחלטת הסירוב**")
BULLET(doc, "השאלה: האם הסירוב נגוע בשיקולים זרים? (עשוי להתייתר אם הבחינה העצמאית מכריעה).")
BULLET(doc, "לאן נוטה: **נתמך בראיות** — תמליל (נספח 19, \"ועדה פוליטית\") + הצוות המקצועי המליץ לאשר (נספח 17). מנגד — בסיס מהותי לא-NIMBY (נספחים 25, 30).")
H2(doc, "ממצאי ניתוח-העומק — מה התחדש לאחר מיצוי 22 נספחי הרקע")
BULLET(doc, "**6.9 (ליבת התיק) נחלש לוועדה** — ועדת המשנה לעררים כבר אישרה הבשלה פתוחה.")
BULLET(doc, "**טעות הסופר בתכסית התחזקה** — שתי ערכאות הורו לתקן את חישוב השטח.")
BULLET(doc, "**פקיעה — הוכרע עובדתית** — סעיף הפקיעה היה בנוסח המופקד (נספח 12) והושמט מהמאושר; נותרה מחלוקת משפטית בלבד (נקר/לויתן מול חמדת הגליל).")
BULLET(doc, "**אופי המתקן מטה למשיבים** — תת\"ל 220 + ויתור על 8 מבני קומפוסט לטובת מתקן תרמי (נספחים 25, 30) → טיעון \"פריסת סלאמי\"/עקיפת תכנון.")
BULLET(doc, "**תיקון עובדתי** — פסה\"ד שדחה את העתירה (נספח 10) ניתן נגד תכנית דרך הגישה, לא נגד מי/1030.")
BULLET(doc, "**פער תחבורתי** — 80-92 משאיות/יום (מוסדות התכנון) מול 320-400 (ד\"ר לינק) — טעון יישוב.")
H2(doc, "פסיקה הדורשת אימות חיצוני (אינה בקורפוס הסמכותי)")
BULLET(doc, "ע\"א 3213/97 **נקר** — \"הדין המתפרסם ברבים מחייב\" (עוגן הוועדה לדחיית הפקיעה).")
BULLET(doc, "עע\"מ 4768/22 **חמדת הגליל** — פקיעת תכנית (עוגן המשיבים).")
BULLET(doc, "ע\"א 482/99 בלפוריה; בג\"ץ 5636/13 מתיישבי תימורים; בג\"ץ 9098/01 גניס; דנ\"א 3993/07 איקאפוד.")
H2(doc, "שאלות פתוחות להכרעת היו\"ר + סדר דיון מומלץ")
BULLET(doc, "(1) סיווג ההבשלה — טיפול (6.9) או אחסון (4.1.1)?")
BULLET(doc, "(2) דין הפקיעה לאור ההשמטה המוכחת מהנוסח המאושר.")
BULLET(doc, "(3) האם התכסית 3.5% היא טעות סופר הניתנת לתיקון פרשני?")
BULLET(doc, "(4) האם תת\"ל 220 + הוויתור על מבני הקומפוסט הופכים את הבקשה ל\"עקיפת תכנון\" (אירוס הגלבוע)?")
BULLET(doc, "(5) הסעד: סירוב מלא / אישור מותנה בהתאמה לשלב א' / החזרה לוועדה המקומית עם הנחיות.")
P(doc, "**סדר דיון מומלץ:** טענות סף (פקיעה → מיצוי/השתק → עמידה/פרסום) ← סוגיה 1 (שלביות) ← סוגיה 2 (6.9) ← סוגיה 3 (הקלות/טעות סופר) ← סוגיה 4 (תצורת בינוי) ← סוגיה 5 (שיקולים זרים).")
H2(doc, "הערכת תרחישים")
P(doc, "התמונה שקולה. לטובת הוועדה: סוגיה 1 (שלביות) וטיעון \"עקיפת תכנון/סלאמי\" נוכח תת\"ל 220 — חזקים. לטובת העוררים: סוגיה 2 (6.9) נחלשה, טעות הסופר התחזקה, ו-NIMBY נתמך בראיות. **התרחיש הסביר ביותר:** קבלה חלקית / החזרה מותנית או דחייה — תלוי בעיקר בשאלת היקף הבינוי בשלב א' ובשאלת \"עקיפת התכנון\". ההכרעה במובהק של יו\"ר הוועדה.")
OUT.parent.mkdir(parents=True, exist_ok=True)
doc.save(str(OUT))
print(f"saved: {OUT}")
if __name__ == "__main__":
main()

View File

@@ -91,15 +91,22 @@ async def main(args: argparse.Namespace) -> int:
# The 3 steps as durable nodes (X16 / INV-DUR1) — shared runtime with # The 3 steps as durable nodes (X16 / INV-DUR1) — shared runtime with
# final_halacha (scripts/_pipeline_runtime.py). A crash/OOM in the long style # final_halacha (scripts/_pipeline_runtime.py). A crash/OOM in the long style
# panel [3] resumes from [3] instead of re-paying the Opus distillation [1]. # panel [3] resumes from [3] instead of re-paying the Opus distillation [2].
#
# Order (#159): enroll FIRST. enroll creates the style_corpus row (fast, no LLM)
# that decision_lessons attach to. Running it before the up-to-30-min distillation
# means the corpus exists within seconds — so the curator agent's §A (which fires
# in parallel on a final_learning_* wake) can record source='curator' findings
# without racing the 30-min step. enroll/ingest are mutually independent; panel
# needs both, so it stays last.
async def step_ingest(results: dict) -> dict: async def step_ingest(results: dict) -> dict:
# [1] distillation (Opus) — skip if already analyzed (idempotent; --force to redo). # [2] distillation (Opus) — skip if already analyzed (idempotent; --force to redo).
status = await _latest_pair_status(case["id"]) status = await _latest_pair_status(case["id"])
if status == "analyzed" and not args.force: if status == "analyzed" and not args.force:
print("[1/3] ingest_final_version — דולג (הזוג כבר analyzed; --force לחידוש)") print("[2/3] ingest_final_version — דולג (הזוג כבר analyzed; --force לחידוש)")
return {"ingest": "skipped:analyzed"} return {"ingest": "skipped:analyzed"}
print("[1/3] ingest_final_version — דיסטילציית טיוטה↔סופי…", flush=True) print("[2/3] ingest_final_version — דיסטילציית טיוטה↔סופי…", flush=True)
raw = await ingest_final_version(case_number, file_path=final_path) raw = await ingest_final_version(case_number, file_path=final_path)
try: try:
env = json.loads(raw) env = json.loads(raw)
@@ -120,8 +127,9 @@ async def main(args: argparse.Namespace) -> int:
return {"ingest": "done"} return {"ingest": "done"}
async def step_enroll(results: dict) -> dict: async def step_enroll(results: dict) -> dict:
# [2] enroll into style_corpus (idempotent) — lessons need a corpus_id. # [1] enroll into style_corpus (idempotent) — lessons need a corpus_id, and
print("[2/3] רישום לקורפוס-הסגנון (idempotent)…", flush=True) # it must exist early (before the long distillation) so curator §A can attach.
print("[1/3] רישום לקורפוס-הסגנון (idempotent)…", flush=True)
if await _has_style_corpus(case_number): if await _has_style_corpus(case_number):
print(" ✓ כבר רשום בקורפוס-הסגנון") print(" ✓ כבר רשום בקורפוס-הסגנון")
return {"enroll": "exists"} return {"enroll": "exists"}
@@ -150,8 +158,8 @@ async def main(args: argparse.Namespace) -> int:
return {"panel_rc": rc or 0} return {"panel_rc": rc or 0}
steps = [ steps = [
_pipeline_runtime.Step("ingest_final_version", step_ingest),
_pipeline_runtime.Step("enroll_style_corpus", step_enroll), _pipeline_runtime.Step("enroll_style_corpus", step_enroll),
_pipeline_runtime.Step("ingest_final_version", step_ingest),
_pipeline_runtime.Step("style_panel", step_panel), _pipeline_runtime.Step("style_panel", step_panel),
] ]
checkpoint_db = config.DATA_DIR / "checkpoints" / "learning.sqlite" checkpoint_db = config.DATA_DIR / "checkpoints" / "learning.sqlite"

View File

@@ -0,0 +1,102 @@
"""Batch ingest of appeals-committee decisions staged in data/precedents/incoming/.
Sequential (NOT concurrent — avoids the 2026-05-31 load-spike incident) ingest of
each .doc/.docx via the canonical internal pipeline, followed by metadata extraction
per case (the internal path does NOT auto-queue metadata — INV-ING3). Halacha is
auto-queued by ingest; drain it separately via MCP precedent_process_pending.
case_number canonical follows the filename/Nevo convention validated against the
corpus + missing_precedents list:
- מרכז/חיפה/ת"א committees number with month: NNNN/MM/YY → NNNN-MM-YY
- ירושלים/צפון committees number without month: NNNN/YY → NNNN-YY
decision_date / summary / subject_tags / appeal_subtype are left empty on purpose —
the metadata extractor fills them from the full text (more reliable than parsing here).
Run: mcp-server/.venv/bin/python scripts/ingest_incoming_batch.py
Config (POSTGRES_URL, VOYAGE_API_KEY, ANTHROPIC_API_KEY) auto-loads from ~/.env.
"""
import asyncio
import os
import sys
import traceback
from pathlib import Path
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "mcp-server", "src"))
from legal_mcp.services import internal_decisions as svc
from legal_mcp.services import precedent_metadata_extractor as meta
INC = "/home/chaim/legal-ai/data/precedents/incoming"
# file, case_number(canonical-with-slashes), chair, district, court, practice_area
DECISIONS = [
("105-07.doc", "105/07", "דרור לביא-אפרת", "צפון", "rishuy_uvniya"),
("ARAR-17-105-44.doc", "105/17", "רונית אלפר", "מרכז", "betterment_levy"),
("ARAR-18-1029.doc", "1029/18", "אליעד וינשל", "ירושלים", "rishuy_uvniya"),
("ARAR-20-1018-44.doc", "1018/20", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-20-1023-55.doc", "1023/20", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-21-1080-55.doc", "1080/21", "נילי בן משה ידגר", "צפון", "rishuy_uvniya"),
("ARAR-21-11-1051.doc", "1051/11/21", "רונית אלפר", "מרכז", "rishuy_uvniya"),
("ARAR-22-01-1015.doc", "1015/01/22", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-22-06-1029.doc", "1029/06/22", "מאיה אשכנזי", "מרכז", "rishuy_uvniya"),
("ARAR-22-08-1044.doc", "1044/08/22", "מאיה אשכנזי", "מרכז", "rishuy_uvniya"),
("ARAR-22-10-1050.doc", "1050/10/22", "מאיה אשכנזי", "מרכז", "rishuy_uvniya"),
("ARAR-22-1079.doc", "1079/22", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-23-04-1010.doc", "1010/04/23", "מאיה אשכנזי", "מרכז", "rishuy_uvniya"),
("ARAR-23-08-1074-9.doc", "1074/08/23", "מיכל הלברשטם דגני", "חיפה", "rishuy_uvniya"),
("ARAR-23-1034.docx", "1006/23", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-23-1073.doc", "1073/23", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-23-1085.doc", "1085/23", "נילי בן משה ידגר", "צפון", "rishuy_uvniya"),
("ARAR-24-01-1009-5.docx","1009/01/24", "מיכל דגני הלברשטם", "תל אביב", "rishuy_uvniya"),
("ARAR-24-05-1044.doc", "1044/05/24", "שרית אריאלי בן שמחון", "ירושלים", "rishuy_uvniya"),
("ARAR-25-01-10072.docx", "1007/01/25", "יפעת בן אריה שטיינברג", "תל אביב", "rishuy_uvniya"),
]
async def main():
results = []
for fname, case_number, chair, district, parea in DECISIONS:
fp = Path(INC) / fname
rec = {"file": fname, "case_number": case_number}
if not fp.exists():
rec["error"] = "file-missing"
print(f"{fname}: file missing", flush=True)
results.append(rec)
continue
try:
out = await svc.ingest_internal_decision(
file_path=fp,
case_number=case_number,
chair_name=chair,
district=district,
court=f"ועדת הערר לתכנון ובנייה — מחוז {district}",
practice_area=parea,
proceeding_type="ערר",
is_binding=False,
)
cid = out.get("case_law_id")
rec["case_law_id"] = cid
rec["chunks"] = out.get("chunks")
print(f"✓ ingest {case_number}: id={cid} chunks={out.get('chunks')}", flush=True)
# metadata (internal path does not auto-queue it)
m = await meta.extract_and_apply(cid)
rec["meta_status"] = m.get("status")
sug = m.get("suggested") or {}
rec["suggested_case_number"] = sug.get("case_number_clean")
rec["citation_formatted"] = sug.get("citation_formatted")
rec["meta_date"] = sug.get("decision_date_iso")
print(f" meta {case_number}: {m.get('status')} | clean={sug.get('case_number_clean')} | {sug.get('citation_formatted')}", flush=True)
except Exception as e:
rec["error"] = f"{type(e).__name__}: {e}"
print(f"{fname} ({case_number}): {e}", flush=True)
traceback.print_exc()
results.append(rec)
print("\n===SUMMARY===", flush=True)
for r in results:
print(r, flush=True)
if __name__ == "__main__":
asyncio.run(main())

View File

@@ -0,0 +1,392 @@
#!/usr/bin/env python3
"""OCR Benchmark: Current system (PyMuPDF + Google Vision) vs Mistral OCR 4.0
Usage:
python scripts/ocr_benchmark_mistral.py [--docs N] [--output PATH]
Downloads PDFs from MinIO, calls Mistral OCR API, compares against
already-extracted text stored in the DB, and writes a Markdown report.
"""
from __future__ import annotations
import argparse
import asyncio
import base64
import json
import re
import subprocess
import sys
import tempfile
import time
from pathlib import Path
from typing import TypedDict
import httpx
# ── Config ───────────────────────────────────────────────────────────────────
MISTRAL_API_KEY = "UsZjLCX30ev6pox0KgvXyuFP3ktsPYpN"
MISTRAL_OCR_MODEL = "mistral-ocr-latest"
MISTRAL_OCR_URL = "https://api.mistral.ai/v1/ocr"
MINIO_ALIAS = "legalminio"
MINIO_BUCKET = "legal-documents"
# Documents to benchmark — selected for diversity:
# - main appeal (40 pp, 4.4 MB, large digital doc)
# - permit (9 pp, 1.7 MB, likely partially scanned)
# - protocol (5 pp, 178 KB, administrative typed)
# - response (12 pp, ~390 KB, digital legal brief)
DOCS_TO_BENCHMARK = [
{
"title": "כתב ערר",
"minio_key": "cases/1027-04-26/documents/originals/כתב ערר - מושב נחם נ׳ ועדה מקומית בית שמש.pdf",
"doc_type": "appeal",
"pages": 40,
},
{
"title": "נספח 1 — היתר הבנייה",
"minio_key": "cases/1027-04-26/documents/originals/נספח 1 - היתר הבנייה (פורסם 17.03.26).pdf",
"doc_type": "permit",
"pages": 9,
},
{
"title": "נספח 14 — פרוטוקול ועדה מחוזית",
"minio_key": "cases/1027-04-26/documents/originals/נספח 14 - פרוטוקול דיון ועדה מחוזית - 06.06.23.pdf",
"doc_type": "protocol",
"pages": 5,
},
{
"title": "תגובת המשיבה 3",
"minio_key": "cases/1027-04-26/documents/originals/תגובת המשיבה 3 לערר ולבקשה להתליית היתר בנייה.pdf",
"doc_type": "response",
"pages": 12,
},
]
# ── DB access (read-only: fetch already-extracted text) ──────────────────────
def _fetch_extracted_texts(case_number: str) -> dict[str, str]:
"""Pull extracted_text from DB for all docs in the case."""
import psycopg2 # type: ignore
conn = psycopg2.connect(
host="localhost", port=5433, dbname="legal_ai",
user="legal_ai", password="od0ASJZFYibOlWK59krLvvETmgqwlXe8",
)
cur = conn.cursor()
cur.execute(
"""
SELECT d.title, d.extracted_text, d.file_path, d.page_count
FROM documents d
JOIN cases c ON c.id = d.case_id
WHERE c.case_number = %s AND d.extracted_text IS NOT NULL
""",
(case_number,),
)
rows = cur.fetchall()
cur.close()
conn.close()
return {row[0]: {"text": row[1], "file_path": row[2], "pages": row[3]} for row in rows}
# ── MinIO download ────────────────────────────────────────────────────────────
def download_from_minio(minio_key: str, dest: Path) -> None:
"""Download a file from MinIO using mcli."""
src = f"{MINIO_ALIAS}/{MINIO_BUCKET}/{minio_key}"
result = subprocess.run(
["mcli", "cp", src, str(dest)],
capture_output=True, text=True,
)
if result.returncode != 0:
raise RuntimeError(f"mcli cp failed: {result.stderr}")
# ── Text quality metrics ─────────────────────────────────────────────────────
_HEBREW_RE = re.compile(r'[֐-׿]')
_WORD_RE = re.compile(r'\S+')
def compute_metrics(text: str) -> dict:
"""Compute quality metrics for extracted text."""
if not text:
return {"chars": 0, "words": 0, "avg_word_len": 0,
"single_char_pct": 0, "words_per_line": 0,
"hebrew_pct": 0, "quality_ok": False}
words = _WORD_RE.findall(text)
n_words = len(words)
if n_words == 0:
return {"chars": len(text), "words": 0, "avg_word_len": 0,
"single_char_pct": 0, "words_per_line": 0,
"hebrew_pct": 0, "quality_ok": False}
avg_len = sum(len(w) for w in words) / n_words
single_char_pct = sum(1 for w in words if len(w) == 1) / n_words
lines = [ln for ln in text.split("\n") if ln.strip()]
words_per_line = n_words / len(lines) if lines else 0
letters = re.findall(r'[a-zA-Z֐-׿]', text)
hebrew_pct = (
sum(1 for c in letters if _HEBREW_RE.match(c)) / len(letters)
if letters else 0
)
quality_ok = (
n_words >= 10
and avg_len >= 2.5
and single_char_pct <= 0.4
and words_per_line >= 3.0
and hebrew_pct >= 0.5
)
return {
"chars": len(text),
"words": n_words,
"avg_word_len": round(avg_len, 2),
"single_char_pct": round(single_char_pct * 100, 1),
"words_per_line": round(words_per_line, 1),
"hebrew_pct": round(hebrew_pct * 100, 1),
"quality_ok": quality_ok,
}
def count_known_abbrev_errors(text: str) -> int:
"""Count common Hebrew abbreviation OCR errors (pre-fix indicators)."""
patterns = ['עוהייד', 'עוייד', 'הנייל', 'ביהמייש', 'עייי', 'בייכ', 'תמייא']
return sum(text.count(p) for p in patterns)
def count_correct_abbrevs(text: str) -> int:
"""Count correctly rendered Hebrew abbreviations."""
correct = ['עו"ד', 'הנ"ל', 'ביהמ"ש', 'ע"י', 'ב"כ', 'תמ"א', 'ס"ק']
return sum(text.count(p) for p in correct)
# ── Mistral OCR ───────────────────────────────────────────────────────────────
def call_mistral_ocr(pdf_path: Path) -> tuple[str, float]:
"""Call Mistral OCR API on a PDF. Returns (extracted_text, elapsed_seconds)."""
with open(pdf_path, "rb") as f:
pdf_b64 = base64.b64encode(f.read()).decode()
payload = {
"model": MISTRAL_OCR_MODEL,
"document": {
"type": "document_url",
"document_url": f"data:application/pdf;base64,{pdf_b64}",
},
"include_image_base64": False,
}
t0 = time.time()
with httpx.Client(timeout=300.0) as client:
resp = client.post(
MISTRAL_OCR_URL,
headers={
"Authorization": f"Bearer {MISTRAL_API_KEY}",
"Content-Type": "application/json",
},
json=payload,
)
elapsed = time.time() - t0
if resp.status_code != 200:
raise RuntimeError(f"Mistral OCR error {resp.status_code}: {resp.text[:500]}")
data = resp.json()
# Response: {"pages": [{"index": 0, "markdown": "..."}, ...]}
pages = data.get("pages", [])
text = "\n\n".join(p.get("markdown", "") for p in pages)
return text, elapsed
# ── Report ────────────────────────────────────────────────────────────────────
def render_metric_table(current: dict, mistral: dict) -> str:
rows = [
("תווים", f"{current['chars']:,}", f"{mistral['chars']:,}"),
("מילים", f"{current['words']:,}", f"{mistral['words']:,}"),
("אורך מילה ממוצע", str(current['avg_word_len']), str(mistral['avg_word_len'])),
("% מילים חד-תוויות", f"{current['single_char_pct']}%", f"{mistral['single_char_pct']}%"),
("מילים לשורה", str(current['words_per_line']), str(mistral['words_per_line'])),
("% תווים עבריים", f"{current['hebrew_pct']}%", f"{mistral['hebrew_pct']}%"),
("איכות כוללת", "" if current['quality_ok'] else "", "" if mistral['quality_ok'] else ""),
]
lines = ["| מדד | OCR נוכחי | Mistral OCR |", "|-----|-----------|-------------|"]
for label, cur_val, mis_val in rows:
lines.append(f"| {label} | {cur_val} | {mis_val} |")
return "\n".join(lines)
def build_report(results: list[dict], output_path: Path) -> None:
lines = [
"# השוואת OCR: מערכת נוכחית מול Mistral OCR",
f"\n**תיק:** 1027-04-26 — בל\"מ מושב נחם מפעל בטון בית שמש ",
f"**תאריך:** {time.strftime('%Y-%m-%d %H:%M')} ",
f"**מודל Mistral:** `{MISTRAL_OCR_MODEL}` ",
f"**מערכת נוכחית:** PyMuPDF (born-digital) + Google Cloud Vision (scanned) ",
"\n---\n",
"## סיכום מנהלים\n",
]
# Summary table
sum_lines = ["| מסמך | עמודים | נוכחי תווים | Mistral תווים | Mistral זמן | עדיפות |",
"|------|--------|-------------|---------------|-------------|--------|"]
for r in results:
if "error" in r:
sum_lines.append(f"| {r['title']} | {r['pages']} | — | שגיאה | — | — |")
continue
cur_chars = r["current_metrics"]["chars"]
mis_chars = r["mistral_metrics"]["chars"]
winner = "🔵 נוכחי" if cur_chars > mis_chars * 1.05 else (
"🟢 Mistral" if mis_chars > cur_chars * 1.05 else "⚖️ שקול")
sum_lines.append(
f"| {r['title']} | {r['pages']} | {cur_chars:,} | {mis_chars:,} | "
f"{r['mistral_elapsed']:.1f}s | {winner} |"
)
lines.extend(sum_lines)
lines.append("\n---\n")
# Per-document detail
for r in results:
lines.append(f"## {r['title']} ({r['pages']} עמודים)\n")
if "error" in r:
lines.append(f"**שגיאה ב-Mistral OCR:** `{r['error']}`\n")
continue
lines.append(render_metric_table(r["current_metrics"], r["mistral_metrics"]))
lines.append("")
cur_abbr_err = count_known_abbrev_errors(r["current_text"])
mis_abbr_err = count_known_abbrev_errors(r["mistral_text"])
cur_abbr_ok = count_correct_abbrevs(r["current_text"])
mis_abbr_ok = count_correct_abbrevs(r["mistral_text"])
lines.append(f"\n**קיצורים עבריים:**")
lines.append(f"- נוכחי: {cur_abbr_ok} נכונים, {cur_abbr_err} שגויים")
lines.append(f"- Mistral: {mis_abbr_ok} נכונים, {mis_abbr_err} שגויים")
lines.append(f"\n**זמן Mistral:** {r['mistral_elapsed']:.1f} שניות\n")
# Side-by-side first 600 chars
lines.append("### דוגמת טקסט — 600 תווים ראשונים\n")
lines.append("**מערכת נוכחית:**")
lines.append("```")
lines.append((r["current_text"] or "")[:600].replace("```", "'''"))
lines.append("```\n")
lines.append("**Mistral OCR:**")
lines.append("```")
lines.append((r["mistral_text"] or "")[:600].replace("```", "'''"))
lines.append("```\n")
lines.append("---\n")
output_path.write_text("\n".join(lines), encoding="utf-8")
print(f"\n✅ דוח נשמר: {output_path}")
# ── Main ──────────────────────────────────────────────────────────────────────
def main() -> None:
parser = argparse.ArgumentParser(description="OCR benchmark: current vs Mistral")
parser.add_argument("--docs", type=int, default=4,
help="כמה מסמכים לבדוק (ברירת מחדל: 4)")
parser.add_argument("--output", type=str,
default="/home/chaim/legal-ai/data/audit/ocr-benchmark-mistral.md",
help="נתיב לדוח הפלט")
args = parser.parse_args()
docs = DOCS_TO_BENCHMARK[: args.docs]
output_path = Path(args.output)
output_path.parent.mkdir(parents=True, exist_ok=True)
print(f"🔍 שולף טקסטים קיימים מה-DB עבור תיק 1027-04-26...")
db_texts = _fetch_extracted_texts("1027-04-26")
print(f" נמצאו {len(db_texts)} מסמכים עם טקסט מחולץ")
results = []
with tempfile.TemporaryDirectory() as tmp_dir:
for doc in docs:
print(f"\n📄 מעבד: {doc['title']} ({doc['pages']} עמודים)")
# --- Current OCR text from DB ---
current_text = ""
for title, info in db_texts.items():
if doc["title"].split("")[0].strip() in title or doc["minio_key"].split("/")[-1] in (info.get("file_path") or ""):
current_text = info["text"] or ""
break
if not current_text:
# Try by file_path suffix match
key_name = doc["minio_key"].split("/")[-1]
for title, info in db_texts.items():
if key_name in (info.get("file_path") or ""):
current_text = info["text"] or ""
break
if not current_text:
print(f" ⚠️ לא נמצא טקסט ב-DB, מחפש לפי שם מסמך...")
# fallback: match by doc_type + rough title
for title, info in db_texts.items():
if doc["doc_type"] in title.lower() or doc["title"][:8] in title:
current_text = info["text"] or ""
break
print(f" נוכחי: {len(current_text):,} תווים")
# --- Download PDF from MinIO ---
pdf_name = doc["minio_key"].split("/")[-1]
pdf_path = Path(tmp_dir) / pdf_name
print(f" מוריד מ-MinIO...")
try:
download_from_minio(doc["minio_key"], pdf_path)
print(f" הורד: {pdf_path.stat().st_size:,} bytes")
except Exception as e:
print(f" ❌ שגיאה בהורדה: {e}")
results.append({"title": doc["title"], "pages": doc["pages"], "error": str(e)})
continue
# --- Mistral OCR ---
print(f" קורא Mistral OCR API...")
try:
mistral_text, elapsed = call_mistral_ocr(pdf_path)
print(f" Mistral: {len(mistral_text):,} תווים ({elapsed:.1f}s)")
except Exception as e:
print(f" ❌ שגיאת Mistral API: {e}")
results.append({
"title": doc["title"], "pages": doc["pages"],
"current_text": current_text,
"current_metrics": compute_metrics(current_text),
"mistral_text": "", "mistral_metrics": compute_metrics(""),
"mistral_elapsed": 0, "error": str(e),
})
continue
results.append({
"title": doc["title"],
"pages": doc["pages"],
"current_text": current_text,
"current_metrics": compute_metrics(current_text),
"mistral_text": mistral_text,
"mistral_metrics": compute_metrics(mistral_text),
"mistral_elapsed": elapsed,
})
print("\n📊 בונה דוח...")
build_report(results, output_path)
# Also save raw texts for manual inspection
raw_dir = output_path.parent / "ocr-benchmark-raw"
raw_dir.mkdir(exist_ok=True)
for r in results:
safe = r["title"].replace("/", "-").replace(" ", "_")[:40]
if "current_text" in r:
(raw_dir / f"{safe}_current.txt").write_text(r["current_text"], encoding="utf-8")
if "mistral_text" in r:
(raw_dir / f"{safe}_mistral.txt").write_text(r["mistral_text"], encoding="utf-8")
print(f"💾 טקסטים גולמיים נשמרו: {raw_dir}")
if __name__ == "__main__":
main()

View File

@@ -19,7 +19,13 @@ migrates atomically per case — is everything that embeds the number as *text*:
4. MinIO keys cases/{old}/… 3 buckets; cp→new then rm old. 4. MinIO keys cases/{old}/… 3 buckets; cp→new then rm old.
legal-immutable (WORM/object-lock) → copy-only, old object stays locked. legal-immutable (WORM/object-lock) → copy-only, old object stays locked.
5. Gitea repo cases/{old} API PATCH name + local .git remote rewrite 5. Gitea repo cases/{old} API PATCH name + local .git remote rewrite
6. Paperclip project name replace(old→new) so case↔project lookup holds 6. Paperclip case↔issue linkage replace(old→new) in THREE places, because the
legal-ai → Paperclip lookup (get_case_issues) keys on the case number as text:
• projects.name so case↔project lookup holds
• plugin_state.legal-case-number the authoritative issue linkage value_json
• issues.title the '[ערר {cn}] …' tag the title-path lookup uses
Without (b)+(c) the issues keep the OLD number and get_case_issues returns [],
so post-final actions (run-learning / run-halacha) silently skip ("no_issue").
Bare occurrences of the old number that are NOT inside a '/cases/{old}/' path Bare occurrences of the old number that are NOT inside a '/cases/{old}/' path
(e.g. prose in notes, a citation) are *reported for review*, never auto-edited. (e.g. prose in notes, a citation) are *reported for review*, never auto-edited.
@@ -273,10 +279,25 @@ async def inspect_paperclip(old: str) -> dict:
try: try:
c = await asyncpg.connect(PAPERCLIP_DSN, timeout=10) c = await asyncpg.connect(PAPERCLIP_DSN, timeout=10)
except Exception as e: except Exception as e:
return {"reachable": False, "error": str(e)[:120], "projects": []} return {"reachable": False, "error": str(e)[:120],
"projects": [], "linkage_rows": 0, "issue_titles": 0}
try: try:
rows = await c.fetch("SELECT id, name FROM projects WHERE name LIKE $1", f"%{old}%") rows = await c.fetch("SELECT id, name FROM projects WHERE name LIKE $1", f"%{old}%")
return {"reachable": True, "projects": [(str(r["id"]), r["name"]) for r in rows]} # The two surfaces get_case_issues actually keys on: the legal-case-number
# plugin_state linkage and the '[ערר {cn}]' tag in issue titles.
linkage = await c.fetchval(
"SELECT count(*) FROM plugin_state "
"WHERE state_key = 'legal-case-number' AND value_json = to_jsonb($1::text)",
old,
)
titles = await c.fetchval(
"SELECT count(*) FROM issues WHERE title LIKE $1", f"%{old}%")
return {
"reachable": True,
"projects": [(str(r["id"]), r["name"]) for r in rows],
"linkage_rows": linkage or 0,
"issue_titles": titles or 0,
}
finally: finally:
await c.close() await c.close()
@@ -369,16 +390,31 @@ async def apply_case(conn, rec: dict, *, skip_minio: bool, skip_gitea: bool,
except urllib.error.HTTPError as e: except urllib.error.HTTPError as e:
log(f" ✗ Gitea rename failed: HTTP {e.code} {e.read()[:160]!r}") log(f" ✗ Gitea rename failed: HTTP {e.code} {e.read()[:160]!r}")
# 6. Paperclip project name # 6. Paperclip case↔issue linkage — project name + legal-case-number value +
if not skip_paperclip and rec["paperclip"].get("reachable") and rec["paperclip"]["projects"]: # issue titles. (b)+(c) are what get_case_issues keys on; without them the
# issues keep the old number and run-learning/run-halacha skip with "no_issue".
if not skip_paperclip and rec["paperclip"].get("reachable"):
import asyncpg import asyncpg
c = await asyncpg.connect(PAPERCLIP_DSN, timeout=10) c = await asyncpg.connect(PAPERCLIP_DSN, timeout=10)
try: try:
if rec["paperclip"]["projects"]:
res = await c.execute(
"UPDATE projects SET name = replace(name, $1, $2), updated_at = now() "
"WHERE name LIKE $3",
old, new, f"%{old}%",
)
log(f" ✓ Paperclip projects: {res}")
res = await c.execute( res = await c.execute(
"UPDATE projects SET name = replace(name, $1, $2), updated_at = now() WHERE name LIKE $3", "UPDATE plugin_state SET value_json = to_jsonb($2::text) "
"WHERE state_key = 'legal-case-number' AND value_json = to_jsonb($1::text)",
old, new,
)
log(f" ✓ Paperclip plugin_state (legal-case-number): {res}")
res = await c.execute(
"UPDATE issues SET title = replace(title, $1, $2) WHERE title LIKE $3",
old, new, f"%{old}%", old, new, f"%{old}%",
) )
log(f" ✓ Paperclip projects: {res}") log(f" ✓ Paperclip issue titles: {res}")
finally: finally:
await c.close() await c.close()
@@ -419,6 +455,8 @@ def print_inspection(rec: dict) -> None:
log(f" pclip: {name}") log(f" pclip: {name}")
if not pc["projects"]: if not pc["projects"]:
log(" pclip: (no matching project)") log(" pclip: (no matching project)")
log(f" pclip: linkage rows={pc.get('linkage_rows', 0)} "
f"issue titles={pc.get('issue_titles', 0)} (→ rewritten to {rec['new']})")
else: else:
log(f" pclip: unreachable ({pc.get('error','')})") log(f" pclip: unreachable ({pc.get('error','')})")
log(" DB path columns to rewrite:") log(" DB path columns to rewrite:")

View File

@@ -110,6 +110,19 @@ def _category(change: dict) -> str:
return "style" return "style"
# Graduated gate (INV-LRN1, chair decision 2026-06-28): a STYLE lesson the panel
# kept by 2/2 consensus flows straight to the writer (review_status='approved'),
# reversibly — the chair can veto it in /training. SUBSTANCE (halacha/precedent/
# fact) never reaches here (it's filtered to `substance` and skipped, and routes
# through the strict 3-judge halacha gate), so every category this panel emits is
# style and auto-approves. The constant keeps the gate explicit and future-proof.
_STYLE_CATEGORIES = frozenset({"style", "structure", "lexicon", "tabular"})
def _review_status_for(category: str) -> str:
return "approved" if category in _STYLE_CATEGORIES else "proposed"
# ── two judges, one signature: (system, user) -> dict|None ── # ── two judges, one signature: (system, user) -> dict|None ──
async def judge_deepseek(client: httpx.AsyncClient, system: str, user: str) -> dict | None: async def judge_deepseek(client: httpx.AsyncClient, system: str, user: str) -> dict | None:
@@ -323,18 +336,24 @@ async def main(args: argparse.Namespace) -> int:
_lesson_text(r["_change"])]) _lesson_text(r["_change"])])
written = 0 written = 0
approved = 0
for r in fresh: for r in fresh:
cat = _category(r["_change"])
rs = _review_status_for(cat)
await db.add_decision_lesson( await db.add_decision_lesson(
UUID(corpus_id), UUID(corpus_id),
lesson_text=_lesson_text(r["_change"]), lesson_text=_lesson_text(r["_change"]),
category=_category(r["_change"]), category=cat,
source="panel:deepseek+gemini", source="panel:deepseek+gemini",
created_by="panel", created_by="panel",
review_status=rs,
) )
written += 1 written += 1
approved += (rs == "approved")
chair = cc["split"] + cc["incomplete"] chair = cc["split"] + cc["incomplete"]
print(f"\nAPPLIED (reversible): wrote {written} decision_lesson proposals " print(f"\nAPPLIED (reversible): wrote {written} decision_lessons "
f"({approved} auto-approved style → writer; graduated gate) "
f"(source=panel:deepseek+gemini) · {skipped_dup} כפילויות דולגו · " f"(source=panel:deepseek+gemini) · {skipped_dup} כפילויות דולגו · "
f"{chair} escalated to chair · {len(substance)} substance skipped") f"{chair} escalated to chair · {len(substance)} substance skipped")
print(f"backup → {backup}") print(f"backup → {backup}")

View File

@@ -0,0 +1,170 @@
# Advanced Features — פיצ'רים מתקדמים
## הערות שוליים (Footnotes)
**שימוש מרכזי:** הפניות לחקיקה ופסיקה.
```javascript
const { FootnoteReferenceRun } = require('docx');
const doc = new Document({
footnotes: {
1: { children: [new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
children: [new TextRun({
text: "חוק החוזים (חלק כללי), התשל״ג-1973, סעיף 12.",
font: "David", size: 20, rightToLeft: true
})]
})] },
},
// ...sections
});
// הפניה בגוף הטקסט:
new Paragraph({
bidirectional: true, alignment: AlignmentType.BOTH,
children: [
new TextRun({ text: "חובת תום הלב", font: "David", size: 24, rightToLeft: true }),
new FootnoteReferenceRun(1),
new TextRun({ text: " חלה על כל שלבי המשא ומתן.", font: "David", size: 24, rightToLeft: true }),
]
})
```
### תיקון RTL בהערות שוליים (post-unpack)
docx-js לא מגדיר RTL מלא. אחרי unpack, תקן ב-`word/footnotes.xml`:
```xml
<w:footnote w:id="1">
<w:p>
<w:pPr>
<w:pStyle w:val="FootnoteText"/>
<w:bidi/>
<w:jc w:val="start"/>
</w:pPr>
<w:r>
<w:rPr>
<w:rStyle w:val="FootnoteReference"/>
<w:rtl/>
</w:rPr>
<w:footnoteRef/>
</w:r>
</w:p>
</w:footnote>
```
---
## תוכן עניינים (TOC)
**⚠️ TOC ידני בלבד** — `TableOfContents` של docx-js מאבד הגדרות RTL בעדכון Word.
```javascript
const { Tab, TabStopType, LeaderType, LineRuleType } = require('docx');
const tocEntry = (text, pageNum, opts = {}) => new Paragraph({
bidirectional: true,
spacing: { after: 60, line: 276, lineRule: LineRuleType.AUTO },
...(opts.indent ? { indent: { right: opts.indent } } : {}),
tabStops: [{ type: TabStopType.RIGHT, position: 9026, leader: LeaderType.DOT }],
children: [
new TextRun({ text, font: "David", size: 24, rightToLeft: true, bold: opts.bold || false }),
new TextRun({ children: [new Tab()], font: "David", rightToLeft: true }),
new TextRun({ text: String(pageNum), font: "David", size: 24, rightToLeft: true }),
]
});
// שימוש:
tocEntry("פרק א׳ — הגדרות כלליות", 2, { bold: true }),
tocEntry("1. הגדרות יסוד", 2, { indent: 400 }),
```
---
## מספר סקשנים (Multiple Sections)
**שימוש:** כותרות שונות לנספחים, שוליים שונים.
```javascript
const doc = new Document({
sections: [
{
properties: {
page: { size: { width: 11906, height: 16838 }, margin: { top: 1417, right: 1417, bottom: 1417, left: 1417 } },
bidi: true,
},
headers: {
default: new Header({ children: [new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "הסכם שירותים", font: "David", size: 20, bold: true, rightToLeft: true })]
})] })
},
children: [ /* גוף ההסכם */ ]
},
{
properties: {
page: { size: { width: 11906, height: 16838 }, margin: { top: 1417, right: 1417, bottom: 1417, left: 1417 } },
bidi: true,
},
headers: {
default: new Header({ children: [new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
children: [new TextRun({ text: "נספח א׳ — לוח תעריפים", font: "David", size: 20, bold: true, rightToLeft: true })]
})] })
},
children: [ /* הנספח */ ]
}
]
});
```
---
## לוגו/תמונה בכותרת (Letterhead)
```javascript
const { ImageRun } = require('docx');
const logoBuffer = fs.readFileSync('/path/to/logo.png');
headers: {
default: new Header({
children: [
new Paragraph({
alignment: AlignmentType.CENTER,
children: [new ImageRun({ data: logoBuffer, transformation: { width: 200, height: 60 }, type: "png" })],
}),
new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "משרד עורכי דין ישראלי ושות׳", font: "David", size: 20, bold: true, rightToLeft: true })],
}),
],
}),
}
```
**הערה:** תמונה חייבת להיות קובץ אמיתי — לבקש מהמשתמש אם אין.
---
## היפרלינקים
```javascript
const { ExternalHyperlink, UnderlineType } = require('docx');
new Paragraph({
bidirectional: true,
children: [
new TextRun({ text: "ראה: ", font: "David", size: 24, rightToLeft: true }),
new ExternalHyperlink({
link: "https://www.nevo.co.il/law_html/law01/073_002.htm",
children: [new TextRun({
text: "חוק החוזים באתר נבו",
font: "David", size: 24, rightToLeft: true,
color: "0563C1",
underline: { type: UnderlineType.SINGLE },
})],
}),
]
})
```
**⚠️ אל תשתמש ב-`style: "Hyperlink"`** — מפריע ל-RTL. הגדר `color` + `underline` ידנית.

View File

@@ -0,0 +1,219 @@
# Document Templates — תבניות מסמכים משפטיים
## תבנית 1: כתב טענות (בקשה, תביעה, הגנה, ערעור)
```javascript
const { Document, Packer, Paragraph, TextRun, Table, TableRow, TableCell,
AlignmentType, LevelFormat, BorderStyle, WidthType } = require('docx');
const PAGE_WIDTH = 11906;
const MARGINS = { top: 1134, right: 1134, bottom: 1134, left: 1134 };
const CONTENT_WIDTH = PAGE_WIDTH - MARGINS.left - MARGINS.right;
const noBorder = { style: BorderStyle.NONE, size: 0, color: "FFFFFF" };
const noBorders = { top: noBorder, bottom: noBorder, left: noBorder, right: noBorder };
// Header בית משפט — טבלה עם שם בית המשפט (ימין) ומספר תיק (שמאל)
function courtHeader(courtName, caseNumber) {
return new Table({
width: { size: CONTENT_WIDTH, type: WidthType.DXA },
columnWidths: [CONTENT_WIDTH / 2, CONTENT_WIDTH / 2],
visuallyRightToLeft: true,
rows: [
new TableRow({
children: [
new TableCell({
width: { size: CONTENT_WIDTH / 2, type: WidthType.DXA },
borders: noBorders,
children: [new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
children: [new TextRun({ text: courtName, bold: true, font: "David", size: 26, rightToLeft: true })]
})]
}),
new TableCell({
width: { size: CONTENT_WIDTH / 2, type: WidthType.DXA },
borders: noBorders,
children: [new Paragraph({
bidirectional: true, alignment: AlignmentType.END,
children: [new TextRun({ text: caseNumber, bold: true, font: "David", size: 26, rightToLeft: true })]
})]
})
]
})
]
});
}
function mainTitle(text) {
return new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
spacing: { before: 300, after: 300 },
children: [new TextRun({ text, bold: true, font: "David", size: 28, rightToLeft: true, underline: {} })]
});
}
function subHeading(text) {
return new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
spacing: { before: 240, after: 120 },
children: [new TextRun({ text, bold: true, font: "David", size: 24, rightToLeft: true, underline: {} })]
});
}
const doc = new Document({
numbering: {
config: [{
reference: "legal-clauses",
levels: [{
level: 0, format: LevelFormat.DECIMAL, text: "%1.",
alignment: AlignmentType.START, suffix: "tab",
style: { paragraph: { indent: { left: 360, hanging: 360 } } }
}]
}]
},
sections: [{
properties: {
page: { size: { width: PAGE_WIDTH, height: 16838 }, margin: MARGINS },
bidi: true
},
children: [
courtHeader("בית המשפט המחוזי בתל אביב", "ת\"א 12345-01-26"),
mainTitle("כתב תביעה"),
// ... פרטי צדדים, סעיפים, חתימה
]
}]
});
```
---
## תבנית 2: מכתב התראה
```javascript
function letterHeader(firmName, address, phone, email) {
return [
new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
children: [new TextRun({ text: firmName, bold: true, font: "David", size: 28, rightToLeft: true })]
}),
new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
children: [new TextRun({ text: address, font: "David", size: 22, rightToLeft: true })]
}),
new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
spacing: { after: 300 },
children: [new TextRun({ text: `טל': ${phone} | ${email}`, font: "David", size: 22, rightToLeft: true })]
}),
];
}
function subjectLine(text) {
return new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
spacing: { before: 200, after: 200 },
children: [
new TextRun({ text: "הנדון: ", bold: true, font: "David", size: 24, rightToLeft: true }),
new TextRun({ text, bold: true, font: "David", size: 24, rightToLeft: true, underline: {} })
]
});
}
// שימוש:
sections: [{
properties: { page: { size: { width: 11906, height: 16838 }, margin: { top: 1417, right: 1417, bottom: 1417, left: 1417 } }, bidi: true },
children: [
...letterHeader("משרד עו\"ד כהן ושות'", "רח' הרצל 1, תל אביב", "03-1234567", "office@cohen-law.co.il"),
new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
children: [new TextRun({ text: "תאריך: 10.2.2026", font: "David", size: 24, rightToLeft: true })]
}),
new Paragraph({
bidirectional: true, alignment: AlignmentType.START,
spacing: { before: 200 },
children: [new TextRun({ text: "לכבוד: [שם הנמען]", font: "David", size: 24, rightToLeft: true })]
}),
subjectLine("התראה בטרם נקיטת הליכים משפטיים"),
]
}]
```
---
## תבנית 3: הסכם/חוזה
```javascript
const CONTENT_WIDTH = 9638; // A4 עם שוליים 2.5 ס
const noBorders = { /* ראה תבנית 1 */ };
function contractTitle(text) {
return new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
spacing: { after: 300 },
children: [new TextRun({ text, bold: true, font: "David", size: 32, rightToLeft: true })]
});
}
function partyClause(label, name, id, address, alias) {
return new Paragraph({
bidirectional: true, alignment: AlignmentType.BOTH,
spacing: { after: 120 },
children: [
new TextRun({ text: `${label}: `, bold: true, font: "David", size: 24, rightToLeft: true }),
new TextRun({ text: `${name}, ח.פ./ת.ז. ${id}, מ${address} (להלן: "`, font: "David", size: 24, rightToLeft: true }),
new TextRun({ text: alias, bold: true, font: "David", size: 24, rightToLeft: true }),
new TextRun({ text: '")', font: "David", size: 24, rightToLeft: true }),
]
});
}
function signatureTable(contentWidth) {
return new Table({
width: { size: contentWidth, type: WidthType.DXA },
columnWidths: [contentWidth / 2, contentWidth / 2],
visuallyRightToLeft: true,
rows: [new TableRow({
children: [
new TableCell({
borders: noBorders,
children: [
new Paragraph({ bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "_________________", font: "David", size: 24, rightToLeft: true })] }),
new Paragraph({ bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "צד א'", font: "David", size: 24, rightToLeft: true })] })
]
}),
new TableCell({
borders: noBorders,
children: [
new Paragraph({ bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "_________________", font: "David", size: 24, rightToLeft: true })] }),
new Paragraph({ bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "צד ב'", font: "David", size: 24, rightToLeft: true })] })
]
})
]
})]
});
}
sections: [{
properties: { page: { size: { width: 11906, height: 16838 }, margin: { top: 1417, right: 1417, bottom: 1417, left: 1417 } }, bidi: true },
children: [
contractTitle("הסכם שירותים"),
new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
children: [new TextRun({ text: "נערך ונחתם בתל אביב ביום __________", font: "David", size: 24, rightToLeft: true })]
}),
partyClause("מצד אחד", "[שם]", "[מספר]", "[כתובת]", "המזמין"),
partyClause("מצד שני", "[שם]", "[מספר]", "[כתובת]", "הספק"),
// הואילים + סעיפים...
new Paragraph({
bidirectional: true, alignment: AlignmentType.CENTER,
spacing: { before: 400, after: 300 },
children: [new TextRun({ text: "ולראיה באו הצדדים על החתום:", bold: true, font: "David", size: 24, rightToLeft: true })]
}),
signatureTable(CONTENT_WIDTH)
]
}]
```

View File

@@ -0,0 +1,57 @@
# Tracked Changes — עקוב אחר שינויים
## שם מחבר בעברית
```xml
<w:del w:id="10" w:author="עו&quot;ד כהן" w:date="2026-02-06T09:00:00Z">
```
## שינוי ערך (סכום, תאריך, תקופה)
פצל את הטקסט ועטוף רק את הערך שמשתנה:
```xml
<w:r><w:rPr>...RTL PROPS...</w:rPr>
<w:t xml:space="preserve">שכר הטרחה יעמוד על סך של </w:t></w:r>
<w:del w:id="10" w:author="עו&quot;ד כהן" w:date="...">
<w:r><w:rPr>...RTL PROPS...</w:rPr><w:delText>750</w:delText></w:r>
</w:del>
<w:ins w:id="11" w:author="עו&quot;ד כהן" w:date="...">
<w:r><w:rPr>...RTL PROPS...</w:rPr><w:t>850</w:t></w:r>
</w:ins>
<w:r><w:rPr>...RTL PROPS...</w:rPr>
<w:t xml:space="preserve"> ש״ח לשעת עבודה</w:t></w:r>
```
## מחיקת סעיף שלם
```xml
<w:p>
<w:pPr>
<w:bidi/>
<w:jc w:val="both"/>
<w:rPr>
<w:del w:id="20" w:author="עו&quot;ד כהן" w:date="..."/>
</w:rPr>
</w:pPr>
<w:del w:id="21" w:author="עו&quot;ד כהן" w:date="...">
<w:r><w:rPr>...RTL PROPS...</w:rPr>
<w:delText>הסעיף שנמחק</w:delText></w:r>
</w:del>
</w:p>
```
## RTL PROPS — בלוק rPr מלא לכל run
```xml
<w:rPr>
<w:rFonts w:ascii="David" w:cs="David" w:eastAsia="David" w:hAnsi="David"/>
<w:sz w:val="24"/>
<w:szCs w:val="24"/>
<w:rtl/>
</w:rPr>
```
## קבלה/דחייה של שינויים
| פעולה | לפני | אחרי |
|-------|------|------|
| קבלת הוספה | `<w:ins ...><w:r>...<w:t>טקסט</w:t></w:r></w:ins>` | `<w:r>...<w:t>טקסט</w:t></w:r>` |
| דחיית הוספה | `<w:ins ...><w:r>...</w:r></w:ins>` | *(מחק הכל)* |
| קבלת מחיקה | `<w:del ...><w:r>...<w:delText>טקסט</w:delText></w:r></w:del>` | *(מחק הכל)* |
| דחיית מחיקה | `<w:del ...><w:r>...<w:delText>טקסט</w:delText></w:r></w:del>` | `<w:r>...<w:t>טקסט</w:t></w:r>` |

View File

@@ -9,10 +9,15 @@ import { Button } from "@/components/ui/button";
import { Skeleton } from "@/components/ui/skeleton"; import { Skeleton } from "@/components/ui/skeleton";
import { SubsectionCard } from "@/components/compose/subsection-card"; import { SubsectionCard } from "@/components/compose/subsection-card";
import { PrecedentsSection } from "@/components/compose/precedents-section"; import { PrecedentsSection } from "@/components/compose/precedents-section";
import { CitationVerificationPanel } from "@/components/compose/citation-verification-panel";
import { DecisionBlocksPanel } from "@/components/cases/decision-blocks-panel"; import { DecisionBlocksPanel } from "@/components/cases/decision-blocks-panel";
import { Markdown } from "@/components/ui/markdown"; import { Markdown } from "@/components/ui/markdown";
import { Tabs, TabsContent, TabsList, TabsTrigger } from "@/components/ui/tabs"; import { Tabs, TabsContent, TabsList, TabsTrigger } from "@/components/ui/tabs";
import { useCase, type CaseStatus } from "@/lib/api/cases"; import { useCase, type CaseStatus } from "@/lib/api/cases";
import {
useCaseLearningStatus,
type CaseLearningStatus,
} from "@/lib/api/learning";
import { useResearchAnalysis } from "@/lib/api/research"; import { useResearchAnalysis } from "@/lib/api/research";
import { useCasePrecedents } from "@/lib/api/precedents"; import { useCasePrecedents } from "@/lib/api/precedents";
import { APPEAL_SUBTYPES } from "@/lib/practice-area"; import { APPEAL_SUBTYPES } from "@/lib/practice-area";
@@ -51,6 +56,45 @@ function subtypeLabel(subtype?: string | null): string | null {
return APPEAL_SUBTYPES.find((s) => s.value === subtype)?.label ?? null; return APPEAL_SUBTYPES.find((s) => s.value === subtype)?.label ?? null;
} }
// ── Staged-pipeline indicator text (mockup 03) — derived from the live
// learning-status, same source as the drafts-panel LearningStatusBadges. ──────
function voiceLearningText(s?: CaseLearningStatus): string {
if (!s?.final_uploaded) return "ממתין להעלאת הסופי";
const v = s.voice_learning;
if (v.outcome === "succeeded") {
const bits = [`${v.lessons_count} לקחים הופקו`];
if (v.lessons_proposed > 0) bits.push(`${v.lessons_proposed} הוצעו לאישור`);
return `✓ הושלם · ${bits.join(" · ")}`;
}
if (v.outcome === "failed") return v.error ? `✗ נכשל — ${v.error}` : "✗ נכשל";
return "ממתין להרצה";
}
function halachaExtractionText(s?: CaseLearningStatus): string {
if (!s?.final_uploaded) return "ממתין להעלאת הסופי";
const h = s.halacha_extraction;
if (!h.enrolled_in_corpus)
return h.not_enrolled_reason ?? "לא נכנס לקורפוס-הפסיקה";
switch (h.status) {
case "completed":
return `✓ הושלם · חולצו ${h.halachot_count} · ${h.approved} אושרו · ${h.rejected} נדחו`;
case "processing":
return "רץ עכשיו…";
case "pending":
case "busy":
return "בתור";
case "partial":
return `חלקי · חולצו ${h.halachot_count}`;
case "failed":
case "extraction_failed":
return "✗ נכשל";
case "no_chunks":
return "אין טקסט לחילוץ";
default:
return "ממתין להרצה";
}
}
function ProseSection({ title, content }: { title: string; content?: string }) { function ProseSection({ title, content }: { title: string; content?: string }) {
if (!content?.trim()) return null; if (!content?.trim()) return null;
return ( return (
@@ -76,6 +120,7 @@ function FinishRail({
const fileRef = useRef<HTMLInputElement>(null); const fileRef = useRef<HTMLInputElement>(null);
const [uploading, setUploading] = useState(false); const [uploading, setUploading] = useState(false);
const [uploadMsg, setUploadMsg] = useState<{ ok: boolean; text: string } | null>(null); const [uploadMsg, setUploadMsg] = useState<{ ok: boolean; text: string } | null>(null);
const learning = useCaseLearningStatus(caseNumber);
async function handleUpload(file: File) { async function handleUpload(file: File) {
setUploading(true); setUploading(true);
@@ -167,10 +212,10 @@ function FinishRail({
{/* mockup 03: stage indicators — informational pointers, not actions */} {/* mockup 03: stage indicators — informational pointers, not actions */}
<div className="mt-3 space-y-0"> <div className="mt-3 space-y-0">
<div className="text-[0.78rem] text-ink-muted pt-2 border-t border-rule-soft"> <div className="text-[0.78rem] text-ink-muted pt-2 border-t border-rule-soft">
<b className="text-navy">הרץ למידת-קול</b> ממתין להעלאת הסופי <b className="text-navy">הרץ למידת-קול</b> {voiceLearningText(learning.data)}
</div> </div>
<div className="text-[0.78rem] text-ink-muted pt-2 mt-2 border-t border-rule-soft"> <div className="text-[0.78rem] text-ink-muted pt-2 mt-2 border-t border-rule-soft">
<b className="text-navy">הרץ אימות-הלכות</b> ממתין להעלאת הסופי <b className="text-navy">הרץ אימות-הלכות</b> {halachaExtractionText(learning.data)}
</div> </div>
</div> </div>
@@ -236,6 +281,12 @@ export default function ComposePage({
<span aria-hidden>·</span> <span aria-hidden>·</span>
<span className="text-navy">עורך החלטה</span> <span className="text-navy">עורך החלטה</span>
</nav> </nav>
<Link
href={`/cases/${caseNumber}`}
className="inline-flex items-center gap-1.5 mb-3 rounded-md border border-rule bg-surface px-3 py-1.5 text-[0.82rem] font-semibold text-gold-deep hover:bg-gold-wash hover:text-navy transition-colors"
>
<span aria-hidden></span> חזרה לדף התיק
</Link>
<div className="flex items-center gap-3 flex-wrap"> <div className="flex items-center gap-3 flex-wrap">
<h1 className="text-navy text-2xl font-bold mb-0">ערר {caseNumber}</h1> <h1 className="text-navy text-2xl font-bold mb-0">ערר {caseNumber}</h1>
<StatusChip status={caseQuery.data?.status} /> <StatusChip status={caseQuery.data?.status} />
@@ -270,6 +321,7 @@ export default function ComposePage({
<TabsList className="bg-rule-soft/60"> <TabsList className="bg-rule-soft/60">
<TabsTrigger value="blocks">עורך הבלוקים</TabsTrigger> <TabsTrigger value="blocks">עורך הבלוקים</TabsTrigger>
<TabsTrigger value="positions">עמדות וטענות</TabsTrigger> <TabsTrigger value="positions">עמדות וטענות</TabsTrigger>
<TabsTrigger value="verify">אימות פסיקה</TabsTrigger>
</TabsList> </TabsList>
{/* Tab 1 — the 12-block decision editor (reused DecisionBlocksPanel) */} {/* Tab 1 — the 12-block decision editor (reused DecisionBlocksPanel) */}
@@ -277,6 +329,11 @@ export default function ComposePage({
<DecisionBlocksPanel caseNumber={caseNumber} /> <DecisionBlocksPanel caseNumber={caseNumber} />
</TabsContent> </TabsContent>
{/* Tab 3 — citation verification: per-argument support + verify gate (#154) */}
<TabsContent value="verify" className="mt-5">
<CitationVerificationPanel caseNumber={caseNumber} />
</TabsContent>
{/* Tab 2 — chair positions on the analyst's threshold-claims + issues */} {/* Tab 2 — chair positions on the analyst's threshold-claims + issues */}
<TabsContent value="positions" className="mt-5"> <TabsContent value="positions" className="mt-5">
{analysis.isPending ? ( {analysis.isPending ? (

View File

@@ -39,14 +39,16 @@ export default function CaseDetailPage({
params: Promise<{ caseNumber: string }>; params: Promise<{ caseNumber: string }>;
}) { }) {
const { caseNumber } = use(params); const { caseNumber } = use(params);
const { data, isPending, error } = useCase(caseNumber); const { data, isPending, error, refetch } = useCase(caseNumber);
const startWorkflow = useStartWorkflow(caseNumber); const startWorkflow = useStartWorkflow(caseNumber);
const canStartWorkflow = data?.status === "new" || data?.status === "documents_ready"; const canStartWorkflow = data?.status === "new" || data?.status === "documents_ready";
const expectedOutcomeLabel = data?.expected_outcome const expectedOutcomeLabel = data?.expected_outcome
? EXPECTED_OUTCOME_LABELS[data.expected_outcome] ?? data.expected_outcome ? EXPECTED_OUTCOME_LABELS[data.expected_outcome] ?? data.expected_outcome
: null; : null;
if (error) { // Only take over the whole page when there is NO data to show. A transient
// 5xx on the 5s background refetch must not blow away an already-loaded page.
if (error && !data) {
return ( return (
<AppShell> <AppShell>
<section className="space-y-6"> <section className="space-y-6">
@@ -54,9 +56,14 @@ export default function CaseDetailPage({
<CardContent className="px-6 py-6 text-center space-y-3"> <CardContent className="px-6 py-6 text-center space-y-3">
<p className="text-danger font-semibold">שגיאה בטעינת התיק</p> <p className="text-danger font-semibold">שגיאה בטעינת התיק</p>
<p className="text-sm text-ink-muted">{error.message}</p> <p className="text-sm text-ink-muted">{error.message}</p>
<Button asChild variant="outline"> <div className="flex items-center justify-center gap-2">
<Link href="/">חזרה לרשימת התיקים</Link> <Button variant="outline" onClick={() => refetch()}>
</Button> נסה שוב
</Button>
<Button asChild variant="ghost">
<Link href="/">חזרה לרשימת התיקים</Link>
</Button>
</div>
</CardContent> </CardContent>
</Card> </Card>
</section> </section>
@@ -91,6 +98,9 @@ export default function CaseDetailPage({
<> <>
{data && <CaseEditDialog data={data} />} {data && <CaseEditDialog data={data} />}
<UploadSheet caseNumber={caseNumber} /> <UploadSheet caseNumber={caseNumber} />
<Button asChild className="bg-gold text-white hover:bg-gold-deep border-transparent">
<Link href={`/cases/${caseNumber}/compose`}>פתח עורך החלטה</Link>
</Button>
{canStartWorkflow && ( {canStartWorkflow && (
<Button <Button
className="bg-gold-deep hover:bg-gold-deep/90 text-parchment" className="bg-gold-deep hover:bg-gold-deep/90 text-parchment"
@@ -135,16 +145,7 @@ export default function CaseDetailPage({
<DocumentsPanel data={data} /> <DocumentsPanel data={data} />
<AgentActivityPreview caseNumber={caseNumber} /> <AgentActivityPreview caseNumber={caseNumber} />
{/* decision-editor CTA moved to the band actions (visible on all tabs) */}
{/* gold CTA — open the decision editor */}
<Button
asChild
className="w-full bg-gold text-white hover:bg-gold-deep border-transparent py-6 text-base font-semibold"
>
<Link href={`/cases/${caseNumber}/compose`}>
פתח עורך החלטה
</Link>
</Button>
</TabsContent> </TabsContent>
<TabsContent value="arguments" className="mt-0"> <TabsContent value="arguments" className="mt-0">

View File

@@ -22,12 +22,12 @@ import { MissingPrecedentsTable } from "@/components/missing-precedents/missing-
type StatusFilter = MissingPrecedentStatus | "all"; type StatusFilter = MissingPrecedentStatus | "all";
// Only the two states the chair acts on: open gaps to fill, and closed gaps for
// reference. "הועלה" (transient) and "לא-רלוונטי" were dropped from the filter,
// and "הכל" with them (chair's request, design 09).
const STATUS_CHIPS: { value: StatusFilter; label: string }[] = [ const STATUS_CHIPS: { value: StatusFilter; label: string }[] = [
{ value: "open", label: "פתוח" }, { value: "open", label: "פתוח" },
{ value: "uploaded", label: "הועלה" },
{ value: "closed", label: "נסגר" }, { value: "closed", label: "נסגר" },
{ value: "irrelevant", label: "לא-רלוונטי" },
{ value: "all", label: "הכל" },
]; ];
export default function MissingPrecedentsPage() { export default function MissingPrecedentsPage() {
@@ -145,15 +145,13 @@ export default function MissingPrecedentsPage() {
<div className="rounded-lg border border-rule bg-parchment px-5 py-3.5 text-[0.82rem] text-ink-muted leading-7"> <div className="rounded-lg border border-rule bg-parchment px-5 py-3.5 text-[0.82rem] text-ink-muted leading-7">
<b className="text-ink-soft">מחזור-חיים:</b>{" "} <b className="text-ink-soft">מחזור-חיים:</b>{" "}
<LifecycleChip tone="open">פתוח</LifecycleChip> {" "} <LifecycleChip tone="open">פתוח</LifecycleChip> {" "}
<LifecycleChip tone="up">הועלה</LifecycleChip> {" "}
<LifecycleChip tone="closed">נסגר</LifecycleChip>. פריט נפתח אוטומטית <LifecycleChip tone="closed">נסגר</LifecycleChip>. פריט נפתח אוטומטית
בעת חילוץ ציטוט שאין לו תקדים בקורפוס; בהעלאת פסק-הדין הוא מקושר לרשומת בעת חילוץ ציטוט שאין לו תקדים בקורפוס; בהעלאת פסק-הדין הוא מקושר לרשומת
הפסיקה דרך{" "} הפסיקה דרך{" "}
<code className="rounded border border-rule bg-surface px-1.5 py-0.5 text-[0.75rem] text-gold-deep" dir="ltr"> <code className="rounded border border-rule bg-surface px-1.5 py-0.5 text-[0.75rem] text-gold-deep" dir="ltr">
linked_case_law_id linked_case_law_id
</code>{" "} </code>{" "}
ונסגר. פריט שאינו רלוונטי מסומן{" "} ונסגר.
<LifecycleChip tone="na">לא-רלוונטי</LifecycleChip> מבלי שתידרש העלאה.
</div> </div>
</section> </section>
</AppShell> </AppShell>

View File

@@ -12,7 +12,6 @@ import { Switch } from "@/components/ui/switch";
import { Input } from "@/components/ui/input"; import { Input } from "@/components/ui/input";
import { Label } from "@/components/ui/label"; import { Label } from "@/components/ui/label";
import { Skeleton } from "@/components/ui/skeleton"; import { Skeleton } from "@/components/ui/skeleton";
import { ScrollArea } from "@/components/ui/scroll-area";
import { import {
Dialog, Dialog,
DialogContent, DialogContent,
@@ -673,11 +672,11 @@ function RunLogDialog({ run, onClose }: { run: AgentRun | null; onClose: () => v
) : error ? ( ) : error ? (
<p className="text-sm text-destructive">שגיאה בטעינת הלוג: {String(error)}</p> <p className="text-sm text-destructive">שגיאה בטעינת הלוג: {String(error)}</p>
) : ( ) : (
<ScrollArea className="h-[60vh] rounded-md border border-rule-soft bg-rule-soft/20 p-3"> <div className="h-[60vh] w-full overflow-y-auto overflow-x-hidden rounded-md border border-rule-soft bg-rule-soft/20 p-3">
<pre dir="ltr" className="text-[0.72rem] leading-relaxed whitespace-pre-wrap break-words text-navy text-start"> <pre dir="ltr" className="w-full min-w-0 max-w-full overflow-x-hidden text-[0.72rem] leading-relaxed whitespace-pre-wrap break-all text-navy text-start">
{text || "אין פלט עדיין."} {text || "אין פלט עדיין."}
</pre> </pre>
</ScrollArea> </div>
)} )}
</DialogContent> </DialogContent>
</Dialog> </Dialog>

View File

@@ -69,8 +69,10 @@ export default function HomePage() {
{/* KPI row — mockup 04 .kpis (4-up, gold-washed "ממתינים לאישור") */} {/* KPI row — mockup 04 .kpis (4-up, gold-washed "ממתינים לאישור") */}
<KPICards cases={data} loading={isPending} /> <KPICards cases={data} loading={isPending} />
{/* two-column body — main flow + narrow gold gate rail (mockup 04 .cols) */} {/* two-column body — main flow + narrow gold gate rail (mockup 04 .cols).
<div className="grid gap-6 lg:grid-cols-[1fr_360px]"> Rail trimmed 360→280 so the cases table gets the width it needs and
no longer needs a horizontal scrollbar (chair request). */}
<div className="grid gap-6 lg:grid-cols-[1fr_280px]">
<div className="space-y-6 min-w-0"> <div className="space-y-6 min-w-0">
{/* "תיקים לפי סטטוס" (פסים אופקיים) הוסר — פיזור-הסטטוסים מוצג {/* "תיקים לפי סטטוס" (פסים אופקיים) הוסר — פיזור-הסטטוסים מוצג
בדונאט "פיזור סטטוסים" בטור-הצד (#1). */} בדונאט "פיזור סטטוסים" בטור-הצד (#1). */}

View File

@@ -236,7 +236,11 @@ export default function PrecedentDetailPage({
{/* side rail — citations + corroboration */} {/* side rail — citations + corroboration */}
<div className="space-y-4"> <div className="space-y-4">
<RelatedCasesSection caseId={id} related={data.related_cases ?? []} /> <RelatedCasesSection
caseId={id}
related={data.related_cases ?? []}
incoming={data.incoming_citations ?? []}
/>
</div> </div>
</div> </div>
</div> </div>

View File

@@ -1,6 +1,6 @@
"use client"; "use client";
import { useRef, useState, useEffect } from "react"; import { useRef, useState, useEffect, useMemo } from "react";
import { Button } from "@/components/ui/button"; import { Button } from "@/components/ui/button";
import { Textarea } from "@/components/ui/textarea"; import { Textarea } from "@/components/ui/textarea";
import { Badge } from "@/components/ui/badge"; import { Badge } from "@/components/ui/badge";
@@ -151,7 +151,7 @@ function CommentCard({
{identifier} {identifier}
</Badge> </Badge>
)} )}
<span className="text-[11px] text-ink-faint mr-auto flex items-center gap-1"> <span className="text-[11px] text-ink-faint me-auto flex items-center gap-1">
<Clock className="w-3 h-3" /> <Clock className="w-3 h-3" />
{timeAgo(comment.created_at)} {timeAgo(comment.created_at)}
</span> </span>
@@ -267,7 +267,7 @@ function AskUserQuestionsForm({
<fieldset key={q.id} className="space-y-2"> <fieldset key={q.id} className="space-y-2">
<legend className="text-sm font-semibold text-navy mb-1"> <legend className="text-sm font-semibold text-navy mb-1">
{q.prompt} {q.prompt}
{(q.required ?? true) && <span className="text-rose-600 mr-1">*</span>} {(q.required ?? true) && <span className="text-rose-600 ms-1">*</span>}
</legend> </legend>
<div className="space-y-1.5"> <div className="space-y-1.5">
{q.options.map((opt) => { {q.options.map((opt) => {
@@ -316,7 +316,7 @@ function AskUserQuestionsForm({
{pending ? ( {pending ? (
<Loader2 className="w-4 h-4 animate-spin" /> <Loader2 className="w-4 h-4 animate-spin" />
) : ( ) : (
<Send className="w-4 h-4 ml-1" /> <Send className="w-4 h-4 me-1" />
)} )}
{interaction.payload.submitLabel || "שלח תשובה"} {interaction.payload.submitLabel || "שלח תשובה"}
</Button> </Button>
@@ -397,14 +397,14 @@ function RequestConfirmationForm({
onClick={handleReject} onClick={handleReject}
disabled={pending || (requireReason && !reason.trim())} disabled={pending || (requireReason && !reason.trim())}
> >
<XCircle className="w-4 h-4 ml-1" /> <XCircle className="w-4 h-4 me-1" />
{rejectLabel} {rejectLabel}
</Button> </Button>
<Button size="sm" onClick={onAccept} disabled={pending}> <Button size="sm" onClick={onAccept} disabled={pending}>
{pending ? ( {pending ? (
<Loader2 className="w-4 h-4 animate-spin" /> <Loader2 className="w-4 h-4 animate-spin" />
) : ( ) : (
<CheckCircle2 className="w-4 h-4 ml-1" /> <CheckCircle2 className="w-4 h-4 me-1" />
)} )}
{acceptLabel} {acceptLabel}
</Button> </Button>
@@ -489,7 +489,7 @@ function SuggestTasksForm({
onClick={() => (showReason ? onReject(reason.trim()) : setShowReason(true))} onClick={() => (showReason ? onReject(reason.trim()) : setShowReason(true))}
disabled={pending} disabled={pending}
> >
<XCircle className="w-4 h-4 ml-1" /> <XCircle className="w-4 h-4 me-1" />
{showReason ? "אישור דחייה" : "דחייה"} {showReason ? "אישור דחייה" : "דחייה"}
</Button> </Button>
<Button <Button
@@ -500,7 +500,7 @@ function SuggestTasksForm({
{pending ? ( {pending ? (
<Loader2 className="w-4 h-4 animate-spin" /> <Loader2 className="w-4 h-4 animate-spin" />
) : ( ) : (
<CheckCircle2 className="w-4 h-4 ml-1" /> <CheckCircle2 className="w-4 h-4 me-1" />
)} )}
אישור משימות נבחרות ({selected.size}) אישור משימות נבחרות ({selected.size})
</Button> </Button>
@@ -575,7 +575,7 @@ function InteractionCard({
{identifier} {identifier}
</Badge> </Badge>
)} )}
<span className="text-[11px] text-ink-faint mr-auto flex items-center gap-1"> <span className="text-[11px] text-ink-faint me-auto flex items-center gap-1">
<Clock className="w-3 h-3" /> <Clock className="w-3 h-3" />
{timeAgo(interaction.resolved_at ?? interaction.created_at)} {timeAgo(interaction.resolved_at ?? interaction.created_at)}
</span> </span>
@@ -635,13 +635,13 @@ export function AgentActivityFeed({
const [body, setBody] = useState(""); const [body, setBody] = useState("");
const endRef = useRef<HTMLDivElement>(null); const endRef = useRef<HTMLDivElement>(null);
// Build issue_id → identifier map // Build issue_id → identifier map (memoized — the feed refetches every 10s,
const issueMap = new Map<string, string>(); // and a fresh Map each render would defeat the child cards' memoization).
if (data?.issues) { const issueMap = useMemo(() => {
for (const iss of data.issues) { const m = new Map<string, string>();
issueMap.set(iss.id, iss.identifier); for (const iss of data?.issues ?? []) m.set(iss.id, iss.identifier);
} return m;
} }, [data?.issues]);
// Auto-scroll on new comments or interactions // Auto-scroll on new comments or interactions
const commentCount = data?.comments?.length ?? 0; const commentCount = data?.comments?.length ?? 0;
@@ -669,7 +669,7 @@ export function AgentActivityFeed({
if (isLoading) { if (isLoading) {
return ( return (
<div className="flex items-center justify-center py-12 text-ink-faint"> <div className="flex items-center justify-center py-12 text-ink-faint">
<Loader2 className="w-5 h-5 animate-spin ml-2" /> <Loader2 className="w-5 h-5 animate-spin me-2" />
<span>טוען פעילות סוכנים...</span> <span>טוען פעילות סוכנים...</span>
</div> </div>
); );
@@ -777,6 +777,7 @@ export function AgentActivityFeed({
value={body} value={body}
onChange={(e) => setBody(e.target.value)} onChange={(e) => setBody(e.target.value)}
placeholder="כתוב הוראה לסוכנים..." placeholder="כתוב הוראה לסוכנים..."
aria-label="הוראה לסוכנים"
className="min-h-[60px] resize-none text-sm" className="min-h-[60px] resize-none text-sm"
dir="rtl" dir="rtl"
onKeyDown={(e) => { onKeyDown={(e) => {
@@ -798,7 +799,7 @@ export function AgentActivityFeed({
{sendComment.isPending ? ( {sendComment.isPending ? (
<Loader2 className="w-4 h-4 animate-spin" /> <Loader2 className="w-4 h-4 animate-spin" />
) : ( ) : (
<Send className="w-4 h-4 ml-1" /> <Send className="w-4 h-4 me-1" />
)} )}
שלח שלח
</Button> </Button>

View File

@@ -1,8 +1,20 @@
"use client"; "use client";
import { useAgentActivity } from "@/lib/api/agents"; import { useState } from "react";
import { useAgentActivity, useResetCaseAgents } from "@/lib/api/agents";
import type { PaperclipAgentStatus } from "@/lib/api/agents"; import type { PaperclipAgentStatus } from "@/lib/api/agents";
import { Bot } from "lucide-react"; import { Bot, RotateCcw, Loader2 } from "lucide-react";
import { toast } from "sonner";
import {
Dialog,
DialogContent,
DialogDescription,
DialogFooter,
DialogHeader,
DialogTitle,
DialogTrigger,
} from "@/components/ui/dialog";
import { Button } from "@/components/ui/button";
/* ── Status dot colors ───────────────────────────────────────── */ /* ── Status dot colors ───────────────────────────────────────── */
@@ -33,13 +45,82 @@ function AgentRow({ agent }: { agent: PaperclipAgentStatus }) {
className={`w-2 h-2 rounded-full flex-shrink-0 ${statusDot(agent.status)}`} className={`w-2 h-2 rounded-full flex-shrink-0 ${statusDot(agent.status)}`}
/> />
<span className="text-xs text-ink truncate">{agent.name}</span> <span className="text-xs text-ink truncate">{agent.name}</span>
<span className="text-[10px] text-ink-faint mr-auto"> <span className="text-[10px] text-ink-faint ms-auto">
{STATUS_LABEL[agent.status] ?? agent.status} {STATUS_LABEL[agent.status] ?? agent.status}
</span> </span>
</div> </div>
); );
} }
/* ── Reset button ─────────────────────────────────────────────── */
function ResetAgentsButton({ caseNumber }: { caseNumber: string }) {
const [open, setOpen] = useState(false);
const reset = useResetCaseAgents(caseNumber);
function handleConfirm() {
reset.mutate(undefined, {
onSuccess: (result) => {
setOpen(false);
const agentNames = result.reset_agents
.filter((a) => a.ok)
.map((a) => a.name)
.join(", ");
const msg = agentNames
? `סוכנים אופסו: ${agentNames}`
: "אופסו בהצלחה";
toast.success(msg);
},
onError: () => {
setOpen(false);
toast.error("שגיאה בביצוע האיפוס");
},
});
}
return (
<Dialog open={open} onOpenChange={setOpen}>
<DialogTrigger asChild>
<button
className="text-[10px] text-ink-faint hover:text-red-600 flex items-center gap-0.5 transition-colors"
title="אפס סוכנים תקועים"
>
<RotateCcw className="w-3 h-3" />
<span>אפס</span>
</button>
</DialogTrigger>
<DialogContent>
<DialogHeader>
<DialogTitle>איפוס סוכנים</DialogTitle>
<DialogDescription>
פעולה זו תאפס את כל הסוכנים במצב שגיאה ותחזיר issues פתוחים לניהול ידני.
פעולה הפיכה ניתן להפעיל את הסוכנים מחדש אחרי האיפוס.
</DialogDescription>
</DialogHeader>
<DialogFooter>
<Button variant="outline" onClick={() => setOpen(false)} disabled={reset.isPending}>
ביטול
</Button>
<Button
variant="destructive"
onClick={handleConfirm}
disabled={reset.isPending}
>
{reset.isPending ? (
<>
<Loader2 className="w-3.5 h-3.5 ms-1.5 animate-spin" />
מאפס...
</>
) : (
"אפס סוכנים"
)}
</Button>
</DialogFooter>
</DialogContent>
</Dialog>
);
}
/* ── Widget ───────────────────────────────────────────────────── */ /* ── Widget ───────────────────────────────────────────────────── */
export function AgentStatusWidget({ export function AgentStatusWidget({
@@ -54,6 +135,7 @@ export function AgentStatusWidget({
const agents = data.agents ?? []; const agents = data.agents ?? [];
const activeCount = agents.filter((a) => a.status === "active" || a.status === "running").length; const activeCount = agents.filter((a) => a.status === "active" || a.status === "running").length;
const hasErrors = agents.some((a) => a.status === "error");
return ( return (
<div className="space-y-2"> <div className="space-y-2">
@@ -62,11 +144,14 @@ export function AgentStatusWidget({
<Bot className="w-3.5 h-3.5" /> <Bot className="w-3.5 h-3.5" />
<span>סוכנים</span> <span>סוכנים</span>
</div> </div>
{agents.length > 0 && ( <div className="flex items-center gap-2">
<span className="text-[10px] text-ink-faint"> {hasErrors && <ResetAgentsButton caseNumber={caseNumber} />}
{activeCount} פעילים מתוך {agents.length} {agents.length > 0 && (
</span> <span className="text-[10px] text-ink-faint">
)} {activeCount} פעילים מתוך {agents.length}
</span>
)}
</div>
</div> </div>
<div className="space-y-0.5"> <div className="space-y-0.5">

View File

@@ -1,6 +1,6 @@
"use client"; "use client";
import { useEffect, useState } from "react"; import { useState } from "react";
import { useForm, Controller } from "react-hook-form"; import { useForm, Controller } from "react-hook-form";
import { zodResolver } from "@hookform/resolvers/zod"; import { zodResolver } from "@hookform/resolvers/zod";
import { toast } from "sonner"; import { toast } from "sonner";
@@ -54,22 +54,26 @@ export function CaseEditDialog({ data }: { data: CaseDetail }) {
}, },
}); });
/* Re-sync the form when the underlying case refetches after save */ /* Reset to the latest case values only on the open→true transition.
useEffect(() => { * Resetting on every `data` change would clobber in-progress edits, because
if (!open) return; * useCase refetches every 5s (refetchInterval) while the dialog is open. */
form.reset({ const handleOpenChange = (next: boolean) => {
title: data.title ?? "", if (next) {
subject: data.subject ?? "", form.reset({
hearing_date: data.hearing_date ?? "", title: data.title ?? "",
notes: "", subject: data.subject ?? "",
expected_outcome: data.expected_outcome ?? "", hearing_date: data.hearing_date ?? "",
appellants: data.appellants ?? [], notes: "",
respondents: data.respondents ?? [], expected_outcome: data.expected_outcome ?? "",
property_address: data.property_address ?? "", appellants: data.appellants ?? [],
permit_number: data.permit_number ?? "", respondents: data.respondents ?? [],
proceeding_type: data.proceeding_type ?? "ערר", property_address: data.property_address ?? "",
}); permit_number: data.permit_number ?? "",
}, [open, data, form]); proceeding_type: data.proceeding_type ?? "ערר",
});
}
setOpen(next);
};
const onSubmit = form.handleSubmit(async (values) => { const onSubmit = form.handleSubmit(async (values) => {
try { try {
@@ -82,7 +86,7 @@ export function CaseEditDialog({ data }: { data: CaseDetail }) {
}); });
return ( return (
<Dialog open={open} onOpenChange={setOpen}> <Dialog open={open} onOpenChange={handleOpenChange}>
<DialogTrigger asChild> <DialogTrigger asChild>
<Button variant="outline" size="sm"> <Button variant="outline" size="sm">
עריכת פרטי תיק עריכת פרטי תיק

View File

@@ -63,8 +63,10 @@ export function CaseHeader({
<span className="text-navy tabular-nums">{data?.case_number ?? "…"}</span> <span className="text-navy tabular-nums">{data?.case_number ?? "…"}</span>
</nav> </nav>
<div className="flex items-start justify-between gap-6 flex-wrap"> {/* title block + metadata on one row — no wrap, so the metadata dl stays
<div className="min-w-0"> parallel to the H1 and a long parties line can't push it below. */}
<div className="flex items-start justify-between gap-6">
<div className="min-w-0 flex-1">
{/* title row — H1 + status/type/blam chips inline (mockup .band h1) */} {/* title row — H1 + status/type/blam chips inline (mockup .band h1) */}
<h1 className="text-navy text-[1.7rem] font-bold leading-tight flex items-center gap-3 flex-wrap mb-0"> <h1 className="text-navy text-[1.7rem] font-bold leading-tight flex items-center gap-3 flex-wrap mb-0">
<span className="tabular-nums"> <span className="tabular-nums">
@@ -107,11 +109,15 @@ export function CaseHeader({
{data.title} {data.title}
</p> </p>
)} )}
{/* parties line (mockup .parties) */} {/* parties line (mockup .parties) — clamped to 2 lines so a long list
of appellants/respondents doesn't grow the band; full text on hover. */}
{parties ? ( {parties ? (
<p className="text-ink-soft text-sm mt-1.5">{parties}</p> <p className="text-ink-soft text-sm mt-1.5 max-w-3xl line-clamp-2" title={parties}>
{parties}
</p>
) : data?.subject ? ( ) : data?.subject ? (
<p className="text-ink-soft text-sm mt-1.5 max-w-3xl leading-relaxed"> <p className="text-ink-soft text-sm mt-1.5 max-w-3xl leading-relaxed line-clamp-2"
title={data.subject}>
{data.subject} {data.subject}
</p> </p>
) : null} ) : null}

View File

@@ -120,7 +120,7 @@ const columns: ColumnDef<Case>[] = [
accessorKey: "title", accessorKey: "title",
header: "כותרת", header: "כותרת",
cell: ({ row }) => ( cell: ({ row }) => (
<div className="text-ink max-w-[420px] truncate flex items-center gap-2" title={row.original.title}> <div className="text-ink max-w-[420px] min-w-0 truncate flex items-center gap-2" title={row.original.title}>
{(row.original.proceeding_type === 'בל"מ' || isBlamSubtype(row.original.appeal_subtype)) && ( {(row.original.proceeding_type === 'בל"מ' || isBlamSubtype(row.original.appeal_subtype)) && (
<Badge <Badge
variant="outline" variant="outline"

View File

@@ -56,7 +56,7 @@ export function DecisionBlocksPanel({ caseNumber }: { caseNumber: string }) {
} }
if (!data) return null; if (!data) return null;
const written = data.blocks.filter((b) => b.word_count > 0).length; const written = data.blocks.filter((b) => (b.word_count ?? 0) > 0).length;
return ( return (
<div className="space-y-4"> <div className="space-y-4">
@@ -101,11 +101,11 @@ export function DecisionBlocksPanel({ caseNumber }: { caseNumber: string }) {
{blockLabel(block)} {blockLabel(block)}
</span> </span>
<Badge <Badge
className={`text-[0.65rem] border ${STATUS_CLASSES[block.status]}`} className={`text-[0.65rem] border ${STATUS_CLASSES[block.status] ?? ""}`}
> >
{STATUS_LABELS[block.status]} {STATUS_LABELS[block.status] ?? block.status}
</Badge> </Badge>
{block.word_count > 0 && ( {(block.word_count ?? 0) > 0 && (
<span className="text-[0.7rem] text-ink-muted tabular-nums"> <span className="text-[0.7rem] text-ink-muted tabular-nums">
{block.word_count} מילים {block.word_count} מילים
</span> </span>
@@ -137,19 +137,22 @@ function BlockEditor({
caseNumber: string; caseNumber: string;
block: DecisionBlock; block: DecisionBlock;
}) { }) {
// The endpoint has no FastAPI response model, so `content` may arrive null
// for an empty block — coerce to "" to keep `.trim()` / state safe.
const content = block.content ?? "";
const [editing, setEditing] = useState(false); const [editing, setEditing] = useState(false);
const [value, setValue] = useState(block.content); const [value, setValue] = useState(content);
const [state, setState] = useState<SaveState>({ kind: "idle" }); const [state, setState] = useState<SaveState>({ kind: "idle" });
/* The last content known to be persisted — used to skip no-op saves. */ /* The last content known to be persisted — used to skip no-op saves. */
const [baseline, setBaseline] = useState(block.content); const [baseline, setBaseline] = useState(content);
const save = useSaveBlock(caseNumber); const save = useSaveBlock(caseNumber);
/* Re-sync when the upstream query refetches (e.g. after another save) while /* Re-sync when the upstream query refetches (e.g. after another save) while
* not actively editing. Adjusting state during render — the documented React * not actively editing. Adjusting state during render — the documented React
* pattern for derived-from-props — avoids a setState-in-effect cascade. */ * pattern for derived-from-props — avoids a setState-in-effect cascade. */
if (!editing && block.content !== baseline) { if (!editing && content !== baseline) {
setBaseline(block.content); setBaseline(content);
setValue(block.content); setValue(content);
} }
async function handleSave() { async function handleSave() {
@@ -171,8 +174,8 @@ function BlockEditor({
if (!editing) { if (!editing) {
return ( return (
<div className="space-y-3"> <div className="space-y-3">
{block.content.trim() ? ( {content.trim() ? (
<Markdown content={block.content} /> <Markdown content={content} />
) : ( ) : (
<p className="text-sm text-ink-muted italic">בלוק ריק.</p> <p className="text-sm text-ink-muted italic">בלוק ריק.</p>
)} )}

View File

@@ -90,13 +90,23 @@ export function DocumentTypeEditor({
// clear it so it doesn't dangle confusingly in metadata. // clear it so it doesn't dangle confusingly in metadata.
if (!isAppraisal && appraiserSide) body.appraiser_side = ""; if (!isAppraisal && appraiserSide) body.appraiser_side = "";
await patch.mutateAsync({ docId, patch: body }); // Swallow the rejection — errors surface via `patch.isError`; an unhandled
setSaved(true); // rejection from this async click handler would otherwise leak.
try {
await patch.mutateAsync({ docId, patch: body });
setSaved(true);
} catch {
/* surfaced via patch.isError */
}
} }
async function handleExtract() { async function handleExtract() {
const result = await extract.mutateAsync(); try {
setExtractResult(result); const result = await extract.mutateAsync();
setExtractResult(result);
} catch {
/* surfaced via extract.isError */
}
} }
// Build the on-badge label: "שומה · שמאי הוועדה" when both present. // Build the on-badge label: "שומה · שמאי הוועדה" when both present.
@@ -301,7 +311,7 @@ function PostSaveView({
נותרו {extractResult.missing.length} שומות ללא תיוג צד. תייג אותן נותרו {extractResult.missing.length} שומות ללא תיוג צד. תייג אותן
לפני הפעלת החילוץ: לפני הפעלת החילוץ:
</p> </p>
<ul className="list-disc pr-4 space-y-0.5"> <ul className="list-disc ps-4 space-y-0.5">
{extractResult.missing.map((m) => ( {extractResult.missing.map((m) => (
<li key={m.document_id} className="truncate"> <li key={m.document_id} className="truncate">
{m.title} {m.title}

View File

@@ -1,6 +1,7 @@
"use client"; "use client";
import { useRef, useState } from "react"; import { useRef, useState } from "react";
import Link from "next/link";
import { useQueryClient } from "@tanstack/react-query"; import { useQueryClient } from "@tanstack/react-query";
import { Badge } from "@/components/ui/badge"; import { Badge } from "@/components/ui/badge";
import { Button } from "@/components/ui/button"; import { Button } from "@/components/ui/button";
@@ -344,9 +345,9 @@ export function DraftsPanel({
</Button> </Button>
<span className="text-[0.7rem] text-ink-muted"> <span className="text-[0.7rem] text-ink-muted">
סטטוס הריצה בדף{" "} סטטוס הריצה בדף{" "}
<a href="/operations" className="underline"> <Link href="/operations" className="underline">
התפעול התפעול
</a> </Link>
</span> </span>
</div> </div>
)} )}
@@ -628,7 +629,7 @@ function CitationsSection({ caseNumber }: { caseNumber: string }) {
<div className="rounded-lg border border-rule overflow-hidden divide-y divide-rule"> <div className="rounded-lg border border-rule overflow-hidden divide-y divide-rule">
{data.linked.map((c) => ( {data.linked.map((c) => (
<a <Link
key={`l-${c.citation}`} key={`l-${c.citation}`}
href={`/precedents/${c.cited_id}`} href={`/precedents/${c.cited_id}`}
className="flex items-center gap-2 px-4 py-2.5 text-sm hover:bg-rule-soft/20" className="flex items-center gap-2 px-4 py-2.5 text-sm hover:bg-rule-soft/20"
@@ -641,7 +642,7 @@ function CitationsSection({ caseNumber }: { caseNumber: string }) {
<Badge className="ms-auto bg-success-bg text-success border-success/40 text-[0.65rem] shrink-0"> <Badge className="ms-auto bg-success-bg text-success border-success/40 text-[0.65rem] shrink-0">
בספרייה בספרייה
</Badge> </Badge>
</a> </Link>
))} ))}
{data.missing.map((c) => ( {data.missing.map((c) => (
<div <div

View File

@@ -1,6 +1,6 @@
"use client"; "use client";
import { useMemo } from "react"; import { useMemo, type ReactNode } from "react";
import { import {
Accordion, Accordion,
AccordionContent, AccordionContent,
@@ -9,7 +9,6 @@ import {
} from "@/components/ui/accordion"; } from "@/components/ui/accordion";
import { Badge } from "@/components/ui/badge"; import { Badge } from "@/components/ui/badge";
import { Button } from "@/components/ui/button"; import { Button } from "@/components/ui/button";
import { Card, CardContent } from "@/components/ui/card";
import { import {
Popover, Popover,
PopoverContent, PopoverContent,
@@ -22,6 +21,7 @@ import {
PRIORITY_ORDER, PRIORITY_ORDER,
useAggregateArguments, useAggregateArguments,
useLegalArguments, useLegalArguments,
type AggregateArgumentsResult,
type LegalArgument, type LegalArgument,
type LegalArgumentParty, type LegalArgumentParty,
type LegalArgumentPriority, type LegalArgumentPriority,
@@ -189,82 +189,167 @@ export function LegalArgumentsPanel({ caseNumber }: LegalArgumentsPanelProps) {
}, [data]); }, [data]);
const handleAggregate = (force: boolean) => { const handleAggregate = (force: boolean) => {
// Status feedback is rendered inline as a banner from `aggregate.data`;
// only hard transport/HTTP errors fall through to a toast.
aggregate.mutate(force, { aggregate.mutate(force, {
onSuccess: () => {
toast.success(
force
? "הופעלה חזרה חישוב טיעונים (force). יסתיים תוך דקה."
: "הופעל חישוב טיעונים. רענן בעוד דקה.",
);
},
onError: (e) => toast.error(`שגיאה: ${(e as Error).message}`), onError: (e) => toast.error(`שגיאה: ${(e as Error).message}`),
}); });
}; };
return ( return (
<Card className="bg-surface border-rule shadow-sm"> <div className="space-y-4">
<CardContent className="px-6 py-5 space-y-4"> <div className="flex items-center justify-between flex-wrap gap-3">
<div className="flex items-center justify-between flex-wrap gap-3"> <div>
<div> <h2 className="text-navy text-base font-semibold">טיעונים משפטיים</h2>
<h2 className="text-navy text-base font-semibold"> <p className="text-ink-muted text-xs mt-0.5">
טיעונים משפטיים טיעונים מאוגדים מתוך הטענות הגולמיות, מקובצים לפי צד וקדימות. החישוב
</h2> רץ אצל המנתח המשפטי ומדווח בחזרה.
<p className="text-ink-muted text-xs mt-0.5"> </p>
טיעונים מאוגדים מתוך הפרופוזיציות הגולמיות, מקובצים לפי צד וקדימות.
</p>
</div>
<div className="flex items-center gap-2">
<Button
variant="outline"
size="sm"
disabled={aggregate.isPending}
onClick={() => handleAggregate(false)}
>
{aggregate.isPending ? (
<Loader2 className="w-3.5 h-3.5 animate-spin me-1.5" />
) : (
<Sparkles className="w-3.5 h-3.5 me-1.5" />
)}
חשב טיעונים
</Button>
<Button
variant="ghost"
size="sm"
disabled={aggregate.isPending || !data?.total}
onClick={() => handleAggregate(true)}
title="חישוב מחדש (מוחק טיעונים קיימים)"
>
<RefreshCw className="w-3.5 h-3.5" />
</Button>
</div>
</div> </div>
<div className="flex items-center gap-2">
<Button
variant="outline"
size="sm"
disabled={aggregate.isPending}
onClick={() => handleAggregate(false)}
>
{aggregate.isPending ? (
<Loader2 className="w-3.5 h-3.5 animate-spin me-1.5" />
) : (
<Sparkles className="w-3.5 h-3.5 me-1.5" />
)}
חשב טיעונים
</Button>
<Button
variant="ghost"
size="sm"
disabled={aggregate.isPending || !data?.total}
onClick={() => handleAggregate(true)}
title="חישוב מחדש (מוחק טיעונים קיימים)"
>
<RefreshCw className="w-3.5 h-3.5" />
</Button>
</div>
</div>
{isPending ? ( <AggregateStatusBanner
<div className="space-y-2"> result={aggregate.data}
<Skeleton className="h-6 w-48" /> force={aggregate.variables ?? false}
<Skeleton className="h-20 w-full" /> />
<Skeleton className="h-20 w-full" />
</div> {isPending ? (
) : isError ? ( <div className="space-y-2">
<p className="text-danger text-sm"> <Skeleton className="h-6 w-48" />
שגיאה בטעינת טיעונים: {(error as Error).message} <Skeleton className="h-20 w-full" />
</p> <Skeleton className="h-20 w-full" />
) : !data?.total ? ( </div>
<p className="text-ink-muted text-sm"> ) : isError ? (
אין טיעונים מאוגדים עדיין. לחץ &ldquo;חשב טיעונים&rdquo; כדי להריץ את ה-aggregator. <p className="text-danger text-sm">
</p> שגיאה בטעינת טיעונים: {(error as Error).message}
) : ( </p>
<div className="space-y-6"> ) : !data?.total ? (
{parties.map((party) => ( <p className="text-ink-muted text-sm">
<PartySection אין טיעונים מאוגדים עדיין. לחץ &ldquo;חשב טיעונים&rdquo; כדי לשלוח את
key={party} החישוב למנתח המשפטי.
party={party} </p>
args={data.by_party[party] ?? []} ) : (
/> <div className="space-y-6">
))} {parties.map((party) => (
</div> <PartySection
)} key={party}
</CardContent> party={party}
</Card> args={data.by_party[party] ?? []}
/>
))}
</div>
)}
</div>
);
}
const BANNER_TONE = {
info: "border-info/30 bg-info-bg",
gold: "border-gold/40 bg-gold-wash",
muted: "border-rule bg-rule-soft",
warn: "border-warn/40 bg-warn-bg",
} as const;
const DOT_TONE = {
info: "bg-info",
gold: "bg-gold-deep",
muted: "bg-ink-light",
warn: "bg-warn",
} as const;
const TITLE_TONE = {
info: "text-info",
gold: "text-gold-deep",
muted: "text-ink-soft",
warn: "text-warn",
} as const;
/**
* Inline status feedback for the aggregation trigger — mirrors the approved
* Claude Design mockup (25-legal-arguments-panel). Shown only after a click;
* the four states map to the endpoint's discriminated-union result.
*/
function AggregateStatusBanner({
result,
force,
}: {
result: AggregateArgumentsResult | undefined;
force: boolean;
}) {
if (!result) return null;
let tone: keyof typeof BANNER_TONE;
let title: string;
let body: ReactNode;
switch (result.status) {
case "queued":
tone = "info";
title = "נשלח לאנליטיקאי.";
body = `${force ? "החישוב-מחדש" : "החישוב"} רץ ברקע אצל המנתח המשפטי; התוצאה תופיע תוך כמה דקות — רענן את הדף.`;
break;
case "exists":
tone = "gold";
title = "כבר חושב.";
body = result.message;
break;
case "no_claims":
tone = "muted";
title = "אין טענות גולמיות.";
body = result.message;
break;
case "skipped":
tone = "warn";
title = "לא ניתן להפעיל אוטומטית";
body = (
<>
{` (${result.reason}). הרץ ידנית מ-Claude Code: `}
<code className="font-mono text-[0.72rem] bg-navy/5 rounded px-1.5 py-0.5 select-all">
mcp__legal-ai__aggregate_claims_to_arguments
</code>
</>
);
break;
default:
return null;
}
return (
<div
className={`flex items-start gap-2.5 rounded-md border px-3.5 py-2.5 text-[0.8rem] leading-relaxed text-ink-soft ${BANNER_TONE[tone]}`}
>
<span
className={`mt-1.5 size-2 flex-none rounded-full ${DOT_TONE[tone]}`}
aria-hidden
/>
<p>
<strong className={`font-semibold ${TITLE_TONE[tone]}`}>{title}</strong>{" "}
{body}
</p>
</div>
); );
} }

View File

@@ -25,14 +25,20 @@ export function StatusChanger({
caseNumber: string; caseNumber: string;
currentStatus?: CaseStatus; currentStatus?: CaseStatus;
}) { }) {
const [selected, setSelected] = useState<CaseStatus | "">(currentStatus ?? ""); // `null` = untouched → the dropdown tracks the live `currentStatus` (which
// arrives async and changes on the 5s poll / external updates). Only an
// explicit pick overrides it, until save resets back to tracking.
const [picked, setPicked] = useState<CaseStatus | null>(null);
const mutate = useUpdateCase(caseNumber); const mutate = useUpdateCase(caseNumber);
const effective: CaseStatus | "" = picked ?? currentStatus ?? "";
const canSave = Boolean(effective) && effective !== currentStatus;
const handleSave = async () => { const handleSave = async () => {
if (!selected || selected === currentStatus) return; if (!canSave || !effective) return;
try { try {
await mutate.mutateAsync({ status: selected }); await mutate.mutateAsync({ status: effective });
toast.success(`הסטטוס עודכן ל${STATUS_LABELS[selected]}`); toast.success(`הסטטוס עודכן ל${STATUS_LABELS[effective]}`);
setPicked(null);
} catch (e) { } catch (e) {
toast.error(e instanceof Error ? e.message : "שגיאה בעדכון הסטטוס"); toast.error(e instanceof Error ? e.message : "שגיאה בעדכון הסטטוס");
} }
@@ -43,8 +49,8 @@ export function StatusChanger({
<label className="text-[0.72rem] text-ink-muted block">שינוי סטטוס ידני</label> <label className="text-[0.72rem] text-ink-muted block">שינוי סטטוס ידני</label>
<div className="flex items-center gap-2"> <div className="flex items-center gap-2">
<Select <Select
value={selected || "__current__"} value={effective || "__current__"}
onValueChange={(v) => setSelected(v === "__current__" ? "" : v as CaseStatus)} onValueChange={(v) => setPicked(v === "__current__" ? null : v as CaseStatus)}
dir="rtl" dir="rtl"
> >
<SelectTrigger className="text-[0.75rem] h-8"> <SelectTrigger className="text-[0.75rem] h-8">
@@ -68,7 +74,7 @@ export function StatusChanger({
size="sm" size="sm"
variant="outline" variant="outline"
className="h-8 text-[0.72rem] px-3 shrink-0" className="h-8 text-[0.72rem] px-3 shrink-0"
disabled={!selected || selected === currentStatus || mutate.isPending} disabled={!canSave || mutate.isPending}
onClick={handleSave} onClick={handleSave}
> >
{mutate.isPending ? "שומר…" : "עדכן"} {mutate.isPending ? "שומר…" : "עדכן"}

View File

@@ -0,0 +1,261 @@
"use client";
/**
* "אימות פסיקה" tab (X11 Phase 2 / #154) — per legal argument, the in-corpus
* supporting precedents with the cumulative authority signal (cited_by), a verify
* gate + chair note, and per-issue radar (unlinked digests). The writer cites only
* verified quotes (INV-AH); the digest is never cited (INV-DIG1, radar only).
*/
import { useState } from "react";
import { toast } from "sonner";
import { Card, CardContent } from "@/components/ui/card";
import { Skeleton } from "@/components/ui/skeleton";
import {
useCitationVerification,
useVerifyCitation,
type ArgumentBlock,
type SupportingPrecedent,
type RadarLead,
} from "@/lib/api/citation-verification";
const PRIORITY_LABEL: Record<string, string> = {
threshold: "סף",
substantive: "מהותית",
procedural: "פרוצדורלית",
relief: "סעד",
};
const RADAR_LABEL: Record<string, string> = {
new_lead: "ליד חדש",
gap_open: "בתור-חסרים",
fetched: "נמשך",
available_link: "בקורפוס — לקשר",
};
export function CitationVerificationPanel({ caseNumber }: { caseNumber: string }) {
const { data, isPending, isError } = useCitationVerification(caseNumber);
const verify = useVerifyCitation(caseNumber);
const [notes, setNotes] = useState<Record<string, string>>({});
if (isPending) {
return (
<Card className="bg-surface border-rule shadow-sm">
<CardContent className="px-6 py-5 space-y-3">
<Skeleton className="h-6 w-64" />
<Skeleton className="h-24 w-full" />
<Skeleton className="h-24 w-full" />
</CardContent>
</Card>
);
}
if (isError || !data || data.status !== "ok") {
return (
<Card className="bg-surface border-rule shadow-sm">
<CardContent className="px-6 py-8 text-center text-ink-muted">
לא ניתן לטעון את אימות-הפסיקה כעת.
</CardContent>
</Card>
);
}
if (!data.arguments.length) {
return (
<Card className="bg-surface border-rule shadow-sm">
<CardContent className="px-6 py-12 text-center space-y-2">
<div className="text-gold text-3xl" aria-hidden></div>
<p className="text-navy font-semibold mb-0">אין עדיין סוגיות מזוקקות לתיק</p>
<p className="text-ink-muted text-sm">
הרץ את ניתוח-הטענות (aggregate) כדי שהמערכת תזהה סוגיות ותתאים להן פסיקה.
</p>
</CardContent>
</Card>
);
}
function runVerify(a: ArgumentBlock, s: SupportingPrecedent, verified: boolean) {
verify.mutate(
{
argument_id: a.argument_id,
case_law_id: s.case_law_id,
quote: s.quote,
citation: s.case_number || s.case_name,
chair_note: notes[s.case_law_id] ?? s.chair_note,
verified,
attached_id: s.attached_id ?? "",
},
{
onSuccess: () => toast.success(verified ? "הציטוט אומת" : "סומן כלא-רלוונטי"),
onError: (e: Error) => toast.error(`שגיאה: ${e.message}`),
},
);
}
function saveNote(a: ArgumentBlock, s: SupportingPrecedent) {
const note = notes[s.case_law_id];
if (note === undefined || note === s.chair_note) return;
verify.mutate(
{
argument_id: a.argument_id,
case_law_id: s.case_law_id,
quote: s.quote,
citation: s.case_number || s.case_name,
chair_note: note,
verified: s.verified,
attached_id: s.attached_id ?? "",
},
{
onSuccess: () => toast.success("הערת-יו״ר נשמרה"),
onError: (e: Error) => toast.error(`שגיאה: ${e.message}`),
},
);
}
const sum = data.summary;
return (
<div className="space-y-4">
{/* summary */}
<div className="flex flex-wrap items-center gap-x-5 gap-y-1 text-sm text-ink-soft bg-parchment border border-rule rounded-lg px-4 py-2.5">
<span>סוגיות עם פסיקה: <b className="text-navy tabular-nums">{sum.arguments_with_support}/{sum.arguments_total}</b></span>
<span>אומתו: <b className="text-success tabular-nums">{sum.verified}</b></span>
<span>לידֵי רדאר: <b className="text-warn tabular-nums">{sum.radar_leads}</b></span>
<span className="text-ink-muted text-[0.78rem] ms-auto">ה-writer מצטט רק מאומתים (INV-AH) · היומון אינו מצוטט (INV-DIG1)</span>
</div>
{data.arguments.map((a) => (
<Card key={a.argument_id} className="bg-surface border-rule shadow-sm overflow-hidden">
<CardContent className="p-0">
{/* issue header */}
<div className="flex items-start gap-3 px-5 py-3.5 border-b border-rule-soft bg-parchment/60">
<div className="min-w-0 flex-1">
<div className="text-navy font-bold text-[0.98rem]">{a.title}</div>
<div className="text-ink-muted text-xs mt-0.5">
סוגיה{a.legal_topic ? ` · ${a.legal_topic}` : ""}
</div>
</div>
{a.priority && (
<span className="shrink-0 rounded-full bg-gold-wash text-gold-deep text-[0.68rem] font-semibold px-2.5 py-0.5">
{PRIORITY_LABEL[a.priority] ?? a.priority}
</span>
)}
</div>
{/* supporting precedents */}
<div className="px-5 py-3">
<div className="text-[0.72rem] font-bold text-ink-muted mb-1">פסיקה תומכת בקורפוס</div>
{a.supporting.length === 0 ? (
<p className="text-ink-muted text-sm py-2">לא נמצאה פסיקה תומכת בקורפוס לסוגיה זו.</p>
) : (
<ul className="list-none p-0 m-0 divide-y divide-rule-soft">
{a.supporting.map((s) => (
<li key={s.case_law_id} className="py-3 first:pt-1">
<div className="flex items-center gap-2 flex-wrap">
<a
href={`/precedents/${s.case_law_id}`}
className="font-bold text-gold-deep text-[0.86rem] hover:text-navy"
>
{s.case_number || s.case_name}
</a>
{s.cited_by.positive > 0 && (
<span className="rounded-full bg-success-bg text-success text-[0.68rem] font-semibold px-2 py-0.5">
אומץ ×{s.cited_by.positive}
</span>
)}
{s.cited_by.negative > 0 && (
<span className="rounded-full bg-danger-bg text-danger text-[0.68rem] font-semibold px-2 py-0.5">
אובחן ×{s.cited_by.negative}
</span>
)}
</div>
{s.quote && (
<blockquote className="border-s-[3px] border-gold bg-gold-wash text-ink-soft text-sm leading-7 rounded-e-md px-3.5 py-2 my-2 mx-0">
{s.quote}
</blockquote>
)}
<div className="flex items-center gap-2 flex-wrap">
{s.verified ? (
<>
<span className="text-success text-xs font-semibold"> אומת ע״י היו״ר</span>
<button
type="button"
disabled={verify.isPending}
onClick={() => runVerify(a, s, false)}
className="text-xs text-ink-muted hover:text-danger border border-rule rounded-md px-2.5 py-1 disabled:opacity-40"
>
בטל אימות
</button>
</>
) : (
<>
<button
type="button"
disabled={verify.isPending}
onClick={() => runVerify(a, s, true)}
className="text-xs font-semibold text-success bg-success-bg border border-success/30 rounded-md px-3 py-1 hover:brightness-95 disabled:opacity-40"
>
מאמת
</button>
<button
type="button"
disabled={verify.isPending}
onClick={() => runVerify(a, s, false)}
className="text-xs font-semibold text-danger border border-rule rounded-md px-3 py-1 hover:bg-danger-bg disabled:opacity-40"
>
לא רלוונטי
</button>
</>
)}
{s.cited_by.negative > 0 && !s.verified && (
<span className="text-warn text-[0.72rem] ms-1"> אובחן בדוק הקשר</span>
)}
</div>
{/* chair note */}
<div className="mt-2 flex items-start gap-2">
<label className="text-[0.72rem] font-bold text-gold-deep whitespace-nowrap pt-1.5">
הערת יו״ר:
</label>
<input
type="text"
defaultValue={s.chair_note}
placeholder="הוסף הערה לציטוט — מדוע תומך / הסתייגות / איך לשלב…"
onChange={(e) => setNotes((n) => ({ ...n, [s.case_law_id]: e.target.value }))}
onBlur={() => saveNote(a, s)}
className="flex-1 text-[0.8rem] text-ink-soft bg-parchment border border-dashed border-rule rounded-md px-2.5 py-1.5 focus:outline-none focus:border-gold"
/>
</div>
</li>
))}
</ul>
)}
</div>
{/* radar — unlinked digests for this issue (INV-DIG1: pointer, not citation) */}
<RadarStrip radar={a.radar} />
</CardContent>
</Card>
))}
</div>
);
}
function RadarStrip({ radar }: { radar: RadarLead[] }) {
if (!radar.length) return null;
return (
<div className="mx-5 mb-4 rounded-lg border border-dashed border-gold bg-parchment px-4 py-2.5">
<div className="text-[0.72rem] font-bold text-gold-deep mb-1">
📡 רדאר פסיקה שאין לנו בקורפוס (מצביע, לא מצוטט)
</div>
<ul className="list-none p-0 m-0 space-y-1">
{radar.map((r) => (
<li key={r.digest_id} className="flex items-center gap-2 flex-wrap text-[0.8rem]">
<span className="font-semibold text-navy">{r.underlying_citation || "(אין מראה-מקום)"}</span>
<span className="rounded-full bg-warn-bg text-warn text-[0.64rem] font-semibold px-2 py-0.5">
{RADAR_LABEL[r.action] ?? r.action}
</span>
<span className="text-ink-muted truncate">{r.headline}</span>
</li>
))}
</ul>
</div>
);
}

View File

@@ -43,7 +43,9 @@ function statusLabel(event: ProgressEvent | null): string {
if (event.status === "processing") if (event.status === "processing")
return event.step ? `בעיבוד · ${event.step}` : "בעיבוד"; return event.step ? `בעיבוד · ${event.step}` : "בעיבוד";
if (event.status === "completed") return "הושלם"; if (event.status === "completed") return "הושלם";
if (event.status === "unknown") return "הושלם"; // TTL expired / subscribed too late — the outcome is genuinely unknown, not
// a confirmed success. The case-detail refetch is the real source of truth.
if (event.status === "unknown") return "הסתיים — רענן לאישור";
if (event.status === "failed") return event.error ?? "נכשל"; if (event.status === "failed") return event.error ?? "נכשל";
return event.status; return event.status;
} }
@@ -62,7 +64,9 @@ function UploadRowView({ row, caseNumber }: { row: UploadRow; caseNumber: string
const progress = useProgress(row.taskId, caseNumber); const progress = useProgress(row.taskId, caseNumber);
const pct = row.error ? 100 : progressPercent(progress); const pct = row.error ? 100 : progressPercent(progress);
const failed = row.error || progress?.status === "failed"; const failed = row.error || progress?.status === "failed";
const done = progress?.status === "completed" || progress?.status === "unknown"; const done = progress?.status === "completed";
// `unknown` = settled-but-indeterminate; render neutral, not green success.
const indeterminate = progress?.status === "unknown";
return ( return (
<li className="rounded-lg border border-rule bg-parchment/40 px-4 py-3 space-y-2"> <li className="rounded-lg border border-rule bg-parchment/40 px-4 py-3 space-y-2">
@@ -71,6 +75,8 @@ function UploadRowView({ row, caseNumber }: { row: UploadRow; caseNumber: string
<CheckCircle2 className="w-4 h-4 text-success shrink-0" /> <CheckCircle2 className="w-4 h-4 text-success shrink-0" />
) : failed ? ( ) : failed ? (
<XCircle className="w-4 h-4 text-danger shrink-0" /> <XCircle className="w-4 h-4 text-danger shrink-0" />
) : indeterminate ? (
<CheckCircle2 className="w-4 h-4 text-ink-muted shrink-0" />
) : ( ) : (
<Loader2 className="w-4 h-4 text-gold animate-spin shrink-0" /> <Loader2 className="w-4 h-4 text-gold animate-spin shrink-0" />
)} )}

View File

@@ -48,7 +48,14 @@ export type GraphControls = {
}; };
const ALL = "__all__"; const ALL = "__all__";
const YEARS = Array.from({ length: 2026 - 1994 + 1 }, (_, i) => 2026 - i); // Floor covers the oldest dated precedent in the corpus (currently ע"א 725/81 →
// 1982); ceiling tracks the current year so the range never ages out.
const YEAR_FLOOR = 1980;
const CURRENT_YEAR = new Date().getFullYear();
const YEARS = Array.from(
{ length: CURRENT_YEAR - YEAR_FLOOR + 1 },
(_, i) => CURRENT_YEAR - i,
);
const COLOR_BY: { value: ColorBy; label: string }[] = [ const COLOR_BY: { value: ColorBy; label: string }[] = [
{ value: "type", label: "סוג נקודה" }, { value: "type", label: "סוג נקודה" },

View File

@@ -1,7 +1,7 @@
"use client"; "use client";
import { useState } from "react"; import { useMemo, useState } from "react";
import { Trash2, Upload, Pencil, ExternalLink } from "lucide-react"; import { Trash2, Upload, Pencil, ExternalLink, ChevronDown } from "lucide-react";
import { toast } from "sonner"; import { toast } from "sonner";
import Link from "next/link"; import Link from "next/link";
import { import {
@@ -63,6 +63,24 @@ function SourceChip({ party }: { party: CitedByParty | null }) {
); );
} }
const COLS = 7;
function TableHeaderRow() {
return (
<TableHeader className="bg-parchment">
<TableRow className="border-rule hover:bg-transparent">
<TableHead className="text-ink-muted text-right font-medium text-xs">פסיקה</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">נושא</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">תיק</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">צוטט ע״י</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">סטטוס</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">נוצר</TableHead>
<TableHead className="text-navy" />
</TableRow>
</TableHeader>
);
}
function TableSkeleton({ cols }: { cols: number }) { function TableSkeleton({ cols }: { cols: number }) {
return ( return (
<> <>
@@ -79,6 +97,38 @@ function TableSkeleton({ cols }: { cols: number }) {
); );
} }
/** Accordion section meta — groups rows by discovery source (chair's request).
* "צוטט ע״י דפנה" (corpus-decision reliance) is the high-value group and opens
* by default; the noisy יומון group and the misc "אחר" group start collapsed. */
type SectionKey = "chair" | "digest" | "other";
const SECTION_META: Record<
SectionKey,
{ label: string; title: string; desc: string; chip: string; defaultOpen: boolean }
> = {
chair: {
label: "צוטט ע״י דפנה",
title: "פסיקה שהוועדה נסמכת עליה",
desc: "מצוטטת בהחלטות-הקורפוס (גרף-הציטוטים)",
chip: "bg-success-bg text-success",
defaultOpen: true,
},
digest: {
label: "יומון",
title: "זוהתה ביומון יומי",
desc: "מצביע, טרם צוטט בהחלטה",
chip: "bg-teal-bg text-teal",
defaultOpen: false,
},
other: {
label: "אחר",
title: "כתב-טענות / ידני",
desc: "צוטט בערר חי או נרשם ידנית",
chip: "bg-info-bg text-info",
defaultOpen: false,
},
};
const SECTION_ORDER: SectionKey[] = ["chair", "digest", "other"];
type Props = { type Props = {
status?: MissingPrecedentStatus | ""; status?: MissingPrecedentStatus | "";
q?: string; q?: string;
@@ -91,7 +141,11 @@ export function MissingPrecedentsTable({ status, q, legalTopic }: Props) {
status: status === "" ? undefined : status, status: status === "" ? undefined : status,
q, q,
legalTopic, legalTopic,
limit: 200, // Load the full set so the accordion section counts (chair/digest/other)
// reflect the real totals — at limit:200 the page showed only the first
// page (e.g. 16 of 106 committee rows) while the header showed the true
// count, so the sections didn't sum to it. Backend caps at 2000.
limit: 1000,
}); });
const del = useDeleteMissingPrecedent(); const del = useDeleteMissingPrecedent();
@@ -108,6 +162,196 @@ export function MissingPrecedentsTable({ status, q, legalTopic }: Props) {
} }
}; };
/* Partition the current result set by discovery source so each accordion
* section renders only its own rows. A row "cited by the committee" (has a
* resolved chair from the citation-graph bridge) wins over its raw
* discovery_source — that's the group the chair cares about. */
const groups = useMemo(() => {
const g: Record<SectionKey, MissingPrecedent[]> = { chair: [], digest: [], other: [] };
for (const mp of data?.items ?? []) {
if (mp.cited_by_chairs?.length) g.chair.push(mp);
else if (mp.discovery_source === "digest") g.digest.push(mp);
else g.other.push(mp);
}
return g;
}, [data]);
const renderRow = (mp: MissingPrecedent) => (
<TableRow
key={mp.id}
className="border-rule hover:bg-rule-soft/30 cursor-pointer"
onClick={() => setOpenId(mp.id)}
>
<TableCell className="max-w-[440px]">
{/* Single line when there's no distinct case_name — the old two-line
layout repeated the citation (name fell back to a truncation of the
same citation). Show the name row only when it adds information. */}
{mp.case_name && mp.case_name.trim() ? (
<>
<div className="text-sm text-navy font-semibold truncate">
{mp.case_name}
</div>
<div className="text-[0.72rem] text-ink-muted truncate" dir="rtl">
{mp.citation}
</div>
</>
) : (
<div className="text-sm text-navy font-semibold truncate" dir="rtl">
{mp.citation}
</div>
)}
</TableCell>
<TableCell>
<span className="text-sm text-ink">{mp.legal_topic || "—"}</span>
</TableCell>
<TableCell>
{mp.cited_by_decisions?.length ? (
/* committee decision(s) that cite this missing ruling
(corpus citation-graph bridge) */
<span className="inline-flex items-baseline gap-1.5">
<span className="text-sm text-navy font-semibold tabular-nums" dir="ltr">
{mp.cited_by_decisions[0]}
</span>
{mp.cited_by_decisions.length > 1 ? (
<span className="text-[0.7rem] text-ink-muted">
+{mp.cited_by_decisions.length - 1}
</span>
) : null}
</span>
) : mp.cited_in_case_number ? (
<Link
href={`/cases/${encodeURIComponent(mp.cited_in_case_number)}`}
onClick={(e) => e.stopPropagation()}
className="text-sm text-navy hover:text-gold-deep inline-flex items-center gap-1"
>
{mp.cited_in_case_number}
<ExternalLink className="w-3 h-3" />
</Link>
) : (
<span className="text-ink-muted text-sm"></span>
)}
</TableCell>
<TableCell className="text-sm text-ink">
{mp.cited_by_chairs?.length ? (
/* cited by a committee DECISION in the corpus → show the chair who
relied on it (e.g. דפנה תמיר). Bridge from
precedent_internal_citations; takes priority over the generic
discovery-source chips. */
<>
<Badge
variant="outline"
className="rounded-full whitespace-nowrap bg-success-bg text-success border-transparent"
>
{mp.cited_by_chairs[0]}
</Badge>
{mp.cited_by_chairs.length > 1 ? (
<div className="text-[0.7rem] text-ink-muted mt-1">
+{mp.cited_by_chairs.length - 1} יו״ר
</div>
) : null}
</>
) : mp.discovery_source === "cited_only" ? (
<>
<Badge
variant="outline"
className="rounded-full whitespace-nowrap bg-plum-bg text-plum border-transparent"
>
פסיקה בקורפוס
</Badge>
{mp.cited_by_precedents?.length ? (
<div className="text-[0.7rem] text-ink-muted truncate max-w-[180px] mt-1">
מצוטט ע״י: {mp.cited_by_precedents.join(", ")}
</div>
) : null}
</>
) : mp.discovery_source === "digest" ? (
<>
<Badge
variant="outline"
className="rounded-full whitespace-nowrap bg-teal-bg text-teal border-transparent"
>
יומון
</Badge>
{mp.yomon_number ? (
<div className="text-[0.7rem] text-ink-muted mt-1">
מס׳ {mp.yomon_number}
</div>
) : null}
</>
) : (
<>
<SourceChip party={mp.cited_by_party} />
{mp.cited_by_party_name ? (
<div className="text-[0.7rem] text-ink-muted truncate max-w-[160px] mt-1">
{mp.cited_by_party_name}
</div>
) : null}
</>
)}
</TableCell>
<TableCell>
<StatusBadge status={mp.status} />
{mp.linked_case_law_number ? (
<div className="text-[0.7rem] text-success mt-1">
{mp.linked_case_law_name || mp.linked_case_law_number}
</div>
) : null}
</TableCell>
<TableCell className="text-[0.78rem] text-ink-muted">
{formatDate(mp.created_at)}
</TableCell>
<TableCell className="text-end">
<div className="flex items-center justify-end gap-2">
{mp.status === "open" ? (
/* gold "העלה והשלם" CTA (mockup 09 `.btn`) */
<Button
size="sm"
onClick={(e) => {
e.stopPropagation();
setOpenId(mp.id);
}}
className="h-7 bg-gold text-white hover:bg-gold-deep border-transparent text-[0.78rem] font-semibold"
>
<Upload className="w-3.5 h-3.5 me-1" />
העלה והשלם
</Button>
) : (
/* passive "done" label for non-open rows */
<span className="inline-flex items-center gap-1 rounded-md bg-rule-soft text-ink-muted text-[0.78rem] font-medium px-2.5 py-1">
{mp.status === "closed" ? "קושר" : STATUS_LABELS[mp.status]}
</span>
)}
{mp.status !== "open" ? (
<Button
variant="ghost"
size="sm"
onClick={(e) => {
e.stopPropagation();
setOpenId(mp.id);
}}
title="פרטים"
>
<Pencil className="w-4 h-4" />
</Button>
) : null}
<Button
variant="ghost"
size="sm"
onClick={(e) => {
e.stopPropagation();
handleDelete(mp);
}}
disabled={del.isPending}
className="text-danger hover:text-danger"
title="מחיקה"
>
<Trash2 className="w-4 h-4" />
</Button>
</div>
</TableCell>
</TableRow>
);
if (error) { if (error) {
return ( return (
<div className="rounded bg-danger-bg border border-danger/40 px-6 py-4 text-danger text-center text-sm"> <div className="rounded bg-danger-bg border border-danger/40 px-6 py-4 text-danger text-center text-sm">
@@ -116,168 +360,61 @@ export function MissingPrecedentsTable({ status, q, legalTopic }: Props) {
); );
} }
return ( if (isPending) {
<> return (
<div className="rounded-lg border border-rule bg-surface shadow-sm overflow-hidden"> <div className="rounded-lg border border-rule bg-surface shadow-sm overflow-hidden">
<Table> <Table>
<TableHeader className="bg-parchment"> <TableHeaderRow />
<TableRow className="border-rule hover:bg-transparent">
<TableHead className="text-ink-muted text-right font-medium text-xs">פסיקה</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">נושא</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">תיק</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">צוטט ע״י</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">סטטוס</TableHead>
<TableHead className="text-ink-muted text-right font-medium text-xs">נוצר</TableHead>
<TableHead className="text-navy" />
</TableRow>
</TableHeader>
<TableBody> <TableBody>
{isPending ? ( <TableSkeleton cols={COLS} />
<TableSkeleton cols={7} />
) : !data?.items.length ? (
<TableRow className="border-rule">
<TableCell colSpan={7} className="text-center text-ink-muted py-8">
אין פסיקות חסרות בקריטריונים הנוכחיים.
</TableCell>
</TableRow>
) : (
data.items.map((mp) => (
<TableRow
key={mp.id}
className="border-rule hover:bg-rule-soft/30 cursor-pointer"
onClick={() => setOpenId(mp.id)}
>
<TableCell className="max-w-[440px]">
<div className="text-sm text-navy font-semibold truncate">
{mp.case_name || mp.citation.split(" ").slice(0, 6).join(" ")}
</div>
<div className="text-[0.72rem] text-ink-muted truncate" dir="rtl">
{mp.citation}
</div>
</TableCell>
<TableCell>
<span className="text-sm text-ink">{mp.legal_topic || "—"}</span>
</TableCell>
<TableCell>
{mp.cited_in_case_number ? (
<Link
href={`/cases/${encodeURIComponent(mp.cited_in_case_number)}`}
onClick={(e) => e.stopPropagation()}
className="text-sm text-navy hover:text-gold-deep inline-flex items-center gap-1"
>
{mp.cited_in_case_number}
<ExternalLink className="w-3 h-3" />
</Link>
) : (
<span className="text-ink-muted text-sm"></span>
)}
</TableCell>
<TableCell className="text-sm text-ink">
{mp.discovery_source === "cited_only" ? (
<>
<Badge
variant="outline"
className="rounded-full whitespace-nowrap bg-plum-bg text-plum border-transparent"
>
פסיקה בקורפוס
</Badge>
{mp.cited_by_precedents?.length ? (
<div className="text-[0.7rem] text-ink-muted truncate max-w-[180px] mt-1">
מצוטט ע״י: {mp.cited_by_precedents.join(", ")}
</div>
) : null}
</>
) : mp.discovery_source === "digest" ? (
<>
<Badge
variant="outline"
className="rounded-full whitespace-nowrap bg-teal-bg text-teal border-transparent"
>
יומון
</Badge>
{mp.yomon_number ? (
<div className="text-[0.7rem] text-ink-muted mt-1">
מס׳ {mp.yomon_number}
</div>
) : null}
</>
) : (
<>
<SourceChip party={mp.cited_by_party} />
{mp.cited_by_party_name ? (
<div className="text-[0.7rem] text-ink-muted truncate max-w-[160px] mt-1">
{mp.cited_by_party_name}
</div>
) : null}
</>
)}
</TableCell>
<TableCell>
<StatusBadge status={mp.status} />
{mp.linked_case_law_number ? (
<div className="text-[0.7rem] text-success mt-1">
{mp.linked_case_law_name || mp.linked_case_law_number}
</div>
) : null}
</TableCell>
<TableCell className="text-[0.78rem] text-ink-muted">
{formatDate(mp.created_at)}
</TableCell>
<TableCell className="text-end">
<div className="flex items-center justify-end gap-2">
{mp.status === "open" ? (
/* gold "העלה והשלם" CTA (mockup 09 `.btn`) */
<Button
size="sm"
onClick={(e) => {
e.stopPropagation();
setOpenId(mp.id);
}}
className="h-7 bg-gold text-white hover:bg-gold-deep border-transparent text-[0.78rem] font-semibold"
>
<Upload className="w-3.5 h-3.5 me-1" />
העלה והשלם
</Button>
) : (
/* passive "done" label for non-open rows */
<span className="inline-flex items-center gap-1 rounded-md bg-rule-soft text-ink-muted text-[0.78rem] font-medium px-2.5 py-1">
{mp.status === "closed" ? "קושר" : STATUS_LABELS[mp.status]}
</span>
)}
{mp.status !== "open" ? (
<Button
variant="ghost"
size="sm"
onClick={(e) => {
e.stopPropagation();
setOpenId(mp.id);
}}
title="פרטים"
>
<Pencil className="w-4 h-4" />
</Button>
) : null}
<Button
variant="ghost"
size="sm"
onClick={(e) => {
e.stopPropagation();
handleDelete(mp);
}}
disabled={del.isPending}
className="text-danger hover:text-danger"
title="מחיקה"
>
<Trash2 className="w-4 h-4" />
</Button>
</div>
</TableCell>
</TableRow>
))
)}
</TableBody> </TableBody>
</Table> </Table>
</div> </div>
);
}
if (!data?.items.length) {
return (
<div className="rounded-lg border border-rule bg-surface shadow-sm px-6 py-10 text-center text-ink-muted text-sm">
אין פסיקות חסרות בקריטריונים הנוכחיים.
</div>
);
}
return (
<>
<div className="space-y-3.5">
{SECTION_ORDER.map((key) => {
const items = groups[key];
if (!items.length) return null;
const meta = SECTION_META[key];
return (
<details
key={key}
open={meta.defaultOpen}
className="group rounded-lg border border-rule bg-surface shadow-sm overflow-hidden"
>
<summary className="flex items-center gap-3 px-4 py-3.5 bg-parchment cursor-pointer select-none list-none [&::-webkit-details-marker]:hidden">
<ChevronDown className="w-4 h-4 text-ink-muted transition-transform -rotate-90 group-open:rotate-0 shrink-0" />
<span className={`rounded-full px-2.5 py-0.5 text-[0.72rem] font-semibold whitespace-nowrap ${meta.chip}`}>
{meta.label}
</span>
<span className="font-bold text-navy text-sm">{meta.title}</span>
<span className="hidden sm:inline text-ink-muted text-[0.78rem]">
{meta.desc}
</span>
<span className="ms-auto tabular-nums text-[0.8rem] text-ink-soft bg-surface border border-rule rounded-full px-2.5 py-0.5 font-semibold">
{items.length}
</span>
</summary>
<Table>
<TableHeaderRow />
<TableBody>{items.map(renderRow)}</TableBody>
</Table>
</details>
);
})}
</div>
<MissingPrecedentDetailDrawer <MissingPrecedentDetailDrawer
id={openId} id={openId}

View File

@@ -18,6 +18,7 @@ import {
useLinkRelatedCase, useLinkRelatedCase,
useUnlinkRelatedCase, useUnlinkRelatedCase,
RelatedCase, RelatedCase,
IncomingCitation,
} from "@/lib/api/precedent-library"; } from "@/lib/api/precedent-library";
const LEVEL_LABELS: Record<string, string> = { const LEVEL_LABELS: Record<string, string> = {
@@ -133,19 +134,28 @@ function LinkDialog({ caseId, currentRelated, open, onOpenChange }: DialogProps)
type SectionProps = { type SectionProps = {
caseId: string; caseId: string;
related: RelatedCase[]; related: RelatedCase[];
/** Decisions that cite THIS ruling — auto-detected from the citation graph
* (read-only). Optional so existing call-sites stay valid. */
incoming?: IncomingCitation[];
}; };
/* Rail-styled citations card (mockup 08 side rail). Renders linked related /* Rail-styled citations card (mockup 08 side rail). Renders linked related
* decisions as a navy-headed card with arrow-prefixed rows; keeps the full * decisions as a navy-headed card with arrow-prefixed rows; keeps the full
* link/unlink logic. Used in the precedent-detail side rail. */ * link/unlink logic. Auto-detected incoming citations (decisions that cite this
export function RelatedCasesSection({ caseId, related }: SectionProps) { * ruling) render in the SAME row style, read-only (no unlink). Used in the
* precedent-detail side rail. */
export function RelatedCasesSection({ caseId, related, incoming = [] }: SectionProps) {
const [dialogOpen, setDialogOpen] = useState(false); const [dialogOpen, setDialogOpen] = useState(false);
// De-dupe: a decision already linked manually shouldn't appear twice.
const manualIds = new Set(related.map((r) => r.id));
const autoCites = incoming.filter((c) => !manualIds.has(c.id));
const total = related.length + autoCites.length;
return ( return (
<div className="rounded-lg border border-rule bg-surface shadow-sm px-4 py-3.5 space-y-2.5"> <div className="rounded-lg border border-rule bg-surface shadow-sm px-4 py-3.5 space-y-2.5">
<div className="flex items-center justify-between gap-2"> <div className="flex items-center justify-between gap-2">
<h3 className="text-navy text-[0.92rem] font-semibold m-0"> <h3 className="text-navy text-[0.92rem] font-semibold m-0">
ציטוטים מקושרים{related.length > 0 ? ` (${related.length})` : ""} ציטוטים מקושרים{total > 0 ? ` (${total})` : ""}
</h3> </h3>
<Button <Button
variant="outline" variant="outline"
@@ -157,10 +167,46 @@ export function RelatedCasesSection({ caseId, related }: SectionProps) {
</Button> </Button>
</div> </div>
{related.length === 0 ? ( {total === 0 ? (
<p className="text-ink-muted text-[0.82rem] m-0">אין החלטות קשורות עדיין</p> <p className="text-ink-muted text-[0.82rem] m-0">אין החלטות קשורות עדיין</p>
) : ( ) : (
<ul className="list-none p-0 m-0"> <ul className="list-none p-0 m-0">
{autoCites.map((c) => (
<li
key={c.id}
className="flex items-start gap-2 py-2 border-b border-rule-soft last:border-b-0"
>
<span className="text-gold font-bold leading-6 shrink-0" aria-hidden></span>
<a
href={`/precedents/${c.id}`}
className="min-w-0 flex-1 group hover:opacity-90 transition-opacity"
>
<div className="text-[0.82rem] text-ink-soft leading-5 group-hover:text-navy">
{c.case_name || c.case_number}
</div>
<div className="flex items-center gap-2 mt-0.5 flex-wrap">
{c.precedent_level && (
<Badge
variant="outline"
className={`text-[0.6rem] ${LEVEL_COLORS[c.precedent_level] ?? ""}`}
>
{LEVEL_LABELS[c.precedent_level] ?? c.precedent_level}
</Badge>
)}
{(c.chair_name || c.court) && (
<span className="text-[0.68rem] text-ink-muted truncate">
{c.chair_name || c.court}
</span>
)}
{c.date && (
<span className="text-[0.68rem] text-ink-muted tabular-nums" dir="ltr">
{c.date.slice(0, 10)}
</span>
)}
</div>
</a>
</li>
))}
{related.map((r) => ( {related.map((r) => (
<li <li
key={r.id} key={r.id}

View File

@@ -45,6 +45,15 @@ export function CuratorPortraitPanel() {
<StatsCard /> <StatsCard />
<RecentFindings /> <RecentFindings />
<div className="bg-info-bg border border-info/40 rounded-lg px-4 py-3 flex items-start gap-2.5 text-[0.78rem] text-ink-soft leading-relaxed">
<span>🔒</span>
<div>
<b className="text-navy">שער-היו״ר (INV-LRN1 / G10):</b> כל הערוצים <b>מציעים</b> בלבד.
ממצא או לקח משפיע על הכתיבה רק אחרי שדפנה אישרה אותו. האוצֵר עצמו read-only על התוכן
מזהה דפוסים ומגיש, לעולם לא משנה קבצים או סגנון בעצמו.
</div>
</div>
<Tabs defaultValue="curator-prompt" dir="rtl"> <Tabs defaultValue="curator-prompt" dir="rtl">
<TabsList className="bg-rule-soft/60"> <TabsList className="bg-rule-soft/60">
<TabsTrigger value="curator-prompt">פרומפט ה-Curator</TabsTrigger> <TabsTrigger value="curator-prompt">פרומפט ה-Curator</TabsTrigger>
@@ -67,38 +76,127 @@ export function CuratorPortraitPanel() {
// ── stats card ───────────────────────────────────────────────────── // ── stats card ─────────────────────────────────────────────────────
function SectionLabel({ children }: { children: React.ReactNode }) {
return (
<div className="text-[0.72rem] uppercase tracking-wider text-gold-deep font-bold mb-3">
{children}
</div>
);
}
/** One of the three learning channels — engine, backing store, count, flow gate. */
function ChannelCard({
letter, title, engine, target, store, value, unit, flows, flowLabel, meta,
}: {
letter: string; title: string; engine: string; target: string; store: string;
value: number; unit: string; flows: boolean; flowLabel: string; meta: string;
}) {
return (
<Card className="bg-surface border-rule">
<CardContent className="px-[18px] py-4 flex flex-col gap-2.5">
<div className="flex items-baseline gap-2">
<span className="w-[22px] h-[22px] rounded-md flex items-center justify-center font-bold text-xs text-parchment shrink-0 bg-navy">
{letter}
</span>
<div>
<h3 className="text-[0.9rem] text-navy font-semibold leading-tight">{title}</h3>
<div className="text-[0.72rem] text-ink-muted">{engine}</div>
</div>
</div>
<div className="text-[0.72rem] text-ink-soft bg-rule-soft rounded-md px-2.5 py-1.5">
{target} <code className="text-[0.68rem] text-gold-deep">{store}</code>
</div>
<div className="flex items-baseline gap-1.5">
<span className="text-[1.9rem] leading-none font-bold text-navy tabular-nums">{value}</span>
<span className="text-[0.78rem] text-ink-muted">{unit}</span>
</div>
<span
className={`inline-flex items-center gap-1.5 rounded-full text-[0.72rem] font-semibold px-2.5 py-0.5 w-fit border ${
flows
? "bg-success-bg text-success border-success/40"
: "bg-gold-wash text-gold-deep border-gold/40"
}`}
>
<span className="text-[0.5rem]"></span>{flowLabel}
</span>
<div className="text-[0.72rem] text-ink-muted leading-relaxed">{meta}</div>
</CardContent>
</Card>
);
}
function StatsCard() { function StatsCard() {
const { data, isPending } = useCuratorStats(); const { data, isPending } = useCuratorStats();
if (isPending) { if (isPending) {
return ( return (
<div className="grid grid-cols-2 md:grid-cols-4 gap-3"> <div className="space-y-4">
{[...Array(4)].map((_, i) => <Skeleton key={i} className="h-20 w-full" />)} <Skeleton className="h-16 w-full" />
<div className="grid grid-cols-1 md:grid-cols-3 gap-4">
{[...Array(3)].map((_, i) => <Skeleton key={i} className="h-44 w-full" />)}
</div>
</div> </div>
); );
} }
if (!data) return null; if (!data) return null;
const { distillation: a, panel: b, curator: c } = data.channels;
return ( return (
<div className="grid grid-cols-2 md:grid-cols-4 gap-3"> <div className="space-y-5">
<Kpi label="ממצאי curator" value={data.total_findings} icon={<Sparkles className="w-4 h-4" />} /> <div className="bg-gold-wash border border-gold rounded-lg px-5 py-4 flex items-start gap-3">
<Kpi label="החלטות שנסקרו" value={`${data.decisions_with_findings}/${data.decisions_total}`} icon={<FileText className="w-4 h-4" />} /> <span className="text-gold-deep text-xl leading-tight"></span>
<Kpi label="ממצאים מאושרים (זורמים לכותב)" value={data.findings_approved} icon={<CheckCircle2 className="w-4 h-4" />} /> <p className="text-ink-soft text-sm leading-relaxed">
<Kpi label="ממוצע ממצאים להחלטה" הלמידה זורמת לכותב ב<b className="text-gold-deep font-semibold">שלושה ערוצים נפרדים</b>, כל אחד
value={ מופעל כשדפנה מסמנת החלטה כסופית. הכרטיס מראה כמה כל ערוץ תרם ומה כבר עובר בפועל לכותב
data.decisions_with_findings > 0 רק ידע ש<b className="text-gold-deep font-semibold">אושר בשער-היו״ר</b> משפיע על הטיוטות (INV-LRN1).
? (data.total_findings / data.decisions_with_findings).toFixed(1) </p>
: "—" </div>
}
icon={<Brain className="w-4 h-4" />} <div>
/> <SectionLabel>שלושת ערוצי-ההזנה לכותב</SectionLabel>
<div className="grid grid-cols-1 md:grid-cols-3 gap-4">
<ChannelCard
letter="א" title="דיסטילציה" engine="Claude · השוואת טיוטה↔סופי"
target="כללי-מתודולוגיה" store="appeal_type_rules"
value={a.items_total} unit="פריטים זורמים"
flows={a.items_total > 0} flowLabel="זורם לכותב"
meta={`${a.discussion_rules} כללי-דיון + ${a.transition_phrases} ביטויי-מעבר · אושרו ב-${a.pairs_folded} זוגות דרך טאב ״למידה״.`}
/>
<ChannelCard
letter="ב" title="פאנל דו-סוכני" engine="DeepSeek + Gemini · הצבעה 2/2"
target="לקחי-סגנון" store="decision_lessons"
value={b.total} unit="לקחים"
flows={b.approved > 0} flowLabel={`${b.approved} אושרו · ${b.proposed} ממתינים`}
meta={`על פני ${b.decisions} החלטות · אישור per-החלטה בטאב הקורפוס.`}
/>
<ChannelCard
letter="ג" title="ממצאי-אוצֵר" engine="האוצֵר · קריאת הסופי + דפוסים"
target="ממצאי-סגנון" store="source=curator"
value={c.total} unit="ממצאים"
flows={c.approved > 0} flowLabel={`${c.approved} אושרו · ${c.proposed} ממתינים`}
meta="35 דפוסים מתויגים להחלטה (סגנון/מבנה/לקסיקון/טבלאי) — נתפסים מבנית, לא רק כהערה."
/>
</div>
</div>
<div>
<SectionLabel>ערוץ האוצֵר מבט מקרוב</SectionLabel>
<div className="grid grid-cols-2 md:grid-cols-4 gap-3">
<Kpi label="סך ממצאי-אוצֵר" value={c.total} icon={<Sparkles className="w-4 h-4" />} />
<Kpi label="החלטות שנסקרו" value={`${c.decisions_reviewed}/${data.decisions_total}`} icon={<FileText className="w-4 h-4" />} />
<Kpi label="מאושרים · זורמים לכותב" value={c.approved} tone="success" icon={<CheckCircle2 className="w-4 h-4" />} />
<Kpi label="ממתינים לאישור" value={c.proposed} tone="gold" icon={<Brain className="w-4 h-4" />} />
</div>
</div>
</div> </div>
); );
} }
function Kpi({ function Kpi({
label, value, icon, label, value, icon, tone,
}: { label: string; value: string | number; icon: React.ReactNode }) { }: { label: string; value: string | number; icon: React.ReactNode; tone?: "success" | "gold" }) {
const valueCls = tone === "success" ? "text-success" : tone === "gold" ? "text-gold-deep" : "text-navy";
return ( return (
<Card className="bg-surface border-rule"> <Card className="bg-surface border-rule">
<CardContent className="px-4 py-3"> <CardContent className="px-4 py-3">
@@ -106,7 +204,7 @@ function Kpi({
{icon} {icon}
<span>{label}</span> <span>{label}</span>
</div> </div>
<p className="text-2xl text-navy font-semibold tabular-nums mt-1">{value}</p> <p className={`text-2xl font-semibold tabular-nums mt-1 ${valueCls}`}>{value}</p>
</CardContent> </CardContent>
</Card> </Card>
); );
@@ -123,10 +221,10 @@ function RecentFindings() {
if (!data || data.recent_findings.length === 0) { if (!data || data.recent_findings.length === 0) {
return ( return (
<Card className="bg-rule-soft/40 border-rule"> <Card className="bg-rule-soft/40 border-rule">
<CardContent className="px-6 py-5 text-center text-ink-muted text-sm"> <CardContent className="px-6 py-5 text-center text-ink-muted text-sm leading-relaxed">
אין עדיין ממצאים של ה-Curator. הוא מופעל אוטומטית כאשר דפנה מסמנת אין עדיין ממצאי-אוצֵר. כשדפנה מסמנת החלטה כסופית, האוצֵר קורא את הסופי, מזהה
החלטה כסופית (mark-final), ושומר את ממצאיו כ-decision_lessons עם 35 דפוסי-סגנון, ושומר אותם כ-<code>decision_lessons</code> (source=&quot;curator&quot;)
source=&quot;curator&quot;. הממתינים לאישורה בנפרד מלקחי-הפאנל ומכללי-המתודולוגיה שלמעלה.
</CardContent> </CardContent>
</Card> </Card>
); );

View File

@@ -197,7 +197,7 @@ export function StyleReportPanel() {
<h3 className="text-navy text-lg m-0">אוצֵר ממצאי-סגנון</h3> <h3 className="text-navy text-lg m-0">אוצֵר ממצאי-סגנון</h3>
<div className="flex items-center gap-4 rounded-lg border border-success bg-success-bg px-[18px] py-3.5"> <div className="flex items-center gap-4 rounded-lg border border-success bg-success-bg px-[18px] py-3.5">
<span className="text-success font-bold text-[1.75rem] leading-none tabular-nums"> <span className="text-success font-bold text-[1.75rem] leading-none tabular-nums">
{curator.data ? curator.data.findings_approved : "—"} {curator.data ? curator.data.channels.curator.approved : "—"}
</span> </span>
<div className="min-w-0"> <div className="min-w-0">
<b className="text-navy text-sm font-semibold block"> <b className="text-navy text-sm font-semibold block">
@@ -214,8 +214,8 @@ export function StyleReportPanel() {
{curator.data {curator.data
? Math.max( ? Math.max(
0, 0,
curator.data.total_findings - curator.data.channels.curator.total -
curator.data.findings_approved, curator.data.channels.curator.approved,
) )
: "—"} : "—"}
</div> </div>
@@ -223,7 +223,7 @@ export function StyleReportPanel() {
</div> </div>
<div className="flex-1 rounded-lg border border-rule bg-rule-soft px-3.5 py-3"> <div className="flex-1 rounded-lg border border-rule bg-rule-soft px-3.5 py-3">
<div className="text-ink-muted font-bold text-[1.35rem] tabular-nums leading-none"> <div className="text-ink-muted font-bold text-[1.35rem] tabular-nums leading-none">
{curator.data ? curator.data.total_findings : "—"} {curator.data ? curator.data.channels.curator.total : "—"}
</div> </div>
<div className="text-[0.78rem] text-ink-soft mt-1">סך ממצאים</div> <div className="text-[0.78rem] text-ink-soft mt-1">סך ממצאים</div>
</div> </div>

View File

@@ -179,3 +179,25 @@ export function useSubmitInteraction(caseNumber: string | undefined) {
}, },
}); });
} }
export type AgentResetResult = {
ok: boolean;
reassigned_issues: { id: string; identifier: string }[];
reset_agents: { id: string; name: string; ok: boolean; error?: string }[];
};
export function useResetCaseAgents(caseNumber: string | undefined) {
const qc = useQueryClient();
return useMutation({
mutationFn: () =>
apiRequest<AgentResetResult>(
`/api/cases/${caseNumber}/agents/reset`,
{ method: "POST" },
),
onSuccess: () => {
if (caseNumber) {
qc.invalidateQueries({ queryKey: agentKeys.activity(caseNumber) });
}
},
});
}

View File

@@ -74,6 +74,7 @@ const LEGACY_STATUS_LABELS: Record<string, string> = {
ready_for_writing: "מוכן לכתיבה", ready_for_writing: "מוכן לכתיבה",
drafting: "בכתיבה", drafting: "בכתיבה",
qa_failed: "בדיקת איכות נכשלה", qa_failed: "בדיקת איכות נכשלה",
qa_passed: "טיוטה",
}; };
/** /**

View File

@@ -176,14 +176,16 @@ export function useUpdateCase(caseNumber: string | undefined) {
body: input, body: input,
}), }),
onSuccess: (data) => { onSuccess: (data) => {
/* Patch cached detail and nudge the list to refetch on next focus */ /* Patch cached detail, then invalidate only the LIST queries — not
* casesKeys.all, which would also invalidate the detail we just patched
* and throw away the optimistic merge. */
if (caseNumber) { if (caseNumber) {
qc.setQueryData<CaseDetail | undefined>( qc.setQueryData<CaseDetail | undefined>(
casesKeys.detail(caseNumber), casesKeys.detail(caseNumber),
(prev) => (prev ? { ...prev, ...data } : prev), (prev) => (prev ? { ...prev, ...data } : prev),
); );
} }
qc.invalidateQueries({ queryKey: casesKeys.all }); qc.invalidateQueries({ queryKey: [...casesKeys.all, "list"] });
}, },
}); });
} }

View File

@@ -0,0 +1,103 @@
/**
* Citation-verification domain (X11 Phase 2 / #154) — the compose "אימות פסיקה" tab.
*
* Per legal argument: in-corpus supporting precedents with the cumulative authority
* signal (cited_by — followed/distinguished), their verify state + chair note, and
* per-issue radar (unlinked digests). The chair verifies before the writer cites
* (INV-AH). Backed by GET/POST /api/cases/{n}/citation-verification[/verify].
*/
import { useMutation, useQuery, useQueryClient } from "@tanstack/react-query";
import { apiRequest } from "./client";
export type CitedBy = {
total: number;
positive: number;
negative: number;
unclassified: number;
by_treatment: Record<string, number>;
};
export type SupportingPrecedent = {
case_law_id: string;
case_number: string;
case_name: string;
quote: string;
score: number;
cited_by: CitedBy;
attached_id: string | null;
verified: boolean;
chair_note: string;
};
export type RadarLead = {
digest_id: string;
yomon_number?: string | number | null;
headline: string;
underlying_citation: string;
underlying_court: string;
score: number;
missing_precedent_id: string | null;
missing_precedent_status: string | null;
action: "new_lead" | "gap_open" | "fetched" | "available_link" | string;
matched_issues?: string[];
};
export type ArgumentBlock = {
argument_id: string;
title: string;
legal_topic: string;
priority: string;
party: string;
supporting: SupportingPrecedent[];
radar: RadarLead[];
};
export type CitationVerificationView = {
status: string;
case_number: string;
arguments: ArgumentBlock[];
summary: {
arguments_total: number;
arguments_with_support: number;
verified: number;
radar_leads: number;
};
};
export type VerifyCitationBody = {
argument_id: string;
case_law_id?: string;
quote?: string;
citation?: string;
chair_note?: string | null;
verified: boolean;
attached_id?: string;
};
export function useCitationVerification(caseNumber: string | undefined) {
return useQuery({
queryKey: ["citation-verification", caseNumber ?? ""],
queryFn: ({ signal }) =>
apiRequest<CitationVerificationView>(
`/api/cases/${caseNumber}/citation-verification`,
{ signal },
),
enabled: Boolean(caseNumber),
staleTime: 10_000,
});
}
export function useVerifyCitation(caseNumber: string | undefined) {
const qc = useQueryClient();
return useMutation({
mutationFn: (body: VerifyCitationBody) =>
apiRequest(`/api/cases/${caseNumber}/citation-verification/verify`, {
method: "POST",
body,
}),
onSuccess: () => {
qc.invalidateQueries({ queryKey: ["citation-verification", caseNumber ?? ""] });
},
});
}

View File

@@ -9,7 +9,8 @@ import { apiRequest } from "./client";
export type LinkedCitation = { export type LinkedCitation = {
citation: string; citation: string;
cited_id: string; cited_id: string;
case_name: string; /** May be null/empty — the panel guards with `c.case_name && …`. */
case_name: string | null;
court: string; court: string;
precedent_level: string; precedent_level: string;
}; };

View File

@@ -71,12 +71,30 @@ export function useLegalArguments(caseNumber: string | undefined) {
}); });
} }
export type AggregateArgumentsResult = { /**
status: "started" | string; * The aggregation runs on the legal-analyst agent (host-side, where the
case_number: string; * `claude` CLI lives) — NOT inline in the FastAPI container. The endpoint
force: boolean; * either delegates via a Paperclip wakeup (`queued`) or short-circuits on a
message: string; * cheap in-container DB pre-check (`no_claims` / `exists`). `skipped` means
}; * no analyst route was available; the chair can run the MCP tool manually.
*/
export type AggregateArgumentsResult =
| {
status: "queued";
sub_issue_id: string;
analyst_id: string;
main_issue_id: string;
}
| {
status: "no_claims" | "exists";
total: number;
message: string;
}
| {
status: "skipped";
reason: "no_api_key" | "no_analyst" | "no_issue" | string;
company_id?: string;
};
export function useAggregateArguments(caseNumber: string | undefined) { export function useAggregateArguments(caseNumber: string | undefined) {
const qc = useQueryClient(); const qc = useQueryClient();

View File

@@ -56,6 +56,10 @@ export type MissingPrecedent = {
discovery_source: string | null; // manual | cited_only | digest | court_fetch discovery_source: string | null; // manual | cited_only | digest | court_fetch
cited_by_precedents: string[] | null; // corpus precedents citing a cited_only stub cited_by_precedents: string[] | null; // corpus precedents citing a cited_only stub
yomon_number: string | null; // digest gap's source yomon number yomon_number: string | null; // digest gap's source yomon number
// Bridge to the corpus citation graph: committee decisions (and their chairs)
// that cite this still-missing ruling — fills "צוטט ע״י <יו״ר>" + "תיק".
cited_by_chairs: string[] | null;
cited_by_decisions: string[] | null;
}; };
export type MissingPrecedentListResponse = { export type MissingPrecedentListResponse = {

View File

@@ -143,10 +143,24 @@ export type RelatedCase = {
relation_type: string; relation_type: string;
}; };
/** A decision that cites THIS ruling — auto-detected from the citation graph
* (precedent_internal_citations incoming edges). Read-only (no unlink). */
export type IncomingCitation = {
id: string;
case_number: string;
case_name: string;
court: string;
precedent_level: string;
chair_name: string;
date: string | null;
confidence: number | null;
};
export type PrecedentDetail = Precedent & { export type PrecedentDetail = Precedent & {
full_text: string; full_text: string;
halachot: Halacha[]; halachot: Halacha[];
related_cases: RelatedCase[]; related_cases: RelatedCase[];
incoming_citations: IncomingCitation[];
}; };
export type SearchHit = export type SearchHit =

View File

@@ -302,14 +302,39 @@ export type CuratorFinding = {
created_at: string; created_at: string;
}; };
// One decision_lessons channel split by the INV-LRN1 review gate.
export type LessonChannelStats = {
total: number;
approved: number;
proposed: number;
rejected: number;
};
// The three honest learning channels that feed the writer (INV-IA2/IA5).
// Earlier this was a single flat count of source='curator' — which nothing
// wrote, so it always read 0 while the panel/methodology channels carried all
// the real learning. See get_curator_stats (web/app.py).
export type CuratorStats = { export type CuratorStats = {
total_findings: number;
decisions_with_findings: number;
decisions_total: number; decisions_total: number;
// LRN-5 (INV-IA5): the count of curator lessons with review_status='approved' channels: {
// (the real INV-LRN1 writer gate) — replaces the old findings_applied, which // A — Claude draft↔final distillation → methodology overrides; flows now.
// counted the informative-only applied_to_skill flag. distillation: {
findings_approved: number; discussion_rules: number;
transition_phrases: number;
items_total: number;
pairs_folded: number;
pairs_total: number;
};
// B — DeepSeek+Gemini panel → decision_lessons; flows when chair-approved.
panel: LessonChannelStats & {
decisions: number;
by_practice_area: Record<string, number>;
};
// C — curator agent's qualitative findings → decision_lessons (source='curator').
curator: LessonChannelStats & {
decisions_reviewed: number;
};
};
recent_findings: CuratorFinding[]; recent_findings: CuratorFinding[];
}; };

View File

@@ -52,10 +52,12 @@ from web.paperclip_client import (
post_comment as pc_post_comment, post_comment as pc_post_comment,
reject_interaction as pc_reject_interaction, reject_interaction as pc_reject_interaction,
reset_agent_session as pc_reset_agent_session, reset_agent_session as pc_reset_agent_session,
reset_case_agents as pc_reset_case_agents,
respond_to_interaction as pc_respond_to_interaction, respond_to_interaction as pc_respond_to_interaction,
restore_project as pc_restore_project, restore_project as pc_restore_project,
update_project_name as pc_update_project_name, update_project_name as pc_update_project_name,
wake_analyst_for_appraiser_facts as pc_wake_analyst_for_appraiser_facts, wake_analyst_for_appraiser_facts as pc_wake_analyst_for_appraiser_facts,
wake_analyst_for_argument_aggregation as pc_wake_analyst_for_argument_aggregation,
wake_ceo_agent as pc_wake_ceo, wake_ceo_agent as pc_wake_ceo,
wake_ceo_for_feedback_fold as pc_wake_ceo_for_feedback_fold, wake_ceo_for_feedback_fold as pc_wake_ceo_for_feedback_fold,
wake_curator_for_final as pc_wake_curator_for_final, wake_curator_for_final as pc_wake_curator_for_final,
@@ -99,6 +101,7 @@ __all__ = [
"pc_wake_curator_for_final", "pc_wake_curator_for_final",
"pc_wake_for_precedent_extraction", "pc_wake_for_precedent_extraction",
"pc_wake_analyst_for_appraiser_facts", "pc_wake_analyst_for_appraiser_facts",
"pc_wake_analyst_for_argument_aggregation",
# comments / interactions # comments / interactions
"pc_post_comment", "pc_post_comment",
"pc_get_issue_comments", "pc_get_issue_comments",
@@ -112,4 +115,5 @@ __all__ = [
"pc_get_run_events", "pc_get_run_events",
"pc_cancel_run", "pc_cancel_run",
"pc_reset_agent_session", "pc_reset_agent_session",
"pc_reset_case_agents",
] ]

View File

@@ -72,9 +72,11 @@ from web.agent_platform_port import (
pc_reject_interaction, pc_reject_interaction,
pc_request, pc_request,
pc_reset_agent_session, pc_reset_agent_session,
pc_reset_case_agents,
pc_respond_to_interaction, pc_respond_to_interaction,
pc_restore_project, pc_restore_project,
pc_wake_analyst_for_appraiser_facts, pc_wake_analyst_for_appraiser_facts,
pc_wake_analyst_for_argument_aggregation,
rename_case_project, rename_case_project,
pc_wake_ceo, pc_wake_ceo,
pc_wake_ceo_for_feedback_fold, pc_wake_ceo_for_feedback_fold,
@@ -1316,31 +1318,77 @@ async def get_style_analyzer_prompt():
@app.get("/api/training/curator/stats") @app.get("/api/training/curator/stats")
async def get_curator_stats(): async def get_curator_stats():
"""Cheap aggregate stats over decision_lessons + style_corpus. """Aggregate over the THREE learning channels that feed the writer, for the
/training "אוצֵר" tab. Each channel is surfaced honestly with its real
backing store and how much of it actually reaches the writer:
Used by the Curator-Portrait tab to show "10 curator findings across 24 A. distillation — Claude draft↔final diff → methodology overrides
decisions". We deliberately keep this server-side and aggregate so the (appeal_type_rules['_global']). Folded items flow to the writer now.
UI can render a single card without fanning out N queries. B. panel — DeepSeek+Gemini 2/2 vote → decision_lessons
(source='panel:deepseek+gemini'). Flow to the writer only once the
chair approves (review_status='approved', INV-LRN1/G10).
C. curator — the agent's Opus-read qualitative findings → decision_lessons
(source='curator'). Same chair gate as the panel.
Earlier this endpoint counted ONLY source='curator' and reported it as the
whole story, so it read 0 while the panel/methodology channels carried all
the real learning. The card was measuring a column nothing wrote (the agent
posted findings as Paperclip comments, never as decision_lessons — a
spec↔impl drift since closed). It now reports all three (INV-IA2/IA5).
""" """
pool = await db.get_pool() pool = await db.get_pool()
async with pool.acquire() as conn: async with pool.acquire() as conn:
total_lessons = await conn.fetchval( total_corpus = await conn.fetchval("SELECT count(*) FROM style_corpus") or 0
"SELECT count(*) FROM decision_lessons WHERE source = 'curator'"
# ── Channel A — distillation → methodology overrides ──────────────
meth_rows = await conn.fetch(
"SELECT rule_category, coalesce(jsonb_array_length(rule_value), 0) AS n "
"FROM appeal_type_rules "
"WHERE appeal_type = '_global' "
" AND rule_category IN ('discussion_rules', 'transition_phrases')"
) )
decisions_with_findings = await conn.fetchval( meth = {r["rule_category"]: r["n"] for r in meth_rows}
discussion_rules = meth.get("discussion_rules", 0)
transition_phrases = meth.get("transition_phrases", 0)
pairs_folded = await conn.fetchval(
"SELECT count(*) FROM draft_final_pairs WHERE status = 'lessons_folded'"
) or 0
pairs_total = await conn.fetchval("SELECT count(*) FROM draft_final_pairs") or 0
# ── Channels B & C — decision_lessons by source × review_status ───
ls_rows = await conn.fetch(
"SELECT source, review_status, count(*) AS n FROM decision_lessons "
"GROUP BY source, review_status"
)
def _channel(src_predicate) -> dict:
buckets = {"approved": 0, "proposed": 0, "rejected": 0}
for r in ls_rows:
if src_predicate(r["source"] or ""):
st = r["review_status"] or "proposed"
buckets[st] = buckets.get(st, 0) + r["n"]
buckets["total"] = sum(buckets.values())
return buckets
panel = _channel(lambda s: s.startswith("panel:"))
curator = _channel(lambda s: s == "curator")
panel_decisions = await conn.fetchval(
"SELECT count(DISTINCT style_corpus_id) FROM decision_lessons "
"WHERE source LIKE 'panel:%'"
) or 0
curator_decisions = await conn.fetchval(
"SELECT count(DISTINCT style_corpus_id) FROM decision_lessons " "SELECT count(DISTINCT style_corpus_id) FROM decision_lessons "
"WHERE source = 'curator'" "WHERE source = 'curator'"
) or 0
pa_rows = await conn.fetch(
"SELECT sc.practice_area, count(*) AS n "
"FROM decision_lessons dl JOIN style_corpus sc ON sc.id = dl.style_corpus_id "
"WHERE dl.source LIKE 'panel:%' GROUP BY sc.practice_area"
) )
total_corpus = await conn.fetchval("SELECT count(*) FROM style_corpus") panel_by_practice = {(r["practice_area"] or ""): r["n"] for r in pa_rows}
# LRN-5 (INV-IA5): count the *real* consumer-mapped gate — review_status
# 'approved' is what flows to the writer (INV-LRN1, #126). The old count # Recent curator findings (channel C) — newest first
# of applied_to_skill was a KPI over an informative-only flag (LRN-1)
# that writes nowhere, so it reported adoption that never happened.
approved = await conn.fetchval(
"SELECT count(*) FROM decision_lessons "
"WHERE source = 'curator' AND review_status = 'approved'"
)
# Last 10 curator findings — newest first
recent_rows = await conn.fetch( recent_rows = await conn.fetch(
""" """
SELECT dl.id, dl.lesson_text, dl.category, dl.review_status, SELECT dl.id, dl.lesson_text, dl.category, dl.review_status,
@@ -1353,11 +1401,27 @@ async def get_curator_stats():
LIMIT 10 LIMIT 10
""" """
) )
return { return {
"total_findings": total_lessons or 0, "decisions_total": total_corpus,
"decisions_with_findings": decisions_with_findings or 0, "channels": {
"decisions_total": total_corpus or 0, "distillation": {
"findings_approved": approved or 0, "discussion_rules": discussion_rules,
"transition_phrases": transition_phrases,
"items_total": discussion_rules + transition_phrases,
"pairs_folded": pairs_folded,
"pairs_total": pairs_total,
},
"panel": {
**panel,
"decisions": panel_decisions,
"by_practice_area": panel_by_practice,
},
"curator": {
**curator,
"decisions_reviewed": curator_decisions,
},
},
"recent_findings": [ "recent_findings": [
{ {
"id": str(r["id"]), "id": str(r["id"]),
@@ -2454,41 +2518,78 @@ async def api_get_claims(case_number: str):
# the FastAPI container it short-circuits with status="llm_unavailable". # the FastAPI container it short-circuits with status="llm_unavailable".
@app.post("/api/cases/{case_number}/aggregate-arguments") @app.post("/api/cases/{case_number}/aggregate-arguments")
async def api_aggregate_arguments( async def api_aggregate_arguments(case_number: str, force: bool = False):
case_number: str, """Queue claim→argument aggregation by waking the legal-analyst agent.
background_tasks: BackgroundTasks,
force: bool = False,
):
"""Aggregate raw claims into distinct legal arguments via Claude.
Runs as a BackgroundTask because the LLM pass can take 30-90 seconds. The aggregation itself calls `claude_session.query_json()`, which shells
out to the local `claude` CLI — present on the agent host, **absent in
this FastAPI container**. Running it inline as a BackgroundTask (the old
behaviour) silently produced nothing, and on `force` it destructively
deleted the existing arguments *before* the doomed LLM call. So we
delegate to the analyst exactly like `extract-appraiser-facts`: create a
child Paperclip issue, assign it to the company's analyst, and trigger a
wakeup. The analyst runs the MCP tool locally and posts results.
Cheap in-container pre-checks (DB only, no LLM) short-circuit before any
agent is spun up:
- `no_claims` — there are no raw claims to aggregate yet.
- `exists` — arguments already computed and `force` is False.
Response shape:
{"status": "queued", "sub_issue_id", "analyst_id", "main_issue_id"}
or {"status": "no_claims"|"exists", "total", "message"}
or {"status": "skipped", "reason": "no_api_key"|"no_analyst"|"no_issue"}
""" """
case = await db.get_case_by_number(case_number) case = await db.get_case_by_number(case_number)
if not case: if not case:
raise HTTPException(404, f"תיק {case_number} לא נמצא") raise HTTPException(404, f"תיק {case_number} לא נמצא")
async def _run() -> None: case_id = UUID(case["id"])
try: pool = await db.get_pool()
from legal_mcp.services import argument_aggregator async with pool.acquire() as conn:
result = await argument_aggregator.aggregate_claims_to_arguments( claim_count = await conn.fetchval(
UUID(case["id"]), force=force, "SELECT COUNT(*) FROM claims WHERE case_id = $1", case_id,
) )
logger.info( existing_args = await conn.fetchval(
"aggregate_arguments[%s] finished: %s", "SELECT COUNT(*) FROM legal_arguments WHERE case_id = $1", case_id,
case_number, result, )
)
except Exception as e: # noqa: BLE001
logger.exception(
"aggregate_arguments[%s] failed: %s", case_number, e,
)
background_tasks.add_task(_run) if not claim_count:
return { return {
"status": "started", "status": "no_claims",
"case_number": case_number, "total": 0,
"force": force, "message": (
"message": "Aggregation started in background. Poll /legal-arguments for results.", "אין טענות גולמיות בתיק. הרץ קודם חילוץ טענות (extract_claims) "
} "ואז חשב טיעונים."
),
}
if existing_args and not force:
return {
"status": "exists",
"total": existing_args,
"message": (
f"כבר קיימים {existing_args} טיעונים מאוגדים. השתמש בכפתור "
"החישוב-מחדש כדי לחשב מחדש (מוחק ובונה מחדש)."
),
}
# Route to the analyst of the correct company by case-number prefix.
prefix = case_number[:1]
company_id = (
PAPERCLIP_COMPANIES["licensing"] if prefix == "1"
else PAPERCLIP_COMPANIES["betterment"] if prefix in ("8", "9")
else ""
)
try:
result = await pc_wake_analyst_for_argument_aggregation(
case_number, company_id=company_id, force=force,
)
except Exception as e:
logger.exception("analyst wakeup failed for argument aggregation %s", case_number)
raise HTTPException(500, f"לא ניתן לשלוח לאנליטיקאי: {e}")
return result
@app.get("/api/cases/{case_number}/legal-arguments") @app.get("/api/cases/{case_number}/legal-arguments")
@@ -3155,6 +3256,62 @@ async def api_precedent_list(case_number: str):
return envelope_unwrap(parsed) return envelope_unwrap(parsed)
# ── Citation-verification panel (X11 Phase 2 / #154) — "אימות פסיקה" tab ──
@app.get("/api/cases/{case_number}/citation-verification")
async def api_citation_verification(case_number: str):
"""Per legal-argument: supporting corpus precedents (with cited_by authority),
their verify state + chair note, and per-issue radar (unlinked digests). Powers
the compose 'אימות פסיקה' tab — the chair verifies before the writer cites."""
from legal_mcp.services import case_citation_verification as ccv
view = await ccv.build_view(case_number)
if view.get("status") == "case_not_found":
raise HTTPException(404, f"תיק {case_number} לא נמצא")
return view
class CitationVerifyRequest(BaseModel):
argument_id: str
case_law_id: str = ""
quote: str = ""
citation: str = ""
chair_note: str | None = None
verified: bool = True
attached_id: str = "" # set when the precedent row already exists
@app.post("/api/cases/{case_number}/citation-verification/verify")
async def api_citation_verify(case_number: str, req: CitationVerifyRequest):
"""Verify / un-verify a precedent for a specific argument (the INV-AH gate the
writer respects). Upsert: PATCH an existing attachment, else attach a new one
linked to the argument + corpus ruling and mark it."""
case = await db.get_case_by_number(case_number)
if not case:
raise HTTPException(404, f"תיק {case_number} לא נמצא")
try:
if req.attached_id:
row = await db.set_case_precedent_verified(
UUID(req.attached_id), req.verified, req.chair_note)
if not row:
raise HTTPException(404, "שיוך-פסיקה לא נמצא")
return row
if not req.quote.strip() or not req.citation.strip():
raise HTTPException(400, "quote ו-citation חובה לצירוף חדש")
cid = case["id"]
row = await db.create_case_precedent(
case_id=UUID(cid) if isinstance(cid, str) else cid,
quote=req.quote, citation=req.citation,
chair_note=req.chair_note or "",
argument_id=UUID(req.argument_id) if req.argument_id else None,
case_law_id=UUID(req.case_law_id) if req.case_law_id else None,
verified=req.verified,
)
return row
except ValueError:
raise HTTPException(400, "מזהה לא תקין")
@app.delete("/api/precedents/{precedent_id}") @app.delete("/api/precedents/{precedent_id}")
async def api_precedent_delete(precedent_id: str): async def api_precedent_delete(precedent_id: str):
"""Delete a precedent attachment. The archived PDF (if any) stays """Delete a precedent attachment. The archived PDF (if any) stays
@@ -3548,9 +3705,11 @@ async def _enroll_final_in_library(
out["error"] = "no final text extracted" out["error"] = "no final text extracted"
return out return out
# Deterministic metadata from the case record — the Gemini metadata extractor is # Deterministic seeds from the case record first — proceeding_type / date /
# tuned for EXTERNAL rulings and returns no_metadata for internal decisions, so we # citation are identity, not inference, so we set them ourselves (no LLM) and the
# populate proceeding_type / date / tags / summary / citation ourselves (no LLM). # later Gemini pass (reextract_metadata, below) is fills-empty-only and won't
# touch them. subject_tags/summary seeded here from the case (chair-curated
# subject_categories win when present); when empty, Gemini fills them inline.
district = "ירושלים" district = "ירושלים"
proceeding_type = (case.get("proceeding_type") or "ערר").strip() proceeding_type = (case.get("proceeding_type") or "ערר").strip()
decision_date = case.get("decision_date") or case.get("hearing_date") decision_date = case.get("decision_date") or case.get("hearing_date")
@@ -3590,6 +3749,42 @@ async def _enroll_final_in_library(
except Exception as e: except Exception as e:
logger.warning("citation build failed for %s: %s", case_number, e) logger.warning("citation build failed for %s: %s", case_number, e)
# Fill the LLM-derived metadata (subject_tags, summary, headnote, key_quote)
# inline via Gemini (REST — container-safe; GOOGLE_GEMINI_API_KEY in Coolify),
# reusing the same reextract_metadata path the UI/drain use (G2 — one path, full
# status lifecycle). apply_to_record fills ONLY empty fields, so the deterministic
# seeds above (proceeding_type / date / citation, plus any chair-curated
# subject_categories) are preserved. Without this the row sat at
# metadata_extraction_status='pending' with empty tags until a drain happened to
# pick it up, so a freshly-enrolled final showed no subject tags in the edit UI.
# Best-effort — surfaced in the return value, never silently swallowed.
try:
from legal_mcp.services import precedent_library as plib_service
meta = await plib_service.reextract_metadata(UUID(case_law_id))
out["metadata"] = {
"status": meta.get("status"),
"fields": meta.get("fields") or [],
}
except Exception as e:
logger.warning("inline metadata extraction failed for %s: %s", case_number, e)
out["metadata_error"] = str(e)
# Grow the style-exemplar corpus (channel B): break this final into block-level
# paragraphs the writer retrieves (07-learning §0.2). Before this, exemplars were
# frozen at the one-time seed backfill — new finals never got exemplarized, so the
# richest style channel never grew. Same extraction the backfill uses (G2). Voyage
# embeds over REST → container-safe. Best-effort; surfaced, never fails the upload.
try:
from legal_mcp.services import style_exemplars as _sx
n_ex = await _sx.extract_and_store(
decision_number=case_number, source="internal_committee",
full_text=final_text, practice_area=case.get("practice_area", ""),
)
out["exemplars"] = n_ex
except Exception as e:
logger.warning("style-exemplar extraction failed for %s: %s", case_number, e)
out["exemplars_error"] = str(e)
# The precedents this decision cites → link to the library; flag the ones not found. # The precedents this decision cites → link to the library; flag the ones not found.
try: try:
await cit_tools.extract_internal_citations(case_law_id=case_law_id, limit=0) await cit_tools.extract_internal_citations(case_law_id=case_law_id, limit=0)
@@ -3718,6 +3913,25 @@ async def api_upload_final_decision(case_number: str, file: UploadFile = File(..
# checking + missing-precedent flagging (07-learning §1.3). Surfaced, not silent. # checking + missing-precedent flagging (07-learning §1.3). Surfaced, not silent.
library = await _enroll_final_in_library(case, case_number, final_text, chair_name) library = await _enroll_final_in_library(case, case_number, final_text, chair_name)
# Path A — prospective held-out snapshot. Right now the draft (decision_blocks)
# was written with only the PRIOR lesson pool — this case's lessons fold later,
# manually, in /training. So measuring style-distance HERE is a clean
# generalization datapoint. Append-only; best-effort (never fails the upload).
try:
from legal_mcp.services import style_distance as _sd
sd = await _sd.style_distance(case_number)
if isinstance(sd, dict) and "error" not in sd:
pool = await db.voice_lesson_pool_sizes()
summ = sd.get("summary", {})
await db.record_style_distance_snapshot(
case_number, pair_id,
pool.get("discussion_rules", 0), pool.get("transition_phrases", 0),
summ.get("anti_pattern_total"), summ.get("ratio_max_deviation_pp"),
summ.get("change_percent"),
)
except Exception as e:
logger.warning("held-out style-distance snapshot failed for %s: %s", case_number, e)
case_dir = config.find_case_dir(case_number) case_dir = config.find_case_dir(case_number)
if case_dir.exists(): if case_dir.exists():
commit_and_push(case_dir, f"החלטה סופית של היו\"ר: {final_name}") commit_and_push(case_dir, f"החלטה סופית של היו\"ר: {final_name}")
@@ -4079,6 +4293,19 @@ async def api_post_interaction_response(
raise HTTPException(502, f"שגיאת Paperclip: {e}") raise HTTPException(502, f"שגיאת Paperclip: {e}")
@app.post("/api/cases/{case_number}/agents/reset")
async def api_reset_case_agents(case_number: str):
"""Reset stuck agents for a case.
Clears writer/QA agents from 'error' status and reassigns any open
issues back to the chair user, stopping Paperclip recovery loops.
"""
result = await pc_reset_case_agents(case_number)
if not result.get("ok"):
raise HTTPException(502, result.get("error", "שגיאה בביצוע האיפוס"))
return result
# ── Settings: MCP Server Configuration ──────────────────────────── # ── Settings: MCP Server Configuration ────────────────────────────
# #
# Source of truth for legal-ai env vars is Coolify (see memory: # Source of truth for legal-ai env vars is Coolify (see memory:
@@ -4630,6 +4857,15 @@ async def api_learning_style_distance(case_number: str):
return await _sd.style_distance(case_number) return await _sd.style_distance(case_number)
@app.get("/api/learning/style-distance-history")
async def api_learning_style_distance_history():
"""Path A — מגמת ה-held-out הפרוספקטיבי: snapshot של מרחק-הסגנון שנלכד בכל
העלאת-סופי *לפני* הטמעת-לקחי-התיק, מול גודל-בריכת-הלקחים באותו רגע. ירידה
ב-anti_pattern_total / change_percent ככל שהבריכה גדלה = הלמידה מכלילה."""
items = await db.get_style_distance_history()
return {"items": items, "count": len(items)}
def _coerce_json(raw): def _coerce_json(raw):
if isinstance(raw, str): if isinstance(raw, str):
try: try:
@@ -4670,36 +4906,13 @@ class PromoteLearningRequest(BaseModel):
async def _append_methodology_override(category: str, key: str, items: list[str]) -> None: async def _append_methodology_override(category: str, key: str, items: list[str]) -> None:
"""Read current (override-or-default) list value, append new items, upsert override. """Thin wrapper over db.append_global_rule (the single locked append impl, G2) —
Shared by the T14 approval gate to fold approved learnings into writer-consumed channels. seeds from _METHODOLOGY_DEFAULTS when no override row exists yet. Shared by the
T14 promote gate; chair-feedback auto-flow calls db.append_global_rule directly."""
MET-2/3 (INV-IA3): the read-modify-write runs inside ONE transaction with the await db.append_global_rule(
existing override row locked FOR UPDATE, so a concurrent promote (or a methodology category, key, items,
PUT) can't interleave between the read and the write and silently drop items. seed_if_missing=list(_METHODOLOGY_DEFAULTS.get(category, {}).get(key, [])),
The methodology PUT overwrites the same row; the MET-1 invalidation (גל-1) makes )
the /methodology editor refetch on promote so it edits post-append state."""
pool = await db.get_pool()
async with pool.acquire() as conn:
async with conn.transaction():
row = await conn.fetchrow(
"SELECT rule_value FROM appeal_type_rules "
"WHERE appeal_type = '_global' AND rule_category = $1 AND rule_key = $2 "
"FOR UPDATE",
category, key,
)
if row:
current = _coerce_json(row["rule_value"]) or []
else:
current = list(_METHODOLOGY_DEFAULTS.get(category, {}).get(key, []))
if not isinstance(current, list):
current = []
merged = current + [s for s in items if s and s not in current]
await conn.execute(
"INSERT INTO appeal_type_rules (id, appeal_type, rule_category, rule_key, rule_value) "
"VALUES (gen_random_uuid(), '_global', $1, $2, $3::text::jsonb) "
"ON CONFLICT (appeal_type, rule_category, rule_key) DO UPDATE SET rule_value = $3::text::jsonb",
category, key, json.dumps(merged, ensure_ascii=False),
)
@app.post("/api/learning/pairs/{pair_id}/promote") @app.post("/api/learning/pairs/{pair_id}/promote")
@@ -6444,6 +6657,17 @@ class DigestLinkRequest(BaseModel):
case_law_id: str case_law_id: str
@app.get("/api/cases/{case_number}/digest-radar")
async def api_case_digest_radar(case_number: str, limit: int = 5, min_score: float = 0.45):
"""Case-contextual digest radar (X12) — UNLINKED digests whose topic is close to
this case (rulings we don't hold yet). Powers the case-page "📡 רדאר יומונים" lead
so a relevant ruling known only via a digest doesn't fall through the cracks while
the case is decided. INV-DIG1: points at the underlying ruling, never cites the
digest."""
return await digest_service.case_digest_radar(
case_number, limit=max(1, min(int(limit), 20)), min_score=float(min_score))
@app.post("/api/digests/upload") @app.post("/api/digests/upload")
async def digest_upload( async def digest_upload(
file: UploadFile = File(...), file: UploadFile = File(...),
@@ -7141,7 +7365,7 @@ async def internal_decisions_upload(
practice_area: str = Form(""), practice_area: str = Form(""),
appeal_subtype: str = Form(""), appeal_subtype: str = Form(""),
subject_tags: str = Form("[]"), subject_tags: str = Form("[]"),
is_binding: bool = Form(True), is_binding: bool = Form(False), # INV-DM7: committee = persuasive (coerced at db too)
summary: str = Form(""), summary: str = Form(""),
): ):
"""Upload a planning appeals-committee decision to the internal corpus. """Upload a planning appeals-committee decision to the internal corpus.
@@ -7877,7 +8101,7 @@ async def missing_precedents_list(
case_id=case_uuid, case_id=case_uuid,
legal_topic=legal_topic.strip() or None, legal_topic=legal_topic.strip() or None,
q=q.strip() or None, q=q.strip() or None,
limit=max(1, min(int(limit), 500)), limit=max(1, min(int(limit), 2000)),
offset=max(0, int(offset)), offset=max(0, int(offset)),
) )
# Counters useful for the sidebar badge. # Counters useful for the sidebar badge.

View File

@@ -791,6 +791,71 @@ async def reset_agent_session(agent_id: str) -> dict:
return resp.json() return resp.json()
CHAIM_USER_ID = "ZpDWXxFweC3MftuF1Ttyu2VUbSPHbKZd"
async def reset_case_agents(case_number: str) -> dict:
"""Reset agent state for a case: clear error status + reassign stuck issues to user.
Two actions:
1. Any non-completed issue still assigned to an agent is reassigned to the chair user,
stopping Paperclip's source_scoped_recovery_action loop.
2. Every agent in the case's company whose global status is 'error' gets a
reset_agent_session call (clears wedged runtime) and its DB status set to 'idle'.
"""
first_digit = case_number.split("-")[0][0] if case_number else "1"
company_id = COMPANIES["betterment"] if first_digit in ("8", "9") else COMPANIES["licensing"]
conn = await asyncpg.connect(PAPERCLIP_DB_URL)
try:
project = await conn.fetchrow(
"SELECT id FROM projects WHERE name LIKE $1 LIMIT 1",
f"%{case_number}%",
)
if not project:
return {"ok": False, "error": f"No Paperclip project found for {case_number}"}
project_id = project["id"]
reassigned = await conn.fetch(
"""UPDATE issues
SET assignee_agent_id = null, assignee_user_id = $1, updated_at = now()
WHERE project_id = $2
AND assignee_agent_id IS NOT NULL
AND status NOT IN ('done', 'cancelled')
RETURNING id, identifier""",
CHAIM_USER_ID, project_id,
)
error_agents = await conn.fetch(
"SELECT id, name FROM agents WHERE company_id = $1::uuid AND status = 'error'",
company_id,
)
if error_agents:
await conn.execute(
"UPDATE agents SET status = 'idle' WHERE id = ANY($1::uuid[])",
[r["id"] for r in error_agents],
)
finally:
await conn.close()
reset_results = []
for agent in error_agents:
try:
await reset_agent_session(str(agent["id"]))
reset_results.append({"id": str(agent["id"]), "name": agent["name"], "ok": True})
except Exception as e:
logger.warning("reset_agent_session failed for %s: %s", agent["id"], e)
reset_results.append({"id": str(agent["id"]), "name": agent["name"], "ok": False, "error": str(e)})
return {
"ok": True,
"reassigned_issues": [{"id": str(r["id"]), "identifier": r["identifier"]} for r in reassigned],
"reset_agents": reset_results,
}
async def respond_to_interaction( async def respond_to_interaction(
issue_id: str, interaction_id: str, payload: dict, issue_id: str, interaction_id: str, payload: dict,
) -> dict: ) -> dict:
@@ -1371,3 +1436,121 @@ async def wake_analyst_for_appraiser_facts(
"analyst_id": analyst_id, "analyst_id": analyst_id,
"main_issue_id": main_issue_id, "main_issue_id": main_issue_id,
} }
async def wake_analyst_for_argument_aggregation(
case_number: str,
company_id: str,
force: bool = False,
) -> dict:
"""Wake the legal-analyst to aggregate raw claims into legal arguments.
Triggered by the chair clicking "חשב טיעונים" in the case page. The
FastAPI container cannot run `aggregate_claims_to_arguments` directly —
the aggregator calls `claude_session.query_json()`, which only works
where the local `claude` CLI is present (the MCP server / agent runner
on the host), **not in this container**. So instead of an in-container
BackgroundTask (which fails silently and, on `force`, destructively
deletes the existing arguments before the doomed LLM call), we create a
child issue under the case's main Paperclip issue, assign it to the
analyst of the correct company, and trigger a wakeup. The analyst's
HEARTBEAT picks up the issue, runs the MCP tool locally, and reports
back via a comment.
Mirrors ``wake_analyst_for_appraiser_facts`` — same delegation shape.
Returns a dict shaped for the FastAPI endpoint to serialize as-is:
{"status": "queued", "sub_issue_id", "analyst_id", "main_issue_id"}
or {"status": "skipped", "reason": "..."} for non-fatal early outs.
"""
if not PAPERCLIP_BOARD_API_KEY:
logger.warning(
"PAPERCLIP_BOARD_API_KEY not set — cannot queue analyst wakeup "
"for argument aggregation on %s",
case_number,
)
return {"status": "skipped", "reason": "no_api_key"}
analyst_id = ANALYST_AGENTS.get(company_id)
if not analyst_id:
logger.info("No analyst configured for company %s — skipping", company_id)
return {"status": "skipped", "reason": "no_analyst", "company_id": company_id}
issues = await get_case_issues(case_number)
if not issues:
logger.warning(
"No Paperclip issues found for case %s — cannot queue analyst", case_number,
)
return {"status": "skipped", "reason": "no_issue"}
main_issue = next((i for i in issues if i.get("status") == "in_progress"), None) or issues[0]
main_issue_id = main_issue["id"]
force_clause = ", force=True" if force else ""
rerun_note = (
"זהו חישוב-מחדש (force) — הטיעונים הקיימים יימחקו ויחושבו מחדש.\n\n"
if force else ""
)
description = (
f"חיים ביקש חישוב טיעונים משפטיים בתיק {case_number}.\n\n"
f"{rerun_note}"
f"הרץ `mcp__legal-ai__aggregate_claims_to_arguments(case_number=\"{case_number}\"{force_clause})` "
f"וכתוב comment בעברית עם תוצאת האיגוד — כמה טיעונים מובחנים נוצרו לכל צד "
f"(עוררים / משיבים / ועדה / מבקשי-היתר). אם אין טענות גולמיות בתיק, דווח "
f"ב-comment שצריך להריץ קודם חילוץ טענות (`extract_claims`) וסגור את ה-issue כ-blocked."
)
child_resp = await pc_request(
"POST",
f"/api/issues/{main_issue_id}/children",
json={
"title": f"[ערר {case_number}] חישוב טיעונים משפטיים",
"description": description,
"status": "in_progress",
# Paperclip ISSUE_PRIORITIES = critical|high|medium|low — "normal"
# is NOT a valid enum value (Zod 400 → surfaces as 500).
"priority": "medium",
"assigneeAgentId": analyst_id,
},
raise_on_error=True,
)
sub_issue = child_resp.json()
sub_issue_id = sub_issue["id"]
# Tag plugin_state so the case page surfaces this sub-issue too.
try:
conn = await asyncpg.connect(PAPERCLIP_DB_URL)
try:
await _link_case_to_issue(conn, sub_issue_id, case_number)
finally:
await conn.close()
except Exception as e:
logger.warning("plugin_state link failed for sub_issue=%s: %s", sub_issue_id, e)
wake_resp = await pc_request(
"POST",
f"/api/agents/{analyst_id}/wakeup",
json={
"source": "on_demand",
"triggerDetail": "manual",
"reason": f"aggregate_arguments_{case_number}",
# "assignment" is the generic mutation the HEARTBEAT recognises;
# task intent lives in the child-issue description, not the payload.
"payload": {
"issueId": sub_issue_id,
"mutation": "assignment",
"caseNumber": case_number,
},
},
raise_on_error=True,
)
logger.info(
"Analyst wakeup for argument aggregation on case %s: sub_issue=%s "
"analyst=%s force=%s wake=%s",
case_number, sub_issue_id, analyst_id, force, wake_resp.status_code,
)
return {
"status": "queued",
"sub_issue_id": sub_issue_id,
"analyst_id": analyst_id,
"main_issue_id": main_issue_id,
}