Compare commits
19 Commits
worktree-c
...
worktree-2
| Author | SHA1 | Date | |
|---|---|---|---|
| 2f50a8cc79 | |||
| 079a489f0e | |||
| 720057bc72 | |||
| d8bb2c7c0c | |||
| 5f5b13c6a4 | |||
| b394177207 | |||
| 3f48dc0e11 | |||
| c218864f36 | |||
| c22935c008 | |||
| 58a2f33187 | |||
| 363485d025 | |||
| eebb193fd3 | |||
| 6aac4a73f6 | |||
| 81ea43a70a | |||
| b8f6226fb1 | |||
| 20a1da0a0e | |||
| 90de9831c3 | |||
| f91cb2a660 | |||
| bf9559c61a |
@@ -15,8 +15,12 @@ tools:
|
||||
- mcp__legal-ai__document_get_text
|
||||
- mcp__legal-ai__extract_claims
|
||||
- mcp__legal-ai__extract_appraiser_facts
|
||||
- mcp__legal-ai__get_appraiser_facts
|
||||
- mcp__legal-ai__get_claims
|
||||
- mcp__legal-ai__aggregate_claims_to_arguments
|
||||
- mcp__legal-ai__get_legal_arguments
|
||||
- mcp__legal-ai__analyze_protocol
|
||||
- mcp__legal-ai__get_protocol_analysis
|
||||
- mcp__legal-ai__search_case_documents
|
||||
- mcp__legal-ai__search_decisions
|
||||
- mcp__legal-ai__search_precedent_library
|
||||
@@ -123,9 +127,9 @@ tools:
|
||||
4. חלץ טענות/תשובות/תגובות (`extract_claims` עם doc_type ו-party_hint מתאימים)
|
||||
- **מסמך גדול (>15,000 תווים):** מאז phase 1 של מערכת הניתוח, ה-chunking הסמנטי + מקבילות + retry מטופל אוטומטית. גם מסמך של 100K+ תווים ירוץ עד הסוף. אם בכל זאת נכשל — דווח ב-issue.
|
||||
- **טיפול בכשל:** אם `extract_claims` החזיר `partial=true` או 0 טענות ממסמך לא ריק — נסה שוב פעם אחת. אם עדיין נכשל — סטטוס issue = `blocked`, פרסם comment עם הפירוט.
|
||||
5. **חלץ עובדות שמאי** — לכל מסמך `doc_type='appraisal'` בתיק, הרץ `extract_appraiser_facts(case_number)` (פעם אחת לתיק, מטפל בכל השומות). **חובה בכל ערר השבחה (8xxx) ופיצויים (9xxx) — בלי זה ה-writer לא יוכל לכתוב את בלוק ז עם מספרים מדויקים.**
|
||||
5. **חלץ עובדות שמאי** — לכל מסמך `doc_type='appraisal'` בתיק, הרץ `extract_appraiser_facts(case_number)` (פעם אחת לתיק, מטפל בכל השומות). **חובה בכל ערר השבחה (8xxx) ופיצויים (9xxx) — בלי זה ה-writer לא יוכל לכתוב את בלוק ז עם מספרים מדויקים.** מיד אחריו הרץ `mcp__legal-ai__get_appraiser_facts(case_number)` כדי **לקרוא בחזרה** את מה שנשמר ולוודא שהחילוץ אכן נחת — אל תדווח על חילוץ שלא אימתת בקריאה.
|
||||
6. וודא שכל פריט מסווג ל-claim_type הנכון
|
||||
7. **קבץ טענות לטיעונים משפטיים** — לאחר שכל הטענות חולצו וסוּוגו, הרץ `aggregate_claims_to_arguments(case_number)` שמקבץ את הפרופוזיציות הגולמיות לטיעונים משפטיים מובחנים (~6-12 לכל צד). זהו קלט מובנה לבלוק ז (טענות הצדדים) ולבלוק י (דיון) — הכותב נשען עליו. אם 0 טענות חולצו — דלג. הפלט עובר שער-אישור (ראה `get_legal_arguments`).
|
||||
7. **קבץ טענות לטיעונים משפטיים** — לאחר שכל הטענות חולצו וסוּוגו, הרץ `aggregate_claims_to_arguments(case_number)` שמקבץ את הפרופוזיציות הגולמיות לטיעונים משפטיים מובחנים (~6-12 לכל צד). זהו קלט מובנה לבלוק ז (טענות הצדדים) ולבלוק י (דיון) — הכותב נשען עליו. אם 0 טענות חולצו — דלג. הפלט עובר שער-אישור — קרא אותו בחזרה עם `mcp__legal-ai__get_legal_arguments(case_number)` ואמת שמספר הטיעונים לכל צד סביר לפני שאתה ממשיך.
|
||||
|
||||
### שלב 2: ניתוח מעמיק
|
||||
הצג במבנה הבא:
|
||||
@@ -270,6 +274,29 @@ search_precedent_library(
|
||||
|
||||
**מינימום:** מספר queries ב-Q1+Q2+Q3 לקורפוס הסמכותי = מספר טענות סף + מספר סוגיות מרכזיות. אם זיהית 5 סוגיות + 2 טענות סף → לפחות 7 queries.
|
||||
|
||||
## משימה על-פי-דרישה: ניתוח פרוטוקול-הדיון
|
||||
|
||||
**זו אינה חלק מהזרימה הרגילה** — היא מגיעה כ-issue ייעודי כשחיים לוחץ "נתח את פרוטוקול הדיון" בפאנל **"מה קרה בדיון"** (טאב טיעונים-ועמדות). ה-issue יאמר במפורש להריץ `analyze_protocol`.
|
||||
|
||||
```
|
||||
mcp__legal-ai__analyze_protocol(case_number="<מספר-התיק>")
|
||||
mcp__legal-ai__analyze_protocol(case_number="<מספר-התיק>", document_id="<uuid>") # כשיש כמה פרוטוקולים
|
||||
```
|
||||
|
||||
הכלי משווה את **פרוטוקול דיון ועדת-הערר** מול **הטיעונים המאוגדים**, ומסווג כל שינוי ל-`strengthened` / `newly_raised` / `dropped`, עם `evidence_quote` מהפרוטוקול לכל שורה (INV-AH — אין שורה בלי ציטוט).
|
||||
|
||||
**שני תנאים מוקדמים — בדוק אותם לפני שאתה מריץ:**
|
||||
1. **פרוטוקול של ועדת-הערר.** הכלי בוחר אוטומטית פרוטוקול שה-`protocol_scope` שלו אינו `lower`; פרוטוקול של הוועדה **המקומית** מזין רקע (בלוק ו) בלבד ולא מוביל את ההשוואה. אם כל הפרוטוקולים בתיק הם `lower` — אין דיון-ערר להשוות אליו.
|
||||
2. **טיעונים מאוגדים.** הרץ `get_legal_arguments(case_number)`. אם ריק — הרץ קודם `aggregate_claims_to_arguments`.
|
||||
|
||||
אם תנאי חסר — **אל תעקוף**: כתוב comment בעברית שמפרט מה חסר, וסגור `blocked`.
|
||||
|
||||
`re-run` מחליף את הניתוח הקודם (idempotent). לקריאה בלבד, בלי ניתוח מחדש: `mcp__legal-ai__get_protocol_analysis(case_number, change_type="")`.
|
||||
|
||||
**דווח ב-comment בעברית**: כמה טענות התחזקו, כמה נטענו לראשונה, כמה ירדו — ומה החידוד המרכזי שעלה בדיון.
|
||||
|
||||
> ⚠️ **הרץ את הכלי בקדמה — לעולם לא ברקע.** שיגור ל-background וסיום התור מסיים את ה-run, וקבוצת-התהליכים נהרגת יחד איתו: הניתוח נקטע ולא נשמר דבר (נצפה ב-CMP-229, 2026-08-04). הכלי לוקח כמה דקות; זה תקין. חכה לו.
|
||||
|
||||
## שלב 6: בדיקת שלמות — לפני שמסיימים!
|
||||
|
||||
**לפני סיום, בצע את הבדיקות הבאות. אם בדיקה נכשלת — אל תסיים כ-"done".**
|
||||
|
||||
@@ -41,6 +41,7 @@ tools:
|
||||
- mcp__legal-ai__halacha_corroboration
|
||||
- mcp__legal-ai__corroboration_rebuild
|
||||
- mcp__legal-ai__extract_appraiser_facts
|
||||
- mcp__legal-ai__get_appraiser_facts
|
||||
- mcp__legal-ai__extract_plans
|
||||
- mcp__legal-ai__plan_get
|
||||
- mcp__legal-ai__plan_search
|
||||
@@ -704,6 +705,8 @@ ls data/cases/$CASE_NUMBER/documents/research/analysis-and-research.md
|
||||
```
|
||||
⚠️ אם מחזיר `status="sides_missing"` → דווח לחיים שאין תיוג `appraiser_side` במסמכי השומה (`document_update` עם `appraiser_side` בערכים `committee`/`appellant`/`deciding`). עצור עד שיתוקן.
|
||||
|
||||
אחרי החילוץ — הרץ `mcp__legal-ai__get_appraiser_facts(case_number="...")` כדי **לקרוא בחזרה** ולוודא שהעובדות אכן נשמרו. חילוץ שדיווח הצלחה אך לא נקרא בחזרה אינו ראיה שהנתונים שם.
|
||||
|
||||
אם הטבלה כבר מלאה — `write_interim_draft` ידלג על ההרצה אוטומטית, אז גם בלי הצעד הזה זה יעבוד.
|
||||
|
||||
3. **כתיבת 5 הבלוקים:**
|
||||
|
||||
@@ -17,6 +17,7 @@ tools:
|
||||
- mcp__legal-ai__search_precedent_library
|
||||
- mcp__legal-ai__search_internal_decisions
|
||||
- mcp__legal-ai__precedent_library_get
|
||||
- mcp__legal-ai__precedent_library_list
|
||||
- mcp__legal-ai__precedent_list
|
||||
- mcp__legal-ai__halacha_review
|
||||
---
|
||||
|
||||
25
.gitea/workflows/agent-tool-grants.yaml
Normal file
25
.gitea/workflows/agent-tool-grants.yaml
Normal file
@@ -0,0 +1,25 @@
|
||||
name: INV-AG3 Agent Tool Grants
|
||||
|
||||
# Hard gate for INV-AG3 (docs/spec/X4-agents.md §2א): a subagent's `tools:`
|
||||
# frontmatter is a CLOSED allow-list, so any MCP tool an agent is TOLD to run —
|
||||
# by the backend delegation in web/, or by its own instructions — must appear
|
||||
# there. Built after analyze_protocol shipped without a grant (2026-06-30) and
|
||||
# the analyst was handed an issue instructing it to run a tool it could not
|
||||
# call (CMP-229, 2026-08-04). Pure-stdlib check (no venv) — fast, runs on every
|
||||
# PR and on push to main.
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
branches: [main]
|
||||
push:
|
||||
branches: [main]
|
||||
|
||||
jobs:
|
||||
agent-tool-grants:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: INV-AG3 — agent tool-grant guard
|
||||
run: python3 scripts/agent_tool_grants_guard.py
|
||||
@@ -17,7 +17,7 @@
|
||||
| ezer-mishpati-web | ממשק העלאת מסמכים (Docker/Coolify) | `legal-ai.nautilus.marcusgroup.org` |
|
||||
| Paperclip | סוכן AI — מריץ Claude Code agents (pm2, מקומי) | `localhost:3100` |
|
||||
| legal-chat-service | גשר claude CLI לטאב הצ'אט ב-/training (pm2, loopback) | `127.0.0.1:8770` |
|
||||
| Infisical | ניהול סודות | `secret.dev.marcus-law.co.il` |
|
||||
| Infisical | ניהול סודות — פרויקט **All Infrastructure** (`2c462576-b125-4279-b0ec-7220dbf51ccf`), env `main`. סודות המערכת מפוזרים על שלוש תיקיות-אחיות: `/apps/legal-ai` · `/apps/paperclip` · `/apps/paperclip-hermes` | `secret.marcus-law.co.il` |
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -146,9 +146,29 @@ another company`, [X2 §2](X2-multi-company.md)).
|
||||
**כלל:** ה-frontmatter `tools:` של כל סוכן מעניק **בדיוק** את הכלים שהוראותיו דורשות — כל כלי שההוראות
|
||||
מצריכות מוענק, וכלי שמוענק-ולא-בשימוש נבחן. מופע של [G10](00-constitution.md#inv-g10-המערכת-מסייעת--שערים-אנושיים-הם-invariant)
|
||||
(שערים מוגדרים) ו-[G2](00-constitution.md#inv-g2-מקור-אמת-יחיד--אין-מסלולים-מקבילים-מתפצלים); מקביל ל-[X9 INV-TOOL6](X9-mcp-tool-contract.md).
|
||||
**מקור-סמכות:** frontmatter `tools:` מול ה-instructions בקבצי-[.claude/agents/](../../.claude/agents/). (פרויקטלי-תפעולי.)
|
||||
**אכיפה:** בדיקת-עקביות tools↔instructions (FU-13 ✅ 2026-06-06). אכיפה אוטומטית עתידית — בתת-פרויקט 5 (spec-guardian).
|
||||
**הפרה ידועה:** — (טופל ב-FU-13: legal-analyst קיבל `aggregate_claims_to_arguments`; researcher כבר היה תקין; `extract_references`/`extract_internal_citations` הם מטלת-researcher, לא analyst — ראה §2א).
|
||||
|
||||
> **ה-frontmatter הוא allow-list סגורה.** כלי הרשום בשרת-ה-MCP אך חסר מהרשימה **אינו ניתן לקריאה**
|
||||
> ע"י הסוכן — גם כשהשרת מחובר לחלוטין. הסוכן חווה זאת כ"הכלים לא נחשפים לסשן", ובלי הבנת המנגנון
|
||||
> הוא נוטה **לעקוף** (SQL ישיר, סקריפט מקומי) במקום לדווח — וזה מסלול מקביל, כלומר הפרת G2.
|
||||
> **"מורים להריץ" כולל את ה-backend:** טקסט של issue שנוצר ב-`web/` ומכיל `mcp__legal-ai__X` הוא
|
||||
> הוראה לכל דבר, ולכן מחייב הענקה.
|
||||
|
||||
**מקור-סמכות:** frontmatter `tools:` מול ה-instructions בקבצי-[.claude/agents/](../../.claude/agents/)
|
||||
**ומול הוראות-ה-backend** ב-`web/`. (פרויקטלי-תפעולי.)
|
||||
**אכיפה:** ✅ **אוטומטית מ-2026-08-04** — [`scripts/agent_tool_grants_guard.py`](../../scripts/agent_tool_grants_guard.py),
|
||||
שער-CI קשיח ([`.gitea/workflows/agent-tool-grants.yaml`](../../.gitea/workflows/agent-tool-grants.yaml)),
|
||||
בדפוס [`leak_guard.py`](../../scripts/leak_guard.py) של G12. ארבעה כללים: (1) כל `mcp__legal-ai__X`
|
||||
ב-`web/` מוענק לסוכן כלשהו · (2) כל `mcp__legal-ai__X` בגוף קובץ-סוכן מוענק **באותו** קובץ ·
|
||||
(3) אין הענקה לכלי שאינו רשום בשרת · (4) שם-כלי בגרשיים-הפוכים ללא תחילית — מוענק, או מסווג
|
||||
מפורשות ב-`CONTRASTIVE_OK` (אזכור ניגודי / מטלת-סוכן-אחר / שם-עמודה מתנגש). לא-סוכנים ולכן
|
||||
מוחרגים: `hermes-curator.md` ו-`legal-analyst-gemini-critique.md` (בלי frontmatter בכוונה —
|
||||
האדפטר שולח פרומפט גולמי) ו-`HEARTBEAT.md` (checklist משותף).
|
||||
**הפרה ידועה:** — (היסטוריה: FU-13 ✅ 2026-06-06 — `aggregate_claims_to_arguments` ל-analyst;
|
||||
`extract_references`/`extract_internal_citations` הם מטלת-researcher, ראה §2א. **הישנות 2026-08-04**
|
||||
— האכיפה הידנית לא החזיקה: `analyze_protocol` (24e3e2f, 2026-06-30) נרשם בשרת בלי הענקה, ו-#226
|
||||
הוסיף delegation שמורה למנתח להריץ אותו → CMP-229 נשרף בשתי הרצות ועקף ל-SQL ידני. נסגר יחד עם
|
||||
`get_protocol_analysis`/`get_legal_arguments`/`get_appraiser_facts` ו-`precedent_library_list` ל-QA,
|
||||
והאכיפה הועברה ל-CI כדי שלא תישען שוב על משמעת ידנית.)
|
||||
|
||||
### INV-AG4: שער שטן-מליץ — red-team לידים לא-סמכותיים תחת אישור-יו"ר
|
||||
**כלל:** אחרי שלב-הניתוח (`analysis-and-research.md` תקין) וב**לפני** הפעלת הכותב, ה-CEO מפעיל
|
||||
|
||||
@@ -6,6 +6,7 @@ Run with: python -m legal_mcp.server
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import os
|
||||
import sys
|
||||
from collections.abc import AsyncIterator
|
||||
from contextlib import asynccontextmanager
|
||||
@@ -41,11 +42,56 @@ async def lifespan(server: FastMCP) -> AsyncIterator[None]:
|
||||
logger.info("Ezer Mishpati MCP server stopped")
|
||||
|
||||
|
||||
# HTTP listener address, used only when MCP_TRANSPORT selects an HTTP transport.
|
||||
#
|
||||
# These MUST be passed to the constructor rather than left to FastMCP's own
|
||||
# FASTMCP_HOST / FASTMCP_PORT environment settings: FastMCP.__init__ declares
|
||||
# `host: str = "127.0.0.1"` and `port: int = 8000` as explicit keyword defaults
|
||||
# and forwards them into Settings(**settings). In pydantic-settings, explicit
|
||||
# init kwargs outrank environment variables — so FASTMCP_PORT is silently
|
||||
# ignored and the server binds 8000 regardless (verified 2026-08-05; on this
|
||||
# host 8000 is already taken, so it failed loudly by luck rather than design).
|
||||
#
|
||||
# Default to loopback. The bearer gate below is the real protection; the narrow
|
||||
# bind is defence in depth, not a substitute for it.
|
||||
MCP_HTTP_HOST = os.environ.get("MCP_HTTP_HOST", "127.0.0.1")
|
||||
MCP_HTTP_PORT = int(os.environ.get("MCP_HTTP_PORT", "8790"))
|
||||
|
||||
# Bearer gate — wired only when an HTTP transport is actually selected (#231.2).
|
||||
#
|
||||
# stdio must never require a token: the pipe is the boundary there, and every
|
||||
# interactive session reaches us that way. Demanding a token on stdio would
|
||||
# break all of them for no security gain.
|
||||
#
|
||||
# Building the verifier at import time (rather than inside main()) is deliberate:
|
||||
# FastMCP takes `token_verifier` and `auth` as constructor arguments, so a
|
||||
# missing token has to fail here — before the listener exists — not after it is
|
||||
# already accepting connections.
|
||||
_http_transport = os.environ.get("MCP_TRANSPORT", "stdio").strip() in ("sse", "streamable-http")
|
||||
_auth_kwargs: dict = {}
|
||||
if _http_transport:
|
||||
from mcp.server.auth.settings import AuthSettings
|
||||
|
||||
from legal_mcp.services.http_auth import StaticTokenVerifier, load_token
|
||||
|
||||
_base_url = f"http://{MCP_HTTP_HOST}:{MCP_HTTP_PORT}"
|
||||
_auth_kwargs = {
|
||||
"token_verifier": StaticTokenVerifier(load_token()),
|
||||
# AuthSettings is what switches on the SDK's BearerAuthBackend. We are a
|
||||
# resource server with a pre-shared token, not an OAuth client, so both
|
||||
# URLs simply point at ourselves — they exist to satisfy the protected-
|
||||
# resource metadata contract, and nothing issues tokens from them.
|
||||
"auth": AuthSettings(issuer_url=_base_url, resource_server_url=_base_url),
|
||||
}
|
||||
|
||||
# Create MCP server
|
||||
mcp = FastMCP(
|
||||
"Ezer Mishpati - עוזר משפטי",
|
||||
instructions="מערכת AI לסיוע בניסוח החלטות משפטיות בסגנון דפנה תמיר",
|
||||
lifespan=lifespan,
|
||||
host=MCP_HTTP_HOST,
|
||||
port=MCP_HTTP_PORT,
|
||||
**_auth_kwargs,
|
||||
)
|
||||
|
||||
# ── Import and register tools ───────────────────────────────────────
|
||||
@@ -1260,7 +1306,44 @@ async def corroboration_rebuild(case_law_id: str = "") -> dict:
|
||||
|
||||
|
||||
def main():
|
||||
mcp.run(transport="stdio")
|
||||
"""Run the server on the transport named by ``MCP_TRANSPORT`` (default stdio).
|
||||
|
||||
ONE server, two transports — deliberately not a second implementation (G2).
|
||||
The tool registry, services and DB pool above are shared verbatim; only the
|
||||
wire protocol differs, so a tool can never exist on one transport and not
|
||||
the other.
|
||||
|
||||
- ``stdio`` (default) — the historical path. Every interactive Claude Code
|
||||
session reaches us this way via the ``legal-ai`` entry in ``~/.claude.json``.
|
||||
Changing this default would break them all, so it stays the default.
|
||||
- ``streamable-http`` — required by the agent-platform port. Agents driven
|
||||
over the Agent Client Protocol receive their MCP servers from the *client*
|
||||
at session start, and that channel accepts HTTP servers only: each is
|
||||
injected as ``{type:"http", name, url, headers:[Bearer]}``. A stdio server
|
||||
has no path into such a session at all, which is why platform-driven agents
|
||||
lost all 108 tools when their execution engine changed — TaskMaster #231.
|
||||
Platform-side wiring lives behind the port (``web/agent_platform_port.py``),
|
||||
not here; this module only has to be reachable over HTTP.
|
||||
|
||||
Host/port come from FastMCP's own ``FASTMCP_HOST`` / ``FASTMCP_PORT`` settings.
|
||||
Bind to loopback only: this transport carries no authentication yet, and the
|
||||
registry includes destructive tools (``case_delete``, ``precedent_library_delete``).
|
||||
The Bearer gate is #231.2 and MUST land before this is reachable off-host.
|
||||
"""
|
||||
transport = os.environ.get("MCP_TRANSPORT", "stdio").strip() or "stdio"
|
||||
valid = ("stdio", "sse", "streamable-http")
|
||||
if transport not in valid:
|
||||
# Fail loudly — a typo must not silently fall back to stdio and leave
|
||||
# the HTTP listener absent while everything "looks" fine (§6).
|
||||
raise SystemExit(
|
||||
f"MCP_TRANSPORT={transport!r} is not one of {valid}",
|
||||
)
|
||||
if transport != "stdio":
|
||||
logger.info(
|
||||
"Serving MCP over %s on %s:%s",
|
||||
transport, mcp.settings.host, mcp.settings.port,
|
||||
)
|
||||
mcp.run(transport=transport)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -147,6 +147,38 @@ def _normalize_argument(raw: dict, fallback_topic: str = "") -> dict | None:
|
||||
}
|
||||
|
||||
|
||||
class AggregationFailed(RuntimeError):
|
||||
"""One side's aggregation call failed. Never swallowed — see #233."""
|
||||
|
||||
|
||||
#: Largest number of propositions sent to Claude in a single aggregation call.
|
||||
#:
|
||||
#: Above roughly this size the model stops returning a JSON array and the whole
|
||||
#: side is lost. The per-brief split added for the respondent side (#224) hid
|
||||
#: this by accident — it kept those calls small — but appellant and committee
|
||||
#: speak with a single voice and are never split by brief, so a big appeal goes
|
||||
#: out in one oversized call. 1069-04-26 sent 310 propositions and got nothing
|
||||
#: back (#233).
|
||||
#:
|
||||
#: Chunking is a real trade-off, not a free win: each chunk is grouped in
|
||||
#: isolation, so a side split across chunks can end up with more (and slightly
|
||||
#: overlapping) arguments than a single pass would have produced. Losing an
|
||||
#: entire litigant is worse, and the alternative — a smaller prompt per
|
||||
#: proposition — would degrade every case to fix the large ones.
|
||||
MAX_PROPS_PER_CALL = 80
|
||||
|
||||
|
||||
def _chunk(propositions: list[dict], size: int) -> list[list[dict]]:
|
||||
"""Split propositions into calls of at most ``size``, keeping claim order.
|
||||
|
||||
Order matters: claims arrive sorted by ``claim_index``, so neighbouring
|
||||
propositions usually belong to the same head of argument. Slicing in order
|
||||
keeps related material together instead of scattering one argument across
|
||||
chunks.
|
||||
"""
|
||||
return [propositions[i:i + size] for i in range(0, len(propositions), size)]
|
||||
|
||||
|
||||
async def _aggregate_party(
|
||||
party: str, propositions: list[dict], party_name: str = "",
|
||||
) -> list[dict]:
|
||||
@@ -155,9 +187,27 @@ async def _aggregate_party(
|
||||
``party_name`` names the specific pleading when this is a split side
|
||||
(respondent / permit_applicant brief), so the prompt scopes to that
|
||||
litigant's position only (#224).
|
||||
|
||||
Sides larger than ``MAX_PROPS_PER_CALL`` are aggregated in several calls and
|
||||
concatenated; a failure in any chunk raises rather than returning a partial
|
||||
side quietly.
|
||||
"""
|
||||
if not propositions:
|
||||
return []
|
||||
|
||||
if len(propositions) > MAX_PROPS_PER_CALL:
|
||||
chunks = _chunk(propositions, MAX_PROPS_PER_CALL)
|
||||
logger.info(
|
||||
"argument_aggregator: party '%s'%s has %d propositions — "
|
||||
"aggregating in %d calls of up to %d",
|
||||
party, f" ({party_name})" if party_name else "",
|
||||
len(propositions), len(chunks), MAX_PROPS_PER_CALL,
|
||||
)
|
||||
out: list[dict] = []
|
||||
for chunk in chunks:
|
||||
out.extend(await _aggregate_party(party, chunk, party_name=party_name))
|
||||
return out
|
||||
|
||||
prompt = _build_prompt(party, propositions, party_name=party_name)
|
||||
|
||||
try:
|
||||
@@ -171,11 +221,18 @@ async def _aggregate_party(
|
||||
) from e
|
||||
|
||||
if not isinstance(raw_result, list):
|
||||
logger.warning(
|
||||
"argument_aggregator: Claude returned non-list (%s) for party '%s'",
|
||||
type(raw_result).__name__, party,
|
||||
# NOT a silent []: returning empty here is indistinguishable from "this
|
||||
# side genuinely has no arguments", and the caller would then report
|
||||
# status=completed while a whole litigant vanished. That is exactly what
|
||||
# happened to the appellant side of 1069-04-26 — 310 claims in, 0
|
||||
# arguments out, "completed" (#233). Raise so the caller records it in
|
||||
# ``errors`` and the status degrades to completed_with_errors (§6).
|
||||
raise AggregationFailed(
|
||||
f"Claude returned {type(raw_result).__name__}, not a list, for party "
|
||||
f"'{party}'{f' ({party_name})' if party_name else ''} with "
|
||||
f"{len(propositions)} propositions — the side would otherwise be "
|
||||
f"dropped without trace",
|
||||
)
|
||||
return []
|
||||
|
||||
out: list[dict] = []
|
||||
for entry in raw_result:
|
||||
@@ -276,6 +333,12 @@ async def aggregate_claims_to_arguments(
|
||||
group_key = f"{party}·{party_name}" if party_name else party
|
||||
try:
|
||||
arguments = await _aggregate_party(party, props, party_name=party_name)
|
||||
except AggregationFailed as e:
|
||||
# A side that failed is NOT a side with no arguments. Record it so
|
||||
# the status degrades and the caller can see which litigant is
|
||||
# missing (#233).
|
||||
errors.append(f"{group_key}: {e}")
|
||||
continue
|
||||
except RuntimeError as e:
|
||||
# Most likely cause: Claude CLI not installed (running from
|
||||
# the container). Don't crash — record the gap and continue.
|
||||
|
||||
@@ -5324,21 +5324,19 @@ async def list_external_case_law(
|
||||
search: str = "",
|
||||
limit: int = 100,
|
||||
offset: int = 0,
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
) -> list[dict]:
|
||||
"""List chair-uploaded precedents, with simple filters.
|
||||
|
||||
source_kind="" (default) = the whole corpus — court rulings *and*
|
||||
appeals-committee decisions. The old ``external_upload`` default hid
|
||||
every committee decision from a plain listing (#232).
|
||||
source_kind="all_committees" expands to: source_kind='internal_committee'
|
||||
OR (source_kind='external_upload' AND source_type='appeals_committee').
|
||||
"""
|
||||
pool = await get_pool()
|
||||
if source_kind == "all_committees":
|
||||
conditions = [
|
||||
"(source_kind = 'internal_committee' OR "
|
||||
"(source_kind = 'external_upload' AND source_type = 'appeals_committee'))"
|
||||
]
|
||||
else:
|
||||
conditions = [f"source_kind = '{source_kind}'"]
|
||||
sk_clause = _source_kind_clause(source_kind)
|
||||
conditions = [sk_clause] if sk_clause else []
|
||||
params: list = []
|
||||
idx = 1
|
||||
if practice_area:
|
||||
@@ -5358,13 +5356,19 @@ async def list_external_case_law(
|
||||
params.append(source_type)
|
||||
idx += 1
|
||||
if search:
|
||||
# Case-number separator normalisation (#232, trap 2): committee numbers
|
||||
# are stored hyphenated ("83-16") while people write them with a slash
|
||||
# ("83/16"). Fold both sides to '/' so either form finds the row —
|
||||
# matching at the point of comparison rather than asking every caller
|
||||
# to guess the stored form.
|
||||
conditions.append(
|
||||
f"(case_number ILIKE ${idx} OR case_name ILIKE ${idx} "
|
||||
f"OR summary ILIKE ${idx} OR headnote ILIKE ${idx})"
|
||||
f"OR summary ILIKE ${idx} OR headnote ILIKE ${idx} "
|
||||
f"OR replace(case_number, '-', '/') ILIKE replace(${idx}, '-', '/'))"
|
||||
)
|
||||
params.append(f"%{search}%")
|
||||
idx += 1
|
||||
where_sql = " AND ".join(conditions)
|
||||
where_sql = " AND ".join(conditions) if conditions else "TRUE"
|
||||
params.extend([limit, offset])
|
||||
sql = f"""
|
||||
SELECT id, case_number, case_name, court, date, practice_area,
|
||||
@@ -7775,6 +7779,45 @@ async def list_corroboration_for_halacha(halacha_id: UUID) -> list[dict]:
|
||||
]
|
||||
|
||||
|
||||
#: Accepted ``source_kind`` selectors. ``""``/``"all"`` mean *no filter* —
|
||||
#: the whole corpus, court rulings and appeals-committee decisions alike.
|
||||
_SOURCE_KIND_SELECTORS = frozenset(
|
||||
{"", "all", "all_committees", "external_upload", "internal_committee", "cited_only"}
|
||||
)
|
||||
|
||||
|
||||
def _source_kind_clause(source_kind: str, column_prefix: str = "") -> str:
|
||||
"""Return the SQL predicate for a ``source_kind`` selector, or "" for none.
|
||||
|
||||
Single definition for every caller (G2) — the selector vocabulary was
|
||||
previously restated at each search/list site, which is how the
|
||||
``external_upload`` default silently hid 104 appeals-committee
|
||||
decisions from ``search_precedent_library`` (#232).
|
||||
|
||||
``column_prefix`` is the table alias plus dot (e.g. ``"cl."``) or "" when
|
||||
the query selects from ``case_law`` directly.
|
||||
|
||||
Raises ValueError on an unknown selector rather than interpolating it —
|
||||
the value reaches SQL by f-string, so the whitelist is also what keeps
|
||||
that safe.
|
||||
"""
|
||||
sk = (source_kind or "").strip()
|
||||
if sk not in _SOURCE_KIND_SELECTORS:
|
||||
raise ValueError(
|
||||
f"source_kind לא מוכר: {source_kind!r}. "
|
||||
f"ערכים חוקיים: {sorted(_SOURCE_KIND_SELECTORS - {''})} או '' (הכל)"
|
||||
)
|
||||
if sk in ("", "all"):
|
||||
return ""
|
||||
p = column_prefix
|
||||
if sk == "all_committees":
|
||||
return (
|
||||
f"({p}source_kind = 'internal_committee' OR "
|
||||
f"({p}source_kind = 'external_upload' AND {p}source_type = 'appeals_committee'))"
|
||||
)
|
||||
return f"{p}source_kind = '{sk}'"
|
||||
|
||||
|
||||
async def search_precedent_library_semantic(
|
||||
query_embedding: list[float],
|
||||
practice_area: str = "",
|
||||
@@ -7785,14 +7828,15 @@ async def search_precedent_library_semantic(
|
||||
subject_tag: str = "",
|
||||
limit: int = 10,
|
||||
include_halachot: bool = True,
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
district: str = "",
|
||||
chair_name: str = "",
|
||||
) -> list[dict]:
|
||||
"""Semantic search over precedents filtered by source_kind.
|
||||
"""Semantic search over precedents, optionally filtered by source_kind.
|
||||
|
||||
source_kind='external_upload' → court rulings (default)
|
||||
source_kind='internal_committee' → appeals-committee decisions
|
||||
source_kind='' → the whole corpus (default, #232)
|
||||
source_kind='external_upload' → court rulings only
|
||||
source_kind='internal_committee' → appeals-committee decisions only
|
||||
|
||||
Returns merged halachot + chunks. Halachot are pre-distilled rules, so
|
||||
they get a small score boost. Only ``approved`` / ``published`` halachot
|
||||
@@ -7800,12 +7844,15 @@ async def search_precedent_library_semantic(
|
||||
of halacha review status.
|
||||
"""
|
||||
pool = await get_pool()
|
||||
sk_clause = _source_kind_clause(source_kind, "cl.")
|
||||
halacha_filters = [
|
||||
"h.review_status <> 'rejected'", # #153: include background; rank verified higher
|
||||
f"cl.source_kind = '{source_kind}'",
|
||||
"cl.searchable = true",
|
||||
]
|
||||
chunk_filters = [f"cl.source_kind = '{source_kind}'", "cl.searchable = true"]
|
||||
chunk_filters = ["cl.searchable = true"]
|
||||
if sk_clause:
|
||||
halacha_filters.append(sk_clause)
|
||||
chunk_filters.append(sk_clause)
|
||||
h_params: list = [query_embedding, limit]
|
||||
c_params: list = [query_embedding, limit]
|
||||
h_idx = 3
|
||||
@@ -8017,7 +8064,7 @@ async def search_precedent_library_lexical(
|
||||
appeal_subtype: str = "",
|
||||
is_binding: bool | None = None,
|
||||
subject_tag: str = "",
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
district: str = "",
|
||||
chair_name: str = "",
|
||||
limit: int = 30,
|
||||
@@ -8043,12 +8090,15 @@ async def search_precedent_library_lexical(
|
||||
return []
|
||||
|
||||
pool = await get_pool()
|
||||
sk_clause = _source_kind_clause(source_kind, "cl.")
|
||||
halacha_filters = [
|
||||
"h.review_status <> 'rejected'", # #153: include background; rank verified higher
|
||||
f"cl.source_kind = '{source_kind}'",
|
||||
"cl.searchable = true",
|
||||
]
|
||||
chunk_filters = [f"cl.source_kind = '{source_kind}'", "cl.searchable = true"]
|
||||
chunk_filters = ["cl.searchable = true"]
|
||||
if sk_clause:
|
||||
halacha_filters.append(sk_clause)
|
||||
chunk_filters.append(sk_clause)
|
||||
# $1 = query, $2 = limit. Filters append starting at $3.
|
||||
h_params: list = [query, limit]
|
||||
c_params: list = [query, limit]
|
||||
|
||||
106
mcp-server/src/legal_mcp/services/http_auth.py
Normal file
106
mcp-server/src/legal_mcp/services/http_auth.py
Normal file
@@ -0,0 +1,106 @@
|
||||
"""Bearer-token gate for the HTTP transport of the MCP server (#231.2).
|
||||
|
||||
Why this exists
|
||||
---------------
|
||||
Over ``stdio`` the protection is the pipe itself: only a process that already
|
||||
runs as this user can speak to the server. ``streamable-http`` removes that
|
||||
property entirely — anything that can reach the socket can call any of the 108
|
||||
registered tools, and the registry includes ``case_delete``,
|
||||
``precedent_library_delete``, ``document_upload`` and every block-writing tool.
|
||||
An unauthenticated listener is therefore a delete-any-case endpoint.
|
||||
|
||||
The agent platform already speaks this dialect: it injects each runtime MCP
|
||||
server as ``{type:"http", name, url, headers:[{name:"Authorization",
|
||||
value:"Bearer <token>"}]}``. So a static bearer token is exactly the shape the
|
||||
caller will present — no negotiation, no OAuth dance.
|
||||
|
||||
Design notes
|
||||
------------
|
||||
- We implement the SDK's own ``TokenVerifier`` protocol and let
|
||||
``BearerAuthBackend`` do the enforcement, rather than adding bespoke
|
||||
middleware. One auth path, the framework's (G2).
|
||||
- The token is read from the environment, which the service unit populates from
|
||||
Infisical. It is never defaulted, never logged, and never embedded here.
|
||||
- Comparison is constant-time: a naive ``==`` leaks the token byte-by-byte to a
|
||||
caller who can time responses.
|
||||
- ``verify_token`` returns ``None`` (not an exception) on mismatch — that is the
|
||||
protocol's "reject" signal and yields a clean 401 instead of a 500 that would
|
||||
read as a server fault.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hmac
|
||||
import logging
|
||||
import os
|
||||
|
||||
from mcp.server.auth.provider import AccessToken, TokenVerifier
|
||||
|
||||
logger = logging.getLogger("legal_mcp.http_auth")
|
||||
|
||||
#: Environment variable carrying the shared bearer token.
|
||||
#:
|
||||
#: Name matches the Infisical key exactly — All Infrastructure / env `main` /
|
||||
#: `/apps/legal-ai` / ``MCP_HTTP_SHARED_SECRET``, tagged ``credentials``. Keeping
|
||||
#: the two identical means nobody has to hold a mapping in their head, and it
|
||||
#: follows the two bridge tokens already in that folder
|
||||
#: (``COURT_FETCH_SHARED_SECRET``, ``LEGAL_CHAT_SHARED_SECRET``).
|
||||
TOKEN_ENV = "MCP_HTTP_SHARED_SECRET"
|
||||
|
||||
#: Minimum acceptable token length. Short tokens are brute-forceable; refusing
|
||||
#: them at startup is cheaper than discovering it from an access log.
|
||||
MIN_TOKEN_LEN = 32
|
||||
|
||||
#: Reported as the authenticated principal. Single shared token today, so this
|
||||
#: is a constant rather than a real identity — kept explicit so that audit rows
|
||||
#: never imply per-agent attribution we cannot actually make.
|
||||
CLIENT_ID = "mcp-http-shared"
|
||||
|
||||
|
||||
class MissingTokenError(RuntimeError):
|
||||
"""Raised when the HTTP transport is requested without a usable token.
|
||||
|
||||
Deliberately fatal. The tempting alternative — start anyway and log a
|
||||
warning — produces a listener that looks healthy and answers every
|
||||
destructive tool call. Refusing to boot is the safe failure (§6: never
|
||||
swallow, never degrade silently).
|
||||
"""
|
||||
|
||||
|
||||
def load_token() -> str:
|
||||
"""Return the configured bearer token, or raise if it is unusable."""
|
||||
token = (os.environ.get(TOKEN_ENV) or "").strip()
|
||||
if not token:
|
||||
raise MissingTokenError(
|
||||
f"{TOKEN_ENV} is not set. The HTTP transport exposes destructive "
|
||||
f"tools and will not start without a bearer token. Set it from "
|
||||
f"Infisical, or use MCP_TRANSPORT=stdio.",
|
||||
)
|
||||
if len(token) < MIN_TOKEN_LEN:
|
||||
raise MissingTokenError(
|
||||
f"{TOKEN_ENV} is shorter than {MIN_TOKEN_LEN} characters — refusing "
|
||||
f"to start. Generate a long random token.",
|
||||
)
|
||||
return token
|
||||
|
||||
|
||||
class StaticTokenVerifier(TokenVerifier):
|
||||
"""Verifies the single shared bearer token presented by the platform.
|
||||
|
||||
Not an identity system: it answers "may this caller in at all", not "who is
|
||||
it". Per-agent attribution would need per-agent tokens, which is a later
|
||||
step once profiles exist (#231.4).
|
||||
"""
|
||||
|
||||
def __init__(self, expected: str) -> None:
|
||||
self._expected = expected
|
||||
|
||||
async def verify_token(self, token: str) -> AccessToken | None:
|
||||
# compare_digest over bytes; it tolerates unequal lengths without
|
||||
# short-circuiting, which is the whole point.
|
||||
if not hmac.compare_digest(token.encode("utf-8"),
|
||||
self._expected.encode("utf-8")):
|
||||
# No token material in the log line — only the fact of a rejection.
|
||||
logger.warning("Rejected MCP HTTP request: bearer token mismatch")
|
||||
return None
|
||||
return AccessToken(token=token, client_id=CLIENT_ID, scopes=[])
|
||||
@@ -98,15 +98,16 @@ async def search_precedent_library_hybrid(
|
||||
is_binding: bool | None = None,
|
||||
subject_tag: str = "",
|
||||
include_halachot: bool = True,
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
district: str = "",
|
||||
chair_name: str = "",
|
||||
max_per_case_law: int = 2,
|
||||
) -> list[dict]:
|
||||
"""Hybrid wrapper for precedent-library search.
|
||||
|
||||
source_kind='external_upload' → court rulings (default)
|
||||
source_kind='internal_committee' → appeals-committee decisions
|
||||
source_kind='' → the whole corpus (default, #232)
|
||||
source_kind='external_upload' → court rulings only
|
||||
source_kind='internal_committee' → appeals-committee decisions only
|
||||
max_per_case_law: MMR-style diversity cap — at most N hits per
|
||||
case_law_id in the final ranked list (default 2). Prevents a
|
||||
single precedent from monopolizing the result list when many of
|
||||
|
||||
@@ -477,7 +477,7 @@ async def list_precedents(
|
||||
precedent_level: str = "",
|
||||
source_type: str = "",
|
||||
search: str = "",
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
limit: int = 100,
|
||||
offset: int = 0,
|
||||
) -> list[dict]:
|
||||
@@ -503,12 +503,18 @@ async def search_library(
|
||||
subject_tag: str = "",
|
||||
limit: int = 10,
|
||||
include_halachot: bool = True,
|
||||
source_kind: str = "",
|
||||
) -> list[dict]:
|
||||
"""Semantic search merging halachot (rule-level) and chunks (passage-level).
|
||||
|
||||
Only ``approved`` / ``published`` halachot are returned, per chair-review
|
||||
policy. Chunks are returned regardless of halacha review status.
|
||||
|
||||
``source_kind=""`` (default) covers the whole corpus — court rulings and
|
||||
appeals-committee decisions together. It used to be hard-wired to
|
||||
``external_upload`` here, which made 104 committee decisions
|
||||
unreachable through this entry point (#232).
|
||||
|
||||
When ``VOYAGE_RERANK_ENABLED`` is set, results are passed through
|
||||
voyage rerank-2 (cross-encoder). The +0.05 halacha boost from
|
||||
``search_precedent_library_semantic`` is preserved before rerank
|
||||
@@ -529,4 +535,5 @@ async def search_library(
|
||||
is_binding=is_binding,
|
||||
subject_tag=subject_tag,
|
||||
include_halachot=include_halachot,
|
||||
source_kind=source_kind,
|
||||
)
|
||||
|
||||
@@ -96,10 +96,13 @@ async def precedent_library_list(
|
||||
precedent_level: str = "",
|
||||
source_type: str = "",
|
||||
search: str = "",
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
limit: int = 100,
|
||||
) -> str:
|
||||
"""רשימה של פסיקה בקורפוס הסמכותי, עם פילטרים."""
|
||||
"""רשימה של פסיקה בקורפוס הסמכותי, עם פילטרים.
|
||||
|
||||
source_kind ריק (ברירת מחדל) = כל הקורפוס, כולל החלטות ועדות ערר.
|
||||
"""
|
||||
rows = await precedent_library.list_precedents(
|
||||
practice_area=practice_area,
|
||||
court=court,
|
||||
@@ -266,8 +269,9 @@ async def search_precedent_library(
|
||||
subject_tag: str = "",
|
||||
limit: int = 10,
|
||||
include_halachot: bool = True,
|
||||
source_kind: str = "",
|
||||
) -> str:
|
||||
"""חיפוש סמנטי בקורפוס הפסיקה הסמכותית.
|
||||
"""חיפוש סמנטי בקורפוס הפסיקה הסמכותית — פסקי דין **והחלטות ועדות ערר**.
|
||||
|
||||
מחזיר תוצאות מעורבות: הלכות (rule-level, מאושרות בלבד) + קטעי טקסט
|
||||
(passage-level). הלכות מקבלות boost קל בדירוג כי הן מזוקקות מראש.
|
||||
@@ -282,6 +286,9 @@ async def search_precedent_library(
|
||||
subject_tag: סינון לפי תגית נושא (לדוגמה "מועד_קביעת_שומה").
|
||||
limit: מספר תוצאות מקסימלי.
|
||||
include_halachot: האם לכלול הלכות (ברירת מחדל: כן).
|
||||
source_kind: ריק (ברירת מחדל) = כל הקורפוס — פסקי דין והחלטות ועדות
|
||||
ערר יחד. "external_upload" = פסקי בתי משפט בלבד;
|
||||
"internal_committee" = החלטות ועדות ערר בלבד.
|
||||
|
||||
Returns: רשימה מדורגת. כל פריט הוא {"type": "halacha"|"passage", "score", ...}.
|
||||
"""
|
||||
@@ -299,6 +306,7 @@ async def search_precedent_library(
|
||||
subject_tag=subject_tag,
|
||||
limit=limit,
|
||||
include_halachot=include_halachot,
|
||||
source_kind=source_kind,
|
||||
)
|
||||
# X11 Phase 2 (#154): attach the incoming-citation authority breakdown so the
|
||||
# research agent can WEIGH and ARGUE authority ("הלכה שאומצה ב-N החלטות ועדת-ערר")
|
||||
|
||||
54
mcp-server/tests/test_argument_aggregator.py
Normal file
54
mcp-server/tests/test_argument_aggregator.py
Normal file
@@ -0,0 +1,54 @@
|
||||
"""Regression tests for argument aggregation (#233).
|
||||
|
||||
Both tests cover the same 2026-08-05 incident from different angles: the
|
||||
appellant side of 1069-04-26 sent 310 propositions in one Claude call, the call
|
||||
came back as something other than a JSON array, and the code logged a warning
|
||||
and returned ``[]``. The caller could not tell that apart from "this side has no
|
||||
arguments", so ``aggregate_claims_to_arguments`` reported ``completed`` with the
|
||||
central litigant of the appeal missing entirely.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from legal_mcp.services.argument_aggregator import (
|
||||
MAX_PROPS_PER_CALL,
|
||||
AggregationFailed,
|
||||
_aggregate_party,
|
||||
_chunk,
|
||||
)
|
||||
|
||||
|
||||
def test_chunk_preserves_every_proposition_and_their_order():
|
||||
"""Chunking must not drop or reorder — losing claims here is invisible."""
|
||||
props = [{"i": i} for i in range(310)]
|
||||
|
||||
chunks = _chunk(props, MAX_PROPS_PER_CALL)
|
||||
|
||||
assert sum(len(c) for c in chunks) == 310, "propositions were lost"
|
||||
assert [p for c in chunks for p in c] == props, "order changed"
|
||||
assert all(len(c) <= MAX_PROPS_PER_CALL for c in chunks)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_non_list_reply_raises_instead_of_dropping_the_side(monkeypatch):
|
||||
"""A malformed reply must surface, never look like an empty side.
|
||||
|
||||
This is the exact 1069-04-26 failure. If this test ever goes back to
|
||||
asserting ``== []``, the silent-drop bug has been reintroduced.
|
||||
"""
|
||||
async def _query_json(prompt, tools=""): # noqa: ARG001
|
||||
return {"error": "not a list"}
|
||||
|
||||
monkeypatch.setattr(
|
||||
"legal_mcp.services.argument_aggregator.claude_session.query_json",
|
||||
_query_json,
|
||||
)
|
||||
|
||||
with pytest.raises(AggregationFailed) as excinfo:
|
||||
await _aggregate_party("appellant", [{"id": "x", "claim_text": "t"}])
|
||||
|
||||
# The message has to name the side, or an operator reading
|
||||
# completed_with_errors cannot tell which litigant went missing.
|
||||
assert "appellant" in str(excinfo.value)
|
||||
67
mcp-server/tests/test_source_kind_selector.py
Normal file
67
mcp-server/tests/test_source_kind_selector.py
Normal file
@@ -0,0 +1,67 @@
|
||||
"""#232 — the source_kind selector must not silently hide a corpus.
|
||||
|
||||
`search_precedent_library` defaulted to source_kind='external_upload' all the
|
||||
way down the stack, so 104 appeals-committee decisions (29% of the corpus)
|
||||
were unreachable through the entry point the writing agents actually call.
|
||||
These tests pin the selector semantics; the end-to-end retrieval check lives
|
||||
in scripts/test_retrieval_by_name.py (needs a live DB).
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
from legal_mcp.services.db import _source_kind_clause
|
||||
|
||||
|
||||
def test_empty_selector_means_no_filter():
|
||||
"""'' and 'all' must produce no predicate — the whole corpus."""
|
||||
assert _source_kind_clause("") == ""
|
||||
assert _source_kind_clause("all") == ""
|
||||
assert _source_kind_clause(" ") == ""
|
||||
|
||||
|
||||
def test_named_kinds_produce_equality_predicate():
|
||||
assert _source_kind_clause("internal_committee", "cl.") == (
|
||||
"cl.source_kind = 'internal_committee'"
|
||||
)
|
||||
assert _source_kind_clause("external_upload") == "source_kind = 'external_upload'"
|
||||
|
||||
|
||||
def test_all_committees_expands_to_both_shapes():
|
||||
"""Committee decisions live under two historical shapes — cover both."""
|
||||
clause = _source_kind_clause("all_committees", "cl.")
|
||||
assert "cl.source_kind = 'internal_committee'" in clause
|
||||
assert "cl.source_type = 'appeals_committee'" in clause
|
||||
assert clause.startswith("(") and clause.endswith(")")
|
||||
|
||||
|
||||
def test_unknown_selector_raises_rather_than_reaching_sql():
|
||||
"""The value is f-string-interpolated, so the whitelist is the guard."""
|
||||
with pytest.raises(ValueError, match="source_kind"):
|
||||
_source_kind_clause("'; DROP TABLE case_law; --")
|
||||
with pytest.raises(ValueError):
|
||||
_source_kind_clause("internal")
|
||||
|
||||
|
||||
def test_default_of_the_search_entry_points_is_whole_corpus():
|
||||
"""A regression guard on the defaults themselves — this is the bug."""
|
||||
import inspect
|
||||
|
||||
from legal_mcp.services import hybrid_search, precedent_library
|
||||
from legal_mcp.services import db as db_mod
|
||||
from legal_mcp.tools import precedent_library as plib_tool
|
||||
|
||||
for fn in (
|
||||
db_mod.search_precedent_library_semantic,
|
||||
db_mod.search_precedent_library_lexical,
|
||||
db_mod.list_external_case_law,
|
||||
hybrid_search.search_precedent_library_hybrid,
|
||||
precedent_library.search_library,
|
||||
precedent_library.list_precedents,
|
||||
plib_tool.search_precedent_library,
|
||||
plib_tool.precedent_library_list,
|
||||
):
|
||||
default = inspect.signature(fn).parameters["source_kind"].default
|
||||
assert default == "", (
|
||||
f"{fn.__module__}.{fn.__qualname__} defaults source_kind to "
|
||||
f"{default!r} — that hides a corpus from every caller (#232)"
|
||||
)
|
||||
@@ -26,6 +26,7 @@
|
||||
| `test_retrieval_by_name.py` | python | בדיקת אחזור-לפי-שם (#52/RC-A) — מאמת ש`search_precedent_library`/`search_internal_decisions` מדרגים את ההחלטה עצמה (אגסי) מעל מי שמצטט אותה, + רגרסיות לשאילתות מהותיות. הרצה: `DOTENV_PATH=/home/chaim/.env DATA_DIR=.../data mcp-server/.venv/bin/python scripts/test_retrieval_by_name.py` (exit 0 = עבר). | ידני אחרי שינוי שכבת חיפוש |
|
||||
| `eval_gold_bootstrap.py` | python | **FU-5 (GAP-11) — bootstrap ל-gold-set** של הערכת-אחזור ל-`data/eval/gold-set.jsonl`. שני מקורות: `--source citations` (cited==relevant מ-`search_relevance_feedback`; ריק עד שייצברו ציטוטים) ו-`--source known_item` (query=שם-תיק → relevant=עצמו; אות אמיתי היום). Idempotent — שומר שורות `source=chair`, מחדש `bootstrap_*`. דורש POSTGRES. | לפני eval; חוזר כשנצבר ground-truth |
|
||||
| `eval_retrieval.py` | python | **FU-5 (GAP-11, INV-RET4/G8) — harness הערכת-אחזור** — מריץ את מסלול-האחזור בייצור (`search_library`/`search_internal`) על ה-gold-set, מחשב precision@k/recall@k/MRR/nDCG@k (k=5,10), מצרף overall+per-corpus+per-PA ל-`data/eval/eval-report-<ts>.{json,md}` + delta מול `data/eval/baseline.json` (מתעד retrieval_config). `--self-test` בודק את המטריקות offline; `--update-baseline` מאמץ snapshot. **שער-CI במשמעת:** הרץ לפני/אחרי כל שינוי בשכבת-האחזור באותו קונפיג. דורש POSTGRES+VOYAGE_API_KEY. | לפני/אחרי שינוי RRF/k/embedder/rerank |
|
||||
| `legal-mcp-http.config.cjs` | pm2/js | **שרת ה-MCP חשוף ב-streamable-http** (#231) — `python -m legal_mcp.server` עם `MCP_TRANSPORT=streamable-http`, bound **`127.0.0.1:8790`**, Bearer `MCP_HTTP_SHARED_SECRET` מ-`~/.legal-mcp-http.env` (מקור-אמת: Infisical → All Infrastructure / `main` / `/apps/legal-ai`, תג `credentials`). **למה:** סוכנים המונעים דרך Agent Client Protocol מקבלים את שרתי-ה-MCP שלהם מהלקוח בפתיחת הסשן, והערוץ הזה נושא שרתי HTTP בלבד — ל-stdio אין מסלול לשם, ומכאן שהסוכנים נותרו בלי 108 הכלים. **אינו מחליף את stdio:** כל סשן אינטראקטיבי ממשיך דרך `~/.claude.json`; אותו קוד, אותו מרשם-כלים, שתי תחבורות (G2). **אבטחה:** loopback בלבד (צר יותר מ-`10.0.1.1` של legal-chat-service — שום קונטיינר לא צריך MCP), והשרת **מסרב לעלות בלי טוקן** (`services/http_auth.py`), כך שתקלת-הגדרה לא יכולה לייצר מאזין לא-מאומת. מראָה לדפוס `legal-chat-service.config.cjs`. התקנה: `pm2 start scripts/legal-mcp-http.config.cjs && pm2 save`. בדיקה: POST ל-`/mcp` → 401 בלי טוקן, 200 עם. | pm2 (host-side) |
|
||||
| `legal-court-fetch-service.config.cjs` | pm2/js | **שירות-מארח Tier-1 לאחזור פסקי-דין מנט המשפט (X13)** — 2 apps: (א) `legal-court-fetch-xvfb` (Xvfb :99, צג-וירטואלי ל-Camoufox); (ב) `legal-court-fetch-service` (`python -m legal_mcp.court_fetch_service.server`, bound `10.0.1.1:8771`, Bearer `COURT_FETCH_SHARED_SECRET` מ-`~/.legal-court-fetch-service.env`, `DISPLAY=:99`). מריץ Camoufox דרך חבילת-הפייתון (in-process) כי הקונטיינר לא יכול דפדפן. תלות: `pip install -e "mcp-server[court-fetch]" && python -m camoufox fetch`. אחזור = ניווט→צופה→`GetImages`(X-Requested-With)→PDF, ללא CAPTCHA; כשל→`ok:false`→orchestrator מסלים ל-fallback אנושי. **אומת על עת"מ 46111-12-22 (34 עמ').** מראָה לדפוס `legal-chat-service.config.cjs`. ספ: `docs/spec/X13-court-fetch.md`. התקנה: `pm2 start scripts/legal-court-fetch-service.config.cjs && pm2 save`. בריאות: `curl http://10.0.1.1:8771/health`. | pm2 (host-side) |
|
||||
| `drain_court_fetch.py` | python | **ריקון תור-אחזור הפסיקה (X13)** — קורא ל-`court_fetch_orchestrator.drain_pending(limit)` שמוריד+קולט כל job ממתין שהיומונים מילאו, וקושר חזרה ליומון. מקומי בלבד (ingest = claude CLI). no-op מהיר כשהתור ריק. הרצה ידנית: `mcp-server/.venv/bin/python scripts/drain_court_fetch.py [limit]`. | דרך `legal-court-fetch-drain.config.cjs` (pm2 cron) |
|
||||
| `legal-court-fetch-drain.config.cjs` | pm2/js | **תזמון שעתי של `drain_court_fetch.py`** (cron `17 * * * *`, `COURT_FETCH_DRAIN_CRON` לעקיפה) — הופך את לולאת יומון→אחזור→קליטה ל-fully-autonomous. `autorestart:false` (one-shot per tick). דורש `legal-court-fetch-service` רץ. התקנה: `pm2 start scripts/legal-court-fetch-drain.config.cjs && pm2 save`. | pm2 cron (host-side) |
|
||||
@@ -98,6 +99,7 @@
|
||||
|--------|------|---------|-----------|
|
||||
| `spec-guard.sh` | bash | **PreToolUse hook לאכיפת "פרוטוקול כתיבת-קוד"** (CLAUDE.md §פרוטוקול כתיבת-קוד) — בכל Edit/Write/MultiEdit על נתיב-קוד (`web/`, `mcp-server/`, `web-ui/src/`, `scripts/`, `adapters/`) מזריק תזכורת ל-Claude לקרוא את `docs/spec/00-constitution.md`+ספ-התחום ולוודא קיום G1–G12 — לפני שכותבים. **+ leak-guard בזמן-אמת (G12):** על כתיבה ל-`mcp-server/src/*` בודק את התוכן-הנכתב (`new_string`/`content`) ומזהיר אם מוזרק מונח-Paperclip לשכבת-האינטליגנציה (לא-deduped). המקבילה האינטראקטיבית ל-INV-AG1. קלט JSON ב-stdin, פלט `hookSpecificOutput.additionalContext` (non-blocking, exit 0). Dedup פעם-בסשן לתזכורת-הספ. רשום ב-`.claude/settings.json`. | נקרא אוטומטית ע"י Claude Code (hook) |
|
||||
| `leak_guard.py` | python | **המאכף הקנוני של INV-G12 (שער-הפלטפורמה / docs/spec/X15 §4 / R4).** שני כללים קשיחים: (1) `mcp-server/src` ללא סמלי-Paperclip (allowlist מנומק לפי substring); (2) רק `web/agent_platform_port.py` (+ קבצי-המעטפת) מייבאים את לקוח-Paperclip. stdlib-בלבד (אין venv). `leak_guard.py` = סריקת-repo (exit 1 על הפרה); `leak_guard.py <file>...` = קבצים נתונים (ל-hook). משותף ל-spec-guard.sh (hook), ל-CI (`.gitea/workflows/leak-guard.yaml`) ול-`mcp-server/tests/test_platform_port_leak_guard.py`. | CI + hook + pytest |
|
||||
| `agent_tool_grants_guard.py` | python | **המאכף הקנוני של INV-AG3 (מפת-הרשאות הסוכנים / docs/spec/X4-agents.md §2א).** ה-frontmatter `tools:` של סוכן הוא **allow-list סגורה** — כלי הרשום בשרת-ה-MCP אך חסר ממנה אינו ניתן לקריאה, גם כשהשרת מחובר. ארבעה כללים קשיחים: (1) כל `mcp__legal-ai__X` המופיע ב-`web/` (delegation שיוצר issue לסוכן) מוענק לסוכן כלשהו; (2) כל `mcp__legal-ai__X` בגוף קובץ-סוכן מוענק ב-frontmatter של **אותו** קובץ; (3) אין הענקה לכלי שאינו רשום ב-`@mcp.tool`; (4) שם-כלי בגרשיים-הפוכים ללא תחילית — מוענק, או מסווג ב-`CONTRASTIVE_OK` עם נימוק (כולל בדיקת-התיישנות לסיווגים). מחריג קבצים שאינם סוכני-claude_local: `hermes-curator.md`, `legal-analyst-gemini-critique.md`, `HEARTBEAT.md`. **כלל 5 (host-only, `--check-bindings`):** allow-list נאכפת רק כשה-runtime בוחר את הסוכן (`--agent <name>`); בלעדיו אותו קובץ נמסר כ-`--append-system-prompt-file` — פרוזה, לא שער — וכל הכלים נשארים נגישים. הקשירה יושבת ב-DB של הפלטפורמה ולכן ה-CI לא רואה אותה; בלי הדגל השער **אומר זאת במפורש** במקום לרמוז על אכיפה שלא אימת. stdlib-בלבד. נבנה אחרי CMP-229 (2026-08-04) — `analyze_protocol` נרשם בשרת ב-2026-06-30 בלי הענקה, ו-#226 הורה למנתח להריץ אותו. CI: `.gitea/workflows/agent-tool-grants.yaml`. | CI |
|
||||
| `check_undefined_names.py` | python | **CI gate ל-undefined names (מחלקת ה-NameError).** מריץ pyflakes על `web`, `mcp-server/src`, `scripts` ומפיל build (exit 1) רק על "undefined name"/"may be undefined" — לא על imports-לא-בשימוש/f-strings (רעש). זו בדיוק מחלקת-הבאג של PR #249 (שינוי-שם תיק → 500): שם שמופנה אך לא מיובא/מוגדר, חבוי בתוך `background_tasks` עד זמן-ריצה. דורש pyflakes (ה-workflow מתקין ל-venv זמני). משותף ל-CI (`.gitea/workflows/lint.yaml`). | CI |
|
||||
| `auto-sync-cases.sh` | bash | סנכרון תיקי ערר ל-Gitea — רץ כל דקה | `* * * * *` (cron) |
|
||||
| `host_sync.sh` | bash | מסנכרן את עץ-המארח `~/legal-ai` ל-origin/main (ff-only) כדי שקוד-המארח (כותב/פאנלים/MCP שרצים מהעץ, לא בקונטיינר) יתעדכן אחרי merge; restart מדויק ל-chat/court-fetch/reaper רק כשקבציהם משתנים. בטוח: אף-פעם לא force; tasks.json הדירטי נשמר. סוגר את פער-פריסת-המארח (TaskMaster #160) | `* * * * *` (cron, flock) |
|
||||
|
||||
319
scripts/agent_tool_grants_guard.py
Executable file
319
scripts/agent_tool_grants_guard.py
Executable file
@@ -0,0 +1,319 @@
|
||||
#!/usr/bin/env python3
|
||||
"""INV-AG3 guard — every MCP tool an agent is TOLD to run must be GRANTED to it.
|
||||
|
||||
The canonical checker for INV-AG3 (docs/spec/X4-agents.md §2א): a Claude-Code
|
||||
subagent's ``tools:`` frontmatter is a CLOSED allow-list. A tool that is
|
||||
registered on the MCP server but absent from that list is *not callable* by the
|
||||
agent, however well the server is connected.
|
||||
|
||||
Why this exists — the failure it is built to catch (2026-08-04):
|
||||
``analyze_protocol`` shipped on 2026-06-30 (24e3e2f) touching 9 files, none of
|
||||
them ``.claude/agents/*``. Later ``wake_analyst_for_protocol_analysis``
|
||||
(web/paperclip_client.py, #226) started writing "הרץ
|
||||
``mcp__legal-ai__analyze_protocol(...)``" straight into the analyst's issue.
|
||||
The analyst therefore received an explicit instruction to run a tool it was
|
||||
never granted, reported "tools exist on the connected server but aren't exposed
|
||||
as callable in this session", and burned two runs working around it via raw
|
||||
psql + a hand-written script. INV-AG3 already covered this on paper; its
|
||||
enforcement was deferred ("אכיפה אוטומטית עתידית"), so the drift went unnoticed
|
||||
for five weeks. This script is that deferred enforcement.
|
||||
|
||||
Three HARD rules:
|
||||
|
||||
1. **Backend delegation.** Every ``mcp__legal-ai__X`` named inside ``web/``
|
||||
(the backend telling an agent what to run) must be granted to at least one
|
||||
agent. This is the rule that catches the 2026-08-04 failure.
|
||||
|
||||
2. **Per-agent instructions.** Every ``mcp__legal-ai__X`` in an agent file's
|
||||
BODY must be granted in that same file's frontmatter. Prefixed mentions are
|
||||
imperative by convention ("הרץ `mcp__legal-ai__…`").
|
||||
|
||||
3. **No phantom grants.** Every granted tool must actually be registered on
|
||||
the MCP server — catches typos and tools deleted out from under an agent.
|
||||
|
||||
Plus one reviewed-exception rule:
|
||||
|
||||
4. **Bare tool names.** An agent body may name a tool in backticks without the
|
||||
``mcp__legal-ai__`` prefix (```get_legal_arguments```). Those are
|
||||
ambiguous: some are real requirements, others are deliberately contrastive
|
||||
("**לא** דרך `precedent_library_upload`"), a pointer at *another* agent's
|
||||
job, or a DB column that merely shares a tool's name. Each is classified
|
||||
once in ``CONTRASTIVE_OK`` below; anything new fails until reviewed.
|
||||
|
||||
And one host-only rule, opt-in via ``--check-bindings``:
|
||||
|
||||
5. **Bindings.** Rules 1–4 compare files to files, which says nothing about
|
||||
whether an allow-list is *enforced*. It is only enforced when the runtime
|
||||
selects that agent (``--agent <name>``); without the flag the same file is
|
||||
delivered as ``--append-system-prompt-file`` — prose, not a gate — and every
|
||||
tool stays reachable. Found on 2026-08-05: one agent declared 41 grants with
|
||||
no ``--agent`` flag, so the largest allow-list in the system was inert while
|
||||
this guard reported OK. The binding lives in the platform DB, so CI cannot
|
||||
see it; without the flag the guard now says so out loud instead of implying
|
||||
enforcement it never verified.
|
||||
|
||||
NOT AGENTS (no frontmatter by design — the adapter sends the file as a raw
|
||||
prompt, so YAML would leak into it): ``hermes-curator.md`` (deepseek_local),
|
||||
``legal-analyst-gemini-critique.md`` (gemini_local). ``HEARTBEAT.md`` is a
|
||||
shared checklist, not an agent. All three are skipped.
|
||||
|
||||
Usage:
|
||||
agent_tool_grants_guard.py # exit 1 on any violation
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO = Path(__file__).resolve().parent.parent
|
||||
|
||||
AGENTS_DIR = REPO / ".claude" / "agents"
|
||||
MCP_SRC = REPO / "mcp-server" / "src"
|
||||
BACKEND_DIR = REPO / "web"
|
||||
|
||||
# Files under .claude/agents/ that are not claude_local subagent definitions.
|
||||
NOT_AGENTS = {
|
||||
"HEARTBEAT.md",
|
||||
"hermes-curator.md",
|
||||
"legal-analyst-gemini-critique.md",
|
||||
}
|
||||
|
||||
# Bare (unprefixed) tool names in an agent body that are NOT requirements.
|
||||
# Each entry is (agent file, tool, why) — reviewed 2026-08-04. Adding to this
|
||||
# map is a deliberate act: it asserts "the agent is not being told to call this".
|
||||
CONTRASTIVE_OK = {
|
||||
("legal-analyst.md", "case_create"): "prose about the cases.practice_area CHECK constraint, not a call",
|
||||
("legal-analyst.md", "search_internal_decisions"): "names the filter surface when contrasting Axis A/B",
|
||||
("legal-ceo.md", "search_decisions"): "contrast — 'search_decisions = only Dafna' vs the granted search_internal_decisions",
|
||||
("legal-ceo.md", "precedent_library_upload"): "explicitly the forbidden path ('לא דרך …', citation guard rejects)",
|
||||
("legal-ceo.md", "document_update"): "describes the tagging chaim must fix, not a CEO call",
|
||||
("legal-proofreader.md", "extraction_status"): "the documents.extraction_status DB column — name collides with a tool",
|
||||
("legal-qa.md", "precedent_attach"): "explicitly the researcher's job ('דרך precedent_attach של ה-researcher')",
|
||||
("legal-writer.md", "revise_draft"): "the CEO calls it ('CEO יקרא ל-revise_draft'), not the writer",
|
||||
("legal-writer.md", "search_case_precedents"): "a do-not-confuse disambiguation note ('שונה! … לא לבלבל')",
|
||||
}
|
||||
|
||||
TOOL_RE = re.compile(r"mcp__legal-ai__(\w+)")
|
||||
REGISTER_RE = re.compile(r"@mcp\.tool\([^)]*\)\s*(?:async\s+)?def\s+(\w+)")
|
||||
# `tool_name(` or `tool_name` inside backticks.
|
||||
BARE_RE = re.compile(r"`(\w+)[(`]")
|
||||
|
||||
|
||||
def server_tools() -> set[str]:
|
||||
"""Tool names registered on the MCP server."""
|
||||
out: set[str] = set()
|
||||
for path in MCP_SRC.rglob("*.py"):
|
||||
out |= set(REGISTER_RE.findall(path.read_text(encoding="utf-8", errors="ignore")))
|
||||
return out
|
||||
|
||||
|
||||
def split_frontmatter(text: str) -> tuple[str, str]:
|
||||
"""Return (frontmatter, body). Empty frontmatter when the file has none."""
|
||||
if not text.startswith("---"):
|
||||
return "", text
|
||||
parts = text.split("---")
|
||||
if len(parts) < 3:
|
||||
return "", text
|
||||
return parts[1], "---".join(parts[2:])
|
||||
|
||||
|
||||
def agent_files() -> list[Path]:
|
||||
return sorted(p for p in AGENTS_DIR.glob("*.md") if p.name not in NOT_AGENTS)
|
||||
|
||||
|
||||
def check_bindings() -> list[tuple[str, str]]:
|
||||
"""Return [(agent file, why)] for agents whose allow-list nothing enforces.
|
||||
|
||||
A ``tools:`` list is only an allow-list when the runtime is told which agent
|
||||
to be. The local adapter enforces it under ``--agent <name>``; without that
|
||||
flag the very same file is delivered as ``--append-system-prompt-file``, i.e.
|
||||
prose the model may follow or ignore, and every tool stays reachable.
|
||||
|
||||
Found the hard way on 2026-08-05: one agent carried 41 grants and no
|
||||
``--agent`` flag, so the largest allow-list in the system was inert — and
|
||||
this guard had been reporting OK on it, because Rules 1–4 only ever compare
|
||||
files to files.
|
||||
|
||||
Host-only. The binding lives in the platform's database, which CI cannot
|
||||
reach, so this shells out to ``psql`` rather than adding a driver dependency
|
||||
that would break the stdlib-only property the CI path relies on. Returns []
|
||||
when the database is unreachable — an unreachable DB is "not checked", not
|
||||
"no violations", and the caller prints that distinction.
|
||||
"""
|
||||
|
||||
sql = (
|
||||
"select adapter_config->>'instructionsEntryFile', "
|
||||
"coalesce(adapter_config->>'extraArgs','') "
|
||||
"from agents where adapter_type='claude_local' "
|
||||
"and adapter_config->>'instructionsEntryFile' is not null;"
|
||||
)
|
||||
try:
|
||||
out = subprocess.run(
|
||||
["psql", "-h", "localhost", "-p", "54329", "-U", "paperclip",
|
||||
"-d", "paperclip", "-X", "-A", "-t", "-F", "\t", "-c", sql],
|
||||
capture_output=True, text=True, timeout=20,
|
||||
env={**os.environ, "PGPASSWORD": os.environ.get("PGPASSWORD", "paperclip")},
|
||||
)
|
||||
except (OSError, subprocess.SubprocessError):
|
||||
return []
|
||||
if out.returncode != 0:
|
||||
return []
|
||||
|
||||
bad: dict[str, str] = {}
|
||||
for line in out.stdout.splitlines():
|
||||
if "\t" not in line:
|
||||
continue
|
||||
entry_file, extra = line.split("\t", 1)
|
||||
entry_file = entry_file.strip()
|
||||
if not entry_file or entry_file in NOT_AGENTS:
|
||||
continue
|
||||
want = entry_file[:-3] if entry_file.endswith(".md") else entry_file
|
||||
try:
|
||||
args = json.loads(extra) if extra.strip() else []
|
||||
except json.JSONDecodeError:
|
||||
args = []
|
||||
# Only a literal ["--agent", "<name>"] pair binds the allow-list.
|
||||
ok = any(
|
||||
a == "--agent" and i + 1 < len(args) and args[i + 1] == want
|
||||
for i, a in enumerate(args)
|
||||
)
|
||||
if not ok:
|
||||
bad[entry_file] = (
|
||||
"extraArgs is empty" if not args
|
||||
else f"extraArgs={extra.strip()} does not select '{want}'"
|
||||
)
|
||||
|
||||
# Only report agents that actually declare grants — an agent with no tools:
|
||||
# list has nothing to enforce and is not a finding.
|
||||
result = []
|
||||
for path in agent_files():
|
||||
if path.name in bad:
|
||||
fm, _ = split_frontmatter(path.read_text(encoding="utf-8", errors="ignore"))
|
||||
if TOOL_RE.findall(fm):
|
||||
result.append((path.name, bad[path.name]))
|
||||
return sorted(result)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
registered = server_tools()
|
||||
if not registered:
|
||||
print("agent-tool-grants: FAIL — no @mcp.tool registrations found; is the tree complete?")
|
||||
return 1
|
||||
|
||||
grants: dict[str, set[str]] = {}
|
||||
bodies: dict[str, str] = {}
|
||||
for path in agent_files():
|
||||
fm, body = split_frontmatter(path.read_text(encoding="utf-8", errors="ignore"))
|
||||
grants[path.name] = set(TOOL_RE.findall(fm))
|
||||
bodies[path.name] = body
|
||||
|
||||
all_granted: set[str] = set().union(*grants.values()) if grants else set()
|
||||
violations: list[str] = []
|
||||
|
||||
# Rule 1 — backend delegation must land on a granted tool.
|
||||
for path in sorted(BACKEND_DIR.rglob("*.py")):
|
||||
text = path.read_text(encoding="utf-8", errors="ignore")
|
||||
for tool in sorted(set(TOOL_RE.findall(text))):
|
||||
if tool not in all_granted:
|
||||
rel = path.relative_to(REPO)
|
||||
violations.append(
|
||||
f"[1 backend] {rel} instructs an agent to run "
|
||||
f"mcp__legal-ai__{tool}, but NO agent grants it.\n"
|
||||
f" fix: add `- mcp__legal-ai__{tool}` to the tools: "
|
||||
f"frontmatter of the agent that receives that issue."
|
||||
)
|
||||
|
||||
# Rule 2 — a prefixed mention in an agent body is an instruction to that agent.
|
||||
for name, body in bodies.items():
|
||||
for tool in sorted(set(TOOL_RE.findall(body))):
|
||||
if tool not in grants[name]:
|
||||
violations.append(
|
||||
f"[2 instructions] .claude/agents/{name} tells the agent to run "
|
||||
f"mcp__legal-ai__{tool}, which its own tools: list omits.\n"
|
||||
f" fix: add `- mcp__legal-ai__{tool}` to that frontmatter."
|
||||
)
|
||||
|
||||
# Rule 3 — no grant may point at a tool the server does not register.
|
||||
for name, granted in grants.items():
|
||||
for tool in sorted(granted - registered):
|
||||
violations.append(
|
||||
f"[3 phantom] .claude/agents/{name} grants mcp__legal-ai__{tool}, "
|
||||
f"which is not registered on the MCP server.\n"
|
||||
f" fix: correct the name, or drop the grant if the tool was removed."
|
||||
)
|
||||
|
||||
# Rule 4 — every bare tool name is either granted or classified as contrastive.
|
||||
for name, body in bodies.items():
|
||||
bare = {m for m in BARE_RE.findall(body) if m in registered}
|
||||
for tool in sorted(bare - grants[name]):
|
||||
if (name, tool) in CONTRASTIVE_OK:
|
||||
continue
|
||||
violations.append(
|
||||
f"[4 bare name] .claude/agents/{name} mentions `{tool}` — a real MCP "
|
||||
f"tool it is not granted.\n"
|
||||
f" fix: grant it if the agent must call it, otherwise add "
|
||||
f"(\"{name}\", \"{tool}\") to CONTRASTIVE_OK with the reason."
|
||||
)
|
||||
|
||||
# Stale exceptions: a classification that no longer matches the text is noise.
|
||||
for (name, tool), _why in sorted(CONTRASTIVE_OK.items()):
|
||||
if name not in bodies:
|
||||
violations.append(
|
||||
f"[4 stale] CONTRASTIVE_OK names {name}, which is not an agent file."
|
||||
)
|
||||
elif tool not in {m for m in BARE_RE.findall(bodies[name])}:
|
||||
violations.append(
|
||||
f"[4 stale] CONTRASTIVE_OK ({name}, {tool}) no longer appears in that "
|
||||
f"file — drop the exception."
|
||||
)
|
||||
|
||||
# Rule 5 — a grant list only binds if the runtime actually selects that agent.
|
||||
unenforced = check_bindings() if "--check-bindings" in sys.argv else None
|
||||
if unenforced:
|
||||
for name, detail in unenforced:
|
||||
violations.append(
|
||||
f"[5 binding] .claude/agents/{name} declares a tools: allow-list, but "
|
||||
f"the runtime does not select that agent — {detail}.\n"
|
||||
f" The list is inert: it is delivered as prompt text only, so "
|
||||
f"every tool remains callable.\n"
|
||||
f" fix: set adapter_config.extraArgs to "
|
||||
f'["--agent", "{name[:-3]}"], or drop tools: and document the agent as '
|
||||
f"unrestricted. Not both."
|
||||
)
|
||||
|
||||
if violations:
|
||||
print(f"INV-AG3 agent-tool-grants guard: {len(violations)} violation(s)\n")
|
||||
for v in violations:
|
||||
print(f" ✗ {v}")
|
||||
print(
|
||||
"\ndocs/spec/X4-agents.md §2א INV-AG3 — the frontmatter tools: list is a "
|
||||
"CLOSED allow-list.\nA tool missing from it is not callable, no matter that "
|
||||
"the MCP server is connected."
|
||||
)
|
||||
return 1
|
||||
|
||||
print(
|
||||
f"INV-AG3 agent-tool-grants guard: OK "
|
||||
f"({len(agent_files())} agents, {len(all_granted)} distinct grants, "
|
||||
f"{len(registered)} tools registered)"
|
||||
)
|
||||
if unenforced is None:
|
||||
# Say plainly what was NOT checked. A guard that prints a bare "OK" invites
|
||||
# the reader to conclude the allow-lists are enforced; this one has only
|
||||
# compared files to files. Enforcement is a runtime property (see Rule 5),
|
||||
# and on 2026-08-05 exactly one agent was found declaring 41 grants that
|
||||
# nothing enforces — while this guard reported OK.
|
||||
print(
|
||||
" note: file-level only. Whether each allow-list is actually ENFORCED "
|
||||
"depends on the\n runtime passing --agent <name>, which needs the platform "
|
||||
"DB — re-run with --check-bindings\n on the host to verify."
|
||||
)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
96
scripts/legal-mcp-http.config.cjs
Normal file
96
scripts/legal-mcp-http.config.cjs
Normal file
@@ -0,0 +1,96 @@
|
||||
/**
|
||||
* pm2 ecosystem entry for legal-mcp-http — the legal-ai MCP server exposed over
|
||||
* streamable-http (TaskMaster #231).
|
||||
*
|
||||
* Why it exists
|
||||
* Agents driven over the Agent Client Protocol get their MCP servers from the
|
||||
* *client* at session start, and that channel carries HTTP servers only. A
|
||||
* stdio server has no path into such a session, which is how platform-driven
|
||||
* agents ended up with none of the 108 tools. This service is the HTTP end
|
||||
* they can actually be pointed at.
|
||||
*
|
||||
* It does NOT replace the stdio path. Every interactive Claude Code session
|
||||
* still reaches the same server through the `legal-ai` entry in
|
||||
* ~/.claude.json, spawned per session. Same code, same tool registry, two
|
||||
* transports (G2) — this is a second *door*, not a second server.
|
||||
*
|
||||
* Security
|
||||
* The registry includes case_delete, precedent_library_delete, document_upload
|
||||
* and every block-writing tool, so an open port here is a delete-any-case
|
||||
* endpoint. Two defences, both required:
|
||||
* 1. Bind 127.0.0.1 — the platform runs on this host, so loopback suffices.
|
||||
* Deliberately narrower than legal-chat-service's 10.0.1.1: nothing in a
|
||||
* container needs to call MCP.
|
||||
* 2. Bearer token — MCP_HTTP_SHARED_SECRET, loaded below. The server
|
||||
* REFUSES TO START without it (services/http_auth.py), so a
|
||||
* misconfiguration cannot silently produce an unauthenticated listener.
|
||||
*
|
||||
* Secret
|
||||
* Source of truth: Infisical, project "All Infrastructure", env `main`,
|
||||
* /apps/legal-ai/MCP_HTTP_SHARED_SECRET (tag: credentials). The file read
|
||||
* below is a chmod-600 runtime copy, same arrangement as
|
||||
* legal-chat-service.config.cjs. Rotate in Infisical first, then refresh the
|
||||
* file and `pm2 restart legal-mcp-http`.
|
||||
*
|
||||
* Install (once):
|
||||
* pm2 start /home/chaim/legal-ai/scripts/legal-mcp-http.config.cjs
|
||||
* pm2 save
|
||||
*
|
||||
* Smoke test — expect 401 without the token, 200 with it:
|
||||
* curl -s -o /dev/null -w '%{http_code}\n' -X POST http://127.0.0.1:8790/mcp \
|
||||
* -H 'Content-Type: application/json' \
|
||||
* -H 'Accept: application/json, text/event-stream' \
|
||||
* -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}'
|
||||
*
|
||||
* Update: pm2 restart legal-mcp-http --update-env
|
||||
* Stop: pm2 stop legal-mcp-http
|
||||
*/
|
||||
const fs = require("fs");
|
||||
|
||||
const ENV_FILE = "/home/chaim/.legal-mcp-http.env";
|
||||
const env = {
|
||||
HOME: "/home/chaim",
|
||||
PATH: "/home/chaim/.local/bin:/usr/local/bin:/usr/bin:/bin",
|
||||
PYTHONUNBUFFERED: "1",
|
||||
// Same DB/data wiring the stdio server gets from ~/.claude.json, so both
|
||||
// transports read exactly the same corpus.
|
||||
DOTENV_PATH: "/home/chaim/.env",
|
||||
DATA_DIR: "/home/chaim/legal-ai/data",
|
||||
MCP_TRANSPORT: "streamable-http",
|
||||
MCP_HTTP_HOST: "127.0.0.1",
|
||||
MCP_HTTP_PORT: "8790",
|
||||
};
|
||||
|
||||
try {
|
||||
const text = fs.readFileSync(ENV_FILE, "utf8");
|
||||
for (const line of text.split("\n")) {
|
||||
if (!line || line.trim().startsWith("#")) continue;
|
||||
const m = line.match(/^\s*([A-Z_][A-Z0-9_]*)\s*=\s*(.*?)\s*$/);
|
||||
if (m) env[m[1]] = m[2];
|
||||
}
|
||||
} catch (e) {
|
||||
// Warn, but do not fabricate a token. The server's own gate turns a missing
|
||||
// secret into a refusal to boot, which is the outcome we want — pm2 will
|
||||
// surface it as a crash loop rather than serve unauthenticated traffic.
|
||||
console.error(`legal-mcp-http: failed to load ${ENV_FILE}: ${e.message}`);
|
||||
console.error("Service will refuse to start without MCP_HTTP_SHARED_SECRET.");
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
apps: [
|
||||
{
|
||||
name: "legal-mcp-http",
|
||||
cwd: "/home/chaim/legal-ai/mcp-server",
|
||||
script: "/home/chaim/legal-ai/mcp-server/.venv/bin/python",
|
||||
args: "-m legal_mcp.server",
|
||||
env,
|
||||
restart_delay: 5000,
|
||||
// Low ceiling on purpose: if the token is missing the process exits
|
||||
// immediately, and we want pm2 to stop retrying and leave an obvious
|
||||
// errored entry rather than loop forever on a config mistake.
|
||||
max_restarts: 10,
|
||||
autorestart: true,
|
||||
max_memory_restart: "800M",
|
||||
},
|
||||
],
|
||||
};
|
||||
@@ -18,6 +18,13 @@ import { useCasePrecedents } from "@/lib/api/precedents";
|
||||
* the main case page (X17 #3). Lives inside the "טיעונים ועמדות" tab above the
|
||||
* collapsible by-party aggregated arguments. The 12-block editor + citation
|
||||
* verification moved to their own top-level tabs; /compose was deleted.
|
||||
*
|
||||
* The threshold-claim / issue cards deliberately stack in a SINGLE column.
|
||||
* They were a two-column CSS grid, but grid rows share a height: expanding one
|
||||
* card grew its row and shoved every card below it down — in *both* columns.
|
||||
* A single column keeps the jump local to what sits underneath, and gives the
|
||||
* expanded body (fields + chair editor + supporting precedents) full width
|
||||
* instead of half. Do not reintroduce `lg:grid-cols-2` here (chair, 2026-08-04).
|
||||
*/
|
||||
|
||||
function ProseSection({ title, content }: { title: string; content?: string }) {
|
||||
@@ -203,7 +210,7 @@ export function PositionsPanel({ caseNumber }: { caseNumber: string }) {
|
||||
{analysis.data.threshold_claims.length}
|
||||
</span>
|
||||
</div>
|
||||
<div className="grid gap-3 lg:grid-cols-2 items-start">
|
||||
<div className="space-y-3">
|
||||
{analysis.data.threshold_claims.map((tc) => (
|
||||
<SubsectionCard
|
||||
key={tc.id}
|
||||
@@ -225,7 +232,7 @@ export function PositionsPanel({ caseNumber }: { caseNumber: string }) {
|
||||
{analysis.data.issues.length}
|
||||
</span>
|
||||
</div>
|
||||
<div className="grid gap-3 lg:grid-cols-2 items-start">
|
||||
<div className="space-y-3">
|
||||
{analysis.data.issues.map((iss) => (
|
||||
<SubsectionCard
|
||||
key={iss.id}
|
||||
|
||||
@@ -6992,7 +6992,7 @@ async def precedent_library_list(
|
||||
precedent_level: str = "",
|
||||
source_type: str = "",
|
||||
search: str = "",
|
||||
source_kind: str = "external_upload",
|
||||
source_kind: str = "",
|
||||
limit: int = 100,
|
||||
offset: int = 0,
|
||||
):
|
||||
@@ -7020,6 +7020,7 @@ async def precedent_library_search(
|
||||
subject_tag: str = "",
|
||||
limit: int = 10,
|
||||
include_halachot: bool = True,
|
||||
source_kind: str = "",
|
||||
):
|
||||
if not q or len(q.strip()) < 2:
|
||||
return {"items": [], "count": 0}
|
||||
@@ -7032,6 +7033,7 @@ async def precedent_library_search(
|
||||
subject_tag=subject_tag,
|
||||
limit=limit,
|
||||
include_halachot=include_halachot,
|
||||
source_kind=source_kind,
|
||||
)
|
||||
return {"items": results, "count": len(results)}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user