Files
legal-ai/mcp-server/tests/test_relevance_scale.py
Chaim 5a2a989e9b
All checks were successful
INV-AG3 Agent Tool Grants / agent-tool-grants (pull_request) Successful in 5s
G12 Leak-Guard / leak-guard (pull_request) Successful in 4s
Lint — undefined names / undefined-names (pull_request) Successful in 11s
fix(retrieval): סף מכויל-קוסינוס סינן פלט RRF — דף אימות-הפסיקה הציג אפס תקדימים
`hybrid_search._merge_sem_lex` דורס את `score` בערך RRF (~0.008–0.02) ברגע
שה-leg הלקסיקלי מחזיר שורות. `score` הוא אות-דירוג לגיטימי, אבל הוא מפסיק
להיות קוסינוס — ואילו `case_citation_verification._SUGGEST_FLOOR = 0.45`
כויל לקוסינוס. התוצאה: **כל שאילתה במונחים משפטיים נפוצים סוננה עד
האחרונה**, והדף הציג "אין תקדים תומך" לכל טיעון בכל תיק.

אותה פונקציה החזירה שתי סקאלות, תלוי בשאילתה:

    "המונח חזית הבניין יפורש..."   leg לקסיקלי ריק   →  0.63–0.71  (קוסינוס)
    "סמכות ועדה מקומית לפי 62א"    leg לקסיקלי מלא   →  0.016      (RRF)

מה שונה
- **עוגן קוסינוס במקור (G1):** `db.search_precedent_library_semantic` מסמן
  `relevance` לצד `score`. **לא** בפונקציה הלקסיקלית — שם `ts_rank_cd`,
  ולסמן אותו כ-relevance היה חוזר על אותה טעות בדיוק.
- **הפיוז'ן משמר:** `_merge_sem_lex` מעביר את הקוסינוס הלאה. שורה
  לקסיקלית-בלבד מקבלת `relevance = None` — "לא נמדד" אינו "נמדד כלא-רלוונטי".
- **הצרכן קורא `relevance`:** `_passes_floor()` במקום השוואה ל-`score`.
  שורות לקסיקליות-בלבד **נשמרות במודע** — הן הגיעו לצמרת בדירוג BM25 בתוך
  top-k זעיר, וסינונן היה מחביא בדיוק את התאמות-הביטוי ומספרי-התיק שהיו"ר
  מחפש בשם.

הדירוג לא השתנה: RRF ממשיך לקבוע סדר. רק הסינון עבר לסקאלה יציבה.

אימות מול הקורפוס החי:
  8124-09-24:  0 → **32 מתוך 32** טיעונים עם תקדים תומך (21.5 שנ', שלם)
  1069-04-26:  0 → **29 מתוך 69** (24.1 שנ', חלקי — תקציב הזמן)
  שורה סמנטית: score=0.0082 · relevance=0.7297

היומונים לא נפגעו ולא נגעתי בהם: `case_digest_radar` עובר דרך
`search_digests_semantic` — סמנטי טהור, בלי RRF, ולכן `min_score=0.45` שלו
מכויל נכון.

invariants: G1 (עוגן במקור, לא תיקון-סף בקריאה) · G2 (הגדרה אחת ל-relevance
לכל הצרכנים) · INV-AH (היעדר-מדידה אינו היעדר-רלוונטיות)

טסטים: 6 חדשים (tests/test_relevance_scale.py) — אחד מהם מוכיח את הבאג
ואת התיקון באותה שורה. 531 עוברים.
2026-08-05 13:31:12 +00:00

69 lines
2.9 KiB
Python

"""`score` and `relevance` are different things — thresholds must use `relevance`.
Retrieval returns cosine similarities (~0.4-0.75) until the lexical leg returns
rows; then `_merge_sem_lex` replaces `score` with an RRF value (~0.008-0.02).
Both are legitimate *ranking* signals, but they are not on the same scale, so a
caller comparing `score >= 0.45` rejected every hit the moment BM25 matched
anything. That is how the citation-verification tab came to show no supporting
precedent for a single argument, on every case.
`relevance` is the fix: always a cosine, or None when the row came from the
lexical leg alone and no cosine was ever computed.
"""
import pytest
from legal_mcp.services.case_citation_verification import _SUGGEST_FLOOR, _passes_floor
from legal_mcp.services.hybrid_search import _merge_sem_lex
def _sem(key: str, score: float) -> dict:
return {"chunk_id": key, "case_law_id": "c1", "score": score, "relevance": score}
def _lex(key: str, score: float) -> dict:
# The lexical leg emits ts_rank_cd, never a cosine — so no `relevance`.
return {"chunk_id": key, "case_law_id": "c1", "score": score}
def test_fusion_replaces_score_but_preserves_the_cosine():
"""The regression in one assertion."""
merged = _merge_sem_lex([_sem("a", 0.73)], [_lex("a", 0.31)], limit=10)
row = merged[0]
assert row["score"] < 0.1, "fused score is an RRF value, not a cosine"
assert row["relevance"] == pytest.approx(0.73), "the cosine must survive fusion"
def test_lexical_only_row_has_no_fabricated_cosine():
"""None means 'not measured' — not 'measured as irrelevant'."""
merged = _merge_sem_lex([], [_lex("b", 0.31)], limit=10)
assert merged[0]["relevance"] is None
def test_semantic_only_row_keeps_its_cosine():
merged = _merge_sem_lex([_sem("c", 0.62)], [], limit=10)
assert merged[0]["relevance"] == pytest.approx(0.62)
def test_floor_would_have_rejected_everything_on_the_fused_score():
"""Guards the exact production symptom: fused scores are ~0.008, the floor
is 0.45, so score-based filtering wipes the result set."""
merged = _merge_sem_lex(
[_sem(k, 0.70) for k in "abcd"], [_lex(k, 0.30) for k in "abcd"], limit=10)
assert all(r["score"] < _SUGGEST_FLOOR for r in merged) # the bug
assert all(_passes_floor(r) for r in merged) # the fix
def test_floor_still_rejects_genuinely_weak_hits():
"""The fix must not become 'accept everything'."""
assert not _passes_floor({"relevance": 0.10})
assert not _passes_floor({"relevance": _SUGGEST_FLOOR - 0.01})
assert _passes_floor({"relevance": _SUGGEST_FLOOR})
def test_lexical_only_hits_are_kept_deliberately():
"""An exact docket/phrase match reaches the top by BM25 rank with no cosine.
Dropping it would hide precisely what a chair searches for by name."""
assert _passes_floor({"relevance": None})
assert _passes_floor({}) # missing key behaves the same as None