tune(learning): lesson-synthesis cluster threshold 0.82 → 0.68 (#158)
כיול אמפירי (dry-run על 81 לקחים): הלקחים מגוונים (חציון-דמיון ~0.48, מקס' ~0.72), כך ש-0.82 מיזג 0 אשכולות. 0.68 תופס 5 קונסולידציות אמיתיות (רישוי+השבחה+מבנה), כולן עברו את שער-ה-drift 0.80 (0.848–0.916) → נאמנות נשמרת. שער-ה-drift נשאר רשת-הביטחון מתחת לסף. Invariants: INV-LRN8 (כיול-סף בלבד; המנגנון/השערים ללא שינוי). depends-on #158. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -286,7 +286,12 @@ LESSON_SYNTH_MODEL = os.environ.get("LESSON_SYNTH_MODEL", HALACHA_EXTRACT_MODEL)
|
||||
LESSON_SYNTH_EFFORT = os.environ.get("LESSON_SYNTH_EFFORT", "high")
|
||||
LESSON_SYNTH_DRIFT_FLOOR = float(os.environ.get("LESSON_SYNTH_DRIFT_FLOOR", "0.80"))
|
||||
# Cosine floor for two lessons to land in the same cluster (greedy, within shard).
|
||||
LESSON_SYNTH_CLUSTER_THRESHOLD = float(os.environ.get("LESSON_SYNTH_CLUSTER_THRESHOLD", "0.82"))
|
||||
# Tuned empirically (#158 dry-run, 81 lessons): the lessons are diverse (median
|
||||
# pairwise cosine ~0.48, max ~0.72), so 0.82 merged nothing. 0.68 captures the real
|
||||
# near-duplicates (5 faithful clusters across rishuy+betterment, all clearing the
|
||||
# 0.80 drift gate) without merging weakly-related lessons; the drift gate is the
|
||||
# quality backstop below this.
|
||||
LESSON_SYNTH_CLUSTER_THRESHOLD = float(os.environ.get("LESSON_SYNTH_CLUSTER_THRESHOLD", "0.68"))
|
||||
|
||||
# Mistral OCR (fallback for scanned PDFs — replaces Google Cloud Vision)
|
||||
MISTRAL_API_KEY = os.environ.get("MISTRAL_API_KEY", "")
|
||||
|
||||
Reference in New Issue
Block a user