A clean held-out test of voice learning was no longer runnable: every final-uploaded case already has its lessons folded, and lessons are stored universal/untagged so leave-one-out is impossible. Path A makes the test prospective instead — capture the generalization datapoint at the one moment it's clean. On final upload, after the draft↔final pair is created but BEFORE this case's lessons are folded (folding is a separate manual /training step), we snapshot style_distance (anti_pattern_total, golden-ratio max-deviation, change_percent) alongside the current voice-lesson pool size. Because the draft was written with only the PRIOR pool, each row is a clean "with N accumulated lessons, our draft on this unseen case scored X" datapoint. As the pool grows over cases, a downward trend = learning generalizes. - db: SCHEMA_V45 style_distance_history (append-only) + helpers voice_lesson_pool_sizes / record_style_distance_snapshot / get_style_distance_history. - app: best-effort capture in api_upload_final_decision (never fails the upload); GET /api/learning/style-distance-history for the trend. Reuses the existing style_distance service + appeal_type_rules pool — no parallel metric path. The 8 existing cases are already folded, so the table starts empty and fills from the next final (their clean window is past). Invariants: G2 (reuse style_distance/appeal_type_rules — one path), INV-LRN4 (measure the draft↔final gap; this is its trend surface). LLM-free (style_distance is deterministic) so it runs in the container. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
322 KiB
322 KiB