feat(question_gen): add lightweight 4-gate quality check

Implement 4 concurrent LLM-based quality gates for generated questions:
- key_verify: validates answer evidence in source material
- blind_answer: rejects questions answerable without video context
- multi_true: detects ambiguous multi-correct options
- leak_test: per-family shortcut detection (5 probe templates)

Includes run_gates orchestrator with verbatim_ratio short-circuit,
JSON response parsing with fallback, and 9 unit tests (all passing).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-07-11 23:30:13 -04:00
parent 9627ac9cf9
commit 271d1682c9
10 changed files with 1030 additions and 0 deletions
@@ -0,0 +1,35 @@
# Key Verify Gate
You are a quality-control judge for video understanding questions.
## Task
Given the source material from a video and a multiple-choice question with its designated correct answer, determine whether the correct answer is **supported by evidence** in the source material.
## Source Material
{source_text}
## Question
{question}
## Options
{options}
## Designated Correct Answer
{answer}
## Instructions
1. Read the source material carefully.
2. Determine if the designated correct answer can be derived or inferred from the source material.
3. If evidence supports the answer, verdict is "pass". If not, verdict is "fail".
## Response Format (strict JSON)
```json
{{"verdict": "pass" or "fail", "reason": "brief explanation"}}
```