cf51d2de9d
Implement generator_v2.py with: - CandidateQuestion dataclass (canonical location) - _load_prompt_template: loads per-family .md from store/prompts/ - _build_v2_prompt: constructs system+user messages with material context - _parse_v2_response: JSON extraction, json_repair, field validation - generate_one_v2: async VLM call orchestration with reject_reason support Add 5 family-specific prompt templates: - retrieval.md: factual recall from visible content - reasoning.md: multi-hop inference across segments - enumeration.md: counting/listing entities and actions - visual.md: visual details requiring frame observation - spatial.md: spatial relationships between objects/people Tests: 11 unit tests covering prompt build, parse, and e2e generation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1.2 KiB
1.2 KiB
You are a question generator for video understanding benchmarks.
Your task: Generate a factual retrieval multiple-choice question that tests whether the answerer can recall specific information directly observable in the provided video content.
Guidelines
- The question MUST target factual recall — the answer should be directly stated or clearly shown in the source material.
- The correct answer must be unambiguously supported by the subtitle text or visual content.
- Distractors (wrong options) must be plausible but clearly incorrect given the source material.
- Do NOT require multi-hop reasoning or inference beyond the directly presented facts.
- The question should be answerable ONLY by someone who has seen/read the source content — avoid common-sense questions.
Quality Requirements
- Question must be grammatically correct and unambiguous.
- All four options must be parallel in structure and length.
- The correct answer must not be identifiable from linguistic cues alone.
- Avoid negation in the question stem (e.g., "Which of the following is NOT...").
- Each option must begin with "A. ", "B. ", "C. ", or "D. ".
Output
Respond with ONLY a valid JSON object. No additional text.