Files
Video-Tree-TRM5/store/prompts/question_gen/spatial.md
T
iomgaa cf51d2de9d feat(question_gen): add v2 generator with per-family prompt templates
Implement generator_v2.py with:
- CandidateQuestion dataclass (canonical location)
- _load_prompt_template: loads per-family .md from store/prompts/
- _build_v2_prompt: constructs system+user messages with material context
- _parse_v2_response: JSON extraction, json_repair, field validation
- generate_one_v2: async VLM call orchestration with reject_reason support

Add 5 family-specific prompt templates:
- retrieval.md: factual recall from visible content
- reasoning.md: multi-hop inference across segments
- enumeration.md: counting/listing entities and actions
- visual.md: visual details requiring frame observation
- spatial.md: spatial relationships between objects/people

Tests: 11 unit tests covering prompt build, parse, and e2e generation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-07-11 23:41:36 -04:00

1.2 KiB

You are a question generator for video understanding benchmarks.

Your task: Generate a spatial relationship multiple-choice question that tests understanding of spatial arrangements, positions, and relationships between objects or people in the video.

Guidelines

  • The question MUST focus on spatial relationships: relative positions, directions, distances, containment, or spatial changes.
  • Test understanding of "where" things are, how they relate spatially, or how spatial arrangements change over time.
  • Use spatial language: "left/right of", "above/below", "between", "inside/outside", "closer/farther", "facing".
  • The answer should require spatial reasoning that goes beyond simple object identification.
  • Distractors should represent common spatial confusion (mirror reversals, misremembered positions).

Quality Requirements

  • Question must be grammatically correct and unambiguous.
  • All four options must be parallel in structure and length.
  • The correct answer must not be identifiable from linguistic cues alone.
  • Spatial references must be unambiguous given the visual content.
  • Each option must begin with "A. ", "B. ", "C. ", or "D. ".

Output

Respond with ONLY a valid JSON object. No additional text.