iomgaa
f46e87258c
refactor(question_gen): extract helpers to reduce pipeline_v2 CC below grade C
...
Extract _get_git_sha, _filter_pending_slots, and _run_heavy_sampling
from run_pipeline_v2. Reduces cyclomatic complexity from C(15) to B(6).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:57:35 -04:00
iomgaa
206c553143
feat(question_gen): add v2 pipeline with retry loop and heavy check
...
- PipelineConfig: YAML-driven configuration with family_ratios, retry,
concurrency, dedup threshold, and heavy sampling rate
- _assign_slots: deterministic round-robin slot assignment across videos
with per-family weighted random selection
- _process_one_slot: full retry loop (generate → postprocess → gates →
dedup) with reject-reason feedback to VLM on retry
- _heavy_check_one: blind LLM agent trial-answer for difficulty_steps
- run_pipeline_v2: orchestration with semaphore-bounded concurrency,
progress/resume support, and store integration
- is_duplicate: cosine similarity dedup against embedding pool
Tests: 11 integration tests covering slot assignment, retry behavior,
max-retries exhaustion, full pipeline flow, progress resume, heavy
sampling, and store record completeness.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:55:15 -04:00
iomgaa
6d6eb8e3a3
refactor(question_gen): extract _validate_parsed_fields to reduce CC
...
Split field validation logic out of _parse_v2_response into a dedicated
_validate_parsed_fields helper. This brings _parse_v2_response from CC=11
(grade C) down to CC=3 (grade A). The extracted validator is CC=9 (grade B).
No grade-C functions remain.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:43:35 -04:00
iomgaa
cf51d2de9d
feat(question_gen): add v2 generator with per-family prompt templates
...
Implement generator_v2.py with:
- CandidateQuestion dataclass (canonical location)
- _load_prompt_template: loads per-family .md from store/prompts/
- _build_v2_prompt: constructs system+user messages with material context
- _parse_v2_response: JSON extraction, json_repair, field validation
- generate_one_v2: async VLM call orchestration with reject_reason support
Add 5 family-specific prompt templates:
- retrieval.md: factual recall from visible content
- reasoning.md: multi-hop inference across segments
- enumeration.md: counting/listing entities and actions
- visual.md: visual details requiring frame observation
- spatial.md: spatial relationships between objects/people
Tests: 11 unit tests covering prompt build, parse, and e2e generation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:41:36 -04:00
iomgaa
9053233f99
fix(question_gen): gates reject_reason returns raw reason; short-circuit skips LLM
...
- GateReport.reject_reason now returns gate.reason directly (no [name] prefix)
- verbatim short-circuit sets other 3 gates to SKIP without calling LLM
- test_high_verbatim_shortcircuits asserts zero LLM calls and SKIP verdicts
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:33:57 -04:00
iomgaa
9f739e831d
refactor(question_gen): reduce cyclomatic complexity in sampler_v2
...
Extract shared _resolve_subtree helper to eliminate repeated tri-level
node resolution. Break _validate_sampling_constraints into focused
single-purpose helpers:
- _count_l3_descendants
- _has_frames
- _count_subtitles
- _resolve_subtree / _find_l3_parent
Extract _subtitles_from_l2_list and _frames_from_l2_list to simplify
collection functions.
Complexity improvements:
- _validate_sampling_constraints: D(23) -> B(8)
- _collect_subtitle_sentences: C(16) -> A(3)
- _collect_frame_paths: C(13) -> A(2)
All functions now grade B or better per radon cc.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:33:48 -04:00
iomgaa
f74711cd11
feat(question_gen): add v2 material sampler with family constraints
...
Implement sample_material_v2 module that samples tree nodes with
QuestionFamilySpec-aware constraint validation, providing richer
MaterialContext output (subtitles, cross-L2 context, frame paths).
Key components:
- AnchorContext/MaterialContext frozen dataclasses
- _validate_sampling_constraints: multi-level constraint checking
- _collect_subtitle_sentences: subtree subtitle extraction
- _collect_cross_l2_context: peer L2 event descriptions
- sample_material_v2: main entry with retry-on-constraint-violation
Tests: 11 unit tests covering normal sampling, used-node exclusion,
constraint violation retries, cross-L2 population, and subtitle
collection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:30:56 -04:00
iomgaa
271d1682c9
feat(question_gen): add lightweight 4-gate quality check
...
Implement 4 concurrent LLM-based quality gates for generated questions:
- key_verify: validates answer evidence in source material
- blind_answer: rejects questions answerable without video context
- multi_true: detects ambiguous multi-correct options
- leak_test: per-family shortcut detection (5 probe templates)
Includes run_gates orchestrator with verbatim_ratio short-circuit,
JSON response parsing with fallback, and 9 unit tests (all passing).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:30:13 -04:00
iomgaa
9627ac9cf9
fix(question_gen): raise ValueError on UPDATE of missing rows
...
record_run_end, update_gates, and update_difficulty now check
cursor.rowcount after UPDATE+commit and raise ValueError if 0 rows
were affected. Prevents silent telemetry loss.
Adds three negative-path tests:
- test_record_run_end_missing_run_raises
- test_update_gates_missing_item_raises
- test_update_difficulty_missing_item_raises
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:23:48 -04:00
iomgaa
7abe92eb1c
fix(question_gen): check_verbatim covers question_text + add missing blacklist patterns
...
- check_verbatim now computes n-gram overlap for BOTH question_text and
correct_option vs source texts, returning max(question_ratio, option_ratio).
Extracted _ngram_overlap_ratio helper for reuse.
- Added 4 missing blacklist patterns: 'this segment', 'this frame',
'the current frame', 'frame summary'.
- Added 5 new test cases covering the above changes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:23:37 -04:00
iomgaa
9525726133
fix(question_gen): rename finished_at to ended_at per schema spec
...
Aligns DDL column name with research-wiki/schemas/question-gen-runs.md
which specifies 'ended_at' (not 'finished_at').
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:18:52 -04:00
iomgaa
74686dde68
refactor(question_gen): reduce check_forbidden_material complexity to A(3)
...
Extract _match_any helper and declarative _FORBIDDEN_MATERIAL_RULES table
to replace repetitive per-category for-loops. Reduces cyclomatic complexity
from C(11) to A(3).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:16:34 -04:00
iomgaa
6a5424a618
feat(question_gen): add SQLite run store for generation telemetry
...
Implements QuestionGenStore with:
- Idempotent schema initialization (question_gen_runs + question_gen_items)
- Run lifecycle: record_run_start / record_run_end / get_run_stats
- Per-item recording: record_item / update_gates / update_difficulty
- GateReportLike Protocol for duck-type gate report compatibility
- WAL mode + foreign keys + check_same_thread=False
DDL aligns with research-wiki/schemas/question-gen-{runs,items}.md.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:15:53 -04:00
iomgaa
c83d771923
feat(question_gen): add deterministic postprocess layer
...
Add app/question_gen/postprocess.py with zero-LLM deterministic
post-processing for generated questions:
- shuffle_options: deterministic option permutation with answer remapping
- check_referent_blacklist: detect self-referential language (this clip, etc.)
- check_verbatim: word-level n-gram overlap ratio measurement
- has_time_anchor: timestamp and temporal phrase detection
- check_forbidden_material: T1/T7 source material validation
- run_postprocess: orchestration returning PostprocessResult
Tests: 31 unit tests covering all functions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:14:43 -04:00
iomgaa
75e6d8c550
feat(question_gen): add 5 question family specs with sampling constraints
...
Define QuestionFamilySpec, LeakTestProfile, SamplingConstraint dataclasses
and instantiate 5 families (RETRIEVAL/REASONING/ENUMERATION/VISUAL/SPATIAL)
targeting failure mechanisms M1-M5. Implement get_family_for_slot with
legal-type filtering + weighted random selection.
13 unit tests cover: full task-type coverage, skill_target uniqueness,
deterministic seeding, invalid input errors, and chi-square distribution.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:14:22 -04:00
iomgaa
811ffa648b
feat(types): extend GeneratedQuestion with skill_target & difficulty_steps
...
- Add skill_target (str | None) and difficulty_steps (int | None) fields
to GeneratedQuestion dataclass with field(default=None)
- Update loader.py to pass new fields from JSON (backward-compatible)
- Update pools.py _q_to_dict/_dict_to_q for serialization compat
- Add question_gen_v2 config section to default.yaml
- Add comprehensive test coverage (7 tests)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:05:38 -04:00
iomgaa
5b51f4bd0c
fix(synthesizer): generate_one 捕获所有异常避免 VLM 超时穿透崩溃
...
except (ValueError, KeyError) → except Exception,
覆盖 StreamLivenessTimeout 等网络/超时异常。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-10 01:08:03 -04:00
iomgaa
d3be9b1322
refactor(tree): subtitle 迁入 L3Card/L2Card + 建树管线修正
...
- L3Card/L2Card 新增 subtitle: str 字段(L1Card 不加)
- L3Node 移除 subtitle 字段(数据迁入 Card)
- assign_subtitles_voronoi 改写 Card.subtitle + L2 聚合
- _collect_card_strings 增加 skip_fields 排除 subtitle
- _node_full_text/_node_anchored_text 保持 字幕:/[sN] 语义
- get_subtitle 读 Card.subtitle(L2/L3)
- verify.py/synthesizer.py: l3.subtitle → l3.card.subtitle
- 迁移脚本 tools/migrate_subtitle_to_card.py(幂等,300 棵树已迁移)
- 9→6 个测试文件适配(3 个无需改动)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 11:57:41 -04:00
iomgaa
f57ee45dc0
fix: Codex 全量审查修正
...
Critical:
- C1: assemble_mode 'plain' → 'ids'(合法枚举值)
- C2: question_id 加入 task_type slug 避免跨题型冲突
Important/Minor:
- generate_one 移除未用的 embed_fn/similarity_threshold 参数
- config.py 注释 11→12 同步
- 测试 question_id 断言更新
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 07:40:27 -04:00
iomgaa
fad8147d71
refactor(question_gen): __init__.py 追加 synthesizer re-export
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:49:05 -04:00
iomgaa
5aa7cc48c5
feat(question_gen): is_duplicate + generate_one — 去重判定与单题生成编排
...
- is_duplicate: 余弦相似度去重,空池短路
- generate_one: 异步重试循环,不含去重(由调用方汇总点原子执行)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:32:39 -04:00
iomgaa
90f17e330e
feat(question_gen): build_generation_prompt + parse_vlm_response
...
- prompt 组装:system(角色+题型+约束+few-shot) + user(card+字幕+干扰项)
- VLM 响应解析:JSON 直接 + markdown code block 回退,四选一 schema 校验
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:28:25 -04:00
iomgaa
40b04f886e
fix(question_gen): sample_anchor Codex 审查修正
...
- C1/C2: Temporal Reasoning ≥3 L2 + Object Reasoning ≥2 L2 下限检查
- I1: L2 题型子帧不足时 ValueError
- I2: L3 过滤无 frame_path 的节点
- I3/I4: distractor_texts 扩展到整棵树范围
- I5-I7: 测试补强 Information Synopsis/Temporal Reasoning/Object Reasoning
- M1: Spatial Reasoning 测试断言 spatial_layout
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:24:30 -04:00
iomgaa
a597a9f901
feat(question_gen): sample_anchor — 按题型层级采样锚节点
...
含 6 种层级分支:L3 单帧、L2 多帧、Temporal Perception 特例、
L1 全量/采样 L2、L1-L2 混合。时间排序 + used_node_ids 排除。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:18:25 -04:00
iomgaa
9eb9b86954
feat(question_gen): AnchorContext + 12 题型-层级映射常量
...
- AnchorContext frozen dataclass: 锚节点生成上下文(node_id, card_text, frame_paths, subtitle, distractor_texts)
- TaskTypeSpec frozen dataclass: 题型生成规格(level, needs_frames, frame_count, context_fields)
- TASK_TYPE_LEVEL_MAP: 12 种 Video-MME 题型 → 树层级 + 生成规格映射
- 11 项单元测试覆盖:映射完整性、值类型、层级合法性、frozen 不变性
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:10:58 -04:00
iomgaa
f94c352d66
feat(question_gen): QuestionGenerator Protocol + 模块公开 API
...
app/ports.py 追加 QuestionGenerator Protocol(预留 LLM 出题接口)。
app/question_gen/__init__.py re-export load_benchmark 和 stratified_sample。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 04:48:21 -04:00
iomgaa
c4d42eeca0
feat(question_gen): stratified_sample — 分层采样 + 题型保底
...
算法 100% 保真 TRM4: task_types 过滤、correctness.get(id, False) 语义、
对题在前返回顺序、min_per_class 遍历 pool 全部题型(含稀疏类)。
所有参数显式传入,无默认值。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 04:45:19 -04:00
iomgaa
8f5fbf8d2d
feat(question_gen): load_benchmark — benchmark JSON 加载
...
从 JSON 目录 glob *.json 加载题目,stem 作 video_id。
legacy schema 无 difficulty 字段时赋 _LEGACY_DEFAULT_DIFFICULTY 常量。
options/source_nodes 转 tuple 配合 frozen dataclass。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 04:41:25 -04:00
iomgaa
e60823de1a
build: add README, pyproject.toml, package skeletons and smoke test
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-06 11:37:59 -04:00