Add a lightweight semantic pre-validation stage for generated video clips before invoking a heavier VLM.
Goal
Use a small video-text model (ViCLIP-L/14 or equivalent compatible implementation) to score semantic alignment between shot_spec text and the generated clip.
Target flow:
shot_spec -> generated clip -> lightweight semantic score -> PASS | ESCALATE -> heavy VLM/temporal validator
Requirements
- provider interface; no hard dependency on one checkpoint;
- cache embeddings/results by content hash;
- configurable confidence thresholds;
- retain evidence/run_id/provenance;
- never auto-approve a clip solely because semantic similarity is high if structural/temporal validators fail;
- escalation to LFM/other heavy VLM on low confidence, mismatch, multi-subject ambiguity or temporal defect signals.
Benchmark
Use a fixed set of existing StoryCore clips including:
- clearly correct shot;
- wrong/missing object;
- identity drift;
- camera mismatch;
- temporally unstable clip.
Measure:
- pre-validator latency;
- GPU/RAM/VRAM use;
- false PASS rate;
- false ESCALATE rate;
- percentage of heavy-VLM calls avoided;
- end-to-end render validation time.
Exit criteria
Demonstrate that the lightweight stage safely removes a meaningful fraction of unnecessary heavy-VLM calls while preserving validation quality on the benchmark.
Add a lightweight semantic pre-validation stage for generated video clips before invoking a heavier VLM.
Goal
Use a small video-text model (ViCLIP-L/14 or equivalent compatible implementation) to score semantic alignment between
shot_spectext and the generated clip.Target flow:
shot_spec -> generated clip -> lightweight semantic score -> PASS | ESCALATE -> heavy VLM/temporal validatorRequirements
Benchmark
Use a fixed set of existing StoryCore clips including:
Measure:
Exit criteria
Demonstrate that the lightweight stage safely removes a meaningful fraction of unnecessary heavy-VLM calls while preserving validation quality on the benchmark.