Skip to content

P1 ViCLIP semantic pre-validator before heavy VLM escalation #41

Description

@zedarvates

Add a lightweight semantic pre-validation stage for generated video clips before invoking a heavier VLM.

Goal

Use a small video-text model (ViCLIP-L/14 or equivalent compatible implementation) to score semantic alignment between shot_spec text and the generated clip.

Target flow:
shot_spec -> generated clip -> lightweight semantic score -> PASS | ESCALATE -> heavy VLM/temporal validator

Requirements

  • provider interface; no hard dependency on one checkpoint;
  • cache embeddings/results by content hash;
  • configurable confidence thresholds;
  • retain evidence/run_id/provenance;
  • never auto-approve a clip solely because semantic similarity is high if structural/temporal validators fail;
  • escalation to LFM/other heavy VLM on low confidence, mismatch, multi-subject ambiguity or temporal defect signals.

Benchmark

Use a fixed set of existing StoryCore clips including:

  • clearly correct shot;
  • wrong/missing object;
  • identity drift;
  • camera mismatch;
  • temporally unstable clip.

Measure:

  • pre-validator latency;
  • GPU/RAM/VRAM use;
  • false PASS rate;
  • false ESCALATE rate;
  • percentage of heavy-VLM calls avoided;
  • end-to-end render validation time.

Exit criteria

Demonstrate that the lightweight stage safely removes a meaningful fraction of unnecessary heavy-VLM calls while preserving validation quality on the benchmark.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions