Skip to content

Repository files navigation

MovideAgent MVP

MovideAgent is a CLI MVP for turning one sentence into a resumable short-film generation run:

  • plan.json and plan.normalized.json
  • assets/ reference images
  • assets/*.json extracted asset prompts
  • slots/*.json extracted slot prompts and ref_images
  • clips/ video slots
  • final.mp4
  • state.json and events.jsonl

The default backend is mock, so the project runs without network access or API keys. Real image and planner APIs are selected with local environment variables. For long real runs, set MOVIDE_IMAGE_RETRY_ATTEMPTS=0 and MOVIDE_VIDEO_RETRY_ATTEMPTS=0 to retry recoverable image/video backend failures without an attempt limit.

Current Pipeline

The current default planner output is the compact style_anchor + assets + slots schema, inspired by movie_understand_v2.py:

{
  "style_anchor": {
    "prompt": "全局美术风格参考图 prompt,只描述风格、色彩、光影、材质和镜头质感,不包含具体角色/道具/场景"
  },
  "assets": [
    {"prompt": "完整资产生图 prompt", "ref_images": ["asset_01"]}
  ],
  "slots": [
    {
      "prompt": "完整 5-15 秒视频 slot prompt",
      "ref_images": ["asset_01", "asset_02"]
    }
  ]
}

Prompt generation now works like this:

  • The text planner receives the user query plus optional user reference images.
  • The planner prompt asks for style_anchor, assets, and slots.
  • style_anchor.prompt generates a global art-style reference image. All assets reference it for art style only, not for content.
  • assets[i].prompt is the exact prompt used for image generation.
  • assets[i].ref_images can reference user images or previously generated assets for variants such as outfit changes, age stages, damage states, or expression/state versions of the same character.
  • Scene and prop assets must explicitly say that people, characters, animals, hands, body parts, faces, silhouettes, or crowds must not appear as extra visible subjects or attached objects. Scene prompts can still describe where characters may later enter, move, or leave traces in the environment.
  • slots[i].prompt is the base prompt used to generate a per-slot storyboard grid and the final video segment.
  • slots[i].ref_images names the asset images that must be passed to storyboard and video generation.
  • User images are passed to asset image generation as style/identity anchors when provided.

The older characters/scenes/props + boundary_keyframes + segments plan shape is still supported as a compatibility path, but it is no longer the default prompt generation strategy.

Documentation

Setup

python3 -m venv .venv
. .venv/bin/activate
pip install -e .

For OpenAI API backends:

pip install -e '.[openai]'
cp .env.example .env

Fill .env locally. Do not commit .env; it is ignored by .gitignore.

Commands

Generate and execute a mock run:

movide-agent run "做一支 1 分钟的 3D 科幻喜剧短片:笨拙机器人误把生日派对变成外星人入侵现场。" --duration 60 --out runs/demo_robot_party

Pass user reference images for asset style/identity anchoring:

movide-agent run "做一支 2 分钟线条小狗生日短片。" \
  --duration 120 \
  --reference-image https://example.com/style_anchor.png \
  --out runs/demo_with_refs

If no --reference-image is provided, image reference inputs are empty for the first asset stage.

For the current assets + slots plan, the planner emits assets and slots directly:

{
  "assets": [
    {"prompt": "小白:白色线条小狗角色资产,全身设定图...", "ref_images": ["user_image_001"]},
    {"prompt": "小白生日帽换装资产:保持小白身份一致,只增加生日帽...", "ref_images": ["asset_01"]},
    {"prompt": "复古相机道具:小白最喜欢的生日礼物...", "ref_images": []}
  ],
  "slots": [
    {
      "prompt": "00:00-00:08;【基础设定】小白收到相机生日礼物...",
      "ref_images": ["asset_01", "asset_02"]
    }
  ]
}

Slot video prompts should be shot-level, not story summaries. A good slot includes:

  • characters_state: character positions, expressions, gaze, posture, held objects.
  • scene_state: spatial layout, background, lighting, time of day.
  • props_state: where each prop is and how it changes.
  • action_beats: 3-5 ordered beat-by-beat actions with timing.
  • emotion: how the emotion changes through the segment.
  • camera: shot size, motion, framing, and screen direction.
  • dialogue_or_voiceover: concrete dialogue, voiceover, or silence/ambient sound.
  • audio_note: music and sound-effect timing.
  • continuity_note: what carries over from the previous segment and what hands off to the next.
  • ref_images: the generated asset images that must be passed to the video model for this slot.

Asset generation also supports assets[i].ref_images. A derived asset, such as the same character in a new outfit, should reference the base character asset so image generation can use the real generated image instead of relying only on text.

Generate plan only:

movide-agent plan "做一支 30 秒的温暖科幻短片:机器人第一次学会道歉。" --duration 30 --out runs/demo_plan.json

Execute an existing plan:

movide-agent execute --plan runs/demo_plan.json --out runs/demo_from_plan

Resume a failed or interrupted run:

movide-agent resume --out runs/demo_robot_party

API Backends

Text planner with OpenAI:

OPENAI_API_KEY=...
MOVIDE_PLANNER_BACKEND=openai
MOVIDE_PLANNER_MODEL=gpt-4.1
MOVIDE_PLANNER_MAX_TOKENS=32768
# Optional for models that support it:
# MOVIDE_PLANNER_REASONING_EFFORT=low

Image generation and image edit:

OPENAI_API_KEY=...
MOVIDE_IMAGE_BACKEND=openai
MOVIDE_IMAGE_MODEL=gpt-image-1
MOVIDE_IMAGE_QUALITY=high
MOVIDE_IMAGE_RETRY_ATTEMPTS=5
MOVIDE_IMAGE_RETRY_DELAY_SEC=20
MOVIDE_IMAGE_RETRY_MAX_DELAY_SEC=120
MOVIDE_IMAGE_RETRY_MODERATION_BLOCKED=false

Video generation with OpenAI:

OPENAI_API_KEY=...
MOVIDE_VIDEO_BACKEND=openai
MOVIDE_VIDEO_MODEL=sora-2
MOVIDE_VIDEO_POLL_INTERVAL_SEC=5
MOVIDE_VIDEO_TIMEOUT_SEC=1800
MOVIDE_VIDEO_RETRY_ATTEMPTS=5
MOVIDE_VIDEO_RETRY_DELAY_SEC=30
MOVIDE_VIDEO_RETRY_MAX_DELAY_SEC=180

OpenAI video calls use the reference_images list, not keyframe start/end images. For each slot, MovideAgent first generates a storyboard grid image for that segment, then passes the storyboard grid plus all referenced assets to the OpenAI video API:

  • image_url is null.
  • last_frame_image_url is null.
  • reference_images[0] is the generated line-sketch storyboard grid for the current slot.
  • reference_images[1:] is populated from slot.ref_images, resolved to generated asset image files or URLs.
  • The text prompt is prefixed with @Image1, @Image2, ... bindings that match the order of reference_images.
  • @Image1 is explicitly described as a line-sketch storyboard design grid for this video segment, not a start frame or end frame; the video model should use it for shot planning only, not copy its sketch style.

Example internal video call shape:

{
  "prompt": "图像绑定(按传入 reference images 顺序):\n@Image1: 当前视频段应该参考的线稿分镜设计宫格图:storyboard_slot_001。它不是首帧,也不是尾帧;只参考分镜设计,不要直接使用线稿画风。\n@Image2: 当前 slot 需要保持一致的资产参考图:asset_01。\n@Image3: 当前 slot 需要保持一致的资产参考图:asset_02。\n\n...slot prompt...",
  "image_url": null,
  "last_frame_image_url": null,
  "reference_images": ["storyboards/storyboard_slot_001.png", "assets/asset_01.png", "assets/asset_02.png"],
  "ratio": "16:9",
  "duration": 8
}

Local reference image paths are opened as files and sent through the OpenAI SDK. Real keys stay in .env only.

Current Limits

  • No Web UI.
  • Video generation supports mock and openai.
  • If ffmpeg is installed, mock video can emit simple MP4 clips; otherwise it writes auditable placeholder .mp4 files and metadata.
  • Resume skips calls already marked ok with existing files.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages