MovideAgent is a CLI MVP for turning one sentence into a resumable short-film generation run:
plan.jsonandplan.normalized.jsonassets/reference imagesassets/*.jsonextracted asset promptsslots/*.jsonextracted slot prompts andref_imagesclips/video slotsfinal.mp4state.jsonandevents.jsonl
The default backend is mock, so the project runs without network access or API keys. Real image and planner APIs are selected with local environment variables. For long real runs, set MOVIDE_IMAGE_RETRY_ATTEMPTS=0 and MOVIDE_VIDEO_RETRY_ATTEMPTS=0 to retry recoverable image/video backend failures without an attempt limit.
The current default planner output is the compact style_anchor + assets + slots schema, inspired by movie_understand_v2.py:
{
"style_anchor": {
"prompt": "全局美术风格参考图 prompt,只描述风格、色彩、光影、材质和镜头质感,不包含具体角色/道具/场景"
},
"assets": [
{"prompt": "完整资产生图 prompt", "ref_images": ["asset_01"]}
],
"slots": [
{
"prompt": "完整 5-15 秒视频 slot prompt",
"ref_images": ["asset_01", "asset_02"]
}
]
}Prompt generation now works like this:
- The text planner receives the user query plus optional user reference images.
- The planner prompt asks for
style_anchor,assets, andslots. style_anchor.promptgenerates a global art-style reference image. All assets reference it for art style only, not for content.assets[i].promptis the exact prompt used for image generation.assets[i].ref_imagescan reference user images or previously generated assets for variants such as outfit changes, age stages, damage states, or expression/state versions of the same character.- Scene and prop assets must explicitly say that people, characters, animals, hands, body parts, faces, silhouettes, or crowds must not appear as extra visible subjects or attached objects. Scene prompts can still describe where characters may later enter, move, or leave traces in the environment.
slots[i].promptis the base prompt used to generate a per-slot storyboard grid and the final video segment.slots[i].ref_imagesnames the asset images that must be passed to storyboard and video generation.- User images are passed to asset image generation as style/identity anchors when provided.
The older characters/scenes/props + boundary_keyframes + segments plan shape is still supported as a compatibility path, but it is no longer the default prompt generation strategy.
- movide_agent_design.md is the current paper-style technical design for the
assets + slotspipeline. - docs/movide_agent_paper_zh.md is the Chinese academic-paper narrative.
- docs/movide_agent_paper_en.md is the English academic-paper narrative.
- docs/design.md is the concise implementation design note.
- docs/course_report_zh.tex is the Chinese LaTeX course project report.
python3 -m venv .venv
. .venv/bin/activate
pip install -e .For OpenAI API backends:
pip install -e '.[openai]'
cp .env.example .envFill .env locally. Do not commit .env; it is ignored by .gitignore.
Generate and execute a mock run:
movide-agent run "做一支 1 分钟的 3D 科幻喜剧短片:笨拙机器人误把生日派对变成外星人入侵现场。" --duration 60 --out runs/demo_robot_partyPass user reference images for asset style/identity anchoring:
movide-agent run "做一支 2 分钟线条小狗生日短片。" \
--duration 120 \
--reference-image https://example.com/style_anchor.png \
--out runs/demo_with_refsIf no --reference-image is provided, image reference inputs are empty for the first asset stage.
For the current assets + slots plan, the planner emits assets and slots directly:
{
"assets": [
{"prompt": "小白:白色线条小狗角色资产,全身设定图...", "ref_images": ["user_image_001"]},
{"prompt": "小白生日帽换装资产:保持小白身份一致,只增加生日帽...", "ref_images": ["asset_01"]},
{"prompt": "复古相机道具:小白最喜欢的生日礼物...", "ref_images": []}
],
"slots": [
{
"prompt": "00:00-00:08;【基础设定】小白收到相机生日礼物...",
"ref_images": ["asset_01", "asset_02"]
}
]
}Slot video prompts should be shot-level, not story summaries. A good slot includes:
characters_state: character positions, expressions, gaze, posture, held objects.scene_state: spatial layout, background, lighting, time of day.props_state: where each prop is and how it changes.action_beats: 3-5 ordered beat-by-beat actions with timing.emotion: how the emotion changes through the segment.camera: shot size, motion, framing, and screen direction.dialogue_or_voiceover: concrete dialogue, voiceover, or silence/ambient sound.audio_note: music and sound-effect timing.continuity_note: what carries over from the previous segment and what hands off to the next.ref_images: the generated asset images that must be passed to the video model for this slot.
Asset generation also supports assets[i].ref_images. A derived asset, such as the same character in a new outfit, should reference the base character asset so image generation can use the real generated image instead of relying only on text.
Generate plan only:
movide-agent plan "做一支 30 秒的温暖科幻短片:机器人第一次学会道歉。" --duration 30 --out runs/demo_plan.jsonExecute an existing plan:
movide-agent execute --plan runs/demo_plan.json --out runs/demo_from_planResume a failed or interrupted run:
movide-agent resume --out runs/demo_robot_partyText planner with OpenAI:
OPENAI_API_KEY=...
MOVIDE_PLANNER_BACKEND=openai
MOVIDE_PLANNER_MODEL=gpt-4.1
MOVIDE_PLANNER_MAX_TOKENS=32768
# Optional for models that support it:
# MOVIDE_PLANNER_REASONING_EFFORT=lowImage generation and image edit:
OPENAI_API_KEY=...
MOVIDE_IMAGE_BACKEND=openai
MOVIDE_IMAGE_MODEL=gpt-image-1
MOVIDE_IMAGE_QUALITY=high
MOVIDE_IMAGE_RETRY_ATTEMPTS=5
MOVIDE_IMAGE_RETRY_DELAY_SEC=20
MOVIDE_IMAGE_RETRY_MAX_DELAY_SEC=120
MOVIDE_IMAGE_RETRY_MODERATION_BLOCKED=falseVideo generation with OpenAI:
OPENAI_API_KEY=...
MOVIDE_VIDEO_BACKEND=openai
MOVIDE_VIDEO_MODEL=sora-2
MOVIDE_VIDEO_POLL_INTERVAL_SEC=5
MOVIDE_VIDEO_TIMEOUT_SEC=1800
MOVIDE_VIDEO_RETRY_ATTEMPTS=5
MOVIDE_VIDEO_RETRY_DELAY_SEC=30
MOVIDE_VIDEO_RETRY_MAX_DELAY_SEC=180OpenAI video calls use the reference_images list, not keyframe start/end images. For each slot, MovideAgent first generates a storyboard grid image for that segment, then passes the storyboard grid plus all referenced assets to the OpenAI video API:
image_urlisnull.last_frame_image_urlisnull.reference_images[0]is the generated line-sketch storyboard grid for the current slot.reference_images[1:]is populated fromslot.ref_images, resolved to generated asset image files or URLs.- The text prompt is prefixed with
@Image1,@Image2, ... bindings that match the order ofreference_images. @Image1is explicitly described as a line-sketch storyboard design grid for this video segment, not a start frame or end frame; the video model should use it for shot planning only, not copy its sketch style.
Example internal video call shape:
{
"prompt": "图像绑定(按传入 reference images 顺序):\n@Image1: 当前视频段应该参考的线稿分镜设计宫格图:storyboard_slot_001。它不是首帧,也不是尾帧;只参考分镜设计,不要直接使用线稿画风。\n@Image2: 当前 slot 需要保持一致的资产参考图:asset_01。\n@Image3: 当前 slot 需要保持一致的资产参考图:asset_02。\n\n...slot prompt...",
"image_url": null,
"last_frame_image_url": null,
"reference_images": ["storyboards/storyboard_slot_001.png", "assets/asset_01.png", "assets/asset_02.png"],
"ratio": "16:9",
"duration": 8
}Local reference image paths are opened as files and sent through the OpenAI SDK. Real keys stay in .env only.
- No Web UI.
- Video generation supports
mockandopenai. - If
ffmpegis installed, mock video can emit simple MP4 clips; otherwise it writes auditable placeholder.mp4files and metadata. - Resume skips calls already marked
okwith existing files.