LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
Shengxiang Ji1,
Boyang Wang2,
Haiyang Xu1,
Bingnan Li1,
Yucheng Mao1,
Zeyuan Chen1,
Xiaojun Shan1,
Xiang Zhang3,
Gang Hua4,
Jianwen Xie5,
Zezhou Cheng2,
Zhuowen Tu1†
1University of California, San Diego 2University of Virginia 3Meta 4Amazon 5Lambda
†Corresponding author
teaser_video.mp4
Given a first frame, users can navigate from the first-frame view along a desired camera path and specify layouts using bounding boxes with local text prompts in the final frame. Then, LIFT generates the intended shot that transitions from the input image to the user-defined last-frame layout following the prescribed camera trajectory.
- Joint Camera and Future-Layout Control: LIFT is a unified video generation framework that enables users to control both camera motion and the semantic-spatial composition of newly revealed regions using only a last-frame layout.
- Dual-mode OPSD Training: We use a dense spatiotemporal layout teacher to train a shared student in both the single-condition mode (conditioned only on the camera trajectory) and the dual-condition mode (conditioned on both the camera trajectory and the last-frame layout).
- LIFT-Vista: We curate a dataset with large viewpoint changes and joint camera and layout annotations.
git clone https://github.com/jsxzs/LIFT.git && cd LIFT
conda create -n lift python=3.12 -y && conda activate lift
pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu126
# Install flash attention
pip install https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.48/flash_attn-2.8.3+cu126torch2.13-cp312-cp312-manylinux_2_34_aarch64.whlgit clone https://github.com/jsxzs/LIFT.git && cd LIFT
apptainer build --fakeroot LIFT.sif env.def
apptainer exec --nv --cleanenv --env PYTHONPATH=$PWD LIFT.sif| Resource | Hugging Face | Content |
|---|---|---|
| LIFT checkpoint | Download | Dense-layout teacher (stage2); and dual-mode OPSD student (stage3, our final model), one set of weights for both the camera-only and the camera + last-frame-layout mode |
| LIFT-Vista dataset | Download | 81-frame, 16 FPS clips with large viewpoint change, per-frame camera trajectories and layout annotation |
| Wan2.1 Base model | Download | VAE, T5 text encoder, CLIP image encoder, tokenizer, and the DiT that LIFT is fine-tuned from |
pip install -U "huggingface_hub[cli]"
mkdir -p models
# base model
hf download alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera --local-dir models/Wan2.1-Fun-V1.1-1.3B-Control-Camera
# LIFT checkpoints: models/LIFT/transformer (the released model) + models/LIFT_dense_layout_teacher (only for OPSD training)
hf download Overdog/LIFT --local-dir models/LIFT
# LIFT-Vista dataset
hf download Overdog/LIFT-Vista --repo-type dataset --local-dir data/LIFT-VistaWe provide one example, examples/example1/.
The example folder has the following files:
| File | Content |
|---|---|
first_frame.png |
the conditioning image |
camera.npz |
extrinsic world-to-camera matrices and intrinsic pixel-unit camera matrices at the image resolution |
layout_lastframe.json |
{"instances": [{"id", "category", "bbox": [x0,y0,x1,y1], "caption"}, ...]}, boxes in the pixel coordinates of first_frame.png |
caption.txt |
the global text prompt |
layout_dense.json |
dense per-frame layout of the same objects, "bboxes": {"<frame>": [x0,y0,x1,y1], ...} (for the dense-layout demo below) |
# give only last-frame layout
python scripts/infer.py \
--transformer_dir models/LIFT/transformer \
--image examples/example1/first_frame.png --camera examples/example1/camera.npz \
--layout examples/example1/layout_lastframe.json \
--prompt_file examples/example1/caption.txt --output outputs/example1.mp4
# give dense per-frame layout
python scripts/infer.py \
--transformer_dir models/LIFT_dense_layout_teacher/transformer \
--image examples/example1/first_frame.png --camera examples/example1/camera.npz \
--layout_dense examples/example1/layout_dense.json \
--prompt_file examples/example1/caption.txt --output outputs/example1_dense.mp4Both trainers read a CSV with one clip per row. The columns that are used:
| Column | Meaning |
|---|---|
UID |
clip id |
Video_Path |
81-frame clip (mp4), the frames the model is trained to reproduce; absolute, or relative to the CSV's directory (or to --data_root) |
resolution |
clip resolution, WxH |
caption |
global text prompt |
Annotation_Path |
directory holding the files below (absolute or relative, like Video_Path) |
CameraFile |
camera .npz inside Annotation_Path (extrinsic (81,3,4) w2c, intrinsic (81,3,3) pixels), same format as examples/example1/camera.npz |
sam3_track_file |
per-object box tracks inside Annotation_Path: {"tracks": {"<id>": {"id", "category", "caption", "bboxes": {"<frame>": [x0,y0,x1,y1], ...}}, ...}} |
sam3_track_resolution |
resolution the track boxes are expressed in, WxH |
TRAIN_CSV=/path/to/train.csv VAL_CSV=/path/to/val.csv \
bash scripts/train_dense_layout.shscripts/train_wan21_camlayout.py is a standard flow-matching fine-tune of the whole transformer.
TEACHER=models/LIFT_dense_layout_teacher/transformer/diffusion_pytorch_model.safetensors \
TRAIN_CSV=/path/to/train.csv VAL_CSV=/path/to/val.csv \
bash scripts/train_opsd.shscripts/train_wan21_camlayout_opd.py keeps the frozen dense-layout teacher and trains a student initialized from it.
The UI/ directory provides an interactive interface for running and visualizing LIFT.
ui_demo.mp4
Install the additional dependencies required by the UI:
pip install -r UI/requirements.txt
pip install git+https://github.com/microsoft/MoGe.git
pip install huggingface-hub==0.30.2Run the viewer on a GPU machine after downloading the LIFT model weights described in the Model Weights section.
# Start the UI and pre-load ../examples/example1
cd UI && CLIP=../examples/example1 ./run_viewer_gpu.sh
# Start with an empty scene and upload inputs from the browser
cd UI && CLIP=none ./run_viewer_gpu.shThe UI is available at http://localhost:8080/.
You can also pre-load a custom example directory:
cd UI
./run_viewer_gpu.sh <dir>The upload panel supports two types of input:
-
Single image
Upload an image directly in the browser. MoGe automatically estimates the point cloud and the first-frame camera.
You can then:
- specify the camera trajectory;
- preview the corresponding keyframe views; and
- draw bounding boxes for the target last-frame layout.
-
Example folder
Upload a complete example folder to directly load the camera trajectory and last-frame layout. In this case, the keyframe views and layout boxes are displayed automatically, without requiring manual drawing.
A folder should follow the
examples/example1format and contain:<dir>/ ├── first_frame.png ├── camera.npz ├── caption.txt └── layout_lastframe.jsonNo point cloud file is needed: MoGe estimates it from the image, both for a browser upload and for a folder pre-loaded from the command line (
./run_viewer_gpu.sh <dir>). A pre-loaded folder may ship a precomputedpcd_moge.npz(asexamples/example1does) to skip that step; it can be made withpython UI/viewer/estimate_pcd.py --clip <dir> --use-clip-intrinsic.
All files read or generated by the editor are stored in a per-run session directory under:
UI/.ui_sessions/
A session may contain:
camera_da3_edited.npz
layout_edited.json
generated/
as well as copies of the pre-loaded inputs and browser uploads.
The original source directory is never modified.
Click Reset to clear the current scene and start again with another image or example folder.
This code base builds on VideoX-Fun and Wan2.1.
Released under the Apache License 2.0 (see LICENSE). The Wan2.1-Fun base model is subject to its own license.
@article{ji2026lift,
title={LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation},
author={Ji, Shengxiang and Wang, Boyang and Xu, Haiyang and Li, Bingnan and Mao, Yucheng and Chen, Zeyuan and Shan, Xiaojun and Zhang, Xiang and Hua, Gang and Xie, Jianwen and Cheng, Zezhou and Tu, Zhuowen},
journal={arXiv preprint arXiv:2609.38146},
year={2026}
}