Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Shengxiang Ji1, Boyang Wang2, Haiyang Xu1, Bingnan Li1, Yucheng Mao1, Zeyuan Chen1,
Xiaojun Shan1, Xiang Zhang3, Gang Hua4, Jianwen Xie5, Zezhou Cheng2, Zhuowen Tu1†

1University of California, San Diego   2University of Virginia   3Meta   4Amazon   5Lambda
†Corresponding author

Project Page arXiv Hugging Face Dataset Hugging Face Model

teaser_video.mp4

LIFT teaser

Given a first frame, users can navigate from the first-frame view along a desired camera path and specify layouts using bounding boxes with local text prompts in the final frame. Then, LIFT generates the intended shot that transitions from the input image to the user-defined last-frame layout following the prescribed camera trajectory.

📌 TL;DR

  • Joint Camera and Future-Layout Control: LIFT is a unified video generation framework that enables users to control both camera motion and the semantic-spatial composition of newly revealed regions using only a last-frame layout.
  • Dual-mode OPSD Training: We use a dense spatiotemporal layout teacher to train a shared student in both the single-condition mode (conditioned only on the camera trajectory) and the dual-condition mode (conditioned on both the camera trajectory and the last-frame layout).
  • LIFT-Vista: We curate a dataset with large viewpoint changes and joint camera and layout annotations.

🛠️ Installation

Option A: Conda Environment

git clone https://github.com/jsxzs/LIFT.git && cd LIFT
conda create -n lift python=3.12 -y && conda activate lift
pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu126

# Install flash attention
pip install https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.48/flash_attn-2.8.3+cu126torch2.13-cp312-cp312-manylinux_2_34_aarch64.whl

Option B: Apptainer Image

git clone https://github.com/jsxzs/LIFT.git && cd LIFT
apptainer build --fakeroot LIFT.sif env.def
apptainer exec --nv --cleanenv --env PYTHONPATH=$PWD LIFT.sif

📦 Model Weights and Dataset

Resource Hugging Face Content
LIFT checkpoint Download Dense-layout teacher (stage2); and dual-mode OPSD student (stage3, our final model), one set of weights for both the camera-only and the camera + last-frame-layout mode
LIFT-Vista dataset Download 81-frame, 16 FPS clips with large viewpoint change, per-frame camera trajectories and layout annotation
Wan2.1 Base model Download VAE, T5 text encoder, CLIP image encoder, tokenizer, and the DiT that LIFT is fine-tuned from
pip install -U "huggingface_hub[cli]"
mkdir -p models
# base model
hf download alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera --local-dir models/Wan2.1-Fun-V1.1-1.3B-Control-Camera
# LIFT checkpoints: models/LIFT/transformer (the released model) + models/LIFT_dense_layout_teacher (only for OPSD training)
hf download Overdog/LIFT --local-dir models/LIFT
# LIFT-Vista dataset
hf download Overdog/LIFT-Vista --repo-type dataset --local-dir data/LIFT-Vista

🎬 Inference

We provide one example, examples/example1/. The example folder has the following files:

File Content
first_frame.png the conditioning image
camera.npz extrinsic world-to-camera matrices and intrinsic pixel-unit camera matrices at the image resolution
layout_lastframe.json {"instances": [{"id", "category", "bbox": [x0,y0,x1,y1], "caption"}, ...]}, boxes in the pixel coordinates of first_frame.png
caption.txt the global text prompt
layout_dense.json dense per-frame layout of the same objects, "bboxes": {"<frame>": [x0,y0,x1,y1], ...} (for the dense-layout demo below)
# give only last-frame layout
python scripts/infer.py \
  --transformer_dir models/LIFT/transformer \
  --image examples/example1/first_frame.png --camera examples/example1/camera.npz \
  --layout examples/example1/layout_lastframe.json \
  --prompt_file examples/example1/caption.txt --output outputs/example1.mp4

# give dense per-frame layout
python scripts/infer.py \
  --transformer_dir models/LIFT_dense_layout_teacher/transformer \
  --image examples/example1/first_frame.png --camera examples/example1/camera.npz \
  --layout_dense examples/example1/layout_dense.json \
  --prompt_file examples/example1/caption.txt --output outputs/example1_dense.mp4

🏋️ Training

Data Format

Both trainers read a CSV with one clip per row. The columns that are used:

Column Meaning
UID clip id
Video_Path 81-frame clip (mp4), the frames the model is trained to reproduce; absolute, or relative to the CSV's directory (or to --data_root)
resolution clip resolution, WxH
caption global text prompt
Annotation_Path directory holding the files below (absolute or relative, like Video_Path)
CameraFile camera .npz inside Annotation_Path (extrinsic (81,3,4) w2c, intrinsic (81,3,3) pixels), same format as examples/example1/camera.npz
sam3_track_file per-object box tracks inside Annotation_Path: {"tracks": {"<id>": {"id", "category", "caption", "bboxes": {"<frame>": [x0,y0,x1,y1], ...}}, ...}}
sam3_track_resolution resolution the track boxes are expressed in, WxH

Dense Layout SFT

TRAIN_CSV=/path/to/train.csv VAL_CSV=/path/to/val.csv \
bash scripts/train_dense_layout.sh

scripts/train_wan21_camlayout.py is a standard flow-matching fine-tune of the whole transformer.

Dual-Mode On-Policy Self-Distillation (OPSD)

TEACHER=models/LIFT_dense_layout_teacher/transformer/diffusion_pytorch_model.safetensors \
TRAIN_CSV=/path/to/train.csv VAL_CSV=/path/to/val.csv \
bash scripts/train_opsd.sh

scripts/train_wan21_camlayout_opd.py keeps the frozen dense-layout teacher and trains a student initialized from it.

🖥️ UI Demo

The UI/ directory provides an interactive interface for running and visualizing LIFT.

ui_demo.mp4

Installation

Install the additional dependencies required by the UI:

pip install -r UI/requirements.txt
pip install git+https://github.com/microsoft/MoGe.git
pip install huggingface-hub==0.30.2

Run the UI

Run the viewer on a GPU machine after downloading the LIFT model weights described in the Model Weights section.

# Start the UI and pre-load ../examples/example1
cd UI && CLIP=../examples/example1 ./run_viewer_gpu.sh

# Start with an empty scene and upload inputs from the browser
cd UI && CLIP=none ./run_viewer_gpu.sh

The UI is available at http://localhost:8080/.

You can also pre-load a custom example directory:

cd UI
./run_viewer_gpu.sh <dir>

Input Options

The upload panel supports two types of input:

  1. Single image

    Upload an image directly in the browser. MoGe automatically estimates the point cloud and the first-frame camera.

    You can then:

    • specify the camera trajectory;
    • preview the corresponding keyframe views; and
    • draw bounding boxes for the target last-frame layout.
  2. Example folder

    Upload a complete example folder to directly load the camera trajectory and last-frame layout. In this case, the keyframe views and layout boxes are displayed automatically, without requiring manual drawing.

    A folder should follow the examples/example1 format and contain:

    <dir>/
    ├── first_frame.png
    ├── camera.npz
    ├── caption.txt
    └── layout_lastframe.json
    

    No point cloud file is needed: MoGe estimates it from the image, both for a browser upload and for a folder pre-loaded from the command line (./run_viewer_gpu.sh <dir>). A pre-loaded folder may ship a precomputed pcd_moge.npz (as examples/example1 does) to skip that step; it can be made with python UI/viewer/estimate_pcd.py --clip <dir> --use-clip-intrinsic.

Session Files

All files read or generated by the editor are stored in a per-run session directory under:

UI/.ui_sessions/

A session may contain:

camera_da3_edited.npz
layout_edited.json
generated/

as well as copies of the pre-loaded inputs and browser uploads.

The original source directory is never modified.

Reset

Click Reset to clear the current scene and start again with another image or example folder.

🙏 Acknowledgements

This code base builds on VideoX-Fun and Wan2.1.

📄 License

Released under the Apache License 2.0 (see LICENSE). The Wan2.1-Fun base model is subject to its own license.

📚 Citation

@article{ji2026lift,
  title={LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation},
  author={Ji, Shengxiang and Wang, Boyang and Xu, Haiyang and Li, Bingnan and Mao, Yucheng and Chen, Zeyuan and Shan, Xiaojun and Zhang, Xiang and Hua, Gang and Xie, Jianwen and Cheng, Zezhou and Tu, Zhuowen},
  journal={arXiv preprint arXiv:2609.38146},
  year={2026}
}

About

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages