Skip to content

Latest commit

 

History

19 Commits

Folders and files

Repository files navigation

InstanceAnimator: Multi-Instance Sketch Video Colorization

   🖥️ GitHub    |   🤖 Model   |    📑 Paper   
📊 Dataset    | 🫰 Website

Overview

We propose InstanceAnimator, a novel Diffusion Transformer framework for multi-instance sketch video colorization. Existing animation colorization methods rely heavily on a single initial reference frame, resulting in fragmented workflows and limited customizability. To eliminate these constraints, we introduce a Canvas Guidance Condition that allows users to freely place reference elements on a blank canvas, enabling flexible user control. To address the misalignment and quality degradation issues of DiT-based approaches, we design an Instance Matching Mechanism that integrates the instances with the sketch and noise channels, ensuring visual consistency across different sequences while maintaining controllability. Additionally, to mitigate the degradation of fine-grained details, we propose an Adaptive Decoupled Control Module that injects semantic features from characters, backgrounds, and text conditions into the diffusion model, significantly enhancing detail fidelity.

Framework

framework

Application

We support a dynamic background rendering generation in some cases during colorization by visual background guidance.

bg.mp4

Set up

Environment

conda create -n InstanceAnimator python=3.12

pip install -r requirements.txt

Repository

git clone https://github.com/YinHan-Zhang/InstanceAnimator.git
    
cd InstanceAnimator

Model Weight

modelscope download PAI/Wan2.1-Fun-14B-Control --local_dir ./Wan2.1-Fun-14B-Control

# so sorry that model weight has been loss, you need to retrain model.
# modelscope download NiceYinHan/InstanceAnimator --local_dir ./ckpt

OpenAnimate Dataset

We fully open-source our training dataset. (Due to the legal copyright issues of the animation dataset, the download of the dataset is restricted. We apologize for any inconvenience.)

modelscope login

modelscope download --dataset NiceYinHan/OpenAnimate --local_dir ./OpenAnimate

Train

Dataset Format:

{
   "file_path": "video.mp4",
    "sketch_file_path": "sketch.mp4",
    "control_file_path": [
        "instance_1.jpg",
        ...
    ],
    "background_path": "background.jpg",
    "text": "",
    "type": "video"
}

After finishing data preparation, you can launch training ...

bash training/train_control_lora.sh

Inference

Modify model path in predict_video_decouple.py,

python  inference/predict_video_decouple.py

Web UI

python interface/web_app.py                # real inference
python interface/web_app.py --mock         # UI test without model
python interface/web_app.py --port 7860 --model_name /path/to/Wan2.1-Fun-14B-Control

Limitation

Due to the limitations of computing resources and data, the amount of data for model training is limited. If you have enough computing resources, you can train the model yourself.

Acknowledgement

Thanks for the reference contributions of these works: - VideoX-Fun

About

The official implementation of "InstanceAnimator: Multi-Instance Sketch Video Colorization"

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages