Project webpage, NeurIPS version, Arxiv
Official implementation of My work PixFoundation 2.0 in NeurIPS 2026 Evaluations and Datasets Track.
- Clone the repository recursively to include the submodules
git clone --recursive https://github.com/MSiam/PixFoundation-2.0/
- Create conda environment
conda create --name pixfoundation2 python=3.10
conda activate pixfoundation2
- Install the base requirements for evaluations
pip install -r requirements.txt
- Setup detectron2 for some utilities used (Optional)
git clone https://github.com/facebookresearch/detectron2.git
python -m pip install -e detectron2
- Follow installation setup for each model you are evaluating, refer to their README and if necessary create its separate conda env.
- Refer to scripts/run_X.sh for each model X script and modify the conda environment if needed or use the common pixfoundation2
- Download MeVIS.
- Generate the multi-video-layout. Layouts l2, l3
python data/create_benchmark_mevis.py --config-file configs/mevis.yaml --dataset_root DATA_ROOT --frames_sel_file data/mevis_keyframes.csv --output_dir OUT_DIR
python data/create_benchmark_mevis.py --config-file configs/mevis.yaml --dataset_root DATA_ROOT --frames_sel_file data/mevis_keyframes.csv --output_dir OUT_DIR --left_flag
python data/create_benchmark_mevis.py --config-file configs/mevis.yaml --dataset_root DATA_ROOT --output_dir OUT_DIR --reverse_flag
python data/create_benchmark_mevis.py --config-file configs/mevis.yaml --dataset_root DATA_ROOT --output_dir OUT_DIR --reverse_flag --left_flag
-
Move our reverse expressions under data/meta_expressions_reverse_filtered.json to the valid_u subset path.
-
The final MeVIS directory is as follows:
|--- MeVIS
|--- train
|--- JPEGImages
|--- mask_dict.json
|--- meta_expressions.json
|--- valid_u
|--- JPEGImages
|--- mask_dict.json
|--- meta_expressions.json
|--- meta_expressions_reverse_filtered.json
|--- valid_u_mocentric_tile_single
|--- JPEGImages
|--- valid_u_mocentric_tile_single_left
|--- JPEGImages
|--- valid_u_mocentric_tile_reverse
|--- JPEGImages
|--- valid_u_mocentric_tile_reverse_left
|--- JPEGImages
- Visualize the synthetic dataset with the segmentation masks
python datasets_/test_loaders.py --config-file configs/mevis.yaml --dataset_root DATA_ROOT --dataset_split mevis_val_mocentric_tile_single --dataset_mask_path DATA_ROOT/valid_u/mask_dict.json --dataset_exp_path DATA_ROOT/valid_u/EXPRS_JSON --out_dir OUT_DIR --save_vis
- You can follow similar procedure to Molmo2Track.
- Run common bash script to run the benchmarking after modifying the paths
cd scripts
bash run_all.sh
The motion centric adaptation can be applied to any MLLM independant of the architecture relying on the synthetic training data. I provide an example with Sa2VA. The forked Sa2VA modified for motion-centric adaptation is provided here.
- Train Sa2VA and convert it to Hugging Face Checkpoint
cd scripts
bash lora_tune_sa2va.sh
-
This is the set of weights that I used in my experiments for the motion-centric Sa2VA. Upon merging with original ByteDance weights you can run the evaluation.
-
Replace Sa2VA_CKPT in run_all.sh with the correct full path to hugging face checkpoint for evaluation.
-
Modifications for the LoRA setup including LLM instead of ViT and other hyperparameters can be added to 'Sa2VA/projects/llava_sam2/configs/sa2va_8b_motion.py'.
- I thank Sa2VA ByteDance authors as I built the experiments for motion-centric adaptation upon their codebase.
Please cite my paper if you find it useful in your research
@article{siam2026pixfoundation2.0,
title={PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?},
author={Siam, Mennatullah},
journal={NeurIPS},
year={2026}
}

