具有多种交互的无限世界
✨ For more high-fidelity and compelling demos, please visit our Project Page.
## News - Sep. 10, 2026: We release the remaining full model variants: the 14B model’s causal-pretrained and bidirectional variants, and the 1.3B model’s causal-fast variant. - Jul. 9, 2026: We release the technical report, inference code, and models for LingBot-World-Infinity. ## TODO - [x] Release the causal-fast inference code and model of the 14B model - [x] Release the causal-pretrained model of the 14B model - [x] Release the bidirectional model of the 14B model - [x] Release the causal-fast model of the 1.3B model ## ⚙️ Quick Start This codebase is built upon [Wan2.2](https://github.com/Wan-Video/Wan2.2). Please refer to their documentation for installation instructions. ### Installation Clone the repo: ```sh git clone https://github.com/robbyant/lingbot-world-v2.git cd lingbot-world-v2 ``` Install dependencies: ```sh # Ensure torch >= 2.4.0 pip install -r requirements.txt ``` Install [`flash_attn`](https://github.com/Dao-AILab/flash-attention): ```sh pip install flash-attn --no-build-isolation ``` ### Model Download | Model | Model Type | Model Size | Download Links | | :--- | :--- | :--- | :--- | | **lingbot-world-v2-14b-causal-fast** | causal-fast | 14B | [HuggingFace](https://huggingface.co/robbyant/lingbot-world-v2-14b-causal-fast) [ModelScope](https://www.modelscope.cn/models/Robbyant/lingbot-world-v2-14b-causal-fast) | | **lingbot-world-v2-14b-causal-pretrain** | causal-pretrain | 14B | [HuggingFace](https://huggingface.co/robbyant/lingbot-world-v2-14b-causal-pretrain) | | **lingbot-world-v2-14b-bid** | bidirectional | 14B | [HuggingFace](https://huggingface.co/robbyant/lingbot-world-v2-14b-bid) | | **lingbot-world-v2-1.3b-causal-fast** | causal-fast | 1.3B | [HuggingFace](https://huggingface.co/robbyant/lingbot-world-v2-1.3b-causal-fast) | Download models using huggingface-cli: ```sh pip install "huggingface_hub[cli]" huggingface-cli download robbyant/lingbot-world-v2-14b-causal-fast --local-dir ./lingbot-world-v2-14b-causal-fast huggingface-cli download robbyant/lingbot-world-v2-1.3b-causal-fast --local-dir ./lingbot-world-v2-1.3b-causal-fast/transformers ``` Download models using modelscope-cli: ```sh pip install modelscope modelscope download robbyant/lingbot-world-v2-14b-causal-fast --local_dir ./lingbot-world-v2-14b-causal-fast ``` The 1.3B Hugging Face package currently contains the DiT weights only. T5, VAE, and the tokenizer are shared with the 14B release — pass them with `--assets_dir` (or the third argument of `run_fast.sh`): ### Inference We provide `generate.py` for causal inference with KV caching, which processes video frames chunk-by-chunk instead of all at once. - `causal_fast` 14B — 480P, 8 GPUs (`ulysses_size` must divide 40 heads): ``` sh torchrun --nproc_per_node=8 generate.py --task i2v-A14B --size 480*832 --ckpt_dir lingbot-world-v2-14b-causal-fast --image examples/03/image.jpg --action_path examples/03 --dit_fsdp --t5_fsdp --ulysses_size 8 --frame_num 361 --local_attn_size 18 --sink_size 6 --prompt "A serene lakeside scene with a lone tree standing in calm water, surrounded by distant snow-capped mountains under a bright blue sky with drifting white clouds — gentle ripples reflect the tree and sky, creating a tranquil, meditative atmosphere." ``` - `causal_fast` 1.3B — 480P, 4 GPUs (`ulysses_size` must divide 12 heads). Reuse T5/VAE from the 14B checkpoint if the 1.3B folder does not include them: ``` sh torchrun --nproc_per_node=4 generate.py --task i2v-1.3B --size 480*832 --ckpt_dir lingbot-world-v2-1.3b-causal-fast --assets_dir lingbot-world-v2-14b-causal-fast --image examples/03/image.jpg --action_path examples/03 --dit_fsdp --t5_fsdp --ulysses_size 4 --frame_num 361 --local_attn_size 18 --sink_size 6 --prompt "A serene lakeside scene with a lone tree standing in calm water, surrounded by distant snow-capped mountains under a bright blue sky with drifting white clouds — gentle ripples reflect the tree and sky, creating a tranquil, meditative atmosphere." ``` - `causal_pretrain` — 480P, multi-GPU: ``` sh torchrun --nproc_per_node=8 generate.py --task i2v-A14B --infer_mode causal_pretrain --size 480*832 --ckpt_dir lingbot-world-v2-14b-causal-pretrain --image examples/03/image.jpg --action_path examples/03 --dit_fsdp --t5_fsdp --ulysses_size 8 --frame_num 81 --prompt "A serene lakeside scene with a lone tree standing in calm water, surrounded by distant snow-capped mountains under a bright blue sky with drifting white clouds — gentle ripples reflect the tree and sky, creating a tranquil, meditative atmosphere." ``` You can also use the provided `run_fast.sh` script. The task and GPU count are inferred from the checkpoint directory name (`*1.3b*` / `*1p3b*` → 1.3B on 4 GPUs, otherwise 14B on 8 GPUs): ``` sh bash run_fast.sh [assets_dir] # e.g. bash run_fast.sh lingbot-world-v2-14b-causal-fast 361 # e.g. bash run_fast.sh lingbot-world-v2-1.3b-causal-fast 361 lingbot-world-v2-14b-causal-fast ``` ### Deployment We do NOT plan to release our deployment code. If you would like to deploy our model yourself, please refer to the LingBot-World deployment in [SGLang](https://docs.sglang.io/cookbook/diffusion/LingBot-World/LingBot-World-2.0) or [flashdreams](https://github.com/NVIDIA/flashdreams). ## Related Projects - [LingBot-World](https://github.com/robbyant/lingbot-world) ## License This project is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). The project is available for non-commercial use only: you may share and adapt it with proper attribution, but derivative works must be distributed under the same license. Please refer to the [LICENSE file](LICENSE.txt) for the full text, including details on rights and restrictions. ## ✨ Acknowledgement We would like to express our gratitude to the Wan Team for open-sourcing their code and models. Their contributions have been instrumental to the development of this project. ## Citation If you find this work useful for your research, please cite our paper: ``` @article{lingbot-world-v2, title={Infinite Worlds with Versatile Interactions}, author={Zelin Gao and Qiuyu Wang and Jiapeng Zhu and Jingye Chen and Zichen Liu and Qingyan Bai and Jiahao Wang and Yufeng Yuan and Hanlin Wang and Yichong Lu and Ka Leong Cheng and Haojie Zhang and Jian Gao and Tianrui Feng and Yuzheng Liu and Yao Yao and Yinghao Xu and Xing Zhu and Yujun Shen and Hao Ouyang}, journal={arXiv preprint arXiv:2607.07534}, year={2026} } ```Released inference code caps out at ~4m16s due to RoPE table size — how was the "over one hour, no perceptible decay" result in the paper achieved?
would a small external action-trajectory sample be useful
What GPU configuration is required for inference?
Does it support generation with multiple prompts?
Any plan to release Agentic Harness code?
Black Screen issue
Release remaining LingBot-World-v2 models on Hugging Face