Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
A pytorch implementation of the text-to-3D model Dreamfusion, powered by the Stable Diffusion text-to-2D model.
ADVERTISEMENT: Please check out threestudio for recent improvements and better implementation in 3D content generation!
NEWS (2023.6.12):
This project is a work-in-progress, and contains lots of differences from the paper. The current generation quality cannot match the results from the original paper, and many prompts still fail badly!
git clone https://github.com/ashawkey/stable-dreamfusion.git
cd stable-dreamfusion
To avoid python package conflicts, we recommend using a virtual environment, e.g.: using conda or venv:
python -m venv venv_stable-dreamfusion
source venv_stable-dreamfusion/bin/activate # you need to repeat this step for every new terminal
pip install -r requirements.txt
To use image-conditioned 3D generation, you need to download some pretrained checkpoints manually:
zero123-xl.ckpt by default, and it is hard-coded in guidance/zero123_utils.py.cd pretrained/zero123
wget https://zero123.cs.columbia.edu/assets/zero123-xl.ckpt
preprocess_image.py.mkdir pretrained/omnidata
cd pretrained/omnidata
# assume gdown is installed
gdown '1Jrh-bRnJEjyMCS7f-WsaFlccfPjJPPHI&confirm=t' # omnidata_dpt_depth_v2.ckpt
gdown '1wNxVO4vVbDEMEpnAi_jwQObf2MFodcBR&confirm=t' # omnidata_dpt_normal_v2.ckpt
To use DeepFloyd-IF, you need to accept the usage conditions from hugging face, and login with huggingface-cli login in command line.
For DMTet, we port the pre-generated 32/64/128 resolution tetrahedron grids under tets.
The 256 resolution one can be found here.
By default, we use load to build the extension at runtime.
We also provide the setup.py to build each extension:
cd stable-dreamfusion
# install all extension modules
bash scripts/install_ext.sh
# if you want to install manually, here is an example:
pip install ./raymarching # install to python path (you still need the raymarching/ folder, since this only installs the built extension.)
Use Taichi backend for Instant-NGP. It achieves comparable performance to CUDA implementation while No CUDA build is required. Install Taichi with pip:
pip install -i https://pypi.taichi.graphics/simple/ taichi-nightly
pip install -U diffusers). If the problem still holds, reporting a bug issue will be appreciated![F glutil.cpp:338] eglInitialize() failed Aborted (core dumped): this usually indicates problems in OpenGL installation. Try to re-install Nvidia driver, or use nvidia-docker as suggested in https://github.com/ashawkey/stable-dreamfusion/issues/131 if you are using a headless server.TypeError: xxx_forward(): incompatible function arguments: this happens when we update the CUDA source and you used setup.py to install the extensions earlier. Try to re-install the corresponding extension (e.g., pip install ./gridencoder).First time running will take some time to compile the CUDA extensions.
…
For example commands, check scripts.
For advanced tips and other developing stuff, check Advanced Tips.
Reproduce the paper CLIP R-precision evaluation
After the testing part in the usage, the validation set containing projection from different angle is generated. Test the R-precision between prompt and the image.(R=1)
python r_precision.py --text "a snake is flying in the sky" --workspace snake_HQ --latest ep0100 --mode depth --clip clip-ViT-B-16
This work is based on an increasing list of amazing research works and open-source projects, thanks a lot to all the authors for sharing!
DreamFusion: Text-to-3D using 2D Diffusion
@article{poole2022dreamfusion,
author = {Poole, Ben and Jain, Ajay and Barron, Jonathan T. and Mildenhall, Ben},
title = {DreamFusion: Text-to-3D using 2D Diffusion},
journal = {arXiv},
year = {2022},
}
Magic3D: High-Resolution Text-to-3D Content Creation
@inproceedings{lin2023magic3d,
title={Magic3D: High-Resolution Text-to-3D Content Creation},
author={Lin, Chen-Hsuan and Gao, Jun and Tang, Luming and Takikawa, Towaki and Zeng, Xiaohui and Huang, Xun and Kreis, Karsten and Fidler, Sanja and Liu, Ming-Yu and Lin, Tsung-Yi},
booktitle={IEEE Conference on Computer Vision and Pattern Recognition ({CVPR})},
year={2023}
}
Zero-1-to-3: Zero-shot One Image to 3D Object
@misc{liu2023zero1to3,
title={Zero-1-to-3: Zero-shot One Image to 3D Object},
author={Ruoshi Liu and Rundi Wu and Basile Van Hoorick and Pavel Tokmakov and Sergey Zakharov and Carl Vondrick},
year={2023},
eprint={2303.11328},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
@article{armandpour2023re,
title={Re-imagine the Negative Prompt Algorithm: Transform 2D Diffusion into 3D, alleviate Janus problem and Beyond},
author={Armandpour, Mohammadreza and Zheng, Huangjie and Sadeghian, Ali and Sadeghian, Amir and Zhou, Mingyuan},
journal={arXiv preprint arXiv:2304.04968},
year={2023}
}
RealFusion: 360° Reconstruction of Any Object from a Single Image
@inproceedings{melaskyriazi2023realfusion,
author = {Melas-Kyriazi, Luke and Rupprecht, Christian and Laina, Iro and Vedaldi, Andrea},
title = {RealFusion: 360 Reconstruction of Any Object from a Single Image},
booktitle={CVPR}
year = {2023},
url = {https://arxiv.org/abs/2302.10663},
}
Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation
@article{chen2023fantasia3d,
title={Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation},
author={Rui Chen and Yongwei Chen and Ningxin Jiao and Kui Jia},
journal={arXiv preprint arXiv:2303.13873},
year={2023}
}
Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior
@article{tang2023make,
title={Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior},
author={Tang, Junshu and Wang, Tengfei and Zhang, Bo and Zhang, Ting and Yi, Ran and Ma, Lizhuang and Chen, Dong},
journal={arXiv preprint arXiv:2303.14184},
year={2023}
}
Stable Diffusion and the diffusers library.
…
The GUI is developed with DearPyGui.
Puppy image from : https://www.pexels.com/photo/high-angle-photo-of-a-corgi-looking-upwards-2664417/
Anya images from : https://www.goodsmile.info/en/product/13301/POP+UP+PARADE+Anya+Forger.html
If you find this work useful, a citation will be appreciated via:
@misc{stable-dreamfusion,
Author = {Jiaxiang Tang},
Year = {2022},
Note = {https://github.com/ashawkey/stable-dreamfusion},
Title = {Stable-dreamfusion: Text-to-3D with Stable-diffusion}
}
No open issues yet, or sync has not completed.