Feature: support precomputed image embeddings
Author: rom1504Created May 26, 2022Updated Nov 28, 2022
Labelsnew feature
Implementing that would make it possible to efficiently train a lit-style clip model. Which is pretty useful when a decent visual encoder is available and the goal is to map it to captions (eg multilingual ones)
https://github.com/lucidrains/DALLE2-pytorch/blob/main/train_diffusion_prior.py has a good data loader for embedding+caption, this can be reused.
We may work on this, not asking for someone else to do it :)
Source: mlfoundations/open_clip