百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
G

GLIP

> 编程语言
开源

基于语言的图像预训练

2.6K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

基于语言的图像预训练

# GLIP: Grounded Language-Image Pre-training ## Updates * 01/17/2023: From image understanding to image generation for open-set grounding? Check out [**GLIGEN (Grounded Language-to-Image Generation)**](https://gligen.github.io/) - GLIGEN: (box, concept) $\rightarrow$ image || GLIP: image $\rightarrow$ (box, concept) * 09/19/2022: GLIPv2 has been accepted to NeurIPS 2022 ([Updated Version](https://arxiv.org/abs/2206.05836)). * 09/18/2022: Organizing ECCV Workshop [*Computer Vision in the Wild (CVinW)*](https://computer-vision-in-the-wild.github.io/eccv-2022/), where two challenges are hosted to evaluate the zero-shot, few-shot and full-shot performance of pre-trained vision models in downstream tasks: - [``*Image Classification in the Wild (ICinW)*''](https://eval.ai/web/challenges/challenge-page/1832/overview) Challenge evaluates on 20 image classification tasks. - [``*Object Detection in the Wild (ODinW)*''](https://eval.ai/web/challenges/challenge-page/1839/overview) Challenge evaluates on 35 object detection tasks. $\qquad$ [ [Workshop]](https://computer-vision-in-the-wild.github.io/eccv-2022/) $\qquad$ [ [IC Challenge] ](https://eval.ai/web/challenges/challenge-page/1832/overview) $\qquad$ [ [OD Challenge] ](https://eval.ai/web/challenges/challenge-page/1839/overview) * 09/13/2022: Updated [HuggingFace Demo](https://huggingface.co/spaces/haotiz/glip-zeroshot-demo)! Feel free to give it a try!!! - Acknowledgement: Many thanks to the help from @[HuggingFace](https://huggingface.co/) for a Space GPU upgrade to host the GLIP demo! * 06/21/2022: GLIP has been selected as a Best Paper Finalist at CVPR 2022! * 06/16/2022: ODinW benchmark released! GLIP-T A&B released! * 06/13/2022: GLIPv2 is on Arxiv https://arxiv.org/abs/2206.05836! * 04/30/2022: Updated [Colab Demo](https://colab.research.google.com/drive/12x7v-_miN7-SRiziK3Cx4ffJzstBJNqb?usp=sharing)! * 04/14/2022: GLIP has been accepted to CVPR 2022 as an oral presentation! First version of code and pre-trained models are released! * 12/06/2021: GLIP paper on arxiv https://arxiv.org/abs/2112.03857. * 11/23/2021: Project page built.
## Introduction This repository is the project page for [GLIP](https://arxiv.org/abs/2112.03857). GLIP demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks. 1. When directly evaluated on COCO and LVIS (without seeing any images in COCO), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines. 2. After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA. 3. When transferred to 13 downstream object detection tasks, a few-shot GLIP rivals with a fully-supervised Dynamic Head. We provide code for: 1. **pre-training** GLIP on detection and grounding data; 2. **zero-shot evaluating** GLIP on standard benchmarks (COCO, LVIS, Flickr30K) and custom COCO-formated datasets; 3. **fine-tuning** GLIP on standard benchmarks (COCO) and custom COCO-formated datasets; 4. **a Colab demo**. 5. Toolkits for **the Object Detection in the Wild Benchmark (ODinW)** with 35 downstream detection tasks. Please see respective sections for instructions. ## Demo Please see a Colab demo at [link](https://colab.research.google.com/drive/12x7v-_miN7-SRiziK3Cx4ffJzstBJNqb?usp=sharing)! ## Installation and Setup ***Environment*** This repo requires Pytorch>=1.9 and torchvision. We recommand using docker to setup the environment. You can use this pre-built docker image ``docker pull pengchuanzhang/maskrcnn:ubuntu18-py3.7-cuda10.2-pytorch1.9`` or this one ``docker pull pengchuanzhang/pytorch:ubuntu20.04_torch1.9-cuda11.3-nccl2.9.9`` depending on your GPU. Then install the following packages: ``` pip install einops shapely timm yacs tensorboardX ftfy prettytable pymongo pip install transformers python setup.py build develop --user ``` ***Backbone Checkpoints.*** Download the ImageNet pre-trained backbone checkpoints into the ``MODEL`` folder. ``` mkdir MODEL wget https://penzhanwu2bbs.blob.core.windows.net/data/GLIPv1_Open/models/swin_tiny_patch4_window7_224.pth -O swin_tiny_patch4_window7_224.pth wget https://penzhanwu2bbs.blob.core.windows.net/data/GLIPv1_Open/models/swin_large_patch4_window12_384_22k.pth -O swin_large_patch4_window12_384_22k.pth ``` ## Model Zoo ***Checkpoint host move.*** The checkpoint links expired. We are moving the checkpoints to https://huggingface.co/harold/GLIP/tree/main. Currently most checkpoints are available. Working to host the remaining checkpoints asap. Model | COCO [1] | LVIS [2] | LVIS [3] | ODinW [4] | Pre-Train Data | Config | Weight -- | -- | -- | -- | -- | -- | -- | -- GLIP-T (**A**) | 42.9 / 52.9 | - | 14.2 | ~28.7 | O365 | [config](configs/pretrain/glip_A_Swin_T_O365.yaml) | [weight](https://huggingface.co/GLIPModel/GLIP/blob/main/glip_a_tiny_o365.pth) GLIP-T (**B**) | 44.9 / 53.8 | - | 13.5 | ~33.2 | O365 | [config](configs/pretrain/glip_Swin_T_O365.yaml) | [weight](https://huggingface.co/GLIPModel/GLIP/blob/main/glip_tiny_model_o365.pth) GLIP-T (**C**) | 46.7 / 55.1 | 14.3 | [17.7](https://penzhanwu2bbs.blob.core.windows.net/data/GLIPv1_Open/models/glip_tiny_model_o365_goldg_lvisbest.pth) | 44.4 | O365,GoldG | [config](configs/pretrain/glip_Swin_T_O365_GoldG.yaml) | [weight](https://huggingface.co/GLIPModel/GLIP/blob/main/glip_tiny_model_o365_goldg.pth) **GLIP-T** [5] | 46.6 / 55.2 | 17.6 | [20.1](https://huggingface.co/GLIPModel/GLIP/blob/main/glip_tiny_model_o365_goldg_cc_sbu_lvisbest.pth) | 42.7 | O365,GoldG,CC3M,SBU | [config](configs/pretrain/glip_Swin_T_O365_GoldG.yaml) [6] | [weight](https://huggingface.co/GLIPModel/GLIP/blob/main/glip_tiny_model_o365_goldg_cc_sbu.pth) **GLIP-L** [7] | 51.4 / 61.7 [8] | 29.3 | [30.1](https://penzhanwu2bbs.blob.core.windows.net/data/GLIPv1_Open/models/glip_large_model_lvisbest.pth) | 51.2 | FourODs,GoldG,CC3M+12M,SBU | [config](configs/pretrain/glip_Swin_L.yaml) [9] | [weight](https://huggingface.co/GLIPModel/GLIP/blob/main/glip_large_model.pth) [1] Zero-shot and fine-tuning performance on COCO val2017. [2] Zero-shot performance on LVIS minival (APr) with the last pre-trained checkpoint. [3] On LVIS, the model could overfit slightly during the pre-training course. Thus we reported two numbers on LVIS: the performance of the last checkpoint (LVIS[2]) and the performance of the best checkpoint during the pre-training course (LVIS[3]). [4] Zero-shot performance on the 13 ODinW datasets. The numbers reported in the GLIP paper is from the best checkpoint during the pre-training course, which may be slightly higher than the numbers for the released last checkpoint, similar to the case of LVIS. [5] GLIP-T released in this repo is pre-trained on Conceptual Captions 3M and SBU captions. It is referred in paper in Table 1 and in Appendix C.3. It differs slightly from the GLIP-T in the main paper in terms of downstream performance. We will release the pre-training support for using CC3M and SBU captions data in the next update. [6] This config is only intended for zero-shot evaluation and fine-tuning. Pre-training config with support for using CC3M and SBU captions data will be updated. [7] GLIP-L released in this repo is pre-trained on Conceptual Captions 3M+12M and SBU captions. It slightly outperforms the GLIP-L in the main paper because the model used to annotate the caption data are improved compared to the main paper. We will release the pre-training support for using CC3M+12M and SBU captions data in the next update. [8] Multi-scale testing used. [9] This config is only intended for zero-shot evaluation and fine-tuning. Pre-training config with support for using CC3M+12M and SBU captions data to be updated. ## Pre-Training ***Required Data.*** Prepare ``Objects365``, ``Flickr30K``, and ``MixedGrounding`` data as in [DATA.md](DATA.md). Support for training using caption data (Conceptual Captions and SBU captions) will be released soon. ***Command.*** Perform pre-training with the following command (please change the config-file accordingly; checkout model zoo for the corresponding config; change the ``{output_dir}`` to your desired output directory): ``` python -m torch.distributed.launch --nnodes 2 --nproc_per_node=16 tools/train_net.py \ --config-file configs/pretrain/glip_Swin_T_O365_GoldG.yaml \ --skip-test --use-tensorboard --override_output_dir {output_dir} ``` For training GLIP-T models, we used `nnodes = 2`, `nproc_per_node=16` on 32GB V100 machines. For training GLIP-L models, we used `nnodes = 4`, `nproc_per_node=16` on 32GB V100 machines. Please setup the environment accordingly based on your local machine. ## (Zero-Shot) Evaluation ### COCO Evaluation Prepare ``COCO/val2017`` data as in [DATA.md](DATA.md). Set ``{config_file}``, ``{model_checkpoint}`` according to the ``Model Zoo``; set ``{output_dir}`` to a folder where the evaluation results will be stored. ``` python tools/test_grounding_net.py --config-file {config_file} --weight {model_checkpoint} \ TEST.IMS_PER_BATCH 1 \ MODEL.DYHEAD.SCORE_AGG "MEAN" \ TEST.EVAL_TASK detection \ MODEL.DYHEAD.FUSE_CONFIG.MLM_LOSS False \ OUTPUT_DIR {output_dir} ``` ### LVIS Evaluation We follow MDETR to evaluate with the [FixedAP](https://arxiv.org/pdf/2102.01066.pdf) criterion. Set ``{config_file}``, ``{model_checkpoint}`` according to the ``Model Zoo``. Prepare ``COCO/val2017`` data as in [DATA.md](DATA.md). ``` python -m torch.distributed.launch --nproc_per_node=4 \ tools/test_grounding_net.py \ --config-file {config_file} \ --task_config configs/lvis/minival.yaml \ --weight {model_checkpoint} \ TEST.EVAL_TASK detection OUTPUT_DIR {output_dir} TEST.CHUNKED_EVALUATION 40 TEST.IMS_PER_BATCH 4 SOLVER.IMS_PER_BATCH 4 TEST.MDETR_STYLE_AGGREGATE_CLASS_NUM 3000 MODEL.RETINANET.DETECTIONS_PER_IMG 300 MODEL.FCOS.DETECTIONS_PER_IMG 300 MODEL.ATSS.DETECTIONS_PER_IMG 300 MODEL.ROI_HEADS.DETECTIONS_PER_IMG 300 ``` If you wish to evaluate on Val 1.0, set ``--task_config`` to ``configs/lvis/val.yaml``. ### ODinW / Custom Dataset Evaluation GLIP supports easy evaluation on a custom dataset. Currently, the code supports evaluation on [COCO-formatted](https://cocodataset.org/#format-data) dataset. We will use the [Aquarium](https://public.roboflow.com/object-detection/aquarium) dataset from ODinW as an example to show how to evaluate on a custom COCO-formatted dataset. 1. Download the raw dataset from RoboFlow in the COCO format into ``DATASET/odinw/Aquarium``. Each train/val/test split has a corresponding ``annotation`` file and a ``image`` folder. 2. Remove the background class from the annotation file. This can be as simple as open "_annotations.coco.json" and remove the entry with "id:0" from "categories". For convenience, we provide the modified annotation files for Aquarium: ``` … ``` 4. Then create a yaml file as in ``configs/odinw_13/Aquarium_Aquarium_Combined.v2-raw-1024.coco.yaml``. A few fields to be noted in the yamls: DATASET.CAPTION_PROMPT allows manually changing the prompt (the default prompt is simply concatnating all the categories); MODELS.\*.NUM_CLASSES need to be set to the number of categories in the dataset (including the background class). E.g., Aquarium has 7 non-background categories thus MODELS.\*.NUM_CLASSES is set to 8; 4. Run the following command to evaluate on the dataset. Set ``{config_file}``, ``{model_checkpoint}`` according to the ``Model Zoo``. Set {odinw_configs} to the path of the task yaml file we just prepared. ``` python tools/test_grounding_net.py --config-file {config_file} --weight {model_checkpoint} \ --task_config {odinw_configs} \ TEST.IMS_PER_BATCH 1 SOLVER.IMS_PER_BATCH 1 \ TEST.EVAL_TASK detection \ DATASETS.TRAIN_DATASETNAME_SUFFIX _grounding \ DATALOADER.DISTRIBUTE_CHUNK_AMONG_NODE False \

GitHub Issues· 0 开放

在 GitHub 查看全部

暂无开放 Issues,或尚未同步最近议题。

核心特点

  • •01/17/2023: From image understanding to image generation for open-set grounding? Check out GLIGEN (Grounded Language-to-Image Generation)
  • •GLIGEN: (box, concept) $\rightarrow$ image || GLIP: image $\rightarrow$ (box, concept)
  • •09/19/2022: GLIPv2 has been accepted to NeurIPS 2022 (Updated Version).
  • •``*Image Classification in the Wild (ICinW)*'' Challenge evaluates on 20 image classification tasks.
  • •``*Object Detection in the Wild (ODinW)*'' Challenge evaluates on 35 object detection tasks.
  • •09/13/2022: Updated HuggingFace Demo! Feel free to give it a try!!!
  • •Acknowledgement: Many thanks to the help from @HuggingFace for a Space GPU upgrade to host the GLIP demo!
  • •06/21/2022: GLIP has been selected as a Best Paper Finalist at CVPR 2022!
  • •06/16/2022: ODinW benchmark released! GLIP-T A&B released!
  • •06/13/2022: GLIPv2 is on Arxiv https://arxiv.org/abs/2206.05836!

> 标签

Python

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言