[ECCV2024] API code for T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
[ECCV2024] API code for T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
[2025/10/15] Rex-Omni: Still using traditional detectors? We've turned object detection into a simple "Next-Token Prediction" task with an MLLM! One model (fully open-sourced), zero-shot SOTA performance, tackling detection, referring, OCR, and GUI grounding all at once. Come see the next generation of perception models
You can get API access here https://cloud.deepdataspace.com/dashboard/usage. Once you get the API key, you can try T-Rex2 by following these example codes: https://github.com/IDEA-Research/T-Rex/tree/trex2/demo_examples
Turn on the music if possible
Object detection, the ability to locate and identify objects within an image, is a cornerstone of computer vision, pivotal to applications ranging from autonomous driving to content moderation. A notable limitation of traditional object detection models is their closed-set nature. These models are trained on a predetermined set of categories, confining their ability to recognize only those specific categories. The training process itself is arduous, demanding expert knowledge, extensive datasets, and intricate model tuning to achieve desirable accuracy. Moreover, the introduction of a novel object category, exacerbates these challenges, necessitating the entire process to be repeated.
T-Rex2 addresses these limitations by integrating both text and visual prompts in one model, thereby harnessing the strengths of both modalities. The synergy of text and visual prompts equips T-Rex2 with robust zero-shot capabilities, making it a versatile tool in the ever-changing landscape of object detection.
T-Rex2 is well-suited for a variety of real-world applications, including but not limited to: agriculture, industry, livstock and wild animals monitoring, biology, medicine, OCR, retail, electronics, transportation, logistics, and more. T-Rex2 mainly supports three major workflows including interactive visual prompt workflow, generic visual prompt workflow and text prompt workflow. It can cover most of the application scenarios that require object detection
We are now opening online demo for T-Rex2. Check our demo here
You can get API access here https://cloud.deepdataspace.com/dashboard/usage. Once you get the API key, you can try T-Rex2 by following these example codes: https://github.com/IDEA-Research/T-Rex/tree/trex2/demo_examples
Install the API package and acquire the API token from the email.
git clone https://github.com/IDEA-Research/T-Rex.git
cd T-Rex
pip install dds-cloudapi-sdk==0.1.1
pip install -v -e .
In interactive visual prompt workflow, users can provide visual prompts in boxes or points format on a given image to specify the object to be detected.
python demo_examples/interactive_inference.py --token <your_token>
demo_vis/In generic visual prompt workflow, users can provide visual prompts on one reference image and detect on the other image.
python demo_examples/generic_inference.py --token <your_token>
demo_vis/In this workflow, you can customize a visual embedding for a object category using multiple images. With this embedding, you can detect on any images.
python demo_examples/customize_embedding.py --token <your_token>
safetensors format. Save it and let's use it for embedding_inference.With the visual prompt embeddings generated from the previous API. You can use it detect on any images.
python demo_examples/embedding_inference.py --token <your_token>
- install gradio and other dependencies
```bash
# install gradio and other dependencies
pip install gradio-image-prompter
python gradio_demo.py --trex2_api_token <your_token>
…
@misc{jiang2024trex2,
title={T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy},
author={Qing Jiang and Feng Li and Zhaoyang Zeng and Tianhe Ren and Shilong Liu and Lei Zhang},
year={2024},
eprint={2403.14610},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
No open issues yet, or sync has not completed.