百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
M

ml-aim

> 编程语言
开源

此仓库提供了 AIMv1 和 AIMv2 研究项目的代码和模型检查点。

1.4K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

此仓库提供了 AIMv1 和 AIMv2 研究项目的代码和模型检查点。

# Autoregressive Pre-training of Large Vision Encoders This repository is the entry point for all things AIM, a family of autoregressive models that push the boundaries of visual and multimodal learning: - **AIMv2**: [`Multimodal Autoregressive Pre-training of Large Vision Encoders`](https://arxiv.org/abs/2411.14402) [[`BibTeX`](#citation)] [**CVPR 2025 (Highlight)**]
Enrico Fini*, Mustafa Shukor*, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor Guilherme Turrisi da Costa, Louis Béthune, Zhe Gan, Alexander T Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, and Alaaeldin El-Nouby* - **AIMv1**: [`Scalable Pre-training of Large Autoregressive Image Models`](https://arxiv.org/abs/2401.08541) [[`BibTeX`](#citation)][**ICML 2024**]
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M Susskind, Armand Joulin. *: Equal technical contribution If you're looking for the original AIM model (AIMv1), please refer to the README [here](aim-v1/README.md). --- ## Overview of AIMv2 We introduce the AIMv2 family of vision models pre-trained with a multimodal autoregressive objective. AIMv2 pre-training is simple and straightforward to train and to scale effectively. Some AIMv2 highlights include: 1. Outperforms OAI CLIP and SigLIP on the majority of multimodal understanding benchmarks. 2. Outperforms DINOv2 on open-vocabulary object detection and referring expression comprehension. 3. Exhibits strong recognition performance with AIMv2-3B achieving *89.5% on ImageNet using a frozen trunk*. ## AIMv2 Model Gallery We share with the community AIMv2 pre-trained checkpoints of varying capacities, pre-training resolutions: + [[`AIMv2 with 224px`]](#aimv2-with-224px) + [[`AIMv2 with 336px`]](#aimv2-with-336px) + [[`AIMv2 with 448px`]](#aimv2-with-448px) + [[`AIMv2 with Native Resolution`]](#aimv2-with-native-resolution) + [[`AIMv2 distilled ViT-Large`]](#aimv2-distilled-vit-large) (*recommended for multimodal applications*) + [[`Zero-shot Adapted AIMv2`]](#zero-shot-adapted-aimv2) ## Installation Please install PyTorch using the official [installation instructions](https://pytorch.org/get-started/locally/). Afterward, install the package as: ```commandline pip install 'git+https://github.com/apple/ml-aim.git#subdirectory=aim-v1' pip install 'git+https://github.com/apple/ml-aim.git#subdirectory=aim-v2' ``` We also offer [MLX](https://ml-explore.github.io/mlx/) backend support for research and experimentation on Apple silicon. To enable MLX support, simply run: ```commandline pip install mlx ``` ## Examples ### Using PyTorch ```python from PIL import Image from aim.v2.utils import load_pretrained from aim.v1.torch.data import val_transforms img = Image.open(...) model = load_pretrained("aimv2-large-patch14-336", backend="torch") transform = val_transforms(img_size=336) inp = transform(img).unsqueeze(0) features = model(inp) ``` ### Using MLX ```python from PIL import Image import mlx.core as mx from aim.v2.utils import load_pretrained from aim.v1.torch.data import val_transforms img = Image.open(...) model = load_pretrained("aimv2-large-patch14-336", backend="mlx") transform = val_transforms(img_size=336) inp = transform(img).unsqueeze(0) inp = mx.array(inp.numpy()) features = model(inp) ``` ### Using JAX ```python from PIL import Image import jax.numpy as jnp from aim.v2.utils import load_pretrained from aim.v1.torch.data import val_transforms img = Image.open(...) model, params = load_pretrained("aimv2-large-patch14-336", backend="jax") transform = val_transforms(img_size=336) inp = transform(img).unsqueeze(0) inp = jnp.array(inp) features = model.apply({"params": params}, inp) ``` ## Pre-trained Checkpoints The pre-trained models can be accessed via [HuggingFace Hub](https://huggingface.co/collections/apple/aimv2-6720fe1558d94c7805f7688c) as: ```python from PIL import Image from transformers import AutoImageProcessor, AutoModel image = Image.open(...) processor = AutoImageProcessor.from_pretrained("apple/aimv2-large-patch14-336") model = AutoModel.from_pretrained("apple/aimv2-large-patch14-336", trust_remote_code=True) inputs = processor(images=image, return_tensors="pt") outputs = model(**inputs) ``` ### AIMv2 with 224px
model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-224 0.3B 86.6 link link
aimv2-huge-patch14-224 0.6B 87.5 link link
aimv2-1B-patch14-224 1.2B 88.1 link link
aimv2-3B-patch14-224 2.7B 88.5 link link
### AIMv2 with 336px
model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-336 0.3B 87.6 link link
aimv2-huge-patch14-336 0.6B 88.2 link link
aimv2-1B-patch14-336 1.2B 88.7 link link
aimv2-3B-patch14-336 2.7B 89.2 link link
### AIMv2 with 448px
model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-448 0.3B 87.9 link link
aimv2-huge-patch14-448 0.6B 88.6 link link
aimv2-1B-patch14-448 1.2B 89.0 link link
aimv2-3B-patch14-448 2.7B 89.5 link link
### AIMv2 with Native Resolution We additionally provide an AIMv2-L checkpoint that is finetuned to process a wide range of image resolutions and aspect ratios. Regardless of the aspect ratio, the image is patchified (patch_size=14) and *a 2D sinusoidal positional embedding* is added to the linearly projected input patches. *This checkpoint supports number of patches in the range of [112, 4096]*.
model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-native 0.3B 87.3 link link
### AIMv2 distilled ViT-Large We provide an AIMv2-L checkpoint distilled from AIMv2-3B that provides a remarkable performance for multimodal understanding benchmarks.
Model VQAv2 GQA OKVQA TextVQA DocVQA InfoVQA ChartQA SciQA MMEp
AIMv2-L 80.2 72.6 60.9 53.9 26.8 22.4 20.3 74.5 1457
AIMv2-L-distilled 81.1 73.0 61.4 53.5 29.2 23.3 24.0 76.3 1627
model_id #params Res. HF Link Backbone
aimv2-large-patch14-224-distilled 0.3B 224px link link
aimv2-large-patch14-336-distilled 0.3B 336px link link
### Zero-shot Adapted AIMv2 We provide the AIMv2-L vision and text encoders after LiT tuning to enable zero-shot recognition.
model #params zero-shot IN1-k Backbone

Issues· 0 开放

查看全部 Issues在 GitHub 打开

暂无开放 Issues,或尚未同步最近议题。

> 标签

Pythonjaxlarge-scale-vision-modelsmlxpytorch

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言