Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
< Back to tools
M

ml-aim

> 编程语言
Open source

This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.

1.4K stars0 likes0 views
WebsiteGitHub

About

This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.

Autoregressive Pre-training of Large Vision Encoders

This repository is the entry point for all things AIM, a family of autoregressive models that push the boundaries of visual and multimodal learning:

  • AIMv2: Multimodal Autoregressive Pre-training of Large Vision Encoders [BibTeX] [CVPR 2025 (Highlight)]
    Enrico Fini*, Mustafa Shukor*, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor Guilherme Turrisi da Costa, Louis Béthune, Zhe Gan, Alexander T Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, and Alaaeldin El-Nouby*
  • AIMv1: Scalable Pre-training of Large Autoregressive Image Models [BibTeX][ICML 2024]
    Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M Susskind, Armand Joulin.

*: Equal technical contribution

If you're looking for the original AIM model (AIMv1), please refer to the README here.


Overview of AIMv2

We introduce the AIMv2 family of vision models pre-trained with a multimodal autoregressive objective. AIMv2 pre-training is simple and straightforward to train and to scale effectively. Some AIMv2 highlights include:

  1. Outperforms OAI CLIP and SigLIP on the majority of multimodal understanding benchmarks.
  2. Outperforms DINOv2 on open-vocabulary object detection and referring expression comprehension.
  3. Exhibits strong recognition performance with AIMv2-3B achieving 89.5% on ImageNet using a frozen trunk.

AIMv2 Model Gallery

We share with the community AIMv2 pre-trained checkpoints of varying capacities, pre-training resolutions:

  • [AIMv2 with 224px]
  • [AIMv2 with 336px]
  • [AIMv2 with 448px]
  • [AIMv2 with Native Resolution]
  • [AIMv2 distilled ViT-Large] (recommended for multimodal applications)
  • [Zero-shot Adapted AIMv2]

Installation

Please install PyTorch using the official installation instructions. Afterward, install the package as:

commandline
pip install 'git+https://github.com/apple/ml-aim.git#subdirectory=aim-v1'
pip install 'git+https://github.com/apple/ml-aim.git#subdirectory=aim-v2'

We also offer MLX backend support for research and experimentation on Apple silicon. To enable MLX support, simply run:

commandline
pip install mlx

Examples

Using PyTorch

python
from PIL import Image

from aim.v2.utils import load_pretrained
from aim.v1.torch.data import val_transforms

img = Image.open(...)
model = load_pretrained("aimv2-large-patch14-336", backend="torch")
transform = val_transforms(img_size=336)

inp = transform(img).unsqueeze(0)
features = model(inp)

Using MLX

python
from PIL import Image
import mlx.core as mx

from aim.v2.utils import load_pretrained
from aim.v1.torch.data import val_transforms

img = Image.open(...)
model = load_pretrained("aimv2-large-patch14-336", backend="mlx")
transform = val_transforms(img_size=336)

inp = transform(img).unsqueeze(0)
inp = mx.array(inp.numpy())
features = model(inp)

Using JAX

python
from PIL import Image
import jax.numpy as jnp

from aim.v2.utils import load_pretrained
from aim.v1.torch.data import val_transforms

img = Image.open(...)
model, params = load_pretrained("aimv2-large-patch14-336", backend="jax")
transform = val_transforms(img_size=336)

inp = transform(img).unsqueeze(0)
inp = jnp.array(inp)
features = model.apply({"params": params}, inp)

Pre-trained Checkpoints

The pre-trained models can be accessed via HuggingFace Hub as:

python
from PIL import Image
from transformers import AutoImageProcessor, AutoModel

image = Image.open(...)
processor = AutoImageProcessor.from_pretrained("apple/aimv2-large-patch14-336")
model = AutoModel.from_pretrained("apple/aimv2-large-patch14-336", trust_remote_code=True)

inputs = processor(images=image, return_tensors="pt")
outputs = model(**inputs)

AIMv2 with 224px

model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-224 0.3B 86.6 link link
aimv2-huge-patch14-224 0.6B 87.5 link link
aimv2-1B-patch14-224 1.2B 88.1 link link
aimv2-3B-patch14-224 2.7B 88.5 link link

AIMv2 with 336px

model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-336 0.3B 87.6 link link
aimv2-huge-patch14-336 0.6B 88.2 link link
aimv2-1B-patch14-336 1.2B 88.7 link link
aimv2-3B-patch14-336 2.7B 89.2 link link

AIMv2 with 448px

model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-448 0.3B 87.9 link link
aimv2-huge-patch14-448 0.6B 88.6 link link
aimv2-1B-patch14-448 1.2B 89.0 link link
aimv2-3B-patch14-448 2.7B 89.5 link link

AIMv2 with Native Resolution

We additionally provide an AIMv2-L checkpoint that is finetuned to process a wide range of image resolutions and aspect ratios. Regardless of the aspect ratio, the image is patchified (patch_size=14) and a 2D sinusoidal positional embedding is added to the linearly projected input patches. This checkpoint supports number of patches in the range of [112, 4096].

model_id #params IN-1k HF Link Backbone
aimv2-large-patch14-native 0.3B 87.3 link link

AIMv2 distilled ViT-Large

We provide an AIMv2-L checkpoint distilled from AIMv2-3B that provides a remarkable performance for multimodal understanding benchmarks.

Model VQAv2 GQA OKVQA TextVQA DocVQA InfoVQA ChartQA SciQA MMEp
AIMv2-L 80.2 72.6 60.9 53.9 26.8 22.4 20.3 74.5 1457
AIMv2-L-distilled 81.1 73.0 61.4 53.5 29.2 23.3 24.0 76.3 1627
model_id #params Res. HF Link Backbone
aimv2-large-patch14-224-distilled 0.3B 224px link link
aimv2-large-patch14-336-distilled 0.3B 336px link link

Zero-shot Adapted AIMv2

We provide the AIMv2-L vision and text encoders after LiT tuning to enable zero-shot recognition.

model #params zero-shot IN1-k Backbone

Issues· 0 open

View all issuesOpen on GitHub

No open issues yet, or sync has not completed.

> Tags

Pythonjaxlarge-scale-vision-modelsmlxpytorch

No comments yet. Be the first to share.

> Details

PublishedAug 1, 2026
UpdatedSep 17, 2026
Category编程语言
PricingOpen source

> Related tools

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言