OpenMMLab 的下一代视频理解工具箱和基准
Documentation | ️Installation | Model Zoo | Update News | Ongoing Projects | Reporting Issues
English | 简体中文
The default branch has been switched to main(previous 1.x) from master(current 0.x), and we encourage users to migrate to the latest version with more supported models, stronger pre-training checkpoints and simpler coding. Please refer to Migration Guide for more details.
Release (2023.10.12): v1.2.0 with the following new features:
MMAction2 is an open-source toolbox for video understanding based on PyTorch. It is a part of the OpenMMLab project.
Modular design: We decompose a video understanding framework into different components. One can easily construct a customized video understanding framework by combining different modules.
Support five major video understanding tasks: MMAction2 implements various algorithms for multiple video understanding tasks, including action recognition, action localization, spatio-temporal action detection, skeleton-based action detection and video retrieval.
Well tested and documented: We provide detailed documentation and API reference, as well as unit tests.
MMAction2 depends on PyTorch, MMCV, MMEngine, MMDetection (optional) and MMPose (optional).
Please refer to install.md for detailed instructions.
Quick instructionsconda create --name openmmlab python=3.8 -y
conda activate openmmlab
conda install pytorch torchvision -c pytorch # This command will automatically install the latest version PyTorch and cudatoolkit, please check whether they match your environment.
pip install -U openmim
mim install mmengine
mim install mmcv
mim install mmdet # optional
mim install mmpose # optional
git clone https://github.com/open-mmlab/mmaction2.git
cd mmaction2
pip install -v -e .Results and models are available in the model zoo.
Supported model| Action Recognition | ||||
| C3D (CVPR'2014) | TSN (ECCV'2016) | I3D (CVPR'2017) | C2D (CVPR'2018) | I3D Non-Local (CVPR'2018) |
| R(2+1)D (CVPR'2018) | TRN (ECCV'2018) | TSM (ICCV'2019) | TSM Non-Local (ICCV'2019) | SlowOnly (ICCV'2019) |
| SlowFast (ICCV'2019) | CSN (ICCV'2019) | TIN (AAAI'2020) | TPN (CVPR'2020) | X3D (CVPR'2020) |
| MultiModality: Audio (ArXiv'2020) | TANet (ArXiv'2020) | TimeSformer (ICML'2021) | ActionCLIP (ArXiv'2021) | VideoSwin (CVPR'2022) |
| VideoMAE (NeurIPS'2022) | MViT V2 (CVPR'2022) | UniFormer V1 (ICLR'2022) | UniFormer V2 (Arxiv'2022) | VideoMAE V2 (CVPR'2023) |
| Action Localization | ||||
| BSN (ECCV'2018) | BMN (ICCV'2019) | TCANet (CVPR'2021) | ||
| Spatio-Temporal Action Detection | ||||
| ACRN (ECCV'2018) | SlowOnly+Fast R-CNN (ICCV'2019) | SlowFast+Fast R-CNN (ICCV'2019) | LFB (CVPR'2019) | VideoMAE (NeurIPS'2022) |
| Skeleton-based Action Recognition | ||||
| ST-GCN (AAAI'2018) | 2s-AGCN (CVPR'2019) | PoseC3D (CVPR'2022) | STGCN++ (ArXiv'2022) | CTRGCN (CVPR'2021) |
| MSG3D (CVPR'2020) | ||||
| Video Retrieval | ||||
| CLIP4Clip (ArXiv'2022) |
| Action Recognition | |||
| HMDB51 (Homepage) (ICCV'2011) | UCF101 (Homepage) (CRCV-IR-12-01) | ActivityNet (Homepage) (CVPR'2015) | Kinetics-[400/600/700] (Homepage) (CVPR'2017) |
| SthV1 (ICCV'2017) | SthV2 (Homepage) (ICCV'2017) | Diving48 (Homepage) (ECCV'2018) | Jester (Homepage) (ICCV'2019) |
| Moments in Time ( |
暂无开放 Issues,或尚未同步最近议题。