百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
C

cleanrl

> 编程语言
开源

深度强化学习算法的高质量单文件实现,具有研究友好功能(如 PPO、DQN、C51、DDPG、TD3、SAC、PPG)

10.2K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

深度强化学习算法的高质量单文件实现,具有研究友好功能(如 PPO、DQN、C51、DDPG、TD3、SAC、PPG)

CleanRL (Clean Implementation of RL Algorithms)

CleanRL is a Deep Reinforcement Learning library that provides high-quality single-file implementation with research-friendly features. The implementation is clean and simple, yet we can scale it to run thousands of experiments using AWS Batch. The highlight features of CleanRL are:

  • Single-file implementation
    • Every detail about an algorithm variant is put into a single standalone file.
    • For example, our ppo_atari.py only has 340 lines of code but contains all implementation details on how PPO works with Atari games, so it is a great reference implementation to read for folks who do not wish to read an entire modular library.
  • Benchmarked Implementation (7+ algorithms and 34+ games at https://benchmark.cleanrl.dev)
  • Tensorboard Logging
  • Local Reproducibility via Seeding
  • Videos of Gameplay Capturing
  • Experiment Management with Weights and Biases
  • Cloud Integration with docker and AWS

You can read more about CleanRL in our JMLR paper and documentation.

Notable CleanRL-related projects:

  • corl-team/CORL: Offline RL algorithm implemented in CleanRL style
  • pytorch-labs/LeanRL: Fast optimized PyTorch implementation of CleanRL RL algorithms using CUDAGraphs.

ℹ️ Support for Gymnasium: Farama-Foundation/Gymnasium is the next generation of openai/gym that will continue to be maintained and introduce new features. Please see their announcement for further detail. We are migrating to gymnasium and the progress can be tracked in vwxyzjn/cleanrl#277.

⚠️ NOTE: CleanRL is not a modular library and therefore it is not meant to be imported. At the cost of duplicate code, we make all implementation details of a DRL algorithm variant easy to understand, so CleanRL comes with its own pros and cons. You should consider using CleanRL if you want to 1) understand all implementation details of an algorithm's variant or 2) prototype advanced features that other modular DRL libraries do not support (CleanRL has minimal lines of code so it gives you great debugging experience and you don't have do a lot of subclassing like sometimes in modular DRL libraries).

Get started

Prerequisites:

  • Python >=3.7.1,<3.11
  • uv 0.7.9+

To run experiments locally, give the following a try:

git clone https://github.com/vwxyzjn/cleanrl.git && cd cleanrl
uv pip install .

# alternatively, you could use `uv venv` and do
# `python run cleanrl/ppo.py`
uv run python cleanrl/ppo.py \
    --seed 1 \
    --env-id CartPole-v0 \
    --total-timesteps 50000

# open another terminal and enter `cd cleanrl/cleanrl`
tensorboard --logdir runs

To use experiment tracking with wandb, run

wandb login # only required for the first time
uv run python cleanrl/ppo.py \
    --seed 1 \
    --env-id CartPole-v0 \
    --total-timesteps 50000 \
    --track \
    --wandb-project-name cleanrltest

If you are not using uv, you can install CleanRL with requirements.txt:

…

To run training scripts in other games:

…

You may also use a prebuilt development environment hosted in Gitpod:

Algorithms Implemented

Algorithm Variants Implemented
✅ Proximal Policy Optimization (PPO) ppo.py, docs
ppo_atari.py, docs
ppo_continuous_action.py, docs
ppo_atari_lstm.py, docs
ppo_atari_envpool.py, docs
ppo_atari_envpool_xla_jax.py, docs
ppo_atari_envpool_xla_jax_scan.py, docs)
ppo_procgen.py, docs
ppo_atari_multigpu.py, docs
ppo_pettingzoo_ma_atari.py, docs
ppo_continuous_action_isaacgym.py, docs
ppo_trxl.py, docs
✅ Deep Q-Learning (DQN) dqn.py, docs
dqn_atari.py, docs
dqn_jax.py, docs
dqn_atari_jax.py, docs
✅ Categorical DQN (C51) c51.py, docs
c51_atari.py, docs
c51_jax.py, docs
c51_atari_jax.py, docs
✅ Soft Actor-Critic (SAC) sac_continuous_action.py, docs
sac_atari.py, docs
✅ Deep Deterministic Policy Gradient (DDPG) ddpg_continuous_action.py, docs
ddpg_continuous_action_jax.py, docs
✅ Twin Delayed Deep Deterministic Policy Gradient (TD3) td3_continuous_action.py, docs
td3_continuous_action_jax.py, docs
✅ Phasic Policy Gradient (PPG) ppg_procgen.py, docs
✅ Random Network Distillation (RND) ppo_rnd_envpool.py, docs
✅ Qdagger qdagger_dqn_atari_impalacnn.py, docs
qdagger_dqn_atari_jax_impalacnn.py, docs

Open RL Benchmark

To make our experimental data transparent, CleanRL participates in a related project called Open RL Benchmark, which contains tracked experiments from popular DRL libraries such as ours, Stable-baselines3, openai/baselines, jaxrl, and others.

Check out https://benchmark.cleanrl.dev/ for a collection of Weights and Biases reports showcasing tracked DRL experiments. The reports are interactive, and researchers can easily query information such as GPU utilization and videos of an agent's gameplay that are normally hard to acquire in other RL benchmarks. In the future, Open RL Benchmark will likely provide an dataset API for researchers to easily access the data (see repo).

Support and get involved

We have a Discord Community for support. Feel free to ask questions. Posting in Github Issues and PRs are also welcome. Also our past video recordings are available at YouTube

Citing CleanRL

If you use CleanRL in your work, please cite our technical paper:

@article{huang2022cleanrl,
  author  = {Shengyi Huang and Rousslan Fernand Julien Dossa and Chang Ye and Jeff Braga and Dipam Chakraborty and Kinal Mehta and João G.M. Araújo},
  title   = {CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms},
  journal = {Journal of Machine Learning Research},
  year    = {2022},
  volume  = {23},
  number  = {274},
  pages   = {1--18},
  url     = {http://jmlr.org/papers/v23/21-1342.html}
}

Acknowledgement

CleanRL is a community-powered by project and our contributors run experiments on a variety of hardware.

  • We thank many contributors for

Issues· 0 开放

查看全部 Issues在 GitHub 打开

暂无开放 Issues,或尚未同步最近议题。

> 标签

Pythona2cactor-criticadvantage-actor-criticale

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言