Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
< Back to tools
C

CORL

> 编程语言
Open source

High-quality single-file implementations of SOTA Offline and Offline-to-Online RL algorithms: AWAC, BC, CQL, DT, EDAC, IQL, SAC-N, TD3+BC, LB-SAC, SPOT, Cal-QL,

1.4K stars0 likes0 views
WebsiteGitHub

About

High-quality single-file implementations of SOTA Offline and Offline-to-Online RL algorithms: AWAC, BC, CQL, DT, EDAC, IQL, SAC-N, TD3+BC, LB-SAC, SPOT, Cal-QL,

CORL (Clean Offline Reinforcement Learning)

CORL is an Offline Reinforcement Learning library that provides high-quality and easy-to-follow single-file implementations of SOTA ORL algorithms. Each implementation is backed by a research-friendly codebase, allowing you to run or tune thousands of experiments. Heavily inspired by cleanrl for online RL, check them out too!

  • Single-file implementation
  • Benchmarked Implementation for N algorithms
  • Weights and Biases integration

  • ⭐ If you're interested in discrete control, make sure to check out our new library — Katakomba. It provides both discrete control algorithms augmented with recurrence and an offline RL benchmark for the NetHack Learning environment.

Getting started

git clone https://github.com/tinkoff-ai/CORL.git && cd CORL
pip install -r requirements/requirements_dev.txt

# alternatively, you could use docker
docker build -t <image_name> .
docker run --gpus=all -it --rm --name <container_name> <image_name>

Algorithms Implemented

Algorithm Variants Implemented Wandb Report
Offline and Offline-to-Online
✅ Conservative Q-Learning for Offline Reinforcement Learning
(CQL)
offline/cql.py
finetune/cql.py
Offline

Offline-to-online
✅ Accelerating Online Reinforcement Learning with Offline Datasets
(AWAC)
offline/awac.py
finetune/awac.py
Offline

Offline-to-online
✅ Offline Reinforcement Learning with Implicit Q-Learning
(IQL)
offline/iql.py
finetune/iql.py
Offline

Offline-to-online
Offline-to-Online only
✅ Supported Policy Optimization for Offline Reinforcement Learning
(SPOT)
finetune/spot.py Offline-to-online
✅ Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
(Cal-QL)
finetune/cal_ql.py Offline-to-online
Offline only
✅ Behavioral Cloning
(BC)
offline/any_percent_bc.py Offline
✅ Behavioral Cloning-10%
(BC-10%)
offline/any_percent_bc.py Offline
✅ A Minimalist Approach to Offline Reinforcement Learning
(TD3+BC)
offline/td3_bc.py Offline
✅ Decision Transformer: Reinforcement Learning via Sequence Modeling
(DT)
offline/dt.py Offline
✅ Uncertainty-Based Offline Reinforcement Learning with Diversified Q-Ensemble
(SAC-N)
offline/sac_n.py Offline
✅ Uncertainty-Based Offline Reinforcement Learning with Diversified Q-Ensemble
(EDAC)
offline/edac.py Offline
✅ Revisiting the Minimalist Approach to Offline Reinforcement Learning
(ReBRAC)
offline/rebrac.py Offline
✅ Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size
(LB-SAC)
offline/lb_sac.py Offline Gym-MuJoCo

D4RL Benchmarks

You can check the links above for learning curves and details. Here, we report reproduced final and best scores. Note that they differ by a significant margin, and some papers may use different approaches, not making it always explicit which reporting methodology they chose. If you want to re-collect our results in a more structured/nuanced manner, see results.

Offline

Last Scores

Gym-MuJoCo
Task-Name BC 10% BC TD3+BC AWAC CQL IQL ReBRAC SAC-N EDAC DT
halfcheetah-medium-v2 42.40 ± 0.19 42.46 ± 0.70 48.10 ± 0.18 49.46 ± 0.62 47.04 ± 0.22 48.31 ± 0.22 64.04 ± 0.68 68.20 ± 1.28 67.70 ± 1.04 42.20 ± 0.26
halfcheetah-medium-replay-v2 35.66 ± 2.33 23.59 ± 6.95 44.84 ± 0.59 44.70 ± 0.69 45.04 ± 0.27 44.46 ± 0.22 51.18 ± 0.31 60.70 ± 1.01 62.06 ± 1.10 38.91 ± 0.50
halfcheetah-medium-expert-v2 55.95 ± 7.35 90.10 ± 2.45 90.78 ± 6.04 93.62 ± 0.41 95.63 ± 0.42 94.74 ± 0.52 103.80 ± 2.95 98.96 ± 9.31 104.76 ± 0.64 91.55 ± 0.95
hopper-medium-v2 53.51 ± 1.76 55.48 ± 7.30 60.37 ± 3.49 74.45 ± 9.14 59.08 ± 3.77 67.53 ± 3.78 102.29 ± 0.17 40.82 ± 9.91 101.70 ± 0.28 65.10 ± 1.61
hopper-medium-replay-v2 29.81 ± 2.07 70.42 ± 8.66 64.42 ± 21.52 96.39 ± 5.28 95.11 ± 5.27 97.43 ± 6.39 94.98 ± 6.53 100.33 ± 0.78 99.66 ± 0.81 81.77 ± 6.87
hopper-medium-expert-v2 52.30 ± 4.01 111.16 ± 1.03 101.17 ± 9.07 52.73 ± 37.47 99.26 ± 10.91 107.42 ± 7.80 109.45 ± 2.34 101.31 ± 11.63 105.19 ± 10.08 110.44 ± 0.33
walker2d-medium-v2 63.23 ± 16.24 67.34 ± 5.17 82.71 ± 4.78 66.53 ± 26.04 80.75 ± 3.28 80.91 ± 3.17 85.82 ± 0.77 87.47 ± 0.66 93.36 ± 1.38 67.63 ± 2.54
walker2d-medium-replay-v2 21.80 ± 10.15 54.35 ± 6.34 85.62 ± 4.01 82.20 ± 1.05 73.09 ± 13.22 82.15 ± 3.03 84.25 ± 2.25 78.99 ± 0.50 87.10 ± 2.78 59.86 ± 2.73
walker2d-medium-expert-v2 98.96 ± 15.98 108.70 ± 0.25 110.03 ± 0.36 49.41 ± 38.16 109.56 ± 0.39 111.72 ± 0.86 111.86 ± 0.43 114.93 ± 0.41 114.75 ± 0.74 107.11 ± 0.96
locomotion average 50.40 69.29 76.45 67.72 78.28 81.63 89.74 83.52 92.92 73.84
Maze2d
Task-Name BC 10% BC TD3+BC AWAC CQL IQL ReBRAC SAC-N EDAC DT
maze2d-umaze-v1 0.36 ± 8.69 12.18 ± 4.29 29.41 ± 12.31 82.67 ± 28.30 -8.90 ± 6.11 42.11 ± 0.58 106.87 ± 22.16 130.59 ± 16.52 95.26 ± 6.39 18.08 ± 25.42
maze2d-medium-v1 0.79 ± 3.25 14.25 ± 2.33 59.45 ± 36.25 52.88 ± 55.12 86.11 ± 9.68 34.85 ± 2.72 105.11 ± 31.67 88.61 ± 18.72 57.04 ± 3.45 31.71 ± 26.33
maze2d-large-v1 2.26 ± 4.39 11.32 ± 5.10 97.10 ± 25.41 209.13 ± 8.19 23.75 ± 36.70 61.72 ± 3.50 78.33 ± 61.77 204.76 ± 1.19 95.60 ± 22.92 35.66 ± 28.20
maze2d average 1.13 12.58 61.99 114.89 33.65 46.23 96.77 141.32 82.64 28.48
Antmaze
Task-Name BC 10% BC TD3+BC AWAC CQL IQL ReBRAC SAC-N EDAC DT
antmaze-umaze-v2 55.25 ± 4.15 65.75 ± 5.26 70.75 ± 39.18 57.75 ± 10.28 92.75 ± 1.92 77.00 ± 5.52 97.75 ± 1.48 0.00 ± 0.00 0.00 ± 0.00 57.00 ± 9.82
antmaze-umaze-diverse-v2 47.25 ± 4.09 44.00 ± 1.00 44.75 ± 11.61 58.00 ± 7.68 37.25 ± 3.70 54.25 ± 5.54 83.50 ± 7.02 0.00 ± 0.00 0.00 ± 0.00 51.75 ± 0.43
antmaze-medium-play-v2 0.00 ± 0.00 2.00 ± 0.71 0.25 ± 0.43 0.00 ± 0.00 65.75 ± 11.61 65.75 ± 11.71 89.50 ± 3.35 0.00 ± 0.00 0.00 ± 0.00 0.00 ± 0.00
antmaze-medium-diverse-v2 0.75 ± 0.83 5.75 ± 9.39 0.25 ± 0.43 0.00 ± 0.00 67.25 ± 3.56 73.75 ± 5.45 83.50 ± 8.20 0.00 ± 0.00 0.00 ± 0.00 0.00 ± 0.00
antmaze-large-play-v2 0.00 ± 0.00 0.00 ± 0.00 0.00 ± 0.00 0.00 ± 0.00 20.75 ± 7.26 42.00 ± 4.53 52.25 ± 29.01 0.00 ± 0.00 0.00 ± 0.00 0.00 ± 0.00
antmaze-large-diverse-v2 0.00 ± 0.00 0.75 ± 0.83 0.00 ± 0.00 0.00 ± 0.00 20.50 ± 13.24 30.25 ± 3.63 64.00 ± 5.43 0.00 ± 0.00 0.00 ± 0.00 0.00 ± 0.00
antmaze average 17.21 19.71 19.33 19.29 50.71 57.17 78.42 0.00 0.00 18.12
Adroit
Task-Name BC 10% BC TD3+BC AWAC CQL IQL ReBRAC SAC-N EDAC DT
pen-human-v1 71.03 ± 6.26 26.99 ± 9.60 -3.88 ± 0.21 81.12 ± 13.47 13.71 ± 16.98 78.49 ± 8.21 103.16 ± 8.49 6.86 ± 5.93 5.07 ± 6.16 67.68 ± 5.48
pen-cloned-v1 51.92 ± 15.15 46.67 ± 14.25 5.13 ± 5.28 89.56 ± 15.57 1.04 ± 6.62 83.42 ± 8.19 102.79 ± 7.84 31.35 ± 2.14 12.02 ± 1.75 64.43 ± 1.43
pen-expert-v1 109.65 ± 7.28 114.96 ± 2.96 122.53 ± 21.27 160.37 ± 1.21 -1.41 ± 2.34 128.05 ± 9.21 152.16 ± 6.33 87.11 ± 48.

Issues· 0 open

View all issuesOpen on GitHub

No open issues yet, or sync has not completed.

> Tags

Pythond4rlgymoffline-reinforcement-learningreinforcement-learning

No comments yet. Be the first to share.

> Details

PublishedAug 1, 2026
UpdatedSep 17, 2026
Category编程语言
PricingOpen source

> Related tools

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言