stable-baselines3 vs tianshou

Our great sponsors

WorkOS - The modern identity platform for B2B SaaS

InfluxDB - Power Real-Time Data Analytics at Scale

SaaSHub - Software Alternatives and Reviews

Our great sponsors

stable-baselines3		tianshou
	Project
46	Mentions	8
7,850	Stars	7,356
4.0%	Growth	2.5%
8.2	Activity	9.5
8 days ago	Latest Commit	4 days ago
Python	Language	Python
MIT License	License	MIT License

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

stable-baselines3

Posts with mentions or reviews of stable-baselines3. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-12-09.

Sim-to-real RL pipeline for open-source wheeled bipeds
2 projects | /r/robotics | 9 Dec 2023

The latest release (v3.0.0) of Upkie's software brings a functional sim-to-real reinforcement learning pipeline based on Stable Baselines3, with standard sim-to-real tricks. The pipeline trains on the Gymnasium environments distributed in upkie.envs (setup: pip install upkie) and is implemented in the PPO balancer. Here is a policy running on an Upkie:
[P] PettingZoo 1.24.0 has been released (including Stable-Baselines3 tutorials)
4 projects | /r/reinforcementlearning | 24 Aug 2023

PettingZoo 1.24.0 is now live! This release includes Python 3.11 support, updated Chess and Hanabi environment versions, and many bugfixes, documentation updates and testing expansions. We are also very excited to announce 3 tutorials using Stable-Baselines3, and a full training script using CleanRL with TensorBoard and WandB.
SB3 - NotImplementedError: Box([-1. -1. -8.], [1. 1. 8.], (3,), <class 'numpy.float32'>) observation space is not supported
2 projects | /r/reinforcementlearning | 19 Jun 2023

Therefore, I debugged this error to the ReplayBuffer that was imported from `SB3`. This is the problem function -
Shimmy 1.0: Gymnasium & PettingZoo bindings for popular external RL environments
10 projects | /r/reinforcementlearning | 25 Apr 2023

Have you ever wanted to use dm-control with stable-baselines3? Within Reinforcement learning (RL), a number of APIs are used to implement environments, with limited ability to convert between them. This makes training agents across different APIs highly difficult, and has resulted in a fractured ecosystem.
Stable-Baselines3 v1.8 Release
2 projects | /r/reinforcementlearning | 12 Apr 2023

Changelog: https://github.com/DLR-RM/stable-baselines3/releases/tag/v1.8.0
[P] Reinforcement learning evolutionary hyperparameter optimization - 10x speed up
3 projects | /r/MachineLearning | 24 Mar 2023

Great project! One question though, is there any reason why you are not using existing RL models instead of creating your own, such as stable baselines?
Is Stable Baselines 3 no longer compatible with PettingZoo?
2 projects | /r/reinforcementlearning | 11 Jan 2023

I was able to get Stable Baselines 3 to work with gymnasium by following the details in this work-in-progress PR: https://github.com/DLR-RM/stable-baselines3/pull/780. I have not used PettingZoo, though.
[OC] Inteligência Artificial jogando Mega Man!
2 projects | /r/brdev | 23 Dec 2022
New to reinforcement learning.
3 projects | /r/reinforcementlearning | 7 Nov 2022

I'd say this is a great path but I'd also look at the basic on-policy gradient actor critic methods like A2C and eventually PPO. Someone recommended SAC which also really good. There are tons of environments in the https://github.com/Farama-Foundation/PettingZoo as well if you want to mess with those. You can also check out stable baselines https://github.com/DLR-RM/stable-baselines3 which is pretty popular. If you want to get into the theory more I recommend reading the Sutton and Barto book on reinforcement learning.
How to proceed further? (Learning RL)
3 projects | /r/reinforcementlearning | 3 Oct 2022

If you want to iterate quickly through different RL methods then it's a good idea to use one of the RL libraries like stable baselines 3. Then you can dig further into the methods that work best for you. Coding RL methods from scratch is very time consuming and error prone even for experienced programmers.

tianshou

Posts with mentions or reviews of tianshou. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-02-02.

Multi-Agent Stable Baselines
2 projects | /r/reinforcementlearning | 2 Feb 2023

https://github.com/thu-ml/tianshou Imho there isn't a library that has it all, RLlib is quite good too, but I think that Tianshou is more similar to Pytorch and that helps to change the internals more intuitively and know what you are doing.
Question about the old policy and new policy in TRPO code
4 projects | /r/reinforcementlearning | 6 Jul 2022

Good point...I'll check in more detail when I get a chance later today! I would suggest looking at a more recent implementation like https://github.com/DLR-RM/stable-baselines3 or https://github.com/thu-ml/tianshou if you're trying to build. https://spinningup.openai.com/en/latest/algorithms/trpo.html is particularly good for understanding
Tensorflow vs PyTorch for A3C
4 projects | /r/reinforcementlearning | 17 Nov 2021

Do you absolutely need A3C? A2C has become more widely used (see, e.g., the comment in https://github.com/ikostrikov/pytorch-a3c, and the fact that both https://github.com/thu-ml/tianshou and https://github.com/facebookresearch/salina have A2C implementations, but no A3C at first glance).
Best PyTorch RL library for doing research
9 projects | /r/reinforcementlearning | 30 Apr 2021

I tried tianshou and thought it was well-designed for modularity, but it was early in development when I tried and missing some basic features

What are some alternatives?

When comparing stable-baselines3 and tianshou you can also consider the following projects:

Ray - Ray is a unified framework for scaling AI and Python applications. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

stable-baselines - A fork of OpenAI Baselines, implementations of reinforcement learning algorithms

Pytorch - Tensors and Dynamic neural networks in Python with strong GPU acceleration

cleanrl - High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)

Super-mario-bros-PPO-pytorch - Proximal Policy Optimization (PPO) algorithm for Super Mario Bros

ElegantRL - Massively Parallel Deep Reinforcement Learning. 🔥

SuperSuit - A collection of wrappers for Gymnasium and PettingZoo environments (being merged into gymnasium.wrappers and pettingzoo.wrappers

Tic-Tac-Toe-Gym - This is the Tic-Tac-Toe game made with Python using the PyGame library and the Gym library to implement the AI with Reinforcement Learning

rl-baselines3-zoo - A training framework for Stable Baselines3 reinforcement learning agents, with hyperparameter optimization and pre-trained agents included.

RL-Adventure - Pytorch Implementation of DQN / DDQN / Prioritized replay/ noisy networks/ distributional values/ Rainbow/ hierarchical RL

agents - TF-Agents: A reliable, scalable and easy to use TensorFlow library for Contextual Bandits and Reinforcement Learning.

stable-baselines3 vs Ray stable-baselines3 vs stable-baselines stable-baselines3 vs Pytorch stable-baselines3 vs cleanrl stable-baselines3 vs Super-mario-bros-PPO-pytorch stable-baselines3 vs ElegantRL stable-baselines3 vs SuperSuit stable-baselines3 vs Tic-Tac-Toe-Gym stable-baselines3 vs rl-baselines3-zoo tianshou vs cleanrl stable-baselines3 vs RL-Adventure stable-baselines3 vs agents

Compare stable-baselines3 vs tianshou and see what are their differences.

stable-baselines3

tianshou

stable-baselines3

tianshou

What are some alternatives?