FasterTransformer vs DeepSpeed

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

FasterTransformer		DeepSpeed
	Project
7	Mentions	51
5,456	Stars	32,739
4.2%	Growth	3.8%
4.3	Activity	9.8
about 1 month ago	Latest Commit	2 days ago
C++	Language	Python
Apache License 2.0	License	Apache License 2.0

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

FasterTransformer

Posts with mentions or reviews of FasterTransformer. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-07-08.

Train Your AI Model Once and Deploy on Any Cloud
3 projects | news.ycombinator.com | 8 Jul 2023

https://docs.nvidia.com/ai-enterprise/overview/0.1.0/platfor...
RIVA: NVIDIA® Riva, a premium edition of NVIDIA AI Enterprise software, is a GPU-accelerated speech and translation AI SDK
FasterTransformer: https://github.com/NVIDIA/FasterTransformer an
Whether the ML computation engineering expertise will be valuable, is the question.
2 projects | /r/LanguageTechnology | 21 Apr 2023

There could be some spectrum of this expertise. For instance, https://github.com/NVIDIA/FasterTransformer, https://github.com/microsoft/DeepSpeed
Optimized implementation of training/fine-tuning of LLMs [D]
1 project | /r/MachineLearning | 5 Mar 2023

Have anyone tried to optimize the forward and backward using custom Cuda code or fused kernel to speed up the training time of current LLMs? I only have seen FasterTransformer ( NVIDIA/FasterTransformer) and other similar tools but they're only focusing on inference.
Exploring Ghostwriter, a GitHub Copilot alternative
3 projects | dev.to | 8 Nov 2022

Replit built Ghostwriter on the open source scene based on Salesforce’s Codegen, using Nvidia’s FasterTransformer and Triton server for highly optimized decoders, and the knowledge distillation process of the CodeGen model from two billion parameters to a faster model of one billion parameters.
Why are self attention not as deployment friendly?
2 projects | /r/deeplearning | 26 Jul 2022
[P] What we learned by making T5-large 2X faster than Pytorch (and any autoregressive transformer)
6 projects | /r/MachineLearning | 23 May 2022

Nvidia FasterTransformer is a mix of Pytorch and CUDA/C++ dedicated code. The performance boost is huge on T5, they report a 10X speedup like TensorRT. However, the speedup is computed on a translation task where sequences are 25 tokens long on average. In our experience, improvement on very short sequences tend to decrease by large margins on longer ones. Still we plan to dig deeper into this project as it implements very interesting ideas.
[P] Python library to optimize Hugging Face transformer for inference: < 0.5 ms latency / 2850 infer/sec
4 projects | /r/MachineLearning | 23 Nov 2021

On the other side of the spectrum, there is Nvidia demos (here or there) showing us how to build manually a full Transformer graph (operator by operator) in TensorRT to get best performance from their hardware. It’s out of reach for many NLP practitioners and it’s time consuming to debug/maintain/adapt to a slightly different architecture (I tried). Plus, there is a secret: the very optimized model only works for specific sequence lengths and batch sizes. Truth is that, so far (and it will improve soon), it’s mainly for MLPerf benchmark (the one used to compare DL hardware), marketing content, and very specialized engineers.

DeepSpeed

Posts with mentions or reviews of DeepSpeed. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-12-06.

Can we discuss MLOps, Deployment, Optimizations, and Speed?
7 projects | /r/LocalLLaMA | 6 Dec 2023

DeepSpeed can handle parallelism concerns, and even offload data/model to RAM, or even NVMe (!?) . I'm surprised I don't see this project used more.
[P][D] A100 is much slower than expected at low batch size for text generation
1 project | /r/MachineLearning | 5 Dec 2023
DeepSpeed-FastGen: High-Throughput for LLMs via MII and DeepSpeed-Inference
1 project | news.ycombinator.com | 4 Nov 2023
DeepSpeed-FastGen: High-Throughput Text Generation for LLMs
1 project | news.ycombinator.com | 3 Nov 2023
Why async gradient update doesn't get popular in LLM community?
3 projects | news.ycombinator.com | 10 Oct 2023
DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models (r/MachineLearning)
1 project | /r/datascienceproject | 29 Aug 2023
[P] DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
1 project | /r/MachineLearning | 28 Aug 2023
A comprehensive guide to running Llama 2 locally
19 projects | news.ycombinator.com | 25 Jul 2023

While on the surface, a 192GB Mac Studio seems like a great deal (it's not much more than a 48GB A6000!), there are several reasons why this might not be a good idea:
* I assume most people have never used llama.cpp Metal w/ large models. It will drop to CPU speeds whenever the context window is full: https://github.com/ggerganov/llama.cpp/issues/1730#issuecomm... - while sure this might be fixed in the future, it's been an issue since Metal support was added, and is a significant problem if you are actually trying to actually use it for inferencing. With 192GB of memory, you could probably run larger models w/o quantization, but I've never seen anyone post benchmarks of their experiences. Note that at that point, the limited memory bandwidth will be a big factor.
* If you are planning on using Apple Silicon for ML/training, I'd also be wary. There are multi-year long open bugs in PyTorch[1], and most major LLM libs like deepspeed, bitsandbytes, etc don't have Apple Silicon support[2][3].
You can see similar patterns w/ Stable Diffusion support [4][5] - support lagging by months, lots of problems and poor performance with inference, much less fine tuning. You can apply this to basically any ML application you want (srt, tts, video, etc)
Macs are fine to poke around with, but if you actually plan to do more than run a small LLM and say "neat", especially for a business, recommending a Mac for anyone getting started w/ ML workloads is a bad take. (In general, for anyone getting started, unless you're just burning budget, renting cloud GPU is going to be the best cost/perf, although on-prem/local obviously has other advantages.)
[1] https://github.com/pytorch/pytorch/issues?q=is%3Aissue+is%3A...
[2] https://github.com/microsoft/DeepSpeed/issues/1580
[3] https://github.com/TimDettmers/bitsandbytes/issues/485
[4] https://github.com/AUTOMATIC1111/stable-diffusion-webui/disc...
[5] https://forums.macrumors.com/threads/ai-generated-art-stable...
Microsoft Research proposes new framework, LongMem, allowing for unlimited context length along with reduced GPU memory usage and faster inference speed. Code will be open-sourced
4 projects | /r/LocalLLaMA | 13 Jun 2023

And https://github.com/microsoft/deepspeed
April 2023
40 projects | /r/dailyainews | 2 Jun 2023

DeepSpeed Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales (https://github.com/microsoft/DeepSpeed/tree/master/blogs/deepspeed-chat)

What are some alternatives?

When comparing FasterTransformer and DeepSpeed you can also consider the following projects:

TensorRT - NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

ColossalAI - Making large AI models cheaper, faster and more accessible

onnxruntime - ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator

Megatron-LM - Ongoing research training transformer models at scale

transformer-deploy - Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models 🚀

fairscale - PyTorch extensions for high performance and large scale training.

parallelformers - Parallelformers: An Efficient Model Parallelization Toolkit for Deployment

wenet - Production First and Production Ready End-to-End Speech Recognition Toolkit

accelerate - 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support

optimum - 🚀 Accelerate training and inference of 🤗 Transformers and 🤗 Diffusers with easy to use hardware optimization tools

fairseq - Facebook AI Research Sequence-to-Sequence Toolkit written in Python.

FasterTransformer vs TensorRT DeepSpeed vs ColossalAI FasterTransformer vs onnxruntime DeepSpeed vs Megatron-LM FasterTransformer vs transformer-deploy DeepSpeed vs fairscale FasterTransformer vs parallelformers DeepSpeed vs TensorRT FasterTransformer vs wenet DeepSpeed vs accelerate FasterTransformer vs optimum DeepSpeed vs fairseq

Compare FasterTransformer vs DeepSpeed and see what are their differences.

FasterTransformer

DeepSpeed

FasterTransformer

DeepSpeed

What are some alternatives?