Transformer-deploy Alternatives

Similar projects and alternatives to transformer-deploy

stable-diffusion-webui

2,808 129,299 9.9 Python transformer-deploy VS stable-diffusion-webui

Stable Diffusion web UI
diffusers

266 22,429 9.9 Python transformer-deploy VS diffusers

🤗 Diffusers: State-of-the-art diffusion models for image and audio generation in PyTorch and FLAX.
WorkOS

workos.com sponsored

The modern identity platform for B2B SaaS. The APIs are flexible and easy-to-use, supporting authentication, user identity, and complex enterprise features like SSO and SCIM provisioning.
onnxruntime

54 12,656 10.0 C++ transformer-deploy VS onnxruntime

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
TensorRT

22 9,065 5.0 C++ transformer-deploy VS TensorRT

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
flash-attention

26 10,773 9.4 Python transformer-deploy VS flash-attention

Fast and memory-efficient exact attention
server

24 7,314 9.5 Python transformer-deploy VS server

The Triton Inference Server provides an optimized cloud and edge inferencing solution. (by triton-inference-server)
deepsparse

21 2,866 9.6 Python transformer-deploy VS deepsparse

Sparsity-aware deep learning inference runtime for CPUs
InfluxDB

www.influxdata.com sponsored

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.
BentoML

16 6,521 9.8 Python transformer-deploy VS BentoML

The most flexible way to serve AI/ML models in production - Build Model Inference Service, LLM APIs, Inference Graph/Pipelines, Compound AI systems, Multi-Modal, RAG as a Service, and more!
FasterTransformer

7 5,436 4.3 C++ transformer-deploy VS FasterTransformer

Transformer related optimization, including BERT, GPT
TensorRT

5 2,328 9.6 Python transformer-deploy VS TensorRT

PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT (by pytorch)
serve

11 3,949 9.6 Java transformer-deploy VS serve

Serve, optimize and scale PyTorch models in production (by pytorch)
optimum

8 2,132 9.5 Python transformer-deploy VS optimum

🚀 Accelerate training and inference of 🤗 Transformers and 🤗 Diffusers with easy to use hardware optimization tools
mmrazor

4 1,361 2.8 Python transformer-deploy VS mmrazor

OpenMMLab Model Compression Toolbox and Benchmark.
kernl

8 1,458 1.5 Jupyter Notebook transformer-deploy VS kernl

Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackable.
torch2trt

5 4,388 3.1 Python transformer-deploy VS torch2trt

An easy to use PyTorch to TensorRT converter
fastT5

5 540 0.0 Python transformer-deploy VS fastT5

⚡ boost inference speed of T5 models by 5x & reduce the model size by 3x.
OpenSeeFace

7 1,312 4.2 Python transformer-deploy VS OpenSeeFace

Robust realtime face and facial landmark tracking on CPU with Unity integration
openai-whisper-cpu

5 206 10.0 Jupyter Notebook transformer-deploy VS openai-whisper-cpu

Improving transcription performance of OpenAI Whisper for CPU based deployment
sparsednn

1 51 0.0 Python transformer-deploy VS sparsednn

Fast sparse deep learning on CPUs
parallelformers

3 748 0.0 Python transformer-deploy VS parallelformers

Parallelformers: An Efficient Model Parallelization Toolkit for Deployment
SaaSHub

www.saashub.com sponsored

SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better transformer-deploy alternative or higher similarity.

Suggest an alternative to transformer-deploy

transformer-deploy reviews and mentions

Posts with mentions or reviews of transformer-deploy. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2022-10-28.

[D] How to get the fastest PyTorch inference and what is the "best" model serving framework?
8 projects | /r/MachineLearning | 28 Oct 2022

For 2), I am aware of a few options. Triton inference server is an obvious one as is the ‘transformer-deploy’ version from LDS. My only reservation here is that they require the model compilation or are architecture specific. I am aware of others like Bento, Ray serving and TorchServe. Ideally I would have something that allows any (PyTorch model) to be used without the extra compilation effort (or at least optionally) and has some convenience things like ease of use, easy to deploy, easy to host multiple models and can perform some dynamic batching. Anyway, I am really interested to hear people's experience here as I know there are now quite a few options! Any help is appreciated! Disclaimer - I have no affiliation or are connected in any way with the libraries or companies listed here. These are just the ones I know of. Thanks in advance.
[P] Up to 12X faster GPU inference on Bert, T5 and other transformers with OpenAI Triton kernels
8 projects | /r/MachineLearning | 25 Oct 2022

We work for Lefebvre Sarrut, a leading European legal publisher. Several of our products include transformer models in latency sensitive scenarios (search, content recommendation). So far, ONNX Runtime and TensorRT served us well, and we learned interesting patterns along the way that we shared with the community through an open-source library called transformer-deploy. However, recent changes in our environment made our needs evolve:
Convert Pegasus model to ONNX [Discussion]
1 project | /r/MachineLearning | 23 Sep 2022

here you will find a notebook for T5 on GPU with some tricks to make it fast: https://github.com/ELS-RD/transformer-deploy/blob/main/demo/generative-model/t5.ipynb
[P] What we learned by benchmarking TorchDynamo (PyTorch team), ONNX Runtime and TensorRT on transformers model (inference)
1 project | /r/MachineLearning | 2 Aug 2022

Check the notebook https://github.com/ELS-RD/transformer-deploy/blob/main/demo/TorchDynamo/benchmark.ipynb for detailed results, but what we will keep in mind:
[P] What we learned by making T5-large 2X faster than Pytorch (and any autoregressive transformer)
6 projects | /r/MachineLearning | 23 May 2022

notebook: https://github.com/ELS-RD/transformer-deploy/blob/main/demo/generative-model/t5.ipynb (Onnx Runtime only)
[P] 4.5 times faster Hugging Face transformer inference by modifying some Python AST
4 projects | /r/MachineLearning | 28 Dec 2021

Regarding CPU inference, quantization is very easy, and supported by Transformer-deploy , however performance on transformer are very low outside corner cases (like no batch, very short sequence and distilled model), and last Intel generation CPU based instance like C6 or M6 on AWS are quite expensive compared to a cheap GPU like Nvidia T4, to say it otherwise, on transformer, until you are ok with slow inference and takes a small instance (for a PoC for instance), CPU inference is probably not a good idea.
[P] First ever tuto to perform *GPU* quantization on 🤗 Hugging Face transformer models -> 2X faster inference
1 project | /r/MachineLearning | 8 Dec 2021

The end to end tutorial: https://github.com/ELS-RD/transformer-deploy/blob/main/demo/quantization_end_to_end.ipynb
[P] Python library to optimize Hugging Face transformer for inference: < 0.5 ms latency / 2850 infer/sec
4 projects | /r/MachineLearning | 23 Nov 2021

Want to try it 👉 https://github.com/ELS-RD/transformer-deploy
A note from our sponsor - SaaSHub
www.saashub.com | 24 Apr 2024

SaaSHub helps you find the best software and product alternatives Learn more →