TensorRT vs ollama

TensorRT

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT. (by NVIDIA)

Source Code

developer.nvidia.com

Suggest alternative

Edit details

ollama

Get up and running with Llama 3, Mistral, Gemma, and other large language models. (by ollama)

Artificial intelligence

Source Code

ollama.com

Suggest alternative

Edit details

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

SaaSHub - Software Alternatives and Reviews

SaaSHub helps you find the best software and product alternatives

www.saashub.com

featured

TensorRT		ollama
	Project
22	Mentions	229
9,961	Stars	72,781
2.3%	Growth	14.0%
6.1	Activity	9.9
1 day ago	Latest Commit	4 days ago
C++	Language	Go
Apache License 2.0	License	MIT License

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

TensorRT

Posts with mentions or reviews of TensorRT. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-09-26.

AMD MI300X 30% higher performance than Nvidia H100, even with optimized stack
1 project | news.ycombinator.com | 17 Dec 2023

> It's not rocket science to implement matrix multiplication in any GPU.
You're right, it's harder. Saying this as someone who's done more work on the former than the latter. (I have, with a team, built a rocket engine. And not your school or backyard project size, but nozzle bigger than your face kind. I've also written CUDA kernels and boy is there a big learning curve to the latter that you gotta fundamentally rethink how you view a problem. It's unquestionable why CUDA devs are paid so much. Really it's only questionable why they aren't paid more)
I know it is easy to think this problem is easy, it really looks that way. But there's an incredible amount of optimization that goes into all of this and that's what's really hard. You aren't going to get away with just N for loops for a tensor rank N. You got to chop the data up, be intelligent about it, manage memory, how you load memory, handle many data types, take into consideration different results for different FMA operations, and a whole lot more. There's a whole lot of non-obvious things that result in high optimization (maybe obvious __after__ the fact, but that's not truthfully "obvious"). The thing is, the space is so well researched and implemented that you can't get away with naive implementations, you have to be on the bleeding edge.
Then you have to do that and make it reasonably usable for the programmer too, abstracting away all of that. Cuda also has a huge head start and momentum is not a force to be reckoned with (pun intended).
Look at TensorRT[0]. The software isn't even complete and it still isn't going to cover all neural networks on all GPUs. I've had stuff work on a V100 and H100 but not an A100, then later get fixed. They even have the "Apple Advantage" in that they have control of the hardware. I'm not certain AMD will have the same advantage. We talk a lot about the difficulties of being first mover, but I think we can also recognize that momentum is an advantage of being first mover. And it isn't one to scoff at.
[0] https://github.com/NVIDIA/TensorRT
Getting SDXL-turbo running with tensorRT
1 project | /r/StableDiffusion | 6 Dec 2023

(python demo_txt2img.py "a beautiful photograph of Mt. Fuji during cherry blossom"). https://github.com/NVIDIA/TensorRT/tree/release/8.6/demo/Diffusion
Show HN: Ollama for Linux – Run LLMs on Linux with GPU Acceleration
14 projects | news.ycombinator.com | 26 Sep 2023

- https://github.com/NVIDIA/TensorRT
TVM and other compiler-based approaches seem to really perform really well and make supporting different backends really easy. A good friend who's been in this space for a while told me llama.cpp is sort of a "hand crafted" version of what these compilers could output, which I think speaks to the craftmanship Georgi and the ggml team have put into llama.cpp, but also the opportunity to "compile" versions of llama.cpp for other model architectures or platforms.
Nvidia Introduces TensorRT-LLM for Accelerating LLM Inference on H100/A100 GPUs
3 projects | news.ycombinator.com | 8 Sep 2023

https://github.com/NVIDIA/TensorRT/issues/982
Maybe? Looks like tensorRT does work, but I couldn't find much.
Train Your AI Model Once and Deploy on Any Cloud
3 projects | news.ycombinator.com | 8 Jul 2023

highly optimized transformer-based encoder and decoder component, supported on pytorch, tensorflow and triton
TensorRT, custom ml framework/ inference runtime from nvidia, https://developer.nvidia.com/tensorrt, but you have to port your models
A1111 just added support for TensorRT for webui as an extension!
5 projects | /r/StableDiffusion | 27 May 2023
WIP - TensorRT accelerated stable diffusion img2img from mobile camera over webrtc + whisper speech to text. Interdimensional cable is here! Code: https://github.com/venetanji/videosd
3 projects | /r/StableDiffusion | 21 Feb 2023

It uses the nvidia demo code from: https://github.com/NVIDIA/TensorRT/tree/main/demo/Diffusion
[P] Get 2x Faster Transcriptions with OpenAI Whisper Large on Kernl
7 projects | /r/MachineLearning | 8 Feb 2023

The traditional way to deploy a model is to export it to Onnx, then to TensorRT plan format. Each step requires its own tooling, its own mental model, and may raise some issues. The most annoying thing is that you need Microsoft or Nvidia support to get the best performances, and sometimes model support takes time. For instance, T5, a model released in 2019, is not yet correctly supported on TensorRT, in particular K/V cache is missing (soon it will be according to TensorRT maintainers, but I wrote the very same thing almost 1 year ago and then 4 months ago so… I don’t know).
Speeding up T5
2 projects | /r/LanguageTechnology | 22 Jan 2023

I've tried to speed it up with TensorRT and followed this example: https://github.com/NVIDIA/TensorRT/blob/main/demo/HuggingFace/notebooks/t5.ipynb - it does give considerable speedup for batch-size=1 but it does not work with bigger batch sizes, which is useless as I can simply increase the batch-size of HuggingFace model.
demoDiffusion on TensorRT - supports 3090, 4090, and A100
1 project | /r/StableDiffusion | 10 Dec 2022

ollama

Posts with mentions or reviews of ollama. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-06-15.

Ollama v0.1.45
3 projects | news.ycombinator.com | 15 Jun 2024

I think the two main maintainers of Ollama have good intentions but suffer from a combination of being far too busy, juggling their forked llama.cpp server and not having enough automation/testing for PRs.
There is a new draft PR up to look at moving away from trying to juggle maintaining a llama.cpp fork to using llama.cpp with cgo bindings which I think will really help: https://github.com/ollama/ollama/pull/5034
SpringAI, llama3 and pgvector: bRAGging rights!
8 projects | dev.to | 15 Jun 2024

To support the exploration, I've developed a simple Retrieval Augmented Generation (RAG) workflow that works completely locally on the laptop for free. If you're interested, you can find the code itself here. Basically, I've used Testcontainers to create a Postgres database container with the pgvector extension to store text embeddings and an open source LLM with which I send requests to: Meta's llama3 through ollama.
RAG with OLLAMA
1 project | dev.to | 13 Jun 2024

Note: Before proceeding further you need to download and run Ollama, you can do so by clicking here.
Ollama 0.1.42
2 projects | news.ycombinator.com | 8 Jun 2024

`file://*` URLs are now allowed => ollama works with simple html files now
https://github.com/ollama/ollama/commit/1a29e9a879433fc55cf1...
How to setup a free, self-hosted AI model for use with VS Code
3 projects | dev.to | 4 Jun 2024

This guide assumes you have a supported NVIDIA GPU and have installed Ubuntu 22.04 on the machine that will host the ollama docker image. AMD is now supported with ollama but this guide does not cover this type of setup.
beginner guide to fully local RAG on entry-level machines
5 projects | dev.to | 2 Jun 2024

Nowadays, running powerful LLMs locally is ridiculously easy when using tools such as ollama. Just follow the installation instructions for your #OS. From now on, we'll assume using bash on Ubuntu.
Codestral: Mistral's Code Model
5 projects | news.ycombinator.com | 29 May 2024
AIM Weekly 27 May 2024
21 projects | dev.to | 28 May 2024
Devoxx Genie Plugin : an Update
6 projects | dev.to | 28 May 2024

I focused on supporting Ollama, GPT4All, and LMStudio, all of which run smoothly on a Mac computer. Many of these tools are user-friendly wrappers around Llama.cpp, allowing easy model downloads and providing a REST interface to query the available models. Last week, I also added "👋🏼 Jan" support because HuggingFace has endorsed this provider out-of-the-box.
Ask HN: Are companies self hosting LLMs?
1 project | news.ycombinator.com | 25 May 2024

What are some alternatives?

When comparing TensorRT and ollama you can also consider the following projects:

DeepSpeed - DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

llama.cpp - LLM inference in C/C++

FasterTransformer - Transformer related optimization, including BERT, GPT

gpt4all - gpt4all: run open-source LLMs anywhere

onnx-tensorrt - ONNX-TensorRT: TensorRT backend for ONNX

text-generation-webui - A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.

vllm - A high-throughput and memory-efficient inference and serving engine for LLMs

private-gpt - Interact with your documents using the power of GPT, 100% privately, no data leaks

openvino - OpenVINO™ is an open-source toolkit for optimizing and deploying AI inference

LocalAI - :robot: The free, Open Source OpenAI alternative. Self-hosted, community-driven and local-first. Drop-in replacement for OpenAI running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. It allows to generate Text, Audio, Video, Images. Also with voice cloning capabilities.

stable-diffusion-webui - Stable Diffusion web UI

llama - Inference code for Llama models

TensorRT vs DeepSpeed ollama vs llama.cpp TensorRT vs FasterTransformer ollama vs gpt4all TensorRT vs onnx-tensorrt ollama vs text-generation-webui TensorRT vs vllm ollama vs private-gpt TensorRT vs openvino ollama vs LocalAI TensorRT vs stable-diffusion-webui ollama vs llama

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

SaaSHub - Software Alternatives and Reviews

SaaSHub helps you find the best software and product alternatives

www.saashub.com

featured

Compare TensorRT vs ollama and see what are their differences.

TensorRT

ollama

TensorRT

ollama

What are some alternatives?