Aria

Codebase for Aria - an Open Multimodal Native MoE (by rhymes-ai)

Aria Alternatives

Similar projects and alternatives to Aria

  1. moshi

    9 Aria VS moshi

    Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec. (by kyutai-labs)

  2. SaaSHub

    SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

    SaaSHub logo
  3. BayLing-Speech

    5 Aria VS BayLing-Speech

    LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.

  4. mini-omni

    3 Aria VS mini-omni

    open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.

  5. ichigo

    3 Aria VS ichigo

    Local realtime voice AI

  6. hertz-dev

    1 Aria VS hertz-dev

    first base model for full-duplex conversational audio

  7. smt

    1 Aria VS smt

    SMT: The Surrogate Modeling Toolbox

  8. LAVIS

    18 Aria VS LAVIS

    LAVIS - A One-stop Library for Language-Vision Intelligence

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better Aria alternative or higher similarity.

Aria discussion

Log in or Post with

Aria reviews and mentions

Posts with mentions or reviews of Aria. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-11-03.
  • Hertz-dev, the first open-source base model for conversational audio
    7 projects | news.ycombinator.com | 3 Nov 2024
    - [LLaMA-Omni](https://github.com/ictnlp/LLaMA-Omni) is a speech-language model built on Llama-3.1-8B-Instruct and trained using just 4 GPUs, offering low-latency, high-quality speech interactions and simultaneous generation of text and speech responses

    - [moshi](https://github.com/kyutai-labs/moshi) a speech-text foundation model that supports low-latency high-quality speech interactions and simultaneous generation of text responses, using Mimi, a SOTA streaming neural audio codec

    - [Mini-Omni](https://github.com/gpt-omni/mini-omni) a multimodal LLM based on Qwen2 offering real-time end-to-end speech input and streaming audio output conversational capabilities

    - [Aria](https://github.com/rhymes-ai/Aria) is a lightweight, multimodal native MoE model with 25B parameters and 3.9B activated parameters per token, offering state-of-the-art performance in multimodal, language, and coding tasks, with a long multimodal context window of 64K tokens and efficient encoding of visual input for fast inference and low fine-tuning cost

    - [Ichigo](https://github.com/homebrewltd/ichigo) an open research project extending a text-based LLM to have native listening ability, using an early fusion technique, with improved multiturn capabilities and refusal to process inaudible queries

  • Aria: Open Multimodal Native Moe
    1 project | news.ycombinator.com | 14 Oct 2024

Stats

Basic Aria repo stats
2
1,086
9.3
over 1 year ago

rhymes-ai/Aria is an open source project licensed under Apache License 2.0 which is an OSI approved license.

The primary programming language of Aria is Jupyter Notebook.


Sponsored
SaaSHub - Software Alternatives and Reviews
SaaSHub helps you find the best software and product alternatives
www.saashub.com

Did you know that Jupyter Notebook is
the 15th most popular programming language
based on number of references?