SaaSHub helps you find the best software and product alternatives Learn more →
Xberg Alternatives
Similar projects and alternatives to xberg
-
-
SaaSHub
SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives
-
-
-
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
-
-
-
-
-
pandas-ai
Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.
-
-
-
code2prompt
A CLI tool to convert your codebase into a single LLM prompt with source tree, prompt templating, and token counting.
-
smolmodels
Discontinued ✨ build ml models in natural language and minimal code [GET https://api.github.com/repos/plexe-ai/smolmodels: 404 - Not Found // See: https://docs.github.com/rest/repos/repos#get-a-repository]
-
-
PyMuPDF
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
-
html-to-markdown
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR. (by xberg-io)
-
Meltano Singer SDK
Write 70% less code by using the SDK to build custom extractors and loaders that adhere to the Singer standard: https://sdk.meltano.com (by meltano)
-
-
-
xberg discussion
xberg reviews and mentions
-
Qwen3.8 is launching and going open-weight soon
An alternative to Docling is Xberg: https://github.com/xberg-io/xberg
"A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server."
Amongst other things, you can use it with Open Webui: https://docs.xberg.io/integrations/openwebui/#choosing-an-en...
fyi, they are releasing their v1.0.0 release shortly and until then you will need to a release candidate docker image tag: https://github.com/xberg-io/xberg/issues/1192
As a layman, I would say Xberg is better than Docling because its one engine for everything while the former feels like a bunch of scripts and glue logic.
-
Show HN: Kreuzberg Cloud – ultra fast content intelligence – in public beta
Hi HN!
I'm the maintainer of Kreuzberg, an open-source document intelligence library (https://github.com/kreuzberg-dev/kreuzberg). Some of you may have used it for RAG ingestion.
We're launching Kreuzberg Cloud, a SAAS API and a self-hosted system. It's in public beta, and I would like to invite you all to give it a try.
What out MVP offers: we offer very fast CPU optimized document and code intelligence. You can extract content from more than 90 document file formats and 300 code file formats into Markdown (or plaintext/djot), with additional features (same pricing tier) including chunking, embeddings, keyword extraction - and various types of intelligence.
The OSS library is used as the base engine of the cloud system. Our initial offering is $0.008/page, and you get the first 10K pages free, no card required.
We also offer our entire system for self-hosting - using helm charts. We are looking for design partners, so if thats relevant - shoot me a line.
-
Why most enterprise AI projects underperform
Do you want to find out where your AI projects are losing accuracy? The quickest test is to take some real content from your pipeline, run it through Kreuzberg, and compare the results to what your current pipeline produces. If there is only a small difference, the problem lies elsewhere. If there is a big difference, you have just found the most cost-effective fix in your stack.
-
Why AI Agents Need Structured Code Intelligence (And How to Stop Managing Parsers)
Kreuzberg Cloud will handle the infrastructure side entirely by spinning up, scaling, and managing both the document and code processing pipelines without requiring teams to run or maintain anything themselves. If you're building at scale and would rather not think about parser infrastructure at all, that's the path.
-
Beyond the Model: Why Document Intelligence Is the Next AI Infrastructure Layer
That's the AI infrastructure gap Kreuzberg Cloud fills. We’re launching very soon. Join the waitlist and follow along as we build it out, and join our Discord server to connect directly with the team.
-
Document Structure Extraction with Kreuzberg
We’re grateful to the Docling team at IBM for the truly great foundation they’ve provided. If you’re running Docling in production today, try Kreuzberg against it on your actual documents and let us know what you think.
- Kreuzbery – Fast RAG Pipeline
-
📰 All Data and AI Weekly #231-02March2026
Kreuzberg: A modern library for text extraction.
-
Building a RAG pipeline with Kreuzberg and LangChain
This is where Kreuzberg plays a central role, covering the entire early-stage data flow: document ingestion, text chunking, and embedding generation. A typical RAG pipeline can combine Kreuzberg for ingestion, chunking, and embeddings with LangChain as the orchestration layer, alongside a vector database and an LLM. While the architecture is fairly standard, the quality of the early steps determines everything that follows.
-
Kreuzberg v4.3.0 and benchmarks
Kreuzberg is an open-source (MIT license) polyglot document intelligence framework written in Rust, with bindings for Python, TypeScript/JavaScript (Node/Bun/WASM), PHP, Ruby, Java, C#, Golang and Elixir. It's also available as a docker image and standalone CLI tool you can install via homebrew.
-
A note from our sponsor - SaaSHub
www.saashub.com | 19 Aug 2026
Stats
xberg-io/xberg is an open source project licensed under MIT License which is an OSI approved license.
The primary programming language of xberg is Rust.