mergestat-lite VS octosql

Compare mergestat-lite vs octosql and see what are their differences.

mergestat-lite

Query git repositories with SQL. Generate reports, perform status checks, analyze codebases. 🔍 📊 (by mergestat)

octosql

OctoSQL is a query tool that allows you to join, analyse and transform data from multiple databases and file formats using SQL. (by cube2222)
InfluxDB - Power Real-Time Data Analytics at Scale
Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.
www.influxdata.com
featured
SaaSHub - Software Alternatives and Reviews
SaaSHub helps you find the best software and product alternatives
www.saashub.com
featured
mergestat-lite octosql
10 34
3,419 4,699
0.3% -
6.3 1.2
3 days ago 3 days ago
Go Go
MIT License Mozilla Public License 2.0
The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

mergestat-lite

Posts with mentions or reviews of mergestat-lite. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2022-09-12.
  • SQLite Doesn't Use Git
    2 projects | /r/programming | 12 Sep 2022
    You can query git with this: https://github.com/mergestat/mergestat if you like the idea.
  • A SQLite extension for reading large files line-by-line
    8 projects | news.ycombinator.com | 30 Jul 2022
    Hey, author here, happy to answer any questions! Also checkout this notebook for a deeper dive into sqlite-lines, along with a slick WASM demonstration and more thoughts on the codebase itself https://observablehq.com/@asg017/introducing-sqlite-lines

    I really dig SQLite, and I believe SQLite extensions will push it to another level. I rarely reach for Pandas or other "traditional" tools and query languages, and instead opt for plain ol' SQLite and other extensions. As a shameless plug, I recently started a blog series on SQLite and related tools and extensions if you want to learn more! Next week I'll be publishing more SQLite extensions for parsing HTML + making HTTP requests https://observablehq.com/@asg017/a-new-sqlite-blog-series

    A few other SQLite extensions:

    - xlite, for reading Excel files, in Rust https://github.com/x2bool/xlite

    - sqlean, several small SQLite extensions in C https://github.com/nalgeon/sqlean

    - mergestat, several SQLite extensions for developers (mainly Github's API) in Go https://github.com/mergestat/mergestat

  • Show HN: Contribution Graph as a Git Command
    3 projects | news.ycombinator.com | 27 May 2022
  • Exploring Git Repos With MergeStat 🔬
    2 projects | dev.to | 15 Feb 2022
    mergestat is an open-source tool that allows users to run SQL queries on the contents and history of git repositories.
  • The world of PostgreSQL wire compatibility
    3 projects | news.ycombinator.com | 10 Feb 2022
    Thanks for this write up! I've been really interested in postgres compatibility in the context of a tool I maintain (https://github.com/mergestat/mergestat) that uses SQLite. I've been looking for a way to expose the SQLite capabilities over a more commonly used wire-protocol like postgres (or mysql) so that existing BI and visualization tools can access the data.

    This project is an interesting one: https://github.com/dolthub/go-mysql-server that provides a MySQL interface (wire and SQL) to arbitrary "backends" implemented in go.

    It's really interesting how compatibility with existing protocols has become an important feature of new databases - there's so much existing tooling that already speaks postgres (or mysql), being able to leverage that is a huge advantage IMO

  • Go library for printing human readable, relative time differences 🕰️
    3 projects | /r/programming | 3 Feb 2022
    timediff is a Go package for printing human readable, relative time differences. Output is based on ranges defined in the Day.js JavaScript library, and can be customized if needed. It's currently used by the mergestat command-line interface.
  • Askgit: Command-line tool for running SQL queries on Git repositories
    1 project | /r/CKsTechNews | 27 Nov 2021
    1 project | news.ycombinator.com | 27 Nov 2021
  • Semantic Git Commit Messages
    2 projects | news.ycombinator.com | 27 Sep 2021
    Assuming committers adhere to it, there could be some interesting use cases when combined with a tool like AskGit (https://github.com/askgitdev/askgit) for understanding what "categories" of work is being done in a codebase.

    Maybe even what directories/files tend to see `fix` or `refactor` more frequently (signs of a poorly design or "hot" area?)

  • Git as a NoSql Database
    9 projects | news.ycombinator.com | 5 Apr 2021
    I've been very curious to explore this type of use case with askgit (https://github.com/augmentable-dev/askgit) which was designed for running simple "slice and dice" queries and aggregations on git history (and change stats) for basic analytical purposes. I've been curious about how this could be applied to a small text+git based "db". Say, for a regular json or CSV dumps.

    This also reminds me of Dolt: https://github.com/dolthub/dolt which I believe has been on HN a couple times

octosql

Posts with mentions or reviews of octosql. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-07-01.
  • Wazero: Zero dependency WebAssembly runtime written in Go
    12 projects | news.ycombinator.com | 1 Jul 2023
    Never got it to anything close to a finished state, instead moving on to doing the same prototype in llvm and then cranelift.

    That said, here's some of the wazero-based code on a branch - https://github.com/cube2222/octosql/tree/wasm-experiment/was...

    It really is just a very very basic prototype.

  • Analyzing multi-gigabyte JSON files locally
    14 projects | news.ycombinator.com | 18 Mar 2023
  • DuckDB: Querying JSON files as if they were tables
    9 projects | news.ycombinator.com | 3 Mar 2023
    This is really cool!

    With their Postgres scanner[0] you can now easily query multiple datasources using SQL and join between them (i.e. Postgres table with JSON file). Something I strived to build with OctoSQL[1] before.

    It's amazing to see how quickly DuckDB is adding new features.

    Not a huge fan of C++, which is right now used for authoring extensions, it'd be really cool if somebody implemented a Rust extension SDK, or even something like Steampipe[2] does for Postgres FDWs which would provide a shim for quickly implementing non-performance-sensitive extensions for various things.

    Godspeed!

    [0]: https://duckdb.org/2022/09/30/postgres-scanner.html

    [1]: https://github.com/cube2222/octosql

    [2]: https://steampipe.io

  • Show HN: ClickHouse-local – a small tool for serverless data analytics
    13 projects | news.ycombinator.com | 5 Jan 2023
    Congrats on the Show HN!

    It's great to see more tools in this area (querying data from various sources in-place) and the Lambda use case is a really cool idea!

    I've recently done a bunch of benchmarking, including ClickHouse Local and the usage was straightforward, with everything working as it's supposed to.

    Just to comment on the performance area though, one area I think ClickHouse could still possibly improve on - vs OctoSQL[0] at least - is that it seems like the JSON datasource is slower, especially if only a small part of the JSON objects is used. If only a single field of many is used, OctoSQL lazily parses only that field, and skips the others, which yields non-trivial performance gains on big JSON files with small queries.

    Basically, for a query like `SELECT COUNT(*), AVG(overall) FROM books.json` with the Amazon Review Dataset, OctoSQL is twice as fast (3s vs 6s). That's a minor thing though (OctoSQL will slow down for more complicated queries, while for ClickHouse decoding the input is and remains the bottleneck).

    [0]: https://github.com/cube2222/octosql

  • Steampipe – Select * from Cloud;
    13 projects | news.ycombinator.com | 30 Sep 2022
    To add somewhat of a counterpoint to the other response, I've tried the Steampipe CSV plugin and got 50x slower performance vs OctoSQL[0], which is itself 5x slower than something like DataFusion[1]. The CSV plugin doesn't contact any external API's so it should be a good benchmark of the plugin architecture, though it might just not be optimized yet.

    That said, I don't imagine this ever being a bottleneck for the main use case of Steampipe - in that case I think the APIs themselves will always be the limiting part. But it does - potentially - speak to what you can expect if you'd like to extend your usage of Steampipe to more than just DevOps data.

    [0]: https://github.com/cube2222/octosql

    [1]: https://github.com/apache/arrow-datafusion

    Disclaimer: author of OctoSQL

  • Go runtime: 4 years later
    11 projects | news.ycombinator.com | 26 Sep 2022
    Actually, folks just use gRPC or Yaegi in Go.

    See Terraform[0], Traefik[1], or OctoSQL[2].

    Although I agree plugins would be welcome, especially for performance reasons, though also to be able to compile and load go code into a running go process (JIT-ish).

    [0]: https://github.com/hashicorp/terraform

    [1]: https://github.com/traefik/traefik

    [2]: https://github.com/cube2222/octosql

    Disclaimer: author of OctoSQL

  • Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
    9 projects | news.ycombinator.com | 24 Sep 2022
  • Beginner interested in learning SQL. Have a few question that I wasn’t able to find on google.
    3 projects | /r/SQL | 6 Aug 2022
    Through more magic, you COULD of course use stuff like Spark, or easier with programs like TextQL, sq, OctoSQL.
  • How I Used DALL·E 2 to Generate The Logo for OctoSQL
    1 project | /r/programming | 2 Aug 2022
    The logo was created for OctoSQL and in the article you can find a lot of sample phrase-image combinations, as it describes the whole path (generation, variation, editing) I went down. Let me know what you think!
  • How I Used DALL·E 2 to Generate the Logo for OctoSQL
    3 projects | news.ycombinator.com | 2 Aug 2022
    Hey, author here, happy to answer any questions!

    The logo was created for OctoSQL[0] and in the article you can find a lot of sample phrase-image combinations, as it describes the whole path (generation, variation, editing) I went down. Let me know what you think!

    [0]:https://github.com/cube2222/octosql

What are some alternatives?

When comparing mergestat-lite and octosql you can also consider the following projects:

git-xargs - git-xargs is a command-line tool (CLI) for making updates across multiple Github repositories with a single command.

duckdb - DuckDB is an in-process SQL OLAP Database Management System

crux - General purpose bitemporal database for SQL, Datalog & graph queries. Backed by @juxt [Moved to: https://github.com/xtdb/xtdb]

q - q - Run SQL directly on delimited files and multi-file sqlite databases

flan - A tasty tool that lets you save, load and share postgres snapshots with ease

trdsql - CLI tool that can execute SQL queries on CSV, LTSV, JSON, YAML and TBLN. Can output to various formats.

sqlite-plus - The ultimate set of SQLite extensions

sqlitebrowser - Official home of the DB Browser for SQLite (DB4S) project. Previously known as "SQLite Database Browser" and "Database Browser for SQLite". Website at:

csv-sql - Command-line tool to load csv and excel (xlsx) files and run sql commands

sqlite-utils - Python CLI utility and library for manipulating SQLite databases

datasette-lite - Datasette running in your browser using WebAssembly and Pyodide

textql - Execute SQL against structured text like CSV or TSV