miller vs ndjson.github.io

Our great sponsors

WorkOS - The modern identity platform for B2B SaaS

InfluxDB - Power Real-Time Data Analytics at Scale

SaaSHub - Software Alternatives and Reviews

Our great sponsors

miller		ndjson.github.io
	Project
63	Mentions	17
8,553	Stars	23
-	Growth	-
9.1	Activity	0.0
6 days ago	Latest Commit	9 months ago
Go	Language	CSS
GNU General Public License v3.0 or later	License	-

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

miller

Posts with mentions or reviews of miller. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-12-22.

Qsv: Efficient CSV CLI Toolkit
8 projects | news.ycombinator.com | 22 Dec 2023
jq 1.7 Released
33 projects | news.ycombinator.com | 6 Sep 2023

jq and miller[1] are essential parts of my toolbelt, right up there with awk and vim.
[1]: https://github.com/johnkerl/miller
Perl first commit: a “replacement” for Awk and sed
3 projects | news.ycombinator.com | 8 Jul 2023

> This works really well if your problem can be solved in one or two liners.
My personal comfort threshold is around the 100-line mark. It's even possible to write maintainable shell scripts up to 500 lines, but it mostly depends on the problem you're trying to solve, and the discipline of the programmer to follow best practices (use sane defaults, ShellCheck, etc.).
> It go bad very quickly when, say, you have two CSV files and want to join them the sql-way.
In that case we're talking about structured data, and, yeah, Perl or Python would be easier to work with. That said, depending on the complexity of the CSV, you can still go a long way with plain Bash with IFS/read(1) or tr(1) to split CSV columns. This wouldn't be very robust, but there are tools that handle CSV specifically[1], which can be composed in a shell script just fine.
So it's always a balancing act of being productive quickly with a shell script, or reaching out for a programming language once the tools aren't a good fit, or maintenance becomes an issue.
[1]: https://miller.readthedocs.io/
Need help on cleaning this data!!
1 project | /r/datacleaning | 13 Jun 2023

where mlr is from https://github.com/johnkerl/miller
Running weekly average
1 project | /r/bash | 10 Jun 2023

if this class of problems (i.e., csv/tsv data) is your main target you may find miller (https://github.com/johnkerl/miller) much more useful in the long run
GQL: A new SQL like query language for .git files written in Rust
2 projects | /r/programming | 9 Jun 2023

That said, you may be interested in Miller (https://github.com/johnkerl/miller) which provides similar capabilities for CSV, JSON, and XML files. It doesn't use a SQL grammar, but that's just the proverbial lipstick on the thing. I'm not the author, but I have used it and I see some parallels in use cases at the very least.
johnkerl/miller: Miller is like awk, sed, cut, join, and sort for name-indexed data such as CSV, TSV, and tabular JSON
1 project | /r/devel | 8 Jun 2023
Any cli utility to create ascii/org mode tables?
3 projects | /r/commandline | 12 Apr 2023

worth giving Miller a shot
I wrote this iCalendar (.ics) command-line utility to turn common calendar exports into more broadly compatible CSV files.
6 projects | /r/commandline | 24 Mar 2023

CSV utilities (still haven't pick a favorite one...): https://github.com/harelba/q https://github.com/BurntSushi/xsv https://github.com/wireservice/csvkit https://github.com/johnkerl/miller
Miller: Like Awk, sed, cut, join, and sort for CSV, TSV, and tabular JSON
1 project | /r/hypeurls | 16 Mar 2023

ndjson.github.io

Posts with mentions or reviews of ndjson.github.io. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-04-11.

What the fuck
2 projects | /r/programminghorror | 11 Apr 2023

However, since every JSON document can be represented in a single line, something like newline-delimited JSON / JSON Lines feels like it would've been more suitable for that kind of data.
The XML spec is 25 years old today
1 project | news.ycombinator.com | 10 Feb 2023
Consider Using CSV
7 projects | news.ycombinator.com | 10 Dec 2022

No one uses that format for streamed json, see ndson and jsonl
http://ndjson.org/
The size complaint is overblown, as repeated fields are compressed away.
As other folks rightfully commented, csv is a mine field. One should assume every CSV file is broken in some way. They also don't enumerate any of the downsides of CSV.
What people should consider is using formats like Avro or Parquet that carry their schema with them so the data can be loaded and analyzed without have to manually deal with column meaning.
DevTool Intro: The Algolia CLI!
2 projects | dev.to | 15 Aug 2022

What is ndjson? Newline delimited JSON is the format the Algolia CLI reads from and writes to files. This means that any command that passes ndjson formatted data as output or accepts it as input can be piped together with an Algolia CLI command! We’ll see more of this in the next example
On read of JSON file it loads the entire JSON into memory.
1 project | /r/learnpython | 19 Jul 2022

You might consider using json-lines format (also known as newline-delimited JSON), in which each line is a separate JSON document so they can be loaded individually.
How to format it as json?
1 project | /r/golang | 27 Jun 2022

The format you're getting is known as Newline-Delimited JSON. Instead of trying to parse the whole input and pass that to the JSON Decoder, you can use something like bufio.Scanner to get and parse it line by line.
Arrow2 0.12.0 released - including almost complete support for Parquet
2 projects | /r/rust | 5 Jun 2022

This is in oposition to NDJSON, which allows to split records without deserializing JSON itself, via e.g. read_lines. fwiw CSV suffers from the same problem as JSON - generally not possible to break into records without deserializing. It is worse than NDJSON because the character \n may appear at any position within an item, thus forbidding read_lines.
Processing large JSON files in Python without running out of memory
1 project | /r/Python | 18 Mar 2022

I've always seen it referred to as ndjson
Speeding up Go's builtin JSON encoder up to 55% for large arrays of objects
2 projects | news.ycombinator.com | 3 Mar 2022

I think this would be fine, as long as the CSV layer was still parsable using the RFC 4180, then you could still use a normal CSV parser to parse the CSV layer and a normal JSON parser to parse the JSON layer. My worry with your example is that it is nether format, so it will need custom serialisation and deserialisation logic as it is essentially a bran new format.
https://datatracker.ietf.org/doc/html/rfc4180
If you’re looking for line-oriented JSON, another option would be ndjson: http://ndjson.org/
IETF should keep XMPP as IM standard, instead of Matrix
7 projects | news.ycombinator.com | 16 Jan 2022

What are some alternatives?

When comparing miller and ndjson.github.io you can also consider the following projects:

visidata - A terminal spreadsheet multitool for discovering and arranging data

ndjson - Streaming line delimited json parser + serializer

xsv - A fast CSV command line toolkit written in Rust.

flatten-tool - Tools for generating CSV and other flat versions of the structured data

jq - Command-line JSON processor [Moved to: https://github.com/jqlang/jq]

babashka - A Clojure babushka for the grey areas of Bash (native fast-starting Clojure scripting environment) [Moved to: https://github.com/babashka/babashka]

dasel - Select, put and delete data from JSON, TOML, YAML, XML and CSV files with a single tool. Supports conversion between formats and can be used as a Go package.

datasette - An open source multi-tool for exploring and publishing data

csvtk - A cross-platform, efficient and practical CSV/TSV toolkit in Golang

grop - helper script for the `gron | grep | gron -u` workflow

yq - yq is a portable command-line YAML, JSON, XML, CSV, TOML and properties processor

csv2sqlite

miller vs visidata ndjson.github.io vs ndjson miller vs xsv ndjson.github.io vs flatten-tool miller vs jq ndjson.github.io vs babashka miller vs dasel ndjson.github.io vs datasette miller vs csvtk ndjson.github.io vs grop miller vs yq ndjson.github.io vs csv2sqlite

Compare miller vs ndjson.github.io and see what are their differences.

miller

ndjson.github.io

miller

ndjson.github.io

What are some alternatives?