ndjson.github.io
cyanide
ndjson.github.io | cyanide | |
---|---|---|
17 | 9 | |
23 | 11 | |
- | - | |
0.0 | 3.2 | |
9 months ago | 11 months ago | |
CSS | Elixir | |
- | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
ndjson.github.io
-
What the fuck
However, since every JSON document can be represented in a single line, something like newline-delimited JSON / JSON Lines feels like it would've been more suitable for that kind of data.
- The XML spec is 25 years old today
-
Consider Using CSV
No one uses that format for streamed json, see ndson and jsonl
http://ndjson.org/
The size complaint is overblown, as repeated fields are compressed away.
As other folks rightfully commented, csv is a mine field. One should assume every CSV file is broken in some way. They also don't enumerate any of the downsides of CSV.
What people should consider is using formats like Avro or Parquet that carry their schema with them so the data can be loaded and analyzed without have to manually deal with column meaning.
-
DevTool Intro: The Algolia CLI!
What is ndjson? Newline delimited JSON is the format the Algolia CLI reads from and writes to files. This means that any command that passes ndjson formatted data as output or accepts it as input can be piped together with an Algolia CLI command! We’ll see more of this in the next example
-
On read of JSON file it loads the entire JSON into memory.
You might consider using json-lines format (also known as newline-delimited JSON), in which each line is a separate JSON document so they can be loaded individually.
-
How to format it as json?
The format you're getting is known as Newline-Delimited JSON. Instead of trying to parse the whole input and pass that to the JSON Decoder, you can use something like bufio.Scanner to get and parse it line by line.
-
Arrow2 0.12.0 released - including almost complete support for Parquet
This is in oposition to NDJSON, which allows to split records without deserializing JSON itself, via e.g. read_lines. fwiw CSV suffers from the same problem as JSON - generally not possible to break into records without deserializing. It is worse than NDJSON because the character \n may appear at any position within an item, thus forbidding read_lines.
-
Processing large JSON files in Python without running out of memory
I've always seen it referred to as ndjson
-
Speeding up Go's builtin JSON encoder up to 55% for large arrays of objects
I think this would be fine, as long as the CSV layer was still parsable using the RFC 4180, then you could still use a normal CSV parser to parse the CSV layer and a normal JSON parser to parse the JSON layer. My worry with your example is that it is nether format, so it will need custom serialisation and deserialisation logic as it is essentially a bran new format.
https://datatracker.ietf.org/doc/html/rfc4180
If you’re looking for line-oriented JSON, another option would be ndjson: http://ndjson.org/
- IETF should keep XMPP as IM standard, instead of Matrix
cyanide
-
Would you recommend JSON/CSV/Other for data storage in games?
No, it stands for "Binary JSON".
-
What is MongoDB ?
BSON specification
-
MUON: Compact and simple binary format, that uses gaps in Unicode encoding for markup
I recommend looking at https://ubjson.org and https://bsonspec.org , this will answer most of your questions.
-
I need a json file that includes all or most of the data types supported by MongoDB.
or https://bsonspec.org/
-
Minimizing the size of JSON by using CodingKey
BJSON Binary JSON, with a Swift Library here and should be one for your other end
-
Basics of MongoDB
Starting this tutorial we specified that data in MongoDB is stored in collections. We also specified that in MongoDB we use syntax similar to JSON. That syntax is called "Binary JSON" or BSON. BSON is similar to JSON; but it's more like an encoded serialization of JSON. We can find useful information in the BSON website.
-
It's Time to Retire the CSV
> I'm saying that when you decode an Avro document, the result that comes out (presuming you don't tell the Avro decoder anything special about custom types your runtime supports and how it should map them) is a JSON document.
Semantic point: it's not a "document".
There are tools which will decode Avro and output the data in JSON (typically using the JSON encoding of Avro: https://avro.apache.org/docs/current/spec.html#json_encoding), but the ADT that is created is by no means a JSON document. The ADT that is created has more complex semantics than JSON; JSON is not the canonical representation.
> By which I don't mean JSON-encoded text, but rather an in-memory ADT that has the exact set of types that exist in JSON, no more and no less.
Except Avro has data types that are not the exact set of types that exist in JSON. The first clue on this might be that the Avro spec includes mappings that list how primitive Avro types are mapped to JSON types.
> Or, to put that another way, Avro is a way to encode JSON-typed data, just as "JSON text", or https://bsonspec.org/, is a way to encode JSON-typed data
BSON, by design, was meant to be a more efficient way to encode JSON data, so yes, it is a way to encode JSON-typed data. Avro, however, was not defined as a way to encode JSON data. It was defined as a way to encode data (with a degree of specialization for the case of Hadoop sequence files, where you are generally storing a large number of small records in one file).
A simple counter example: Avro has a "float" type, which is a 32-bit IEEE 754 floating point number. Neither JSON nor BSON have that type.
Technically, JSON doesn't really have types, it has values, but even if you pretend that JavaScript's types are JSON's types, there's nothing "canonical" about JavaScript's types for Avro.
Yes, you can represent JSON data in Avro, and Avro in JSON, much as you can represent data in two different serialization formats. Avro's data model is very much defined independently of JSON's data model (as you'd expect).
-
Sending 😀 in Go
For me, this exploration started when I was attempting to improve handling of Unicode surrogate pair values in the MongoDB Go driver's Extended JSON unmarshaler. The Extended JSON format is an extension to the standard JSON format that adds type information and allows deterministic conversion to and from BSON.
What are some alternatives?
ndjson - Streaming line delimited json parser + serializer
BSONMap - Elixir package that applies a function to each document in a BSON file.
flatten-tool - Tools for generating CSV and other flat versions of the structured data
naya - A fast streaming JSON parser written in Python
miller - Miller is like awk, sed, cut, join, and sort for name-indexed data such as CSV, TSV, and tabular JSON
json - Strongly typed JSON library for Rust
babashka - A Clojure babushka for the grey areas of Bash (native fast-starting Clojure scripting environment) [Moved to: https://github.com/babashka/babashka]
json - JSON for Modern C++
datasette - An open source multi-tool for exploring and publishing data
csvz - The hot new standard in open databases
grop - helper script for the `gron | grep | gron -u` workflow
Mongoose - MongoDB object modeling designed to work in an asynchronous environment.