goawk vs zsv

goawk

A POSIX-compliant AWK interpreter written in Go, with CSV support (by benhoyt)

zsv

zsv+lib: world's fastest (simd) CSV parser, bare metal or wasm, with an extensible CLI for SQL querying, format conversion and more (by liquidaty)

CSV JSON Simd SQL Parser Sqlite3 Tsv flatten Serialize Txt Fixed Markdown Compare Fast WASM web-assembly

Source Code

Suggest alternative

Edit details

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

goawk		zsv
	Project
19	Mentions	25
1,885	Stars	170
-	Growth	-
7.1	Activity	7.4
8 days ago	Latest Commit	11 days ago
Go	Language	C
MIT License	License	MIT License

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

goawk

Posts with mentions or reviews of goawk. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-06-29.

GoAWK, an Awk interpreter written in Go (2018)
1 project | news.ycombinator.com | 3 Feb 2024
The Awk Programming Language, Second Edition
18 projects | news.ycombinator.com | 29 Jun 2023

TIL: GoAWK [1] - A POSIX-compliant AWK interpreter written in Go, with CSV support.
[1]: https://github.com/benhoyt/goawk
Looking for a script for csv file
1 project | /r/awk | 20 Mar 2023
Anyone else doing compiler work in Golang?
10 projects | /r/golang | 28 Feb 2023

Another nice project that I have used from time to time (and a very good source for insight) is the awk interpreter written in go https://github.com/benhoyt/goawk
Tool to interact with CSV
9 projects | /r/commandline | 27 Feb 2023

No, I want exactly the opposite - it should be a , b,c as a single string field containing a literal comma, and c. For example, https://github.com/benhoyt/goawk has csv support. https://github.com/benhoyt/goawk/blob/master/docs/csv.md - more info.
Why does awk parse '1&&x=1' as '1&&(x=1)' not '(1&&x)=1' when '&&' is high precedence than '='?
1 project | /r/ProgrammingLanguages | 11 Feb 2023

I've had a go at solving this in this PR -- feedback welcome. I don't love it, but oh well, it solves the problem at hand. Your comment pointed me in the right direction, thanks again.
Looking for programming languages created with Go
23 projects | /r/golang | 6 Nov 2022

There are quite a few re-implementations of scripting languages like Lua in Go. I've written an AWK interpreter in Go.
Oracle DB support in Benthos
8 projects | /r/golang | 7 Oct 2022

github.com/benhoyt/goawk -> this library lets you embed an AWK runtime in your applications, very easy to use and useful for enabling some powerful scripting in things you build
Brian Kernighan adds Unicode support to Awk (May, 2022)
13 projects | news.ycombinator.com | 20 Aug 2022

Yes, that's right. With my simplistic UTF-8-based implementation it turned length() -- for example -- from O(1) to O(N), turning O(N) algorithms which use length() into O(N^2). See this issue: https://github.com/benhoyt/goawk/issues/93
Similar with substr() and other string functions, which when operating as bytes are O(1), but become O(N) when trying to count the number of codepoints as UTF-8.
GNU Gawk has a fancier approach, which stores strings as UTF-8 as long as it can, but converts to UTF-32 if it needs to (eg: the string is non-ASCII and you call substr).
It looks like Brian Kernighan's code has the same issue with length() and substr(). I'm going to try to email him about this, as I think it's kind of a performance blocker.
Ask HN: Is having a Personal blog/brand worth it for you?
7 projects | news.ycombinator.com | 18 Jul 2022

I'm not sure if it was via my personal website or just my GitHub profile, but I got my current job at Canonical due to the CTO there reaching out about my GoAWK project (https://github.com/benhoyt/goawk). I get regular recruitment emails because I have my CV/resume online: most of them are very low-effort, but 1 in 20 or something are interesting emails where the recruiter has actually looked at my website and will tailor it personally. I also just enjoy technical writing, and get joy out of sharing it on HN. So it's "worth it" for me.

zsv

Posts with mentions or reviews of zsv. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-03-18.

Analyzing multi-gigabyte JSON files locally
14 projects | news.ycombinator.com | 18 Mar 2023

If it could be tabular in nature, maybe convert to sqlite3 so you can make use of indexing, or CSV to make use of high-performance tools like xsv or zsv (the latter of which I'm an author).
https://github.com/BurntSushi/xsv
https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql...
Show HN: Up to 100x Faster FastAPI with simdjson and io_uring on Linux 5.19
20 projects | news.ycombinator.com | 6 Mar 2023

Parsing CSV doesn't have to be slow if you use something like xsv or zsv (https://github.com/liquidaty/zsv) (disclaimer: I'm an author). The speed of CSV parsers is fast enough that unless you are doing something ultra-trivial such as "count rows", your bottleneck will be elsewhere.
The benefits of CSV are:
- human readable
- does not need to be typed (sometimes, data in the raw such as date-formatted data is not amenable to typing without introducing a pre-processing layer that gets you further from the original data)
- accessible to anyone: you don't need to be a data person to dbl-click and open in Excel or similar
The main drawback is that if your data is already typed, CSV does not communicate what the type is. You can alleviate this through various approaches such as is described at https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql..., though I wouldn't disagree that if you can be assured that your starting data conforms to non-text data types, there are probably better formats than CSV.
The main benefit of Arrow, IMHO, is less as a format for transmitting / communicating but rather as a format for data at rest, that would benefit from having higher performance column-based read and compression
Yq is a portable yq: command-line YAML, JSON, XML, CSV and properties processor
11 projects | news.ycombinator.com | 4 Feb 2023
csvkit: Command-line tools for working with CSV
1 project | news.ycombinator.com | 20 Jan 2023

I wanted so much to use csvkit and all the features it had, but its horrendous performance made it unscalable and therefore the more I used it, the more technical debt I accumulated.
This was one of the reasons I wrote zsv (https://github.com/liquidaty/zsv). Maybe csvkit could incorporate the zsv engine and we could get the best of both worlds?
Examples (using majestic million csv):
---
Ask HN: Programs that saved you 100 hours? (2022 edition)
69 projects | news.ycombinator.com | 20 Dec 2022
Show HN: Split CSV into multiple files to avoid the Excel's 1M row limitation
2 projects | news.ycombinator.com | 17 Oct 2022

}
```
This of course assumes that each line is a single record, so you'll need some preprocessing if your CSV might contain embedded line-ends. For the preprocessing, you can use something like the `2tsv` command of https://github.com/liquidaty/zsv (disclaimer: I'm its author), which converts CSV to TSV and replaces newline with \n.
You can also use something like `xsv split` (see https://lib.rs/crates/xsv) which frankly is probably your best option as of today (though zsv will be getting its own shard command soon)
Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
9 projects | news.ycombinator.com | 24 Sep 2022
Ask HN: Best way to find help creating technical doc (open- or closed-source)?
1 project | news.ycombinator.com | 23 Sep 2022

Am looking for one-time help creating documentation (e.g. man pages, tutorials) for open source project (e.g. https://github.com/liquidaty/zsv) as well as product documentation for commercial products, but not enough need for a full-time job. Requires familiarity with, for lack of better term, data janitorial work, and preferably with methods of auto-generating documentation. Any suggestions as to forums or other ways to find folks who might fit the bill for ad-hoc or part-time work of this nature?
Q – Run SQL Directly on CSV or TSV Files
13 projects | news.ycombinator.com | 21 Sep 2022

Nice work. I am a fan of tools like this and look forward to giving this a try.
However, in my first attempted query (version 3.1.6 on MacOS), I ran into significant performance limitations and more importantly, it did not give correct output.
In particular, running on a narrow table with 1mm rows (the same one used in the xsv examples) using the command "select country, count() from worldcitiespop_mil.csv group by country" takes 12 seconds just to get an incorrect error 'no such column: country'.
using sqlite3, it takes two seconds or so to load, and less than a second to run, and gives me the correct result.
Using https://github.com/liquidaty/zsv (disclaimer, I'm one of its authors), I get the correct results in 0.95 seconds with the one-liner `zsv sql 'select country, count() from data group by country' worldcitiespop_mil.csv`.
I look forward to trying it again sometime soon
A Trillion Prices
5 projects | news.ycombinator.com | 6 Sep 2022

All this banter arguing over CSV, JSON, sqlite seems unnecessary when you can just push format X through a pipe and get whichever format Y you want back out: https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql...
(disclaimer: I'm one of the zsv authors)

What are some alternatives?

When comparing goawk and zsv you can also consider the following projects:

bytehound - A memory profiler for Linux.

visidata - A terminal spreadsheet multitool for discovering and arranging data

tsv-utils - eBay's TSV Utilities: Command line tools for large, tabular data files. Filtering, statistics, sampling, joins and more.

duckdb - DuckDB is an in-process SQL OLAP Database Management System

awka - Revive awka - Awk to C Compiler

lnav - Log file navigator

intellij-awk - The missing IntelliJ IDEA language support plugin for AWK

tumblelog - A static tumblelog generator available as both a Perl and Python version

ClickHouse - ClickHouse® is a free analytics DBMS for big data

awk - One true awk

nio - Low Overhead Numerical/Native IO library & tools

goawk vs bytehound zsv vs visidata goawk vs tsv-utils zsv vs duckdb goawk vs awka zsv vs lnav goawk vs intellij-awk zsv vs tsv-utils goawk vs tumblelog zsv vs ClickHouse goawk vs awk zsv vs nio

Compare goawk vs zsv and see what are their differences.

goawk

zsv

goawk

zsv

What are some alternatives?