json_benchmark
data-analysis
json_benchmark | data-analysis | |
---|---|---|
2 | 6 | |
26 | 48 | |
- | - | |
3.7 | 7.3 | |
over 1 year ago | over 1 year ago | |
Python | Jupyter Notebook | |
- | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
json_benchmark
-
Show HN: Up to 100x Faster FastAPI with simdjson and io_uring on Linux 5.19
If you're primarily targeting Python as an application layer, you may also want to check out my msgspec library[1]. All the perf benefits of e.g. yyjson, but with schema validation like pydantic. It regularly benchmarks[2] as the fastest JSON library for Python. Much of the overhead of decoding JSON -> Python comes from the python layer, and msgspec employs every trick I know to minimize that overhead.
[1]: https://github.com/jcrist/msgspec
[2]: https://github.com/TkTech/json_benchmark
-
Sunday Daily Thread: What's everyone working on this week?
- Adding nvme drive support to SMARTie, https://github.com/tktech/smartie, which is a pure-python cross-platform library for getting disk information like serial number, SMART attributes (like disk temperature) - json_benchmark, https://github.com/tktech/json_benchmark, which is a new benchmark and correctness test for the more modern Python JSON libraries - py_yyjson, https://github.com/tktech/py_yyjson, which is still a WIP and provides Python bindings to the yyjson library, which offers comparable speed to simdjson but more flexibility when parsing (comments, arbitrary sized numbers, Inf/Nan, etc) - And some fixes to https://github.com/TkTech/humanmark, which is a markdown library used to edit the README.md in json_benchmark above.
data-analysis
- Why a public database of hospital prices doesn't exist yet
-
Open Database of Hospital Prices
https://github.com/dolthub/data-analysis/tree/main/transpare...
-
Show HN: Up to 100x Faster FastAPI with simdjson and io_uring on Linux 5.19
Absolutely interested, on my end at least. I wrote this to manage the transparency in coverage files: https://github.com/dolthub/data-analysis/tree/main/transpare... but I'm always looking for better techniques.
Oh wow, I see you used it on those exact files. How about that.
- Healthcare datasets with multiple continuous variables
-
Beyond the trillion prices: pricing C-sections in America
Details: data repository, code repository, and notebook. The linked GitHub repo gives you the tools you need to reproduce this analysis or create your own.
- I wrote some tools to find the prices of C-sections in America. Context in README
What are some alternatives?
japronto - Screaming-fast Python 3.5+ HTTP toolkit integrated with pipelining HTTP server based on uvloop and picohttpparser.
synthea - Synthetic Patient Population Simulator
simdjson-go - Golang port of simdjson: parsing gigabytes of JSON per second
jsplit - A Go program to split large JSON files into many jsonl files
json-buffet
typedload - Python library to load dynamically typed data into statically typed data structures
is2 - embedded RESTy http(s) server library from Edgio
search-dw - search-dw is a Python utility to automate "search and download" via the command line. It might be useful if you need to download the results of a Google search for a certain type of topic at the same time