pysimdjson
conda
Our great sponsors
pysimdjson | conda | |
---|---|---|
6 | 30 | |
629 | 6,078 | |
- | 1.3% | |
5.3 | 9.8 | |
2 months ago | 1 day ago | |
Python | Python | |
GNU General Public License v3.0 or later | GNU General Public License v3.0 or later |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
pysimdjson
- Analyzing multi-gigabyte JSON files locally
-
I Use C When I Believe in Memory Safety
Its magic function wrapping comes at a cost, trading ease of use for runtime performance. When you have a single C++ function to call that will run for a "long" time, pybind all the way. But pysimdjson tends to call a single function very quickly, and the overhead of a single function call is orders of magnitude slower than with cython when being explit with types and signatures. Wrap a class in pybind11 and cython and compare the stack trace between the two, and the difference is startling.
-
Processing JSON 2.5x faster than simdjson with msgspec
simdjson
-
[package-find] lsp-bridge
You are aware of simdjson being available in python if you really need some json crunching, albeit json module in Python is implemented in C itself, so I don't think understand why do you think Python is slow there?
-
The fastest tool for querying large JSON files is written in Python (benchmark)
json: 113.79130696877837 ms
While `orjson`, is faster than `ujson`/`json` here, it's only ~6% faster (in this benchmark). `simdjson` and `msgspec` (my library, see https://jcristharif.com/msgspec/) are much faster due to them avoiding creating PyObjects for fields that are never used.
If spyql's query engine can determine the fields it will access statically before processing, you might find using `msgspec` for JSON gives a nice speedup (it'll also type check the JSON if you know the type of each field). If this information isn't known though, you may find using `pysimdjson` (https://pysimdjson.tkte.ch/) gives an easy speed boost, as it should be more of a drop-in for `orjson`.
-
How I cut GTA Online loading times by 70%
I don't think JSON is really the problem - parsing 10MB of JSON is not so slow. For example, using Python's json.load takes about 800ms for a 47MB file on my system, using something like simdjson cuts that down to ~70ms.
conda
-
How to Create Virtual Environments in Python
Python's venv module is officially recommended for creating virtual environments since Python 3.5 comes packaged with your Python installation. While there still are additional older tools available, such as conda and virtualenv, if you are new to virtual environments, it is best to use venv now.
- Why does creating my conda environment use so much memory?
- Installing Anaconda on ChromeOS using Linux
-
PSA: conda-libmamba-solver can cut two hours off of your Anaconda install, but has only 47 GitHub stars. It deserves more praise.
conda's dependency solver solves a harder problem than pip's. This quote alludes to it "Conda will never be as fast as pip, so long as we're doing real environment solves and pip satisfies itself only for the current operation." (from https://github.com/conda/conda/issues/7239). Thus mamba was created to improve performance and now conda is bringing in that performance boost.
- Is Anaconda still open source?
-
How to get the best Conda environment experience in Codespaces
The other challenge I ran into sometimes was that if I was running a lower memory/storage Codespace instance, when I tried to use Conda from the command line to modify environments, the process would be killed after a few seconds. This turns out to be related to some performance issues Conda has that make it consume a lot of memory when trying to work with the conda-forge installation channel. You can always then just increase the size of the Codespace your are working with (just go to your Codespaces list and use the triple dots to change the settings for a Codespace).
-
What is the status of Python 3.11?
It's worth noting that [ana]conda isn't even fully compatible yet with 3.11 (you can use it to create 3.11 environments--and you really should rather than waiting on relying on the system python--but conda itself can only run on 3.10.
-
Miniconda finally released for Python 3.10
It took some time but as great Christmas present Miniconda was finally released with Python 3.10!
-
TW: ZSH (and BASH?) does not show current working dir etc anymore
The September update broke it.
-
Python 3.11.0 is now available
According to this this issue is high on their priority list (whatever that means).
What are some alternatives?
orjson - Fast, correct Python JSON library supporting dataclasses, datetimes, and numpy
mamba - The Fast Cross-Platform Package Manager
cysimdjson - Very fast Python JSON parsing library
Poetry - Python packaging and dependency management made easy
ultrajson - Ultra fast JSON decoder and encoder written in C with Python bindings
miniforge - A conda-forge distribution.
Fast JSON schema for Python - Fast JSON schema validator for Python.
PDM - A modern Python package and dependency manager supporting the latest PEP standards
lupin is a Python JSON object mapper - Python document object mapper (load python object from JSON and vice-versa)
pip-tools - A set of tools to keep your pinned Python dependencies fresh.
PyValico - Small python wrapper around https://github.com/rustless/valico
pip - The Python package installer