dplyr
polars
Our great sponsors
dplyr | polars | |
---|---|---|
40 | 144 | |
4,652 | 26,043 | |
0.7% | 6.1% | |
7.4 | 10.0 | |
21 days ago | 4 days ago | |
R | Rust | |
GNU General Public License v3.0 or later | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
dplyr
-
Show HN: Open-source, browser-local data exploration using DuckDB-WASM and PRQL
That's great feedback, thanks!
This tool definitely comes from a place of personal need - beyond just handling large files, I've also never really gelled well with the Excel/Google Sheet model of changing data in place as if you were editing text. I'm a Data Scientist and always preferred the chained data transforms you see in things like dplyr (https://dplyr.tidyverse.org/) or Polars (https://pola.rs/) and I feel this tool maps very closely to the chained model.
Also, thank you for the feature requests! Those would all be very useful - we'll put them on the roadmap.
-
IS it possible for a R package to set an R option that only affects that package?
There's an example of how to use zzz.R with a .onload() function to set options in the dplyr code base: https://github.com/tidyverse/dplyr/blob/bbcfe99e29fe737d456b0d7adc33d3c445a32d9d/R/zzz.r
-
Calculation within a data table by calling on specific values in two columns
Look at the tidyverse, especially the case_when or mutate functions.
-
PSA: You don't need fancy stuff to do good work.
Before diving into advanced machine learning algorithms or statistical models, we need to start with the basics: collecting and organizing data. Fortunately, both Python and R offer a wealth of libraries that make it easy to collect data from a variety of sources, including web scraping, APIs, and reading from files. Key libraries in Python include requests, BeautifulSoup, and pandas, while R has httr, rvest, and dplyr.
-
Creating data frame
It looks like your syntax is wrong. I think you’re trying to calculate a new variables in your data frame, or alter an existing column in a data frame. Have a look at the select() function in this reference for the proper syntax to use. https://dplyr.tidyverse.org/ Does that help?
-
I'm designing a shirt for a friend, it has 4 embroidered images of things they like/do. One thing is coding, they use R... I'm wondering two things. 1) What's a good image or piece of code or something that I should use? and 2) should I even add it to the design the shirt?
A lot of populat libraries have their own logos. Maybe one of them would be good. Check out dplyr for example: https://dplyr.tidyverse.org/
-
Anyone use Python for statistics, particularly DOE or QA/QC? What are your thoughts?
I hope you give it a try when you get a chance: https://dplyr.tidyverse.org/
-
Rstudio tidyverse help!
You can read up on the dplyr-verbs here, which I strongly suggest for your exam! In the code examples, you can simply click on any function you don't understand and it will take you directly to the documentation. Good Luck!
- Beginner question
- osdc-2023-assignment1
polars
-
Why Python's Integer Division Floors (2010)
This is because 0.1 is in actuality the floating point value value 0.1000000000000000055511151231257827021181583404541015625, and thus 1 divided by it is ever so slightly smaller than 10. Nevertheless, fpround(1 / fpround(1 / 10)) = 10 exactly.
I found out about this recently because in Polars I defined a // b for floats to be (a / b).floor(), which does return 10 for this computation. Since Python's correctly-rounded division is rather expensive, I chose to stick to this (more context: https://github.com/pola-rs/polars/issues/14596#issuecomment-...).
-
Polars
https://github.com/pola-rs/polars/releases/tag/py-0.19.0
-
Stuff I Learned during Hanukkah of Data 2023
That turned out to be related to pola-rs/polars#11912, and this linked comment provided a deceptively simple solution - use PARSE_DECLTYPES when creating the connection:
- Polars 0.20 Released
- Segunda linguagem
- Polars: Dataframes powered by a multithreaded query engine, written in Rust
- Summing columns in remote Parquet files using DuckDB
- Polars 0.34 is released. (A query engine focussing on DataFrame front ends)
What are some alternatives?
worldfootballR - A wrapper for extracting world football (soccer) data from FBref, Transfermark, Understat and fotmob
vaex - Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per second 🚀
Rustler - Safe Rust bridge for creating Erlang NIF functions
modin - Modin: Scale your Pandas workflows by changing a single line of code
ggplot2 - An implementation of the Grammar of Graphics in R
arrow-datafusion - Apache DataFusion SQL Query Engine
nx - Multi-dimensional arrays (tensors) and numerical definitions for Elixir
DataFrames.jl - In-memory tabular data in Julia
explorer - Series (one-dimensional) and dataframes (two-dimensional) for fast and elegant data exploration in Elixir
datatable - A Python package for manipulating 2-dimensional tabular data structures
rmarkdown - Dynamic Documents for R
Apache Arrow - Apache Arrow is a multi-language toolbox for accelerated data interchange and in-memory processing