How Easy It Is to Re-use Old Pandas Code in Spark 3.2

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

db-benchmark

91 319 0.0 R

reproducible benchmark of database-like ops

It seems to me that the Spark model is much more sensible in terms of performance. In Spark, individual tasks are finally compiled into optimized Java code. As I understand Dusk works, a separate Python process is run for each data subset. So because of this architecture, Dusk is unlikely to ever get Spark performance. By the way, both systems build and optimize the operation graph. This is confirmed by benchmarks: https://h2oai.github.io/db-benchmark/

InfluxDB

www.influxdata.com sponsored

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a more popular project.

Suggest a related project

How to generate a great website and reference manual for your R package
1 project | dev.to | 10 Apr 2024
Array Languages: R vs. APL
1 project | news.ycombinator.com | 21 Mar 2024
Data.table: R's data.table package extends data.frame
1 project | news.ycombinator.com | 15 Mar 2024
Database-Like Ops Benchmark
1 project | news.ycombinator.com | 9 Mar 2024
Fable: Forecasting Models for Tidy Time Series
1 project | news.ycombinator.com | 3 Mar 2024

How Easy It Is to Re-use Old Pandas Code in Spark 3.2

This page summarizes the projects mentioned and recommended in the original post on /r/programming Post date: 4 Feb 2022

db-benchmark

InfluxDB

Related posts