tooling-talks
Apache Spark
tooling-talks | Apache Spark | |
---|---|---|
9 | 101 | |
43 | 38,469 | |
- | 0.8% | |
5.0 | 10.0 | |
3 months ago | 4 days ago | |
Scala | Scala | |
- | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
tooling-talks
- Ask HN: What is your favorite Tech Podcasts these days?
- Tooling Talks Podcast – New Episode with Rebecca Mark Talking about Unison
- Tooling Talks Podcast - New episode with Rebecca Mark exploring Unison.
- Guillaume Martres: An Interactive Compiler -- New Tooling Talks episode out!
- New Tooling Talks episode with Eugune Yokota: Coding with Friends and sbt
-
Tooling Talks Podcast
Recently I got a lot of comments that Tooling Talks watcher weren't really watching the live stream, but were instead listening later on. So it seemed like a fitting choice to just migrate the entire thing to podcast form. Enjoy, and listen in for the next episode where I'll be chatting with Eugene Yokota!
-
Tooling Talks Episode 3 - Justin Kaeser
A big thanks to Justin for sitting down with me last weekend! For any of you that have been listening/watching, the future of Tooling Talks will change a bit as I migrate away from the live stream to a podcast. This next episode with Eugene Yokota will instead be podcasted. However, you can submit some questions that might be discussed here on GitHub.
-
Tooling Talks Episode 2 - Meriam Lachkar
Big thanks to Meriam for coming and hanging out for this episode. You can get all the info you want about Tooling Talkings here on GitHub.
-
Tooling Talks Episode 1 - Ólafur Páll Geirsson
Soon there will also be a full transcript here.
Apache Spark
- "xAI will open source Grok"
-
Groovy 🎷 Cheat Sheet - 01 Say "Hello" from Groovy
Recently I had to revisit the "JVM languages universe" again. Yes, language(s), plural! Java isn't the only language that uses the JVM. I previously used Scala, which is a JVM language, to use Apache Spark for Data Engineering workloads, but this is for another post 😉.
-
🦿🛴Smarcity garbage reporting automation w/ ollama
Consume data into third party software (then let Open Search or Apache Spark or Apache Pinot) for analysis/datascience, GIS systems (so you can put reports on a map) or any ticket management system
-
Go concurrency simplified. Part 4: Post office as a data pipeline
also, this knowledge applies to learning more about data engineering, as this field of software engineering relies heavily on the event-driven approach via tools like Spark, Flink, Kafka, etc.
-
Five Apache projects you probably didn't know about
Apache SeaTunnel is a data integration platform that offers the three pillars of data pipelines: sources, transforms, and sinks. It offers an abstract API over three possible engines: the Zeta engine from SeaTunnel or a wrapper around Apache Spark or Apache Flink. Be careful, as each engine comes with its own set of features.
-
Apache Spark VS quix-streams - a user suggested alternative
2 projects | 7 Dec 2023
-
Integrate Pyspark Structured Streaming with confluent-kafka
Apache Spark - https://spark.apache.org/
-
Spark – A micro framework for creating web applications in Kotlin and Java
A JVM based framework named "Spark", when https://spark.apache.org exists?
- Rest in Peas: The Unrecognized Death of Speech Recognition (2010)
-
PySpark SparkSession Builder with Kubernetes Master
I recently saw a pull request that was merged to the Apache/Spark repository that apparently adds initial Python bindings for PySpark on K8s. I posted a comment to the PR asking a question about how to use spark-on-k8s in a Python Jupyter notebook, and was told to ask my question here.
What are some alternatives?
kafka-manager - CMAK is a tool for managing Apache Kafka clusters
Trino - Official repository of Trino, the distributed SQL query engine for big data, former
Play - The Community Maintained High Velocity Web Framework For Java and Scala.
Pytorch - Tensors and Dynamic neural networks in Python with strong GPU acceleration
Lila - ♞ lichess.org: the forever free, adless and open source chess server ♞ [Moved to: https://github.com/lichess-org/lila]
Airflow - Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
learnxinyminutes-docs - Code documentation written as code! How novel and totally my idea!
Scalding - A Scala API for Cascading
PredictionIO - PredictionIO, a machine learning server for developers and ML engineers.
mrjob - Run MapReduce jobs on Hadoop or Amazon Web Services
scala - Scala 2 compiler and standard library. Bugs at https://github.com/scala/bug; Scala 3 at https://github.com/scala/scala3
luigi - Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.