seatunnel
Apache Hive
seatunnel | Apache Hive | |
---|---|---|
31 | 14 | |
7,388 | 5,335 | |
1.0% | 0.6% | |
9.8 | 9.6 | |
about 17 hours ago | about 15 hours ago | |
Java | Java | |
Apache License 2.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
seatunnel
- SeaTunnel – super high-performance, distributed data integration tool
- Apache SeaTunnel: Next-generation high-performance, distributed integration tool
- FLaNK Weekly 31 December 2023
-
Five Apache projects you probably didn't know about
Apache SeaTunnel is a data integration platform that offers the three pillars of data pipelines: sources, transforms, and sinks. It offers an abstract API over three possible engines: the Zeta engine from SeaTunnel or a wrapper around Apache Spark or Apache Flink. Be careful, as each engine comes with its own set of features.
-
SymmetricDS: Open-Source, cross platform database replication software
looks that way. there is an other project that does similar things Apache SeaTunnel: https://seatunnel.apache.org/
- Breakthrough in the book search field! Use Apache SeaTunnel to improve the efficiency of book title similarity search
-
Questions Regarding design DW
https://seatunnel.apache.org/ Might be an overkill though...
-
SeaTunnel Zeta engine, the first choice for massive data synchronization, is officially released!
See the specific Change log: https://github.com/apache/incubator-seatunnel/releases/tag/2.3.0
-
The Ultimate Beginner’s Guide to Open Source Contribution
Apache SeaTunnel (Incubating) SeaTunnel is a very easy-to-use ultra-high-performance distributed data integration platform that supports real-time synchronization of massive data. It can synchronize tens of billions of data stably and efficiently every day, and has been used in the production of nearly 100 companies. Official website https://seatunnel.apache.org/ GitHub projects https://github.com/apache/incubator-seatunnel
- Major Release! SeaTunnel 2.3.0-beta supports the self-innovate SeaTunnel Engine and more connectors!
Apache Hive
-
Apache Iceberg as storage for on-premise data store (cluster)
Trino or Hive for SQL querying. Get Trino/Hive to talk to Nessie.
-
In One Minute : Hadoop
Hive, A data warehouse infrastructure that provides data summarization and ad hoc querying.
- Visionary French entrepreneur, David Gurle, launches new venture – Hive
-
DeWitt Clause, or Can You Benchmark %DATABASE% and Get Away With It
Apache Drill, Druid, Flink, Hive, Kafka, Spark
-
Apache Spark, Hive, and Spring Boot — Testing Guide
In this article, I'm showing you how to create a Spring Boot app that loads data from Apache Hive via Apache Spark to the Aerospike Database. More than that, I'm giving you a recipe for writing integration tests for such scenarios that can be run either locally or during the CI pipeline execution. The code examples are taken from this repository.
- Apache Hive in the vein!
-
Jinja2 not formatting my text correctly. Any advice?
ListItem(name='Apache Hive', website='https://hive.apache.org/', category='Interactive Query', short_description='Apache Hive is a data warehouse software project built on top of Apache Hadoop for providing data query and analysis. Hive gives an SQL-like interface to query data stored in various databases and file systems that integrate with Hadoop.'),
-
Understanding SQL Dialects
Apache Hive takes in a specific SQL dialect and converts it to map-reduce.
-
The Data Engineer Roadmap 🗺
Apache Hive
-
Open Source SQL Parsers
Apache Calcite is a popular parser/optimizer that is used in popular databases and query engines like Apache Hive, BlazingSQL and many others.
What are some alternatives?
airbyte - The leading data integration platform for ETL / ELT data pipelines from APIs, databases & files to data warehouses, data lakes & data lakehouses. Both self-hosted and Cloud-hosted.
superset - Apache Superset is a Data Visualization and Data Exploration Platform
kestra - Infinitely scalable, event-driven, language-agnostic orchestration and scheduling platform to manage millions of workflows declaratively in code.
ObjectBox Java (Kotlin, Android) - Java and Android Database - fast and lightweight without any ORM
Leetcode - Solutions to LeetCode problems; updated daily. Subscribe to my YouTube channel for more.
HikariCP - 光 HikariCP・A solid, high-performance, JDBC connection pool at last.
hudi - Upserts, Deletes And Incremental Processing on Big Data.
Apache Phoenix - Apache Phoenix
com.openai.unity - A Non-Official OpenAI Rest Client for Unity (UPM)
Flyway - Flyway by Redgate • Database Migrations Made Easy.
apisix-helm-chart - Apache APISIX Helm Chart
Presto - The official home of the Presto distributed SQL query engine for big data