Which lakehouse table format do you expect your organization will be using by the end of 2023?

This page summarizes the projects mentioned and recommended in the original post on /r/dataengineering

Our great sponsors
  • InfluxDB - Power Real-Time Data Analytics at Scale
  • WorkOS - The modern identity platform for B2B SaaS
  • SaaSHub - Software Alternatives and Reviews
  • kafka-delta-ingest

    A highly efficient daemon for streaming data from Kafka into Delta Lake

  • This independence from a catalog allows for path based reads and writes. This is handy when writing from Kafka directly to Delta Lake for the first layer of ingestion. You don’t need a catalog (or even Spark). https://github.com/delta-io/kafka-delta-ingest/tree/main/src

  • nessie

    Nessie: Transactional Catalog for Data Lakes with Git-like semantics

  • Project Nessie (https://projectnessie.org/) will be the catalog that eventually decouples Iceberg from Hive. At that point, I think it will be a no brainer to go Iceberg over Delta.

  • InfluxDB

    Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

    InfluxDB logo
NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a more popular project.

Suggest a related project

Related posts