Apache Hudi - Streaming Data Lake Platform

Apache Hudi (pronounced Hoodie) stands for Hadoop Upserts Deletes and Incrementals. Hudi manages the storage of large analytical datasets on DFS (Cloud stores, HDFS or any Hadoop FileSystem compatible storage). As an organization, Hudi can help you build an efficient data lake, solving some of the most complex, low-level storage management problems, while putting data into hands of your data analysts, engineers and scientists much quicker.

Features:

Upsert support with fast, pluggable indexing
Atomically publish data with rollback support
Snapshot isolation between writer & queries
Savepoints for data recovery
Manages file sizes, layout using statistics
Async compaction of row & columnar data
Timeline metadata to track lineage
Optimize data lake layout with clustering

https://hudi.apache.org/

https://github.com/apache/hudi

License:

Tech:

Tags:

Apache Hudi - Streaming Data Lake Platform

Newsletter

Related Projects

Suggested keywords:

Apache Hudi - Streaming Data Lake Platform

Newsletter

Related Projects