Modern data platforms are built in layers—from how data is stored, to how it is managed, and finally how different tools discover and access it. In this workshop, we’ll walk through that journey together.
We’ll start by looking at modern columnar file formats like Apache Parquet and Apache Arrow, understanding why they became the standard for analytics and how they differ from traditional row-based formats.
From there, we’ll explore open table formats such as Apache Iceberg and Delta Tables, and see how they solve challenges like ACID transactions, schema evolution, partition evolution, and time travel.
Finally, we’ll discuss the role of catalogs, why they matter, and how solutions like Apache Polaris, Hive Catalog, and Unity Catalog enable governance and interoperability across multiple compute engines.
The workshop combines concepts with practical demonstrations so participants understand not only what these technologies are, but also why they exist and how they work together in modern lakehouse architectures.
By the end of this workshop, participants will:
- Understand the differences between row-based and columnar storage formats, and why columnar storage is preferred for analytical workloads.
- Learn how Apache Parquet and Apache Arrow are designed for efficient storage and processing.
- Understand why table formats like Apache Iceberg and Delta Tables exist and the challenges they solve.
- Explore features such as:
- ACID transactions
- Schema evolution
- Partition evolution
- Time travel
- Learn the purpose of catalogs and how they enable:
- Metadata management
- Governance
- Interoperability across compute engines
- Leave with a clear understanding of how files, table formats, and catalogs fit together in a modern lakehouse architecture.
This workshop is intended for:
- Data Engineers
- Analytics Engineers
- Data Platform Engineers
- Backend and Software Engineers working with data infrastructure
- Database Engineers and Architects
- Students and professionals interested in learning about modern data platforms
- Basic SQL
- Familiarity with files and tabular data
- Comfort using the command line
- Exposure to Spark, Trino, DuckDB, or similar query engines
- Basic understanding of data lakes or cloud object storage
No prior experience with table formats or catalogs is required.
Participants should install the following before the workshop:
- Docker Desktop (or Docker Engine)
Detailed setup instructions and the workshop repository will be shared before the session.
- Row-based vs. columnar storage
- Apache Parquet
- Apache Arrow
- Compression and encoding
- Why managing data as files isn’t enough
- Apache Iceberg
- Delta Tables
- ACID transactions
- Snapshots
- Schema evolution
- Partition evolution
- Time travel
- Hands-on examples
- Why catalogs are needed
- Metadata management and governance
- Apache Polaris
- Hive Catalog
- Unity Catalog
- Bringing everything together with an end-to-end workflow
We’ll conclude the workshop with:
- Q&A session
- Discussion on real-world adoption
- Best practices for building modern lakehouse architectures
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}