Modern data platforms are built in layers — from how data is stored, to how it is managed, and finally how different tools discover and access it. In this workshop, we’ll walk through that journey together.
We’ll start by looking at modern columnar file formats like Apache Parquet and Apache Arrow, understanding why they became the standard for analytics and how they differ from traditional row-based formats.
From there, we’ll explore open table formats such as Apache Iceberg and Delta Tables, and see how they solve challenges like ACID transactions, schema evolution, partition evolution, and time travel.
Finally, we’ll discuss the role of catalogs, why they matter, and how solutions like Apache Polaris, Hive Catalog, and Unity Catalog enable governance and interoperability across multiple compute engines.
The workshop combines concepts with practical demonstrations so participants understand not only what these technologies are, but also why they exist and how they work together in modern lakehouse architectures.
Level: Beginner to Intermediate
Duration: 3 hours (with a break in-between)
This workshop is intended for:
- Data Engineers
- Analytics Engineers
- Data Platform Engineers
- Backend and Software Engineers working with data infrastructure
- Database Engineers and Architects
- Students and professionals interested in learning about modern data platforms
By the end of this workshop, participants will:
- Understand the differences between row-based and columnar storage formats, and why columnar storage is preferred for analytical workloads
- Learn how Apache Parquet and Apache Arrow are designed for efficient storage and processing
- Understand why table formats like Apache Iceberg and Delta Tables exist and the challenges they solve
- Explore features such as:
- ACID transactions
- Schema evolution
- Partition evolution
- Time travel
- Learn the purpose of catalogs and how they enable:
- Metadata management
- Governance
- Interoperability across compute engines
- Leave with a clear understanding of how files, table formats, and catalogs fit together in a modern lakehouse architecture
Required:
- Basic SQL
- Familiarity with files and tabular data
- Comfort using the command line
Nice to have (not mandatory):
- Exposure to Spark, e6data, Trino, DuckDB, or similar query engines
- Basic understanding of data lakes or cloud object storage
No prior experience with table formats or catalogs is required.
Participants should install the following before the workshop:
- Docker Desktop (or Docker Engine)
https://github.com/rajathc93/lakehouse-lab
Detailed setup instructions and the workshop repository will be shared before the session.
1. Columnar File Formats
- Row-based vs. columnar storage
- Apache Parquet
- Apache Arrow
- Compression and encoding
2. Open Table Formats
- Why managing data as files isn’t enough
- Apache Iceberg
- Delta Tables
- ACID transactions
- Snapshots
- Schema evolution
- Partition evolution
- Time travel
- Hands-on examples
3. Catalogs
- Why catalogs are needed
- Metadata management and governance
- Apache Polaris
- Hive Catalog
- Unity Catalog
- Bringing everything together with an end-to-end workflow
Wrap-up
- Q&A session
- Discussion on real-world adoption
- Best practices for building modern lakehouse architectures
Rajath Gowda and Lakshmi Narayana G are founding engineers at e6data.
This workshop is open to:
🎟️ Fifth Elephant community members — https://hasgeek.com/fifthelephant#memberships
🎟️ Ticket holders for The Fifth Elephant annual conference — https://hasgeek.com/fifthelephant/enterprise-ai-in-production-meetup#tickets
This workshop is open to 30 participants (in-person) & hybrid access for remote attendees. Seats for in-person participants will be available on first-come-first-served basis. 🎟️
☎️ Call: (91) 7676332020
📧 Email: info@hasgeek.com