Unavailable

This livestream is restricted

Already a member? Login with your membership email address

The Fifth Elephant 2026 Annual Conference

Built for humans. Now rebuilding for agents.

Tickets

Loading…

Rajath

Rajath

@rajath93

Lakshmi Narayana G

Lakshmi Narayana G

@lakshminarayanag

From Files to Catalogs

Submitted Jul 13, 2026

Workshop: From Files to Catalogs — Modern Data Foundations with Parquet, Iceberg, and Polaris

Workshop Overview

Modern data platforms are built in layers—from how data is stored, to how it is managed, and finally how different tools discover and access it. In this workshop, we’ll walk through that journey together.

We’ll start by looking at modern columnar file formats like Apache Parquet and Apache Arrow, understanding why they became the standard for analytics and how they differ from traditional row-based formats.

From there, we’ll explore open table formats such as Apache Iceberg and Delta Tables, and see how they solve challenges like ACID transactions, schema evolution, partition evolution, and time travel.

Finally, we’ll discuss the role of catalogs, why they matter, and how solutions like Apache Polaris, Hive Catalog, and Unity Catalog enable governance and interoperability across multiple compute engines.

The workshop combines concepts with practical demonstrations so participants understand not only what these technologies are, but also why they exist and how they work together in modern lakehouse architectures.


Key Takeaways

By the end of this workshop, participants will:

  • Understand the differences between row-based and columnar storage formats, and why columnar storage is preferred for analytical workloads.
  • Learn how Apache Parquet and Apache Arrow are designed for efficient storage and processing.
  • Understand why table formats like Apache Iceberg and Delta Tables exist and the challenges they solve.
  • Explore features such as:
    • ACID transactions
    • Schema evolution
    • Partition evolution
    • Time travel
  • Learn the purpose of catalogs and how they enable:
    • Metadata management
    • Governance
    • Interoperability across compute engines
  • Leave with a clear understanding of how files, table formats, and catalogs fit together in a modern lakehouse architecture.

Who Should Participate

This workshop is intended for:

  • Data Engineers
  • Analytics Engineers
  • Data Platform Engineers
  • Backend and Software Engineers working with data infrastructure
  • Database Engineers and Architects
  • Students and professionals interested in learning about modern data platforms

Background Knowledge Requirements

Required

  • Basic SQL
  • Familiarity with files and tabular data
  • Comfort using the command line

Nice to Have (Not Mandatory)

  • Exposure to Spark, Trino, DuckDB, or similar query engines
  • Basic understanding of data lakes or cloud object storage

No prior experience with table formats or catalogs is required.


Software Installation Prerequisites

Participants should install the following before the workshop:

  • Docker Desktop (or Docker Engine)

Detailed setup instructions and the workshop repository will be shared before the session.


Workshop Plan

1. Columnar File Formats

  • Row-based vs. columnar storage
  • Apache Parquet
  • Apache Arrow
  • Compression and encoding

2. Open Table Formats

  • Why managing data as files isn’t enough
  • Apache Iceberg
  • Delta Tables
  • ACID transactions
  • Snapshots
  • Schema evolution
  • Partition evolution
  • Time travel
  • Hands-on examples

3. Catalogs

  • Why catalogs are needed
  • Metadata management and governance
  • Apache Polaris
  • Hive Catalog
  • Unity Catalog
  • Bringing everything together with an end-to-end workflow

Wrap-up

We’ll conclude the workshop with:

  • Q&A session
  • Discussion on real-world adoption
  • Best practices for building modern lakehouse architectures

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Get your hybrid access ticket

Hosted by

Jumpstart better data engineering and AI futures

Supported by

Platinum Sponsor

Atlassian unleashes the potential of every team. Our agile & DevOps, IT service management and work management software helps teams organize, discuss, and compl

Platinum Sponsor

Sahaj is an artisanal technology services company crafting purpose-built AI and data-led solutions for businesses.

Gold Sponsor

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive

Gold Sponsor

Bronze Sponsor

Internet infrastructure APIs for IP geolocation and more

Bronze Sponsor

Open Source Analytical Database for the AI era.

Community sponsor

Real-time Observability & Governance layer for AI agents