+1 (726) 227-3549

Managed Data Lake & Iceberg Solutions with Fivetran

Land data once as open tables, query it from anywhere

Managed Data Lake & Iceberg Solutions with Fivetran

Fivetran's Managed Data Lake Service writes your synced data to object storage as Apache Iceberg or Delta Lake tables, with Fivetran handling the parts that make self-built lakes painful: file compaction, snapshot expiry, schema evolution, and catalog registration. The result is data that lives in your own S3, ADLS, OneLake (Microsoft Fabric), or GCS buckets and can be queried by Snowflake, Databricks, BigQuery, Athena, DuckDB, or anything else that reads Iceberg.

For a growing number of our clients this is the right architecture. For others it is a fashionable way to make simple things harder. We help you decide, and then we build it.

When a lake destination makes sense

  • More than one query engine. If analytics runs in Snowflake but data science runs in Databricks, landing once as Iceberg removes a second copy and a second Fivetran destination.
  • Large, cold, or raw data. Event streams, logs, and raw history that are queried rarely are cheaper to hold in object storage than in warehouse storage, and MAR is unaffected by where the data lands.
  • Avoiding lock-in. Open table formats mean the warehouse becomes a compute choice rather than the place your data is trapped.
  • Regulatory or residency requirements that favour data staying in your own cloud account.

When your team has one warehouse, modest volumes, and no lake expertise, we will usually recommend staying warehouse-only. Simplicity has value.

What we do

Architecture and decision. A short assessment of your sources, query engines, and volumes, ending in a written recommendation: warehouse, lake, or a mix, and why.

Storage and catalog setup. Bucket layout, IAM, and encryption; catalog integration with AWS Glue, Databricks Unity Catalog, or Apache Polaris / Snowflake Open Catalog, so that tables written by Fivetran are discoverable by every engine you use.

Destination configuration. The Managed Data Lake destination in Fivetran, table format choice (Iceberg or Delta), and connection-by-connection decisions about what lands in the lake versus the warehouse.

Query layer. External Iceberg tables in Snowflake, Unity Catalog federation in Databricks, BigLake in BigQuery, and ad-hoc access from DuckDB for engineers.

Transformations. dbt or SQLMesh models reading Iceberg directly, so the lake is a modeled asset rather than a pile of files. See our transformations consulting.

Migration. For warehouse-only accounts, a phased plan: new sources land in the lake first, then cold history, then hot tables where the query engines are proven.

Cost profile

The lake service does not change how MAR is counted, but it usually changes what you pay for storage and compute. We model both before recommending a move, and we fold the result into our cost optimization work where relevant.

Deliverables

A written architecture, a working destination with at least one production source landing as Iceberg, catalog integration verified from each query engine, and a runbook covering access, maintenance, and how to add the next source.

Contact us to discuss whether a managed data lake fits your stack.

Planning a data lake? Let's talk.
Contact Us Now