<!-- Generated from the rendered page by scripts/write-llm-mirrors.mjs. Do not edit by hand. -->
Canonical: https://molo17.com/solutions/replicate-to-azure-data-lake/
Markdown mirror: https://molo17.com/solutions/replicate-to-azure-data-lake/index.md
Title: Replicate to Azure Data Lake Gen2 in real time | MOLO17
Description: Real-time replication to Azure Data Lake Storage Gen2 from Oracle, SQL Server, PostgreSQL, MongoDB, and Azure Cosmos DB as Parquet files. Start a trial.

[Solutions](/solutions/)  Replicate to Azure Data Lake

 SOLUTION · REPLICATE TO AZURE DATA LAKE

# Real-time data replication to Azure Data Lake Storage Gen2

**Committed changes from your systems of record land in your data lake as Parquet files continuously, instead of arriving in a nightly extract.**

Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MongoDB, Azure Cosmos DB, and other [heterogeneous sources](/integrations/?target=Azure+Data+Lake+Storage+Gen2#integration-finder) with a dedicated agent per database. The Azure Data Lake Storage Gen2 target agent writes each entity into your container through the native Azure Storage SDK, after an initial snapshot and then continuously, as Parquet by default. Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

[Start a Gluesync trial](/get-gluesync/) [Talk to us](/contacts/)

WHO THIS IS FOR

## Teams that keep the Azure data lake current with the operational systems

-   Analytics engineers who build lake tables and Spark or SQL models on ADLS Gen2 and need every change tagged with its operation and transaction identifiers, so current-state tables can be rebuilt from the history
-   Data platform leads who feed one lake from several databases and want one Core Hub for every source, backed by [best-in-class enterprise support, rated 4.9/5 by customers](/support/#customer-ratings)
-   Architects in Microsoft-centric estates who need a landing tier that uses Microsoft Entra ID service principals, hierarchical namespace storage, and Private Link where the network requires it
-   Engineers running copy pipelines or Kafka sinks into the lake who want [every source](/integrations/?target=Azure+Data+Lake+Storage+Gen2#integration-finder) delivered by one product, with snapshots and monitoring built in

THE PROBLEM

## Lake files loaded on a schedule are always a step behind

Most data lakes on Azure are fed by scheduled copy pipelines: a job queries each database on a timer, writes files, and moves on. The lake reflects the business as of the last run, every run adds read load to production, and each source needs its own pipeline, schedule, and failure mode. Teams searching for real-time Azure Data Lake ingestion or a CDC pipeline into ADLS Gen2 usually weigh the native copy and orchestration services, managed ELT connectors such as Fivetran or Airbyte, a Debezium and Kafka Connect sink, or scripts they run themselves.

Gluesync addresses that with **per-agent CDC into Azure Data Lake Storage Gen2**. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the Azure target agent writes them to your container as files. The pipeline model, the snapshot, and the operations stay the same across Oracle, SQL Server, and Azure Cosmos DB.

HOW IT WORKS

## How Gluesync writes to Azure Data Lake Storage Gen2

### The write path: the Azure Storage SDK on the hierarchical namespace

The ADLS Gen2 agent writes through Microsoft's native Azure Storage SDK against the hierarchical namespace of your storage account, with Apache Arrow building the Parquet data. It is a target agent: it receives changes from any Gluesync source agent and uploads them as files into the container you name.

-   **Optimized batches, never row by row:** Core Hub hands the agent changes in batches, and each batch becomes a file in the container, so the lake grows in analytics-friendly files rather than one blob per change. A large insert batch can be split across several files that Spark reads in parallel.
-   **Parquet first:** files are written by MOLO17 ParquetKt with Snappy compression by default; ZSTD, GZIP, or uncompressed output is a per-entity choice.
-   **Configurable file sizing:** the Parquet file size threshold and the polling interval are configurable per entity, so you choose fresher files or fewer, larger ones.
-   **JSON or CSV when a consumer needs it:** `Use JSON file` or `Use CSV file` switches a single entity to that format while the rest stay on Parquet.

### Snapshot, then a continuous trail of change files

-   **Initial load:** each entity's snapshot is exported into the hierarchical namespace under Gluesync's folder structure, built from the transaction type, the table, the year, the month, and a timestamp.
-   **CDC files:** once the snapshot is in place, changes stream into files whose `_operation` column marks each row I, U, or D, with the full row image on updates and deletes.
-   **History you can replay:** files are append-only, so Synapse, Fabric, or Databricks jobs merge them into current-state tables while the raw trail stays available for audit.
-   **Resumable snapshots:** an interrupted snapshot picks up from its last saved state when the entity starts again.

### Lineage, manifests, and storage events

-   **Transaction context in every file:** `_transaction_id`, `_timestamp`, and the source checkpoint fields travel with the rows, so any change can be traced back to its source transaction.
-   **Sidecar manifest:** a JSON companion beside each Parquet file records its timestamp, row count, operation, schema version, and transaction identifiers. It is on by default.
-   **Event-driven downstream jobs:** each file and manifest is uploaded with a `PutBlob` operation, so Azure storage events on new blobs can start Data Factory pipelines or functions as data lands.
-   **Governed on the way in:** Custom Field Functions mask or reshape sensitive fields before they reach the lake, and Allowed Operations decides which operations each entity writes. See [data transformation](/data-transformation/).

### What your Azure admin sets up

The agent signs in as a Microsoft Entra ID service principal and writes to a storage account with the hierarchical namespace enabled. When the account sits behind Private Link, create both the `blob` and the `dfs` private endpoints, because the SDK uses both. Full fields are in the [Azure Data Lake Storage Gen2 target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/azure-datalake-gen2-target.html).

1.  Register a single-tenant application in Microsoft Entra ID, create a client secret, and record the application (client) ID, the directory (tenant) ID, and the secret.
2.  Create a storage account with the hierarchical namespace enabled, and a container for the lake, for example `datalake`.
3.  Grant the application Storage Blob Data Contributor on the account for production writes.
4.  Set POSIX ACLs on the container for the application, since role assignments alone do not reach file-level permissions.
5.  In Core Hub, enter `https://<storage-account>.dfs.core.windows.net` as the host, the client ID and client secret as the credentials, and the container as the database name. Over the REST API, the tenant ID goes in `customHostCredentials`.

### Core Hub around the lake

A pipeline groups a source agent, the ADLS Gen2 agent, and the entities they replicate. Core Hub orchestrates every agent through its web UI and REST APIs, and the agents deploy with Docker, Docker Compose, or Kubernetes, on-premises, on Azure, or in any other cloud. The [data lake solution](/solutions/data-lake/) covers the wider lake patterns.

The same capture agent that feeds the lake can also feed Azure SQL, Cosmos DB, or a warehouse in a second pipeline, so raw history and current state come from one change stream under one Core Hub.

[Explore the general CDC streaming architecture →](/solutions/cdc-streaming/)

WRITE OPTIONS

## The Azure Data Lake Storage Gen2 target agent

Azure Data Lake Storage Gen2 has one Gluesync target agent. It writes grouped change files through the native Azure Storage SDK, Parquet by default, with JSON or CSV selectable per entity.

| Agent | Write technique | Versions | Best for |
| --- | --- | --- | --- |
| [Azure Data Lake Storage Gen2 agent ↗](https://docs.molo17.com/gluesync/latest/agents/azure-datalake-gen2-intro.html) | Azure Storage SDK on the hierarchical namespace, grouped change files, Apache Arrow for Parquet, JSON or CSV per entity | All Azure regions, with the hierarchical namespace enabled on the storage account | Microsoft-centric estates landing operational data in ADLS Gen2 for Synapse, Fabric, or Databricks, with Entra ID service principal authentication, Private Link networking, and any Gluesync source agent feeding it under the same Core Hub. |

SOURCES AND TOPOLOGIES

## Land changes from the databases you already run in ADLS Gen2

Any Gluesync source agent can write to Azure Data Lake Storage Gen2, each with its own native capture method. Open the [integrations finder with Azure Data Lake pre-selected](/integrations/?target=Azure+Data+Lake+Storage+Gen2#integration-finder) to see every source you can pair with it, from Azure Cosmos DB and Azure SQL to Oracle and IBM i.

Gluesync keeps pace with your change volume at any scale. File size threshold and polling interval are configurable, so each entity can favor freshness or fewer files; [MOLO17 Professional Services](/solutions/professional-services/) can help lay out the container with your team.

-   Azure Cosmos DB to Azure Data Lake Storage Gen2 for analytics on operational documents: see [Azure Cosmos DB CDC](/solutions/cosmos-db-cdc/)
-   SQL Server and Azure SQL to Azure Data Lake Storage Gen2 through Change Data Capture or Change Tracking: see [SQL Server CDC](/solutions/sql-server-cdc/)
-   PostgreSQL to Azure Data Lake Storage Gen2 from the write-ahead log: see [PostgreSQL CDC](/solutions/postgresql-cdc/)
-   Oracle to Azure Data Lake Storage Gen2 from the redo logs through LogMiner or XStream: see [Oracle CDC](/solutions/oracle-cdc/)

FAIR, HIGH-LEVEL COMPARISON

## Where Gluesync fits among Azure lake ingestion approaches

| Approach | What buyers usually get | Where Gluesync fits |
| --- | --- | --- |
| Azure Data Factory copy pipelines | Native orchestration with copy activities and triggers, usually on a schedule, with change detection designed into each pipeline | Log-based agents that write changes continuously, and storage events from each upload that can still trigger your Data Factory jobs; see the [data lake solution](/solutions/data-lake/) |
| Fivetran and Airbyte | Managed and open-source ELT connectors that load cloud storage and warehouses, with sync frequency set per connector | A dedicated capture agent per database, Parquet with transaction lineage by default, and MOLO17 enterprise support behind every pipeline |
| Debezium with a Kafka Connect Azure Blob sink | Open-source capture into Kafka topics, then a sink connector writes blobs; you run Kafka, Connect, offsets, and the file layout | Changes land in ADLS Gen2 with no Kafka cluster in the path, and Kafka remains available as a separate target; see the [Debezium alternative](/solutions/debezium-alternative/) |
| Qlik Replicate and Striim | Commercial replication products with Azure storage among many targets | Native capture agents for every source, including IBM i and SAP HANA, writing to the lake from one Core Hub |
| Custom scripts and Azure Functions | Full control; your team owns the extract queries, file naming, retries, and the load each run puts on production | Log-based capture, snapshot resume, and Core Hub monitoring with no loader code to maintain; read [log-based CDC](/blog/real-time-replication-log-based-cdc/) |

RELATED CONTENT

## Azure Data Lake ingestion: research and implementation detail

-   [Log-based CDC explained](/blog/real-time-replication-log-based-cdc/)
-   [MOLO17 ParquetKt: the Parquet library behind lake files](/blog/molo17-parquetkt-kotlin-apache-parquet-library/)
-   [Mask data in flight before it lands in the lake](/blog/data-masking-transformation-edge-gluesync-field-functions/)
-   [Data lake solution: land operational changes in object storage](/solutions/data-lake/)
-   [Azure Cosmos DB CDC](/solutions/cosmos-db-cdc/)
-   [Azure Data Lake Storage Gen2 agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/azure-datalake-gen2-intro.html)
-   [Azure Data Lake Storage Gen2 target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/azure-datalake-gen2-target.html)
-   [Parquet files support ↗](https://docs.molo17.com/gluesync/latest/core-hub/parquet-files-support.html)

FAQ

## Azure Data Lake Storage Gen2 replication questions

What does the Azure Data Lake Storage Gen2 target do?

It writes replicated data into an ADLS Gen2 container through the native Azure Storage SDK. Each entity is snapshotted first, then changes are written continuously as grouped files, Parquet by default.

Which sources can replicate to Azure Data Lake Storage Gen2?

Any Gluesync source agent, including Oracle, SQL Server, PostgreSQL, MySQL, MongoDB, and Azure Cosmos DB. The integrations finder on our website lists every pairing.

How does the agent authenticate?

As a Microsoft Entra ID service principal, using the application client ID, client secret, and tenant ID. Private Link works when both the blob and dfs private endpoints exist.

What does the storage account need?

The hierarchical namespace enabled, which makes the account Data Lake Storage Gen2, plus Storage Blob Data Contributor and POSIX ACLs on the container for the service principal.

How are inserts, updates, and deletes recorded in the lake?

As rows in change files with an \_operation value of I, U, or D. Updates and deletes carry the full row, and every file holds the transaction identifier and timestamp of each change.

Can new files trigger Azure Data Factory pipelines?

Yes. Gluesync uploads each data file and manifest with a PutBlob operation, so storage events for newly created blobs can start Data Factory pipelines or other downstream jobs.

Does Gluesync write one file per change?

No. Changes are grouped into batches and each batch becomes a file. Parquet output fills up to a file size threshold, and both that threshold and the polling interval are configurable per entity.

Which Azure regions are supported?

All Azure regions, as long as the storage account has the hierarchical namespace enabled.

REPLICATE TO A TARGET

## Other targets Gluesync delivers to

-    [Replicate to Aerospike](/solutions/replicate-to-aerospike/)
-    [Replicate to DynamoDB](/solutions/replicate-to-dynamodb/)
-    [Replicate to Redshift](/solutions/replicate-to-redshift/)
-    [Replicate to Amazon S3 & S3-compatible](/solutions/replicate-to-amazon-s3/)
-    [Replicate to Cassandra](/solutions/replicate-to-cassandra/)
-    [Replicate to Kafka](/solutions/replicate-to-kafka/)
-    [Replicate to Cosmos DB](/solutions/replicate-to-cosmos-db/)
-    [Replicate to ClickHouse](/solutions/replicate-to-clickhouse/)
-    [Replicate to CockroachDB](/solutions/replicate-to-cockroachdb/)
-    [Replicate to Couchbase](/solutions/replicate-to-couchbase/)
-    [Replicate to file stores](/solutions/replicate-to-file-stores/)
-    [Replicate to BigQuery](/solutions/replicate-to-bigquery/)
-    [Replicate to Google Cloud Storage](/solutions/replicate-to-google-cloud-storage/)
-    [Replicate to Google Pub/Sub](/solutions/replicate-to-google-pubsub/)
-    [Replicate to GridGain](/solutions/replicate-to-gridgain/)
-    [Replicate to Db2](/solutions/replicate-to-db2/)
-    [Replicate to Informix](/solutions/replicate-to-informix/)
-    [Replicate to MariaDB](/solutions/replicate-to-mariadb/)
-    [Replicate to SQL Server](/solutions/replicate-to-sql-server/)
-    [Replicate to MongoDB](/solutions/replicate-to-mongodb/)
-    [Replicate to MySQL](/solutions/replicate-to-mysql/)
-    [Replicate to Oracle](/solutions/replicate-to-oracle/)
-    [Replicate to PostgreSQL](/solutions/replicate-to-postgresql/)
-    [Replicate to RavenDB](/solutions/replicate-to-ravendb/)
-    [Replicate to Redis](/solutions/replicate-to-redis/)
-    [Replicate to SAP ASE](/solutions/replicate-to-sap-ase/)
-    [Replicate to SAP HANA](/solutions/replicate-to-sap-hana/)
-    [Replicate to ScyllaDB](/solutions/replicate-to-scylladb/)
-    [Replicate to SingleStore](/solutions/replicate-to-singlestore/)
-    [Replicate to Snowflake](/solutions/replicate-to-snowflake/)
-    [Replicate to Solace PubSub+](/solutions/replicate-to-solace/)
-    [Replicate to Vertica](/solutions/replicate-to-vertica/)
-    [Replicate to YugabyteDB](/solutions/replicate-to-yugabytedb/)

## Evaluate Gluesync with your Azure Data Lake Storage Gen2 account

Start a trial on your own Azure subscription, or talk to MOLO17 about your sources, container layout, and the downstream jobs that read the lake.

[Start a Gluesync trial](/get-gluesync/) [Talk to MOLO17](/contacts/)
