SOLUTION · REPLICATE TO AZURE DATA LAKE

Real-time data replication to Azure Data Lake Storage Gen2

Committed changes from your systems of record land in your data lake as Parquet files continuously, instead of arriving in a nightly extract.

Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MongoDB, Azure Cosmos DB, and other heterogeneous sources with a dedicated agent per database. The Azure Data Lake Storage Gen2 target agent writes each entity into your container through the native Azure Storage SDK, after an initial snapshot and then continuously, as Parquet by default. Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

WHO THIS IS FOR

Teams that keep the Azure data lake current with the operational systems

  • Analytics engineers who build lake tables and Spark or SQL models on ADLS Gen2 and need every change tagged with its operation and transaction identifiers, so current-state tables can be rebuilt from the history
  • Data platform leads who feed one lake from several databases and want one Core Hub for every source, backed by best-in-class enterprise support, rated 4.9/5 by customers
  • Architects in Microsoft-centric estates who need a landing tier that uses Microsoft Entra ID service principals, hierarchical namespace storage, and Private Link where the network requires it
  • Engineers running copy pipelines or Kafka sinks into the lake who want every source delivered by one product, with snapshots and monitoring built in

THE PROBLEM

Lake files loaded on a schedule are always a step behind

Most data lakes on Azure are fed by scheduled copy pipelines: a job queries each database on a timer, writes files, and moves on. The lake reflects the business as of the last run, every run adds read load to production, and each source needs its own pipeline, schedule, and failure mode. Teams searching for real-time Azure Data Lake ingestion or a CDC pipeline into ADLS Gen2 usually weigh the native copy and orchestration services, managed ELT connectors such as Fivetran or Airbyte, a Debezium and Kafka Connect sink, or scripts they run themselves.

Gluesync addresses that with per-agent CDC into Azure Data Lake Storage Gen2. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the Azure target agent writes them to your container as files. The pipeline model, the snapshot, and the operations stay the same across Oracle, SQL Server, and Azure Cosmos DB.

HOW IT WORKS

How Gluesync writes to Azure Data Lake Storage Gen2

The write path: the Azure Storage SDK on the hierarchical namespace

The ADLS Gen2 agent writes through Microsoft's native Azure Storage SDK against the hierarchical namespace of your storage account, with Apache Arrow building the Parquet data. It is a target agent: it receives changes from any Gluesync source agent and uploads them as files into the container you name.

  • Optimized batches, never row by row: Core Hub hands the agent changes in batches, and each batch becomes a file in the container, so the lake grows in analytics-friendly files rather than one blob per change. A large insert batch can be split across several files that Spark reads in parallel.
  • Parquet first: files are written by MOLO17 ParquetKt with Snappy compression by default; ZSTD, GZIP, or uncompressed output is a per-entity choice.
  • Configurable file sizing: the Parquet file size threshold and the polling interval are configurable per entity, so you choose fresher files or fewer, larger ones.
  • JSON or CSV when a consumer needs it: Use JSON file or Use CSV file switches a single entity to that format while the rest stay on Parquet.

Snapshot, then a continuous trail of change files

  • Initial load: each entity's snapshot is exported into the hierarchical namespace under Gluesync's folder structure, built from the transaction type, the table, the year, the month, and a timestamp.
  • CDC files: once the snapshot is in place, changes stream into files whose _operation column marks each row I, U, or D, with the full row image on updates and deletes.
  • History you can replay: files are append-only, so Synapse, Fabric, or Databricks jobs merge them into current-state tables while the raw trail stays available for audit.
  • Resumable snapshots: an interrupted snapshot picks up from its last saved state when the entity starts again.

Lineage, manifests, and storage events

  • Transaction context in every file: _transaction_id, _timestamp, and the source checkpoint fields travel with the rows, so any change can be traced back to its source transaction.
  • Sidecar manifest: a JSON companion beside each Parquet file records its timestamp, row count, operation, schema version, and transaction identifiers. It is on by default.
  • Event-driven downstream jobs: each file and manifest is uploaded with a PutBlob operation, so Azure storage events on new blobs can start Data Factory pipelines or functions as data lands.
  • Governed on the way in: Custom Field Functions mask or reshape sensitive fields before they reach the lake, and Allowed Operations decides which operations each entity writes. See data transformation.

What your Azure admin sets up

The agent signs in as a Microsoft Entra ID service principal and writes to a storage account with the hierarchical namespace enabled. When the account sits behind Private Link, create both the blob and the dfs private endpoints, because the SDK uses both. Full fields are in the Azure Data Lake Storage Gen2 target setup guide ↗.

  1. Register a single-tenant application in Microsoft Entra ID, create a client secret, and record the application (client) ID, the directory (tenant) ID, and the secret.
  2. Create a storage account with the hierarchical namespace enabled, and a container for the lake, for example datalake.
  3. Grant the application Storage Blob Data Contributor on the account for production writes.
  4. Set POSIX ACLs on the container for the application, since role assignments alone do not reach file-level permissions.
  5. In Core Hub, enter https://<storage-account>.dfs.core.windows.net as the host, the client ID and client secret as the credentials, and the container as the database name. Over the REST API, the tenant ID goes in customHostCredentials.

Core Hub around the lake

A pipeline groups a source agent, the ADLS Gen2 agent, and the entities they replicate. Core Hub orchestrates every agent through its web UI and REST APIs, and the agents deploy with Docker, Docker Compose, or Kubernetes, on-premises, on Azure, or in any other cloud. The data lake solution covers the wider lake patterns.

The same capture agent that feeds the lake can also feed Azure SQL, Cosmos DB, or a warehouse in a second pipeline, so raw history and current state come from one change stream under one Core Hub.

Explore the general CDC streaming architecture →

WRITE OPTIONS

The Azure Data Lake Storage Gen2 target agent

Azure Data Lake Storage Gen2 has one Gluesync target agent. It writes grouped change files through the native Azure Storage SDK, Parquet by default, with JSON or CSV selectable per entity.

AgentWrite techniqueVersionsBest for
Azure Data Lake Storage Gen2 agent ↗ Azure Storage SDK on the hierarchical namespace, grouped change files, Apache Arrow for Parquet, JSON or CSV per entity All Azure regions, with the hierarchical namespace enabled on the storage account Microsoft-centric estates landing operational data in ADLS Gen2 for Synapse, Fabric, or Databricks, with Entra ID service principal authentication, Private Link networking, and any Gluesync source agent feeding it under the same Core Hub.

SOURCES AND TOPOLOGIES

Land changes from the databases you already run in ADLS Gen2

Any Gluesync source agent can write to Azure Data Lake Storage Gen2, each with its own native capture method. Open the integrations finder with Azure Data Lake pre-selected to see every source you can pair with it, from Azure Cosmos DB and Azure SQL to Oracle and IBM i.

Gluesync keeps pace with your change volume at any scale. File size threshold and polling interval are configurable, so each entity can favor freshness or fewer files; MOLO17 Professional Services can help lay out the container with your team.

  • Azure Cosmos DB to Azure Data Lake Storage Gen2 for analytics on operational documents: see Azure Cosmos DB CDC
  • SQL Server and Azure SQL to Azure Data Lake Storage Gen2 through Change Data Capture or Change Tracking: see SQL Server CDC
  • PostgreSQL to Azure Data Lake Storage Gen2 from the write-ahead log: see PostgreSQL CDC
  • Oracle to Azure Data Lake Storage Gen2 from the redo logs through LogMiner or XStream: see Oracle CDC

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among Azure lake ingestion approaches

ApproachWhat buyers usually getWhere Gluesync fits
Azure Data Factory copy pipelines Native orchestration with copy activities and triggers, usually on a schedule, with change detection designed into each pipeline Log-based agents that write changes continuously, and storage events from each upload that can still trigger your Data Factory jobs; see the data lake solution
Fivetran and Airbyte Managed and open-source ELT connectors that load cloud storage and warehouses, with sync frequency set per connector A dedicated capture agent per database, Parquet with transaction lineage by default, and MOLO17 enterprise support behind every pipeline
Debezium with a Kafka Connect Azure Blob sink Open-source capture into Kafka topics, then a sink connector writes blobs; you run Kafka, Connect, offsets, and the file layout Changes land in ADLS Gen2 with no Kafka cluster in the path, and Kafka remains available as a separate target; see the Debezium alternative
Qlik Replicate and Striim Commercial replication products with Azure storage among many targets Native capture agents for every source, including IBM i and SAP HANA, writing to the lake from one Core Hub
Custom scripts and Azure Functions Full control; your team owns the extract queries, file naming, retries, and the load each run puts on production Log-based capture, snapshot resume, and Core Hub monitoring with no loader code to maintain; read log-based CDC

FAQ

Azure Data Lake Storage Gen2 replication questions

What does the Azure Data Lake Storage Gen2 target do?

It writes replicated data into an ADLS Gen2 container through the native Azure Storage SDK. Each entity is snapshotted first, then changes are written continuously as grouped files, Parquet by default.

Which sources can replicate to Azure Data Lake Storage Gen2?

Any Gluesync source agent, including Oracle, SQL Server, PostgreSQL, MySQL, MongoDB, and Azure Cosmos DB. The integrations finder on our website lists every pairing.

How does the agent authenticate?

As a Microsoft Entra ID service principal, using the application client ID, client secret, and tenant ID. Private Link works when both the blob and dfs private endpoints exist.

What does the storage account need?

The hierarchical namespace enabled, which makes the account Data Lake Storage Gen2, plus Storage Blob Data Contributor and POSIX ACLs on the container for the service principal.

How are inserts, updates, and deletes recorded in the lake?

As rows in change files with an _operation value of I, U, or D. Updates and deletes carry the full row, and every file holds the transaction identifier and timestamp of each change.

Can new files trigger Azure Data Factory pipelines?

Yes. Gluesync uploads each data file and manifest with a PutBlob operation, so storage events for newly created blobs can start Data Factory pipelines or other downstream jobs.

Does Gluesync write one file per change?

No. Changes are grouped into batches and each batch becomes a file. Parquet output fills up to a file size threshold, and both that threshold and the polling interval are configurable per entity.

Which Azure regions are supported?

All Azure regions, as long as the storage account has the hierarchical namespace enabled.

Evaluate Gluesync with your Azure Data Lake Storage Gen2 account

Start a trial on your own Azure subscription, or talk to MOLO17 about your sources, container layout, and the downstream jobs that read the lake.