<!-- Generated from the rendered page by scripts/write-llm-mirrors.mjs. Do not edit by hand. -->
Canonical: https://molo17.com/solutions/replicate-to-cosmos-db/
Markdown mirror: https://molo17.com/solutions/replicate-to-cosmos-db/index.md
Title: Replicate to Azure Cosmos DB in real time | MOLO17
Description: Real-time replication to Azure Cosmos DB from SQL Server, Oracle, PostgreSQL, and MongoDB, with batched Cosmos SDK writes and upsert by item id. Start a trial.

[Solutions](/solutions/)  Replicate to Cosmos DB

 SOLUTION · REPLICATE TO COSMOS DB

# Real-time data replication to Azure Cosmos DB from your operational databases

**Committed changes from SQL Server, Oracle, and your other systems of record land in Cosmos DB containers continuously, through the Cosmos DB Java SDK, instead of waiting for the next copy pipeline.**

Gluesync by MOLO17 captures changes from SQL Server, Oracle, PostgreSQL, MySQL, MongoDB, and other [heterogeneous sources](/integrations/?target=Azure+Cosmos+DB#integration-finder) with a dedicated agent per database, then writes them to Azure Cosmos DB through a target agent built on the Azure Cosmos DB Java SDK and the SQL API, over TLS. A snapshot seeds each container, CDC follows in real time, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

[Start a Gluesync trial](/get-gluesync/) [Talk to us](/contacts/)

WHO THIS IS FOR

## Teams that serve application data from Cosmos DB and need it to follow the source

-   Application engineers who read Cosmos DB items directly and need them to reflect current rows from SQL Server, Oracle, or PostgreSQL, without writing a consumer for each container
-   Azure platform leads who run Cosmos DB beside relational and document systems and want one Core Hub for every source, backed by [best-in-class enterprise support, rated 4.9/5 by customers](/support/#customer-ratings)
-   Architects modernizing a SQL Server or Oracle estate onto Cosmos DB, who need the snapshot, the CDC handoff, and the duplicate-id policy decided per entity before cutover
-   Engineers replacing a Kafka Connect Cosmos DB sink or nightly copy jobs, who want [every source](/integrations/?target=Azure+Cosmos+DB#integration-finder) delivered to Cosmos DB by one product, with snapshots, checkpoints, and monitoring built in

THE PROBLEM

## Cosmos DB containers loaded on a schedule trail the system of record

Cosmos DB containers that mirror operational data are usually loaded by scheduled copy jobs: a pipeline queries the source, writes documents, and reruns on a timer. Application reads trail the system of record between runs, every run adds load to the production database, and each source ends up with its own job, schedule, and failure mode. Teams looking for a Cosmos DB CDC pipeline or real-time Cosmos DB ingestion usually weigh Azure Data Factory copy pipelines, Kafka Connect with a Cosmos DB sink, Debezium with a custom sink, or consumer code they write and run themselves.

Gluesync addresses that with **per-agent CDC into Cosmos DB**. A source agent reads each database's native change mechanism, Core Hub routes the changes, and the Cosmos DB target agent writes them through the SDK. SQL Server to Cosmos DB and MongoDB to Cosmos DB run on the same pipeline model, the same snapshot, and the same operations.

HOW IT WORKS

## How Gluesync writes to Cosmos DB

### The write path: the Cosmos DB Java SDK, in optimized batches

The Cosmos DB agent embeds the Azure Cosmos DB Java SDK (`azure-cosmos`) and writes through the SQL API. It is a target agent: it receives changes from any Gluesync source agent and writes them to the database and container you name in Core Hub.

-   **Optimized batches, never row by row:** Core Hub groups incoming changes into chunks, and the agent streams them to Cosmos DB through SDK batching, with no staging container in between. Fewer round-trips keep request load predictable, and the batch size is configurable on the target agent.
-   **Secure and failover-aware:** connections use TLS encryption, and the SDK is aware of automatic failover in your Cosmos DB account.
-   **Direct writes:** items are flushed straight to Cosmos DB through the SDK, with no local disk buffer on the agent. [Azure Cosmos DB agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/azure-cosmosdb-intro.html)

### Snapshot first, then continuous CDC

-   **Seed, then stream:** a full snapshot task seeds each container, then the entity switches to CDC from the source agent's change mechanism and applies changes as they arrive.
-   **INSERT or UPSERT:** INSERT is the fast path for empty containers; UPSERT merges the snapshot with items already in Cosmos DB.
-   **Parallel snapshot writes:** snapshot writing concurrency is configurable per entity, and logical partitioning splits large source tables into ranges read in parallel.
-   **Resume:** an interrupted snapshot resumes from its last saved state when you start the entity again.
-   **Scheduled refreshes:** the Chronos Scheduler runs snapshots, or a snapshot followed by CDC, on a schedule for containers you prefer to reload on a cadence. See [schedules and events](/schedules-and-events/).

### Duplicate item ids: Upsert, Skip, or Fail per entity

Cosmos DB rejects a second item with the same id, so each entity carries an **On duplicate key** setting with three behaviors. **Upsert** is the default: a rejected insert is written again as an upsert, so the item with that id holds the latest values. **Skip** keeps the item already in Cosmos DB and raises a warning in the Notifications Hub that names the batch index and key values of each skipped row, so you know exactly what to reconcile. **Fail** stops the entity on the colliding transaction and leaves the source untouched, so a person can resolve the conflict before data moves again.

Allowed Operations then decides per entity which operations reach the container, and Unlock Schema, Custom Field Functions, and UDFs reshape records into the document form your application reads. See [data transformation](/data-transformation/).

### What your Azure admin sets up

The agent needs a Cosmos DB account it can reach over TLS, an account key, and the database and container that receive the items. Full steps are in the [Azure Cosmos DB target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/azure-cosmosdb-target.html).

1.  Note the hostname of your Azure Cosmos DB account, for example `your-account.documents.azure.com`.
2.  In the Azure portal, copy a service key for the account that the agent will use.
3.  Choose the database and the container that receive the replicated items.
4.  In Core Hub, add the Cosmos DB agent and enter the hostname, the port (443 by default), the database name, the Azure key, and the container name. The same values can be set through the Core Hub REST API.

### Architecture around Core Hub

Lightweight agents sit close to each source; Core Hub orchestrates them through its web UI and REST APIs and routes changes to the Cosmos DB agent. A pipeline groups a source agent, the Cosmos DB agent, and the entities they replicate, so a SQL Server pipeline and an Oracle pipeline can feed containers in the same account side by side.

Core Hub and the agents deploy with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud, including Azure next to your Cosmos DB account. See [CDC streaming](/solutions/cdc-streaming/).

[Explore the general CDC streaming architecture →](/solutions/cdc-streaming/)

WRITE OPTIONS

## The Azure Cosmos DB target agent: one agent, one write path

Cosmos DB has one Gluesync target agent. Every entity writes through the Cosmos DB SDK in optimized batches, so the choices per entity are the source agent, the duplicate-id behavior, and the operations that reach the container.

| Agent | Write technique | Versions | Best for |
| --- | --- | --- | --- |
| [Azure Cosmos DB agent ↗](https://docs.molo17.com/gluesync/latest/agents/azure-cosmosdb-intro.html) | Azure Cosmos DB Java SDK (`azure-cosmos`) over the SQL API, with SDK batching and TLS | Azure Cosmos DB as a service, all versions, any region, SQL API | Application data that must be served from Cosmos DB as it changes: migrations off SQL Server or Oracle, read models for Azure services, and offloads from operational databases. Pair any Gluesync source agent with it under the same Core Hub. |

SOURCES AND TOPOLOGIES

## Feed Cosmos DB from the systems of record you already run

Any Gluesync source agent can feed Cosmos DB, each with its own native capture technique. Open the [integrations finder with Cosmos DB pre-selected](/integrations/?target=Azure+Cosmos+DB#integration-finder) to see every source you can pair with it. To capture changes out of Cosmos DB instead, see [Cosmos DB CDC](/solutions/cosmos-db-cdc/).

Freshness follows the capture method of each source agent and the throughput you provision on each container, and Gluesync keeps pace with your change volume at any scale. [MOLO17 Professional Services](/solutions/professional-services/) can plan the topology and the container throughput with your team.

-   SQL Server to Cosmos DB through Change Data Capture or Change Tracking: see [SQL Server CDC](/solutions/sql-server-cdc/)
-   Oracle to Cosmos DB from the redo logs through LogMiner or XStream: see [Oracle CDC](/solutions/oracle-cdc/)
-   PostgreSQL and MySQL to Cosmos DB from the WAL and the binlog: see [PostgreSQL CDC](/solutions/postgresql-cdc/) and [MySQL CDC](/solutions/mysql-cdc/)
-   MongoDB to Cosmos DB from Change Streams: see [MongoDB CDC](/solutions/mongodb-cdc/)

FAIR, HIGH-LEVEL COMPARISON

## Where Gluesync fits among Cosmos DB ingestion approaches

| Approach | What buyers usually get | Where Gluesync fits |
| --- | --- | --- |
| Azure Data Factory and Synapse pipelines | Managed batch pipelines with copy activities into Cosmos DB, scheduled and orchestrated inside Azure | Continuous CDC from each source's own change mechanism, a snapshot first, and Core Hub monitoring across pipelines; see [cloud migration](/solutions/cloud-migration/) |
| Azure Synapse Link for Cosmos DB | An analytical store that Azure keeps in step with Cosmos DB for analytics queries; it reads from Cosmos DB rather than writing into it | Gluesync writes source changes into Cosmos DB itself, so applications read current items where they already live |
| Kafka Connect with a Cosmos DB sink | Sink connectors that write Kafka records into containers; you run Kafka, Connect, offsets, and schemas, plus a capture tool on each source | Changes go to Cosmos DB with no Kafka cluster in the path, and Kafka stays available as another target; read the [Debezium alternative](/solutions/debezium-alternative/) comparison |
| Qlik Replicate, Striim, and similar commercial platforms | Commercial replication platforms with broad source and target catalogs; packaging and licensing differ by product | A dedicated capture agent per database, one Core Hub, and [MOLO17 enterprise support](/support/#customer-ratings) behind every Cosmos DB pipeline |
| Custom export scripts and consumer code | Full control of mapping and ordering; your team owns checkpoints, retries, throttling, and the load each run puts on production | Continuous CDC from the source log, batched SDK writes, snapshot resume, and Core Hub monitoring with no consumer code to maintain; read [batch ETL vs real-time replication](/blog/batch-etl-vs-real-time-data-replication/) |

RELATED CONTENT

## Cosmos DB replication background and deployment detail

-   [Azure Cosmos DB CDC and target connector for bi-directional integration](/blog/azure-cosmos-db-cdc-target-connector-unlock-bi-directional-data-integration-with-gluesync/)
-   [Batch ETL vs real-time data replication: how to choose](/blog/batch-etl-vs-real-time-data-replication/)
-   [Log-based CDC explained](/blog/real-time-replication-log-based-cdc/)
-   [Cloud migration without a long cutover](/solutions/cloud-migration/)
-   [Database offload](/solutions/database-offload/)
-   [Azure Cosmos DB agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/azure-cosmosdb-intro.html)
-   [Azure Cosmos DB target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/azure-cosmosdb-target.html)
-   [On duplicate key: Upsert, Skip, or Fail ↗](https://docs.molo17.com/gluesync/latest/core-hub/insert-conflict-strategy.html)

FAQ

## Cosmos DB replication questions

What does replicating to Azure Cosmos DB with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the Cosmos DB target agent writes them to your container continuously after a snapshot seeds it.

How does Gluesync write to Azure Cosmos DB?

Through the Azure Cosmos DB Java SDK over the SQL API, with TLS. Changes are grouped into optimized batches and streamed with SDK batching, never one row at a time, and applied as they arrive.

Which sources can replicate to Azure Cosmos DB?

Any Gluesync source agent, including SQL Server, Oracle, PostgreSQL, MySQL, MariaDB, IBM Db2, SAP HANA, MongoDB, and Couchbase. The integrations finder on our website lists every pairing.

How are duplicate item ids handled?

Each entity chooses Upsert, Skip, or Fail. Upsert is the default and overwrites the existing item, Skip keeps it and raises a warning naming the skipped keys, and Fail stops the entity so the collision can be resolved.

What does the Azure admin need to set up?

A Cosmos DB account with its hostname, a service key for the agent, and the database and container that receive the items. Core Hub connects on port 443 unless you change it.

Does the initial load run before CDC starts?

Yes. A full snapshot task seeds each container first, in INSERT or UPSERT mode with configurable writing concurrency, then changes stream in from the source's change mechanism. An interrupted snapshot resumes from its last saved state.

Which Azure Cosmos DB deployments are supported?

Azure Cosmos DB as a service, all versions, in any region, through the SQL API.

REPLICATE TO A TARGET

## Other targets Gluesync delivers to

-    [Replicate to Aerospike](/solutions/replicate-to-aerospike/)
-    [Replicate to DynamoDB](/solutions/replicate-to-dynamodb/)
-    [Replicate to Redshift](/solutions/replicate-to-redshift/)
-    [Replicate to Amazon S3 & S3-compatible](/solutions/replicate-to-amazon-s3/)
-    [Replicate to Cassandra](/solutions/replicate-to-cassandra/)
-    [Replicate to Kafka](/solutions/replicate-to-kafka/)
-    [Replicate to Azure Data Lake](/solutions/replicate-to-azure-data-lake/)
-    [Replicate to ClickHouse](/solutions/replicate-to-clickhouse/)
-    [Replicate to CockroachDB](/solutions/replicate-to-cockroachdb/)
-    [Replicate to Couchbase](/solutions/replicate-to-couchbase/)
-    [Replicate to file stores](/solutions/replicate-to-file-stores/)
-    [Replicate to BigQuery](/solutions/replicate-to-bigquery/)
-    [Replicate to Google Cloud Storage](/solutions/replicate-to-google-cloud-storage/)
-    [Replicate to Google Pub/Sub](/solutions/replicate-to-google-pubsub/)
-    [Replicate to GridGain](/solutions/replicate-to-gridgain/)
-    [Replicate to Db2](/solutions/replicate-to-db2/)
-    [Replicate to Informix](/solutions/replicate-to-informix/)
-    [Replicate to MariaDB](/solutions/replicate-to-mariadb/)
-    [Replicate to SQL Server](/solutions/replicate-to-sql-server/)
-    [Replicate to MongoDB](/solutions/replicate-to-mongodb/)
-    [Replicate to MySQL](/solutions/replicate-to-mysql/)
-    [Replicate to Oracle](/solutions/replicate-to-oracle/)
-    [Replicate to PostgreSQL](/solutions/replicate-to-postgresql/)
-    [Replicate to RavenDB](/solutions/replicate-to-ravendb/)
-    [Replicate to Redis](/solutions/replicate-to-redis/)
-    [Replicate to SAP ASE](/solutions/replicate-to-sap-ase/)
-    [Replicate to SAP HANA](/solutions/replicate-to-sap-hana/)
-    [Replicate to ScyllaDB](/solutions/replicate-to-scylladb/)
-    [Replicate to SingleStore](/solutions/replicate-to-singlestore/)
-    [Replicate to Snowflake](/solutions/replicate-to-snowflake/)
-    [Replicate to Solace PubSub+](/solutions/replicate-to-solace/)
-    [Replicate to Vertica](/solutions/replicate-to-vertica/)
-    [Replicate to YugabyteDB](/solutions/replicate-to-yugabytedb/)

## Evaluate Gluesync with your Azure Cosmos DB account

Start a trial on your infrastructure, or talk to MOLO17 about your sources, your containers, and the throughput each one needs.

[Start a Gluesync trial](/get-gluesync/) [Talk to MOLO17](/contacts/)
