<!-- Generated from the rendered page by scripts/write-llm-mirrors.mjs. Do not edit by hand. -->
Canonical: https://molo17.com/solutions/replicate-to-kafka/
Markdown mirror: https://molo17.com/solutions/replicate-to-kafka/index.md
Title: Replicate to Apache Kafka in real time | MOLO17
Description: Real-time replication into Apache Kafka topics from Oracle, SQL Server, PostgreSQL, and MongoDB as keyed JSON change events, after a snapshot. Start a trial.

[Solutions](/solutions/)  Replicate to Kafka

 SOLUTION · REPLICATE TO KAFKA

# Real-time data replication to Apache Kafka topics

**Every committed change in your systems of record is published to a Kafka topic as it happens, as a JSON change event, instead of waiting for a scheduled export.**

Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MySQL, IBM i, MongoDB, and other [heterogeneous sources](/integrations/?target=Apache+Kafka#integration-finder) with a dedicated agent per database. The Kafka target agent publishes each change with the Apache Kafka Java client, after a snapshot loads every entity, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

[Start a Gluesync trial](/get-gluesync/) [Talk to us](/contacts/)

WHO THIS IS FOR

## Teams that want Kafka fed from operational databases without a CDC stack of their own

-   Platform engineers who run shared Kafka clusters or Confluent Cloud and want [every source](/integrations/?target=Apache+Kafka#integration-finder) to publish through one product, with a topic per entity and one envelope format across all of them
-   Architects replacing Debezium connectors and Kafka Connect clusters that someone must upgrade, rebalance, and watch, with Core Hub monitoring and [best-in-class enterprise support, rated 4.9/5 by customers](/support/#customer-ratings)
-   Application teams that consume topics and need a message contract they can rely on: a stable JSON envelope, the row's key on every record, and an operation type on every change event
-   Data engineers maintaining nightly extracts and dump jobs that feed Kafka today, who want change events instead, with a snapshot for the first load and no polling of production tables

THE PROBLEM

## Kafka topics fed by batch extracts show yesterday's state

Most Kafka topics fed from operational databases start as a scheduled export or a polling job. Consumers then read a copy that trails the source by the length of the batch window, every poll adds read load to production tables, and each database brings its own connector, offsets, and failure mode. Teams weighing a fix usually compare Debezium with Kafka Connect, Confluent's managed source connectors, cloud migration services, or a custom producer they write and run themselves.

Gluesync addresses that with **per-agent CDC into Kafka**. A source agent reads each database's native change mechanism, Core Hub routes the changes, and the Kafka target agent publishes them to topics as they arrive. The message model stays the same whatever the source is: one topic per entity, one JSON change event per change, and the operation type in every event.

HOW IT WORKS

## How Gluesync publishes to Kafka

### The write path: the Apache Kafka Java client, publishing in optimized batches

The Kafka agent publishes through the official Apache Kafka Java client, which also connects to Confluent Cloud. It is a target agent: it receives changes from any Gluesync source agent and produces one record per change to the topic for that entity, with no staging layer and no Connect cluster in between.

-   **Optimized batches, never row by row:** Core Hub groups changes into pages and the producer publishes each page as a batch, so the brokers receive a few large requests instead of one per change, and each topic receives its records in the order they were grouped. The batch size is configurable per agent.
-   **Transactions:** Kafka transactions can be enabled for the agent's writes. [Apache Kafka agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/apache-kafka-intro.html)

### Snapshot first, then change events on the same topic

-   **Seed, then stream:** each entity is loaded from its full table as snapshot records, then the same topic receives the change stream from the source agent in real time.
-   **Snapshot flag:** the `source` block of each event records whether it came from a snapshot or from the change stream, so a consumer can rebuild state from the first record on.
-   **Parallel and resumable snapshots:** logical partitioning reads large source tables in parallel ranges, snapshot writing concurrency is configurable per entity, and an interrupted snapshot resumes from its last saved state.
-   **Append-only by design:** the agent only produces records and never deletes them from a topic, so retention and compaction stay under your cluster's own policies.

### Topics, keys, and the JSON envelope

-   **One topic per entity:** the topics your pipeline needs can be created ahead of time, or Gluesync creates them when automatic topic creation is on.
-   **Record key:** each record is keyed by the row's document key, so every change to a row lands with the same key.
-   **Envelope:** each value is a JSON change event with the `before` and `after` row images, a `source` block with the database, schema, table, and transaction identifier, the operation type, and a timestamp. The shape follows the Debezium-style envelope many consumers already parse.
-   **Delete events:** a delete is published with the deleted row's document key and an empty value, the convention log-compacted topics and Kafka consumers already use to drop a key from their state.
-   **Shaping before publish:** the document key builder, Custom Field Functions, and UDFs decide the record key and reshape each payload, and Allowed Operations chooses which operations reach a topic. See [data transformation](/data-transformation/).

### Setup: publish rights, authentication, and the broker address

Each Kafka pipeline needs a user that can write to its topics and the connection settings below. The full field list and REST calls are in the [target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/apache-kafka-target.html).

1.  Create the topics your entities need, or turn on automatic topic creation, and give the Gluesync user write and publish rights on them.
2.  Choose the authentication: SASL with `SCRAM-SHA-256` (the default) or `SCRAM-SHA-512` over `SASL_SSL`, the default security protocol, or an unauthenticated cluster.
3.  Upload the PEM or JKS certificate for TLS, and enter its password when it is protected.
4.  Enter the broker host, a DNS name or the address of any node, which discovers the other nodes automatically, and the port, which defaults to `9092`.

### Architecture around Core Hub

Lightweight agents sit close to each source. Core Hub orchestrates them through its web UI and REST APIs and routes every change to the Kafka agent. A pipeline groups the source agents, the Kafka agent, and the entities they replicate, and Core Hub and the agents deploy with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud. See [CDC streaming](/solutions/cdc-streaming/) and [snapshot tasks ↗](https://docs.molo17.com/gluesync/latest/core-hub/snapshot-tasks.html).

Because Kafka is one target among many on the same Core Hub, a source table can publish to a topic for event consumers and land in a database or warehouse in another pipeline, from the same capture agent and with the same monitoring.

[Explore the general CDC streaming architecture →](/solutions/cdc-streaming/)

WRITE OPTIONS

## The Kafka target agent: one producer, one topic per entity

Kafka has one Gluesync target agent. Every source agent feeds it through Core Hub, so the choice that shapes the pipeline is the source, while the Kafka side stays the same.

| Agent | Write technique | Versions | Best for |
| --- | --- | --- | --- |
| [Apache Kafka agent ↗](https://docs.molo17.com/gluesync/latest/agents/apache-kafka-intro.html) | Apache Kafka Java client producer with optimized batched publishing, keyed JSON change events, and optional Kafka transactions | Any Apache Kafka version, including Confluent Cloud | Shared Kafka clusters and Confluent Cloud that need change events from several databases. Pick it when each entity should publish to its own topic, with snapshots loaded first and the same envelope across every source. |

SOURCES AND TOPOLOGIES

## Publish changes from the databases you already run

Any Gluesync source agent can feed Kafka. Open the [integrations finder with Apache Kafka pre-selected](/integrations/?target=Apache+Kafka#integration-finder) to see every source you can pair with it, from Oracle and PostgreSQL to IBM i and Couchbase.

One Core Hub can run Oracle to Kafka, PostgreSQL to Kafka, and MongoDB to Kafka side by side, with one set of snapshot and monitoring controls. Gluesync keeps pace with your change volume at any scale, and the batch size is configurable per agent; [MOLO17 Professional Services](/solutions/professional-services/) can tune it with your platform team.

-   Oracle to Kafka from the redo logs through LogMiner or XStream: see [Oracle CDC](/solutions/oracle-cdc/)
-   SQL Server to Kafka through Change Data Capture or Change Tracking: see [SQL Server CDC](/solutions/sql-server-cdc/)
-   IBM i (AS/400) to Kafka through the native journal APIs: see [IBM i CDC](/solutions/ibm-i-cdc/)
-   PostgreSQL, MySQL, and MongoDB to Kafka from the WAL, the binlog, and change streams: see [PostgreSQL CDC](/solutions/postgresql-cdc/), [MySQL CDC](/solutions/mysql-cdc/), and [MongoDB CDC](/solutions/mongodb-cdc/)

FAIR, HIGH-LEVEL COMPARISON

## Where Gluesync fits among Kafka ingestion approaches

| Approach | What buyers usually get | Where Gluesync fits |
| --- | --- | --- |
| Debezium with Kafka Connect | Open-source connectors that run inside Kafka Connect. Your team runs the Connect cluster, manages offsets and schema history, and handles upgrades and restarts | Agents next to each source, Core Hub for operations, and MOLO17 enterprise support, with no Connect cluster to run. Read the [Debezium alternative](/solutions/debezium-alternative/) comparison |
| Confluent managed source connectors | Managed source connectors on Confluent Cloud, configured one connector at a time inside the Confluent platform | One Core Hub for every source, including IBM i and SAP HANA, and the same pipeline model for Kafka and for every other target |
| Kafka Connect JDBC or other polling connectors | Quick to start on tables with a timestamp or incrementing column. Each poll runs a query against the source, and the query defines what counts as a change | Log-based capture from the transaction log, journal, or change stream, so production is not polled; see [CDC streaming](/solutions/cdc-streaming/) |
| Cloud-managed replication services | Managed replication into a cloud platform, with destination lists and an operating model that often sit apart from the team that runs Kafka | Kafka is one of the targets on the same Core Hub as every other target, with the same snapshot and CDC model; see [cloud migration](/solutions/cloud-migration/) |
| Custom producers written into applications | Full control of the message contract. Your team owns the publish calls, retries, schema changes, and the query or log reader behind each topic | Change events captured from the database log, with snapshots and operation types handled by Gluesync, so producers stop being code to maintain. Read [bridging relational and NoSQL databases](/blog/it-doesnt-have-to-be-scary-bridging-the-divide-between-relational-and-nosql-databases/) |

RELATED CONTENT

## Kafka replication background and implementation detail

-   [Now available: CDC for Apache Kafka via Gluesync](/blog/general-availability-cdc-for-apache-kafka-ne-replication-target-gluesync/)
-   [CDC streaming without source overhead](/solutions/cdc-streaming/)
-   [Debezium alternative: managed CDC versus Kafka + Debezium](/solutions/debezium-alternative/)
-   [Data transformation before targets](/data-transformation/)
-   [G-Able: multi-source sync to Aerospike and Kafka](/g-able-gluesync-success-story/)
-   [Indian Bank: Oracle changes delivered to Aerospike and Kafka](/indian-bank-gluesync-success-story/)
-   [Apache Kafka agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/apache-kafka-intro.html)
-   [Apache Kafka target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/apache-kafka-target.html)

FAQ

## Kafka replication questions

What does replicating to Kafka with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the Kafka agent publishes each change to the topic for its entity as it arrives, after a snapshot loads the entity.

How are changes laid out in Kafka topics?

Each entity writes to its own topic, and each change is one record keyed by the row's document key. The value is a JSON change event with the before and after row images, a source block, the operation type, and a timestamp.

Which Kafka deployments does the agent support?

Any Apache Kafka version, including Confluent Cloud. The agent connects with the official Apache Kafka Java client, using SASL SCRAM authentication over TLS or an unauthenticated cluster.

Can Gluesync create the topics?

Yes. Turn on automatic topic creation and Gluesync creates each entity's topic. You can also create the topics ahead of time and give the Gluesync user publish rights on them.

Which sources can publish to Kafka?

Any Gluesync source agent, including Oracle, SQL Server, PostgreSQL, MySQL, IBM i, MongoDB, and Couchbase. The integrations finder lists every pairing with Apache Kafka.

How are deletes represented in Kafka?

A delete is published with the deleted row's document key and an empty value, so log-compacted topics and consumers drop that key from their state.

Does the Kafka agent support transactions?

Yes. Kafka transactions can be enabled on the agent for its writes.

How do I size the writes?

Gluesync publishes in optimized batches, never one row at a time. Core Hub groups changes into pages, the producer publishes each page as a batch, and the batch size is configurable per agent.

REPLICATE TO A TARGET

## Other targets Gluesync delivers to

-    [Replicate to Aerospike](/solutions/replicate-to-aerospike/)
-    [Replicate to DynamoDB](/solutions/replicate-to-dynamodb/)
-    [Replicate to Redshift](/solutions/replicate-to-redshift/)
-    [Replicate to Amazon S3 & S3-compatible](/solutions/replicate-to-amazon-s3/)
-    [Replicate to Cassandra](/solutions/replicate-to-cassandra/)
-    [Replicate to Cosmos DB](/solutions/replicate-to-cosmos-db/)
-    [Replicate to Azure Data Lake](/solutions/replicate-to-azure-data-lake/)
-    [Replicate to ClickHouse](/solutions/replicate-to-clickhouse/)
-    [Replicate to CockroachDB](/solutions/replicate-to-cockroachdb/)
-    [Replicate to Couchbase](/solutions/replicate-to-couchbase/)
-    [Replicate to file stores](/solutions/replicate-to-file-stores/)
-    [Replicate to BigQuery](/solutions/replicate-to-bigquery/)
-    [Replicate to Google Cloud Storage](/solutions/replicate-to-google-cloud-storage/)
-    [Replicate to Google Pub/Sub](/solutions/replicate-to-google-pubsub/)
-    [Replicate to GridGain](/solutions/replicate-to-gridgain/)
-    [Replicate to Db2](/solutions/replicate-to-db2/)
-    [Replicate to Informix](/solutions/replicate-to-informix/)
-    [Replicate to MariaDB](/solutions/replicate-to-mariadb/)
-    [Replicate to SQL Server](/solutions/replicate-to-sql-server/)
-    [Replicate to MongoDB](/solutions/replicate-to-mongodb/)
-    [Replicate to MySQL](/solutions/replicate-to-mysql/)
-    [Replicate to Oracle](/solutions/replicate-to-oracle/)
-    [Replicate to PostgreSQL](/solutions/replicate-to-postgresql/)
-    [Replicate to RavenDB](/solutions/replicate-to-ravendb/)
-    [Replicate to Redis](/solutions/replicate-to-redis/)
-    [Replicate to SAP ASE](/solutions/replicate-to-sap-ase/)
-    [Replicate to SAP HANA](/solutions/replicate-to-sap-hana/)
-    [Replicate to ScyllaDB](/solutions/replicate-to-scylladb/)
-    [Replicate to SingleStore](/solutions/replicate-to-singlestore/)
-    [Replicate to Snowflake](/solutions/replicate-to-snowflake/)
-    [Replicate to Solace PubSub+](/solutions/replicate-to-solace/)
-    [Replicate to Vertica](/solutions/replicate-to-vertica/)
-    [Replicate to YugabyteDB](/solutions/replicate-to-yugabytedb/)

## Evaluate Gluesync with your Kafka cluster

Start a trial on your infrastructure, or talk to MOLO17 about your sources, your topic design, and the change events your consumers need.

[Start a Gluesync trial](/get-gluesync/) [Talk to MOLO17](/contacts/)
