SOLUTION · REPLICATE TO KAFKA

Real-time data replication to Apache Kafka topics

Every committed change in your systems of record is published to a Kafka topic as it happens, as a JSON change event, instead of waiting for a scheduled export.

Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MySQL, IBM i, MongoDB, and other heterogeneous sources with a dedicated agent per database. The Kafka target agent publishes each change with the Apache Kafka Java client, after a snapshot loads every entity, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

WHO THIS IS FOR

Teams that want Kafka fed from operational databases without a CDC stack of their own

  • Platform engineers who run shared Kafka clusters or Confluent Cloud and want every source to publish through one product, with a topic per entity and one envelope format across all of them
  • Architects replacing Debezium connectors and Kafka Connect clusters that someone must upgrade, rebalance, and watch, with Core Hub monitoring and best-in-class enterprise support, rated 4.9/5 by customers
  • Application teams that consume topics and need a message contract they can rely on: a stable JSON envelope, the row's key on every record, and an operation type on every change event
  • Data engineers maintaining nightly extracts and dump jobs that feed Kafka today, who want change events instead, with a snapshot for the first load and no polling of production tables

THE PROBLEM

Kafka topics fed by batch extracts show yesterday's state

Most Kafka topics fed from operational databases start as a scheduled export or a polling job. Consumers then read a copy that trails the source by the length of the batch window, every poll adds read load to production tables, and each database brings its own connector, offsets, and failure mode. Teams weighing a fix usually compare Debezium with Kafka Connect, Confluent's managed source connectors, cloud migration services, or a custom producer they write and run themselves.

Gluesync addresses that with per-agent CDC into Kafka. A source agent reads each database's native change mechanism, Core Hub routes the changes, and the Kafka target agent publishes them to topics as they arrive. The message model stays the same whatever the source is: one topic per entity, one JSON change event per change, and the operation type in every event.

HOW IT WORKS

How Gluesync publishes to Kafka

The write path: the Apache Kafka Java client, publishing in optimized batches

The Kafka agent publishes through the official Apache Kafka Java client, which also connects to Confluent Cloud. It is a target agent: it receives changes from any Gluesync source agent and produces one record per change to the topic for that entity, with no staging layer and no Connect cluster in between.

  • Optimized batches, never row by row: Core Hub groups changes into pages and the producer publishes each page as a batch, so the brokers receive a few large requests instead of one per change, and each topic receives its records in the order they were grouped. The batch size is configurable per agent.
  • Transactions: Kafka transactions can be enabled for the agent's writes. Apache Kafka agent overview ↗

Snapshot first, then change events on the same topic

  • Seed, then stream: each entity is loaded from its full table as snapshot records, then the same topic receives the change stream from the source agent in real time.
  • Snapshot flag: the source block of each event records whether it came from a snapshot or from the change stream, so a consumer can rebuild state from the first record on.
  • Parallel and resumable snapshots: logical partitioning reads large source tables in parallel ranges, snapshot writing concurrency is configurable per entity, and an interrupted snapshot resumes from its last saved state.
  • Append-only by design: the agent only produces records and never deletes them from a topic, so retention and compaction stay under your cluster's own policies.

Topics, keys, and the JSON envelope

  • One topic per entity: the topics your pipeline needs can be created ahead of time, or Gluesync creates them when automatic topic creation is on.
  • Record key: each record is keyed by the row's document key, so every change to a row lands with the same key.
  • Envelope: each value is a JSON change event with the before and after row images, a source block with the database, schema, table, and transaction identifier, the operation type, and a timestamp. The shape follows the Debezium-style envelope many consumers already parse.
  • Delete events: a delete is published with the deleted row's document key and an empty value, the convention log-compacted topics and Kafka consumers already use to drop a key from their state.
  • Shaping before publish: the document key builder, Custom Field Functions, and UDFs decide the record key and reshape each payload, and Allowed Operations chooses which operations reach a topic. See data transformation.

Setup: publish rights, authentication, and the broker address

Each Kafka pipeline needs a user that can write to its topics and the connection settings below. The full field list and REST calls are in the target setup guide ↗.

  1. Create the topics your entities need, or turn on automatic topic creation, and give the Gluesync user write and publish rights on them.
  2. Choose the authentication: SASL with SCRAM-SHA-256 (the default) or SCRAM-SHA-512 over SASL_SSL, the default security protocol, or an unauthenticated cluster.
  3. Upload the PEM or JKS certificate for TLS, and enter its password when it is protected.
  4. Enter the broker host, a DNS name or the address of any node, which discovers the other nodes automatically, and the port, which defaults to 9092.

Architecture around Core Hub

Lightweight agents sit close to each source. Core Hub orchestrates them through its web UI and REST APIs and routes every change to the Kafka agent. A pipeline groups the source agents, the Kafka agent, and the entities they replicate, and Core Hub and the agents deploy with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud. See CDC streaming and snapshot tasks ↗.

Because Kafka is one target among many on the same Core Hub, a source table can publish to a topic for event consumers and land in a database or warehouse in another pipeline, from the same capture agent and with the same monitoring.

Explore the general CDC streaming architecture →

WRITE OPTIONS

The Kafka target agent: one producer, one topic per entity

Kafka has one Gluesync target agent. Every source agent feeds it through Core Hub, so the choice that shapes the pipeline is the source, while the Kafka side stays the same.

AgentWrite techniqueVersionsBest for
Apache Kafka agent ↗ Apache Kafka Java client producer with optimized batched publishing, keyed JSON change events, and optional Kafka transactions Any Apache Kafka version, including Confluent Cloud Shared Kafka clusters and Confluent Cloud that need change events from several databases. Pick it when each entity should publish to its own topic, with snapshots loaded first and the same envelope across every source.

SOURCES AND TOPOLOGIES

Publish changes from the databases you already run

Any Gluesync source agent can feed Kafka. Open the integrations finder with Apache Kafka pre-selected to see every source you can pair with it, from Oracle and PostgreSQL to IBM i and Couchbase.

One Core Hub can run Oracle to Kafka, PostgreSQL to Kafka, and MongoDB to Kafka side by side, with one set of snapshot and monitoring controls. Gluesync keeps pace with your change volume at any scale, and the batch size is configurable per agent; MOLO17 Professional Services can tune it with your platform team.

  • Oracle to Kafka from the redo logs through LogMiner or XStream: see Oracle CDC
  • SQL Server to Kafka through Change Data Capture or Change Tracking: see SQL Server CDC
  • IBM i (AS/400) to Kafka through the native journal APIs: see IBM i CDC
  • PostgreSQL, MySQL, and MongoDB to Kafka from the WAL, the binlog, and change streams: see PostgreSQL CDC, MySQL CDC, and MongoDB CDC

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among Kafka ingestion approaches

ApproachWhat buyers usually getWhere Gluesync fits
Debezium with Kafka Connect Open-source connectors that run inside Kafka Connect. Your team runs the Connect cluster, manages offsets and schema history, and handles upgrades and restarts Agents next to each source, Core Hub for operations, and MOLO17 enterprise support, with no Connect cluster to run. Read the Debezium alternative comparison
Confluent managed source connectors Managed source connectors on Confluent Cloud, configured one connector at a time inside the Confluent platform One Core Hub for every source, including IBM i and SAP HANA, and the same pipeline model for Kafka and for every other target
Kafka Connect JDBC or other polling connectors Quick to start on tables with a timestamp or incrementing column. Each poll runs a query against the source, and the query defines what counts as a change Log-based capture from the transaction log, journal, or change stream, so production is not polled; see CDC streaming
Cloud-managed replication services Managed replication into a cloud platform, with destination lists and an operating model that often sit apart from the team that runs Kafka Kafka is one of the targets on the same Core Hub as every other target, with the same snapshot and CDC model; see cloud migration
Custom producers written into applications Full control of the message contract. Your team owns the publish calls, retries, schema changes, and the query or log reader behind each topic Change events captured from the database log, with snapshots and operation types handled by Gluesync, so producers stop being code to maintain. Read bridging relational and NoSQL databases

FAQ

Kafka replication questions

What does replicating to Kafka with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the Kafka agent publishes each change to the topic for its entity as it arrives, after a snapshot loads the entity.

How are changes laid out in Kafka topics?

Each entity writes to its own topic, and each change is one record keyed by the row's document key. The value is a JSON change event with the before and after row images, a source block, the operation type, and a timestamp.

Which Kafka deployments does the agent support?

Any Apache Kafka version, including Confluent Cloud. The agent connects with the official Apache Kafka Java client, using SASL SCRAM authentication over TLS or an unauthenticated cluster.

Can Gluesync create the topics?

Yes. Turn on automatic topic creation and Gluesync creates each entity's topic. You can also create the topics ahead of time and give the Gluesync user publish rights on them.

Which sources can publish to Kafka?

Any Gluesync source agent, including Oracle, SQL Server, PostgreSQL, MySQL, IBM i, MongoDB, and Couchbase. The integrations finder lists every pairing with Apache Kafka.

How are deletes represented in Kafka?

A delete is published with the deleted row's document key and an empty value, so log-compacted topics and consumers drop that key from their state.

Does the Kafka agent support transactions?

Yes. Kafka transactions can be enabled on the agent for its writes.

How do I size the writes?

Gluesync publishes in optimized batches, never one row at a time. Core Hub groups changes into pages, the producer publishes each page as a batch, and the batch size is configurable per agent.

Evaluate Gluesync with your Kafka cluster

Start a trial on your infrastructure, or talk to MOLO17 about your sources, your topic design, and the change events your consumers need.