SOLUTION · REPLICATE TO CASSANDRA

Real-time data replication to Apache Cassandra from your operational databases

Committed changes from your relational and document systems reach Cassandra continuously, as size-capped CQL batches, instead of waiting for the next Spark job or bulk load.

Gluesync by MOLO17 captures changes from Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, DynamoDB, and other heterogeneous sources with a dedicated agent per database. The Cassandra target agent writes them through the official DataStax Java driver, with TLS and multi-datacenter routing. Each table is seeded with a snapshot, then follows the change stream, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

WHO THIS IS FOR

Teams that need Cassandra tables to reflect the systems of record now

  • Platform engineers running Cassandra clusters for high-write or read-heavy services who want the data behind those services kept current without writing and operating a change consumer
  • Data platform leads feeding Cassandra from Oracle, PostgreSQL, MongoDB, or DynamoDB who want one Core Hub for every source, backed by best-in-class enterprise support, rated 4.9/5 by customers
  • Architects designing read models and event-driven services who need to decide which source feeds which keyspace, where each agent runs, and how the datacenters are routed
  • Engineers replacing a Kafka Connect Cassandra sink, a custom consumer, or a scheduled load script with one product for every source, with snapshots, checkpoints, and monitoring built in

THE PROBLEM

Cassandra read models that are loaded in batches fall behind

Cassandra tables that serve an application are often filled by dual writes in application code, by periodic bulk jobs, or by a custom consumer that someone has to keep running. Each option either adds write logic to every service, leaves the table stale between runs, or adds a Cassandra CDC pipeline that the team builds and operates by hand. Buyers comparing options usually weigh Debezium with Kafka and a Kafka Connect sink, managed ELT services, scheduled Spark or ETL jobs, and their own consumers.

Gluesync addresses that with per-agent CDC into Cassandra. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the Cassandra target agent applies them through the DataStax driver in batches sized for your cluster. Oracle to Cassandra, PostgreSQL to Cassandra, or MongoDB to Cassandra all share the same snapshot, pipeline model, and operations.

HOW IT WORKS

How Gluesync writes to Cassandra

The write path: the DataStax Java driver and CQL

The Cassandra agent connects through the official DataStax Java driver, bundled with the agent, and writes CQL straight to your keyspace tables. It is a target agent: it receives changes from any Gluesync source agent, relational or document, and lands them as rows keyed the way your Cassandra tables expect.

  • Connections: TLS is configured on the connection.
  • Datacenters: set the Datacenter property to the datacenter the agent writes to. The driver routes requests across a multi-datacenter deployment from there.
  • Cassandra agent overview ↗

Optimized batches, never row by row

Gluesync never writes one row at a time to Cassandra. Changes are grouped into batches before they reach the cluster, which cuts round-trips from the agent and keeps the write load on your nodes predictable. On Cassandra the batch is capped by payload size rather than row count, and the cap is configurable, so you can set it to match, or sit slightly under, the batch_size_warn_threshold of your cluster and keep every batch inside the limit your nodes enforce.

Snapshot first, then continuous CDC

  • Seed: snapshot batches are persisted to the Cassandra keyspace tables, so each entity starts from a full copy of its source table.
  • Stream: after the snapshot, inserts, updates, and deletes from the source change stream are applied in real time.
  • Resume: an interrupted snapshot resumes from its last saved state when you start the entity again.
  • Scheduled refreshes: the Chronos Scheduler runs snapshots on a cadence for tables you prefer to reload. See schedules and events.

Keys, identifiers, and table creation

  • Identifier case: Cassandra folds unquoted identifiers to lower case, so Core Hub normalizes target column names and filter clauses to lower case when an entity is saved.
  • Table creation: when a target table does not exist, Core Hub generates the CREATE TABLE statement from the source columns and primary key, using the lower-case names Cassandra expects.
  • Keys: a duplicate key cannot arise on Cassandra, because a write replaces the row stored under the same primary key. Replayed changes converge on the latest state of the source row with no conflict policy to manage.

What your Cassandra admin sets up

The agent needs a user with read and write access to the target tables and keyspace, plus the connection details of one node or contact point. The Cassandra target setup guide ↗ lists every field.

  1. Create or choose a user with read and write permission on the target keyspace and its tables.
  2. Note the host or IP address, the port (9042 by default), the keyspace name, and the name of the datacenter the agent writes to.
  3. When the cluster requires TLS, switch it on for the connection and provide the certificate path (enableTls and certificatePath in the REST configuration).
  4. In Core Hub, add the Cassandra target agent with those credentials, then attach it to a pipeline with one or more source agents.

Architecture around Core Hub

Source agents sit close to each database, the Cassandra agent sits close to the cluster, and Core Hub coordinates them through its web UI and REST API. Pipelines group a source agent, the Cassandra agent, and the entities they replicate, and every component deploys with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud. See CDC streaming without source overhead for the wider pattern.

Explore the general CDC streaming architecture →

WRITE OPTIONS

The Cassandra target agent: one agent for every entity

Cassandra has one Gluesync target agent. Every entity written to it uses the same DataStax driver path: a snapshot to seed the table, then the change stream applied in batches.

AgentWrite techniqueVersionsBest for
Cassandra agent ↗ DataStax Java driver; optimized CQL batches capped by payload size, for snapshot inserts and real-time changes Apache Cassandra 2.1 and later, with TLS and multi-datacenter routing Cassandra clusters that serve applications and need relational or document changes applied continuously. Pick it when the cluster spans datacenters and each agent should write to its local one.

SOURCES AND TOPOLOGIES

Feed Cassandra from the systems of record you already run

Any Gluesync source agent can feed Cassandra, each with its own native capture technique. Open the integrations finder with Cassandra pre-selected to see every source you can pair with it.

Changes from several sources can land in one cluster through one Core Hub, each pipeline with its own snapshot and monitoring. Gluesync keeps pace with your change volume at any scale, and the batch size and target datacenter are configurable per agent. MOLO17 Professional Services can help plan keyspaces and agent placement.

  • Oracle to Cassandra from the redo logs through LogMiner or XStream: see Oracle CDC
  • PostgreSQL and MySQL to Cassandra from the write-ahead log and the binlog: see PostgreSQL CDC and MySQL CDC
  • MongoDB to Cassandra from Change Streams: see MongoDB CDC
  • DynamoDB to Cassandra from DynamoDB Streams: see DynamoDB CDC

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among Cassandra ingestion approaches

ApproachWhat buyers usually getWhere Gluesync fits
Debezium, Kafka, and a Kafka Connect Cassandra sink Open-source capture into Kafka topics, then a sink connector writes to Cassandra; you run Kafka and Connect, manage offsets, and decide how change events map to rows Changes applied to Cassandra tables from each source agent, with no Kafka cluster in the path. Read the Debezium alternative comparison
Fivetran and Fivetran HVR Managed ELT and log-based replication with a broad source catalog, oriented toward warehouses and lakes Agents installed next to each source, native capture per engine, and one Core Hub for every pipeline; see CDC streaming
Airbyte Open-source and cloud ELT connectors; incremental and CDC modes vary by connector, and you run the platform or use the managed service A commercial product with a dedicated capture agent per database and MOLO17 enterprise support behind every pipeline
Qlik Replicate, AWS DMS, and similar replication services Mature replication across many engines; target lists and load behavior differ by product Agents run on-premises or in any cloud with Docker, Docker Compose, or Kubernetes and write straight to Cassandra under one Core Hub. See migrating to Gluesync
Custom consumers and scheduled scripts Full control; your team owns the consumer, its offsets, retries, batching, and the load each run puts on production Log-based capture and batched writes without consumer code to maintain, with Core Hub monitoring each pipeline

FAQ

Cassandra replication questions

What does replicating to Cassandra with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the Cassandra target agent writes them to your keyspace tables. Each table is seeded with a snapshot first, then follows the change stream continuously.

How does Gluesync write data to Cassandra?

Through the official DataStax Java driver, which is bundled with the agent. Changes are grouped into optimized CQL batches, each capped by a payload size that is configurable on the agent.

Which sources can replicate to Cassandra?

Any Gluesync source agent, including Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB. The integrations finder on our website lists every source you can pair with Cassandra.

What happens when a row already exists in Cassandra?

The incoming write replaces the row stored under the same primary key, so the table converges on the latest state of the source row and a replayed change never fails on a duplicate.

What does the Cassandra admin need to set up?

A user with read and write permission on the target keyspace and tables, the host or IP address, the port (9042 by default), the keyspace name, and the datacenter the agent writes to. Enable TLS with a certificate path when the cluster requires it.

Which Cassandra versions does the agent support?

Apache Cassandra 2.1 and later, connected through the DataStax Java driver with TLS and multi-datacenter routing.

Does Gluesync create the tables in Cassandra?

Yes. Core Hub generates the CREATE TABLE statement from the source columns and primary key, with lower-case column names to match how Cassandra folds unquoted identifiers.

Evaluate Gluesync with your Cassandra cluster

Start a trial on your infrastructure, or talk to MOLO17 about your sources, keyspaces, datacenters, and batch sizing.