SOLUTION · REPLICATE TO CASSANDRA
Real-time data replication to Apache Cassandra from your operational databases
Committed changes from your relational and document systems reach Cassandra continuously, as size-capped CQL batches, instead of waiting for the next Spark job or bulk load.
Gluesync by MOLO17 captures changes from Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, DynamoDB, and other heterogeneous sources with a dedicated agent per database. The Cassandra target agent writes them through the official DataStax Java driver, with TLS and multi-datacenter routing. Each table is seeded with a snapshot, then follows the change stream, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.
WHO THIS IS FOR
Teams that need Cassandra tables to reflect the systems of record now
- Platform engineers running Cassandra clusters for high-write or read-heavy services who want the data behind those services kept current without writing and operating a change consumer
- Data platform leads feeding Cassandra from Oracle, PostgreSQL, MongoDB, or DynamoDB who want one Core Hub for every source, backed by best-in-class enterprise support, rated 4.9/5 by customers
- Architects designing read models and event-driven services who need to decide which source feeds which keyspace, where each agent runs, and how the datacenters are routed
- Engineers replacing a Kafka Connect Cassandra sink, a custom consumer, or a scheduled load script with one product for every source, with snapshots, checkpoints, and monitoring built in
THE PROBLEM
Cassandra read models that are loaded in batches fall behind
Cassandra tables that serve an application are often filled by dual writes in application code, by periodic bulk jobs, or by a custom consumer that someone has to keep running. Each option either adds write logic to every service, leaves the table stale between runs, or adds a Cassandra CDC pipeline that the team builds and operates by hand. Buyers comparing options usually weigh Debezium with Kafka and a Kafka Connect sink, managed ELT services, scheduled Spark or ETL jobs, and their own consumers.
Gluesync addresses that with per-agent CDC into Cassandra. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the Cassandra target agent applies them through the DataStax driver in batches sized for your cluster. Oracle to Cassandra, PostgreSQL to Cassandra, or MongoDB to Cassandra all share the same snapshot, pipeline model, and operations.
HOW IT WORKS
How Gluesync writes to Cassandra
The write path: the DataStax Java driver and CQL
The Cassandra agent connects through the official DataStax Java driver, bundled with the agent, and writes CQL straight to your keyspace tables. It is a target agent: it receives changes from any Gluesync source agent, relational or document, and lands them as rows keyed the way your Cassandra tables expect.
- Connections: TLS is configured on the connection.
- Datacenters: set the
Datacenterproperty to the datacenter the agent writes to. The driver routes requests across a multi-datacenter deployment from there. - Cassandra agent overview ↗
Optimized batches, never row by row
Gluesync never writes one row at a time to Cassandra. Changes are grouped into batches before they reach the cluster, which cuts round-trips from the agent and keeps the write load on your nodes predictable. On Cassandra the batch is capped by payload size rather than row count, and the cap is configurable, so you can set it to match, or sit slightly under, the batch_size_warn_threshold of your cluster and keep every batch inside the limit your nodes enforce.
Snapshot first, then continuous CDC
- Seed: snapshot batches are persisted to the Cassandra keyspace tables, so each entity starts from a full copy of its source table.
- Stream: after the snapshot, inserts, updates, and deletes from the source change stream are applied in real time.
- Resume: an interrupted snapshot resumes from its last saved state when you start the entity again.
- Scheduled refreshes: the Chronos Scheduler runs snapshots on a cadence for tables you prefer to reload. See schedules and events.
Keys, identifiers, and table creation
- Identifier case: Cassandra folds unquoted identifiers to lower case, so Core Hub normalizes target column names and filter clauses to lower case when an entity is saved.
- Table creation: when a target table does not exist, Core Hub generates the
CREATE TABLEstatement from the source columns and primary key, using the lower-case names Cassandra expects. - Keys: a duplicate key cannot arise on Cassandra, because a write replaces the row stored under the same primary key. Replayed changes converge on the latest state of the source row with no conflict policy to manage.
What your Cassandra admin sets up
The agent needs a user with read and write access to the target tables and keyspace, plus the connection details of one node or contact point. The Cassandra target setup guide ↗ lists every field.
- Create or choose a user with read and write permission on the target keyspace and its tables.
- Note the host or IP address, the port (9042 by default), the keyspace name, and the name of the datacenter the agent writes to.
- When the cluster requires TLS, switch it on for the connection and provide the certificate path (
enableTlsandcertificatePathin the REST configuration). - In Core Hub, add the Cassandra target agent with those credentials, then attach it to a pipeline with one or more source agents.
Architecture around Core Hub
Source agents sit close to each database, the Cassandra agent sits close to the cluster, and Core Hub coordinates them through its web UI and REST API. Pipelines group a source agent, the Cassandra agent, and the entities they replicate, and every component deploys with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud. See CDC streaming without source overhead for the wider pattern.
WRITE OPTIONS
The Cassandra target agent: one agent for every entity
Cassandra has one Gluesync target agent. Every entity written to it uses the same DataStax driver path: a snapshot to seed the table, then the change stream applied in batches.
| Agent | Write technique | Versions | Best for |
|---|---|---|---|
| Cassandra agent ↗ | DataStax Java driver; optimized CQL batches capped by payload size, for snapshot inserts and real-time changes | Apache Cassandra 2.1 and later, with TLS and multi-datacenter routing | Cassandra clusters that serve applications and need relational or document changes applied continuously. Pick it when the cluster spans datacenters and each agent should write to its local one. |
SOURCES AND TOPOLOGIES
Feed Cassandra from the systems of record you already run
Any Gluesync source agent can feed Cassandra, each with its own native capture technique. Open the integrations finder with Cassandra pre-selected to see every source you can pair with it.
Changes from several sources can land in one cluster through one Core Hub, each pipeline with its own snapshot and monitoring. Gluesync keeps pace with your change volume at any scale, and the batch size and target datacenter are configurable per agent. MOLO17 Professional Services can help plan keyspaces and agent placement.
- Oracle to Cassandra from the redo logs through LogMiner or XStream: see Oracle CDC
- PostgreSQL and MySQL to Cassandra from the write-ahead log and the binlog: see PostgreSQL CDC and MySQL CDC
- MongoDB to Cassandra from Change Streams: see MongoDB CDC
- DynamoDB to Cassandra from DynamoDB Streams: see DynamoDB CDC
FAIR, HIGH-LEVEL COMPARISON
Where Gluesync fits among Cassandra ingestion approaches
| Approach | What buyers usually get | Where Gluesync fits |
|---|---|---|
| Debezium, Kafka, and a Kafka Connect Cassandra sink | Open-source capture into Kafka topics, then a sink connector writes to Cassandra; you run Kafka and Connect, manage offsets, and decide how change events map to rows | Changes applied to Cassandra tables from each source agent, with no Kafka cluster in the path. Read the Debezium alternative comparison |
| Fivetran and Fivetran HVR | Managed ELT and log-based replication with a broad source catalog, oriented toward warehouses and lakes | Agents installed next to each source, native capture per engine, and one Core Hub for every pipeline; see CDC streaming |
| Airbyte | Open-source and cloud ELT connectors; incremental and CDC modes vary by connector, and you run the platform or use the managed service | A commercial product with a dedicated capture agent per database and MOLO17 enterprise support behind every pipeline |
| Qlik Replicate, AWS DMS, and similar replication services | Mature replication across many engines; target lists and load behavior differ by product | Agents run on-premises or in any cloud with Docker, Docker Compose, or Kubernetes and write straight to Cassandra under one Core Hub. See migrating to Gluesync |
| Custom consumers and scheduled scripts | Full control; your team owns the consumer, its offsets, retries, batching, and the load each run puts on production | Log-based capture and batched writes without consumer code to maintain, with Core Hub monitoring each pipeline |
RELATED CONTENT
Cassandra replication background and implementation detail
- Introduction to data integration: a comprehensive guide
- Batch ETL vs real-time data replication: how to choose
- Debezium alternative: managed CDC versus Kafka + Debezium
- CDC streaming without source overhead
- Move read traffic off the system of record
- Cassandra agent overview ↗
- Cassandra target setup guide ↗
- Write strategies: optimized batches ↗
FAQ
Cassandra replication questions
What does replicating to Cassandra with Gluesync involve?
A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the Cassandra target agent writes them to your keyspace tables. Each table is seeded with a snapshot first, then follows the change stream continuously.
How does Gluesync write data to Cassandra?
Through the official DataStax Java driver, which is bundled with the agent. Changes are grouped into optimized CQL batches, each capped by a payload size that is configurable on the agent.
Which sources can replicate to Cassandra?
Any Gluesync source agent, including Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB. The integrations finder on our website lists every source you can pair with Cassandra.
What happens when a row already exists in Cassandra?
The incoming write replaces the row stored under the same primary key, so the table converges on the latest state of the source row and a replayed change never fails on a duplicate.
What does the Cassandra admin need to set up?
A user with read and write permission on the target keyspace and tables, the host or IP address, the port (9042 by default), the keyspace name, and the datacenter the agent writes to. Enable TLS with a certificate path when the cluster requires it.
Which Cassandra versions does the agent support?
Apache Cassandra 2.1 and later, connected through the DataStax Java driver with TLS and multi-datacenter routing.
Does Gluesync create the tables in Cassandra?
Yes. Core Hub generates the CREATE TABLE statement from the source columns and primary key, with lower-case column names to match how Cassandra folds unquoted identifiers.
REPLICATE TO A TARGET
Other targets Gluesync delivers to
- Replicate to Aerospike
- Replicate to DynamoDB
- Replicate to Redshift
- Replicate to Amazon S3 & S3-compatible
- Replicate to Kafka
- Replicate to Cosmos DB
- Replicate to Azure Data Lake
- Replicate to ClickHouse
- Replicate to CockroachDB
- Replicate to Couchbase
- Replicate to file stores
- Replicate to BigQuery
- Replicate to Google Cloud Storage
- Replicate to Google Pub/Sub
- Replicate to GridGain
- Replicate to Db2
- Replicate to Informix
- Replicate to MariaDB
- Replicate to SQL Server
- Replicate to MongoDB
- Replicate to MySQL
- Replicate to Oracle
- Replicate to PostgreSQL
- Replicate to RavenDB
- Replicate to Redis
- Replicate to SAP ASE
- Replicate to SAP HANA
- Replicate to ScyllaDB
- Replicate to SingleStore
- Replicate to Snowflake
- Replicate to Solace PubSub+
- Replicate to Vertica
- Replicate to YugabyteDB
Evaluate Gluesync with your Cassandra cluster
Start a trial on your infrastructure, or talk to MOLO17 about your sources, keyspaces, datacenters, and batch sizing.