SOLUTION · SCYLLADB CDC

ScyllaDB change data capture and real-time data replication

Real-time ScyllaDB CDC from native CDC log tables, read partition by partition for incremental changes.

Gluesync by MOLO17 captures changes from ScyllaDB through a dedicated source agent that reads the CDC log tables ScyllaDB keeps for each CDC-enabled table. The agent connects with the official ScyllaDB Java driver, then delivers inserts, updates, and deletes continuously to the databases, warehouses, lakes, and event streams your teams already use. Manage every pipeline from the Core Hub web UI, the Gluesync control plane.

WHO THIS IS FOR

ScyllaDB teams moving operational data into analytics and applications

  • ScyllaDB DBAs who need to see which tables get CDC enabled, which user reads the CDC log tables, and what the connection needs (TLS and datacenter) before approving a pipeline
  • Data platform leads feeding Snowflake, BigQuery, a data lake, Kafka, or a relational database from ScyllaDB, who want one control plane for sources and targets instead of a Kafka Connect mesh
  • Engineers who prototyped a Debezium-based connector or a custom CQL consumer on ScyllaDB CDC and now own the consumer, the offsets, and the on-call rotation
  • Architects who want ScyllaDB changes in SQL, warehouse, or lake targets without assembling a connector stack, with best-in-class enterprise support from MOLO17, rated 4.9/5 by customers

THE PROBLEM

ScyllaDB data that only moves in batches becomes stale

ScyllaDB is often chosen for high-throughput, low-latency workloads, and the data that matters most stays in that cluster. Batch extracts and table scans leave analytics and downstream applications working from a copy that is already behind, and repeated reads compete with production for cluster resources. Buyers searching for ScyllaDB CDC usually weigh three routes: a Kafka connector built on Debezium that they run themselves, CDC client libraries they wrap in their own consumer, or custom code that reads the log tables directly.

Gluesync addresses that pattern with per-agent CDC. A ScyllaDB source agent reads the CDC log tables, and target agents write to the destination you choose. The agent runs in the same pipeline model as every other Gluesync source, so you can reach a warehouse, a SQL database, or a stream without first running a Kafka Connect cluster.

HOW IT WORKS

How Gluesync does ScyllaDB CDC

Capture from CDC log tables

ScyllaDB CDC is enabled per table, either when the table is created with WITH cdc = {'enabled': true} or later with an ALTER statement. ScyllaDB then maintains a log table named <table>_scylla_cdc_log in the same keyspace, with the same replication strategy, and each key column of the base table appears in it with the same name and type. The agent reads the CDC table partitions for incremental changes.

One write can produce several log rows. The count depends on the CDC options, the write action (insert, update, row deletion, range deletion, or partition deletion), and the column types, since collections need extra handling.

What your DBA will be asked to enable

  1. Enable CDC on each source table, at creation with WITH cdc = {'enabled': true} or on an existing table with an ALTER statement, using the options in the source setup guide.
  2. Give the Gluesync user read permission on the source tables. The source setup guide lists read permission as the prerequisite.
  3. Note the keyspace and datacenter names, which the connection setup asks for.

Connections and TLS

  • Connection fields: host or IP address, port (default 9042), keyspace, username, password, and datacenter name.
  • TLS: set enableTls and a certificatePath in the Core Hub credentials call, which is also available through the REST API.
  • Cluster access: the official ScyllaDB Java driver supports TLS and load-balanced connections, and connects with cluster awareness.

Architecture around Core Hub

The ScyllaDB agent runs close to the cluster, and Core Hub coordinates it with the target agents through its web UI and REST APIs. A pipeline groups the source agent, its target agents, and the CDC-enabled tables it replicates as entities. Source progress is tracked in Core Hub metadata and CDC checkpoints. Core Hub and agents deploy with Docker, Docker Compose, or Kubernetes.

Explore the general CDC streaming architecture →

CAPTURE OPTIONS

The ScyllaDB source agent

One agent covers ScyllaDB as a source. It reads the CDC log tables through the official ScyllaDB Java driver, on ScyllaDB Cloud, on-premises, or a public cloud provider.

AgentCapture techniqueVersionsBest for
ScyllaDB agent ↗ ScyllaDB CDC log tables read through the official ScyllaDB Java driver Tested from ScyllaDB 6.0; ScyllaDB Cloud, on-premises, or hosted in a public cloud Any ScyllaDB table where CDC can be enabled, when you want the change stream delivered to SQL, warehouse, lake, or stream targets under Core Hub.

TARGETS AND TOPOLOGIES

Use ScyllaDB as a source, a target, or both

Use the integrations directory to pair ScyllaDB with relational engines, warehouses, lakes, and event streams, subject to each agent's documented source and target role. ScyllaDB is supported in both roles, and the target agent writes through the same official Java driver with cluster awareness.

Gluesync keeps pace with your change volume at any scale. MOLO17 Professional Services can size the deployment with your team.

Target agents write in optimized batches, never row by row, and switch to native bulk load for both snapshots and CDC on targets such as Snowflake, Google BigQuery, Amazon Redshift, Microsoft SQL Server, and PostgreSQL.

  • Offload reads from ScyllaDB to another database for reporting: see database offload
  • Feed a cloud warehouse or data lake continuously from ScyllaDB changes: see warehouse sync and data lake
  • Publish ScyllaDB changes to Apache Kafka for event-driven consumers
  • Use ScyllaDB as the destination, where the Core Hub TRUNCATE before snapshot option is honored

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among ScyllaDB CDC approaches

ApproachWhat buyers usually getWhere Gluesync fits
Kafka connector built on Debezium Change events published to Kafka topics; you run Kafka and Kafka Connect and own the connector, its offsets, and its upgrades A dedicated ScyllaDB agent with Core Hub operations and targets beyond Kafka. Read the Debezium alternative comparison
ScyllaDB CDC client libraries Library access to the CDC log for teams that want to write the consumer; you own parsing, checkpointing, and delivery The log reading, state, and delivery are productized in the source agent, and the same pipeline model covers SQL, warehouse, lake, and stream targets
Spark-based migration jobs Strong for bulk moves and one-off migrations; ongoing change handling is a separate build Snapshot plus CDC in one pipeline, so an initial load and the changes after it reach the target without a second tool. See cloud migration
Custom consumers that read CDC log tables Full control; you handle multi-row writes, collections, retries, checkpoints, and schema changes in your own code The agent reads the log rows and delivers the changes, with MOLO17 enterprise support instead of an internal on-call rotation

FAQ

ScyllaDB CDC questions

What is ScyllaDB CDC with Gluesync?

Change data capture from ScyllaDB by reading the CDC log tables that ScyllaDB maintains for each CDC-enabled table. The ScyllaDB agent connects through the official ScyllaDB Java driver and delivers inserts, updates, and deletes to configured targets through Core Hub.

Which ScyllaDB versions are supported?

The integration matrix lists ScyllaDB as tested from version 6.0, on ScyllaDB Cloud, on-premises, or a public cloud provider.

How is CDC enabled on a ScyllaDB table?

CDC is enabled per table with WITH cdc = {'enabled': true} when the table is created, or with an ALTER statement on an existing table. ScyllaDB then creates a log table named after the base table with the suffix _scylla_cdc_log in the same keyspace, and the agent reads that table.

Which changes does the ScyllaDB agent capture?

Inserts, updates, row deletions, range deletions, and partition deletions appear as CDC log rows, and the agent reads them partition by partition for incremental delivery.

What does the DBA have to enable on ScyllaDB?

Enable CDC on each source table, and give the Gluesync user read permission on those tables. The connection also needs the host, port (9042 by default), keyspace, username, password, and datacenter name. TLS is available with a certificate path.

Does Gluesync support ScyllaDB Cloud?

Yes. ScyllaDB Cloud, on-premises ScyllaDB, and ScyllaDB hosted under a public cloud provider are all supported deployments.

Can Gluesync write to ScyllaDB, not only read from it?

Yes. ScyllaDB is supported as a target over the official Java driver with cluster awareness, and the Core Hub TRUNCATE before snapshot option is honored when writing to ScyllaDB.

Where should we start?

Start a Gluesync trial on your infrastructure, then follow the ScyllaDB source setup guide. For help with prerequisites, sizing, and targets, talk to MOLO17.

Evaluate Gluesync with your ScyllaDB CDC log tables

Start a trial on your infrastructure, or talk to MOLO17 about prerequisites, versions, and targets for your ScyllaDB cluster.