SOLUTION · REPLICATE TO SCYLLADB
Real-time data replication to ScyllaDB from your operational databases
Committed changes from your systems of record reach ScyllaDB continuously, written in optimized batches through a cluster-aware driver, instead of waiting for the next migrator run or load job.
Gluesync by MOLO17 captures changes from Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, and other heterogeneous sources with a dedicated agent per database. The ScyllaDB target agent writes them through the official ScyllaDB Java driver, with TLS and load-balanced connections across nodes, in ScyllaDB Cloud, on-premises, or on a public cloud. Each table is seeded with a snapshot and then follows the change stream, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.
WHO THIS IS FOR
Teams that keep ScyllaDB current with the systems of record behind it
- Backend engineers building latency-sensitive services on ScyllaDB who need the operational data behind those services mirrored into it without a batch window
- Data platform leads who want one Core Hub for every source feeding ScyllaDB, backed by best-in-class enterprise support, rated 4.9/5 by customers
- Architects placing ScyllaDB in a topology with relational and document sources, deciding which agent runs next to which database and how the cluster is reached
- Engineers replacing a Kafka Connect sink, a custom CQL consumer, or a scheduled load script with one product for every source, with snapshots, checkpoints, and monitoring built in
THE PROBLEM
ScyllaDB data that moves in batches is already behind
ScyllaDB is chosen for throughput and low latency, and the data that matters most often lives in relational or document systems. Batch extracts leave the cluster working from a copy that is already behind, and bulk reads against production compete with the workload that relies on it. Buyers looking for a ScyllaDB CDC pipeline or real-time ScyllaDB ingestion usually weigh a Debezium and Kafka stack with a CQL sink connector, Spark-based migration jobs, or custom CQL consumers that they write and run themselves.
Gluesync addresses that with per-agent CDC into ScyllaDB. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the ScyllaDB target agent applies them through the ScyllaDB Java driver. Oracle to ScyllaDB, PostgreSQL to ScyllaDB, or MongoDB to ScyllaDB all share the same snapshot, pipeline model, and operations.
HOW IT WORKS
How Gluesync writes to ScyllaDB
The write path: the cluster-aware ScyllaDB Java driver
The ScyllaDB agent connects through the official ScyllaDB Java driver. The driver is cluster-aware, so requests reach the nodes that own the data.
- Topology: the driver supports TLS and load-balanced connections, and the
Datacenterproperty selects the datacenter the agent connects to in a multi-datacenter deployment. - Deployments: the same agent connects to ScyllaDB Cloud, to self-managed clusters, and to ScyllaDB hosted by a public cloud provider.
- ScyllaDB agent overview ↗
Optimized batches, never row by row
Gluesync never writes one row at a time to ScyllaDB. Changes are grouped into chunks before they reach the cluster, and the cluster-aware driver sends each request to the nodes that own the data, so round-trips stay few and the write load stays predictable next to the latency-sensitive traffic your cluster serves. The chunk size is configurable on the target agent, so you can push throughput during the initial load and keep requests small once production traffic peaks.
Snapshot first, then continuous CDC
- Seed: snapshot batches are inserted through the driver before the change stream starts, so each entity begins from a full copy of its source table.
- Stream: after the snapshot, inserts, updates, and deletes from the source change stream are applied in real time.
- Clean reloads: with the Core Hub setting for TRUNCATE before snapshot, the target table is cleared before the initial load, so a reload starts from an empty table.
- Resume: an interrupted snapshot resumes from its last saved state when you start the entity again.
Keys, identifiers, and table creation
- Identifier case: ScyllaDB folds unquoted identifiers to lower case, so Core Hub normalizes target column names and filter clauses to lower case when an entity is saved.
- Table creation: when a target table does not exist, Core Hub generates the
CREATE TABLEstatement from the source columns and primary key, using the lower-case names ScyllaDB expects. - Keys: a duplicate key cannot arise on ScyllaDB, because a write replaces the row stored under the same primary key. The table converges on the latest state of the source row, and replayed changes land as the same row.
What your ScyllaDB admin sets up
The agent needs a user with read and write access to the target tables and keyspace. The ScyllaDB target setup guide ↗ lists every field.
- Create or choose a user with read and write permission on the target keyspace and its tables.
- Note the host or IP address, the port (9042 by default), the keyspace name, and the datacenter the agent connects to.
- When the cluster requires TLS, switch it on for the connection and provide the certificate path (
enableTlsandcertificatePathin the REST configuration). - In Core Hub, add the ScyllaDB target agent with those credentials, then attach it to a pipeline with one or more source agents.
Architecture around Core Hub
Source agents sit close to each database, the ScyllaDB agent sits close to the cluster, and Core Hub coordinates them through its web UI and REST API. Every component deploys with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud. See CDC streaming without source overhead for the wider pattern.
WRITE OPTIONS
The ScyllaDB target agent: one agent for every entity
ScyllaDB has one Gluesync target agent. Every entity written to it uses the same driver path: a snapshot to seed the table, then the change stream applied as it arrives.
| Agent | Write technique | Versions | Best for |
|---|---|---|---|
| ScyllaDB agent ↗ | Official ScyllaDB Java driver, cluster-aware; optimized batches for snapshot inserts and real-time changes | Tested from ScyllaDB 6.0; ScyllaDB Cloud, on-premises, or on a public cloud | Latency-sensitive ScyllaDB clusters that need changes from relational or document systems applied continuously. Pick it when the cluster needs TLS and load-balanced connections across its nodes. |
SOURCES AND TOPOLOGIES
Feed ScyllaDB from the systems of record you already run
Any Gluesync source agent can feed ScyllaDB, each with its own native capture technique. Open the integrations finder with ScyllaDB pre-selected to see every source you can pair with it.
Several sources can feed one cluster through one Core Hub, each pipeline with its own snapshot and monitoring. Gluesync keeps pace with your change volume at any scale, and the batch size and target datacenter are configurable per agent. MOLO17 Professional Services can help plan keyspaces and agent placement.
- Oracle to ScyllaDB from the redo logs through LogMiner or XStream: see Oracle CDC
- PostgreSQL to ScyllaDB from the write-ahead log: see PostgreSQL CDC
- MongoDB to ScyllaDB from Change Streams: see MongoDB CDC
- SQL Server to ScyllaDB through Change Data Capture or Change Tracking: see SQL Server CDC
FAIR, HIGH-LEVEL COMPARISON
Where Gluesync fits among ScyllaDB ingestion approaches
| Approach | What buyers usually get | Where Gluesync fits |
|---|---|---|
| Debezium, Kafka, and a Kafka Connect sink | Open-source capture into Kafka topics, then a sink connector writes to the cluster; you run Kafka and Connect, manage offsets, and decide how change events map to rows | Changes applied to ScyllaDB tables from each source agent, with no Kafka cluster in the path. Read the Debezium alternative comparison |
| Fivetran and Fivetran HVR | Managed ELT and log-based replication with a broad source catalog, oriented toward warehouses and lakes | Agents installed next to each source, native capture per engine, and one Core Hub for every pipeline; see CDC streaming |
| Airbyte | Open-source and cloud ELT connectors; incremental and CDC modes vary by connector, and you run the platform or use the managed service | A commercial product with a dedicated capture agent per database and MOLO17 enterprise support behind every pipeline |
| Qlik Replicate, AWS DMS, and similar replication services | Mature replication across many engines; target lists and load behavior differ by product | Agents run on-premises or in any cloud with Docker, Docker Compose, or Kubernetes and write straight to ScyllaDB under one Core Hub. See migrating to Gluesync |
| Custom CQL consumers and scheduled scripts | Full control; your team owns the consumer, its offsets, retries, batching, and the load each run puts on the cluster | Log-based capture and batched, cluster-aware writes without consumer code to maintain, with Core Hub monitoring each pipeline |
RELATED CONTENT
ScyllaDB replication background and implementation detail
- New ScyllaDB agent for Gluesync
- Introduction to data integration: a comprehensive guide
- ScyllaDB CDC: capture changes from ScyllaDB
- CDC streaming without source overhead
- Debezium alternative: managed CDC versus Kafka + Debezium
- ScyllaDB agent overview ↗
- ScyllaDB target setup guide ↗
- Write strategies: optimized batches ↗
FAQ
ScyllaDB replication questions
What does replicating to ScyllaDB with Gluesync involve?
A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the ScyllaDB target agent applies them to your tables. Each table is seeded with a snapshot first, then follows the change stream continuously.
How does Gluesync write data to ScyllaDB?
Through the official ScyllaDB Java driver, which is cluster-aware and supports TLS and load-balanced connections. Snapshot rows and live changes are grouped into optimized batches whose size is configurable on the target agent.
Which sources can replicate to ScyllaDB?
Any Gluesync source agent, including Oracle, PostgreSQL, MySQL, SQL Server, MongoDB, and YugabyteDB. The integrations finder on our website lists every source you can pair with ScyllaDB.
Which ScyllaDB versions does the agent support?
ScyllaDB is tested from version 6.0. The agent connects to ScyllaDB Cloud, self-managed clusters, and ScyllaDB hosted by a public cloud provider.
What does the ScyllaDB admin need to set up?
A user with read and write permission on the target keyspace and tables, the host or IP address, the port (9042 by default), the keyspace name, and the datacenter the agent connects to. Enable TLS with a certificate path when the cluster requires it.
What happens when a row already exists in ScyllaDB?
The incoming write replaces the row stored under the same primary key, so the table converges on the latest state of the source row and a replayed change never fails on a duplicate.
Does Gluesync create the tables in ScyllaDB?
Yes. Core Hub generates the CREATE TABLE statement from the source columns and primary key, with lower-case column names to match how ScyllaDB folds unquoted identifiers.
Does TRUNCATE before snapshot work with ScyllaDB?
Yes. When ScyllaDB is the target, Core Hub's TRUNCATE before snapshot behavior clears the target table before the initial load.
REPLICATE TO A TARGET
Other targets Gluesync delivers to
- Replicate to Aerospike
- Replicate to DynamoDB
- Replicate to Redshift
- Replicate to Amazon S3 & S3-compatible
- Replicate to Cassandra
- Replicate to Kafka
- Replicate to Cosmos DB
- Replicate to Azure Data Lake
- Replicate to ClickHouse
- Replicate to CockroachDB
- Replicate to Couchbase
- Replicate to file stores
- Replicate to BigQuery
- Replicate to Google Cloud Storage
- Replicate to Google Pub/Sub
- Replicate to GridGain
- Replicate to Db2
- Replicate to Informix
- Replicate to MariaDB
- Replicate to SQL Server
- Replicate to MongoDB
- Replicate to MySQL
- Replicate to Oracle
- Replicate to PostgreSQL
- Replicate to RavenDB
- Replicate to Redis
- Replicate to SAP ASE
- Replicate to SAP HANA
- Replicate to SingleStore
- Replicate to Snowflake
- Replicate to Solace PubSub+
- Replicate to Vertica
- Replicate to YugabyteDB
Evaluate Gluesync with your ScyllaDB cluster
Start a trial on your infrastructure, or talk to MOLO17 about your sources, keyspaces, datacenters, and agent placement.