SOLUTION · REPLICATE TO CLICKHOUSE
Real-time data replication to ClickHouse from your operational databases
Committed changes from your operational databases reach ClickHouse continuously, written as optimized multi-row INSERT batches, instead of waiting for the next extract.
Gluesync by MOLO17 captures changes from MySQL, PostgreSQL, MongoDB, Oracle, SQL Server, and other heterogeneous sources with a dedicated agent per database. The ClickHouse agent writes them through the ClickHouse JDBC driver over the HTTP endpoint, creates each table from the source definition when it does not exist, and keeps changes in commit order per entity. Core Hub, the Gluesync control plane, runs snapshots and CDC for every pipeline from one web UI and REST API.
WHO THIS IS FOR
Teams that need ClickHouse to reflect operations now, not after the next poll
- Analytics engineers building dashboards on ClickHouse who need tables keyed like the source, with column types generated from the source definition and
FINALreads that return the current row of each key - Platform leads running ClickHouse on-premises or on AWS, Google Cloud, or Azure who want one Core Hub for every source, backed by best-in-class enterprise support, rated 4.9/5 by customers
- Architects placing ClickHouse next to transactional systems: an agent on each source, one ClickHouse agent, and TLS to the HTTP endpoint
- Engineers replacing Kafka sinks or polling scripts that feed ClickHouse, who want every source delivered to ClickHouse by one product with snapshots and CDC under one control plane
THE PROBLEM
ClickHouse data that is polled on a timer is already stale
Most ClickHouse pipelines start as a script: a job reads the source on a timer, writes batches, and inserts them into ClickHouse. Dashboards then trail the last run, every poll adds read load to production, and each source adds another job to maintain. Teams looking for real-time ClickHouse ingestion or a ClickHouse CDC pipeline usually weigh ClickPipes, a Debezium and Kafka stack with a ClickHouse sink, managed ELT services such as Airbyte or Fivetran, or scripts they keep running themselves.
Gluesync addresses that with per-agent CDC into ClickHouse. A source agent reads each database's native change log or change stream, Core Hub routes the changes, and the ClickHouse agent writes them as optimized multi-row INSERT batches. MySQL to ClickHouse and PostgreSQL to ClickHouse follow the same model, with the same snapshot and operations.
HOW IT WORKS
How Gluesync writes to ClickHouse
Optimized batches, never row by row: multi-row INSERTs over HTTP
The ClickHouse agent connects with the ClickHouse JDBC driver, bundled with the agent, to the HTTP endpoint. It is a target agent: it applies the change stream from any Gluesync source agent to your database in real time, after a snapshot seeds each table. Gluesync never writes one row at a time. Changes are grouped into highly optimized batches, and each batch goes out as a multi-row INSERT ... VALUES statement, which is the shape ClickHouse ingests best. Transactions are written in source order, so a row inserted again later in the same transaction keeps its final state.
- Batch size: the rows per
INSERTare configurable per entity. - Compression: HTTP compression for requests and responses is on by default.
- Async inserts: optional, with a configurable flush interval.
- Truncate before snapshot: on by default, and it can be disabled in the target settings. ClickHouse agent overview ↗
Snapshot first, then continuous CDC
- Seed, then stream: snapshot batches populate each table, then changes follow from the source agent as they arrive, in commit order per entity.
- Scheduled reloads: the Chronos scheduler runs a snapshot, or a snapshot followed by CDC, on a cadence for tables you refresh on a timetable. See schedules and events.
Tables, keys, and column types
When a target table does not exist, the agent creates it from the source definition. A table with a primary key uses ReplacingMergeTree ordered by those keys, so FINAL returns the current row of each key. A table without a key uses MergeTree. A table you create with a ReplacingMergeTree engine and the reserved version columns receives versioned rows, and a table with another engine receives appended rows.
- Integers and floats map to
Int16,Int32,Int64,Float32, andFloat64, and decimals keep their source precision and scale asDecimal(P, S), sent as plain text rather than scientific notation. - Dates and times map to
Date32,DateTime64with the source's fractional precision, andDateTime64(6, 'UTC')for offset timestamps, which keep the instant. - Booleans, strings, and enums map to
BoolandString, with source enums written asLowCardinality(String); arrays and maps becomeArray(Nullable(T))andMap(String, Nullable(String)).
Setup in your ClickHouse deployment
The agent needs a database user with read and write access to the target database. Full statements are in the ClickHouse target setup guide ↗.
- Create a database user with read and write access to the target database and its tables, with rights to insert, truncate, read metadata, and create tables when Gluesync generates them.
- Note the HTTP endpoint, which listens on port
8123by default, and the database name. - In Core Hub, enter the host, port, database name, username, and password. The username and password are required.
- Turn on TLS for the connection when your endpoint serves it, and the agent connects over
https.
Architecture around Core Hub
Source agents sit close to each database, and Core Hub orchestrates them through its web UI and REST APIs. A pipeline groups a source agent with the ClickHouse agent and the tables it replicates. Query Studio opens ClickHouse on a dedicated connection that shows the current row of each key, so you can check a table's state in the same console. See Query Studio, and CDC streaming for the wider pattern.
WRITE OPTIONS
The ClickHouse target agent: one agent, batched INSERTs
ClickHouse has one Gluesync target agent. Snapshots and CDC both write optimized multi-row INSERT batches over the HTTP endpoint, so one agent covers every table in the pipeline.
| Agent | Write technique | Versions | Best for |
|---|---|---|---|
| ClickHouse agent ↗ | ClickHouse JDBC driver writing optimized multi-row INSERT ... VALUES batches over HTTP, with compression and optional async inserts | Any ClickHouse deployment, on-premises or DBaaS on AWS, Google Cloud, or Azure | Teams feeding ClickHouse from transactional databases who want ReplacingMergeTree tables keyed like the source, optimized INSERT batches with a configurable size and TLS under one Core Hub. |
SOURCES AND TOPOLOGIES
Feed ClickHouse from the databases you already run
Any Gluesync source agent can feed ClickHouse, each with its own native capture technique. Open the integrations finder with ClickHouse pre-selected to see every source you can pair with it.
One Core Hub runs MySQL to ClickHouse, PostgreSQL to ClickHouse, and MongoDB to ClickHouse side by side, with the same snapshot and write settings for each. Gluesync keeps pace with your change volume at any scale, and MOLO17 Professional Services can help with table design and batch settings.
- MySQL to ClickHouse from the binlog: see MySQL CDC
- PostgreSQL to ClickHouse from the write-ahead log: see PostgreSQL CDC
- MongoDB to ClickHouse from change streams: see MongoDB CDC
- Oracle and SQL Server to ClickHouse: see Oracle CDC and SQL Server CDC
FAIR, HIGH-LEVEL COMPARISON
Where Gluesync fits among ClickHouse ingestion approaches
| Approach | What buyers usually get | Where Gluesync fits |
|---|---|---|
| ClickPipes in ClickHouse Cloud | Ingestion managed by ClickHouse Cloud, scoped to that service's connector list and deployment model | Agents you run next to each database, with one Core Hub across every source and target; see CDC streaming |
| Kafka with a ClickHouse sink | Event streams consumed into ClickHouse; your team runs the brokers, the connectors, and the schema handling | Changes applied straight from each database agent to ClickHouse, with no Kafka cluster in the path. Kafka stays available as another target when you need it |
| Debezium with Kafka Connect | Open-source log-based capture, usually run through Kafka Connect; keys, offsets, and restarts are yours to operate | A native capture agent per database under one Core Hub, so operations sit in one control plane; read the Debezium alternative comparison |
| Airbyte and other ELT tools | Open-source and cloud ELT connectors that load ClickHouse on a sync schedule; change capture varies by source connector | Log-based capture with a commercial agent per source and MOLO17 enterprise support, writing continuously rather than on a sync schedule |
| Commercial replication suites | Mature replication products with many targets; ClickHouse coverage and capture methods vary by product | Native capture per engine and optimized INSERT batches into ClickHouse under one Core Hub; read batch ETL vs real-time replication |
| DIY polling scripts | Full control; your team owns polling queries, batching, retries, and the load each run puts on the source database | Log-based capture with no polling queries against production and no batching code to maintain; see migrating to Gluesync |
RELATED CONTENT
ClickHouse replication research and implementation detail
- Batch ETL vs real-time data replication: how to choose
- Log-based CDC: how real-time replication works
- CDC streaming without source overhead
- Debezium alternative: managed CDC versus Kafka + Debezium
- Keep cloud warehouses synchronized with operations
- ClickHouse agent overview ↗
- ClickHouse target setup guide ↗
- Query Studio ↗
FAQ
ClickHouse replication questions
What does replicating to ClickHouse with Gluesync involve?
A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the ClickHouse agent writes them to ClickHouse tables. Each table is seeded with a snapshot first and then follows the change stream.
How does Gluesync write data into ClickHouse?
Through the ClickHouse JDBC driver over the HTTP endpoint, in highly optimized batches: each batch is one multi-row INSERT statement, with a batch size you configure per entity. Snapshots populate each table first, then changes are applied as they arrive, in commit order per entity.
Which sources can replicate to ClickHouse?
Any Gluesync source agent. MySQL, PostgreSQL, MongoDB, Oracle, and SQL Server can all feed ClickHouse, and each one uses its own native capture technique.
Does Gluesync create the ClickHouse tables?
Yes. When a target table does not exist, the agent creates it from the source definition. A table with a primary key uses ReplacingMergeTree ordered by those keys, so FINAL returns the current row of each key.
Which ClickHouse deployments are supported?
Any ClickHouse deployment, whether it runs on-premises or as a managed service on AWS, Google Cloud, or Azure. The agent connects to the HTTP endpoint, port 8123 by default.
How do updates show up in ClickHouse tables?
Keyed tables that Gluesync creates use ReplacingMergeTree ordered by the source keys. Each insert or update is written as a new row with a strictly increasing _version, and a FINAL query returns the latest version of each key.
Can the connection to ClickHouse use TLS?
Yes. Turn on TLS for the connection in Core Hub and the agent connects over https. TLS is off by default.
REPLICATE TO A TARGET
Other targets Gluesync delivers to
- Replicate to Aerospike
- Replicate to DynamoDB
- Replicate to Redshift
- Replicate to Amazon S3 & S3-compatible
- Replicate to Cassandra
- Replicate to Kafka
- Replicate to Cosmos DB
- Replicate to Azure Data Lake
- Replicate to CockroachDB
- Replicate to Couchbase
- Replicate to file stores
- Replicate to BigQuery
- Replicate to Google Cloud Storage
- Replicate to Google Pub/Sub
- Replicate to GridGain
- Replicate to Db2
- Replicate to Informix
- Replicate to MariaDB
- Replicate to SQL Server
- Replicate to MongoDB
- Replicate to MySQL
- Replicate to Oracle
- Replicate to PostgreSQL
- Replicate to RavenDB
- Replicate to Redis
- Replicate to SAP ASE
- Replicate to SAP HANA
- Replicate to ScyllaDB
- Replicate to SingleStore
- Replicate to Snowflake
- Replicate to Solace PubSub+
- Replicate to Vertica
- Replicate to YugabyteDB
Evaluate Gluesync with your ClickHouse cluster
Start a trial against your self-managed or cloud ClickHouse, or talk to MOLO17 about your sources, table design, and insert volume.