SOLUTION · REPLICATE TO CLICKHOUSE

Real-time data replication to ClickHouse from your operational databases

Committed changes from your operational databases reach ClickHouse continuously, written as optimized multi-row INSERT batches, instead of waiting for the next extract.

Gluesync by MOLO17 captures changes from MySQL, PostgreSQL, MongoDB, Oracle, SQL Server, and other heterogeneous sources with a dedicated agent per database. The ClickHouse agent writes them through the ClickHouse JDBC driver over the HTTP endpoint, creates each table from the source definition when it does not exist, and keeps changes in commit order per entity. Core Hub, the Gluesync control plane, runs snapshots and CDC for every pipeline from one web UI and REST API.

WHO THIS IS FOR

Teams that need ClickHouse to reflect operations now, not after the next poll

  • Analytics engineers building dashboards on ClickHouse who need tables keyed like the source, with column types generated from the source definition and FINAL reads that return the current row of each key
  • Platform leads running ClickHouse on-premises or on AWS, Google Cloud, or Azure who want one Core Hub for every source, backed by best-in-class enterprise support, rated 4.9/5 by customers
  • Architects placing ClickHouse next to transactional systems: an agent on each source, one ClickHouse agent, and TLS to the HTTP endpoint
  • Engineers replacing Kafka sinks or polling scripts that feed ClickHouse, who want every source delivered to ClickHouse by one product with snapshots and CDC under one control plane

THE PROBLEM

ClickHouse data that is polled on a timer is already stale

Most ClickHouse pipelines start as a script: a job reads the source on a timer, writes batches, and inserts them into ClickHouse. Dashboards then trail the last run, every poll adds read load to production, and each source adds another job to maintain. Teams looking for real-time ClickHouse ingestion or a ClickHouse CDC pipeline usually weigh ClickPipes, a Debezium and Kafka stack with a ClickHouse sink, managed ELT services such as Airbyte or Fivetran, or scripts they keep running themselves.

Gluesync addresses that with per-agent CDC into ClickHouse. A source agent reads each database's native change log or change stream, Core Hub routes the changes, and the ClickHouse agent writes them as optimized multi-row INSERT batches. MySQL to ClickHouse and PostgreSQL to ClickHouse follow the same model, with the same snapshot and operations.

HOW IT WORKS

How Gluesync writes to ClickHouse

Optimized batches, never row by row: multi-row INSERTs over HTTP

The ClickHouse agent connects with the ClickHouse JDBC driver, bundled with the agent, to the HTTP endpoint. It is a target agent: it applies the change stream from any Gluesync source agent to your database in real time, after a snapshot seeds each table. Gluesync never writes one row at a time. Changes are grouped into highly optimized batches, and each batch goes out as a multi-row INSERT ... VALUES statement, which is the shape ClickHouse ingests best. Transactions are written in source order, so a row inserted again later in the same transaction keeps its final state.

  • Batch size: the rows per INSERT are configurable per entity.
  • Compression: HTTP compression for requests and responses is on by default.
  • Async inserts: optional, with a configurable flush interval.
  • Truncate before snapshot: on by default, and it can be disabled in the target settings. ClickHouse agent overview ↗

Snapshot first, then continuous CDC

  • Seed, then stream: snapshot batches populate each table, then changes follow from the source agent as they arrive, in commit order per entity.
  • Scheduled reloads: the Chronos scheduler runs a snapshot, or a snapshot followed by CDC, on a cadence for tables you refresh on a timetable. See schedules and events.

Tables, keys, and column types

When a target table does not exist, the agent creates it from the source definition. A table with a primary key uses ReplacingMergeTree ordered by those keys, so FINAL returns the current row of each key. A table without a key uses MergeTree. A table you create with a ReplacingMergeTree engine and the reserved version columns receives versioned rows, and a table with another engine receives appended rows.

  • Integers and floats map to Int16, Int32, Int64, Float32, and Float64, and decimals keep their source precision and scale as Decimal(P, S), sent as plain text rather than scientific notation.
  • Dates and times map to Date32, DateTime64 with the source's fractional precision, and DateTime64(6, 'UTC') for offset timestamps, which keep the instant.
  • Booleans, strings, and enums map to Bool and String, with source enums written as LowCardinality(String); arrays and maps become Array(Nullable(T)) and Map(String, Nullable(String)).

Setup in your ClickHouse deployment

The agent needs a database user with read and write access to the target database. Full statements are in the ClickHouse target setup guide ↗.

  1. Create a database user with read and write access to the target database and its tables, with rights to insert, truncate, read metadata, and create tables when Gluesync generates them.
  2. Note the HTTP endpoint, which listens on port 8123 by default, and the database name.
  3. In Core Hub, enter the host, port, database name, username, and password. The username and password are required.
  4. Turn on TLS for the connection when your endpoint serves it, and the agent connects over https.

Architecture around Core Hub

Source agents sit close to each database, and Core Hub orchestrates them through its web UI and REST APIs. A pipeline groups a source agent with the ClickHouse agent and the tables it replicates. Query Studio opens ClickHouse on a dedicated connection that shows the current row of each key, so you can check a table's state in the same console. See Query Studio, and CDC streaming for the wider pattern.

Explore the general CDC streaming architecture →

WRITE OPTIONS

The ClickHouse target agent: one agent, batched INSERTs

ClickHouse has one Gluesync target agent. Snapshots and CDC both write optimized multi-row INSERT batches over the HTTP endpoint, so one agent covers every table in the pipeline.

AgentWrite techniqueVersionsBest for
ClickHouse agent ↗ ClickHouse JDBC driver writing optimized multi-row INSERT ... VALUES batches over HTTP, with compression and optional async inserts Any ClickHouse deployment, on-premises or DBaaS on AWS, Google Cloud, or Azure Teams feeding ClickHouse from transactional databases who want ReplacingMergeTree tables keyed like the source, optimized INSERT batches with a configurable size and TLS under one Core Hub.

SOURCES AND TOPOLOGIES

Feed ClickHouse from the databases you already run

Any Gluesync source agent can feed ClickHouse, each with its own native capture technique. Open the integrations finder with ClickHouse pre-selected to see every source you can pair with it.

One Core Hub runs MySQL to ClickHouse, PostgreSQL to ClickHouse, and MongoDB to ClickHouse side by side, with the same snapshot and write settings for each. Gluesync keeps pace with your change volume at any scale, and MOLO17 Professional Services can help with table design and batch settings.

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among ClickHouse ingestion approaches

ApproachWhat buyers usually getWhere Gluesync fits
ClickPipes in ClickHouse Cloud Ingestion managed by ClickHouse Cloud, scoped to that service's connector list and deployment model Agents you run next to each database, with one Core Hub across every source and target; see CDC streaming
Kafka with a ClickHouse sink Event streams consumed into ClickHouse; your team runs the brokers, the connectors, and the schema handling Changes applied straight from each database agent to ClickHouse, with no Kafka cluster in the path. Kafka stays available as another target when you need it
Debezium with Kafka Connect Open-source log-based capture, usually run through Kafka Connect; keys, offsets, and restarts are yours to operate A native capture agent per database under one Core Hub, so operations sit in one control plane; read the Debezium alternative comparison
Airbyte and other ELT tools Open-source and cloud ELT connectors that load ClickHouse on a sync schedule; change capture varies by source connector Log-based capture with a commercial agent per source and MOLO17 enterprise support, writing continuously rather than on a sync schedule
Commercial replication suites Mature replication products with many targets; ClickHouse coverage and capture methods vary by product Native capture per engine and optimized INSERT batches into ClickHouse under one Core Hub; read batch ETL vs real-time replication
DIY polling scripts Full control; your team owns polling queries, batching, retries, and the load each run puts on the source database Log-based capture with no polling queries against production and no batching code to maintain; see migrating to Gluesync

FAQ

ClickHouse replication questions

What does replicating to ClickHouse with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the ClickHouse agent writes them to ClickHouse tables. Each table is seeded with a snapshot first and then follows the change stream.

How does Gluesync write data into ClickHouse?

Through the ClickHouse JDBC driver over the HTTP endpoint, in highly optimized batches: each batch is one multi-row INSERT statement, with a batch size you configure per entity. Snapshots populate each table first, then changes are applied as they arrive, in commit order per entity.

Which sources can replicate to ClickHouse?

Any Gluesync source agent. MySQL, PostgreSQL, MongoDB, Oracle, and SQL Server can all feed ClickHouse, and each one uses its own native capture technique.

Does Gluesync create the ClickHouse tables?

Yes. When a target table does not exist, the agent creates it from the source definition. A table with a primary key uses ReplacingMergeTree ordered by those keys, so FINAL returns the current row of each key.

Which ClickHouse deployments are supported?

Any ClickHouse deployment, whether it runs on-premises or as a managed service on AWS, Google Cloud, or Azure. The agent connects to the HTTP endpoint, port 8123 by default.

How do updates show up in ClickHouse tables?

Keyed tables that Gluesync creates use ReplacingMergeTree ordered by the source keys. Each insert or update is written as a new row with a strictly increasing _version, and a FINAL query returns the latest version of each key.

Can the connection to ClickHouse use TLS?

Yes. Turn on TLS for the connection in Core Hub and the agent connects over https. TLS is off by default.

Evaluate Gluesync with your ClickHouse cluster

Start a trial against your self-managed or cloud ClickHouse, or talk to MOLO17 about your sources, table design, and insert volume.