<!-- Generated from the rendered page by scripts/write-llm-mirrors.mjs. Do not edit by hand. -->
Canonical: https://molo17.com/solutions/replicate-to-clickhouse/
Markdown mirror: https://molo17.com/solutions/replicate-to-clickhouse/index.md
Title: Replicate to ClickHouse in real time with Gluesync | MOLO17
Description: Real-time replication to ClickHouse from MySQL, PostgreSQL, MongoDB, Oracle, and SQL Server, with batched multi-row INSERTs and per-agent CDC. Start a trial.

[Solutions](/solutions/)  Replicate to ClickHouse

 SOLUTION · REPLICATE TO CLICKHOUSE

# Real-time data replication to ClickHouse from your operational databases

**Committed changes from your operational databases reach ClickHouse continuously, written as optimized multi-row INSERT batches, instead of waiting for the next extract.**

Gluesync by MOLO17 captures changes from MySQL, PostgreSQL, MongoDB, Oracle, SQL Server, and other [heterogeneous sources](/integrations/?target=ClickHouse#integration-finder) with a dedicated agent per database. The ClickHouse agent writes them through the ClickHouse JDBC driver over the HTTP endpoint, creates each table from the source definition when it does not exist, and keeps changes in commit order per entity. Core Hub, the Gluesync control plane, runs snapshots and CDC for every pipeline from one web UI and REST API.

[Start a Gluesync trial](/get-gluesync/) [Talk to us](/contacts/)

WHO THIS IS FOR

## Teams that need ClickHouse to reflect operations now, not after the next poll

-   Analytics engineers building dashboards on ClickHouse who need tables keyed like the source, with column types generated from the source definition and `FINAL` reads that return the current row of each key
-   Platform leads running ClickHouse on-premises or on AWS, Google Cloud, or Azure who want one Core Hub for every source, backed by [best-in-class enterprise support, rated 4.9/5 by customers](/support/#customer-ratings)
-   Architects placing ClickHouse next to transactional systems: an agent on each source, one ClickHouse agent, and TLS to the HTTP endpoint
-   Engineers replacing Kafka sinks or polling scripts that feed ClickHouse, who want [every source](/integrations/?target=ClickHouse#integration-finder) delivered to ClickHouse by one product with snapshots and CDC under one control plane

THE PROBLEM

## ClickHouse data that is polled on a timer is already stale

Most ClickHouse pipelines start as a script: a job reads the source on a timer, writes batches, and inserts them into ClickHouse. Dashboards then trail the last run, every poll adds read load to production, and each source adds another job to maintain. Teams looking for real-time ClickHouse ingestion or a ClickHouse CDC pipeline usually weigh ClickPipes, a Debezium and Kafka stack with a ClickHouse sink, managed ELT services such as Airbyte or Fivetran, or scripts they keep running themselves.

Gluesync addresses that with **per-agent CDC into ClickHouse**. A source agent reads each database's native change log or change stream, Core Hub routes the changes, and the ClickHouse agent writes them as optimized multi-row INSERT batches. MySQL to ClickHouse and PostgreSQL to ClickHouse follow the same model, with the same snapshot and operations.

HOW IT WORKS

## How Gluesync writes to ClickHouse

### Optimized batches, never row by row: multi-row INSERTs over HTTP

The ClickHouse agent connects with the ClickHouse JDBC driver, bundled with the agent, to the HTTP endpoint. It is a target agent: it applies the change stream from any Gluesync source agent to your database in real time, after a snapshot seeds each table. Gluesync never writes one row at a time. Changes are grouped into highly optimized batches, and each batch goes out as a multi-row `INSERT ... VALUES` statement, which is the shape ClickHouse ingests best. Transactions are written in source order, so a row inserted again later in the same transaction keeps its final state.

-   **Batch size:** the rows per `INSERT` are configurable per entity.
-   **Compression:** HTTP compression for requests and responses is on by default.
-   **Async inserts:** optional, with a configurable flush interval.
-   **Truncate before snapshot:** on by default, and it can be disabled in the target settings. [ClickHouse agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/clickhouse-intro.html)

### Snapshot first, then continuous CDC

-   **Seed, then stream:** snapshot batches populate each table, then changes follow from the source agent as they arrive, in commit order per entity.
-   **Scheduled reloads:** the Chronos scheduler runs a snapshot, or a snapshot followed by CDC, on a cadence for tables you refresh on a timetable. See [schedules and events](/schedules-and-events/).

### Tables, keys, and column types

When a target table does not exist, the agent creates it from the source definition. A table with a primary key uses `ReplacingMergeTree` ordered by those keys, so `FINAL` returns the current row of each key. A table without a key uses `MergeTree`. A table you create with a `ReplacingMergeTree` engine and the reserved version columns receives versioned rows, and a table with another engine receives appended rows.

-   **Integers and floats** map to `Int16`, `Int32`, `Int64`, `Float32`, and `Float64`, and **decimals** keep their source precision and scale as `Decimal(P, S)`, sent as plain text rather than scientific notation.
-   **Dates and times** map to `Date32`, `DateTime64` with the source's fractional precision, and `DateTime64(6, 'UTC')` for offset timestamps, which keep the instant.
-   **Booleans, strings, and enums** map to `Bool` and `String`, with source enums written as `LowCardinality(String)`; arrays and maps become `Array(Nullable(T))` and `Map(String, Nullable(String))`.

### Setup in your ClickHouse deployment

The agent needs a database user with read and write access to the target database. Full statements are in the [ClickHouse target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/clickhouse-target.html).

1.  Create a database user with read and write access to the target database and its tables, with rights to insert, truncate, read metadata, and create tables when Gluesync generates them.
2.  Note the HTTP endpoint, which listens on port `8123` by default, and the database name.
3.  In Core Hub, enter the host, port, database name, username, and password. The username and password are required.
4.  Turn on TLS for the connection when your endpoint serves it, and the agent connects over `https`.

### Architecture around Core Hub

Source agents sit close to each database, and Core Hub orchestrates them through its web UI and REST APIs. A pipeline groups a source agent with the ClickHouse agent and the tables it replicates. Query Studio opens ClickHouse on a dedicated connection that shows the current row of each key, so you can check a table's state in the same console. See [Query Studio](/query-studio/), and [CDC streaming](/solutions/cdc-streaming/) for the wider pattern.

[Explore the general CDC streaming architecture →](/solutions/cdc-streaming/)

WRITE OPTIONS

## The ClickHouse target agent: one agent, batched INSERTs

ClickHouse has one Gluesync target agent. Snapshots and CDC both write optimized multi-row INSERT batches over the HTTP endpoint, so one agent covers every table in the pipeline.

| Agent | Write technique | Versions | Best for |
| --- | --- | --- | --- |
| [ClickHouse agent ↗](https://docs.molo17.com/gluesync/latest/agents/clickhouse-intro.html) | ClickHouse JDBC driver writing optimized multi-row INSERT ... VALUES batches over HTTP, with compression and optional async inserts | Any ClickHouse deployment, on-premises or DBaaS on AWS, Google Cloud, or Azure | Teams feeding ClickHouse from transactional databases who want `ReplacingMergeTree` tables keyed like the source, optimized INSERT batches with a configurable size and TLS under one Core Hub. |

SOURCES AND TOPOLOGIES

## Feed ClickHouse from the databases you already run

Any Gluesync source agent can feed ClickHouse, each with its own native capture technique. Open the [integrations finder with ClickHouse pre-selected](/integrations/?target=ClickHouse#integration-finder) to see every source you can pair with it.

One Core Hub runs MySQL to ClickHouse, PostgreSQL to ClickHouse, and MongoDB to ClickHouse side by side, with the same snapshot and write settings for each. Gluesync keeps pace with your change volume at any scale, and [MOLO17 Professional Services](/solutions/professional-services/) can help with table design and batch settings.

-   MySQL to ClickHouse from the binlog: see [MySQL CDC](/solutions/mysql-cdc/)
-   PostgreSQL to ClickHouse from the write-ahead log: see [PostgreSQL CDC](/solutions/postgresql-cdc/)
-   MongoDB to ClickHouse from change streams: see [MongoDB CDC](/solutions/mongodb-cdc/)
-   Oracle and SQL Server to ClickHouse: see [Oracle CDC](/solutions/oracle-cdc/) and [SQL Server CDC](/solutions/sql-server-cdc/)

FAIR, HIGH-LEVEL COMPARISON

## Where Gluesync fits among ClickHouse ingestion approaches

| Approach | What buyers usually get | Where Gluesync fits |
| --- | --- | --- |
| ClickPipes in ClickHouse Cloud | Ingestion managed by ClickHouse Cloud, scoped to that service's connector list and deployment model | Agents you run next to each database, with one Core Hub across every source and target; see [CDC streaming](/solutions/cdc-streaming/) |
| Kafka with a ClickHouse sink | Event streams consumed into ClickHouse; your team runs the brokers, the connectors, and the schema handling | Changes applied straight from each database agent to ClickHouse, with no Kafka cluster in the path. Kafka stays available as another target when you need it |
| Debezium with Kafka Connect | Open-source log-based capture, usually run through Kafka Connect; keys, offsets, and restarts are yours to operate | A native capture agent per database under one Core Hub, so operations sit in one control plane; read the [Debezium alternative](/solutions/debezium-alternative/) comparison |
| Airbyte and other ELT tools | Open-source and cloud ELT connectors that load ClickHouse on a sync schedule; change capture varies by source connector | Log-based capture with a commercial agent per source and MOLO17 enterprise support, writing continuously rather than on a sync schedule |
| Commercial replication suites | Mature replication products with many targets; ClickHouse coverage and capture methods vary by product | Native capture per engine and optimized INSERT batches into ClickHouse under one Core Hub; read [batch ETL vs real-time replication](/blog/batch-etl-vs-real-time-data-replication/) |
| DIY polling scripts | Full control; your team owns polling queries, batching, retries, and the load each run puts on the source database | Log-based capture with no polling queries against production and no batching code to maintain; see [migrating to Gluesync](/migrate-to-gluesync/) |

RELATED CONTENT

## ClickHouse replication research and implementation detail

-   [Batch ETL vs real-time data replication: how to choose](/blog/batch-etl-vs-real-time-data-replication/)
-   [Log-based CDC: how real-time replication works](/blog/real-time-replication-log-based-cdc/)
-   [CDC streaming without source overhead](/solutions/cdc-streaming/)
-   [Debezium alternative: managed CDC versus Kafka + Debezium](/solutions/debezium-alternative/)
-   [Keep cloud warehouses synchronized with operations](/solutions/warehouse-sync/)
-   [ClickHouse agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/clickhouse-intro.html)
-   [ClickHouse target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/clickhouse-target.html)
-   [Query Studio ↗](https://docs.molo17.com/gluesync/latest/gs-modules/query-studio.html)

FAQ

## ClickHouse replication questions

What does replicating to ClickHouse with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the ClickHouse agent writes them to ClickHouse tables. Each table is seeded with a snapshot first and then follows the change stream.

How does Gluesync write data into ClickHouse?

Through the ClickHouse JDBC driver over the HTTP endpoint, in highly optimized batches: each batch is one multi-row INSERT statement, with a batch size you configure per entity. Snapshots populate each table first, then changes are applied as they arrive, in commit order per entity.

Which sources can replicate to ClickHouse?

Any Gluesync source agent. MySQL, PostgreSQL, MongoDB, Oracle, and SQL Server can all feed ClickHouse, and each one uses its own native capture technique.

Does Gluesync create the ClickHouse tables?

Yes. When a target table does not exist, the agent creates it from the source definition. A table with a primary key uses ReplacingMergeTree ordered by those keys, so FINAL returns the current row of each key.

Which ClickHouse deployments are supported?

Any ClickHouse deployment, whether it runs on-premises or as a managed service on AWS, Google Cloud, or Azure. The agent connects to the HTTP endpoint, port 8123 by default.

How do updates show up in ClickHouse tables?

Keyed tables that Gluesync creates use ReplacingMergeTree ordered by the source keys. Each insert or update is written as a new row with a strictly increasing \_version, and a FINAL query returns the latest version of each key.

Can the connection to ClickHouse use TLS?

Yes. Turn on TLS for the connection in Core Hub and the agent connects over https. TLS is off by default.

REPLICATE TO A TARGET

## Other targets Gluesync delivers to

-    [Replicate to Aerospike](/solutions/replicate-to-aerospike/)
-    [Replicate to DynamoDB](/solutions/replicate-to-dynamodb/)
-    [Replicate to Redshift](/solutions/replicate-to-redshift/)
-    [Replicate to Amazon S3 & S3-compatible](/solutions/replicate-to-amazon-s3/)
-    [Replicate to Cassandra](/solutions/replicate-to-cassandra/)
-    [Replicate to Kafka](/solutions/replicate-to-kafka/)
-    [Replicate to Cosmos DB](/solutions/replicate-to-cosmos-db/)
-    [Replicate to Azure Data Lake](/solutions/replicate-to-azure-data-lake/)
-    [Replicate to CockroachDB](/solutions/replicate-to-cockroachdb/)
-    [Replicate to Couchbase](/solutions/replicate-to-couchbase/)
-    [Replicate to file stores](/solutions/replicate-to-file-stores/)
-    [Replicate to BigQuery](/solutions/replicate-to-bigquery/)
-    [Replicate to Google Cloud Storage](/solutions/replicate-to-google-cloud-storage/)
-    [Replicate to Google Pub/Sub](/solutions/replicate-to-google-pubsub/)
-    [Replicate to GridGain](/solutions/replicate-to-gridgain/)
-    [Replicate to Db2](/solutions/replicate-to-db2/)
-    [Replicate to Informix](/solutions/replicate-to-informix/)
-    [Replicate to MariaDB](/solutions/replicate-to-mariadb/)
-    [Replicate to SQL Server](/solutions/replicate-to-sql-server/)
-    [Replicate to MongoDB](/solutions/replicate-to-mongodb/)
-    [Replicate to MySQL](/solutions/replicate-to-mysql/)
-    [Replicate to Oracle](/solutions/replicate-to-oracle/)
-    [Replicate to PostgreSQL](/solutions/replicate-to-postgresql/)
-    [Replicate to RavenDB](/solutions/replicate-to-ravendb/)
-    [Replicate to Redis](/solutions/replicate-to-redis/)
-    [Replicate to SAP ASE](/solutions/replicate-to-sap-ase/)
-    [Replicate to SAP HANA](/solutions/replicate-to-sap-hana/)
-    [Replicate to ScyllaDB](/solutions/replicate-to-scylladb/)
-    [Replicate to SingleStore](/solutions/replicate-to-singlestore/)
-    [Replicate to Snowflake](/solutions/replicate-to-snowflake/)
-    [Replicate to Solace PubSub+](/solutions/replicate-to-solace/)
-    [Replicate to Vertica](/solutions/replicate-to-vertica/)
-    [Replicate to YugabyteDB](/solutions/replicate-to-yugabytedb/)

## Evaluate Gluesync with your ClickHouse cluster

Start a trial against your self-managed or cloud ClickHouse, or talk to MOLO17 about your sources, table design, and insert volume.

[Start a Gluesync trial](/get-gluesync/) [Talk to MOLO17](/contacts/)
