<!-- Generated from the rendered page by scripts/write-llm-mirrors.mjs. Do not edit by hand. -->
Canonical: https://molo17.com/solutions/replicate-to-google-cloud-storage/
Markdown mirror: https://molo17.com/solutions/replicate-to-google-cloud-storage/index.md
Title: Replicate to Google Cloud Storage in real time | MOLO17
Description: Real-time CDC replication to Google Cloud Storage from Oracle, PostgreSQL, SQL Server, MySQL, and MongoDB as Parquet, JSON, or CSV files. Start a trial.

[Solutions](/solutions/)  Replicate to Google Cloud Storage

 SOLUTION · REPLICATE TO GOOGLE CLOUD STORAGE

# Real-time data replication to Google Cloud Storage

**Committed changes from your databases land in your Google Cloud buckets as files continuously, instead of waiting for the next batch export.**

Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MySQL, MongoDB, and other [heterogeneous sources](/integrations/?target=Google+Cloud+Storage#integration-finder) with a dedicated agent per database. The Google Cloud Storage target agent writes each entity to your bucket through the native Google Cloud SDK, after an initial snapshot and then continuously, as Parquet by default with JSON or CSV selectable per entity. Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

[Start a Gluesync trial](/get-gluesync/) [Talk to us](/contacts/)

WHO THIS IS FOR

## Teams that want Google Cloud storage to reflect operations now

-   Analytics engineers who load Cloud Storage files into BigQuery or Spark and need every change tagged with its operation and transaction identifiers, so current-state tables can be rebuilt from the history
-   Data platform leads who feed one Google Cloud estate from several databases and want one Core Hub for every source, backed by [best-in-class enterprise support, rated 4.9/5 by customers](/support/#customer-ratings)
-   Architects designing a landing zone who need service account authentication, Parquet by default, and per-entity JSON or CSV where downstream tools read flat files
-   Engineers running export scripts or Kafka sinks into Cloud Storage who want [every source](/integrations/?target=Google+Cloud+Storage#integration-finder) delivered by one product, with snapshots, resume, and monitoring built in

THE PROBLEM

## Cloud Storage fed by exports is always behind the source

Most Google Cloud landing zones are fed by scheduled exports: a job queries each database on a timer, writes files to a bucket, and a loader picks them up. The bucket reflects the business as of the last export, every export adds read load to the production database, and each source has its own script and failure mode. Teams searching for real-time Cloud Storage ingestion or a CDC pipeline to GCS usually weigh Google's own change capture services, managed ELT connectors such as Fivetran or Airbyte, a Debezium and Kafka Connect sink, or scripts they keep running themselves.

Gluesync addresses that with **per-agent CDC into Google Cloud Storage**. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the Google Cloud Storage target agent writes them to your bucket as files. The pipeline model, the snapshot, and the operations stay the same whether the source is Oracle, SQL Server, or MongoDB.

HOW IT WORKS

## How Gluesync writes to Google Cloud Storage

### The write path: Google Cloud's native SDK, in grouped uploads

The Cloud Storage agent streams snapshot and CDC batches to your bucket through Google Cloud's native SDK, with Apache Arrow building the Parquet data. It is a target agent: it receives changes from any Gluesync source agent.

-   **Optimized batches, never row by row:** Core Hub groups changes into batches and every batch is uploaded as a file, so the bucket holds objects sized for BigQuery and Dataproc rather than one object per change. A large insert batch can fan out into several files for parallel reads.
-   **Parquet with your codec:** MOLO17 ParquetKt writes Snappy-compressed Parquet by default, and ZSTD, GZIP, or uncompressed output can be chosen per entity.
-   **Configurable file sizing:** the Parquet file size threshold and the polling interval are configurable per entity, and a longer interval suits readers that prefer fewer, larger objects.
-   **Flat files per entity:** `Use JSON file` or `Use CSV file` moves one entity to that format, and JSON and CSV documents are named by primary key.

### A snapshot in the bucket, then change files as they happen

-   **Snapshot:** each entity's initial load is exported into the bucket's folder hierarchy, with paths built from the transaction type, the table, the year, the month, and a timestamp.
-   **Change files:** CDC then writes files whose `_operation` column tags each row I, U, or D, and updates and deletes carry the complete row.
-   **Current state where you query it:** objects are never rewritten, so a BigQuery, Dataproc, or Spark job builds current-state tables from the files while the bucket keeps the full change history.
-   **Resume:** an interrupted snapshot restarts from its last saved state, so a large initial load never begins again from zero.

### Metadata that travels with every object

-   **Transaction lineage:** `_transaction_id`, `_timestamp`, and the source checkpoint fields are written into the files, so audits and replays start from the source transaction.
-   **Sidecar manifest:** a JSON companion next to each Parquet file lists its timestamp, row count, operation, schema version, and transaction identifiers, and it is on by default.
-   **Source column order:** the Parquet writer keeps fields in source order and handles nullable columns automatically, so external table definitions stay predictable.
-   **Masking and filtering in flight:** Custom Field Functions mask or reshape fields before they reach the bucket, and Allowed Operations picks which operations each entity writes. See [data transformation](/data-transformation/).

### What your Google Cloud admin sets up

The agent authenticates as a Google Cloud service account with a JSON key and connects on port 443. The field list and a REST example are in the [Google Cloud Storage target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/google-cloud-storage-target.html).

1.  Create the landing bucket in your Google Cloud project.
2.  Create a service account and grant it Storage Object Admin, Storage Object Creator, and Storage Object Viewer.
3.  Create a JSON key for the service account in the Google Cloud console.
4.  In Core Hub, enter the Project ID, keep port 443, and upload the JSON key, or mount it into the agent container as a volume so the key never passes through the UI.

### Core Hub around the landing zone

A pipeline groups a source agent, the Cloud Storage agent, and the entities they replicate. Core Hub orchestrates the agents through its web UI and REST APIs, and they deploy with Docker, Docker Compose, or Kubernetes, on Google Cloud, on-premises, or in another cloud. The [data lake solution](/solutions/data-lake/) covers the broader landing-zone patterns.

Because Cloud Storage, BigQuery, and Pub/Sub are all Gluesync targets, one capture agent per source can land history in the bucket, current state in the warehouse, and events on a topic, each in its own pipeline under one Core Hub.

[Explore the general CDC streaming architecture →](/solutions/cdc-streaming/)

WRITE OPTIONS

## The Google Cloud Storage target agent

Google Cloud Storage has one Gluesync target agent. It uploads grouped change files through Google Cloud's native SDK, Parquet by default, with JSON or CSV selectable per entity.

| Agent | Write technique | Versions | Best for |
| --- | --- | --- | --- |
| [Google Cloud Storage agent ↗](https://docs.molo17.com/gluesync/latest/agents/google-cloud-storage-intro.html) | Native Google Cloud SDK uploads of grouped snapshot and CDC files; Parquet with Apache Arrow; JSON or CSV per entity | All versions; any Google Cloud project, authenticated with a service account key | Google Cloud landing zones that feed BigQuery external tables, Dataproc, or Spark, with transaction lineage in every file and any Gluesync source agent feeding it under the same Core Hub. |

SOURCES AND TOPOLOGIES

## Land changes from the databases you already run in Cloud Storage

Any Gluesync source agent can write to Google Cloud Storage, each with its own native capture method. Open the [integrations finder with Google Cloud Storage pre-selected](/integrations/?target=Google+Cloud+Storage#integration-finder) to see every source you can pair with it, including Cloud SQL for PostgreSQL and Cloud SQL for MySQL.

Gluesync keeps pace with your change volume at any scale, and the polling interval and file size threshold are configurable, so each entity balances freshness against object count. [MOLO17 Professional Services](/solutions/professional-services/) can design the bucket layout with your team.

-   PostgreSQL and Cloud SQL for PostgreSQL to Google Cloud Storage from the write-ahead log: see [PostgreSQL CDC](/solutions/postgresql-cdc/)
-   Oracle to Google Cloud Storage from the redo logs through LogMiner or XStream: see [Oracle CDC](/solutions/oracle-cdc/)
-   SQL Server to Google Cloud Storage through Change Data Capture or Change Tracking: see [SQL Server CDC](/solutions/sql-server-cdc/)
-   MongoDB to Google Cloud Storage as JSON or Parquet from Change Streams: see [MongoDB CDC](/solutions/mongodb-cdc/)

FAIR, HIGH-LEVEL COMPARISON

## Where Gluesync fits among Google Cloud landing approaches

| Approach | What buyers usually get | Where Gluesync fits |
| --- | --- | --- |
| Google Datastream | Google-managed change capture that writes to Cloud Storage and BigQuery for the sources it supports | Agents you deploy next to each source, including IBM i and SAP HANA, with operation and transaction metadata in every file; see the [data lake solution](/solutions/data-lake/) |
| Fivetran and Airbyte | Managed and open-source ELT connectors that load cloud storage and warehouses, with sync frequency set per connector | Log-based agents writing to the bucket continuously, with MOLO17 enterprise support behind every pipeline |
| Debezium with a Kafka Connect GCS sink | Open-source capture into Kafka topics, then a sink connector writes objects to Cloud Storage; you run Kafka, Connect, and the file layout | Changes land in Cloud Storage with no Kafka cluster in the path, and Kafka remains available as a separate target; see the [Debezium alternative](/solutions/debezium-alternative/) |
| Qlik Replicate and Striim | Commercial replication products with Google Cloud Storage among many targets | One Core Hub that also feeds BigQuery and Pub/Sub from the same capture agents, so every Google Cloud target shares one operating model |
| Dataflow jobs, Cloud Functions, and scripts | Full control; your team owns the extract queries, object naming, retries, and the load each run puts on production | Log-based capture, snapshot resume, and Core Hub monitoring with no loader code to maintain; read [log-based CDC](/blog/real-time-replication-log-based-cdc/) |

RELATED CONTENT

## Cloud Storage ingestion: research and implementation detail

-   [Log-based CDC explained](/blog/real-time-replication-log-based-cdc/)
-   [MOLO17 ParquetKt: the Parquet library behind lake files](/blog/molo17-parquetkt-kotlin-apache-parquet-library/)
-   [Data lake solution: land operational changes in object storage](/solutions/data-lake/)
-   [Warehouse sync: keep analytical stores aligned with operations](/solutions/warehouse-sync/)
-   [Mask data in flight before it lands in the lake](/blog/data-masking-transformation-edge-gluesync-field-functions/)
-   [Google Cloud Storage agent overview ↗](https://docs.molo17.com/gluesync/latest/agents/google-cloud-storage-intro.html)
-   [Google Cloud Storage target setup guide ↗](https://docs.molo17.com/gluesync/latest/agents/google-cloud-storage-target.html)
-   [Parquet files support ↗](https://docs.molo17.com/gluesync/latest/core-hub/parquet-files-support.html)

FAQ

## Google Cloud Storage replication questions

What does the Google Cloud Storage target do?

It writes replicated data to a Cloud Storage bucket through Google Cloud's native SDK. Each entity is snapshotted first, then changes are uploaded continuously as grouped files, Parquet by default.

Which sources can replicate to Google Cloud Storage?

Any Gluesync source agent, including Cloud SQL for PostgreSQL, Cloud SQL for MySQL, Oracle, SQL Server, and MongoDB. The integrations finder on our website lists every pairing.

How does the agent authenticate?

As a Google Cloud service account. You create a JSON key for it in the Google Cloud console and give the agent that file, uploaded in the Web UI or mounted as a volume.

Which roles does the service account need?

Storage Object Admin, Storage Object Creator, and Storage Object Viewer on the project or bucket that receives the files.

Which file formats does Gluesync write to Cloud Storage?

Snappy-compressed Parquet by default, with ZSTD, GZIP, or no compression available. Any entity can switch to JSON or CSV, and each Parquet file gets a JSON manifest unless you turn it off.

How can BigQuery or Spark rebuild current tables from the files?

Every row carries an \_operation value of I, U, or D plus its transaction identifier and timestamp, and updates and deletes include the full row. A merge job applies them in transaction order to produce current state.

Does Gluesync upload one object per change?

No. Changes are grouped into batches and each batch is written as a file. Parquet output fills up to a file size threshold, and both that threshold and the polling interval are configurable per entity.

REPLICATE TO A TARGET

## Other targets Gluesync delivers to

-    [Replicate to Aerospike](/solutions/replicate-to-aerospike/)
-    [Replicate to DynamoDB](/solutions/replicate-to-dynamodb/)
-    [Replicate to Redshift](/solutions/replicate-to-redshift/)
-    [Replicate to Amazon S3 & S3-compatible](/solutions/replicate-to-amazon-s3/)
-    [Replicate to Cassandra](/solutions/replicate-to-cassandra/)
-    [Replicate to Kafka](/solutions/replicate-to-kafka/)
-    [Replicate to Cosmos DB](/solutions/replicate-to-cosmos-db/)
-    [Replicate to Azure Data Lake](/solutions/replicate-to-azure-data-lake/)
-    [Replicate to ClickHouse](/solutions/replicate-to-clickhouse/)
-    [Replicate to CockroachDB](/solutions/replicate-to-cockroachdb/)
-    [Replicate to Couchbase](/solutions/replicate-to-couchbase/)
-    [Replicate to file stores](/solutions/replicate-to-file-stores/)
-    [Replicate to BigQuery](/solutions/replicate-to-bigquery/)
-    [Replicate to Google Pub/Sub](/solutions/replicate-to-google-pubsub/)
-    [Replicate to GridGain](/solutions/replicate-to-gridgain/)
-    [Replicate to Db2](/solutions/replicate-to-db2/)
-    [Replicate to Informix](/solutions/replicate-to-informix/)
-    [Replicate to MariaDB](/solutions/replicate-to-mariadb/)
-    [Replicate to SQL Server](/solutions/replicate-to-sql-server/)
-    [Replicate to MongoDB](/solutions/replicate-to-mongodb/)
-    [Replicate to MySQL](/solutions/replicate-to-mysql/)
-    [Replicate to Oracle](/solutions/replicate-to-oracle/)
-    [Replicate to PostgreSQL](/solutions/replicate-to-postgresql/)
-    [Replicate to RavenDB](/solutions/replicate-to-ravendb/)
-    [Replicate to Redis](/solutions/replicate-to-redis/)
-    [Replicate to SAP ASE](/solutions/replicate-to-sap-ase/)
-    [Replicate to SAP HANA](/solutions/replicate-to-sap-hana/)
-    [Replicate to ScyllaDB](/solutions/replicate-to-scylladb/)
-    [Replicate to SingleStore](/solutions/replicate-to-singlestore/)
-    [Replicate to Snowflake](/solutions/replicate-to-snowflake/)
-    [Replicate to Solace PubSub+](/solutions/replicate-to-solace/)
-    [Replicate to Vertica](/solutions/replicate-to-vertica/)
-    [Replicate to YugabyteDB](/solutions/replicate-to-yugabytedb/)

## Evaluate Gluesync with your Google Cloud Storage buckets

Start a trial in your own Google Cloud project, or talk to MOLO17 about your sources, bucket layout, and the BigQuery or Dataproc jobs that read it.

[Start a Gluesync trial](/get-gluesync/) [Talk to MOLO17](/contacts/)
