SOLUTION · REPLICATE TO AMAZON S3 & S3-COMPATIBLE

Real-time data replication to Amazon S3 and S3-compatible storage

Committed changes from your databases land as Parquet files in Amazon S3, MinIO, Dell ECS, or any S3-compatible store as they happen, instead of arriving in a nightly export.

Gluesync by MOLO17 captures changes from Oracle, PostgreSQL, MySQL, MongoDB, and other heterogeneous sources with a dedicated agent per database. The S3 target agent writes each entity to your bucket, on Amazon S3 or on any S3-compatible server such as MinIO or Dell ECS, through the AWS SDK, after an initial snapshot and then continuously, as Parquet by default with JSON or CSV selectable per entity. Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.

VENDOR COMPATIBILITY

Battle-tested on Amazon S3 and every S3-compatible server

One Gluesync S3 agent is tested against Amazon S3, MinIO, and Dell ECS, so your files, formats, snapshots, and CDC behave the same on AWS, on premises, or in any S3-compatible cloud.

  • AWS S3 & S3-like Target Tested
  • MinIO Target Tested
  • Dell ECS Target Tested

WHO THIS IS FOR

Teams that want S3 to hold current operational data, not last night's export

  • Analytics engineers who build lake tables on S3 and need every change tagged with its insert, update, or delete operation and its transaction identifiers, so current-state tables can be rebuilt from the history
  • Data platform leads who feed one bucket from several databases and want one Core Hub for every source, backed by best-in-class enterprise support, rated 4.9/5 by customers
  • Architects designing a data lake or cold-storage tier who need Parquet by default, JSON or CSV per entity where downstream tools read flat files, and AWS or S3-compatible storage such as MinIO or Dell ECS
  • Engineers running Kafka Connect S3 sinks or export scripts who want every source landing in S3 with snapshots and monitoring built in, and no export schedule to maintain

THE PROBLEM

Object storage loaded in batches goes stale between exports

Most S3 lakes are fed by scheduled exports: a nightly job dumps tables, writes delimited files, and a loader picks them up. The lake reflects the business as it was at the last dump, every export adds read load to the production database, and each source carries its own script, schedule, and failure mode. Teams searching for real-time S3 ingestion or a CDC pipeline to S3 usually weigh managed ELT services such as Fivetran or Airbyte, a Debezium and Kafka Connect sink, AWS Database Migration Service, or scripts they keep running themselves.

Gluesync addresses that with per-agent CDC into Amazon S3. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the S3 target agent writes them as files in your bucket. The pipeline model, the snapshot, and the operations stay the same whether the source is Oracle, MongoDB, or DynamoDB.

HOW IT WORKS

How Gluesync writes to Amazon S3

The write path: the AWS SDK, writing grouped change files

The S3 target agent connects with the AWS SDK for Java and Kotlin and writes changes as objects in your bucket. It works with buckets in any AWS region and with S3-compatible storage such as MinIO and Dell ECS.

  • Optimized batches, never row by row: Core Hub groups changes into change batches, and each batch is written as a file, so the bucket receives a steady flow of well-sized objects instead of one object per change. Large insert batches can split into several files for parallel processing downstream.
  • Parquet by default: the MOLO17 ParquetKt writer produces the files, with Snappy compression. ZSTD, GZIP, and uncompressed output are selectable per entity.
  • File size and polling interval: both are configurable per entity. A longer interval puts more changes into each file and produces fewer objects.
  • Per-entity formats: turning on Use JSON file or Use CSV file for an entity writes that entity as JSON or CSV instead.

Any S3-compatible storage, one agent

The same S3 target agent works with every S3-compatible object store, on premises or in any cloud. Gluesync is tested against MinIO and Dell ECS, so the files, formats, snapshot and CDC behavior are identical wherever your bucket lives, and moving from AWS to your own storage, or the other way round, is a connection change rather than a new pipeline.

  • Point it at your endpoint: enter the store's https:// or http:// address, leave the AWS Region empty, and use the storage user's credentials.
  • Same files everywhere: Parquet by default with JSON or CSV per entity, the same operation tags, the same file-size and polling controls.
  • Same Core Hub: one control plane runs AWS buckets and S3-compatible servers side by side, with every Gluesync source feeding either.

Snapshot first, then change files

  • Snapshot: the initial load is exported into the snapshots folder of each entity, following the keyspace convention of its schema and table.
  • Change files: after the snapshot, CDC runs continuously and each change batch is written as a file with an _operation column set to I, U, or D. Updates and deletes include the full updated row.
  • Merge downstream: files are never edited in place, so a Spark, Athena, or SQL job merges the change files into current-state tables, with the full history kept in the bucket.
  • Resume: an interrupted snapshot resumes from its last saved state when you start the entity again.

Folders, files, and metadata

Each entity lands in its own folder path, built from its schema and table name. JSON and CSV documents are named by primary key, so each row stays identifiable in the bucket.

  • Lineage columns: each file carries _transaction_id and _timestamp plus the source checkpoint fields, for auditing and replay.
  • Sidecar manifest: each Parquet file gets a JSON companion with its timestamp, row count, operation, schema version, and transaction identifiers. It is on by default and can be switched off.
  • Column order and nulls: the Parquet writer keeps the source field order and handles nullable columns automatically.
  • Shaping before landing: Custom Field Functions mask or reshape fields in flight, and Allowed Operations chooses which operations reach the bucket, for example inserts only for an append-only archive. See data transformation.

Bucket and credentials

Create the bucket in the AWS region that holds your lake, then give Core Hub an IAM access key with read and write access to it. The full Web UI field list and a REST example are in the AWS S3 target setup guide ↗.

  1. Create an IAM user with an access key and grant it read and write access to the bucket.
  2. In Core Hub, enter the bucket's DNS name as the host on port 443, the bucket name, the access key ID as the username, the secret access key as the password, and the AWS Region.
  3. For S3-compatible storage, enter the endpoint with https:// or http://, leave the AWS Region empty, and enter the storage user's credentials.

Core Hub around the bucket

A pipeline groups a source agent, the S3 target agent, and the entities they replicate. Core Hub orchestrates the agents through its web UI and REST APIs, and agents deploy with Docker, Docker Compose, or Kubernetes, on premises or in any cloud. The data lake solution covers the lake patterns.

Because S3 is one target among many on the same Core Hub, a source table can land in the bucket as history and in a warehouse or operational database as current state, from the same capture agent and with the same monitoring.

Explore the general CDC streaming architecture →

WRITE OPTIONS

The Amazon S3 target agent: one agent for AWS and S3-compatible storage

Amazon S3 has one Gluesync target agent. It writes Parquet by default, with JSON or CSV selectable per entity, to AWS buckets or to S3-compatible endpoints.

AgentWrite techniqueVersionsBest for
AWS S3 agent ↗ AWS SDK uploads of grouped change files; Parquet with Apache Arrow batching; JSON or CSV per entity Amazon S3 in any region; S3-compatible storage, tested on MinIO and Dell ECS Data lakes, archives, and cold storage on Amazon S3 or S3-compatible storage. Choose Parquet for Spark, Athena, and other analytics engines, and JSON or CSV for tools that read flat files, with any Gluesync source agent feeding it under the same Core Hub.

SOURCES AND TOPOLOGIES

Land changes from the databases you already run in S3

Any Gluesync source agent can write to Amazon S3, each with its own native capture method. Open the integrations finder with Amazon S3 pre-selected to see every source you can pair with it.

Gluesync keeps pace with your change volume at any scale, and the polling interval and Parquet file size threshold are configurable per entity and agent, so you choose between fresher files and fewer, larger ones. MOLO17 Professional Services can tune both with your team.

  • Oracle to Amazon S3 as Parquet files, from the redo logs through LogMiner or XStream: see Oracle CDC
  • PostgreSQL to Amazon S3 from the write-ahead log: see PostgreSQL CDC
  • MySQL to Amazon S3 from the binary log: see MySQL CDC
  • MongoDB and DynamoDB to Amazon S3 from change streams and DynamoDB Streams: see MongoDB CDC and DynamoDB CDC

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among S3 ingestion approaches

ApproachWhat buyers usually getWhere Gluesync fits
Fivetran Managed connectors with scheduled syncs into lake and warehouse destinations, and a large connector catalog Log-based agents next to each database, continuous writes to S3, and one Core Hub for every source; see the data lake solution
Airbyte Open-source and cloud connectors with S3 destinations; you run the platform or use its managed service A commercial product with a dedicated capture agent per database, Parquet output by default, one pipeline model across sources, and MOLO17 enterprise support
Debezium with a Kafka Connect S3 sink Open-source capture into Kafka topics, then a sink connector writes objects to S3; you run Kafka, Connect, offsets, and the file layout Changes written to S3 with no Kafka cluster in the path, and Kafka remains available as a separate target; see the Debezium alternative
AWS Database Migration Service AWS-managed replication with S3 among its targets, suited to AWS-centric migrations Agents you deploy on premises or in any cloud, writing to any S3 or S3-compatible endpoint under one Core Hub; see migrating to Gluesync
Qlik Replicate and Striim Commercial replication products with S3 among many targets A commercial product with a dedicated capture agent per database, Parquet output, and one Core Hub for the whole estate
DIY export scripts Full control; your team owns the export queries, file naming, retries, and the load each run puts on production Log-based capture, snapshot resume, and Core Hub monitoring with no pipeline code to maintain; read log-based CDC

FAQ

Amazon S3 replication questions

What does the Amazon S3 target do?

It writes replicated data to Amazon S3 buckets and S3-compatible storage such as MinIO and Dell ECS. After an initial snapshot of each entity, changes from the source agent are written continuously as files in your bucket.

Does Gluesync work with MinIO, Dell ECS, and other S3-compatible storage?

Yes. The S3 target agent works with any S3-compatible object store and is tested on MinIO and Dell ECS. Enter the store's endpoint with https:// or http://, leave the AWS Region empty, and use the storage user's credentials. Parquet, JSON, and CSV output, the initial snapshot, and continuous CDC behave exactly as they do on Amazon S3.

Which file formats does Gluesync write to S3?

Parquet by default, compressed with Snappy. Each entity can also be set to JSON or CSV, and each Parquet file gets a sidecar JSON manifest unless that option is turned off.

Which sources can replicate to Amazon S3?

Any Gluesync source agent, including Oracle, SQL Server, PostgreSQL, MySQL, MongoDB, and DynamoDB. The integrations finder on our website lists every pairing with Amazon S3.

How are inserts, updates, and deletes stored in S3?

Each change is a row in a file with an _operation column: I for insert, U for update, and D for delete. Updates and deletes include the full row, and a downstream job merges the files into current-state tables.

Does Gluesync create tables or schemas in S3?

No tables are needed. Each entity is written to a folder path built from its schema and table name, and the Parquet writer keeps the source column order and nullable columns.

What does the AWS account need?

A bucket, and an IAM access key and secret with read and write access to it, plus the bucket's AWS Region in the agent. For S3-compatible storage, the endpoint and the storage user's credentials are used instead.

How do I control file size and file count?

Parquet file size and polling interval are configurable per entity, so you choose fresher files or fewer, larger ones. The polling interval sets how many changes go into each file.

Does Gluesync write one object per change?

No. Changes are grouped into change batches and each batch is written as a file, with Parquet output buffered up to the file size threshold, so S3 receives well-sized objects rather than one object per row.

Evaluate Gluesync with your Amazon S3 buckets

Start a trial on your own infrastructure, or talk to MOLO17 about your sources, file formats, and expected change volume.