SOLUTION · REPLICATE TO REDSHIFT

Real-time data replication to Amazon Redshift from your operational databases

Committed changes from your systems of record land in Redshift continuously, staged as CSV in S3 and loaded with COPY, instead of waiting for the next batch.

Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MySQL, SAP HANA, MongoDB, and other heterogeneous sources with a dedicated agent per database. The RedShift agent writes them through the RedShift JDBC driver: each batch is uploaded as CSV to an S3 staging bucket, copied into staging tables with COPY, and merged into your tables. Snapshots and CDC run from the same pipeline, and Core Hub, the Gluesync control plane, manages every pipeline from one web UI and REST API.

WHO THIS IS FOR

Teams that need Redshift to reflect operations now, not after the nightly load

  • Analytics engineers modeling in Redshift who need current-state tables that follow the source key by key, refreshed every mirroring cycle instead of every night
  • Data platform leads feeding Redshift from Oracle, SQL Server, PostgreSQL, or SAP HANA who want one Core Hub for every source instead of one connector product per source, backed by best-in-class enterprise support, rated 4.9/5 by customers
  • Architects placing Redshift loads: Serverless or provisioned clusters, a dedicated S3 staging bucket in the cluster's region, and COPY options set once on the agent
  • Engineers replacing AWS DMS tasks, Kafka sinks, or ELT scripts that load Redshift, who want every source delivered to Redshift by one product with snapshots and CDC under one control plane

THE PROBLEM

Redshift data that arrives in scheduled loads is already behind

Most Redshift estates are fed by scheduled jobs: an ELT run extracts from the source, writes files, and runs COPY on a timetable. Reports then reflect the last run, every extract adds read load to the production database, and each source needs its own connector and schedule. Teams searching for real-time Redshift ingestion or a Redshift CDC pipeline usually weigh AWS Database Migration Service, managed ELT such as Fivetran or Airbyte, AWS zero-ETL integrations for Aurora and RDS, a Debezium and Kafka stack with a Redshift sink, or scripts they keep running themselves.

Gluesync addresses that with per-agent CDC into Redshift. A source agent reads each database's native change log, Core Hub routes the changes, and the RedShift agent stages them as CSV in S3, loads them with COPY, and merges them into your tables. Oracle to Redshift, SQL Server to Redshift, and PostgreSQL to Redshift follow the same model, with the same snapshot and operations.

HOW IT WORKS

How Gluesync writes to Redshift

The write path: JDBC connections, S3 staging, and COPY

The RedShift agent connects with the RedShift JDBC driver, bundled with the agent. It writes to Amazon Redshift Serverless and provisioned clusters, and it uses AWS S3 as the staging area for every bulk load. It is a target agent: it applies the change stream from any Gluesync source agent to your schema in real time, after a snapshot seeds each table.

  • Optimized batches, never row by row: every write is grouped. Changes are gathered into highly optimized batches, with a batch size configurable on the agent, so the cluster receives a few large statements instead of single-row DML.
  • Native bulk load: snapshots and changes are uploaded to S3 as CSV and loaded with the native COPY command, which runs in parallel across the cluster. RedShift agent overview ↗

Native bulk load for snapshots and CDC

Bulk load is switched on per entity with two independent settings: Bulk for Snapshot for the initial load and Bulk for CDC for the ongoing change stream. Both can be changed on an existing entity without recreating it.

In each mirroring cycle Core Hub collects the change events and collapses those on the same primary key: an insert followed by a delete is skipped, and consecutive updates become one update. The agent writes the batch as CSV under the gluesync-staging/ prefix of your S3 bucket, runs COPY into a staging table in parallel, optionally backfills columns the source log omitted, then applies deletes and inserts to the target table in one pass and removes the staged files. Each changed key lands once per cycle as one current row. COPY options are configurable on the agent.

Snapshot first, then continuous CDC

  • Seed, then stream: each table loads in full first, then follows changes from its source agent as they arrive, in commit order per entity.
  • Truncate before snapshot: runs as Redshift SQL before a reload. It is on by default and can be disabled in the target settings.
  • Scheduled reloads: the Chronos scheduler runs a snapshot, or a snapshot followed by CDC, on a cadence for tables you refresh on a timetable. See schedules and events.

Schemas, grants, and setup in your AWS account

The agent needs a database user with rights on the target schema, a staging bucket in the cluster's region, and an AWS access key with S3 access to that bucket. Full statements are in the Redshift target setup guide ↗.

  1. Create a database user with USAGE and CREATE on the target schema, and SELECT, INSERT, UPDATE, and DELETE on its tables.
  2. Create an S3 bucket for staging in the same AWS region as the cluster, and consider a VPC endpoint for S3 to shorten load times.
  3. Create an IAM user or role with s3:PutObject, s3:GetObject, and s3:DeleteObject on the staging bucket, and generate its access key ID and secret.
  4. In Core Hub, enter the cluster endpoint, port 5439, the database name, and the database user. TLS is enabled by default.
  5. Enter the S3 bucket, its region, the AWS access key ID and secret, and the schema. For Redshift Serverless, set the workgroup name; provisioned clusters need no workgroup.

Architecture around Core Hub

Source agents sit close to each database, and Core Hub orchestrates them through its web UI and REST APIs. A pipeline groups a source agent with the RedShift agent and the tables it replicates. Query Studio, the SQL workbench in Core Hub, connects to Redshift with autocomplete and text explain, so you can check landed rows in the same console. See Query Studio, and CDC streaming for the wider pattern.

Explore the general CDC streaming architecture →

WRITE OPTIONS

The RedShift target agent: S3 staging for every bulk load

Redshift has one Gluesync target agent. It writes in optimized batches, and native bulk load through S3 and COPY covers both the initial snapshot and ongoing CDC.

AgentWrite techniqueVersionsBest for
RedShift agent ↗ RedShift JDBC driver in optimized batches; native bulk load with CSV staged in S3, loaded with COPY into staging tables, then applied to the targets Amazon Redshift Serverless and provisioned clusters, all versions Analytics teams standardizing on Redshift. Native bulk load through S3 and COPY for snapshots and CDC, switched on per entity, with the same pipeline model as every other source under Core Hub.

SOURCES AND TOPOLOGIES

Feed Redshift from the systems of record you already run

Any Gluesync source agent can feed Redshift, each with its own native capture technique. Open the integrations finder with Amazon Redshift pre-selected to see every source you can pair with it.

One Core Hub runs Oracle to Redshift, SQL Server to Redshift, and SAP HANA to Redshift side by side, with the same snapshot and write settings for each. Gluesync keeps pace with your change volume at any scale, and MOLO17 Professional Services can help plan staging and load cadence with your team.

  • Oracle to Redshift through its native capture: see Oracle CDC
  • SQL Server to Redshift through Change Data Capture or Change Tracking: see SQL Server CDC
  • PostgreSQL and MySQL to Redshift from the WAL and the binlog: see PostgreSQL CDC and MySQL CDC
  • SAP HANA to Redshift from its change capture: see SAP HANA CDC

FAIR, HIGH-LEVEL COMPARISON

Where Gluesync fits among Redshift ingestion approaches

ApproachWhat buyers usually getWhere Gluesync fits
AWS Database Migration Service A managed migration and replication service with Redshift among its targets, run as tasks in your AWS account Agents you deploy near each source, native capture per engine, one Core Hub for every pipeline, and S3 staging with COPY into Redshift; see migrating to Gluesync
Fivetran and Fivetran HVR Managed ELT with a broad connector catalog and scheduled syncs into Redshift; HVR adds log-based database replication. Packaging differs by product Native capture per engine and staged COPY loads into Redshift under one Core Hub, with MOLO17 enterprise support; see warehouse sync
Airbyte Open-source and cloud ELT connectors; incremental and CDC modes vary by source connector, and you run the platform or use the managed service A commercial product with a dedicated capture agent per database and MOLO17 enterprise support behind every pipeline
AWS zero-ETL integrations Managed replication from supported Aurora and RDS sources into Redshift, set up inside AWS for those sources One pipeline model for sources on premises, in other clouds, and in AWS, with the same agent and Core Hub for each source
Debezium, Kafka Connect, and a Redshift sink connector Open-source capture into Kafka topics, loaded by a sink connector; you run Kafka, Connect, offsets, and schemas, and usually add a step that merges change events into current-state tables Changes merged into current-state Redshift tables with no Kafka cluster in the path; read the Debezium alternative comparison
DIY scheduled ELT and COPY scripts Full control; your team owns extract queries, S3 staging, merge logic, retries, and the load each run puts on production Log-based capture and staged COPY loads without extract queries or merge scripts to maintain; read batch ETL vs real-time replication

FAQ

Redshift replication questions

What does replicating to Amazon Redshift with Gluesync involve?

A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the RedShift agent applies them to Redshift tables. Each table is seeded with a snapshot first and then follows the change stream.

Does Gluesync bulk load into Amazon Redshift?

Yes, for the initial snapshot and for ongoing CDC, with a separate switch for each on every entity. Batches are collapsed per primary key, uploaded to an S3 staging bucket as CSV, loaded into staging tables with COPY in parallel, and applied to the target tables in one pass.

Which sources can replicate to Redshift?

Any Gluesync source agent. Oracle, SQL Server, PostgreSQL, MySQL, SAP HANA, and MongoDB can all feed Redshift, and each one uses its own native capture technique.

Does Gluesync merge changes into Redshift tables?

Yes. Each mirroring cycle copies its changes into staging tables and merges them into the target tables with optimized SQL, so each changed row is current in Redshift after the cycle.

Which Redshift deployments are supported?

Amazon Redshift Serverless and provisioned clusters, all versions. For Serverless, the workgroup name is set on the agent; provisioned clusters need no workgroup.

What does the database user need on Redshift?

USAGE and CREATE on the target schema, plus SELECT, INSERT, UPDATE, and DELETE on its tables. The user connects with a username and password over TLS, which is enabled by default.

Can I control how COPY loads the staged files?

Yes. The copy options setting on the agent is passed to the COPY command, so you can tune the format of the staged CSV files.

Evaluate Gluesync with your Redshift cluster

Start a trial against your provisioned cluster or Serverless workgroup, or talk to MOLO17 about your sources, staging bucket, and load cadence.