SOLUTION · DYNAMODB CDC
Amazon DynamoDB change data capture and real-time data replication
Real-time DynamoDB CDC from DynamoDB Streams, the native change stream, with no scans to find the changes.
Gluesync by MOLO17 captures item changes from Amazon DynamoDB tables through a dedicated source agent that reads DynamoDB Streams, the native change stream API for DynamoDB. The agent delivers those changes continuously to the databases, warehouses, lakes, and event streams your teams already use, and it can also write to DynamoDB as a target. Manage every pipeline from the Core Hub web UI, the Gluesync control plane.
WHO THIS IS FOR
DynamoDB teams moving operational data into analytics, databases, and streams
- Cloud and platform engineers who own DynamoDB tables in AWS and must approve an IAM access key, enable DynamoDB Streams on each table, and accept the AWS charges that Streams can add
- Data platform leads feeding Google BigQuery, Google Cloud Storage, Amazon S3 data lakes, or event streams from DynamoDB, who want one control plane for every source and target
- Teams migrating off DynamoDB or running it alongside Aerospike, Couchbase, or MongoDB, who need a snapshot for the initial load and CDC until cutover
- Engineers who built a DynamoDB Streams consumer in Lambda, or a Kinesis export, and now maintain the consumer, its retries, and recovery after a gap themselves
THE PROBLEM
Batch copies of DynamoDB tables go stale between runs
DynamoDB is usually the operational store behind an application, so a scheduled scan or export hands analytics and downstream systems a copy that is already behind. A full-table scan reads every item on each run, and it does not record which items changed in between. Buyers searching for DynamoDB CDC, DynamoDB replication, or DynamoDB to a warehouse usually compare four options: a Lambda consumer they write and run, Kinesis Data Streams for DynamoDB, AWS-native integrations such as AWS Glue or zero-ETL, and managed connectors whose capture method they need to confirm before they compare.
Gluesync addresses that with per-agent CDC. The DynamoDB source agent reads DynamoDB Streams for each table you enable, and target agents write to the destination you choose. Both run under Core Hub, so one pipeline model covers DynamoDB as a source, as a target, or both.
HOW IT WORKS
How Gluesync does DynamoDB CDC
Read DynamoDB Streams, the native change API
CDC is implemented on the native DynamoDB Streams APIs. The agent connects through the AWS SDK for Java and Kotlin, which Gluesync embeds, and it forwards changes from the stream in near real time.
- Source role: consumes DynamoDB Streams to capture changes from tables in any AWS region.
- Initial load: Core Hub snapshot tasks scan DynamoDB tables, and the stream carries ongoing changes.
- State: pipeline state is tracked in Core Hub metadata.
- DynamoDB CDC streams setup docs ↗
What your team enables in AWS before the first pipeline
- Create an IAM access key ID and secret access key for the runtime user. The AWS managed policy
AmazonDynamoDBFullAccesscovers both the source and target roles. - Enable DynamoDB Streams on each table you replicate. Streams are off by default. In the AWS console, open the table's Exports and streams tab, enable DynamoDB stream details, and choose New and old images. Gluesync does not turn streams on for you.
- Enter the region, access key ID, and secret access key in Core Hub, or set them through the credentials REST endpoint.
Regions, connectivity, and TLS
- Regions: the agent works in any AWS region where DynamoDB is deployed.
- TLS: encryption is always on for the source, and TLS is enforced by default for the target.
- Target writes: the target role writes through the AWS SDK.
Architecture around Core Hub
The DynamoDB agent runs next to your AWS estate, and Core Hub coordinates it with every target agent through its web UI and REST APIs. A pipeline groups the source agent, its target agents, and the entities (tables) they replicate. Core Hub and agents deploy with Docker, Docker Compose, or Kubernetes.
CAPTURE OPTIONS
One DynamoDB source agent, built on DynamoDB Streams
DynamoDB has one Gluesync source agent. It captures changes from DynamoDB Streams, so there is no log or journal choice to make; the decisions are IAM and per-table stream enablement. The agent also covers the target role.
| Agent | Capture technique | Versions | Best for |
|---|---|---|---|
| Amazon DynamoDB agent ↗ | DynamoDB Streams with new and old images, enabled per table | All AWS regions where DynamoDB is deployed | Teams that run DynamoDB in AWS and need its changes delivered to databases, warehouses, lakes, or streams, with source and target roles under one Core Hub. |
TARGETS AND TOPOLOGIES
Keep DynamoDB as hot storage and feed the rest of your estate
Use the integrations directory to pair Amazon DynamoDB as source or target with databases, cloud warehouses, object and lake storage, or event streams, subject to the documented role of each agent. The DynamoDB agent also supports the target role. It writes through the AWS SDK, and a write to an existing key overwrites the item. Bulk load and TRUNCATE do not apply to DynamoDB targets.
Gluesync keeps pace with your change volume at any scale. MOLO17 Professional Services can size the deployment with your team.
Target agents write in optimized batches, never row by row, and switch to native bulk load for both snapshots and CDC on targets such as Snowflake, Google BigQuery, Amazon Redshift, Microsoft SQL Server, and PostgreSQL.
- Keep DynamoDB as hot storage and send its changes to Amazon S3 or Google Cloud Storage for a lake: see data lake
- Feed a cloud warehouse such as Google BigQuery continuously, without a custom pipeline per target: see warehouse sync
- Migrate off DynamoDB to Aerospike, Couchbase, or MongoDB with a snapshot and CDC until cutover: see cloud migration
- Publish DynamoDB changes to an event-streaming platform for downstream services, as in our Solace PubSub+ CDC post
FAIR, HIGH-LEVEL COMPARISON
Where Gluesync fits among DynamoDB CDC approaches
| Approach | What buyers usually get | Where Gluesync fits |
|---|---|---|
| Lambda consumers on DynamoDB Streams | AWS-native and close to the data. You write, deploy, and operate the consumer, its retries, and its target logic. | A productized source agent with snapshots for the initial load, Core Hub monitoring, and targets beyond AWS, with no consumer code to maintain. See CDC streaming. |
| Kinesis Data Streams for DynamoDB | AWS-managed delivery of item-level changes into a Kinesis stream, which you then consume and transform yourself. | One control plane for capture, snapshots, and delivery to the target, instead of a Kinesis stream plus your own consumers. |
| AWS Glue and zero-ETL integrations | Managed inside AWS, with the destinations AWS supports. A strong fit when the destination is one AWS service. | Agent-based delivery to the targets in the integrations directory, on-premises or in any cloud, managed from one Core Hub. |
| Managed SaaS connectors | Quick to set up for common warehouses. Pricing, schedules, and capture behavior differ by vendor, so confirm how DynamoDB changes are captured before comparing. | Streams-based capture from agents you deploy on your own infrastructure, with Core Hub operations and best-in-class MOLO17 enterprise support (rated 4.9/5 by customers). |
| Scheduled scans or DynamoDB exports | Simple for periodic refreshes. Each run reads the whole table or export, and changes made between runs wait for the next run. | Stream-based CDC for changes between runs, with snapshots for the initial load. See CDC streaming. |
RELATED CONTENT
DynamoDB CDC research and implementation detail
FAQ
DynamoDB CDC questions
What is DynamoDB CDC with Gluesync?
Change data capture from Amazon DynamoDB through the DynamoDB Streams API. A Gluesync source agent reads the stream for each table you enable and delivers the changes to the configured targets through Core Hub.
Does Gluesync turn on DynamoDB Streams for us?
No. Streams are off by default, and you turn them on per table in the AWS console, choosing New and old images. Gluesync does not enable streams through the SDK.
How long does DynamoDB keep stream records?
DynamoDB keeps stream records for 24 hours, the maximum available value, so the agent always reads from a window that covers a full day of changes. Core Hub snapshot tasks give you a fresh baseline for a table whenever you want one.
Which regions does the DynamoDB agent support?
The agent works in any AWS region where DynamoDB is deployed, and it reads tables in any region as a source, with the same agent covering the target role.
What IAM permissions does the agent need?
The AWS managed policy AmazonDynamoDBFullAccess for both roles, along with an IAM access key ID and secret access key.
Can Gluesync write to DynamoDB, not only read from it?
Yes. The target role writes through the AWS SDK with TLS enforced by default. A write to an existing key overwrites the item, and bulk load and TRUNCATE do not apply to DynamoDB targets.
Is the connection to DynamoDB encrypted?
Yes. TLS is always on for the source agent, and it is enforced by default for the target agent.
How do we run the initial load?
Core Hub snapshot tasks scan DynamoDB tables for the initial load, and the Streams capture carries ongoing changes.
CDC BY SOURCE DATABASE
Other sources Gluesync captures from
Evaluate Gluesync with your DynamoDB streams
Start a trial on your infrastructure, or talk to MOLO17 about stream setup, IAM, snapshots, and the targets you need.