SOLUTION · REPLICATE TO FILE STORES
Real-time data replication to file and object stores
Committed changes land as Parquet, JSON, or CSV files on the file servers and object stores you run, continuously, instead of waiting for the next export job.
Gluesync by MOLO17 captures changes from Oracle, SQL Server, PostgreSQL, MySQL, IBM i, MongoDB, and other heterogeneous sources with a dedicated agent per database. The universal file store agent writes them into folders over FTP, FTPS, SFTP, WebDAV, SMB, CIFS, or NFS, as Parquet by default, with JSON and CSV selectable per entity. A snapshot seeds each table as file batches, CDC follows in real time, and Core Hub, the Gluesync control plane, runs every pipeline from one web UI and REST API.
VENDOR COMPATIBILITY
Battle-tested on every File and object stores you deliver to
The same Gluesync File and object stores agent is tested against each vendor offering below, self-managed or fully managed, so delivery behaves the same wherever File and object stores runs.
- Apache Parquet Target Tested
- CSV Target Tested
- FTP/FTPS Target Tested
- JSON Target Tested
- NFS Target Tested
- SFTP Target Tested
- SMB/CIFS Target Tested
- WebDAV Target Tested
WHO THIS IS FOR
Teams that need files on their own storage to reflect operations now, not after the nightly export
- Data engineers who load an on-premises data lake or a Spark job from Parquet, and want folders that follow the source schema and table names, with a sidecar metadata file beside every batch
- Data platform leads in regulated organizations who must keep data inside their own network, with one Core Hub for every source and best-in-class enterprise support, rated 4.9/5 by customers
- Architects designing a private lake or archive who choose the transport per pipeline, FTP or FTPS, SFTP, WebDAV, SMB, or NFS, and decide which tables land as Parquet, JSON, or CSV
- Engineers running nightly export scripts or file drops who want every source writing files continuously, with snapshots and monitoring built in
THE PROBLEM
Files built from nightly exports describe yesterday's business
Most on-premises data lakes and archives are fed by scheduled exports: a script queries each source, writes CSV, and copies the result to a share. Every export adds read load to the production database, the files describe the state at the last run, and each source needs its own script, schedule, and failure handling. Teams weighing a fix usually compare CDC services that land files in object storage, Kafka Connect with a file sink, ELT tools with file destinations, or the scripts they already run.
Gluesync addresses that with per-agent CDC into files. A source agent reads each database's native change log, journal, or change stream; Core Hub routes the changes; the universal file store agent writes them as Parquet, JSON, or CSV batches to the file servers you already run. Whether the data comes from Oracle or MongoDB, the snapshot, the folder layout, and the operations stay the same.
HOW IT WORKS
How Gluesync writes to file and object stores
The write path: grouped files over the protocol your storage already speaks
The universal file store agent is a target agent. It receives changes from any Gluesync source agent and writes them as files over FTP, FTPS, SFTP, WebDAV, WebDAVS, SMB, CIFS, or NFS, with no cloud service or intermediate store in the path.
- Optimized batches, never row by row: Core Hub groups changes into batches, and each batch is written as a file, with Parquet data buffered through Apache Arrow before upload. Shares receive a steady flow of complete files rather than a trickle of single-row writes.
- Transports: FTP or FTPS, SFTP, WebDAV or WebDAVS, SMB or CIFS with the share name as the base path, and NFS through a mount point on the agent host
- Formats: Parquet by default, with JSON and CSV selectable per entity
- Parquet writer: the MOLO17 ParquetKt library builds the files, with SNAPPY, ZSTD, GZIP, or uncompressed output chosen per entity
- Concurrent uploads: configurable per agent for your bandwidth and file server
Snapshot first, then changes, in one folder tree
- Seed: each table is written as file batches during the initial snapshot, with a snapshot write method of INSERT or UPSERT.
- Changes: after the snapshot, CDC writes inserts, updates, and deletes as change rows in real time. Each row carries an
_operationcolumn set to I, U, or D with the full row image, so downstream jobs can merge the files into current-state tables. - Folders: files are grouped under the remote root path, in a folder per entity and a timestamp, with a sequence number in each file name.
- Sidecar files: a small JSON companion file beside each Parquet upload records the row count, the operation, the schema version, and the transaction identifiers. It is on by default.
Schema and types in the files
- Types: each source value maps to a Gluesync data type, covering strings, booleans, binary, 16 to 64 bit integers, single and double precision floats, big decimals, dates, times, timestamps, timezone-aware values, arrays, and maps.
- Column order and nulls: Parquet files keep the source field order and the nullable columns of each entity.
- Technical fields: Gluesync adds columns prefixed with an underscore, such as transaction identifiers and timestamps, so they never collide with business columns.
- Rows per file: the row count at which Parquet files roll over is set per entity.
- Shaping on the way in: UDFs and data filtering reshape and select records before they are written, and Allowed Operations decides which operations reach each entity's files, for example inserts only for an archive. See data transformation.
What your storage admin sets up
Create the destination before the pipeline starts, then give the agent a login that can write to it. Field-by-field settings and REST examples for every protocol are in the target setup guide ↗.
- Create the destination directory, share, or mount point. For NFS, mount the export on the agent host first, since the agent writes to the local path.
- Create a user with read and write permission on that directory. NFS needs no login, because the agent uses the local mount.
- Enter the host, port, and base path in Core Hub. For SMB or CIFS the base path is the share name, and for WebDAV it is
/. - Choose the transport and the file type, set the remote root path, and enable TLS to move FTP to FTPS or WebDAV to WebDAVS.
Architecture around Core Hub
Lightweight agents sit close to each source, and Core Hub orchestrates them through its web UI and REST APIs and routes changes to the file store agent. A pipeline groups a source agent, the file store agent, and the entities they replicate, and Core Hub and its agents deploy with Docker, Docker Compose, or Kubernetes, on-premises or in any cloud. See CDC streaming.
Because every component can run inside your own network, data never has to leave it: the source agent, Core Hub, and the file store agent sit in your data center and write to the NAS, SFTP server, or NFS export you already operate.
WRITE OPTIONS
The universal file store agent: one agent, every file protocol
One target agent covers every supported file protocol and format. Choose the transport per pipeline and the format per entity.
| Agent | Write technique | Versions | Best for |
|---|---|---|---|
| Universal file store agent ↗ | Grouped file writes over FTP, FTPS, SFTP, WebDAV, WebDAVS, SMB, CIFS, or NFS; Parquet through the MOLO17 ParquetKt writer, JSON and CSV as line-based text files | Protocols: FTP, FTPS, SFTP, WebDAV, WebDAVS, SMB, CIFS, and NFS through a locally mounted path | Private data lakes, archives, and legacy ETL that reads files from a share. Pick Parquet for analytics engines and BigQuery or Snowflake ingestion, JSON for document tools, and CSV for spreadsheets and older ETL pipelines, with any Gluesync source agent feeding it. |
SOURCES AND TOPOLOGIES
Feed your file stores from the systems of record you already run
Any Gluesync source agent can write into a file store, each with its own native capture technique. Open the integrations finder with the file store target pre-selected to see every source you can pair with it.
One Core Hub can run Oracle to SFTP, SQL Server to SMB, and MongoDB to NFS side by side, each with its own folder layout and format. Gluesync keeps pace with your change volume at any scale, and rows per file, polling interval, and concurrent uploads are all configurable; MOLO17 Professional Services can tune them with your storage team.
- Oracle to Parquet on SFTP, from the redo logs through LogMiner or XStream: see Oracle CDC
- SQL Server to CSV on SMB or CIFS, through Change Data Capture or Change Tracking: see SQL Server CDC
- IBM i (AS/400) to JSON on NFS, through the native journal APIs: see IBM i CDC
- PostgreSQL, MySQL, and MongoDB to Parquet on WebDAVS, from the WAL, the binlog, and Change Streams: see PostgreSQL CDC, MySQL CDC, and MongoDB CDC
FAIR, HIGH-LEVEL COMPARISON
Where Gluesync fits among file and lake ingestion approaches
| Approach | What buyers usually get | Where Gluesync fits |
|---|---|---|
| Cloud replication services that land files in object storage | Managed change capture into cloud object storage, usually followed by a separate load into a warehouse or lake. Packaging and targets differ by service | Writes to your own file servers and object stores over standard protocols, with agents on-premises or in any cloud; see data lake ingestion |
| Kafka Connect with a file or object sink | Change events flow through Kafka topics into a sink that writes files; you run Kafka, Connect, offsets, and the file layout rules yourself | Files written straight from the source agent through Core Hub, with no Kafka cluster in the path. Read the Debezium alternative comparison |
| ELT tools with file or cloud storage destinations | Managed or open-source connectors with scheduled syncs; sync modes and destinations vary by connector and product | A commercial product with a dedicated capture agent per database, delivering Parquet, JSON, or CSV files continuously, with MOLO17 enterprise support behind every pipeline |
| Managed file transfer and scheduled exports | Reliable file movement with scheduling, encryption, and audit trails. They move files that already exist, so an export job still has to create them on the source | Creates the files from the change stream itself, continuously, so no export job queries production on a schedule. Snapshots and changes come from the same pipeline |
| DIY export scripts | Full control; your team owns the extract queries, file naming, retries, and the load each run puts on production | Log-based capture, Parquet, JSON, and CSV output, and Core Hub monitoring without export code to maintain; read batch ETL vs real-time data replication |
RELATED CONTENT
Building on-premises data lakes with file stores
- How to build an on-premise data lake with the Universal File Store Agent
- Batch ETL vs real-time data replication: how to choose
- MOLO17 ParquetKt: a high-performance Apache Parquet library for Kotlin
- MOLO17 ParquetKt lands on GitHub
- Data lake ingestion: land operational changes in the data lake
- Universal file store agent overview ↗
- Universal file store target setup guide ↗
- Parquet files support ↗
FAQ
File and object store replication questions
What does replicating to file stores with Gluesync involve?
A source agent captures committed changes from your database through its native change mechanism, Core Hub routes them, and the universal file store agent writes them as Parquet, JSON, or CSV files after a snapshot seeds each table.
Which protocols can Gluesync write to?
FTP, FTPS, SFTP, WebDAV, WebDAVS, SMB, CIFS, and NFS through a locally mounted path. TLS can be enabled for FTP and WebDAV connections, which upgrades them to FTPS and WebDAVS.
Which file formats does Gluesync produce?
Parquet by default, plus line-based JSON and CSV, selectable per entity. Parquet files are compressed with SNAPPY by default, and ZSTD, GZIP, and uncompressed output are also available per entity.
How are inserts, updates, and deletes marked in the files?
Each change row carries an _operation column with I for insert, U for update, or D for delete, along with the full row, so a downstream job can merge the files into current-state tables.
Which sources can write to file stores?
Any Gluesync source agent, including Oracle, SQL Server, PostgreSQL, MySQL, IBM i, and MongoDB. The integrations finder on our website shows every pairing.
What does the storage side need before the pipeline starts?
The destination directory, share, or mount point must already exist, and the login needs read and write permission on it. For NFS, the export must be mounted on the agent host at the base path.
Does Gluesync write one file per change?
No. Changes are grouped into batches and written as files, with the rows per file set per entity and the number of concurrent uploads configurable per agent.
Does each file come with metadata?
Yes, by default. A sidecar JSON companion file beside each Parquet upload records the row count, operation, schema version, and transaction identifiers, and the setting can be switched off in the agent configuration.
REPLICATE TO A TARGET
Other targets Gluesync delivers to
- Replicate to Aerospike
- Replicate to DynamoDB
- Replicate to Redshift
- Replicate to Amazon S3 & S3-compatible
- Replicate to Cassandra
- Replicate to Kafka
- Replicate to Cosmos DB
- Replicate to Azure Data Lake
- Replicate to ClickHouse
- Replicate to CockroachDB
- Replicate to Couchbase
- Replicate to BigQuery
- Replicate to Google Cloud Storage
- Replicate to Google Pub/Sub
- Replicate to GridGain
- Replicate to Db2
- Replicate to Informix
- Replicate to MariaDB
- Replicate to SQL Server
- Replicate to MongoDB
- Replicate to MySQL
- Replicate to Oracle
- Replicate to PostgreSQL
- Replicate to RavenDB
- Replicate to Redis
- Replicate to SAP ASE
- Replicate to SAP HANA
- Replicate to ScyllaDB
- Replicate to SingleStore
- Replicate to Snowflake
- Replicate to Solace PubSub+
- Replicate to Vertica
- Replicate to YugabyteDB
Evaluate Gluesync with your file and object storage
Start a trial against your own SFTP server, NFS mount, or SMB share, or talk to MOLO17 about your sources, folder layout, and file formats.