Gluesync Core Hub · Built into every pipeline

Data transformation

Reshape data while it moves. Pick a pre-built field function per column, write a Java UDF when a record needs real logic, publish reusable Custom Field Functions at pipeline level, keep only the rows that match your filters, and decide how documents and messages are keyed on NoSQL, event-streaming and object-store targets. Spark helps you write and review all of it.

Data transformation
  1. 01 Map and convert fields without code
  2. 02 Add Java logic when the catalog is not enough
  3. 03 Decide which rows, and which keys, reach the target

Why transform in transit

The target rarely wants the source as it is

01 / Shape mismatch

Different engines, different shapes

Oracle dates, SQL Server decimals and MySQL booleans do not land cleanly in Couchbase documents, Kafka messages or BigQuery columns. Fixing them after the fact means another job, another schedule and another copy.

02 / Compliance at the edge

Sensitive data should not travel

If a target team only needs the last four digits of a card, the full number should never reach them. Masking, filtering and anonymization belong in the pipeline, not in a downstream view somebody may forget to apply.

03 / The platform outcome

One pipeline, snapshot and CDC alike

Gluesync applies the same field functions, UDFs, filters and document keys to the initial snapshot and to every change captured afterwards, from the Control Plane, with no extra component to run.

How it works

From a mapped column to a shaped record

  1. 01

    Map

    Pick the column, pick the function

    In the Fields Editor, map source to target columns and attach a pre-built field function where the shape has to change: casts, date patterns, trims, constants, technical fields and masking. Parameters come with ready-to-use defaults.

  2. 02

    Extend

    Write code only where it earns its keep

    When a record needs business logic, add a Java UDF to the entity. When several pipelines need the same custom conversion, publish it once as a Custom Field Function and bind it to target columns like any catalog entry.

  3. 03

    Scope

    Filter rows and shape keys

    Add inclusive filter clauses that decide which rows are copied on snapshot and kept in sync on CDC, and compose the document or message key your NoSQL, Kafka or object-store target expects.

Capabilities

What ships with Core Hub

Every capability on this page is configured in the Control Plane and stored with the pipeline, so Bootstrapper exports and REST calls carry the same configuration. The simulations below mirror the Core Hub screens where each one lives.

Field functions

A catalog of no-code transformations

Eight categories of pre-built functions, applied per column in the Fields Editor

Field functions are applied in transit to a mapped column. Regular transformations change business data on its way to the target; technical field functions populate target-only columns with timestamps or constants, or override the incoming value, which is how data masking and anonymization work without code.

  • Date and time to string, string to date and time, and advanced date and time operations such as zone-aware timestamp conversions
  • Numeric and decimal conversions, boolean conversions, and text and binary operations
  • Technical fields and constants for execution timestamps, record IDs built from the primary key, or fixed values
  • Data masking and anonymization: Mask Credit Card, Mask String, Mask Email, Mask Phone and Mask IBAN
  • Each function has a YAML expression type, so the same configuration can be exported and applied with Bootstrapper

View field functions docs ↗

Data transformation
Data pipelines / P-MSSQL-SNOW Customer 360 / dbo.orders / Fields
Fieldsdbo.orders → ORDERS
TargetField function
ORDER_TSorder_ts · DATETIME2(6) → VARCHAR(32)Convert Date and Time to String
IBANiban · NVARCHAR(34) → VARCHAR(34)Direct mapping
IS_ACTIVEis_active · BIT → CHAR(1)Direct mapping
SYNCED_ATTarget only → TIMESTAMP_TZ—
Select Expression73 functions · 8 categories

Masking pre-built, not hand-rolled

Pick the mask function, keep the defaults or set the mask character, and the target only ever receives the redacted value. Discovery of which columns need it is covered on the PII discovery and masking page.

User Defined Functions

Java logic on every change event

An onChange handler per entity, with new and old values, the operation and a logger

A UDF receives each insert, update and delete after it is read from the source and before it is written to the target. It returns the record to write, or null to skip it, and can raise an error that goes through the pipeline's error handling. Code is checked and compiled in Core Hub before it can run.

  • Signature onChange(newValues, oldValues, operation, logger): mutate, enrich, route or drop a record
  • Old values are available for updates and deletes on sources that capture before images
  • Java today; Kotlin, JavaScript and Python are listed as available soon in the docs
  • Checked and compiled before execution, run sandboxed with execution timeouts, with a logger that writes to Core Hub logs
  • Spark helper in the editor: ask it to explain or improve the UDF, insert the result and compile

View UDF docs ↗

Data transformation
Data pipelines / P-MSSQL-SNOW Customer 360 / dbo.orders / UDF
UDF_ordersJavaInsert · Update · Delete
  1. public class UDF_orders {
  2. public Pair<MappingFunctionOperation, Map<String, Object>> onChange(
  3. Map<String, Object> newValues, Map<String, Object> oldValues,
  4. MappingFunctionOperation operation, Logger logger) {
  5. Map<String, Object> row = operation == MappingFunctionOperation.Delete ? oldValues : newValues;
  6. if ("TEST".equals(row.get("CHANNEL"))) return null; // skip the operation
  7. if (operation == MappingFunctionOperation.Delete) {
  8. Map<String, Object> soft = new HashMap<>(oldValues);
  9. soft.put("IS_DELETED", true);
  10. soft.put("DELETED_AT", OffsetDateTime.now(ZoneOffset.UTC).toString());
  11. return new Pair<>(MappingFunctionOperation.Update, soft);
  12. }
  13. return new Pair<>(operation, newValues);
  14. }
  15. }
Compiled as UDF_orders

Custom Field Functions

Reusable functions at pipeline level

Java methods exported with @GsFunction, published, versioned and bound to columns

When the catalog does not have the conversion you need, or a target column has to be built from several source columns at once, write it once as a Custom Field Function. It belongs to the pipeline, shows up in the Fields Editor next to the pre-built catalog, and can be bound to as many target columns and entities as you like.

  • Exported methods are public static and annotated with @GsFunction, with typed parameters mapped from source columns
  • Compile and publish checks the code, recompiles every UDF in the pipeline and re-checks every bound column before storing a version
  • Install is a separate step: entities are stopped, the executable replaced and restarted, and Gluesync reports which ones
  • Versions can be restored; Try a function runs the installed version on sample values before you bind it
  • Repeatedly failing functions are suspended instead of failing on every record

View Custom Field Functions docs ↗

Data transformation
Data pipelines / P-ORA-PG Plant data / Custom Field Functions

Custom Field Functions

  1. package com.molo17.gluesync.collections;
  2. import java.time.*;
  3. import com.molo17.gluesync.commons.function.GsFunction;
  4. public class Format {
  5. @GsFunction(name = "compose_timestamp",
  6. description = "A date column and a time column into one timestamp")
  7. public static OffsetDateTime composeTimestamp(LocalDate date, LocalTime time, String zone) {
  8. if (date == null) return null;
  9. LocalTime t = time == null ? LocalTime.MIDNIGHT : time;
  10. return date.atTime(t).atZone(ZoneId.of(zone)).toOffsetDateTime();
  11. }
  12. }
v3 · installed

Ask Spark to explain or improve these custom functions… · grounded on this pipeline's sources, columns and compiled signatures

Filters

Inclusive filters on snapshots and on CDC

Per-entity clauses evaluated in Core Hub, on the initial copy and on every change

Filters decide which rows belong on the target. A row is replicated only if it matches all clauses. During the snapshot only matching rows are copied; during CDC a row that stops matching is converted into a DELETE on the target, so the target always reflects the filter. Filters are evaluated in Core Hub after the read, so they add no load on the source.

  • Operators: =, !=, >, <, >=, <=, In, Not in, Is null, Not null and Regex
  • Combine clauses on several columns; all of them must match for the row to be kept
  • Enabled per entity from Filter out incoming data in the entity settings dialog
  • Selective deletion ahead of snapshot reuses the same operators to prune target rows right before a snapshot task runs

View data filtering docs ↗

Data transformation
Data pipelines / P-MSSQL-SNOW Customer 360 / dbo.orders / Settings

Filter out incoming data

2 clauses

Applied to CDC changes and to the initial snapshot. Clauses behave like a SQL WHERE and a row is kept only when all of them match.

Selective deletion ahead of snapshot

Right before a snapshot, Gluesync runs DELETE FROM ORDERS WHERE REGION != 'EU' on the target. Snapshot only; CDC is untouched.

EU

Clean the target before the snapshot

Selective deletion ahead of snapshot removes rows from the target table immediately before a snapshot starts, with the same filter controls. It is a destructive operation on the target and is documented as such.

Custom document keys

Keys shaped for NoSQL, event streaming and data lakes

Compose the document or message key from source fields, separators, prefix and suffix

Document stores, Kafka topics and object-store targets do not have a primary key the way a table does. Gluesync lets you build the key each record is written with: choose the source fields, order them, pick a separator, add a prefix or a suffix, and preview the result before saving. The same key is used for inserts, updates and deletes, so a change on the source always finds its document or message.

  • Default key combines schema, table and primary key; customize it to match the lookups your application performs
  • Order fields, choose the separator, and add a prefix or suffix such as order_ or _gluesync
  • Shown for NoSQL, event-streaming and object-store targets; on Kafka it becomes the message key, and deletes are sent as tombstone records that keep that key
  • Document Key Preview renders the key from a sample row as you edit
  • Changing the format for existing data may require a full resync, and the key must fit the target's length limits

View custom document keys docs ↗

Data transformation
Data pipelines / P-PG-CB Retail / sales.transactions / Settings

Document key

Customize document key
Fields
  1. region · USA-NY
  2. store_id · 1001
  3. transaction_id · 998877
Document Key Preview tx_ - region - store_id - transaction_id - _gluesync Sample · tx_-USA-NY-1001-998877-_gluesync

Spark across all of it

Ask, review, approve

Spark helpers in the editors and the pipeline helper in AI Studio, over the same configuration

Spark reads the pipeline the same way the Control Plane does. In the UDF and Custom Field Functions editors it explains or improves the code in front of you. In AI Studio it can read an entity's mapping, filters and keys, propose a change such as a Mask Email field function on a target column, and apply it only after you approve the plan.

  • Grounded on your pipeline through Core Hub MCP tools: get_pipeline, list_field_functions, compile_mapping_function and upsert_entities
  • Every proposal lists the tool calls it will make; nothing is written without approval
  • Runs with the permissions of the signed-in user; PII-tagged columns are referenced by label, never by value
  • Requires a bring-your-own LLM endpoint configured in AI Studio

View AI Studio docs ↗

Data transformation
AI Studio / Spark / Pipeline helper
AI Studio · SparkPipeline helper · tools run with your permissions
Add a field function that masks the email column on the target of Customer 360.
EMAIL maps dbo.customers.email to CUSTOMERS.EMAIL with a direct mapping and carries a PII:email tag. I propose Mask Email with *: john.doe@example.com → j*******@example.com. This changes the entity, so it needs your approval.
upsert_entities · dbo.customersEMAIL: Direct mapping → Mask Email (char *)

Resources

Data transformation documentation and resources

Roadmap

Public Roadmap

What we're building next. Submit feature requests.

View roadmap ↗
Support

Support

Access technical support and operational assistance.

Get support →
Status

Service Status

Real-time platform availability and incident history.

Check status ↗

Shape the data where it moves.

Request a free trial or talk with the MOLO17 team about field functions, UDFs, filters and document keys on your own sources and targets.