<!-- Generated from the rendered page by scripts/write-llm-mirrors.mjs. Do not edit by hand. -->
Canonical: https://molo17.com/data-transformation/
Markdown mirror: https://molo17.com/data-transformation/index.md
Title: Data transformation · Gluesync · MOLO17
Description: Gluesync data transformation: pre-built field functions and masking, Java User Defined Functions, reusable Custom Field Functions, inclusive filters applied on snapshots and CDC, custom document keys for NoSQL, Kafka and data lakes, and Spark AI helpers in every editor.

Gluesync Core Hub · Built into every pipeline

# Data transformation

Reshape data while it moves. Pick a pre-built field function per column, write a Java UDF when a record needs real logic, publish reusable Custom Field Functions at pipeline level, keep only the rows that match your filters, and decide how documents and messages are keyed on NoSQL, event-streaming and object-store targets. Spark helps you write and review all of it.

[See how it works](#how-it-works) [Talk to our team](/contacts/?topic=gluesync-platform#contact-form-section)

 Data transformation

1.  01 Map and convert fields without code
2.  02 Add Java logic when the catalog is not enough
3.  03 Decide which rows, and which keys, reach the target

Why transform in transit

## The target rarely wants the source as it is

01 / Shape mismatch

### Different engines, different shapes

Oracle dates, SQL Server decimals and MySQL booleans do not land cleanly in Couchbase documents, Kafka messages or BigQuery columns. Fixing them after the fact means another job, another schedule and another copy.

02 / Compliance at the edge

### Sensitive data should not travel

If a target team only needs the last four digits of a card, the full number should never reach them. Masking, filtering and anonymization belong in the pipeline, not in a downstream view somebody may forget to apply.

03 / The platform outcome

### One pipeline, snapshot and CDC alike

Gluesync applies the same field functions, UDFs, filters and document keys to the initial snapshot and to every change captured afterwards, from the Control Plane, with no extra component to run.

How it works

## From a mapped column to a shaped record

1.  01
    
    Map
    
    ### Pick the column, pick the function
    
    In the Fields Editor, map source to target columns and attach a pre-built field function where the shape has to change: casts, date patterns, trims, constants, technical fields and masking. Parameters come with ready-to-use defaults.
    
2.  02
    
    Extend
    
    ### Write code only where it earns its keep
    
    When a record needs business logic, add a Java UDF to the entity. When several pipelines need the same custom conversion, publish it once as a Custom Field Function and bind it to target columns like any catalog entry.
    
3.  03
    
    Scope
    
    ### Filter rows and shape keys
    
    Add inclusive filter clauses that decide which rows are copied on snapshot and kept in sync on CDC, and compose the document or message key your NoSQL, Kafka or object-store target expects.
    

Capabilities

## What ships with Core Hub

Every capability on this page is configured in the Control Plane and stored with the pipeline, so Bootstrapper exports and REST calls carry the same configuration. The simulations below mirror the Core Hub screens where each one lives.

Field functions

### A catalog of no-code transformations

#### Eight categories of pre-built functions, applied per column in the Fields Editor

Field functions are applied in transit to a mapped column. Regular transformations change business data on its way to the target; technical field functions populate target-only columns with timestamps or constants, or override the incoming value, which is how data masking and anonymization work without code.

-   Date and time to string, string to date and time, and advanced date and time operations such as zone-aware timestamp conversions
-   Numeric and decimal conversions, boolean conversions, and text and binary operations
-   Technical fields and constants for execution timestamps, record IDs built from the primary key, or fixed values
-   Data masking and anonymization: Mask Credit Card, Mask String, Mask Email, Mask Phone and Mask IBAN
-   Each function has a YAML expression type, so the same configuration can be exported and applied with Bootstrapper

[View field functions docs ↗](https://docs.molo17.com/gluesync/latest/core-hub/field-functions.html)

 Data transformation

### Masking pre-built, not hand-rolled

Pick the mask function, keep the defaults or set the mask character, and the target only ever receives the redacted value. Discovery of which columns need it is covered on the PII discovery and masking page.

User Defined Functions

### Java logic on every change event

#### An onChange handler per entity, with new and old values, the operation and a logger

A UDF receives each insert, update and delete after it is read from the source and before it is written to the target. It returns the record to write, or null to skip it, and can raise an error that goes through the pipeline's error handling. Code is checked and compiled in Core Hub before it can run.

-   Signature onChange(newValues, oldValues, operation, logger): mutate, enrich, route or drop a record
-   Old values are available for updates and deletes on sources that capture before images
-   Java today; Kotlin, JavaScript and Python are listed as available soon in the docs
-   Checked and compiled before execution, run sandboxed with execution timeouts, with a logger that writes to Core Hub logs
-   Spark helper in the editor: ask it to explain or improve the UDF, insert the result and compile

[View UDF docs ↗](https://docs.molo17.com/gluesync/latest/core-hub/user-defined-functions.html)

 Data transformation

Custom Field Functions

### Reusable functions at pipeline level

#### Java methods exported with @GsFunction, published, versioned and bound to columns

When the catalog does not have the conversion you need, or a target column has to be built from several source columns at once, write it once as a Custom Field Function. It belongs to the pipeline, shows up in the Fields Editor next to the pre-built catalog, and can be bound to as many target columns and entities as you like.

-   Exported methods are public static and annotated with @GsFunction, with typed parameters mapped from source columns
-   Compile and publish checks the code, recompiles every UDF in the pipeline and re-checks every bound column before storing a version
-   Install is a separate step: entities are stopped, the executable replaced and restarted, and Gluesync reports which ones
-   Versions can be restored; Try a function runs the installed version on sample values before you bind it
-   Repeatedly failing functions are suspended instead of failing on every record

[View Custom Field Functions docs ↗](https://docs.molo17.com/gluesync/latest/core-hub/custom-field-functions.html)

 Data transformation

Filters

### Inclusive filters on snapshots and on CDC

#### Per-entity clauses evaluated in Core Hub, on the initial copy and on every change

Filters decide which rows belong on the target. A row is replicated only if it matches all clauses. During the snapshot only matching rows are copied; during CDC a row that stops matching is converted into a DELETE on the target, so the target always reflects the filter. Filters are evaluated in Core Hub after the read, so they add no load on the source.

-   Operators: =, !=, >, <, >=, <=, In, Not in, Is null, Not null and Regex
-   Combine clauses on several columns; all of them must match for the row to be kept
-   Enabled per entity from Filter out incoming data in the entity settings dialog
-   Selective deletion ahead of snapshot reuses the same operators to prune target rows right before a snapshot task runs

[View data filtering docs ↗](https://docs.molo17.com/gluesync/latest/core-hub/data-filtering.html)

 Data transformation

### Clean the target before the snapshot

Selective deletion ahead of snapshot removes rows from the target table immediately before a snapshot starts, with the same filter controls. It is a destructive operation on the target and is documented as such.

Custom document keys

### Keys shaped for NoSQL, event streaming and data lakes

#### Compose the document or message key from source fields, separators, prefix and suffix

Document stores, Kafka topics and object-store targets do not have a primary key the way a table does. Gluesync lets you build the key each record is written with: choose the source fields, order them, pick a separator, add a prefix or a suffix, and preview the result before saving. The same key is used for inserts, updates and deletes, so a change on the source always finds its document or message.

-   Default key combines schema, table and primary key; customize it to match the lookups your application performs
-   Order fields, choose the separator, and add a prefix or suffix such as order\_ or \_gluesync
-   Shown for NoSQL, event-streaming and object-store targets; on Kafka it becomes the message key, and deletes are sent as tombstone records that keep that key
-   Document Key Preview renders the key from a sample row as you edit
-   Changing the format for existing data may require a full resync, and the key must fit the target's length limits

[View custom document keys docs ↗](https://docs.molo17.com/gluesync/latest/core-hub/custom-document-keys.html)

 Data transformation

Spark across all of it

### Ask, review, approve

#### Spark helpers in the editors and the pipeline helper in AI Studio, over the same configuration

Spark reads the pipeline the same way the Control Plane does. In the UDF and Custom Field Functions editors it explains or improves the code in front of you. In AI Studio it can read an entity's mapping, filters and keys, propose a change such as a Mask Email field function on a target column, and apply it only after you approve the plan.

-   Grounded on your pipeline through Core Hub MCP tools: get\_pipeline, list\_field\_functions, compile\_mapping\_function and upsert\_entities
-   Every proposal lists the tool calls it will make; nothing is written without approval
-   Runs with the permissions of the signed-in user; PII-tagged columns are referenced by label, never by value
-   Requires a bring-your-own LLM endpoint configured in AI Studio

[View AI Studio docs ↗](https://docs.molo17.com/gluesync/latest/gs-modules/ai-studio.html)

 Data transformation

Explore the platform

## Continue exploring Gluesync

Transformation is one part of the pipeline. The pages below cover what it connects to: PII discovery so you know what to mask, Spark and the MCP server that drive the helpers, and Query Studio to check the result.

[### PII discovery & masking

Columns classified as PII carry a tag in the Fields Editor, UDF, filter and document key screens, so you know what to mask before it leaves the pipeline.

Explore PII discovery & masking →](/pii-discovery-and-masking/) [### Spark Agents

The agents behind the Spark helpers in the UDF and Custom Field Functions editors, and the pipeline helper in AI Studio.

Explore Spark Agents →](/spark/) [### Core Hub MCP Server

The same field functions, mapping compiler and entity upserts, exposed as tools to Spark and to your own AI clients.

Explore Core Hub MCP Server →](/corehub-mcp/) [

### Query Studio

Check source values before you write a filter, and read the target after the transformation lands, from the same Control Plane.

Explore Query Studio →](/query-studio/)

Resources

## Data transformation documentation and resources

Documentation 

### Documentation

Technical guides, API reference, deployment tutorials.

[Read documentation ↗](https://docs.molo17.com)

Roadmap 

### Public Roadmap

What we're building next. Submit feature requests.

[View roadmap ↗](https://roadmap.molo17.com)

Support 

### Support

Access technical support and operational assistance.

[Get support →](/support/)

Status 

### Service Status

Real-time platform availability and incident history.

[Check status ↗](https://status.molo17.com)

## Shape the data where it moves.

Request a free trial or talk with the MOLO17 team about field functions, UDFs, filters and document keys on your own sources and targets.

[Request a free trial](/get-gluesync/) [Talk to our team](/contacts/?topic=gluesync-platform#contact-form-section) [Docs](https://docs.molo17.com/gluesync/latest/core-hub/fields-editor.html)
