GeneFlow · Data engineeringGENEDATA / 01

Build the flow.Keep the context.

Connect your sources, shape your data, and orchestrate your pipelines in one workspace. Every transformation keeps its lineage.

Shared context · Lineage · Governance
Connected
Your sources
Database CDC
Streaming events
Files & APIs
Connected intelligenceGeneFlow pipeline
Governance
Business impactTrusted data products
Shared contextLineageGovernance
+Illustrative workflow01 / 03
1 / 3
01

Visual + Code

Drag-and-drop pipeline builder for analysts, native SQL/Python/Scala for engineers — the same execution graph runs both.

02

Streaming + Batch

Unified compute for real-time CDC, micro-batch, and overnight backfills. One pipeline, one set of metrics, one bill.

03

Built-in Observability

Per-task lineage, data-quality assertions, and SLA tracking — alerts fire before downstream dashboards go stale.

The GeneFlow Family

Four members, and the two products they hand off to.

GeneFlow is a family rather than a single tool. Each member is independently adoptable and shares the same catalog, control plane, and lineage — and the tables it produces are governed by GeneCatalog and consumed by Cortex SQL and Cortex AI without an export step in between.

See It Working

Five workloads, one surface.

Build visually or in SQL, orchestrate, govern, and monitor — without leaving the platform or switching vendors. Move between the tabs to see the surface each workload actually uses.

Build
Pipelines / revenue_dailyValidateDeploy
postgres.ordersCDC · 12.4k/skafka.clicksstream · 88k/sjoin_sessionswindowed 5mdedupe_ordersincrementalgold.revenuematerialised

Compose a pipeline visually — or write it in SQL

GeneFlow Designer lets analysts drag sources and transforms onto a canvas; engineers express the same pipeline in SQL, Python, or Scala. Both compile to one execution graph, so a visual pipeline and hand-written code sit in the same DAG with the same lineage.

Explore the pipeline builder
  1. GeneFlow Connect supplies 33 live connectors from a catalogue of 84, catalogued on first read
  2. Windowed stream joins alongside batch transforms in one graph
  3. Assertions attach to any node and halt publication on failure
Compose a pipeline visually — or write it in SQL
Execution Graph

One DAG, from change capture to serving layer.

Every pipeline compiles to a single execution graph. Streaming and batch tasks sit in the same DAG, share the same lineage, and are scheduled by the same dependency resolver — so a late-arriving source delays its dependants instead of silently publishing stale output.

PostgresCDC logKafkaclickstreamS3batch dropIngestschema inferenceStream joinwindowedTransformincrementalQuality gateassertionsWarehouseFeature store
How It Works

Connect, model, validate, ship.

The same four steps whether you are moving one table or rebuilding a warehouse — and each step is versioned, so a pipeline can be reviewed and rolled back like any other code change.

  1. Connect the source

    Pick from 33 live connectors, or point at any JDBC, REST, or object-store endpoint. Schema is inferred on first read and registered in the catalog.

  2. Model the transform

    Compose in the visual builder or write SQL, Python, or Scala. Both produce the same execution graph, so teams can mix approaches in one pipeline.

  3. Assert the contract

    Attach freshness, volume, uniqueness, and referential assertions. A failing gate halts publication rather than propagating bad data downstream.

  4. Schedule and observe

    Deploy on cron, event triggers, or continuous streaming. Per-task metrics, lineage, and SLA tracking are on from the first run.

Capabilities

What the engine handles for you.

Pipeline work is mostly the unglamorous parts — backfills, schema drift, retries, and late data. These are handled in the runtime rather than left as an exercise for each team.

Change data capture

Log-based CDC for Postgres, MySQL, Oracle, and SQL Server with transactional ordering preserved end to end.

Incremental transforms

Only changed partitions recompute. A one-row upstream change does not trigger a full-table rebuild.

Schema-drift handling

New columns propagate automatically; breaking changes raise a review rather than failing at 3am.

Backfill and replay

Re-run any window against current logic with full idempotency, without hand-written catch-up jobs.

Self-healing retries

Transient failures retry with backoff against the last known-good checkpoint instead of restarting the DAG.

Column-level lineage

Every field traces to its sources through each transform, so impact analysis is a query rather than an archaeology project.

In Practice

Who this is for.

The same engine serves very different jobs — the difference is which surface each role works in.

Data Engineer

Retire the orchestration zoo

Replace separate ingestion, transformation, and scheduling tools with one graph. Dependencies resolve across what used to be three systems and three on-call rotations.

One runtime, one alerting surface, one bill.

Data Analyst

Ship a pipeline without filing a ticket

Build in the visual editor against governed sources. Output lands in the catalog with lineage attached, so it is reviewable rather than a shadow dataset.

Self-service that governance can actually approve.

Platform / SRE

Know before the business does

Freshness and volume assertions fire on the pipeline, not on the dashboard. Incidents surface at the failing task with lineage showing exactly what is downstream.

Detection ahead of the first user complaint.

What Changes

What changes when the pipeline stops leaving the platform.

Assembled stacks lose context at every hop between tools. A single execution graph keeps it.

DimensionBefore GenedataWith Genedata
ToolingSeparate ingestion, transformation, and orchestration vendorsOne engine covering the full lifecycle
LineageBreaks at each tool boundary; impact analysis is manualColumn-level and continuous across every task
BackfillsBespoke catch-up scripts, frequently non-idempotentReplay any window against current logic
Bad dataDetected downstream, in a dashboard, by a stakeholderQuality gate halts publication at the failing task
Schema driftOvernight failure, then a manual reconciliationAdditive changes propagate; breaking ones raise a review
33 live · 84 cataloguedConnectors
CDC + micro-batchStream Mode
UnlimitedPipeline DAGs
Per-taskSLA Tracking
The next step

From source to production, on one engine.

Stop paying three vendors for ingestion, transformation, and orchestration. GeneFlow covers the full lifecycle — and because it's the same engine running your warehouse and AI, your pipelines never leave the platform.

Interactive walkthrough

From handoff to contract

Six scenes on GeneFlow: watch a pipeline become a contract between steps, break it on purpose, heal it with a review, price a backfill before it runs, and move production by changing a pointer.

FAQ

GeneFlow, answered.

The questions that come up most often in evaluations. If yours isn't here, the architecture guide goes deeper.

What is GeneFlow?

GeneFlow is the data engineering product line on the Genedata platform. It covers ingestion, transformation, and orchestration as one product rather than three tools — GeneFlow Connect brings data in, GeneFlow Pipelines transforms it, and GeneFlow Jobs schedules the whole graph. Because they share a catalog and control plane, lineage and governance are continuous rather than reassembled at each boundary.

How is GeneFlow Connect different from a standard connector library?

Most connector libraries move bytes and stop there. GeneFlow Connect registers each dataset in the catalog on first read, infers and versions its schema, and emits lineage from the moment data lands — so a newly connected source is governed before anyone queries it, rather than after someone notices it isn't.

Do I have to choose between the visual builder and code?

No. GeneFlow Designer and hand-written SQL, Python, or Scala compile to the same execution graph, so one pipeline can contain both. An analyst's visual transform and an engineer's code sit in the same DAG, share the same lineage, and are scheduled by the same resolver.

What happens when a pipeline fails partway through?

Transient failures retry with backoff from the last known-good checkpoint rather than restarting the DAG. If a data-quality assertion fails, publication halts at that task instead of propagating bad data downstream, and the alert names every downstream asset that would have been affected.

How does GeneFlow handle schema drift?

Additive changes — a new column upstream — propagate automatically. Breaking changes raise a review rather than failing overnight, and because every schema version is retained you can answer what a table looked like on the date a given report was produced.

Can I run GeneFlow in my own cloud account?

Yes. GeneFlow deploys in the Genedata enterprise cloud, inside your own AWS, Azure, or GCP account, or into a jurisdiction-bound sovereign region. The APIs, governance model, and security posture are identical across all three; only the perimeter changes.

Take the next step

Start building on GeneFlow.

Walk through a live pipeline with an engineer, or go straight to the architecture guide and connector catalog.