Visual + Code
Drag-and-drop pipeline builder for analysts, native SQL/Python/Scala for engineers — the same execution graph runs both.
Connect your sources, shape your data, and orchestrate your pipelines in one workspace. Every transformation keeps its lineage.
Drag-and-drop pipeline builder for analysts, native SQL/Python/Scala for engineers — the same execution graph runs both.
Unified compute for real-time CDC, micro-batch, and overnight backfills. One pipeline, one set of metrics, one bill.
Per-task lineage, data-quality assertions, and SLA tracking — alerts fire before downstream dashboards go stale.
GeneFlow is a family rather than a single tool. Each member is independently adoptable and shares the same catalog, control plane, and lineage — and the tables it produces are governed by GeneCatalog and consumed by Cortex SQL and Cortex AI without an export step in between.
Managed ingestion from 33 live connectors — databases via log-based CDC, streaming topics, object stores, and SaaS applications — with schema inferred and catalogued on first read.
Learn moreDeclarative batch and streaming transforms with incremental materialisation, so a one-row upstream change rebuilds one partition rather than the whole table.
Learn moreDependency-aware orchestration across every pipeline, with self-healing retries from the last known-good checkpoint and per-task SLA tracking from the first run.
Learn moreVisual authoring for analysts. Drag sources and transforms onto a canvas; the result compiles to the same execution graph an engineer would write in SQL.
Learn moreThe governance product GeneFlow writes into. Classification and column-level lineage are emitted as transforms execute, so every pipeline is covered without separate instrumentation.
Learn moreThe intelligence product GeneFlow feeds. Experiments, feature definitions, and managed endpoints read the governed tables your pipelines produce — no export, no second copy.
Learn moreBuild visually or in SQL, orchestrate, govern, and monitor — without leaving the platform or switching vendors. Move between the tabs to see the surface each workload actually uses.
Sources
Transforms
join_sessions
GeneFlow Designer lets analysts drag sources and transforms onto a canvas; engineers express the same pipeline in SQL, Python, or Scala. Both compile to one execution graph, so a visual pipeline and hand-written code sit in the same DAG with the same lineage.
Explore the pipeline builder-- Incremental: only changed partitions recompute
CREATE MATERIALIZED VIEW gold.revenue_daily
PARTITIONED BY (order_date)
WITH (freshness = '5 minutes') AS
SELECT
o.order_date,
o.region,
SUM(o.amount) AS gross_revenue,
COUNT(DISTINCT o.customer_id) AS buyers
FROM silver.orders o
WHERE o.status = 'settled'
GROUP BY 1, 2;
-- Contract: halts publication on failure
ASSERT gross_revenue >= 0 ON VIOLATION FAIL;
| order_date | region | gross_revenue |
|---|---|---|
| 2026-08-01 | EMEA | 1,284,410 |
| 2026-08-01 | AMER | 2,910,338 |
| 2026-08-01 | APAC | 884,102 |
| 2026-07-31 | EMEA | 1,190,776 |
| 2026-07-31 | AMER | 2,744,015 |
Declare the result you want and a freshness target; the engine works out what to recompute. A one-row upstream change rebuilds one partition, not the whole table.
Read the transformation API| Run | Pipeline | Status | Duration | Rows | Timeline |
|---|---|---|---|---|---|
| run_8f21 | revenue_daily | Succeeded | 4m 12s | 18.2M | |
| run_8f20 | customer_360 | Running | 2m 03s | 6.1M | |
| run_8f19 | inventory_sync | Failed | 0m 48s | — | |
| run_8f18 | clickstream_raw | Succeeded | 11m 30s | 241M | |
| run_8f17 | fx_rates | Queued | — | — |
GeneFlow Jobs runs pipelines on cron, event triggers, or continuously. A late-arriving source delays its dependants rather than letting them publish stale output, and transient failures retry from the last known-good checkpoint.
See orchestration options| Column | Type | Classification | Null % |
|---|---|---|---|
| order_id | bigint | — | 0.00 |
| customer_email | string | PII | 0.01 |
| card_last4 | string | PCI | 2.40 |
| amount | decimal(12,2) | — | 0.00 |
| region | string | — | 0.00 |
Lineage
Owner
data-platform@
Every dataset is registered as it lands, with column contents profiled for PII and PCI before anyone queries them. Lineage is emitted by the engine as transforms execute, so impact analysis is a query rather than an archaeology project.
Explore governanceSignals are emitted for every table without manual instrumentation, and compared against a learned baseline rather than a static threshold. Alerts name every downstream asset affected, so triage starts with known impact.
Explore observabilityEvery pipeline compiles to a single execution graph. Streaming and batch tasks sit in the same DAG, share the same lineage, and are scheduled by the same dependency resolver — so a late-arriving source delays its dependants instead of silently publishing stale output.
The same four steps whether you are moving one table or rebuilding a warehouse — and each step is versioned, so a pipeline can be reviewed and rolled back like any other code change.
Pick from 33 live connectors, or point at any JDBC, REST, or object-store endpoint. Schema is inferred on first read and registered in the catalog.
Compose in the visual builder or write SQL, Python, or Scala. Both produce the same execution graph, so teams can mix approaches in one pipeline.
Attach freshness, volume, uniqueness, and referential assertions. A failing gate halts publication rather than propagating bad data downstream.
Deploy on cron, event triggers, or continuous streaming. Per-task metrics, lineage, and SLA tracking are on from the first run.
Pipeline work is mostly the unglamorous parts — backfills, schema drift, retries, and late data. These are handled in the runtime rather than left as an exercise for each team.
Log-based CDC for Postgres, MySQL, Oracle, and SQL Server with transactional ordering preserved end to end.
Only changed partitions recompute. A one-row upstream change does not trigger a full-table rebuild.
New columns propagate automatically; breaking changes raise a review rather than failing at 3am.
Re-run any window against current logic with full idempotency, without hand-written catch-up jobs.
Transient failures retry with backoff against the last known-good checkpoint instead of restarting the DAG.
Every field traces to its sources through each transform, so impact analysis is a query rather than an archaeology project.
The same engine serves very different jobs — the difference is which surface each role works in.
Replace separate ingestion, transformation, and scheduling tools with one graph. Dependencies resolve across what used to be three systems and three on-call rotations.
One runtime, one alerting surface, one bill.
Build in the visual editor against governed sources. Output lands in the catalog with lineage attached, so it is reviewable rather than a shadow dataset.
Self-service that governance can actually approve.
Freshness and volume assertions fire on the pipeline, not on the dashboard. Incidents surface at the failing task with lineage showing exactly what is downstream.
Detection ahead of the first user complaint.
Assembled stacks lose context at every hop between tools. A single execution graph keeps it.
| Dimension | Before Genedata | With Genedata |
|---|---|---|
| Tooling | Separate ingestion, transformation, and orchestration vendors | One engine covering the full lifecycle |
| Lineage | Breaks at each tool boundary; impact analysis is manual | Column-level and continuous across every task |
| Backfills | Bespoke catch-up scripts, frequently non-idempotent | Replay any window against current logic |
| Bad data | Detected downstream, in a dashboard, by a stakeholder | Quality gate halts publication at the failing task |
| Schema drift | Overnight failure, then a manual reconciliation | Additive changes propagate; breaking ones raise a review |
Stop paying three vendors for ingestion, transformation, and orchestration. GeneFlow covers the full lifecycle — and because it's the same engine running your warehouse and AI, your pipelines never leave the platform.
Interactive walkthrough
Six scenes on GeneFlow: watch a pipeline become a contract between steps, break it on purpose, heal it with a review, price a backfill before it runs, and move production by changing a pointer.
Also related
The questions that come up most often in evaluations. If yours isn't here, the architecture guide goes deeper.
GeneFlow is the data engineering product line on the Genedata platform. It covers ingestion, transformation, and orchestration as one product rather than three tools — GeneFlow Connect brings data in, GeneFlow Pipelines transforms it, and GeneFlow Jobs schedules the whole graph. Because they share a catalog and control plane, lineage and governance are continuous rather than reassembled at each boundary.
Most connector libraries move bytes and stop there. GeneFlow Connect registers each dataset in the catalog on first read, infers and versions its schema, and emits lineage from the moment data lands — so a newly connected source is governed before anyone queries it, rather than after someone notices it isn't.
No. GeneFlow Designer and hand-written SQL, Python, or Scala compile to the same execution graph, so one pipeline can contain both. An analyst's visual transform and an engineer's code sit in the same DAG, share the same lineage, and are scheduled by the same resolver.
Transient failures retry with backoff from the last known-good checkpoint rather than restarting the DAG. If a data-quality assertion fails, publication halts at that task instead of propagating bad data downstream, and the alert names every downstream asset that would have been affected.
Additive changes — a new column upstream — propagate automatically. Breaking changes raise a review rather than failing overnight, and because every schema version is retained you can answer what a table looked like on the date a given report was produced.
Yes. GeneFlow deploys in the Genedata enterprise cloud, inside your own AWS, Azure, or GCP account, or into a jurisdiction-bound sovereign region. The APIs, governance model, and security posture are identical across all three; only the perimeter changes.
Execution model, connector catalog, and the transformation API in full.
RunbookOperational patterns for backfills, replay, and incident response.
CatalogThe full list of managed sources with CDC and streaming support.
StatusLive health and incident history across every Genedata region.
Walk through a live pipeline with an engineer, or go straight to the architecture guide and connector catalog.