Automatic Lineage
Column-level lineage captured by the engine itself — no manual annotation, no broken links, always current with your pipelines.
Connect datasets, definitions, and lineage into a shared map of your business. Find the right data and understand where it came from.
Column-level lineage captured by the engine itself — no manual annotation, no broken links, always current with your pipelines.
ML-driven discovery tags PII, regulated data, and business-critical assets the moment they enter the platform.
Connect platform capabilities, business definitions, catalog assets, analytics, pipelines, and solution packs in a tenant-scoped graph.
Lineage is emitted by the execution engine as queries run, not reconstructed afterwards from logs or hand-maintained diagrams. That makes it complete by construction — including the ad-hoc transform someone wrote last Tuesday.
Assets move through the same lifecycle whether they arrive from a managed connector, a notebook, or an analyst's ad-hoc table — so nothing sits outside the catalog by accident.
Every dataset is registered as it lands, with schema, owner, and source captured automatically. There is no separate crawl to fall behind.
Column contents are profiled and tagged — PII, PHI, payment data, business-critical fields — before anyone queries them.
The engine records column-level provenance as transforms execute, linking each derived field back through every hop to its origin.
Retention and purge rules execute across every tier and replica, leaving a signed record of what was deleted and when.
A catalog is only useful if it is complete and current. These are the mechanisms that keep it both.
Find any table, column, dashboard, model, or pipeline across every environment from a single search surface.
Content-based classification catches sensitive data even when column names give nothing away.
Before changing a column, see every downstream table, metric, dashboard, and model that depends on it.
Promote trusted datasets to certified status with an owner, a data contract, and a review cadence attached.
Traverse the relationships among datasets, definitions, pipelines, analytics, solution packs, owners, and platform capabilities from one governed view.
Every asset has an accountable owner. Orphaned datasets surface in a report rather than lingering unnoticed.
The catalog is the layer everyone touches, usually without thinking about it.
Search by business concept rather than guessing at table names, and see which datasets are certified before building on them.
Less time hunting, fewer reports built on the wrong source.
Query the catalog for every location holding a subject's data, including derived tables and downstream copies.
Requests answered from evidence, not from memory.
Run impact analysis before a migration to see every dependant, then notify owners directly from the catalog.
Breaking changes caught in review, not in production.
Bolt-on catalogs describe the platform from the outside, and drift the moment someone works around them.
| Dimension | Before Genedata | With Genedata |
|---|---|---|
| Coverage | Whatever the last crawl happened to find | Every asset, registered as it is created |
| Lineage | Hand-maintained diagrams, stale within weeks | Emitted by the engine at execution time |
| Sensitive data | Found by column-name heuristics | Detected by profiling actual contents |
| Impact analysis | Grep the codebase and hope | A query over the dependency graph |
| Retention | Scripts per storage system, unevenly applied | One policy executed across every tier and replica |
When every analyst, engineer, AI agent, and business owner works from the same connected knowledge layer, discovery becomes faster and every decision retains the evidence behind it.
What evaluators ask most about ingestion and the catalog.
GeneCatalog is the metadata and governance layer for the platform. Every dataset is registered as it is written — including anything GeneFlow Connect ingests from its catalogued sources — with schema inferred, contents classified, and lineage emitted from the moment it lands.
Assets are registered as they are written rather than found by a periodic crawl. That includes datasets created by a notebook or an analyst's ad-hoc table, so nothing sits outside the catalog because a scan had not run yet.
Column contents are profiled rather than column names matched, so PII, PHI, and payment data are caught even when the column is called col_7. Classifications become the tags that governance policies match against, which is why a newly landed sensitive column is covered on arrival.
Yes. Every schema version is retained, so you can answer what a table's structure was on the date a given report was produced — which is usually what an auditor is actually asking.
Impact analysis queries the dependency graph to show every downstream table, metric, dashboard, and model that depends on it, and lets you notify the owners directly from the catalog. Breaking changes get caught in review rather than in production.
Every managed source with CDC, streaming, and incremental support.
SpecificationCatalog schema, lineage capture, and classification policy model.
GuideCertifying datasets with owners, contracts, and review cadence.
RelatedWhat happens to data after Connect has landed and catalogued it.
See GeneCatalog classify and trace a live source end to end in a working session — or start with the metadata layer specification.