GeneCatalog · Data managementGENEDATA / 01

Know your data.Understand its impact.

Connect datasets, definitions, and lineage into a shared map of your business. Find the right data and understand where it came from.

Shared context · Lineage · Governance
Connected
Your sources
Datasets
Business definitions
Relationships
Connected intelligenceGeneCatalog
Governance
Business impactConnected knowledge
Shared contextLineageGovernance
+Illustrative workflow01 / 03
1 / 3
01

Automatic Lineage

Column-level lineage captured by the engine itself — no manual annotation, no broken links, always current with your pipelines.

02

Smart Classification

ML-driven discovery tags PII, regulated data, and business-critical assets the moment they enter the platform.

03

Enterprise Knowledge Graph

Connect platform capabilities, business definitions, catalog assets, analytics, pipelines, and solution packs in a tenant-scoped graph.

Lineage Graph

Every column traced to its origin, automatically.

Lineage is emitted by the execution engine as queries run, not reconstructed afterwards from logs or hand-maintained diagrams. That makes it complete by construction — including the ad-hoc transform someone wrote last Tuesday.

CRMcontactsBillinginvoicesEventsclickstreamClassifyPII taggingModelconformedcustomer_360governedrevenue_dailycertifiedExec dashboard
How It Works

Discover, classify, govern, retire.

Assets move through the same lifecycle whether they arrive from a managed connector, a notebook, or an analyst's ad-hoc table — so nothing sits outside the catalog by accident.

  1. Discover on write

    Every dataset is registered as it lands, with schema, owner, and source captured automatically. There is no separate crawl to fall behind.

  2. Classify the contents

    Column contents are profiled and tagged — PII, PHI, payment data, business-critical fields — before anyone queries them.

  3. Attach lineage

    The engine records column-level provenance as transforms execute, linking each derived field back through every hop to its origin.

  4. Retire on policy

    Retention and purge rules execute across every tier and replica, leaving a signed record of what was deleted and when.

Capabilities

What the catalog gives you.

A catalog is only useful if it is complete and current. These are the mechanisms that keep it both.

Federated search

Find any table, column, dashboard, model, or pipeline across every environment from a single search surface.

PII and PHI detection

Content-based classification catches sensitive data even when column names give nothing away.

Impact analysis

Before changing a column, see every downstream table, metric, dashboard, and model that depends on it.

Certification workflow

Promote trusted datasets to certified status with an owner, a data contract, and a review cadence attached.

Knowledge graph

Traverse the relationships among datasets, definitions, pipelines, analytics, solution packs, owners, and platform capabilities from one governed view.

Ownership and stewardship

Every asset has an accountable owner. Orphaned datasets surface in a report rather than lingering unnoticed.

In Practice

Who this is for.

The catalog is the layer everyone touches, usually without thinking about it.

Data Analyst

Find the right table the first time

Search by business concept rather than guessing at table names, and see which datasets are certified before building on them.

Less time hunting, fewer reports built on the wrong source.

Compliance

Answer a subject request without a fire drill

Query the catalog for every location holding a subject's data, including derived tables and downstream copies.

Requests answered from evidence, not from memory.

Data Engineer

Change a schema without breaking things

Run impact analysis before a migration to see every dependant, then notify owners directly from the catalog.

Breaking changes caught in review, not in production.

What Changes

What changes when the catalog is built in.

Bolt-on catalogs describe the platform from the outside, and drift the moment someone works around them.

DimensionBefore GenedataWith Genedata
CoverageWhatever the last crawl happened to findEvery asset, registered as it is created
LineageHand-maintained diagrams, stale within weeksEmitted by the engine at execution time
Sensitive dataFound by column-name heuristicsDetected by profiling actual contents
Impact analysisGrep the codebase and hopeA query over the dependency graph
RetentionScripts per storage system, unevenly appliedOne policy executed across every tier and replica
Real-timeCatalog Refresh
84Catalogued Sources
AI-DrivenPII Detection
Column-levelLineage Depth
The next step

Know the data, the meaning, and the impact.

When every analyst, engineer, AI agent, and business owner works from the same connected knowledge layer, discovery becomes faster and every decision retains the evidence behind it.

FAQ

GeneCatalog, answered.

What evaluators ask most about ingestion and the catalog.

What is GeneCatalog?

GeneCatalog is the metadata and governance layer for the platform. Every dataset is registered as it is written — including anything GeneFlow Connect ingests from its catalogued sources — with schema inferred, contents classified, and lineage emitted from the moment it lands.

How does catalogue coverage stay complete?

Assets are registered as they are written rather than found by a periodic crawl. That includes datasets created by a notebook or an analyst's ad-hoc table, so nothing sits outside the catalog because a scan had not run yet.

How is sensitive data detected?

Column contents are profiled rather than column names matched, so PII, PHI, and payment data are caught even when the column is called col_7. Classifications become the tags that governance policies match against, which is why a newly landed sensitive column is covered on arrival.

Can I see what a table looked like last quarter?

Yes. Every schema version is retained, so you can answer what a table's structure was on the date a given report was produced — which is usually what an auditor is actually asking.

What happens before I change a column?

Impact analysis queries the dependency graph to show every downstream table, metric, dashboard, and model that depends on it, and lets you notify the owners directly from the catalog. Breaking changes get caught in review rather than in production.

Take the next step

Put every dataset on the map.

See GeneCatalog classify and trace a live source end to end in a working session — or start with the metadata layer specification.