Data science & MLOpsGENEDATA / 01

From explorationto everyday impact.

Keep notebooks, experiments, and deployed models connected to the same data and governance. Carry context through the model lifecycle.

Shared context · Lineage · Governance
Connected
Your sources
Notebooks
Training datasets
Experiments
Connected intelligenceModel lifecycle
Governance
Business impactServing & monitoring
Shared contextLineageGovernance
+Illustrative workflow01 / 03
1 / 3
01

Notebook Workspaces

Auto-scaling Jupyter and VS Code environments with first-class access to warehouse tables, vector stores, and the model registry.

02

Experiment Tracking

Every run, parameter, metric, and artifact captured automatically — search, compare, and promote models with full reproducibility.

03

One-Click Deploy

Promote any model from notebook to managed serving endpoint with autoscaling, A/B testing, and integrated observability.

Model Lifecycle

A loop, not a handoff.

Models degrade because the world moves. Treating the lifecycle as a closed loop — where monitored production behaviour feeds the next round of features and training — is the difference between a model that keeps earning its place and one nobody dares retire.

Governed dataone source1Explorenotebooks on live tables2Feature engineeringversioned, shared3Traintracked runs4Evaluateheld-out + fairness5Deploymanaged endpoint6Monitordrift + performance
How It Works

Explore, train, promote, watch.

Each stage writes its inputs and outputs to the registry, so any production model can be traced back to the exact data and code that produced it.

  1. Explore on governed data

    Notebooks query warehouse tables directly under the same policy as everyone else — no extracts to a laptop, no shadow copies.

  2. Train with tracking on

    Parameters, metrics, code version, and dataset snapshot are captured per run automatically rather than by convention.

  3. Promote through the registry

    Models move from staging to production behind approval gates, with evaluation results attached to the promotion.

  4. Watch for drift

    Input distributions and prediction quality are monitored against training baselines, triggering retraining when they diverge.

Capabilities

What the workspace provides.

Most of the friction in shipping a model is environment, access, and reproducibility rather than modelling.

Managed environments

Auto-scaling Jupyter and VS Code with pinned images, so a notebook that ran last quarter still runs today.

Feature store

Versioned, shared feature definitions used identically in training and serving, eliminating training-serving skew.

Experiment comparison

Search and diff runs across parameters and metrics, with artefacts retained for every candidate.

Model registry

Lineage from a deployed endpoint back through training run, dataset snapshot, and code commit.

Fairness evaluation

Subgroup performance reported alongside aggregate metrics, recorded as part of the promotion evidence.

Managed serving

Autoscaling endpoints with canary rollout, A/B comparison, and request-level observability included.

In Practice

Who this is for.

The workspace has to satisfy the people building models and the people accountable for them.

Data Scientist

Stop rebuilding the environment

Start from a managed workspace with governed data already accessible, instead of spending the first week on access and dependencies.

Time spent modelling rather than provisioning.

ML Engineer

Ship without a rewrite

The features used in training are the features served in production, defined once, so promotion is not a reimplementation.

No training-serving skew to debug.

Risk / Compliance

Explain a model in production

Trace any prediction back to the model version, training data snapshot, and evaluation results that justified deployment.

Model risk documentation from the record.

What Changes

What changes when the workspace sits on the platform.

Data science outside the platform means copies, and copies mean governance gaps.

DimensionBefore GenedataWith Genedata
Data accessExtracts copied to laptops and notebooksGoverned queries under the same policy as everyone
ReproducibilityReconstructed from a notebook and hopeRun, dataset snapshot, and commit captured automatically
FeaturesReimplemented for serving, subtly differentlyOne definition used in training and production
PromotionA handoff ticket to a platform teamRegistry promotion with evaluation attached
DriftNoticed when a business metric movesMonitored against the training baseline continuously
PyTorch · TF · JAXFrameworks
Bring your own GPUCompute
Versioned + pinnedModel Registry
Batch + onlineServing
The next step

From hypothesis to production endpoint.

Data scientists shouldn't need to leave the platform to ship a model. Genedata Data Science integrates natively with your warehouse, governance, and CI/CD — so the path from notebook to production is days, not quarters.

FAQ

Cortex AI, answered.

The questions data science and model risk teams ask before committing.

What is Cortex AI?

Cortex AI is the intelligence product on the Genedata platform — notebooks, experiment tracking, a feature store, the model registry, managed serving endpoints, and the agent runtime. It reads the same governed tables GeneFlow pipelines write and honours GeneCatalog policy, so training data is under the same rules as everything else.

Do data scientists get their own copy of the data?

No, and that is the point. Notebooks query warehouse tables directly under the same access policy as any other user. There are no extracts on laptops, which removes the most common route by which governed data becomes ungoverned.

How is training-serving skew prevented?

Feature definitions live in the feature store and are used identically in training and serving. A model promoted to production consumes the same feature computation it was trained on, so there is no reimplementation step to diverge.

Can I reproduce a model that shipped a year ago?

Yes. Each run captures parameters, metrics, code version, and a dataset snapshot automatically, and the registry links a deployed endpoint back through its training run to the exact data behind it. That trace is what model risk documentation usually needs.

What happens when a model drifts?

Input distributions and prediction quality are monitored against the training baseline continuously, rather than being noticed when a business metric moves. Divergence triggers a retraining signal with the affected features identified.

Take the next step

Take a model from notebook to endpoint.

Walk the full lifecycle with an ML engineer, or start with the MLOps reference architecture.