Workbook

Data Quality Engineer

You own: ensuring data + model quality contracts hold. GeneFlow's drift baselines + eval sets are core to your toolkit.

What's new for you

WasNow in GeneFlow
Bespoke drift scripts per modelserving.save_drift_baselines() + automatic PSI checks
Eval results in spreadsheetsgf_eval_runs rows w/ per-example results
No record of "passed gate at deploy time"Drift baseline + eval run tied to model version

Day-1 setup

pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io

Workflow 1 — Define a drift baseline at training time

import numpy as np
from geneflow import serving

# For each numeric feature
features = []
for col in NUMERIC_FEATURES:
    hist, edges = np.histogram(df[col].dropna(), bins=10)
    features.append({
        "feature_name": col,
        "feature_type": "numeric",
        "histogram": {"buckets": hist.tolist(), "edges": edges.tolist()},
        "mean":   float(df[col].mean()),
        "stddev": float(df[col].std()),
        "min":    float(df[col].min()),
        "max":    float(df[col].max()),
        "sample_size": len(df),
    })

# Categorical:
for col in CATEGORICAL_FEATURES:
    counts = df[col].value_counts().to_dict()
    features.append({
        "feature_name": col,
        "feature_type": "categorical",
        "histogram": {"cats": counts},
        "sample_size": len(df),
    })

serving.save_drift_baselines(
    model_name="fraud-detector",
    model_version=3,
    features=features,
)

Workflow 2 — Confirm a baseline exists before promoting

baselines = serving.get_drift_baselines("fraud-detector", model_version=3)
assert len(baselines) >= len(REQUIRED_FEATURES), \
    "Missing drift baselines — refuse to promote"

Wire this into your CI step that runs before models.transition(..., "Production").

Workflow 3 — PSI thresholds

Default per endpoint is 0.2. Convention:

  • < 0.1 — no drift
  • 0.1 – 0.25 — warn
  • > 0.25 — critical

Tune per-feature criticality:

  • Identity features (e.g. country_code) → 0.1 (very tight)
  • Continuous behavioral features → 0.25 (looser)

Right now the threshold is per-endpoint, not per-feature — extending this is a coming-round task.

Workflow 4 — Eval gating

Add a CI step that fails if the eval run doesn't pass:

from geneflow import prompts

ev = prompts.run_eval(eval_set_id, prompt_version_id, judge_method="llm_as_judge")
# Poll
import time
while True:
    r = prompts.get_eval_run(ev["id"])
    if r["status"] in ("succeeded", "failed"): break
    time.sleep(2)

if r["pass_rate"] < 0.9:
    raise SystemExit("Eval pass rate too low — block promotion")

Workflow 5 — Periodic baseline refresh

Drift baselines decay — the data may legitimately shift over months. Plan to recompute:

CadenceTrigger
MonthlySchedule a Project run that recomputes from the last 30d of inferences
On critical drift alertRecompute is a candidate response — if the new dist is legitimate
On model retrainAlways recompute against new training data

Common gotchas

  • Sample size matters — drift check needs at least 30 samples in the window or it's skipped.
  • Numeric features only for PSI — categorical drift uses a separate metric (Jensen-Shannon, coming).
  • Baseline edges must match scoring-time edges — same np.histogram bins. If you bin differently at score time, PSI is nonsense.

Where to go next