Data Quality Engineer
You own: ensuring data + model quality contracts hold. GeneFlow's drift baselines + eval sets are core to your toolkit.
What's new for you
| Was | Now in GeneFlow |
|---|---|
| Bespoke drift scripts per model | serving.save_drift_baselines() + automatic PSI checks |
| Eval results in spreadsheets | gf_eval_runs rows w/ per-example results |
| No record of "passed gate at deploy time" | Drift baseline + eval run tied to model version |
Day-1 setup
pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
Workflow 1 — Define a drift baseline at training time
import numpy as np
from geneflow import serving
# For each numeric feature
features = []
for col in NUMERIC_FEATURES:
hist, edges = np.histogram(df[col].dropna(), bins=10)
features.append({
"feature_name": col,
"feature_type": "numeric",
"histogram": {"buckets": hist.tolist(), "edges": edges.tolist()},
"mean": float(df[col].mean()),
"stddev": float(df[col].std()),
"min": float(df[col].min()),
"max": float(df[col].max()),
"sample_size": len(df),
})
# Categorical:
for col in CATEGORICAL_FEATURES:
counts = df[col].value_counts().to_dict()
features.append({
"feature_name": col,
"feature_type": "categorical",
"histogram": {"cats": counts},
"sample_size": len(df),
})
serving.save_drift_baselines(
model_name="fraud-detector",
model_version=3,
features=features,
)
Workflow 2 — Confirm a baseline exists before promoting
baselines = serving.get_drift_baselines("fraud-detector", model_version=3)
assert len(baselines) >= len(REQUIRED_FEATURES), \
"Missing drift baselines — refuse to promote"
Wire this into your CI step that runs before models.transition(..., "Production").
Workflow 3 — PSI thresholds
Default per endpoint is 0.2. Convention:
- < 0.1 — no drift
- 0.1 – 0.25 — warn
- > 0.25 — critical
Tune per-feature criticality:
- Identity features (e.g.
country_code) → 0.1 (very tight) - Continuous behavioral features → 0.25 (looser)
Right now the threshold is per-endpoint, not per-feature — extending this is a coming-round task.
Workflow 4 — Eval gating
Add a CI step that fails if the eval run doesn't pass:
from geneflow import prompts
ev = prompts.run_eval(eval_set_id, prompt_version_id, judge_method="llm_as_judge")
# Poll
import time
while True:
r = prompts.get_eval_run(ev["id"])
if r["status"] in ("succeeded", "failed"): break
time.sleep(2)
if r["pass_rate"] < 0.9:
raise SystemExit("Eval pass rate too low — block promotion")
Workflow 5 — Periodic baseline refresh
Drift baselines decay — the data may legitimately shift over months. Plan to recompute:
| Cadence | Trigger |
|---|---|
| Monthly | Schedule a Project run that recomputes from the last 30d of inferences |
| On critical drift alert | Recompute is a candidate response — if the new dist is legitimate |
| On model retrain | Always recompute against new training data |
Common gotchas
- Sample size matters — drift check needs at least 30 samples in the window or it's skipped.
- Numeric features only for PSI — categorical drift uses a separate metric (Jensen-Shannon, coming).
- Baseline edges must match scoring-time edges — same
np.histogrambins. If you bin differently at score time, PSI is nonsense.
Where to go next
- 01-ml-engineer.md — handoff
- 07-data-engineer.md — feature group contracts
- api-reference.md#model-serving