Machine Learning Engineer
You own: model training, evaluation, registry hygiene, and shipping models to production with GeneFlow.
What's new for you
| Was | Now in GeneFlow |
|---|---|
mlflow.start_run() with no $ tracking | geneflow.start_run() — same shape, plus cost_usd per run |
| Hand-rolled docker for serving | serving.create_endpoint(...) → K8s Deployment + HPA + drift |
MLflow model-versions/transition with no audit | models.transition(...) writes hash-chained audit + approval gate |
| Manual drift detection scripts | serving.save_drift_baselines() + automatic PSI checks |
Day-1 setup
pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
export GENEDATA_PAT=$(genedata auth token)
export GENEFLOW_TENANT_ID=ACME
Confirm: gfctl experiments list should return without errors.
Workflow 1 — Train + log a run
import geneflow
geneflow.set_tracking_uri("https://api.genedata.io")
geneflow.set_tenant("ACME")
with geneflow.start_run(experiment_name="fraud_v3", run_name="xgb-tuned") as run:
geneflow.log_params({"max_depth": 7, "lr": 0.05, "subsample": 0.8})
for epoch in range(1, 11):
geneflow.log_metric("auc", 0.7 + epoch * 0.02, step=epoch)
geneflow.log_metric("log_loss", 0.5 - epoch * 0.03, step=epoch)
geneflow.log_metric("cost_usd", 4.31)
geneflow.set_tag("dataset_sha", "2026-05-12-train.parquet:sha256=...")
Verify in UI: /dashboard/geneflow/ml → experiments → click your run.
Workflow 2 — Register a model version
from geneflow import models, artifacts
artifacts.upload(run.run_id, "/tmp/model.pkl", dst_path="model/model.pkl")
mv = models.create_version(
name="fraud-detector",
source=f"runs:/{run.run_id}/model",
run_id=run.run_id,
description="XGBoost v3 — tuned",
tags={"framework": "xgboost", "language": "python"},
)
print(mv["model_version"]["version"]) # → 3
Workflow 3 — Compare candidates ($ + metrics)
diff = models.compare("fraud-detector", a=2, b=3)
# Or in the UI: /dashboard/geneflow/ml/compare?model=fraud-detector&a=2&b=3
diff has metricsDiff and costDiff — use these to decide which version to promote.
Workflow 4 — Save a drift baseline (do this BEFORE serving)
from geneflow import serving
import numpy as np
# Compute the histogram on your training data
hist, edges = np.histogram(X_train["transaction_amount"], bins=10)
serving.save_drift_baselines(
model_name="fraud-detector",
model_version=3,
features=[{
"feature_name": "transaction_amount",
"feature_type": "numeric",
"histogram": {"buckets": hist.tolist(), "edges": edges.tolist()},
"mean": float(X_train["transaction_amount"].mean()),
"stddev": float(X_train["transaction_amount"].std()),
"sample_size": len(X_train),
}],
)
Workflow 5 — Deploy + monitor
ep = serving.create_endpoint(
name="fraud-prod",
model_name="fraud-detector",
model_version=3,
instance_type="cpu-large",
min_replicas=2,
max_replicas=10,
drift_check_enabled=True,
drift_psi_threshold=0.2,
)
# Poll live metrics
m = serving.get_metrics(ep["id"], window_minutes=60)
print(f"QPS={m['qps']:.2f} p95={m['p95LatencyMs']}ms err={m['errorRatePct']:.2f}%")
# Run a drift check (also runs automatically per `drift_check_window`)
serving.check_drift(ep["id"], window_minutes=60)
alerts = serving.list_drift_alerts(ep["id"], only_open=True)
Workflow 6 — Ship a new version (rolling)
serving.update_endpoint("fraud-prod", model_version=4, reason="precision win on holdout")
# Endpoint URL stays the same; K8s does a rolling update with maxUnavailable=0
If a Production stage transition is required:
models.transition("fraud-detector", 4, to_stage="Production", reason="canary green")
# If approval is required, an audit row records the request; an approver must call
# the approve endpoint before the transition becomes active.
Common gotchas
- Cost not showing? You must call
geneflow.log_metric("cost_usd", ...)explicitly — runs don't infer it. - Drift alerts never fire? Did you save a baseline first? Check
gfctl drift baselines fraud-detector --version 3. - Endpoint stuck in CREATING? Look at
gfctl endpoints show fraud-prod— the most recent revision row has the failure reason. mlflowCLI still works — pointing it at GeneFlow works, but you lose the GeneFlow-native features (cost, drift). Prefergfctl.
Where to go next
- python-sdk.md — full SDK
- api-reference.md — REST shape
- 05-prompt-engineer.md — if you also work on LLM prompts
- 14-data-quality-engineer.md — for drift baseline curation