Workbook

MLOps Engineer

You own: the production reliability of GeneFlow — runner availability, serving SLOs, drift triage, on-call.

What's new for you

WasNow in GeneFlow
Hand-managed K8s manifests per modeljobs-service applyServingDeploy() reconciles gf_endpoints
Custom drift detection job per teamBuilt-in PSI check + alert table
Logs spread across servicesPrometheus + Grafana + OTel + audit log all wired
No standard runner imageregistry.genedata.io/geneflow-runner:<sha> from Helm

Day-1 setup

helm upgrade --install genedata ./infrastructure/k8s/helm/genedata \
  --set global.agentRuntime.enabled=true \
  --set global.modelServing.enabled=true \
  --set observability.prometheusRules.enabled=true \
  --set observability.grafanaDashboards.enabled=true

Smoke test:

kubectl get sa geneflow-runner -n genedata
kubectl get role,rolebinding,networkpolicy -n genedata | grep geneflow

Workflow 1 — Watch production health

Open Grafana → folder GeneFlow:

  • WebSocket Gateway — collab editor connections, snapshot p95
  • (coming next round) Endpoints — Drift + QPS — per-endpoint QPS, p95, drift firings

Or via Prometheus:

sum by (endpoint) (rate(geneflow_inference_logs_total{status_code=~"5.."}[5m]))
  /
sum by (endpoint) (rate(geneflow_inference_logs_total[5m]))

Workflow 2 — Investigate a drift alert

  1. Page fires from rule GeneFlowDriftCritical (see observability.md)
  2. Open /dashboard/geneflow/ml/endpoints — banner shows the open alerts with feature name + PSI
  3. Click into the endpoint to see live metrics and version history
  4. Decide:
    • Was the input dist actually shifted? → coordinate with DE/DS on a retrain
    • False positive? → click "Acknowledge" or gfctl drift ack <id>
  5. If retraining: open a run, register a new version, and serving.update_endpoint(name, model_version=N+1)

Workflow 3 — Roll back a serving deploy

gfctl endpoints revisions fraud-prod
# Pick a previous READY revision's model_version, e.g. 2

gfctl endpoints update fraud-prod --version 2 --reason "rollback: 99th-pct latency regression"

Endpoint URL stays the same; the K8s Deployment rolls forward (yes, "rollback" is a forward roll to an earlier image).

Workflow 4 — Quota + cost triage

-- Top endpoints by cost in the last 24h
SELECT name, total_cost_usd, hourly_cost_usd, replicas
FROM gf_endpoints
WHERE tenant_id='ACME' AND status='READY'
ORDER BY total_cost_usd DESC LIMIT 20;

-- Endpoints with high QPS — possible scale-up candidates
SELECT endpoint_id, COUNT(*) FILTER (WHERE ts >= NOW() - INTERVAL '1 hour') AS qph
FROM gf_inference_logs WHERE tenant_id='ACME'
GROUP BY endpoint_id ORDER BY qph DESC LIMIT 20;

Workflow 5 — Capacity tuning

# Scale up
gfctl endpoints update fraud-prod --min-replicas 5 --max-replicas 20

# Change instance type (requires rebuild of image)
gfctl endpoints update fraud-prod --instance gpu-t4 --reason "GPU inference for v5"

The HPA scales between min and max on 70% CPU util by default; if your model is latency- not CPU-bound, set up a custom-metric HPA via Prometheus adapter.

Workflow 6 — Audit a stage transition

SELECT ts, payload->>'action', payload->>'from', payload->>'to', actor
FROM genedata_audit_log
WHERE tenant_id='ACME' AND action='geneflow.stage.transition'
  AND payload->>'name' = 'fraud-detector'
ORDER BY ts DESC LIMIT 20;

Verify chain integrity:

SELECT * FROM verify_audit_chain('ACME');

On-call cheat sheet

SymptomFirst check
Endpoint 5xx spikekubectl logs deploy/gf-ep-<id> -n <ns>
Endpoint stuck CREATINGgf_endpoint_revisions latest row .status + .reason
Drift alerts floodingRecompute baseline (data dist may have shifted legitimately)
WS gateway disconnectsGrafana WebSocket Gateway dashboard, snapshot failure panel
Project run failuresgfctl runs show <run_id> + kubectl logs job/<name> -n <ns>
Tracking writes failinggeneflow-service /metrics → db_pool_in_use_connections

Where to go next