MLOps Engineer
You own: the production reliability of GeneFlow — runner availability, serving SLOs, drift triage, on-call.
What's new for you
| Was | Now in GeneFlow |
|---|---|
| Hand-managed K8s manifests per model | jobs-service applyServingDeploy() reconciles gf_endpoints |
| Custom drift detection job per team | Built-in PSI check + alert table |
| Logs spread across services | Prometheus + Grafana + OTel + audit log all wired |
| No standard runner image | registry.genedata.io/geneflow-runner:<sha> from Helm |
Day-1 setup
helm upgrade --install genedata ./infrastructure/k8s/helm/genedata \
--set global.agentRuntime.enabled=true \
--set global.modelServing.enabled=true \
--set observability.prometheusRules.enabled=true \
--set observability.grafanaDashboards.enabled=true
Smoke test:
kubectl get sa geneflow-runner -n genedata
kubectl get role,rolebinding,networkpolicy -n genedata | grep geneflow
Workflow 1 — Watch production health
Open Grafana → folder GeneFlow:
- WebSocket Gateway — collab editor connections, snapshot p95
- (coming next round) Endpoints — Drift + QPS — per-endpoint QPS, p95, drift firings
Or via Prometheus:
sum by (endpoint) (rate(geneflow_inference_logs_total{status_code=~"5.."}[5m]))
/
sum by (endpoint) (rate(geneflow_inference_logs_total[5m]))
Workflow 2 — Investigate a drift alert
- Page fires from rule
GeneFlowDriftCritical(see observability.md) - Open
/dashboard/geneflow/ml/endpoints— banner shows the open alerts with feature name + PSI - Click into the endpoint to see live metrics and version history
- Decide:
- Was the input dist actually shifted? → coordinate with DE/DS on a retrain
- False positive? → click "Acknowledge" or
gfctl drift ack <id>
- If retraining: open a run, register a new version, and
serving.update_endpoint(name, model_version=N+1)
Workflow 3 — Roll back a serving deploy
gfctl endpoints revisions fraud-prod
# Pick a previous READY revision's model_version, e.g. 2
gfctl endpoints update fraud-prod --version 2 --reason "rollback: 99th-pct latency regression"
Endpoint URL stays the same; the K8s Deployment rolls forward (yes, "rollback" is a forward roll to an earlier image).
Workflow 4 — Quota + cost triage
-- Top endpoints by cost in the last 24h
SELECT name, total_cost_usd, hourly_cost_usd, replicas
FROM gf_endpoints
WHERE tenant_id='ACME' AND status='READY'
ORDER BY total_cost_usd DESC LIMIT 20;
-- Endpoints with high QPS — possible scale-up candidates
SELECT endpoint_id, COUNT(*) FILTER (WHERE ts >= NOW() - INTERVAL '1 hour') AS qph
FROM gf_inference_logs WHERE tenant_id='ACME'
GROUP BY endpoint_id ORDER BY qph DESC LIMIT 20;
Workflow 5 — Capacity tuning
# Scale up
gfctl endpoints update fraud-prod --min-replicas 5 --max-replicas 20
# Change instance type (requires rebuild of image)
gfctl endpoints update fraud-prod --instance gpu-t4 --reason "GPU inference for v5"
The HPA scales between min and max on 70% CPU util by default; if your model is latency- not CPU-bound, set up a custom-metric HPA via Prometheus adapter.
Workflow 6 — Audit a stage transition
SELECT ts, payload->>'action', payload->>'from', payload->>'to', actor
FROM genedata_audit_log
WHERE tenant_id='ACME' AND action='geneflow.stage.transition'
AND payload->>'name' = 'fraud-detector'
ORDER BY ts DESC LIMIT 20;
Verify chain integrity:
SELECT * FROM verify_audit_chain('ACME');
On-call cheat sheet
| Symptom | First check |
|---|---|
| Endpoint 5xx spike | kubectl logs deploy/gf-ep-<id> -n <ns> |
| Endpoint stuck CREATING | gf_endpoint_revisions latest row .status + .reason |
| Drift alerts flooding | Recompute baseline (data dist may have shifted legitimately) |
| WS gateway disconnects | Grafana WebSocket Gateway dashboard, snapshot failure panel |
| Project run failures | gfctl runs show <run_id> + kubectl logs job/<name> -n <ns> |
| Tracking writes failing | geneflow-service /metrics → db_pool_in_use_connections |
Where to go next
- observability.md — SLOs, alerts, dashboards
- security.md — RBAC, NetworkPolicy
- docs/runbooks/ — wider platform runbooks
- 04-ai-platform-engineer.md — adjacent role