Site Reliability Engineer
You own: GeneFlow's SLOs, on-call rotations, incident response, capacity, and the runbook quality.
SLOs
| Service | SLO | Window | Burn budget |
|---|---|---|---|
| GeneFlow tracking writes | 99.9% availability | 30d | ~43 min |
| Endpoint serving 5xx | < 1% | 1h | 0.6 min/hr |
| WS gateway snapshot success | 99% | 30d | |
| Project run completion | 99% | 24h | |
| Drift baseline write success | 99.5% | 30d |
Alert pages
Configured in infrastructure/k8s/helm/genedata/templates/observability/prometheus-rules.yaml. Add new ones for serving in this round:
- alert: GeneFlowEndpointHighErrorRate
expr: |
sum by (endpoint) (rate(geneflow_inference_logs_total{status_code=~"5.."}[5m]))
/
sum by (endpoint) (rate(geneflow_inference_logs_total[5m])) > 0.01
for: 5m
- alert: GeneFlowEndpointStuckUpdating
expr: |
count by (endpoint) (geneflow_endpoint_status{status="UPDATING"}) > 0 unless on (endpoint)
count by (endpoint) (changes(geneflow_endpoint_status_change_total[10m])) > 0
for: 15m
- alert: GeneFlowDriftCritical
expr: increase(geneflow_drift_alerts_fired_total{severity="critical"}[15m]) > 0
for: 0m
On-call cheat sheet
| Symptom | First check | Quick mitigation |
|---|---|---|
| Endpoint 5xx spike | kubectl logs deploy/gf-ep-<id> | Rollback: gfctl endpoints update <name> --version <prev> |
| Endpoint stuck CREATING | gf_endpoint_revisions .reason | Delete + recreate or fix image |
| Drift alerts flooding | /dashboard/geneflow/ml/endpoints banner | Page MLE — usually a real shift |
| WS gateway slow | Grafana WebSocket Gateway → snapshot p95 | Scale StatefulSet replicas |
| Project runs failing | kubectl get jobs -n genedata + logs | Often image-pull or PAT-expired |
| pg_stat_activity blocked | DB pool exhaustion from geneflow-service | Restart pod (graceful) |
| Audit chain broken | SELECT * FROM verify_audit_chain('TENANT') | Page Steward + DGL immediately — security incident |
Incident response
Follow docs/INCIDENT-RESPONSE.md. Severities:
- SEV1: Audit chain broken; multi-tenant data leak; serving global outage
- SEV2: Single-tenant serving outage; tracking outage; drift not firing
- SEV3: Single endpoint unhealthy; gateway connection thrash
- SEV4: Degraded dashboards; missing metrics
Capacity reviews (monthly)
Run docs/CAPACITY-BASELINE.md checklist. New items for GeneFlow:
gf_inference_logstable size + retention policy- Active endpoints per tenant vs. quota
- WS gateway concurrent connections trend
- pgvector
prompt_embeddingindex size (rebuild when > 2× sample data)
Runbooks
- docs/runbooks/00-postmortem-template.md
- docs/runbooks/01-disaster-recovery.md
- docs/runbooks/02-error-rate-spike.md
- docs/runbooks/03-postgres-saturation.md