Workbook

Site Reliability Engineer

You own: GeneFlow's SLOs, on-call rotations, incident response, capacity, and the runbook quality.

SLOs

ServiceSLOWindowBurn budget
GeneFlow tracking writes99.9% availability30d~43 min
Endpoint serving 5xx< 1%1h0.6 min/hr
WS gateway snapshot success99%30d
Project run completion99%24h
Drift baseline write success99.5%30d

Alert pages

Configured in infrastructure/k8s/helm/genedata/templates/observability/prometheus-rules.yaml. Add new ones for serving in this round:

- alert: GeneFlowEndpointHighErrorRate
  expr: |
    sum by (endpoint) (rate(geneflow_inference_logs_total{status_code=~"5.."}[5m]))
      /
    sum by (endpoint) (rate(geneflow_inference_logs_total[5m])) > 0.01
  for: 5m

- alert: GeneFlowEndpointStuckUpdating
  expr: |
    count by (endpoint) (geneflow_endpoint_status{status="UPDATING"}) > 0 unless on (endpoint)
    count by (endpoint) (changes(geneflow_endpoint_status_change_total[10m])) > 0
  for: 15m

- alert: GeneFlowDriftCritical
  expr: increase(geneflow_drift_alerts_fired_total{severity="critical"}[15m]) > 0
  for: 0m

On-call cheat sheet

SymptomFirst checkQuick mitigation
Endpoint 5xx spikekubectl logs deploy/gf-ep-<id>Rollback: gfctl endpoints update <name> --version <prev>
Endpoint stuck CREATINGgf_endpoint_revisions .reasonDelete + recreate or fix image
Drift alerts flooding/dashboard/geneflow/ml/endpoints bannerPage MLE — usually a real shift
WS gateway slowGrafana WebSocket Gateway → snapshot p95Scale StatefulSet replicas
Project runs failingkubectl get jobs -n genedata + logsOften image-pull or PAT-expired
pg_stat_activity blockedDB pool exhaustion from geneflow-serviceRestart pod (graceful)
Audit chain brokenSELECT * FROM verify_audit_chain('TENANT')Page Steward + DGL immediately — security incident

Incident response

Follow docs/INCIDENT-RESPONSE.md. Severities:

  • SEV1: Audit chain broken; multi-tenant data leak; serving global outage
  • SEV2: Single-tenant serving outage; tracking outage; drift not firing
  • SEV3: Single endpoint unhealthy; gateway connection thrash
  • SEV4: Degraded dashboards; missing metrics

Capacity reviews (monthly)

Run docs/CAPACITY-BASELINE.md checklist. New items for GeneFlow:

  • gf_inference_logs table size + retention policy
  • Active endpoints per tenant vs. quota
  • WS gateway concurrent connections trend
  • pgvector prompt_embedding index size (rebuild when > 2× sample data)

Runbooks

Where to go next