Observability
Metrics, traces, and alerts for pipelines, runs, and serving endpoints.
Signals at a glance
| Signal | Where | Default sample |
|---|---|---|
Prometheus /metrics | every service | 15s scrape via ServiceMonitor |
| Grafana — WebSocket Gateway | infrastructure/observability/grafana/dashboards/websocket-gateway.json | auto-loaded via Helm sidecar |
| OpenTelemetry traces | OTEL_EXPORTER_OTLP_ENDPOINT env | per-request |
| Audit log | genedata_audit_log table | hash-chained, immutable |
| Logs | stdout (JSON) | shipped to whatever your stack uses |
Prometheus metrics
geneflow-service
| Metric | Type | Description |
|---|---|---|
geneflow_runs_started_total{tenant} | counter | Run creations |
geneflow_runs_finished_total{tenant,status} | counter | Run terminations |
geneflow_run_cost_usd_sum{tenant} | counter | Cumulative $ from runs |
geneflow_model_versions_created_total{tenant} | counter | New versions |
geneflow_stage_transitions_total{tenant,from,to} | counter | Stage changes |
geneflow_prompt_versions_total{tenant} | counter | Prompt registrations |
geneflow_inference_logs_total{endpoint} | counter | Sidecar log POSTs received |
geneflow_drift_alerts_fired_total{severity} | counter | New drift alerts |
websocket-gateway (this round added the bottom 8 metrics)
| Metric | Type |
|---|---|
geneflow_ws_rooms | gauge |
geneflow_ws_connections | gauge |
geneflow_ws_connects_total | counter |
geneflow_ws_disconnects_total | counter |
geneflow_ws_auth_failures_total | counter |
geneflow_ws_doc_updates_total | counter |
geneflow_ws_awareness_updates_total | counter |
geneflow_ws_snapshot_success_total | counter |
geneflow_ws_snapshot_failure_total | counter |
geneflow_ws_bytes_sent_total | counter |
geneflow_ws_bytes_received_total | counter |
geneflow_ws_snapshot_seconds | histogram (10 buckets, 10ms–10s) |
jobs-service
| Metric | Type | Description |
|---|---|---|
geneflow_project_runs_total{mode,status} | counter | LOCAL / DOCKER / K8S_JOB |
geneflow_project_run_duration_seconds | histogram | Wall-clock per run |
geneflow_serving_deploys_total{op,status} | counter | apply / delete |
geneflow_serving_deploy_duration_seconds | histogram | k8s apply latency |
Grafana dashboards
| Dashboard | File | Folder |
|---|---|---|
| WebSocket Gateway | infrastructure/observability/grafana/dashboards/websocket-gateway.json | GeneFlow |
Auto-provisioning is wired via templates/observability/grafana-dashboards.yaml — the kube-prometheus-stack sidecar (grafana.sidecar.dashboards.enabled=true) watches for ConfigMaps with grafana_dashboard: "1" and loads them.
To add a new dashboard:
- Drop JSON into
infrastructure/observability/grafana/dashboards/ - Add a
dataentry intemplates/observability/grafana-dashboards.yaml helm upgrade— it appears in the GeneFlow folder
Recommended alerts
# Prometheus rules — add to prometheus-rules.yaml
- alert: GeneFlowEndpointHighErrorRate
expr: |
sum by (endpoint) (rate(geneflow_inference_logs_total{status_code=~"5.."}[5m]))
/
sum by (endpoint) (rate(geneflow_inference_logs_total[5m])) > 0.01
for: 5m
labels: { severity: page }
annotations:
summary: "Endpoint {{ $labels.endpoint }} > 1% 5xx for 5m"
- alert: GeneFlowDriftCritical
expr: increase(geneflow_drift_alerts_fired_total{severity="critical"}[15m]) > 0
for: 0m
labels: { severity: page }
- alert: GeneFlowWSGatewaySnapshotFailures
expr: |
sum(rate(geneflow_ws_snapshot_failure_total[5m]))
/
clamp_min(sum(rate(geneflow_ws_snapshot_success_total[5m])) +
sum(rate(geneflow_ws_snapshot_failure_total[5m])), 0.001) > 0.05
for: 5m
labels: { severity: ticket }
- alert: GeneFlowProjectRunsBacking
expr: sum(rate(geneflow_project_runs_total{status="failed"}[10m])) > 0.5
for: 10m
labels: { severity: ticket }
SLOs (suggested)
| Service | SLO | Burn-rate budget |
|---|---|---|
| Tracking write availability | 99.9% over 30d | ~43 min |
| Endpoint serving 5xx | < 1% over 1h windows | 0.6 min/hr |
| Drift baseline write success | 99.5% over 30d | |
| WS gateway snapshot success | 99% over 30d |
OpenTelemetry traces
Each service is instrumented with span names that follow the convention <service>.<entity>.<action>:
geneflow.run.creategeneflow.endpoint.deploygeneflow.drift.checkws.connection.lifecyclews.snapshot.post
Set OTEL_EXPORTER_OTLP_ENDPOINT to ship to Tempo / Jaeger. Enable exemplar storage on Prometheus (--enable-feature=exemplar-storage) to get trace links from the latency panels.
Audit log
Every state change (run lifecycle, stage transition, endpoint deploy, prompt transition) appends a row to genedata_audit_log:
SELECT ts, action, resource_type, resource_id, payload
FROM genedata_audit_log
WHERE tenant_id = 'ACME' AND action LIKE 'geneflow.%'
ORDER BY ts DESC LIMIT 50;
The row is prev_hash + row_hash chained — tampering breaks the chain. Verify:
SELECT * FROM verify_audit_chain('ACME');
Logs
JSON to stdout, parsed by your stack:
{
"level": "info",
"msg": "endpoint.deploy.apply",
"ts": "2026-05-12T18:42:11Z",
"tenant_id": "ACME",
"endpoint_id": "ep_abc...",
"model": "fraud-detector",
"version": 3,
"duration_ms": 4218
}
Trace IDs propagated via traceparent header on inbound + outbound HTTP.