Reference

Observability

Metrics, traces, and alerts for pipelines, runs, and serving endpoints.

Signals at a glance

SignalWhereDefault sample
Prometheus /metricsevery service15s scrape via ServiceMonitor
Grafana — WebSocket Gatewayinfrastructure/observability/grafana/dashboards/websocket-gateway.jsonauto-loaded via Helm sidecar
OpenTelemetry tracesOTEL_EXPORTER_OTLP_ENDPOINT envper-request
Audit loggenedata_audit_log tablehash-chained, immutable
Logsstdout (JSON)shipped to whatever your stack uses

Prometheus metrics

geneflow-service

MetricTypeDescription
geneflow_runs_started_total{tenant}counterRun creations
geneflow_runs_finished_total{tenant,status}counterRun terminations
geneflow_run_cost_usd_sum{tenant}counterCumulative $ from runs
geneflow_model_versions_created_total{tenant}counterNew versions
geneflow_stage_transitions_total{tenant,from,to}counterStage changes
geneflow_prompt_versions_total{tenant}counterPrompt registrations
geneflow_inference_logs_total{endpoint}counterSidecar log POSTs received
geneflow_drift_alerts_fired_total{severity}counterNew drift alerts

websocket-gateway (this round added the bottom 8 metrics)

MetricType
geneflow_ws_roomsgauge
geneflow_ws_connectionsgauge
geneflow_ws_connects_totalcounter
geneflow_ws_disconnects_totalcounter
geneflow_ws_auth_failures_totalcounter
geneflow_ws_doc_updates_totalcounter
geneflow_ws_awareness_updates_totalcounter
geneflow_ws_snapshot_success_totalcounter
geneflow_ws_snapshot_failure_totalcounter
geneflow_ws_bytes_sent_totalcounter
geneflow_ws_bytes_received_totalcounter
geneflow_ws_snapshot_secondshistogram (10 buckets, 10ms–10s)

jobs-service

MetricTypeDescription
geneflow_project_runs_total{mode,status}counterLOCAL / DOCKER / K8S_JOB
geneflow_project_run_duration_secondshistogramWall-clock per run
geneflow_serving_deploys_total{op,status}counterapply / delete
geneflow_serving_deploy_duration_secondshistogramk8s apply latency

Grafana dashboards

DashboardFileFolder
WebSocket Gatewayinfrastructure/observability/grafana/dashboards/websocket-gateway.jsonGeneFlow

Auto-provisioning is wired via templates/observability/grafana-dashboards.yaml — the kube-prometheus-stack sidecar (grafana.sidecar.dashboards.enabled=true) watches for ConfigMaps with grafana_dashboard: "1" and loads them.

To add a new dashboard:

  1. Drop JSON into infrastructure/observability/grafana/dashboards/
  2. Add a data entry in templates/observability/grafana-dashboards.yaml
  3. helm upgrade — it appears in the GeneFlow folder
# Prometheus rules — add to prometheus-rules.yaml
- alert: GeneFlowEndpointHighErrorRate
  expr: |
    sum by (endpoint) (rate(geneflow_inference_logs_total{status_code=~"5.."}[5m]))
      /
    sum by (endpoint) (rate(geneflow_inference_logs_total[5m])) > 0.01
  for: 5m
  labels: { severity: page }
  annotations:
    summary: "Endpoint {{ $labels.endpoint }} > 1% 5xx for 5m"

- alert: GeneFlowDriftCritical
  expr: increase(geneflow_drift_alerts_fired_total{severity="critical"}[15m]) > 0
  for: 0m
  labels: { severity: page }

- alert: GeneFlowWSGatewaySnapshotFailures
  expr: |
    sum(rate(geneflow_ws_snapshot_failure_total[5m]))
      /
    clamp_min(sum(rate(geneflow_ws_snapshot_success_total[5m])) +
              sum(rate(geneflow_ws_snapshot_failure_total[5m])), 0.001) > 0.05
  for: 5m
  labels: { severity: ticket }

- alert: GeneFlowProjectRunsBacking
  expr: sum(rate(geneflow_project_runs_total{status="failed"}[10m])) > 0.5
  for: 10m
  labels: { severity: ticket }

SLOs (suggested)

ServiceSLOBurn-rate budget
Tracking write availability99.9% over 30d~43 min
Endpoint serving 5xx< 1% over 1h windows0.6 min/hr
Drift baseline write success99.5% over 30d
WS gateway snapshot success99% over 30d

OpenTelemetry traces

Each service is instrumented with span names that follow the convention <service>.<entity>.<action>:

  • geneflow.run.create
  • geneflow.endpoint.deploy
  • geneflow.drift.check
  • ws.connection.lifecycle
  • ws.snapshot.post

Set OTEL_EXPORTER_OTLP_ENDPOINT to ship to Tempo / Jaeger. Enable exemplar storage on Prometheus (--enable-feature=exemplar-storage) to get trace links from the latency panels.

Audit log

Every state change (run lifecycle, stage transition, endpoint deploy, prompt transition) appends a row to genedata_audit_log:

SELECT ts, action, resource_type, resource_id, payload
FROM genedata_audit_log
WHERE tenant_id = 'ACME' AND action LIKE 'geneflow.%'
ORDER BY ts DESC LIMIT 50;

The row is prev_hash + row_hash chained — tampering breaks the chain. Verify:

SELECT * FROM verify_audit_chain('ACME');

Logs

JSON to stdout, parsed by your stack:

{
  "level": "info",
  "msg": "endpoint.deploy.apply",
  "ts": "2026-05-12T18:42:11Z",
  "tenant_id": "ACME",
  "endpoint_id": "ep_abc...",
  "model": "fraud-detector",
  "version": 3,
  "duration_ms": 4218
}

Trace IDs propagated via traceparent header on inbound + outbound HTTP.