Workbook

Performance Optimization Engineer

You own: making GeneFlow workloads fast and cheap — query plans, k8s right-sizing, model serving latency, inference cost.

What's new for you

WasNow in GeneFlow
Latency was a guessp50/p95/p99 per endpoint from gf_inference_logs
Cost per call was unknowncost_usd on every inference row
No baseline to optimize againstDrift baseline = "what does normal look like?"

Day-1 setup

pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io

Workflow 1 — Diagnose a latency regression

# 1. Spot the spike on the dashboard:
#    /dashboard/geneflow/ml/endpoints — p95 column
# 2. Compare versions
gfctl endpoints revisions fraud-prod

# 3. Drill into the spike window
psql -c "
  SELECT model_version,
         percentile_cont(0.95) WITHIN GROUP (ORDER BY latency_ms) AS p95,
         COUNT(*)
  FROM gf_inference_logs WHERE endpoint_id='ep_…'
    AND ts BETWEEN '2026-05-12 10:00' AND '2026-05-12 11:00'
  GROUP BY model_version;
"

If a specific version regressed: rollback (gfctl endpoints update ... --version <prev>).

Workflow 2 — Right-size serving pods

-- Average + p95 CPU across hour of day
SELECT date_trunc('hour', ts) AS hour,
       AVG(latency_ms),
       PERCENTILE_CONT(0.95) WITHIN GROUP (ORDER BY latency_ms) AS p95
FROM gf_inference_logs WHERE endpoint_id='ep_…'
GROUP BY 1 ORDER BY 1 DESC LIMIT 24;

If p95 latency is well below SLO target and pod CPU is low, drop one instance class size — saves ~50%.

Workflow 3 — pgvector query tuning

EXPLAIN ANALYZE
SELECT name, content, 1 - (embedding <=> $1::vector) AS sim
FROM gf_prompt_versions WHERE tenant_id='ACME'
ORDER BY embedding <=> $1::vector LIMIT 5;

If slow, check:

  • Index exists: \d gf_prompt_versions should show idx_…_embedding USING ivfflat
  • ivfflat.probes setting — bump from 10 to 100 for accuracy, down to 1 for speed

Workflow 4 — Training run cost

-- Most expensive runs in the last week
SELECT id, name, cost_usd, tokens_used,
       EXTRACT(epoch FROM (end_time - start_time))::INT AS duration_s
FROM gf_runs WHERE tenant_id='ACME' AND start_time >= now() - INTERVAL '7 days'
ORDER BY cost_usd DESC LIMIT 20;

Top candidates to optimize:

  • Long-running runs that hit timeouts → break into stages
  • High-token-count runs without checkpoints → add geneflow.log_artifact("ckpt_<step>.pt")

Workflow 5 — Inference cost vs. latency Pareto

Plot per-endpoint (avg_cost, p95_latency) over the last 7 days. Endpoints that are both cheap and fast = winners; the others might benefit from quantization, distillation, or smaller models.

Workflow 6 — Snapshot pipeline tuning (WS gateway)

If editor users complain about save lag:

# Grafana → WebSocket Gateway → Snapshot Latency panel
# If p95 > 1s, look at:
kubectl logs deploy/websocket-gateway | grep "draft POST failed"

Usually it's the API side (geneflow-service) slow on the draft endpoint — check its /metrics.

Where to go next