Performance Optimization Engineer
You own: making GeneFlow workloads fast and cheap — query plans, k8s right-sizing, model serving latency, inference cost.
What's new for you
| Was | Now in GeneFlow |
|---|---|
| Latency was a guess | p50/p95/p99 per endpoint from gf_inference_logs |
| Cost per call was unknown | cost_usd on every inference row |
| No baseline to optimize against | Drift baseline = "what does normal look like?" |
Day-1 setup
pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
Workflow 1 — Diagnose a latency regression
# 1. Spot the spike on the dashboard:
# /dashboard/geneflow/ml/endpoints — p95 column
# 2. Compare versions
gfctl endpoints revisions fraud-prod
# 3. Drill into the spike window
psql -c "
SELECT model_version,
percentile_cont(0.95) WITHIN GROUP (ORDER BY latency_ms) AS p95,
COUNT(*)
FROM gf_inference_logs WHERE endpoint_id='ep_…'
AND ts BETWEEN '2026-05-12 10:00' AND '2026-05-12 11:00'
GROUP BY model_version;
"
If a specific version regressed: rollback (gfctl endpoints update ... --version <prev>).
Workflow 2 — Right-size serving pods
-- Average + p95 CPU across hour of day
SELECT date_trunc('hour', ts) AS hour,
AVG(latency_ms),
PERCENTILE_CONT(0.95) WITHIN GROUP (ORDER BY latency_ms) AS p95
FROM gf_inference_logs WHERE endpoint_id='ep_…'
GROUP BY 1 ORDER BY 1 DESC LIMIT 24;
If p95 latency is well below SLO target and pod CPU is low, drop one instance class size — saves ~50%.
Workflow 3 — pgvector query tuning
EXPLAIN ANALYZE
SELECT name, content, 1 - (embedding <=> $1::vector) AS sim
FROM gf_prompt_versions WHERE tenant_id='ACME'
ORDER BY embedding <=> $1::vector LIMIT 5;
If slow, check:
- Index exists:
\d gf_prompt_versionsshould showidx_…_embedding USING ivfflat ivfflat.probessetting — bump from 10 to 100 for accuracy, down to 1 for speed
Workflow 4 — Training run cost
-- Most expensive runs in the last week
SELECT id, name, cost_usd, tokens_used,
EXTRACT(epoch FROM (end_time - start_time))::INT AS duration_s
FROM gf_runs WHERE tenant_id='ACME' AND start_time >= now() - INTERVAL '7 days'
ORDER BY cost_usd DESC LIMIT 20;
Top candidates to optimize:
- Long-running runs that hit timeouts → break into stages
- High-token-count runs without checkpoints → add
geneflow.log_artifact("ckpt_<step>.pt")
Workflow 5 — Inference cost vs. latency Pareto
Plot per-endpoint (avg_cost, p95_latency) over the last 7 days. Endpoints that are both cheap and fast = winners; the others might benefit from quantization, distillation, or smaller models.
Workflow 6 — Snapshot pipeline tuning (WS gateway)
If editor users complain about save lag:
# Grafana → WebSocket Gateway → Snapshot Latency panel
# If p95 > 1s, look at:
kubectl logs deploy/websocket-gateway | grep "draft POST failed"
Usually it's the API side (geneflow-service) slow on the draft endpoint — check its /metrics.