GeneAI Developer
You own: the quality, governance, and lifecycle of every production prompt. You write, eval, and stage prompts to Production through GeneFlow.
What's new for you
| Was | Now in GeneFlow |
|---|---|
| Prompts in git diffs | Versions + 4 stages + hash-chained audit |
| "Send the prompt change for review on Slack" | Real-time collab editor at /dashboard/geneflow/ml/prompts/[name]/edit |
| No metrics for prompt changes | Eval sets + LLM-as-judge + cost-per-call tracked |
| Hard to find similar prompts | prompts.search_similar() (pgvector) |
Day-1 setup
pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
export GENEDATA_PAT=$(genedata auth token)
UI: /dashboard/geneflow/ml → Prompts tab.
Workflow 1 — Author a new prompt in the collab editor
- Open
/dashboard/geneflow/ml/prompts/customer_support/edit(or click "New prompt" first) - Editor opens with CodeMirror 6 + live cursors + presence avatars
- Type your template using
{{var}}placeholders - Click Save as new version — creates an immutable v2
Drafts auto-save every 30s. Multiple people can edit simultaneously; CRDT handles merging.
Workflow 2 — Build an eval set
eval.json (sample):
[
{
"input": { "tone": "polite", "question": "Refund please" },
"expected_output": "I understand"
},
{
"input": { "tone": "friendly", "question": "How do I login?" },
"expected_output": "Click 'sign in'"
}
]
gfctl eval-sets create support-baseline --from eval.json
Workflow 3 — Run an eval
gfctl eval-runs start \
--set support-baseline \
--prompt customer_support --version 2 \
--judge llm_as_judge
gfctl eval-runs show run_eval_xyz # poll
judge_method:
- exact — exact string match
- bleu — BLEU score
- llm_as_judge — Claude scores the candidate against expected
- custom — call your scorer URL
Workflow 4 — Compare two prompt versions
gfctl prompts diff customer_support 1 2
Or in UI: open one version, hit Compare → pick another. You'll see content diff + per-example eval scores side-by-side.
Workflow 5 — Stage to Production
from geneflow import prompts
prompts.transition("customer_support", 2, to_stage="Staging", reason="passed eval @ 0.92")
# Wait for approval if tenant policy requires it
prompts.transition("customer_support", 2, to_stage="Production", reason="A/B winner")
The transition writes:
- A row in
gf_stage_transitions - An immutable audit row in
genedata_audit_logwithprev_hashchain - Optionally fires a webhook (if configured)
Workflow 6 — Find prior art
hits = prompts.search_similar(
"guide a user through a 2FA reset",
k=5,
)
for h in hits:
print(f"{h['name']} v{h['version']} ({h['similarity']:.3f})")
Use this before authoring — there may be a vetted prompt already.
Workflow 7 — Track real-world cost / tokens
Wire your inference code:
geneflow.log_metric("prompt_cost_usd", response.usage.cost_usd)
geneflow.log_metric("prompt_tokens", response.usage.total_tokens)
geneflow.set_tag("prompt_version", str(active_version_id))
Now in the UI you can answer: "what does Production v2 cost us per day?"
Governance checkpoints
| Stage | Who can transition | Audit |
|---|---|---|
| None → Staging | Prompt Engineer | yes |
| Staging → Production | requires approval (default) | yes, with requested_by/approved_by |
| Production → Archived | Prompt Engineer or DGL | yes |
| Anything → Archived | anyone with geneflow:transition | yes |
Common gotchas
- Don't put secrets in prompts. Use
{{SECRET_FOO}}and fill server-side. {{var}}is the only template syntax — Jinja blocks are NOT parsed as variables, they're treated as literal content.- Editor lag? Use a smaller editor surface (split big system prompts into multiple registered prompts).
Where to go next
- 03-ai-engineer.md — wiring prompts into apps
- api-reference.md#prompts — REST
- 20-data-governance-lead.md — for governance policy