Workbook

GeneAI Developer

You own: the quality, governance, and lifecycle of every production prompt. You write, eval, and stage prompts to Production through GeneFlow.

What's new for you

WasNow in GeneFlow
Prompts in git diffsVersions + 4 stages + hash-chained audit
"Send the prompt change for review on Slack"Real-time collab editor at /dashboard/geneflow/ml/prompts/[name]/edit
No metrics for prompt changesEval sets + LLM-as-judge + cost-per-call tracked
Hard to find similar promptsprompts.search_similar() (pgvector)

Day-1 setup

pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
export GENEDATA_PAT=$(genedata auth token)

UI: /dashboard/geneflow/ml → Prompts tab.

Workflow 1 — Author a new prompt in the collab editor

  1. Open /dashboard/geneflow/ml/prompts/customer_support/edit (or click "New prompt" first)
  2. Editor opens with CodeMirror 6 + live cursors + presence avatars
  3. Type your template using {{var}} placeholders
  4. Click Save as new version — creates an immutable v2

Drafts auto-save every 30s. Multiple people can edit simultaneously; CRDT handles merging.

Workflow 2 — Build an eval set

eval.json (sample):

[
  {
    "input": { "tone": "polite",   "question": "Refund please" },
    "expected_output": "I understand"
  },
  {
    "input": { "tone": "friendly", "question": "How do I login?" },
    "expected_output": "Click 'sign in'"
  }
]
gfctl eval-sets create support-baseline --from eval.json

Workflow 3 — Run an eval

gfctl eval-runs start \
  --set support-baseline \
  --prompt customer_support --version 2 \
  --judge llm_as_judge

gfctl eval-runs show run_eval_xyz   # poll

judge_method:

  • exact — exact string match
  • bleu — BLEU score
  • llm_as_judge — Claude scores the candidate against expected
  • custom — call your scorer URL

Workflow 4 — Compare two prompt versions

gfctl prompts diff customer_support 1 2

Or in UI: open one version, hit Compare → pick another. You'll see content diff + per-example eval scores side-by-side.

Workflow 5 — Stage to Production

from geneflow import prompts
prompts.transition("customer_support", 2, to_stage="Staging", reason="passed eval @ 0.92")
# Wait for approval if tenant policy requires it
prompts.transition("customer_support", 2, to_stage="Production", reason="A/B winner")

The transition writes:

  • A row in gf_stage_transitions
  • An immutable audit row in genedata_audit_log with prev_hash chain
  • Optionally fires a webhook (if configured)

Workflow 6 — Find prior art

hits = prompts.search_similar(
    "guide a user through a 2FA reset",
    k=5,
)
for h in hits:
    print(f"{h['name']} v{h['version']} ({h['similarity']:.3f})")

Use this before authoring — there may be a vetted prompt already.

Workflow 7 — Track real-world cost / tokens

Wire your inference code:

geneflow.log_metric("prompt_cost_usd", response.usage.cost_usd)
geneflow.log_metric("prompt_tokens",    response.usage.total_tokens)
geneflow.set_tag("prompt_version", str(active_version_id))

Now in the UI you can answer: "what does Production v2 cost us per day?"

Governance checkpoints

StageWho can transitionAudit
None → StagingPrompt Engineeryes
Staging → Productionrequires approval (default)yes, with requested_by/approved_by
Production → ArchivedPrompt Engineer or DGLyes
Anything → Archivedanyone with geneflow:transitionyes

Common gotchas

  • Don't put secrets in prompts. Use {{SECRET_FOO}} and fill server-side.
  • {{var}} is the only template syntax — Jinja blocks are NOT parsed as variables, they're treated as literal content.
  • Editor lag? Use a smaller editor surface (split big system prompts into multiple registered prompts).

Where to go next