Workbook

AI Engineer

You own: building GeneAI-powered features — prompts, evals, embeddings, RAG, agent tools. You bridge product and the GeneFlow LLM lifecycle.

What's new for you

WasNow in GeneFlow
Prompts in git or NotionFirst-class prompt registry with versions + stages
One-off eval scriptsprompts.create_eval_set() + run_eval() with LLM-as-judge
No similarity search across promptsprompts.search_similar() (pgvector)
Engineers editing prompts in PRsReal-time collab editor at /dashboard/geneflow/ml/prompts/[name]/edit
Token usage in OpenAI dashboards onlycost_usd + tokens_used per run, rolled to tenant

Day-1 setup

pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
export GENEDATA_PAT=$(genedata auth token)
export GENEFLOW_TENANT_ID=ACME

Workflow 1 — Register your first prompt

import geneflow
from geneflow import prompts

p = prompts.register(
    name="customer_support",
    content="You are a {{tone}} support agent. Question: {{question}}",
    description="V1 — baseline tone-templated agent",
)
print(p.version)       # 1
print(p.variables)     # {'tone', 'question'}

The variable scan is server-side ({{var}} regex). Variables are saved on the version row.

Workflow 2 — Iterate in the collab editor

Open /dashboard/geneflow/ml/prompts/customer_support/edit — anyone with access to the prompt gets a live cursor + presence avatar. "Save" creates a new immutable version. Drafts auto-save every 30s.

Behind the scenes the editor runs Y.js CRDT over the websocket gateway (:4033). Multiple AIE on the same prompt is fine — last-writer-wins is not how it works; updates merge.

Workflow 3 — Eval set + LLM-as-judge

es = prompts.create_eval_set(
    name="support-baseline",
    examples=[
        {"input": {"tone": "polite", "question": "Refund please"},
         "expected_output": "I understand"},
        {"input": {"tone": "friendly", "question": "How do I login?"},
         "expected_output": "Click 'sign in'"},
    ],
)

# Run eval on a specific prompt version
ev = prompts.run_eval(
    eval_set_id=es["id"],
    prompt_version_id=p.id,
    judge_method="llm_as_judge",       # or "exact" | "bleu" | "custom"
)
print(prompts.get_eval_run(ev["id"]))  # poll until status='succeeded'

Workflow 4 — Search the prompt registry semantically

hits = prompts.search_similar("how do I respond to angry users", k=5)
for h in hits:
    print(h["name"], h["version"], h["similarity"])

Embeddings are pgvector-backed. Useful when starting a new product feature: don't reinvent.

Workflow 5 — Stage a prompt to Production

prompts.transition("customer_support", version=2, to_stage="Staging", reason="passed eval")
prompts.transition("customer_support", version=2, to_stage="Production", reason="A/B winner")

Production transitions write to the hash-chained audit log. If the tenant policy requires approval, you'll get an APPROVAL_REQUIRED error — coordinate with DGL / Steward.

Workflow 6 — Wire a prompt into your agent code

from geneflow import prompts

# Fetch the Production version
v = prompts.get_latest("customer_support", stage="Production")
template = v["content"]
filled = template.replace("{{tone}}", "friendly").replace("{{question}}", user_question)

# Call your LLM
response = call_claude(filled)

# Log usage back into the run
geneflow.log_metric("cost_usd", response["usage"]["cost_usd"])
geneflow.log_metric("tokens_used", response["usage"]["total_tokens"])

Better: use the geneflow.agents helper (coming) that auto-tags spans + logs cost.

Common gotchas

  • Variable mismatch on register? GeneFlow auto-extracts {{vars}} but if your template uses Jinja {% %} blocks, those don't count as variables — strip them or move logic to client code.
  • Eval run_eval returns immediately — it's async. Poll get_eval_run(id).
  • Collab editor not loading? Check that WEBSOCKET_GATEWAY_URL is set in the platform env, and the gateway pod is healthy.
  • Search returns nothing? Embeddings are computed at register time — older prompts created before pgvector landed may need to be re-registered or backfilled.

Where to go next