Workbook

Solutions Architect

You own: customer-facing design — propose, scope, and sequence GeneFlow rollouts inside customer environments. You map their existing stack to ours.

What's new for you

  • Round 6 closed the lifecycle: train → register → serve → monitor. You can now sell GeneFlow as a complete MLflow + SageMaker + LangChain + LangSmith replacement, not just tracking.
  • The MLflow importer (Round 5) is the conversation-opener for "we already have MLflow."
  • The K8s deployer is portable across EKS / GKE / AKS — no cloud lock-in.

Discovery questions

  1. Existing ML stack: MLflow OSS / hosted MLflow / SageMaker / homegrown / nothing?
  2. Existing GeneAI stack: LangChain / LangSmith / Helicone / no observability?
  3. Cloud + K8s: AWS+EKS / GCP+GKE / Azure+AKS / on-prem?
  4. Compliance: SOC2 / HIPAA / GDPR / FedRAMP?
  5. Tenants: single-team, multi-team-in-one-org, or true multi-tenant SaaS?
  6. Volumes: runs/day, inferences/sec, prompts deployed?
  7. Skill profile: who'll use GeneFlow on their side? Map to our role workbooks

Stack mapping cheat sheet

Their toolGeneFlow equivalent
MLflow OSS / hosted MLflowValidate tracking API and artifact access, then test the importer
SageMaker TrainingMLflow Projects via K8S_JOB
SageMaker Model Registrygf_registered_models
SageMaker Endpointsgf_endpoints
LangSmith / HeliconeTracking (cost_usd + tokens_used) + audit
LangChain prompt strings in codePrompt registry + versions + stages
WhyLabs / Arize for driftBuilt-in PSI w/ baselines
dbt for metric definitionsKeep dbt — wire GeneFlow gf_* as a source
AirByte / Fivetran into warehouseSame; we publish gf_* via CDC

Reference architectures

Pattern A — Single-cloud customer (AWS)

  • EKS cluster with our Helm chart
  • IRSA-enabled geneflow-runner SA → IAM role for S3 artifact access
  • Aurora Postgres + pgvector
  • S3 artifact bucket per tenant with KMS CMK
  • External Secrets Operator → AWS Secrets Manager
  • Their CI: Bitbucket / GHA — points at GeneFlow on every commit

Pattern B — Multi-region (regulatory)

  • One GeneFlow stack per residency region
  • Cross-region audit-log replication (read-only)
  • Tenant data_residency_region enforced at the proxy

Pattern C — Hybrid (training on-prem, serving in cloud)

  • Project runner on-prem K8s
  • Serving on EKS via separate Helm install
  • Artifact bucket in the cloud, mirrored to on-prem for hot training data

Sequencing a deal

QuarterDeliverable
Q1Helm install + tracking migration via importer. DS/MLE point clients at GeneFlow.
Q2Prompt registry + collab editor for AIE team. Eval sets for top 5 prompts.
Q3Model serving for 1–2 production models. Drift baselines.
Q4Full lifecycle: project runners, all prompts in registry, all models served, drift alerts on-call.

Common gotchas

  • Don't underestimate the prompt-registry sell — most customers don't realize they NEED one until you show them how often prompts change silently.
  • Don't sell drift as "AI Observability" — it's PSI, not magic. Be honest about what it detects vs. doesn't.
  • MLflow drop-in is the foot-in-door — they can adopt without committing.

Where to go next