GeneFlow overview
What GeneFlow is, what's inside, and how it relates to MLflow.
GeneFlow is GeneData's ML & GeneAI lifecycle platform. It's MLflow-compatible at the wire level so existing mlflow.start_run() code works unchanged — and it adds the features MLflow doesn't have: a first-class prompt registry, real-time collaborative editing, native cost-per-run, hash-chained audit, model-serving with built-in drift detection, and a one-click migration path from existing MLflow installs.
What's inside the platform
| Capability | Doc | Status |
|---|---|---|
| Experiments, runs, metrics, params, tags | api-reference.md | ✅ |
| Artifact upload / download (S3 or local) | api-reference.md | ✅ |
| Model registry + approval gates | api-reference.md | ✅ |
| Model serving (Deployments + HPA + drift) | api-reference.md | ✅ |
| Prompt registry + semantic search | api-reference.md | ✅ |
| Eval sets + LLM-as-judge | api-reference.md | ✅ |
| Real-time collab prompt editor (Y.js / CRDT) | architecture.md | ✅ |
| MLflow Projects compat (LOCAL / DOCKER / K8S_JOB) | api-reference.md | ✅ |
| Lineage (feature group ↔ run ↔ model version) | api-reference.md | ✅ |
| MLflow → GeneFlow importer | migration-from-mlflow.md | ✅ |
Quickstart
pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
export GENEDATA_PAT=$(genedata auth token)
export GENEFLOW_TENANT_ID=ACME
import geneflow
geneflow.set_tracking_uri("https://api.genedata.io")
geneflow.set_tenant("ACME")
with geneflow.start_run(experiment_name="fraud_v3") as run:
geneflow.log_param("max_depth", 7)
geneflow.log_metric("auc", 0.91)
geneflow.log_metric("cost_usd", 2.34)
That's literally MLflow code with the tracking URI pointed at GeneFlow. Everything else is additive.
Where to read next
- Build & operate:
- architecture.md — services, schema, message flows
- api-reference.md — every REST endpoint
- python-sdk.md — the
geneflowpackage - cli-reference.md — the
gfctl/genedataCLI - observability.md — Grafana, Prometheus, alerts
- security.md — pod hardening, RBAC, network policies
- Migrate in:
- Role-specific how-to:
- workbooks/ — 26 role workbooks (MLE, MLOps, DS, AIE, …)
- Releases:
How GeneFlow fits with the rest of the platform
┌─────────────────────────────────────────────────┐
│ /dashboard/geneflow/ml — control plane UI │
│ • experiments • compare • prompts (collab) │
│ • endpoints • drift • import-from-mlflow│
└────────────────────────┬────────────────────────┘
│ Hono REST
┌────────────────────────▼────────────────────────┐
│ Monolith /api/2.0/mlflow/* + /api/v2.1/... │
│ (transparent proxy to extracted services) │
└─────┬──────────────────────────────────────┬────┘
│ │
┌─────────────▼──────────────┐ ┌────────────▼────────────┐
│ geneflow-service :4032 │ │ jobs-service │
│ (tracking, registry, │◄────────│ (K8s project runner + │
│ prompts, lineage, │ │ serving deployer) │
│ serving, drift, importer) │ └─────────────────────────┘
└─────────┬──────────────────┘
│
┌─────────▼──────────────────┐ ┌─────────────────────────┐
│ websocket-gateway :4033 │ │ PostgreSQL + pgvector │
│ (Y.js CRDT for prompts + │ │ (gf_* schema) │
│ notebook collab) │ └─────────────────────────┘
└────────────────────────────┘
Three positioning sentences
- vs. MLflow OSS — same wire protocol, plus prompts-as-first-class, cost-per-run, hash-chained audit, real-time collab, K8s-native runners, model serving with drift detection, and a one-click migration path.
- A unified lifecycle workspace — bring experiment tracking, the model registry, prompts, and serving into one GeneFlow dashboard, with self-hosted deployment options.
- vs. SageMaker — model lifecycle is in one place, prompts are first-class, costs roll up per tenant in real time, and the K8s deployer is portable across clouds.
See docs/marketing/GENEFLOW-vs-MLFLOW.md for the full comparison.