Migrating from MLflow
Point existing mlflow code at GeneFlow, then import the history.
GeneFlow is wire-compatible with MLflow, so you don't have to migrate to use it — you can flip MLFLOW_TRACKING_URI to your GeneFlow base URL and existing code keeps working. But to bring history with you (past experiments, runs, registered models), use the built-in TS-native importer.
TL;DR
gfctl import-mlflow \
--source https://mlflow.your-company.com \
--token "$OLD_MLFLOW_TOKEN" \
--prefix "mlflow:" \
--dry-run
# If the counts look right:
gfctl import-mlflow \
--source https://mlflow.your-company.com \
--token "$OLD_MLFLOW_TOKEN" \
--prefix "mlflow:"
Or use the UI: /dashboard/geneflow/ml/import — same engine, with live progress.
What gets imported
| MLflow entity | → GeneFlow target | Notes |
|---|---|---|
| Experiment | gf_experiments | Renamed with experiment_prefix if set |
| Run | gf_runs | mlflow.original_run_id stored as tag |
| Metric history | gf_metrics (full time-series) | Per-step values, not just final |
| Params | gf_params | Set-once |
| Tags | gf_tags | Mutable |
| Registered model | gf_registered_models | Idempotent on (tenant_id, name) |
| Model version | gf_model_versions | runs:/…/model source URI rewritten to new GeneFlow run |
| Artifacts | (skipped by default — large) | Pass --copy-artifacts to include |
What does NOT migrate
- Stage history before today — MLflow stages become
Noneon import; re-transition explicitly so the hash-chained audit starts clean. - Webhooks / event listeners — register them again in GeneFlow.
- Custom RBAC — set up via tenant settings in GeneFlow.
Pre-flight
- Confirm the source is reachable from the GeneFlow API process:
``bash curl -H "Authorization: Bearer $OLD_TOKEN" \ https://mlflow.your-company.com/api/2.0/mlflow/experiments/search?max_results=1 ``
- Check quota — the import will fan-out one
gf_runsinsert per run plus N metric inserts. Default tenant quota = 1k runs/hr; for very large MLflow installs, ask SRE to raise it temporarily. - Dry-run — count entities first. The phase will reach
doneinstantly without writes.
Running the import
Via UI
- Open
/dashboard/geneflow/ml/import - Fill in tracking URI + optional token
- Optional: scope by experiment names or registered models
- Set
experiment_prefix(defaultmlflow:) to avoid collisions - Click Start dry run → check counts
- Click Start import → live progress polls every 2s
Via SDK
from geneflow.import_mlflow import import_from
job = import_from(
source_uri="https://mlflow.your-company.com",
source_token=os.environ["OLD_MLFLOW_TOKEN"],
experiment_prefix="mlflow:",
dry_run=False,
)
# Poll
while job["phase"] not in ("done", "failed"):
time.sleep(2)
job = get_import_status(job["id"])
Via REST
curl -X POST https://api.genedata.io/api/v2.1/geneflow/import-mlflow \
-H "Authorization: Bearer $GENEDATA_PAT" \
-H "X-Tenant-Id: ACME" \
-H "Content-Type: application/json" \
-d '{
"source_uri": "https://mlflow.your-company.com",
"source_token": "...",
"experiment_prefix": "mlflow:",
"dry_run": false
}'
Phases
The importer steps through six phases — each is visible in the dashboard:
- queued → just created
- connecting → probing source
/api/2.0/mlflow/experiments/search - experiments → creating GeneFlow experiments (idempotent on name)
- runs → for each experiment, fetch runs + full metric history + params + tags
- models → registered models
- model_versions → versions, with
sourcerewritten to the new GeneFlow run
Final state: done or failed (errors visible in dashboard / job.errors).
Idempotency
Re-running the import is safe:
- Experiments are looked up by name first — existing ones are reused
- Model versions are added to existing registered models (not duplicated, but new gf_model_versions rows may be created)
- Metric points may be duplicated on re-run — to be safe, drop and re-import a specific experiment
After import
- Spot-check a few runs: open
/dashboard/geneflow/mland click into a couple of imported experiments - Re-establish staging/production stages: imported versions land at
None— promote the ones you care about so the audit chain is clean - Build drift baselines for any models you're going to serve through GeneFlow — see api-reference.md → drift baselines
- Delete the old MLflow tracker once you're confident (or keep it read-only for a quarter)
FAQ
- Will my client code break? No — point
MLFLOW_TRACKING_URIat GeneFlow and it'll keep working. The importer is for history only. - Can I run it incrementally? Yes — scope by
experiment_names. Re-run as the source acquires new runs. - What if my MLflow has 100k runs? Use the SDK with a long-running session; the importer batches metric history in 1000-point chunks.
- What about hosted MLflow services? Confirm that the source exposes the MLflow tracking API, then validate its endpoint, authentication method, and artifact access with a small import before migrating the full history.