Reference

Migrating from MLflow

Point existing mlflow code at GeneFlow, then import the history.

GeneFlow is wire-compatible with MLflow, so you don't have to migrate to use it — you can flip MLFLOW_TRACKING_URI to your GeneFlow base URL and existing code keeps working. But to bring history with you (past experiments, runs, registered models), use the built-in TS-native importer.

TL;DR

gfctl import-mlflow \
  --source https://mlflow.your-company.com \
  --token "$OLD_MLFLOW_TOKEN" \
  --prefix "mlflow:" \
  --dry-run

# If the counts look right:
gfctl import-mlflow \
  --source https://mlflow.your-company.com \
  --token "$OLD_MLFLOW_TOKEN" \
  --prefix "mlflow:"

Or use the UI: /dashboard/geneflow/ml/import — same engine, with live progress.

What gets imported

MLflow entity→ GeneFlow targetNotes
Experimentgf_experimentsRenamed with experiment_prefix if set
Rungf_runsmlflow.original_run_id stored as tag
Metric historygf_metrics (full time-series)Per-step values, not just final
Paramsgf_paramsSet-once
Tagsgf_tagsMutable
Registered modelgf_registered_modelsIdempotent on (tenant_id, name)
Model versiongf_model_versionsruns:/…/model source URI rewritten to new GeneFlow run
Artifacts(skipped by default — large)Pass --copy-artifacts to include

What does NOT migrate

  • Stage history before today — MLflow stages become None on import; re-transition explicitly so the hash-chained audit starts clean.
  • Webhooks / event listeners — register them again in GeneFlow.
  • Custom RBAC — set up via tenant settings in GeneFlow.

Pre-flight

  1. Confirm the source is reachable from the GeneFlow API process:

``bash curl -H "Authorization: Bearer $OLD_TOKEN" \ https://mlflow.your-company.com/api/2.0/mlflow/experiments/search?max_results=1 ``

  1. Check quota — the import will fan-out one gf_runs insert per run plus N metric inserts. Default tenant quota = 1k runs/hr; for very large MLflow installs, ask SRE to raise it temporarily.
  2. Dry-run — count entities first. The phase will reach done instantly without writes.

Running the import

Via UI

  1. Open /dashboard/geneflow/ml/import
  2. Fill in tracking URI + optional token
  3. Optional: scope by experiment names or registered models
  4. Set experiment_prefix (default mlflow:) to avoid collisions
  5. Click Start dry run → check counts
  6. Click Start import → live progress polls every 2s

Via SDK

from geneflow.import_mlflow import import_from

job = import_from(
    source_uri="https://mlflow.your-company.com",
    source_token=os.environ["OLD_MLFLOW_TOKEN"],
    experiment_prefix="mlflow:",
    dry_run=False,
)
# Poll
while job["phase"] not in ("done", "failed"):
    time.sleep(2)
    job = get_import_status(job["id"])

Via REST

curl -X POST https://api.genedata.io/api/v2.1/geneflow/import-mlflow \
  -H "Authorization: Bearer $GENEDATA_PAT" \
  -H "X-Tenant-Id: ACME" \
  -H "Content-Type: application/json" \
  -d '{
    "source_uri": "https://mlflow.your-company.com",
    "source_token": "...",
    "experiment_prefix": "mlflow:",
    "dry_run": false
  }'

Phases

The importer steps through six phases — each is visible in the dashboard:

  1. queued → just created
  2. connecting → probing source /api/2.0/mlflow/experiments/search
  3. experiments → creating GeneFlow experiments (idempotent on name)
  4. runs → for each experiment, fetch runs + full metric history + params + tags
  5. models → registered models
  6. model_versions → versions, with source rewritten to the new GeneFlow run

Final state: done or failed (errors visible in dashboard / job.errors).

Idempotency

Re-running the import is safe:

  • Experiments are looked up by name first — existing ones are reused
  • Model versions are added to existing registered models (not duplicated, but new gf_model_versions rows may be created)
  • Metric points may be duplicated on re-run — to be safe, drop and re-import a specific experiment

After import

  1. Spot-check a few runs: open /dashboard/geneflow/ml and click into a couple of imported experiments
  2. Re-establish staging/production stages: imported versions land at None — promote the ones you care about so the audit chain is clean
  3. Build drift baselines for any models you're going to serve through GeneFlow — see api-reference.md → drift baselines
  4. Delete the old MLflow tracker once you're confident (or keep it read-only for a quarter)

FAQ

  • Will my client code break? No — point MLFLOW_TRACKING_URI at GeneFlow and it'll keep working. The importer is for history only.
  • Can I run it incrementally? Yes — scope by experiment_names. Re-run as the source acquires new runs.
  • What if my MLflow has 100k runs? Use the SDK with a long-running session; the importer batches metric history in 1000-point chunks.
  • What about hosted MLflow services? Confirm that the source exposes the MLflow tracking API, then validate its endpoint, authentication method, and artifact access with a small import before migrating the full history.