Workbook

Migration Engineer

You own: moving customers from legacy stacks (MLflow OSS, hosted MLflow, SageMaker, custom) onto GeneFlow without losing history or breaking client code.

What's new for you

The TS-native MLflow importer (Round 5, commit eb21849) does most of the heavy lift. UI: /dashboard/geneflow/ml/import.

Day-1 setup

pip install geneflow
export GENEFLOW_TRACKING_URI=https://api.genedata.io
export GENEDATA_PAT=$(genedata auth token)

Migration playbook — MLflow OSS / hosted MLflow

See migration-from-mlflow.md for the full procedure. Summary:

  1. Pre-flight
    • Confirm source reachable from GeneFlow API process
    • Check tenant quota; bump temporarily if needed
    • Dry-run to count entities
  1. Run import
    • Via UI: /dashboard/geneflow/ml/import
    • Or via CLI: gfctl import-mlflow --source … --token … --prefix "mlflow:"
  1. Post-import
    • Spot-check 5 random runs
    • Re-establish Production stages explicitly (audit chain starts clean)
    • Compute drift baselines for any models you'll serve through GeneFlow
    • Tell engineers to point MLFLOW_TRACKING_URI=https://api.genedata.io — their code keeps working
    • Mark the old tracker read-only (run for a quarter as fallback)

Migration playbook — SageMaker

SageMaker doesn't speak MLflow REST. Two paths:

  1. Export jobs + register manually

``python import boto3, geneflow sm = boto3.client("sagemaker") for job in sm.list_training_jobs(...)["TrainingJobSummaries"]: j = sm.describe_training_job(TrainingJobName=job["TrainingJobName"]) with geneflow.start_run(experiment_name=f"sagemaker:{j['TrainingJobName']}"): for k, v in j["HyperParameters"].items(): geneflow.log_param(k, v) for m in j.get("FinalMetricDataList", []): geneflow.log_metric(m["MetricName"], m["Value"]) ``

  1. Custom export script — sits in apps/api/src/engine/geneflow/import-from-sagemaker.ts (coming round).

Migration playbook — Custom ML stack

If they have a homegrown system:

  1. Phase 1 — point new training runs at GeneFlow (mlflow.start_run() → GeneFlow). Keep legacy parallel for a quarter.
  2. Phase 2 — batch-import historical runs via a custom script using geneflow.tracking.start_run().
  3. Phase 3 — deprecate legacy.

Workflow 1 — Dry-run a large import

gfctl import-mlflow \
  --source https://mlflow.legacy.example.com \
  --token "$LEGACY_TOKEN" \
  --prefix "legacy:" \
  --dry-run

The dashboard shows counts without writing. If counts look right, drop --dry-run.

Workflow 2 — Scope by experiment

gfctl import-mlflow \
  --source https://mlflow.legacy.example.com \
  --token "$LEGACY_TOKEN" \
  --experiments fraud_v1,fraud_v2,fraud_v3 \
  --prefix "legacy:"

Workflow 3 — Verify post-import counts

-- Imported via experiment_prefix
SELECT name, COUNT(*) AS runs
FROM gf_experiments e LEFT JOIN gf_runs r ON r.experiment_id = e.id
WHERE e.tenant_id='ACME' AND e.name LIKE 'legacy:%'
GROUP BY e.name;

Compare to the source counts captured during dry-run.

Workflow 4 — Re-stage Production models after import

# All imported model versions land at stage='None' — re-promote consciously
from geneflow import models

models.transition("fraud-detector", 3, to_stage="Production",
                  reason="re-promote after MLflow → GeneFlow migration")

The hash-chained audit log starts clean here.

Common gotchas

  • Artifacts are NOT copied by default — pass --copy-artifacts for that, but it's slow.
  • runs:/<id>/model URIs are rewritten by the importer to point at new GeneFlow run IDs. Code that hard-codes old MLflow run IDs needs updating.
  • MLflow stages don't carry over — by design. Re-promote.
  • Re-running is safe but creates duplicate metric points — to fully re-do an experiment, archive it first.

Where to go next