Workbook

AI Platform Engineer

You own: the GeneFlow platform itself — capacity, multi-tenancy, cost attribution, internal APIs, and the runtime path for any GeneAI workload at scale.

What's new for you

WasNow in GeneFlow
Glue scripts between MLflow, eval, servingSingle GeneFlow data model with tenant_id everywhere
Custom cost rollups in BIcost_usd first-class on runs + endpoints
K8s manifests applied by handHelm + jobs-service deployer auto-reconciles gf_endpoints
Tenant quotas in SlackCustomerTenantSettings.geneflow enforced in middleware

Capacity-shaping levers

LeverFile / tableWhat it controls
Tenant quotagf_tenant_quotas (via tenant settings)runs/hr, endpoints, GiB
Instance pricingserving.ts HOURLY_COST_USD$/hr per instance type
Snapshot intervalWS_SNAPSHOT_INTERVAL_MS envWS gateway → API write frequency
K8s job TTLttlSecondsAfterFinished in specwhen finished jobs are cleaned
Drift PSI thresholdgf_endpoints.drift_psi_thresholdper-endpoint sensitivity
HPA target CPU%k8s-serving-deployer.tsdefault 70%, change to 50% for latency-bound

Workflow 1 — Onboard a new tenant

genedata tenants create ACME \
  --quota-runs-per-hour 5000 \
  --quota-endpoints 200 \
  --quota-artifact-gib 1000

# Wire IAM (AWS IRSA) so the runner pod can `aws s3 sync`
helm upgrade --reuse-values genedata \
  --set global.agentRuntime.runnerIamRoleArn=arn:aws:iam::123:role/geneflow-ACME

# Issue tenant admin PAT
genedata auth issue-token --tenant ACME --scope geneflow:admin --user admin@acme.com

Workflow 2 — Per-tenant cost rollup

SELECT
  date_trunc('day', start_time) AS day,
  SUM(cost_usd) AS run_cost,
  COUNT(*) AS runs
FROM gf_runs WHERE tenant_id='ACME'
GROUP BY 1 ORDER BY 1 DESC;

-- Endpoint hourly burn
SELECT name, hourly_cost_usd, total_cost_usd, replicas
FROM gf_endpoints WHERE tenant_id='ACME' AND status='READY';

Workflow 3 — Bring up a new GeneFlow region

  1. New K8s cluster + image registry mirror
  2. Helm install with global.region=eu-west-1
  3. Storage: per-tenant prefix in regional bucket
  4. Federate audit chain via cross-region replication (read-only)
  5. Update Tenant data_residency_region for tenants requiring EU-only

Workflow 4 — Capacity guardrails for inference

-- Find tenants near their endpoint quota
SELECT t.tenant_id,
       t.quota_endpoints,
       COUNT(*) FILTER (WHERE status='READY') AS in_use,
       (COUNT(*) FILTER (WHERE status='READY')::FLOAT / t.quota_endpoints) AS usage
FROM gf_endpoints e JOIN tenant_quotas t USING (tenant_id)
WHERE tenant_id='ACME'
GROUP BY t.tenant_id, t.quota_endpoints;

Alert at 80% usage so SRE doesn't get paged at 100%.

Workflow 5 — Bring up a new instance class

  1. Add it to HOURLY_COST_USD and RESOURCE_PRESETS in serving.ts / k8s-serving-deployer.ts
  2. Bump migration: ALTER TABLE gf_endpoints DROP CONSTRAINT … ADD CHECK (instance_type IN (...))
  3. Update SDK type literal in serving.py
  4. Add to UI dropdown in endpoints/page.tsx
  5. Document in api-reference.md

Workflow 6 — Strangler-fig: extract more from the monolith

The platform was originally monolithic. We've already extracted: agent-runtime, embedding, stream-processor, automl, plugin-runner, federated, explainer, geneflow-service, websocket-gateway, jobs-service.

To extract another:

  1. Mirror the engine code into services/<name>-service/src/
  2. Use the Dockerfile pattern that COPYs both apps/api/src and services/<name>-service/src
  3. Mount the same router in services/<name>-service/index.ts
  4. Add to service-proxy.ts with env toggle (SERVICE_X_BACKEND=external)
  5. Add to helm/genedata/values.yaml services array → ServiceMonitor auto-generated

Common gotchas

  • A new column with no tenant_id leading the index is a bug. Every gf_ migration MUST.
  • Quota enforcement must happen in middleware, not in the engine — engine is the per-call enforcer, middleware is the rate-limiter.
  • Serving images need to live in the customer's registry namespace if data-residency requires it — don't hard-code registry.genedata.io.

Where to go next