Reference

Security

Authentication, scopes, secrets, audit, and the controls behind them.

GeneFlow is tight by default: every runner / serving pod runs non-root, drops all capabilities, has seccomp RuntimeDefault, and is constrained by a NetworkPolicy. PATs are scoped, audit is hash-chained, and Production stage transitions require approval.

Pod security context

Both the project runner and the serving pod use the same hardened defaults:

securityContext:
  runAsNonRoot: true
  runAsUser: 1000
  fsGroup: 1000
  seccompProfile: { type: RuntimeDefault }

containers:
  - securityContext:
      allowPrivilegeEscalation: false
      capabilities: { drop: ["ALL"] }
      readOnlyRootFilesystem: false  # model libs need /tmp; mount emptyDir read-write

RBAC

Project runner (geneflow-runner ServiceAccount)

Granted only what runner pods actually need:

rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log"]
    verbs: ["get", "list", "watch"]                    # diagnostics only
  - apiGroups: [""]
    resources: ["configmaps"]
    verbs: ["get", "list"]                              # parameters
  - apiGroups: [""]
    resources: ["secrets"]
    resourceNames: ["geneflow-runner-secrets"]
    verbs: ["get"]                                       # only the system-token secret

No exec. No port-forward. No write access to any resource.

Serving deployer (jobs-service ServiceAccount)

Granted only resources required for serving:

rules:
  - apiGroups: ["apps"]
    resources: ["deployments", "deployments/scale", "deployments/status"]
    verbs: ["get","list","watch","create","update","patch","delete"]
  - apiGroups: [""]
    resources: ["services"]
    verbs: ["get","list","watch","create","update","patch","delete"]
  - apiGroups: ["autoscaling"]
    resources: ["horizontalpodautoscalers"]
    verbs: ["get","list","watch","create","update","patch","delete"]
  - apiGroups: [""]
    resources: ["pods","pods/log","events"]
    verbs: ["get","list","watch"]                        # diagnostics

Cannot read secrets. Cannot exec.

NetworkPolicy

Project runner pods

egress:
  - DNS (UDP/TCP 53 in kube-system)
  - GeneFlow API (port 4032, in-cluster)
  - Monolith   (port 4003, in-cluster)
  - External HTTPS 443 + HTTP 80
    except: 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16

No ingress (runners only push).

Serving pods (app.kubernetes.io/component: model-serving)

ingress:
  - from: { API gateway, monolith, geneflow }
    ports: [8080]
egress:
  - DNS
  - GeneFlow API (4032) — for logging inferences back
  - Monolith   (4003)
  - HTTPS 443 — for artifact pulls
    except: RFC1918 ranges

A PodDisruptionBudget with minAvailable: 1 covers the serving fleet.

Secrets

SecretContainsSync
geneflow-runner-secretsJOBS_SYSTEM_TOKEN, GENEFLOW_API_URLExternal Secrets Operator → AWS Secrets Manager / GCP Secret Manager (recommended)
Image-pull secretsper-registryprovisioned outside Helm

Production deployments must not ship the placeholder JOBS_SYSTEM_TOKEN baked in via Helm values — wire it through External Secrets. The Helm template has an opt-out:

global:
  agentRuntime:
    runnerSecretExternal: true   # don't create the Secret; ESO will

Tokens & scopes

Personal Access Tokens (PATs) issued via /api/v2/auth/tokens carry scopes; the engine enforces them on every call:

ScopeGrants
geneflow:readList + fetch everything
geneflow:writeTracking writes, model registration
geneflow:transitionStage changes (Production = approval)
geneflow:serveCreate / update / delete endpoints
geneflow:promptsPrompt CRUD + collab
geneflow:adminTenant-level config

System tokens used by runners / serving pods are issued separately and carry only internal:ingest (POST inference logs, POST run status). They cannot read.

Approval gates

Production transitions for model versions and prompt versions require an approver if requireApproval=true (the default). Approval lives in:

  • gf_stage_transitions.requested_by / approved_by
  • genedata_audit_log with action="geneflow.stage.transition" and prev_hash chain

Bypass: register the version with requireApproval=false (logged separately) or mark the tenant as auto_approve_prod=true (governance review required).

Hash-chained audit

Every important change writes to genedata_audit_log:

row_hash_N = sha256(prev_hash_{N-1} || canonical_json(payload))

Tampering with a single row breaks all rows after. Verify:

SELECT * FROM verify_audit_chain('ACME');

Coverage today:

  • run lifecycle (create, finish, fail)
  • model version create + transitions
  • prompt version create + transitions
  • endpoint create / update / delete
  • drift alert ack

Storage

  • Artifacts: per-tenant prefix s3://bucket/<tenant_id>/... with bucket policy denying cross-tenant access.
  • Server-side encryption: KMS-managed keys (one CMK per tenant for high-sensitivity tenants).
  • Versioning enabled. Lifecycle: 90 days warm, 1 year glacier, 7 year retention.

Data residency

Per-tenant data_residency_region field on the tenant settings record. Importer + serving deployer both honor it — if set to eu-west-1, the serving image registry, artifact bucket, and K8s namespace must all be in-region.

Threat model summary

ThreatMitigation
Compromised runner pod escapesNon-root + cap-drop + seccomp + NetworkPolicy
Compromised pod reads other tenant datatenant_id in every WHERE; SA can only read its own ConfigMap + named Secret
Compromised pod calls model registrySystem token can't transition stages; needs human approval
Stolen PATScope-limited; can be rotated; audit row per use
Audit tamperingHash chain detects
Supply-chain on serving imageSigned images via cosign; image-pull restricted to registry.genedata.io
Drift not noticedPSI alerts at endpoint threshold; can page on critical
Approval bypassrequireApproval=false is logged separately as a tenant-level event

Operational checklist (per release)

  • [ ] All new endpoints / fields ship with tenant_id in the leading index
  • [ ] New audit-relevant actions wired through appendImmutableAudit()
  • [ ] New REST endpoints behind authMiddleware and call getTenantId(c)
  • [ ] New Pod templates inherit the shared securityContext
  • [ ] Any new NetworkPolicy is additive, not replacement
  • [ ] Threat model reviewed; new threats added to the table above