Workbook

DataLake Architect

You own: the storage tier — S3 / GCS layout, lifecycle policies, partitioning, retention, archival, recovery. GeneFlow's artifact store + audit-log archive live here.

Storage layout

s3://genedata-geneflow-{region}/
  {tenant_id}/
    experiments/
      {experiment_id}/
        {run_id}/
          artifacts/
            model/                # the registered model bits
            logs/                 # training stdout/err
            checkpoints/
            metadata/
    serving-images-cache/         # rarely used; per-tenant image cache
    audit-archive/                # WORM archive of audit log
    prompt-snapshots/             # collab editor periodic snapshots
    inference-logs-cold/          # off-loaded gf_inference_logs (> 90d)

Per-tenant prefix is non-negotiable for the bucket policy to deny cross-tenant access.

Lifecycle

Prefix0–30d30–90d90d–1y1y+
experiments/*/artifacts/StandardStd-IAGlacier IRDeep Archive
audit-archive/Standard + Object Lock7 year retention
inference-logs-cold/Std-IAGlacier IRDeep Archive
prompt-snapshots/StandardStd-IAGlacier IRDelete (1y)

Object Lock on audit-archive/ in compliance mode = nothing, including admins, can delete.

Bucket policy template

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:*",
      "Resource": "arn:aws:s3:::genedata-geneflow-{region}/*",
      "Condition": {
        "StringNotEquals": {
          "s3:ExistingObjectTag/tenant_id": "${aws:PrincipalTag/tenant_id}"
        }
      }
    },
    {
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:DeleteObject*",
      "Resource": "arn:aws:s3:::genedata-geneflow-{region}/*/audit-archive/*"
    }
  ]
}

Cross-region replication

Source: genedata-geneflow-us-east-1/{tenant_id}/audit-archive/ Destination: genedata-geneflow-compliance/

Replicate only audit-archive/ and experiments/.../model/ (the bits you might restore in DR).

Restore drill (quarterly)

  1. Pick a random run_id from prod 90 days ago
  2. Restore its model artifact from Glacier
  3. Spin up an endpoint from it in staging
  4. Verify a sample inference works

This validates: glacier retrieval is configured, KMS keys are accessible, image registry has the matching tag.

Cold inference log offload

Round 6 introduces gf_inference_logs. Default: keep 90 days hot in Postgres. Beyond that, run a nightly job:

COPY (SELECT * FROM gf_inference_logs WHERE ts < now() - INTERVAL '90 days')
  TO PROGRAM 'aws s3 cp - s3://…/inference-logs-cold/$(date +%Y/%m/%d).csv'
  WITH CSV HEADER;

DELETE FROM gf_inference_logs WHERE ts < now() - INTERVAL '90 days';

Query cold logs via Athena.

Common gotchas

  • One bucket per region — never share buckets across regions; data-residency violations are silent and bad.
  • Don't enable bucket versioning on inference-logs-cold/ — explosive cost.
  • DR plan must include re-creating KMS CMKs — if the key is gone, the data is gone.

Where to go next