DataLake Architect
You own: the storage tier — S3 / GCS layout, lifecycle policies, partitioning, retention, archival, recovery. GeneFlow's artifact store + audit-log archive live here.
Storage layout
s3://genedata-geneflow-{region}/
{tenant_id}/
experiments/
{experiment_id}/
{run_id}/
artifacts/
model/ # the registered model bits
logs/ # training stdout/err
checkpoints/
metadata/
serving-images-cache/ # rarely used; per-tenant image cache
audit-archive/ # WORM archive of audit log
prompt-snapshots/ # collab editor periodic snapshots
inference-logs-cold/ # off-loaded gf_inference_logs (> 90d)
Per-tenant prefix is non-negotiable for the bucket policy to deny cross-tenant access.
Lifecycle
| Prefix | 0–30d | 30–90d | 90d–1y | 1y+ |
|---|---|---|---|---|
experiments/*/artifacts/ | Standard | Std-IA | Glacier IR | Deep Archive |
audit-archive/ | Standard + Object Lock | 7 year retention | ||
inference-logs-cold/ | Std-IA | Glacier IR | Deep Archive | |
prompt-snapshots/ | Standard | Std-IA | Glacier IR | Delete (1y) |
Object Lock on audit-archive/ in compliance mode = nothing, including admins, can delete.
Bucket policy template
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Deny",
"Principal": "*",
"Action": "s3:*",
"Resource": "arn:aws:s3:::genedata-geneflow-{region}/*",
"Condition": {
"StringNotEquals": {
"s3:ExistingObjectTag/tenant_id": "${aws:PrincipalTag/tenant_id}"
}
}
},
{
"Effect": "Deny",
"Principal": "*",
"Action": "s3:DeleteObject*",
"Resource": "arn:aws:s3:::genedata-geneflow-{region}/*/audit-archive/*"
}
]
}
Cross-region replication
Source: genedata-geneflow-us-east-1/{tenant_id}/audit-archive/ Destination: genedata-geneflow-compliance/
Replicate only audit-archive/ and experiments/.../model/ (the bits you might restore in DR).
Restore drill (quarterly)
- Pick a random run_id from prod 90 days ago
- Restore its model artifact from Glacier
- Spin up an endpoint from it in staging
- Verify a sample inference works
This validates: glacier retrieval is configured, KMS keys are accessible, image registry has the matching tag.
Cold inference log offload
Round 6 introduces gf_inference_logs. Default: keep 90 days hot in Postgres. Beyond that, run a nightly job:
COPY (SELECT * FROM gf_inference_logs WHERE ts < now() - INTERVAL '90 days')
TO PROGRAM 'aws s3 cp - s3://…/inference-logs-cold/$(date +%Y/%m/%d).csv'
WITH CSV HEADER;
DELETE FROM gf_inference_logs WHERE ts < now() - INTERVAL '90 days';
Query cold logs via Athena.
Common gotchas
- One bucket per region — never share buckets across regions; data-residency violations are silent and bad.
- Don't enable bucket versioning on
inference-logs-cold/— explosive cost. - DR plan must include re-creating KMS CMKs — if the key is gone, the data is gone.