Cloud Data Architect
You own: the cloud topology — region selection, network, IAM, encryption, storage tiers — that GeneFlow runs in.
Reference topology (AWS)
┌─────────────────────────┐
│ Route53 │
│ api.acme.genedata.io │
└────────┬────────────────┘
│
┌────────▼────────┐
│ ALB │
│ (TLS 1.3) │
└────────┬────────┘
│
┌──────────────▼───────────────┐
│ EKS — genedata namespace │
│ • monolith │
│ • geneflow-service │
│ • jobs-service │
│ • websocket-gateway │
│ • <serving pods per ep> │
└──┬─────────────┬─────────┬──┘
│ │ │
┌───────▼──────┐ ┌────▼───┐ ┌───▼──────┐
│ RDS Aurora │ │ S3 │ │ Secrets │
│ PG + pgvec │ │ KMS │ │ Manager │
└──────────────┘ └────────┘ └──────────┘
IAM — what each role needs
| Role | Permissions |
|---|---|
geneflow-runner (via IRSA) | s3:GetObject + s3:PutObject on arn:aws:s3:::genedata-geneflow-{region}/* only |
jobs-service | eks:DescribeCluster, secretsmanager:GetSecretValue for runner secrets |
geneflow-service | RDS Aurora connect, S3 read on its tenant's prefix |
| App user PATs | Scope-limited (see security.md) |
Networking
- VPC with private subnets for pods
- Aurora reachable only from the EKS cluster security group
- NetworkPolicies on pods enforce egress (see security.md)
- Optional VPC endpoints for S3 + Secrets Manager to keep traffic off the internet
Encryption
| Data | At rest | In transit |
|---|---|---|
| Postgres | KMS-encrypted Aurora (per-tenant CMK for high-sensitivity) | TLS to client |
| S3 artifacts | SSE-KMS per-tenant CMK | TLS |
| Secrets | KMS-encrypted Secrets Manager | TLS |
| Audit log | Same Postgres + S3 archive (Object Lock) | TLS |
| Pod-to-pod | mTLS via service mesh (recommended) |
Multi-region
For data-residency-strict tenants, deploy a separate stack per region. Cross-region:
- Audit log replicates read-only to a global compliance bucket (Object Lock)
- Container registry mirrored across regions for fast image pulls
Cost tiers (rough — actual depends on workload)
| Tier | Monthly cost | Capacity |
|---|---|---|
| Dev | $500–$1500 | 10 endpoints, 1M inferences |
| Production small | $3k–$8k | 50 endpoints, 100M inferences |
| Production large | $15k+ | 500+ endpoints, 10B inferences |
Drivers: GPU instances for serving, Aurora size, S3 storage class.