Professional learning path

DataLake Architect

Design a data-lake layout that supports discovery, retention, and recovery.

Prepare

Map the data and control boundaries

Trace the reconciliation kit from source files to review output. Sketch the production sources, trust boundaries, integration path, storage, and downstream consumers before choosing a deployment.

Start with the browser demo or download the local kit. For the workspace exercise, you need approved access to the relevant Genedata capabilities and a reviewer for your output.

Practice tasks

0 / 4 completed

1
2
3
4
Knowledge checkpoint

A source system must stay inside an approved network boundary. What comes first?

Notes and progress are saved on this device.

Apply in your workspace

Make storage a usable data foundation.

Design a data-lake layout that supports discovery, retention, and recovery.

1

Inventory the data domains, artifact types, consumers, and lifecycle requirements.

2

Define a consistent storage layout, naming scheme, partition strategy, and ownership model.

3

Review retention, archival, access, and recovery procedures with governance and operations.

4

Publish the design and validate a representative ingestion and retrieval workflow before expanding it.

Expected handoff

A storage design with lifecycle rules, ownership, and recovery evidence.

Explore your platform capabilities

Measure your progress

  • Time to locate the correct dataset or artifact
  • Storage domains with reviewed lifecycle and recovery rules

Use AI to summarize layout alternatives; validate cost, access, and retrieval behavior against the real workload.