Data Engineer / Data & ModelingGENEDATA / 01

Build once. Deliver trusted data.

Owns source connections, pipelines, transformations, contracts, backfills, and release evidence for reliable data products.

Shared context · Lineage · Governance
Connected
Your sources
Connect trusted sources
Build the workload
Test contracts and behavior
Connected intelligenceData Engineer
Governance
Business impactRelease with evidence
Shared contextLineageGovernance
+Illustrative workflow01 / 03
1 / 3
Role context

The context behind the work.

Your work is the connection between a source system and the decisions built on it. A successful pipeline is not just a completed run: its consumers understand the data, its owner can explain a change, and operations can recover it when something fails.

In practice / 01

A source schema changes

A source team renames a customer field. Inspect the downstream datasets and model consumers, agree a compatible transition, test the revised transformation, and publish the change with the schema reference and migration notes.

In practice / 02

A new feature dataset is needed

A data scientist needs a reproducible training input. Agree the feature definitions and grain, publish a versioned dataset, record the upstream references, and share the run and schema artifacts rather than an untracked file.

Your workflow

A practical path from task to outcome.

Publish a feature dataset that downstream analysts and model builders can trace and reuse.

  1. 01

    Connect trusted sources

    Choose a source and confirm its owner, access scope, schema, and expected freshness before creating the pipeline.

  2. 02

    Build the workload

    Build the transformation in the visual workflow or in code. Keep the dataset version and schema artifact with the run.

  3. 03

    Test contracts and behavior

    Check the output and record upstream dataset and feature-group references. Coordinate with consumers before changing a shared field.

  4. 04

    Release with evidence

    Publish the output and give the next team the run reference, lineage, expected refresh schedule, and recovery instructions.

What you take forward

A versioned dataset, its schema, and a traceable run reference.

Work more effectively

Less repeated effort. More useful work.

Explore the habits and platform connections that can make this role easier, more consistent, and easier to collaborate with.

A common friction

Reconstructing dependencies during an incident

Maintain upstream references with the run so downstream impact can be inspected before a change.

A common friction

Copying transformation logic between teams

Publish a reusable data product with a clear owner and contract instead of repeating the same preparation.

A common friction

Repeating investigation after a failed handoff

Share validation evidence and a recovery path alongside the dataset.

Measure your own improvement

Choose a baseline before you begin. Review these signals with your team; results depend on your data, process, and implementation.

  • Time from a source change to a validated publication
  • Rework caused by missing contracts or unclear ownership
Get started

Build confidence with a first task.

Publish a feature dataset that downstream analysts and model builders can trace and reuse.

Use AI with judgment

Use AI to draft transformation logic or explain a failure, then review the code and validate the output before release.

Your practice checklist

0 / 4 complete
Your toolkit

The right surfaces. The right people.

Continue into the product, deepen your knowledge, or follow the next role in the handoff.

Go deeper

Technical workbookDocumentation

Technical workbooks are maintained in English. Workspace access and available capabilities depend on your deployment and permissions.

Data Engineer

Bring your own workflow.

Explore how these practices could fit your team, your data, and your operating requirements.