Anomaly cockpit
Row counts, freshness, and run duration are baselined per pipeline. A deviation is annotated on the series with the step, the run, and the downstream datasets it would have reached.
On a fragmented stack the first sign of a broken pipeline is a stakeholder asking why the dashboard is stale. On Genedata the engine that runs the pipeline is the engine that watches it: volume and freshness anomalies are detected at the failing step, transient failures retry from the last good checkpoint, and only what needs a human reaches the inbox.
Row counts, freshness, and run duration are baselined per pipeline. A deviation is annotated on the series with the step, the run, and the downstream datasets it would have reached.
Transient failures — a source timeout, a rate limit, a late partition — retry with backoff from the last checkpoint without a page. What remains is a short list of decisions, each with the context attached.
Set who is paged, for what severity, on which pipelines, once. Policies follow the graph, so a new downstream dashboard inherits its upstream coverage instead of needing its own monitor.
A representative row-count series for a nightly orders pipeline. The drop is flagged on the run where it happened, with the step and its downstream datasets attached — before the dashboard that reads it refreshes.
The monitoring surface as it appears in the product. Healed retries are logged under the run; only open decisions sit in the inbox.
Because monitoring runs inside the same graph as the contracts and lineage, an anomaly arrives with its cause and its blast radius already known. The uptime objective we publish is 99.9% per service; we state the objective rather than a marketing figure, and the status page shows the measurement.
What the person on call asks before they trust the inbox to stay quiet.
Transient failures with a known recovery — source timeouts, rate limits, a late-arriving partition, a worker restart — retry with backoff from the last good checkpoint and are logged, not paged. A contract failure, a repeated retry exhaustion, or a volume anomaly outside the baseline reaches the healing inbox with its context, and your alerting policy decides whether it also pages.
Each pipeline's row counts, freshness, and run duration are baselined from its own recent history, with seasonality per weekday where there is enough data. A point outside the expected band is annotated on the series at the step where it occurred, and lineage lists what is downstream. The series on this page is a representative example, not live data.
It is the per-service uptime objective we publish and measure against, recorded in the service level agreement and tracked on the status page. We state the objective rather than quote a historical availability number, and the SLA sets out what happens if it is missed.
Yes. Alerting policies deliver to email, chat, and paging integrations, with severity and routing set per policy. The policy is attached to the graph, so a new pipeline or downstream dashboard inherits coverage rather than needing a separately configured monitor.
The security proof shows the controls around the data — row and column policy, secrets, and the audit trail — and how they map to the frameworks you are asked about.