Runtimezruntimez
runtimez / product / production-reliability
product · production reliability

Find the workload that fails next — while it still looks healthy.

Every running workload is scored continuously against 200+ rules. 49 correlation engines check its config against live telemetry, which a manifest scanner can't do.

one helm install · get, list, watch only · first findings in minutes

200+live risk rules, scored continuously
49correlation engines that check config against telemetry
3failure scenarios per workload: node, AZ, maintenance
4add-on workspaces: GitOps, ingress, autoscaling, mesh

Live workload risk

The daily value: a ranked list, not a one-shot scan.

Continuous scoring

Every running workload is scored on runtime health, spec hygiene, resource sizing, security, change history and add-ons.

200+ rulesper workload

Config checked against telemetry

Examples: an HPA denominator mismatch, a PDB that blocks its own drain, a probe timeout versus the real p95, a grace period versus in-flight requests, a limit cut below observed usage.

49 engines

Failure-scenario verdicts

PASS, AT_RISK, WILL_FAIL or WILL_STALL under node failure, AZ failure and maintenance, each with a plain-English reason.

nodeAZmaintenance

Dependency blast radius

Uses observed call volume to show who else goes down with this workload.

observed calls

Autoscaler verdicts, verbatim

Shows FailedGetResourceMetric and similar HPA errors before the autoscaler silently stops scaling.

HPA

Every finding explained

Each finding comes with its evidence, the rule that fired and a fix. Findings are ranked fix-first across reliability, security and upgrade risk.

evidencefix-first
Runtimez cluster view: pods, deployments, restarts, node CPU and memory, and the workloads that need attention
Runtimez cluster view: pods, deployments, restarts, node CPU and memory, and the workloads that need attention

Add-on workspaces

One tab per detected add-on. Each uses the add-on's own vocabulary and verdicts and links straight into root-cause analysis.

GitOps · Argo CD

Every Application with its sync and health state. NOT_SYNCING is explained in 31 deterministic rungs read from Argo CD's own messages. 39 rules catch problems that would hit at the next sync.

39 rules31 rungs

Ingress · Traefik

Every route with Traefik's own router verdict, the backends it resolves to, and live request counters. 44 rules, including routes that no backend serves.

44 rules

Autoscaling · Karpenter + Cluster Autoscaler

NodePools with limits vs. usage, NodeClaims and disruption events. Cluster Autoscaler node groups get a banner when the scaling loop goes stale. 19 rules, including nodes that fail to register.

19 rules

Service mesh · Istio

Routed workloads, traffic counters that a sidecar restart can't inflate, and Envoy response flags per edge. 24 rules, including fault injection left on and mTLS mismatches.

24 rules

The evidence layer

Uses your telemetry as evidence. It is not a Datadog replacement.

Tracing, logs and metrics

Distributed traces in waterfall, DAG and scatter views; an observed service graph; logs with AI summaries; a metrics explorer with tag search.

tracesservice graph

Zero-code instrumentation

OTel SDKs injected by policy with no operator, a Go eBPF sidecar, and eBPF traces as a fallback for everything else.

no code changes

Right-sizing

Compares requests and limits to observed usage per container. Under-requested containers are your next OOM; over-requested ones are your bill.

requests vs usage

Telemetry cost controls

Log drop and sample rules with a live preview, trace sampling in shadow mode, and a metrics cardinality view, so the bill doesn't surprise you.

shadow modecardinality

How it compares

vs Polaris / Popeye

They check config against best practice. Runtimez checks config against your live telemetry.

vs Datadog / Grafana

They tell you what broke. Runtimez tells you what will break, and why it broke, with citations.

See it on your own cluster in under an hour.

Free for your first cluster. Read-only by default. Uninstall is one helm command.