Runtimezruntimez
runtimez / product / ai-sre
product · ai sre & developer productivity

An AI SRE that shows its evidence — and says when it doesn't know.

Runtimez runs root-cause analysis for any workload. It gathers evidence deterministically, makes one bounded model call, and every sentence cites evidence you can click. You can ask your fleet questions from Slack, the console, the terminal or your own AI assistant.

one helm install · get, list, watch only · first findings in minutes

57MCP tools for Claude, Cursor or any MCP client
7symptoms you can start debugging from
1bounded model call per RCA, every claim cited

Root-cause analysis

It turns a pager alert into a narrative you can check.

On-demand RCA for any workload

Gathers events, logs, restarts, probes, rollouts, HPA, metrics, traces, add-on findings and mesh traffic, then makes one bounded model call. Every sentence cites its evidence.

citeddeterministic gather

Auto-RCA with causal chain

Root cause, then propagation, then symptom, hop by hop.

causal chain

Named failure rungs

Specific causes instead of a generic "unhealthy": a liveness probe killing a container that never finished booting, scaled to zero, image pull failure, missing config reference, a PDB blocking a drain.

specific causes

Hypotheses with refuting evidence

States what would disprove each hypothesis, and reports INDETERMINATE instead of making something up.

no confabulation

"Why is no node coming?"

For Pending replicas, names the Karpenter NodePool that refused the pod and why.

Karpenter

Start from the symptom

Restarting, slow, 5xx, cannot connect, memory growth, stuck rollout, not syncing. Ingress and mesh runtime evidence is scoped to the workload's own routes.

7 symptoms

Trace bottlenecks and profiles

Critical path, self-time and per-edge latency for each trace, plus CPU and allocation profiles captured during a debug session.

critical pathprofiling

Repeat incidents are free

Failure signatures are cached, so a repeat incident gets an answer instantly at no cost.

signature cache
Runtimez AI root-cause analysis: ranked hypotheses with evidence references and a recommendation
Runtimez AI root-cause analysis: ranked hypotheses with evidence references and a recommendation

Developer productivity

Your fleet becomes a tool your team and its AI assistants can call.

Ask Runtimez

Ask the cluster a question in plain English and get a structured verdict. A scope guardrail refuses off-topic questions instead of making up an answer.

plain Englishguardrail

Runtimez MCP server

57 read-only tools for Claude, Cursor or any MCP client, covering inventory, changes, risk, the readiness ledger, observability and RCA. Every answer carries its syncId. Rate-limited and audited.

57 toolsread-only

rtz CLI with TUI

Fleet, cluster, workload, pod and events from the terminal, plus an rtz mcp bridge.

terminal

Slack

Ask Runtimez answers and alerts arrive in the channel you already use.

Slack

AI insights feed

A ranked list of what needs attention: security, reliability, cost, hygiene, anomalies, top movers and forecasts.

ranked

How it compares

vs Datadog / Grafana

They tell you what broke. Runtimez tells you what will break, and why it broke, with citations.

See it on your own cluster in under an hour.

Free for your first cluster. Read-only by default. Uninstall is one helm command.