Skip to main content
Production intelligence / evidence-first incident response Self-hosted · approval-gated

Every conclusion carries its evidence

Production Master investigates your production incidents and returns a report where every claim cites the log, metric or diff it came from — and nothing is changed without your approval.

Self-host via Helm — available now to closed-beta teams.

Evidence workspace · illustrative data
PM-SAMPLE-001checkout-api — p99 latency regressionAwaiting human approval

Conclusion

The connection-pool ceiling was cut from 80 to 20 in deploy a83f9c, starving checkout-api under normal traffic.

p99 410 ms 2.9 s · onset 14:23 UTC · seconds after deploy

Evidence

E-01
p99 request latency stepped from 410 ms to 2.9 s and held.
metrics · http_request_duration_seconds · 14:23:11 UTC
E-02
DB_POOL_MAX reduced 80 → 20 in the deployed values file.
deploy diff · checkout-api@a83f9c · 14:23:14 UTC
E-03
Pool exhaustion in application logs: 0 idle connections at the new limit, timeout_ms=2000.
logs · checkout-api · 2,140 matching lines · 14:23:19 UTC
E-04
Cache hit rate within baseline and upstream latency unchanged across the window.
metrics · cache and upstream telemetry · 14:27:42 UTC
Considered · rejectedCache degradation

Cache degradation does not match the evidence — the hit rate held within baseline through the entire regression window [E-04], and no cache-layer change was deployed. Discarded before the conclusion was formed.

Confidence

87%

4 of 4 supporting claims independently sourced. 3 alternatives tested and rejected.

Proposed action

Restore the previously reviewed DB_POOL_MAX value, then observe p99 and pool errors for ten minutes.

Human approval required

Production Master will not apply this change itself. It waits here.

Sources consulted

  • deploy history1 commit
  • metrics3 series
  • application logs2,140 lines
  • cluster events18 events
Reads the evidence you already produce
deploy diffsmetric seriesapplication logspod & node eventsconfig changestraces

An answer you can audit, not one you have to trust

Most tools hand you a conclusion. Production Master hands you the conclusion, the evidence under it, and the alternatives it ruled out on the way.

01

Gathers, then cites

Every log line, metric window, deploy diff and config change it reads becomes a numbered exhibit. Nothing enters the report uncited.

claim → E-0n → raw source

02

Argues against itself

Competing hypotheses are tested and the losers stay in the report, with the evidence that killed them. A rejected theory is a result, not a deletion.

considered · rejected · why

03

Stops at the gate

It proposes the change and waits. Remediation is a decision a human makes, holding a report they can check line by line.

proposed → awaiting approval

The hypotheses it rejected ship with the report

An investigation that only shows you the winning theory is asking for trust.
Production Master keeps the whole hypothesis ledger — what it considered, what the evidence did to each one, and why the survivor survived.

Every verdict links back to the exhibit that produced it.

PM-SAMPLE-001Hypothesis ledger4 considered · 1 held
rejected
Cache degradation
Cache hit rate stays within baseline across the whole incident window, and no cache-layer change was deployed. [E-04]
rejected
Upstream provider latency
Upstream latency does not change during the window, and checkout-api is slow on requests that never reach the provider. [E-04]
rejected
Traffic spike
The latency series steps rather than ramps, and it steps at the deployment boundary rather than with load. [E-01]
held
Connection-pool exhaustion after deploy a83f9c
The pool ceiling drops 80 to 20, latency steps immediately after, and the logs show zero idle connections at the new limit. [E-02, E-01, E-03]

Run it yourself, or let us run it

The investigation engine is the same either way.

Community
$0

Self-hosted, bring your own LLM keys. The full evidence model, no seat limits, no investigation cap.

Helm chart available now to closed-beta teams.

Read the self-host guide
Cloud Team
$199 /mo

We host it. Includes 25 investigations per month, then $30 per seat, per month for your responders.

Managed upgrades and retention. Additional investigations are $8 per investigation with your own keys, or $14 per investigation on ours.

Start with Cloud Team
Enterprise
from $60k /yr

Private deployment in your own account, SSO, custom evidence connectors, and a named engineer on the account.

Annual agreement.

Talk to sales
Closed beta terms

An 8-week pilot run against your real incidents. $7,500, credited in full toward the first year. No pre-committed conversion discount.

Request access

Ship the fix because you checked, not because you guessed

Bring one real incident.
The first cited report lands in week 1 of a 8-week pilot, so you are judging evidence you can check rather than a promise you have to take.

Self-host via Helm — available now to closed-beta teams.