Failure Rate Distributions & Suite Architecture Debt

Most teams track a single aggregate metric: overall pass rate. That number is almost useless for diagnosing architecture problems. A suite that passes 94% of scenarios on every run can still be hiding a cluster of 40 scenarios that fail 60% of the time, offset by 800 scenarios that never fail — because they test nothing that ever changes. The aggregate masks the distribution.

Failure rate distributions — the per-scenario, per-tag, and per-module breakdown of how often each scenario fails across N consecutive CI runs — expose a different class of problem: structural debt baked into the suite itself. Overcoupled fixtures, shared step libraries with blast radius, and Scenario Outline tables that silently multiply brittle step bindings all leave distinctive fingerprints in the distribution shape.

By the end of this article you'll know how to collect per-scenario failure rates at scale, model the resulting distribution to identify architectural anti-patterns, and prioritize remediation using a signal that's already in your CI data. The tooling is Pytest + Behave or Cucumber-JVM 7 with a lightweight aggregation layer — no third-party observability vendor required.

Manage All Your AI API Keys in One Place

Securely manage keys for 60+ AI providers in one encrypted vault instead of juggling them across apps.

Learn more

What Failure Rate Distributions Actually Measure

A scenario's failure rate over a rolling window of CI runs is a proxy for coupling density. A scenario that fails 5% of the time on a stable codebase is noisy — its fixture setup is touching shared state, its step definitions depend on execution order, or its data assumptions are leaking from an adjacent scenario. A scenario that fails 0% of the time for six months is either perfectly isolated or perfectly inert. Neither extreme is inherently good; the distribution shape tells you which.

In a well-architected suite, failure rates approximate a right-skewed distribution: most scenarios cluster near zero, a small tail reflects genuinely volatile product areas, and almost nothing sits in the 20–60% band. That middle band is the danger zone. Scenarios there are failing often enough to erode trust but not reliably enough to be treated as broken. They consume triage time without producing signal. When you plot your suite and see a fat middle band, you're looking at structural suite debt — not flakiness, not product bugs, but architectural decisions that made coupling the path of least resistance.

Collecting and Modeling the Distribution in CI

The first step is persisting per-scenario outcomes across runs. Behave's JSON formatter and Cucumber-JVM's --plugin json both emit scenario-level pass/fail with a stable URI (feature file path + scenario name + example row index). Write a post-run step in your pipeline that appends those outcomes to a time-series store — a Postgres table or even a flat Parquet file in S3 is sufficient for suites under 10,000 scenarios.

# GitHub Actions — append scenario outcomes after every run
- name: Aggregate scenario outcomes
  run: |
    python scripts/aggregate_outcomes.py \
      --input reports/cucumber.json \
      --run-id "${{ github.run_id }}" \
      --output s3://test-metrics/outcomes/
  env:
    AWS_DEFAULT_REGION: us-east-1

With 30–50 runs of history, compute a failure rate per scenario URI and bucket them into bands: 0–5%, 5–20%, 20–60%, 60–100%. The 20–60% band is your primary target. A Python snippet using Pandas makes this trivial:

import pandas as pd

df = pd.read_parquet("s3://test-metrics/outcomes/")
rates = (
    df.groupby("scenario_uri")["failed"]
    .agg(failure_rate="mean", run_count="count")
    .reset_index()
)
rates["band"] = pd.cut(
    rates["failure_rate"],
    bins=[0, 0.05, 0.20, 0.60, 1.0],
    labels=["stable", "noisy", "toxic", "broken"],
    include_lowest=True,
)
print(rates.groupby("band").size())

The "toxic" band (20–60%) is where architectural debt concentrates. Cross-reference those scenario URIs against your feature file structure. In one platform team's 3,200-scenario Behave suite, 87 scenarios in the toxic band all shared two step definition modules that wrote to a singleton Redis fixture. Extracting those fixtures into scenario-scoped setup via @pytest.fixture(scope="function") equivalents in Behave's before_scenario hook dropped the toxic-band count from 87 to 11 in the next 30-run window — and mean CI run time fell from 22 minutes to 14 because retry loops were no longer masking the failures. That's a measurable outcome from a distribution analysis, not a gut-feel refactor.

Once you have band assignments, join them against your tag taxonomy. Tags that correlate strongly with the toxic band are unintentional test selection policy — teams start skipping @smoke runs to avoid noise, which means tags have become an unintended test selection policy rather than a classification tool. Surfacing this in a Grafana dashboard (failure rate by tag, updated per CI run via a Prometheus pushgateway) makes the debt visible to engineering managers without requiring them to read JSON reports.

Where Senior Engineers Misread the Distribution Signal

The most common mistake is conflating the toxic band with flakiness and routing it to a flake-suppression workflow — retries, quarantine tags, --rerun-failures in Pytest. Retries reduce visible red in the dashboard while the underlying coupling gets worse. Flaky tests have a random failure pattern with no structural correlation; toxic-band scenarios cluster by feature file, by shared fixture, or by execution order. If your 20–60% failures correlate spatially in the suite, they're not flaky — they're structurally coupled. Treating them as flaky delays the architectural conversation by quarters. Note also that shard rebalancing can silently skew your flake rate baselines, making toxic-band scenarios appear to improve when the underlying coupling hasn't changed at all.

A second mistake is measuring failure rates over too short a window. Thirty runs is a minimum; fewer than 20 gives you a binomial estimate with confidence intervals so wide that a scenario with 2 failures in 10 runs looks identical to one with 6 failures in 10 runs. Teams that pull a single week of CI history and act on it are optimizing noise. Use a rolling 60-run window and weight recent runs more heavily only if you've shipped a major fixture refactor — otherwise recency bias hides regression.

Myths That Keep Suite Debt Invisible

Myth 1: A high overall pass rate means the suite is healthy. As shown above, aggregate pass rate is a weighted average that hides distribution shape. A suite with 95% pass rate and a fat toxic band is less trustworthy than one with 91% pass rate and a clean right-skewed distribution, because the former has dozens of scenarios that produce unpredictable signal. Myth 2: Scenario count is a proxy for coverage quality. Suites that grow via copy-paste Scenario Outline expansion — without validating that each example row exercises a distinct code path — inflate count while concentrating risk in shared step bindings. The distribution will show this as a cluster of toxic-band scenarios with identical URIs differing only by row index. Reviewing scenario-level risk scores alongside failure rates separates coverage breadth from coverage depth.

Myth 3: Fixing flaky tests fixes the distribution. Flake remediation and architecture remediation are separate workstreams targeting different root causes. Flake comes from environment instability — timing, network, test isolation at the process level. Architecture debt comes from design decisions: shared fixtures, monolithic step libraries, implicit execution-order dependencies. A quarantine tag on a toxic-band scenario stops the bleeding but doesn't pay down the debt. The distribution will keep regenerating new toxic-band scenarios from the same structural source until the coupling is broken at the fixture or step-definition layer.

Failure rate distributions give you a repeatable, data-driven entry point into suite architecture conversations that would otherwise stall on opinion. Start by computing 60-run rolling rates for every scenario URI, plot the band breakdown in Grafana, and treat any toxic-band cluster that correlates spatially in your feature tree as an architectural finding — not a flake ticket. The next measurement worth adding once the toxic band shrinks: mean-time-to-detect on genuine product regressions, which should improve as the noise floor drops.

Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.

Understanding how systems actually work is the first step toward navigating them effectively.

Browse all articles