Shard Timeouts That Silently Drop Hook Execution
Distributed test sharding is table stakes for any suite that's grown past a 10-minute wall. What most teams don't account for is what happens when a shard agent hits its timeout ceiling mid-scenario: the runner process is killed, the after-scenario hook never fires, and the CI job reports a timeout failure — not a hook failure. The fixture teardown, the database rollback, the browser session close, the token revocation — all silently dropped.
The symptom is usually discovered two steps removed from the cause: a poisoned shared resource causes a flaky failure in a completely different scenario on the next run, or a staging environment accumulates orphaned sessions until it falls over at 2 AM. By then, the shard timeout that triggered it is buried in a log no one reads.
This article explains the exact mechanism by which shard timeouts suppress hook execution in Cucumber-JVM 7, Behave, and SpecFlow, how to reproduce it locally, and what instrumentation to add so you catch it before it corrupts your environment.
Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.
Why the After-Scenario Hook Is the First Casualty of a Timeout
When a CI runner — GitHub Actions, Jenkins, or an Argo workflow step — enforces a hard wall-clock timeout on a shard process, it sends SIGTERM (or on Windows, terminates the JVM/Python process outright). The BDD framework's internal scheduler is not given the chance to drain its hook queue. In Cucumber-JVM 7, the @After hook is registered per-scenario in an event bus; when the process exits, that bus is torn down without flushing pending callbacks. Behave's after_scenario function in environment.py faces the same fate — it's called by the runner loop, and a killed loop never reaches that call.
This is distinct from a scenario failure, where the framework does execute after-hooks (that's the whole point of teardown). A timeout is a process-level event, not a framework-level one. Understanding that distinction matters for where you instrument: you cannot catch this inside the framework. You need to detect it at the infrastructure layer — the agent, the container, or the orchestrator — and design your fixture lifecycle to be resilient to abrupt termination. Hook execution order determines suite-wide fixture reliability precisely because the order in which hooks are registered affects which ones are most exposed when a process is cut short.
Reproducing the Drop and Building a Detection Layer
Start by reproducing it deterministically. The following Behave setup creates a scenario that sleeps past the shard timeout, then checks whether teardown ran:
# environment.py
import time, os
def after_scenario(context, scenario):
# Write a sentinel file; absence after the run means hook was dropped
with open(f"/tmp/teardown_{scenario.name.replace(' ', '_')}.ok", "w") as f:
f.write("clean")
# steps/slow_steps.py
from behave import given
import time
@given("the scenario exceeds the shard timeout")
def step_impl(context):
time.sleep(120) # force SIGTERM from a 60s shard timeout
Set your GitHub Actions job timeout to 1 minute, run the shard, then check for the sentinel file. It won't be there. That absence is your proof of drop. The same pattern works in Cucumber-JVM 7 using a JVM shutdown hook outside the Cucumber lifecycle:
// In your test runner or a dedicated lifecycle class
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
Path sentinel = Path.of("/tmp/jvm_shutdown_marker");
try { Files.writeString(sentinel, Instant.now().toString()); }
catch (IOException ignored) {}
}));
If the shutdown hook fires but the Cucumber @After sentinel does not, you have confirmed the framework hook was dropped by the kill signal. With that baseline established, the fix has two layers. First, move critical teardown out of the BDD hook and into a JVM shutdown hook or Python atexit handler for anything that must run regardless of how the process ends:
# Python: register cleanup outside Behave's lifecycle
import atexit
def _emergency_teardown():
# revoke tokens, close DB connections, release locks
cleanup_shared_resources()
atexit.register(_emergency_teardown)
Second, add a CI-level health check that scans for orphaned resources after each shard completes. A simple shell step in your GitHub Actions workflow catches the gap:
# .github/workflows/test-shards.yml (excerpt)
- name: Verify teardown sentinels
if: always()
run: |
missing=$(find /tmp -name "teardown_*.ok" -newer /tmp/shard_start | wc -l)
expected=${{ env.SCENARIO_COUNT }}
if [ "$missing" -lt "$expected" ]; then
echo "::warning::Shard teardown incomplete — $missing of $expected sentinels found"
fi
Teams that added this check to a 400-scenario suite running across 8 shards found that roughly 3–5% of after-hooks were being silently dropped per week — not enough to fail a build, but enough to accumulate 20–30 orphaned DB rows per day. After moving token revocation and connection pool release to atexit, that number dropped to zero without touching shard timeout values. Also worth noting: shard rebalancing shifts which agent runs which scenario, which can redistribute the timeout exposure unevenly — a scenario that ran safely on a fast agent may hit the ceiling when reassigned to a slower one.
Where Senior Engineers Still Get Caught
The most common mistake is treating the BDD after-hook as the single source of truth for teardown. This works perfectly in local development and in sequential CI runs where the process exits cleanly. It breaks specifically under sharded, time-bounded execution — a condition that only exists in production CI pipelines. Engineers who wrote the hooks in a pre-sharding era haven't revisited the assumption because the hooks appear to work in every run that doesn't time out.
A subtler problem is hook scope creep: after-hooks that were originally written for one tag context accumulate logic over time until they're doing work that belongs to suite-level teardown. When a shard timeout drops that hook, the blast radius is larger than anyone realizes. This is the same org-level drift that causes tag inheritance to silently widen hook scope — the hook does more than its name suggests, and no one audits it until something breaks. Audit your after-hooks annually; if any single hook is doing more than three distinct teardown operations, it's overloaded and the risk surface is too wide.
Myths That Keep Teams Exposed
Myth 1: "If the CI job fails with a timeout, no test results are recorded, so no damage is done." This conflates reporting with side effects. The test results may be absent, but the fixtures the scenario was setting up — database rows, S3 objects, auth tokens, feature flag overrides — are still live. The damage is environmental, not in the JUnit XML. Myth 2: "Increasing the shard timeout is the fix." It delays the problem. A scenario that takes 90 seconds today will take 120 seconds after the next dependency upgrade. Timeout tuning is a symptom treatment; defensive teardown registration is the cure.
Myth 3: "This only matters for integration tests." Browser-based scenarios in Playwright or Selenium 4 that manage their own session state are equally exposed. A Playwright browser context that isn't closed leaks a headless Chromium process on the agent. Run 50 shards a day with a 3% drop rate and you're accumulating zombie processes that degrade agent performance over a sprint. Scenario-level failure rates that trend upward without obvious cause are often a signal of this kind of environmental contamination — the failures look random because the root cause is non-deterministic resource exhaustion, not code defects.
The immediate next step is instrumentation: add sentinel files or a lightweight telemetry counter (OpenTelemetry works well here) to your after-hooks, run your full shard suite, and measure actual drop rate. If it's above 1%, move critical teardown to atexit or JVM shutdown hooks before the next sprint. Once you've closed the gap, the next thing worth measuring is mean-time-to-detect on environment contamination — how long between a dropped hook and the first downstream failure it causes.
Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.