Step Definition Scope Leaks in Cucumber

Cucumber-JVM 7 scans the classpath for step definitions at startup and registers every match it finds — across every module, every JAR, every shared library on that path. That design is intentional and, for small monorepos, harmless. At scale, it becomes the source of some of the most confusing test failures you'll see: ambiguous step matches, silent overrides, and hook side-effects that fire in scenarios that never asked for them. The problem isn't unique to JVM; Behave, SpecFlow 3, and Cypress-Cucumber all have analogous registration models with the same failure modes.

The technical problem is scope leakage: step definitions declared in one module bleed into the execution context of another. This happens through shared Gradle/Maven subprojects, fat-test JARs, Python environment.py imports, and SpecFlow's assembly-scanning hooks. The symptoms range from AmbiguousStepDefinitionsException at runtime to subtler issues — a step that matches in module A silently shadows the intended step in module B, and your scenario passes for the wrong reason.

By the end of this article you'll know exactly where scope boundaries are drawn in each major framework, how to detect leaks before they reach CI, and the structural patterns that prevent them. The relevance is immediate: as teams extract shared step libraries into internal packages (a pattern accelerating with monorepo tooling like Nx and Bazel), scope collisions are becoming a first-class architectural problem, not an occasional nuisance.

Build an API Automation Framework in Python

Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.

Learn more

How Cucumber's Step Registry Actually Works Across Modules

In Cucumber-JVM 7, step definitions are discovered via a glue path — a package prefix passed to the runner. Every class reachable under that prefix is scanned for @Given, @When, @Then annotations. If two modules share the same glue path (common when a shared-steps library lives under com.example.steps and so does your feature module), both sets of definitions are loaded into the same registry. There is no module boundary at the registry level; the JVM classloader doesn't care about your Gradle subproject graph. This is also why step definition registries fragment across shared libraries in ways that aren't obvious from the build file alone.

Behave (Python) uses a different mechanism but reaches the same failure mode. It imports all steps/*.py files in the active context directory and any directory listed in paths. SpecFlow 3 scans assemblies referenced in the test project — add a NuGet reference to a shared steps package and its bindings are unconditionally registered. The characteristics and advantages each framework advertises (Cucumber's natural-language flexibility, Behave's zero-config Python-native setup, SpecFlow's deep .NET IDE integration) all rest on this open-registry model. That openness is the feature; the scope leak is the cost.

Detecting and Containing Scope Leaks in JVM, Behave, and SpecFlow

The fastest diagnostic in Cucumber-JVM is a dry run with --dry-run and --plugin json:out/report.json. Parse the JSON for duplicate step text across different source locations. A one-liner in Python does the job:

import json, collections
report = json.load(open("out/report.json"))
steps = [s["name"] for f in report for e in f["elements"] for s in e["steps"]]
dupes = [s for s, c in collections.Counter(steps).items() if c > 1]
print(dupes)

If you see the same step text appearing under two different match.location values, you have a registry collision. Fix it at the glue path level first: give each module a unique package prefix and pass only that prefix to its runner. In Gradle, that means separate CucumberOptions per subproject rather than a shared integration-test runner that wildcards the whole classpath.

// build.gradle (module-specific runner)
test {
    systemProperty "cucumber.glue", "com.example.billing.steps"
    systemProperty "cucumber.features", "src/test/resources/billing"
}

In Behave, the equivalent is an explicit behave.ini per test suite root that restricts the steps path:

[behave]
paths = features/billing
steps_dir = features/billing/steps

Without this, Behave's default discovery walks up to the nearest environment.py and pulls in every sibling steps directory it finds — a common source of cross-domain step collisions in monorepos where multiple product areas share a single repo root. This is a distinct issue from the World object state leakage problem, but the two often co-occur when shared infrastructure is set up carelessly.

For SpecFlow 3, the containment pattern is assembly-level isolation: put shared bindings in a dedicated assembly and use [Binding] attribute filtering with scoped bindings (SpecFlow's Scope attribute on step classes). This restricts a binding class to scenarios tagged with a specific feature or tag, preventing it from matching across the entire test suite:

[Binding]
[Scope(Tag = "billing")]
public class BillingSteps
{
    [Given(@"a customer with balance (.*)")]
    public void GivenCustomerBalance(decimal balance) { ... }
}

The measurable payoff is real. One platform team running ~1,400 Cucumber-JVM scenarios across four Gradle submodules reduced their ambiguous-match failures from 23 per sprint to zero after enforcing per-module glue paths and adding a CI lint step that fails the build if any step text appears in more than one match.location. Their suite initialization time also dropped from 18 seconds to 4 because the registry no longer loaded 600+ irrelevant step definitions for each module's run. As step definition count grows, registry bloat compounds initialization cost in ways that are easy to miss until the number is already large.

Where Senior Engineers Still Get Burned by Scope Leaks

The most common mistake is treating the shared steps library as a convenience rather than a contract. Teams extract common steps into a shared module to reduce duplication — a reasonable goal — but they copy the glue path alongside the code. Now every consuming module loads every shared step, including ones written for a different domain context. The fix isn't to stop sharing; it's to version the shared library properly and expose only the steps each consumer explicitly opts into, using package-level separation. Regex precision matters here too: vague patterns like ^the user (.*)$ in a shared library will match steps that were never intended for it. The Java Cucumber step regex guide covers how to write patterns that are precise enough to survive a multi-module classpath.

The second mistake is discovering scope leaks only in CI. By the time the pipeline reports an AmbiguousStepDefinitionsException, the collision has often existed for weeks — it just wasn't triggered until a new scenario exercised the ambiguous path. Add a pre-commit or local dry-run hook that runs the registry collision check locally. GitHub Actions makes this cheap: a 30-second dry-run job as a required check catches the problem before the branch merges, not after.

Myths About Module Isolation That Lead Teams Astray

Myth 1: Separate Maven/Gradle modules mean separate step registries. They don't. The registry is a runtime construct, not a build-time one. As long as the compiled classes end up on the same test classpath — which they do whenever you declare a testImplementation dependency on another subproject — their steps are all visible to the runner. Build-tool module boundaries give you compilation isolation, not execution isolation. Myth 2: Behave is safer than Cucumber-JVM because Python doesn't have classpath hell. Behave's file-system-based discovery is simpler, but "simpler" doesn't mean "scoped." A monorepo with a shared steps/ directory at the root and multiple feature suites beneath it will exhibit the same collision behavior. The mechanism differs; the failure mode is identical.

Myth 3: Tagging scenarios is enough to control which steps load. Tags control which scenarios execute; they do not filter which step definitions are registered. Every step in every loaded file is in the registry regardless of tags, which is also why tag inheritance silently widens hook scope in ways that surprise teams who assume tags provide execution isolation. The only reliable isolation boundary is the glue path (JVM), the steps directory path (Behave), or the scoped binding attribute (SpecFlow) — not tags, not file naming conventions, not folder structure alone.

Scope leaks are an architectural problem that manifests as a test reliability problem — which is why they're easy to misdiagnose. Start with the dry-run collision check, enforce per-module glue paths in CI, and treat your shared steps library as a versioned API rather than a shared folder. Once registry hygiene is in place, the next thing worth measuring is whether your ambiguous-match rate correlates with step definition count growth over time — that trend tells you whether your module boundaries are holding.

Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.

Understanding how systems actually work is the first step toward navigating them effectively.

Browse all articles