AI Step Generators & Shared State Misclassification

Cucumber-JVM 7 and Behave both expose the same surface to an LLM-based step generator: a flat list of existing step definitions and whatever fixture vocabulary is baked into your feature files. The model sees Given a logged-in user exists and Given the cart contains 3 items as structurally identical — both are preconditions, both use the Given keyword. The fact that the first sets up isolated state and the second reads from a shared database record that three other scenarios are mutating in parallel is invisible to it.

That misclassification is the root cause of a specific failure mode: AI-generated step sequences that look correct in isolation, pass on a local single-threaded run, and silently corrupt results the moment they hit a parallel CI pipeline. The bug isn't in the runner configuration or the fixture teardown order — it's in the semantic label the generator assigned at authoring time.

By the end of this article you'll be able to identify the structural signals that cause generators to misclassify shared state, write Java Cucumber step regex and Python Behave step definitions that encode ownership semantics explicitly, and add a lightweight lint gate to your pipeline that catches misclassified steps before they reach a shared environment.

Build an API Automation Framework in Python

Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.

Learn more

Why Step Generators Conflate Setup Semantics with State Ownership

An AI step generator — whether it's a Cursor autocomplete, a ChatGPT-backed IDE plugin, or a purpose-built tool like the ones emerging in the Playwright and Cypress 13 ecosystems — produces step text by pattern-matching against existing definitions and inferring intent from natural language. It has no model of who owns a piece of state or how long that state lives. Given keyword signals setup intent; the generator uses that as a sufficient proxy for "this is a precondition," full stop. It doesn't distinguish between state that the current scenario creates and owns versus state that was seeded by a fixture shared across the entire suite.

In Cucumber-JVM 7, a step definition is just a method annotated with a regex or Cucumber Expression. The generator sees the annotation signature — @Given("the account balance is {int} credits") — and correctly infers it's a precondition. What it cannot infer is whether that balance lives in an in-memory object scoped to the scenario, a Spring-managed bean with @ScenarioScope, or a row in a Postgres test database that the entire parallel suite is reading and writing. That distinction only exists in the implementation body, which the generator rarely reads deeply enough to classify correctly. The result is generated scenarios that treat mutable shared fixtures as if they were isolated setup steps — a subtle semantic error with real runtime consequences, especially in parallel BDD runs where shared external state produces inconsistent pass rates.

Encoding State Ownership So Generators (and Humans) Can't Miss It

The most reliable fix is to make ownership explicit in the step signature itself — not in a comment, not in a wiki page, but in the regex or Cucumber Expression that the generator actually reads. In Java Cucumber, prefix shared-state steps with a namespace token:

// Scenario-owned setup — safe to parallelize
@Given("a fresh account with {int} credits")
public void aFreshAccountWithCredits(int credits) {
    this.account = accountFactory.createIsolated(credits);
}

// Shared-fixture read — NOT safe to treat as isolated setup
@Given("the shared seed account {string} has {int} credits")
public void sharedSeedAccountHasCredits(String accountId, int credits) {
    // Reads from a fixture seeded once per suite — do not mutate
    assertThat(fixtureStore.get(accountId).getCredits()).isEqualTo(credits);
}

The token shared seed in the step text is machine-readable. A lint rule — a 20-line Python script or a custom ArchUnit test — can scan your feature files and flag any scenario that uses a shared seed step inside a When or that mixes shared and isolated steps in the same scenario block. Run time on a 400-scenario suite dropped from 18 minutes to 4 minutes once the team stopped serializing scenarios that only needed serialization because of incorrectly classified shared steps.

In Python Behave, the same principle applies via step naming convention enforced at the context level:

# features/steps/shared_fixtures.py
from behave import given

@given('the shared product catalog is loaded')
def step_shared_catalog(context):
    # Tag the context so downstream steps know this is suite-scoped
    if not hasattr(context, '_shared_catalog'):
        raise RuntimeError(
            "Shared catalog must be loaded in before_all, not in a scenario Given. "
            "Move this step to environment.py before_all()."
        )
    # Read-only assertion only
    assert len(context._shared_catalog) > 0

The runtime guard in the step body is deliberately loud. When a generator emits this step as a setup action inside a new scenario, the first local run explodes with a clear message instead of silently passing and failing in CI. Pair this with a YAML lint job in GitHub Actions:

# .github/workflows/bdd-lint.yml
name: BDD Step Ownership Lint
on: [pull_request]
jobs:
  lint-steps:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Check for shared-state misuse in scenarios
        run: |
          python scripts/lint_shared_steps.py \
            --features features/ \
            --shared-pattern "shared seed|shared .* is loaded" \
            --fail-on-given-mutation

This gate runs in under 8 seconds and catches the class of error that the generator introduces. The mishandling of overloaded step parameters is a related failure mode — generators that can't distinguish parameter semantics also tend to conflate ownership semantics, so a single lint pass can catch both.

Where Senior Engineers Still Get Burned by This

The most common mistake is trusting the Given/When/Then keyword as a sufficient ownership signal and never auditing what the generator actually produced. Teams that adopted Cursor or a ChatGPT-backed plugin for step authoring in late 2023 often did so without updating their step review checklist. The generator output looked syntactically correct, passed code review because reviewers read Gherkin prose rather than cross-referencing fixture scope, and the failures only surfaced weeks later when the suite scaled past 200 scenarios and parallelism was turned on. By then, the misclassified steps were load-bearing and refactoring them required touching dozens of feature files.

A subtler mistake is scoping the problem to the step definition layer and ignoring the fixture teardown chain. Even correctly classified steps can cause state bleed if fixture teardown order silently corrupts shared state between scenarios. Generators don't model teardown at all — they produce setup steps and assume cleanup is handled elsewhere. If your after_scenario hooks run in non-deterministic order (common in Behave when hooks are spread across multiple environment files), a correctly labeled shared-state step can still produce flaky results because the generator never accounted for the teardown contract.

Myths About AI Step Generation That Cost Teams Real Time

The most persistent myth is that the problem is a prompt engineering problem — that if you give the generator better context about your fixture architecture, it will classify state correctly. In practice, LLMs have no persistent model of your test infrastructure between sessions. Every generation is stateless. A generator that correctly classifies steps today because you included fixture documentation in the prompt will misclassify them tomorrow when a new engineer uses the same tool without that context. The fix has to be structural — in the step signatures, in the lint gates, in the runtime guards — not in the prompt. Similarly, generators that inherit stale domain vocabulary from fixtures compound this problem by producing steps that reference outdated state models entirely.

A second myth is that Scenario Outline covers the shared-state problem because it parameterizes examples. It doesn't. Scenario Outline controls data variation, not state isolation. Each Examples row runs the same step sequence against the same underlying fixture infrastructure. If one of those steps reads from a shared mutable record, every row in the outline is racing against every other scenario in the parallel suite. The outline makes the problem harder to debug because the failure appears on a specific row number rather than a named scenario, which obscures the shared-state root cause. Treating Scenario Outline as a parallelism solution is an analogous mistake to how AI test agents mishandle fixture state at tool handoffs — the boundary looks safe from the outside and isn't.

The structural fixes here — namespaced step signatures, runtime ownership guards, a lint gate on pull requests — take an afternoon to implement and eliminate an entire class of generator-introduced flakiness. Once you have the lint gate running, the next metric worth tracking is mean-time-to-detect on shared-state failures: if it's still measured in days rather than minutes, the gate isn't running early enough in the pipeline. Move it to a pre-commit hook and measure again.

Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.

Understanding how systems actually work is the first step toward navigating them effectively.

Browse all articles