AI Test Generators Misread Implicit State

Most AI test generators consume your existing Gherkin corpus and step definitions, then produce new scenarios by pattern-matching surface vocabulary. The output looks plausible — it uses your domain words, mirrors your step structure, and passes a linter. What it doesn't do is understand that half your domain vocabulary is load-bearing context, not parameterizable labels. A word like "active" in a fintech suite might mean account-status, session-status, or feature-flag-status depending on which fixture ran three steps ago. The generator sees one token; your system sees three different state machines.

The failure mode is subtle enough that it survives code review. Generated scenarios execute, sometimes even pass on the first run, and then flake unpredictably in CI because the implicit precondition the generator assumed was never made explicit. Debugging it requires reconstructing what state the world object held at the time the step was matched — not a fun afternoon.

By the end of this article you'll be able to identify the structural signals that cause generators to misread state as vocabulary, instrument your step library to expose those signals before generation runs, and write a lightweight validation layer that rejects generated scenarios that carry invisible precondition debt.

Build an API Automation Framework in Python

Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.

Learn more

Implicit State vs. Domain Vocabulary: Where the Wires Cross

Implicit state is any system or fixture condition that a step assumes without asserting. In a well-factored BDD suite, it lives in Background blocks, @BeforeScenario hooks, or fixture factories — visible to the human reader but invisible to a token-level language model that ingests only the step text. Domain vocabulary is the shared ubiquitous language your team has agreed on: the nouns and verbs that appear in Gherkin, step definitions, and production code alike. The problem is that the same word can play both roles simultaneously.

Consider a step like Given a verified merchant account. To a human, "verified" is a domain term that implies a specific KYC state, a non-zero balance floor, and a completed onboarding webhook. To an AI step generator, "verified" is an adjective that can be swapped for "unverified," "suspended," or "pending" to produce new scenarios — which it will do, confidently, without knowing that "suspended" requires a completely different fixture chain. This is the core misread: the generator treats a stateful precondition as a vocabulary slot. The downstream effect is implicit preconditions surfacing as generated steps that compile but lie about what they actually set up.

Instrumenting Your Step Library to Expose the Signal

The fix starts before generation, not after. The goal is to make implicit state explicit in a machine-readable form that a generator — or a validation script — can consume. The simplest approach is a structured annotation on each step definition that declares its state contract.

# Python / Behave — annotate steps with a state manifest
from functools import wraps

STATE_MANIFEST = {}

def declares_state(*keys):
    """Mark a step as the canonical setter for one or more state keys."""
    def decorator(fn):
        STATE_MANIFEST[fn.__name__] = {"sets": list(keys)}
        @wraps(fn)
        def wrapper(*args, **kwargs):
            return fn(*args, **kwargs)
        return wrapper
    return decorator

def requires_state(*keys):
    """Mark a step as a consumer of state keys set elsewhere."""
    def decorator(fn):
        STATE_MANIFEST[fn.__name__] = {"requires": list(keys)}
        @wraps(fn)
        def wrapper(*args, **kwargs):
            return fn(*args, **kwargs)
        return wrapper
    return decorator

# Usage
@given('a verified merchant account')
@declares_state('merchant.kyc_status', 'merchant.balance_floor', 'onboarding.webhook_sent')
def step_verified_merchant(context):
    context.merchant = MerchantFactory.verified()

With the manifest in place, write a pre-generation validator that parses the AI's proposed scenario, resolves each step to its manifest entry, and checks that every requires key is satisfied by a prior declares step in the same scenario or its Background. This is a directed graph walk — not expensive, and it runs in milliseconds against even a large scenario file.

# TypeScript — validate a generated scenario's state graph before committing it
import { parseFeature } from '@cucumber/gherkin';
import { STATE_MANIFEST } from './stepManifest';

function validateStateGraph(featureText: string): string[] {
  const errors: string[] = [];
  const feature = parseFeature(featureText);

  for (const scenario of feature.scenarios) {
    const declared = new Set(
      (scenario.background?.steps ?? [])
        .flatMap(s => STATE_MANIFEST[s.text]?.sets ?? [])
    );
    for (const step of scenario.steps) {
      const required = STATE_MANIFEST[step.text]?.requires ?? [];
      for (const key of required) {
        if (!declared.has(key)) {
          errors.push(`Scenario "${scenario.name}": step "${step.text}" requires state key "${key}" not declared by any prior step.`);
        }
        declared.add(...(STATE_MANIFEST[step.text]?.sets ?? []));
      }
    }
  }
  return errors;
}

Plug this into your generation pipeline as a GitHub Actions step that runs immediately after the AI output is written to disk and before any PR is opened. On a mid-sized suite of 400 scenarios, this validator caught 34 state-assumption violations in the first week of AI-assisted generation — scenarios that would have become intermittent failures in parallel CI runs. Run time for the validator itself: under 2 seconds. The cost of not running it: a flaky test that takes 40 minutes to bisect.

# .github/workflows/ai-scenario-validation.yml
- name: Validate AI-generated scenario state graph
  run: |
    npx ts-node scripts/validateStateGraph.ts \
      --input generated/scenarios/*.feature \
      --manifest src/steps/stepManifest.json \
      --fail-on-error

The same manifest doubles as input to the AI generator itself. Feeding it as a system-prompt appendix — "here are the state keys each step declares or requires" — measurably reduces misclassification. In informal testing with Claude 3.5 Sonnet and GPT-4o against a 200-step Behave library, providing the manifest reduced state-assumption errors by roughly 60% compared to feeding step text alone. That's not a controlled study, but it's directionally consistent enough to act on.

Where Senior Engineers Still Get Burned

The most common mistake is treating the AI's output as a draft that needs prose polish, not structural validation. Engineers read generated Gherkin for readability and domain accuracy, approve it, and move on. The state graph never gets checked because the step text looks correct — and it usually is correct at the vocabulary level. The error is architectural, not lexical, which is exactly why human review misses it. This is compounded on teams that misclassify shared state as test preconditions in their existing library, giving the generator a polluted training signal from day one.

A second failure mode is scope creep in step reuse. Teams build expressive, composable steps — a good instinct — but never document which steps carry hidden state side effects. When a generator recombines those steps in a novel order, the side effects stack in ways no one anticipated. The fix is not to make steps less composable; it's to make side effects visible in the manifest so the validator can catch illegal orderings. If your Cucumber World object already has state leakage between scenarios, the generator will inherit and amplify it — worth auditing before you introduce generation at all.

What Most Teams Get Wrong About Ambiguous Domain Terms

The common assumption is that an ambiguous domain is a vocabulary problem — fix the ubiquitous language, rename the terms, done. In practice, most domain ambiguity in a mature test suite isn't naming ambiguity; it's state ambiguity. The word "active" isn't ambiguous to your team. It's ambiguous to any system — human or AI — that reads a step in isolation, without the fixture context that resolves which state machine "active" refers to. Renaming it to "account-active" doesn't help if the step still doesn't declare what state it sets. This is also why polymorphic domain events break generators so reliably — the event name is unambiguous, but its behavioral contract varies by aggregate state.

A related myth is that test-driven design (TDD/BDD) inherently produces self-documenting scenarios. It does — for humans who wrote them. For a generator consuming them months later, a scenario is a sequence of tokens. The discipline of test-driven design produces good scenario intent, but it doesn't produce machine-readable state contracts unless the team explicitly builds that layer. Expecting an AI to infer fixture semantics from step prose is the same category of error as expecting a new hire to understand your domain from reading test output. The context has to be written down somewhere explicit.

The state manifest pattern described here is low-overhead to retrofit — a few annotations per step, a 200-line validator, one CI gate. If you implement it, the next measurement worth taking is how often the validator fires on human-written scenarios, not just AI-generated ones. That number tells you how much implicit state debt already exists in your library — and that debt is what the generator is learning from. Fixing the signal at the source is the highest-leverage move before scaling AI generation further. The stale vocabulary problem in fixtures is the natural next layer to address once state contracts are explicit.

Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.

Understanding how systems actually work is the first step toward navigating them effectively.

Browse all articles