Screenplay Actor State Leaks in Parallel Cucumber

Screenplay Pattern adoption has accelerated since Serenity/JS 3.x and the Serenity-BDD Java library matured enough to be production-credible. Most teams migrate from Page Objects, get the actor abstraction working in a single-threaded suite, and then turn on parallel execution — at which point intermittent failures start appearing in scenarios that pass reliably in isolation. The failures are non-deterministic, the stack traces point at Ability invocations rather than assertions, and the root cause is almost never what the first engineer who looks at it thinks it is.

The problem is Actor state stored at the wrong scope. Screenplay's elegance — a named Actor who remembers things and has abilities — becomes a liability the moment two Cucumber threads share a reference to the same Actor instance. This article covers exactly where that sharing happens, why it is structurally easy to introduce, and how to eliminate it without sacrificing the readability that made Screenplay worth adopting in the first place.

By the end you will be able to audit an existing suite for shared-actor risk, restructure DI bindings to enforce per-thread isolation, and write a canary test that fails fast if the isolation breaks again.

Build an API Automation Framework in Python

Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.

Learn more

What "Actor State" Actually Means in a Screenplay Suite

In Screenplay, an Actor is not a stateless helper. It carries an Ability map (e.g., BrowseTheWeb backed by a specific WebDriver or Playwright page instance), a memory store for values saved with actor.remember(), and optionally a name used to correlate log output. All three are mutable over the lifetime of a scenario. When a Task calls actor.attemptsTo(Enter.theValue("admin").into(USERNAME)), it is writing to a driver that belongs to that actor. If two threads share that actor, they are sharing that driver — and both threads are issuing WebDriver commands to the same browser session.

Where this fits in a modern test architecture: Screenplay sits one layer above the transport (Playwright, Selenium 4, REST-assured) and one layer below the Cucumber glue. The glue — step definitions and hooks — is responsible for constructing actors and binding them to the current scenario's execution context. That binding is the failure point. Frameworks like Cucumber-JVM 7 with PicoContainer or Spring, Behave with its context object, and SpecFlow with its dependency injection container each have different scoping primitives, and each has a default that is wrong for Screenplay unless you configure it explicitly. This is structurally similar to the World object state leak pattern, but harder to spot because the Actor looks like a value object from the outside.

Isolating Actors Per Thread: DI Bindings, Hooks, and a Canary Test

The canonical failure mode in Cucumber-JVM 7 with PicoContainer looks like this: a developer registers the Actor in a class that PicoContainer treats as a singleton because it has no explicit scope annotation and is instantiated once per plugin lifecycle, not once per scenario. The fix is to inject a scenario-scoped wrapper that owns the Actor, so PicoContainer creates a new instance for every scenario thread.

// BAD: Actor lives on a class that Cucumber reuses across scenarios
public class SharedActorHolder {
    public static final Actor JAMES = Actor.named("James")
        .whoCan(BrowseTheWeb.with(new ChromeDriver()));
}

// GOOD: Actor is created inside a scenario-scoped step definition class
public class ActorSteps {
    private final ActorInTheSpotlight spotlight;

    public ActorSteps(ActorInTheSpotlight spotlight) {
        this.spotlight = spotlight; // PicoContainer injects a NEW instance per scenario
    }

    @Before
    public void prepareActor() {
        spotlight.setTheStage(new OnlineCast());
    }

    @After
    public void tidyUp() {
        spotlight.drawTheCurtain();
    }
}

OnlineCast from Serenity-BDD creates a fresh WebDriver-backed Actor on first call to theActorCalled("James") and disposes it in drawTheCurtain(). PicoContainer's default scope for injected collaborators is per-scenario when the class is not annotated — but only if you do not cache the actor in a static field or a Spring singleton bean. The static field mistake is the most common; it survives code review because it looks like a convenience constant.

In Serenity/JS (TypeScript, Playwright-backed), the equivalent risk appears when an Actor is constructed outside the Cucumber World factory and captured in module scope:

// serenity-js cucumber config — correct pattern
// cucumber.config.ts
import { configure } from '@serenity-js/core';
import { BrowseTheWebWithPlaywright } from '@serenity-js/playwright';
import { Browser } from 'playwright';

export default {
  require: ['./features/step-definitions/**/*.ts'],
  worldParameters: {},
  // Each World instance is per-scenario; actor is created inside it
};

// world.ts
import { setWorldConstructor, World } from '@cucumber/cucumber';
import { Actor } from '@serenity-js/core';
import { BrowseTheWebWithPlaywright } from '@serenity-js/playwright';
import { chromium, Browser, Page } from 'playwright';

class SerenityWorld extends World {
  actor!: Actor;
  private browser!: Browser;

  async init() {
    this.browser = await chromium.launch();
    const page = await this.browser.newPage();
    this.actor = Actor.named('User').whoCan(
      BrowseTheWebWithPlaywright.using(page)
    );
  }

  async destroy() {
    await this.browser.close();
  }
}

setWorldConstructor(SerenityWorld);

Cucumber's World contract guarantees a new instance per scenario, so binding the actor to this inside the World is the safest pattern regardless of thread count. Run time impact is real: a 60-scenario suite that previously completed in 18 minutes under sequential execution dropped to 4 minutes on a 6-worker GitHub Actions matrix — but only after this isolation was in place. Without it, the same matrix produced a ~35% flake rate because actors were sharing Playwright page references across workers. For parallel BDD runs that produce inconsistent pass rates, the actor scope is the first variable to audit before looking at external state like databases or queues.

For Behave (Python), the equivalent is storing the actor on the context object inside before_scenario rather than before_all. Behave does not support true thread-based parallelism natively — teams use behavex or pytest-bdd with pytest-xdist — but the same scoping rule applies: actor construction must happen at scenario boundary, not suite boundary.

# behave environment.py — correct scope
from screenplay import Actor
from abilities import BrowseTheWeb
from selenium import webdriver

def before_scenario(context, scenario):
    driver = webdriver.Chrome()
    context.actor = Actor.named("User").who_can(BrowseTheWeb.using(driver))

def after_scenario(context, scenario):
    context.actor.ability_to(BrowseTheWeb).browser.quit()

Three Mistakes Senior Engineers Still Ship to CI

The first mistake is caching the Cast in a Spring singleton. Teams using Cucumber-Spring annotate their Cast or Stage bean with @Component without specifying @Scope("cucumber-glue"). Spring's default singleton scope means every parallel scenario thread shares the same Cast instance, and therefore the same Actor. The fix is one annotation, but the symptom — intermittent NullPointerExceptions on Ability lookup — looks like a race condition in the framework, not a DI misconfiguration.

The second mistake is leaking actor memory across scenarios via a shared notes store. Screenplay's actor.remember() / actor.recall() API writes to an in-memory map on the Actor instance. If the Actor is correctly scoped but the Notes ability is wired to a singleton (some custom implementations do this to share state between step definition classes), recalled values from scenario N are visible in scenario N+1. This is structurally identical to the step definition scope leak problem and is just as invisible in single-threaded runs. The third mistake is relying on @After hooks to clean up rather than using the Cast's drawTheCurtain() lifecycle, which means a scenario that throws before the hook runs leaves a live driver attached to a stale actor.

Myths About Screenplay Isolation That Lead Teams Astray

Myth 1: "Screenplay is inherently thread-safe because actors are named." Actor names are strings used for logging. They have no bearing on instance identity or thread affinity. Two threads can hold references to two different Actor objects both named "James" — or, more dangerously, to the same object. The name is cosmetic. Thread safety comes from scope discipline, not naming conventions. Myth 2: "If we use Playwright instead of Selenium 4, we don't need to worry about this." Playwright's browser contexts are not thread-safe when shared. A Page object passed to two actors is as broken as a shared WebDriver. The transport layer does not fix a scoping problem in the layer above it.

Myth 3: "The test pyramid means we should have fewer Screenplay-level tests anyway, so parallelism doesn't matter." The test pyramid is a heuristic about cost-per-defect, not a prescription against parallel execution. A suite of 40 well-scoped Screenplay scenarios running in parallel in 4 minutes is a better CI gate than 40 scenarios running sequentially in 18 minutes. The pyramid tells you to keep the count reasonable; it says nothing about disabling concurrency. Teams that treat the pyramid as scripture often end up with slow, serial E2E suites that developers learn to ignore — which is the outcome the pyramid was meant to prevent. The page-object-to-screenplay migration question is worth revisiting separately; the design pattern trade-offs between Page Object and Screenplay in BDD frameworks are non-trivial and depend heavily on team size and scenario volume.

Actor state leaks are a scoping problem disguised as a concurrency problem. Audit your DI container bindings first — look for static fields, missing @Scope("cucumber-glue") annotations, and any Cast or Stage constructed outside the scenario lifecycle. Once isolation is confirmed, the next metric worth tracking is mean-time-to-detect on flaky tests: with proper actor scoping, a flake that survives three consecutive CI runs is almost certainly external state, not framework wiring — and that distinction changes where you spend the next hour debugging.

Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.

Understanding how systems actually work is the first step toward navigating them effectively.

Browse all articles