Flaky Tests Are Lying to You: How We Help Clients Build Stable Automation

By Eduard Dublyer, CTO of Skipper Soft

“It worked on my machine.”

That’s usually when I know we’re dealing with a flaky test.

As CTO at Skipper Soft, I’ve worked with dozens of companies — from early-stage startups to enterprise platforms — helping them stabilize their automated test suites. And in nearly every case, the same root problem appears: tests that sometimes pass, sometimes fail, and always drain the team’s confidence.

The automation isn’t broken — it’s just flaky.

In this article, I want to share five common patterns that lead to flaky tests and show how we’ve helped clients eliminate them using Python, Pytest, and Playwright.


1. Sleep Is Not a Synchronization Strategy

One of our fintech clients was running Playwright-based tests to validate PDF report downloads. Occasionally, the test would fail even though the report was generated.

Here’s what their original test looked like:

def test_report_download():
    page.click("text='Download Report'")
    time.sleep(5)  # This was the problem
    assert os.path.exists("downloads/report.pdf")

The test relied on a hardcoded sleep, assuming the report would always be ready in 5 seconds. But in production environments under load, 5 seconds wasn’t always enough — and that’s how the test failed randomly.

The Fix: Active Waiting

We replaced the hard sleep with a proper wait function that actively checks for the file:

import time
from pathlib import Path

def wait_for_file(filepath, timeout=10):
    start = time.time()
    while time.time() - start < timeout:
        if Path(filepath).exists():
            return True
        time.sleep(0.5)
    raise TimeoutError(f"{filepath} not found after {timeout} seconds")

Then the test became:

def test_report_download():
    page.click("text='Download Report'")
    assert wait_for_file("downloads/report.pdf")

This one change eliminated the intermittent failures entirely.


2. Every Test Deserves Its Universe

In one project, the test suite relied on implicit ordering and shared state between tests:

def test_create_user():
    create_user("alice")

def test_user_login():
    assert login("alice", "password123")  # Depends on test_create_user

When tests ran out of order or in parallel, failures appeared.

The Fix: Use Fixtures and Mocks

We redesigned the suite using Pytest fixtures and isolation principles:

import pytest

@pytest.fixture
def mock_user(monkeypatch):
    monkeypatch.setattr("app.db.get_user_by_name", lambda name: {"name": name, "role": "tester"})

def test_user_login(mock_user):
    user = app.db.get_user_by_name("alice")
    assert user["role"] == "tester"

Each test runs in its environment, independent of test order or external setup.


3. Don’t Trust the Outside World

A retail client had tests that sometimes failed on Fridays. The cause? Their product recommendations adjusted based on real-time weather data from an external API. If the API was down or slow, the UI changed — and the test failed.

The Fix: Mock External APIs

Using Playwright’s route interception, we created a consistent response:

def test_weather_mock(playwright):
    browser = playwright.chromium.launch()
    context = browser.new_context()
    page = context.new_page()

    page.route("**/weather", lambda route, request: route.fulfill(
        status=200,
        content_type="application/json",
        body='{"temp": 22}'
    ))

    page.goto("https://shop.com")
    assert page.locator(".recommendations").is_visible()

This made the UI predictable and the test stable.


4. Randomness Is Not for Testing

A SaaS platform we worked with had a test that occasionally failed because of randomness in OTP generation:

import random

def test_otp_length():
    code = str(random.randint(0, 999999))
    assert len(code) == 6

Sometimes the test failed with values like 123 or 42 — still valid integers, but not six digits long.

The Fix: Control the Random

def test_otp_generation(monkeypatch):
    monkeypatch.setattr("random.randint", lambda a, b: 123456)
    assert str(random.randint(0, 999999)) == "123456"

In testing, predictability is everything. Don’t let randomness undermine your confidence in the suite.


5. Retrying Is a Patch, Not a Strategy

Several teams we’ve worked with applied reruns as a band-aid to keep pipelines green. Pytest even has native support for this:

# pytest.ini

[pytest]

addopts = –reruns 2 –reruns-delay 1 import pytest @pytest.mark.flaky(reruns=2) def test_intermittent_case(): assert perform_operation()

This may help temporarily, but rerunning flaky tests only masks the issue. We recommend using retries to buy time, not as a long-term solution. I saw many instances when the first run was an actual bug, and it created a false positive situation during the retry.


Summary: Five Practices to Eliminate Flaky Tests

Article content
Five Practices to Eliminate Flaky Tests

Final Thoughts

At Skipper Soft , our mission is to help teams build automation they can trust. Flaky tests don’t just waste time — they erode confidence, slow releases, and delay product value.

Reliable automation is possible. It starts with discipline, a good testing strategy, and an understanding of the real-world behaviors of the systems we test.

If your test suite has become unreliable or you want a second set of expert eyes on it, we’re here to help.

You can visit skipper-soft.com or reach out directly. Let’s make your automation stable, predictable, and CI/CD-ready.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *