Flaky Tests Are Lying to You: How We Help Clients Build Stable Automation
By Eduard Dublyer, CTO of Skipper Soft
“It worked on my machine.”
That’s usually when I know we’re dealing with a flaky test.
As CTO at Skipper Soft, I’ve worked with dozens of companies — from early-stage startups to enterprise platforms — helping them stabilize their automated test suites. And in nearly every case, the same root problem appears: tests that sometimes pass, sometimes fail, and always drain the team’s confidence.
The automation isn’t broken — it’s just flaky.
In this article, I want to share five common patterns that lead to flaky tests and show how we’ve helped clients eliminate them using Python, Pytest, and Playwright.
1. Sleep Is Not a Synchronization Strategy
One of our fintech clients was running Playwright-based tests to validate PDF report downloads. Occasionally, the test would fail even though the report was generated.
Here’s what their original test looked like:
def test_report_download():
page.click("text='Download Report'")
time.sleep(5) # This was the problem
assert os.path.exists("downloads/report.pdf")
The test relied on a hardcoded sleep, assuming the report would always be ready in 5 seconds. But in production environments under load, 5 seconds wasn’t always enough — and that’s how the test failed randomly.
The Fix: Active Waiting
We replaced the hard sleep with a proper wait function that actively checks for the file:
import time
from pathlib import Path
def wait_for_file(filepath, timeout=10):
start = time.time()
while time.time() - start < timeout:
if Path(filepath).exists():
return True
time.sleep(0.5)
raise TimeoutError(f"{filepath} not found after {timeout} seconds")
Then the test became:
def test_report_download():
page.click("text='Download Report'")
assert wait_for_file("downloads/report.pdf")
This one change eliminated the intermittent failures entirely.
2. Every Test Deserves Its Universe
In one project, the test suite relied on implicit ordering and shared state between tests:
def test_create_user():
create_user("alice")
def test_user_login():
assert login("alice", "password123") # Depends on test_create_user
When tests ran out of order or in parallel, failures appeared.
The Fix: Use Fixtures and Mocks
We redesigned the suite using Pytest fixtures and isolation principles:
import pytest
@pytest.fixture
def mock_user(monkeypatch):
monkeypatch.setattr("app.db.get_user_by_name", lambda name: {"name": name, "role": "tester"})
def test_user_login(mock_user):
user = app.db.get_user_by_name("alice")
assert user["role"] == "tester"
Each test runs in its environment, independent of test order or external setup.
3. Don’t Trust the Outside World
A retail client had tests that sometimes failed on Fridays. The cause? Their product recommendations adjusted based on real-time weather data from an external API. If the API was down or slow, the UI changed — and the test failed.
The Fix: Mock External APIs
Using Playwright’s route interception, we created a consistent response:
def test_weather_mock(playwright):
browser = playwright.chromium.launch()
context = browser.new_context()
page = context.new_page()
page.route("**/weather", lambda route, request: route.fulfill(
status=200,
content_type="application/json",
body='{"temp": 22}'
))
page.goto("https://shop.com")
assert page.locator(".recommendations").is_visible()
This made the UI predictable and the test stable.
4. Randomness Is Not for Testing
A SaaS platform we worked with had a test that occasionally failed because of randomness in OTP generation:
import random
def test_otp_length():
code = str(random.randint(0, 999999))
assert len(code) == 6
Sometimes the test failed with values like 123 or 42 — still valid integers, but not six digits long.
The Fix: Control the Random
def test_otp_generation(monkeypatch):
monkeypatch.setattr("random.randint", lambda a, b: 123456)
assert str(random.randint(0, 999999)) == "123456"
In testing, predictability is everything. Don’t let randomness undermine your confidence in the suite.
5. Retrying Is a Patch, Not a Strategy
Several teams we’ve worked with applied reruns as a band-aid to keep pipelines green. Pytest even has native support for this:
# pytest.ini
[pytest]
addopts = –reruns 2 –reruns-delay 1 import pytest @pytest.mark.flaky(reruns=2) def test_intermittent_case(): assert perform_operation()
This may help temporarily, but rerunning flaky tests only masks the issue. We recommend using retries to buy time, not as a long-term solution. I saw many instances when the first run was an actual bug, and it created a false positive situation during the retry.
Summary: Five Practices to Eliminate Flaky Tests
Final Thoughts
At Skipper Soft , our mission is to help teams build automation they can trust. Flaky tests don’t just waste time — they erode confidence, slow releases, and delay product value.
Reliable automation is possible. It starts with discipline, a good testing strategy, and an understanding of the real-world behaviors of the systems we test.
If your test suite has become unreliable or you want a second set of expert eyes on it, we’re here to help.
You can visit skipper-soft.com or reach out directly. Let’s make your automation stable, predictable, and CI/CD-ready.