fixing-flaky-tests — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited fixing-flaky-tests (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Three non-negotiables, in order:
Only fall back to analytical fixes when the escalation ladder below is exhausted.
Size N to the observed failure rate.
Before any of these: measure, don't assume. Flaky-vs-deterministic, and the rate, are facts to establish from verifiable GitHub run data (step 1) — never inherited from a Slack alert, a teammate's guess, or a ci:insights label.
For triaging a red CI run (finding and classifying the failure), use the debugging-ci-failures skill first — this skill takes over once the failure is classified as a flaky test. For writing new Playwright tests that aren't flaky, use the playwright-test skill.
The GitHub Actions API (or GitHub MCP) is the source of truth. hogli ci:insights is a digest, not an oracle — it can mislabel flaky-vs-deterministic, misstate the rate, and lag the API. Use it to _validate_ a hypothesis or pull historical context, never as the first move or the classification authority.
Establish the timeline yourself from raw run data:
# Pass/fail history for the workflow on the branch under suspicion (master shown):
gh run list --repo PostHog/posthog --workflow=ci-backend.yml --branch master \
--status completed --limit 60 --json conclusion,headSha,createdAt,databaseIdThe run-level conclusion is not enough — a red run may be failing on a _different_ job/test. Confirm it is the same test each time by reading the failing shard's log; and for green runs, confirm the test actually ran and passed in that shard (it may have been sharded elsewhere, not fixed):
gh api repos/PostHog/posthog/actions/jobs/{job_id}/logs \
| grep -E "<exact::test::id>|short test summary"Read the timeline before you classify:
p^30 ≈ 0. That is a deterministic regression: go to debugging-ci-failures and find the introducing commit (step 4).If a failure is reported as (or you suspect it is) consistent, don't serialize — measure the rate and attempt a repro in parallel.
Record the measured rate (failures / total runs, from the run data). You need it to size the validation loop in step 6.
Then confirm it is not already handled:
git log --oneline -10 -- <test_file_path> # recently fixed already?
hogli ci:insights search "<test name or error>" # historical context / existing fix — corroborate against the run data, do not trust blindlyAn insight with a merged fix means the flake may already be resolved — read the fix, confirm _against the run data_ that it covers this failure, and report instead of re-fixing.
gh run view <run-id> --log-failed
# If the job was re-run and the latest attempt passed, the flaky failure
# lives in a previous attempt:
gh api "repos/PostHog/posthog/actions/runs/{run_id}/attempts/1/jobs"
gh api "repos/PostHog/posthog/actions/jobs/{job_id}/logs"Capture before moving on:
[MSW] Unhandled, async leak warnings, teardown errors — these are often the actual cause, printed before the symptom.Escalate through these conditions until the failure appears. Stop at the first level that reproduces it; that level is your validation environment for step 6.
hogli test <path>::<test> — confirms the test runs at all.frontend/package.json's test script and .github/workflows/ci-frontend.yml before running.Example (with the flags as of this writing), running the test alongside its shard neighbors:
pnpm --filter=@posthog/frontend jest <test_file> <neighbor_file> --maxWorkers=2 --forceExitFor pytest, run the whole file or class rather than the single test, so module-level fixtures and ordering match CI.
Flakes that vanish in isolation are ordering bugs.
Timeout-class flakes often only show here.
The loop harness — judge by exit code, not by grepping output:
N=20; PASS=0; FAIL=0
for i in $(seq 1 $N); do
if <test command> >/tmp/flake-run.log 2>&1; then
PASS=$((PASS+1))
else
FAIL=$((FAIL+1)); cp /tmp/flake-run.log /tmp/flake-fail-$i.log; echo "run $i: FAIL"
fi
done
echo "$PASS passed, $FAIL failed out of $N"Two cost notes for the loop:
break after the first failure — one captured failure log is enough.Complete all N runs only when measuring the failure rate or validating in step 6.
pnpm --filter=@posthog/frontend jest script runs pnpm build:products before every invocation.Inside a loop, build once, then iterate with pnpm --filter=@posthog/frontend exec jest ..., which skips the rebuild.
If nothing reproduces after the full ladder, the flake is CI-environment-specific. Proceed with a fix grounded in the CI evidence and root-cause analysis, and say so explicitly in the report — the validation in step 6 is then analytical, not empirical.
Match the symptom to a cause class; never patch the symptom.
| Symptom | Likely cause class |
|---|---|
| Timeout waiting for promise/listener/element | Unawaited async work, missing mock, hidden pending request |
| Passes alone, fails with neighbors (or vice versa) | Shared state: module cache, DB rows, global config, ordering |
| Fails near midnight/UTC boundaries, or on slow runners | Real clock usage — missing freeze_time / fake timers |
| Assertion on list order or generated IDs | Nondeterministic ordering/IDs asserted as deterministic |
| Query can't see just-written data | Eventual consistency (ClickHouse), missing flush/commit |
Only fails under --maxWorkers=2 / contention | Race condition surfaced by scheduling, too-tight timeout |
If the symptom table doesn't point at a clear cause and the test file itself is unchanged (git log -- <test_file> is stale), the trigger is elsewhere — a neighbor test, a dependency bump, or a product change. Find _when_ it started instead of guessing:
git log <good>..<bad>). That short list often names the culprit outright. git bisect start <bad-sha> <good-sha>
git bisect run bash -c '<repro command>' # exit 0 = good, non-zero = badCaveat: for an intermittent flake, a _lucky pass_ at a step sends git bisect down the wrong path. Trust code-bisect only when the failure is deterministic; otherwise run the repro N times per step (fail if any iteration fails), or just use the CI run-history boundary.
PostHog-specific patterns:
afterMount loaders fire API calls through the connect() chain — a logic three levels deep can trigger an unmocked fetch.Unhandled requests currently resolve with a benign empty paginated 200, so the symptom is a loader succeeding with empty or wrong data, not a network error. The [MSW] Unhandled GET ... warning in the log names the missing mock — add the useMocks entry. The unhandled-request behavior has changed before (it used to hang); if symptoms don't match, read frontend/src/mocks/jest.ts for what unmocked requests do today.
LISTENER_FINISH_WAIT_TIMEOUT in kea-test-utils).Any connected logic with a pending loader blocks it. Fix the pending work; do not raise the timeout.
mocksToHandlers strips trailing slashes, but query params, @current-style segments, and :param patterns must match the real request URL.Compare against the [MSW] Unhandled line.
frontend/src/mocks/jest.ts registers a global afterEach(() => mswServer.resetHandlers()).Each it.each case is a separate test, so runtime mocks must be (re-)registered in beforeEach.
beforeEach and never unmounted leak async work into later tests.@pytest.mark.django_db(transaction=True).freeze_time; never assert on now()-derived values.| Tempting masking move | Do instead |
|---|---|
sleep(2) / setTimeout before asserting | Await the specific condition (waitFor, expectLogic, explicit flush) |
| Raise the test/listener timeout | Find what is hanging; the timeout is the messenger |
Add retries (pytest-rerunfailures --reruns, jest.retryTimes) | Reserve for genuinely nondeterministic external infra, with a comment and a linked issue — never for product code under test |
| Skip / quarantine the test | Only with explicit user approval, with a linked issue |
| Loosen the assertion | Make the data deterministic (sort, freeze, seed), keep the assertion strict |
Keep the fix minimal and inside the test or its fixtures when possible. If the race is in product code, the flake found a real bug — fix the product code and say so in the report.
Run the step-3 harness on the fixed code under the same conditions that reproduced the failure (same neighbors, worker count, contention).
Size N from the observed pre-fix failure rate: if it failed about 1 in k runs, you need roughly N ≥ 3k consecutive passes for ~95% confidence the flake is gone — (1 - 1/k)^(3k) ≈ 5%. So use N = max(3k, 20). Without a usable rate estimate, run 50 and note the reduced confidence in the report. If the flake was never reproducible locally, run N = 20 as a regression check and label the validation as analytical.
Any failure in the loop → back to step 4; the root cause was wrong or incomplete. Finish with one normal run of the surrounding file/suite to confirm the fix didn't break sibling tests.
Test: <file path>::<test name>
Observed in CI: <measured rate from run data, e.g. 8/45 runs over 3h (gh run list); ci:insights corroborates>
Local repro: <command + conditions, e.g. 3/20 failures with neighbor X, maxWorkers=2 | not reproducible locally>
Root cause: <one or two sentences>
Fix: <what changed and why it removes the cause>
Validation: <N>/<N> passes under repro conditions | analytical only (CI-specific)
Follow-ups: <product bug found, related tests with the same pattern, or none>.github/workflows/ as part of a flake fix.~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.