soak-test-99520e — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited soak-test-99520e (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
A soak test (also called an endurance test) is an extended play session run with specific observation goals. Unlike a smoke check (broad critical path, ~10 min) or a single-feature playtest (~30 min), a soak test runs for 30 minutes to several hours to surface:
of a mechanic (inventory full, score overflow, AI state corruption)
repetitive over extended play
This skill generates the observation protocol and analysis harness — the human does the actual playing.
Output: production/qa/soak-test-[date]-[duration].md
When to run:
/gate-check releaseDuration (default: 1h):
30m — short soak; suitable for testing a single mechanic or scene1h — standard soak; covers most common leak categories2h — extended soak; recommended for first full Polish soak4h — deep soak; required for games with long session design (RPGs, sims)Focus (default: all):
memory — focus on heap size, object count, leak patternsstability — focus on crash/freeze/hang detectionbalance — focus on fun fatigue, content exhaustion, difficulty perceptionall — all of the aboveRead:
.claude/docs/technical-preferences.md — engine (for engine-specific memorymonitoring guidance), performance budgets (memory ceiling, target FPS)
design/gdd/game-concept.md — intended session length (for comparison againstsoak duration), core loop description
production/playtests/ — prior playtest findings(to avoid re-documenting known issues)
production/qa/qa-plan-*.md — current sprint test coverage(to understand what has been formally tested vs. what the soak covers)
Note any performance budget targets from technical-preferences.md:
Based on duration, generate timed checkpoints:
30m soak: T+0, T+10, T+20, T+30 1h soak: T+0, T+15, T+30, T+45, T+60 2h soak: T+0, T+20, T+40, T+60, T+80, T+100, T+120 4h soak: T+0, T+30, T+60, T+90, T+120, T+180, T+240
At each checkpoint, the observer records the observation items defined in Phase 4.
Engine-specific monitoring guidance:
Godot 4:
Memory → Static Memory andObject Count → Objects across checkpoints
(some growth on load is expected; sustained growth indicates a leak)
Performance.get_monitor(Performance.MEMORY_STATIC) returns bytesin Godot 4.6
Unity:
Unreal Engine:
stat memory console command at each checkpointAt each checkpoint, note:
Collect subjective observations at each checkpoint:
# Soak Test Protocol
> **Date**: [date]
> **Duration**: [duration]
> **Focus**: [memory | stability | balance | all]
> **Engine**: [engine]
> **Generated by**: /soak-test
---
## Pre-Session Setup
Before starting the soak:
- [ ] Game is running from a **fresh launch** (not resumed from a prior session)
- [ ] All background applications closed (minimise OS memory interference)
- [ ] Performance monitoring tool open and recording:
- **Godot**: Debugger → Monitors tab → Memory section visible
- **Unity**: Memory Profiler window open
- **Unreal**: `stat memory` ready in console
- [ ] Soak target confirmed: [session design intent from game concept]
- [ ] Prior known issues to watch for: [from most recent playtest / qa-plan]
---
## Baseline (T+0) — Record Before Playing
| Metric | Baseline Value |
|--------|---------------|
| Memory / Heap | [record before first frame of gameplay] |
| Object Count | [record] |
| FPS (first 30 seconds) | [record] |
| [Engine-specific metric] | [record] |
---
## Checkpoint Log
### T+[N] minutes
**Memory / Stability** *(if applicable)*:
| Metric | Value | Δ from Baseline | Alert? |
|--------|-------|-----------------|--------|
| Memory / Heap | | | |
| Object Count | | | |
| FPS | | | |
| Crashes / Hangs | | | |
**Stability checks**:
- [ ] No crash or hang since last checkpoint
- [ ] Frame rate within budget ([N] fps target)
- [ ] Audio correct
- [ ] HUD rendering correctly
- [ ] Input responding correctly
**Balance / Fatigue** *(if applicable)*:
- Core mechanic still rewarding: Y / N
- Difficulty perception: too easy / appropriate / too hard
- Notable moments: [note any peak engagement or frustration]
- Content exhaustion signs: Y / N — [describe]
**Free observations**:
*(Note anything unexpected observed since the last checkpoint)*
---
[Repeat Checkpoint Log section for each timed checkpoint]
---
## Post-Session Analysis
### Memory Trend
| Checkpoint | Memory | Δ/hr extrapolated |
|------------|--------|-------------------|
| T+0 | | |
| [T+N] | | |
**Leak detected?** Y / N
**Estimated time to OOM at current rate**: [N hours / not applicable]
### Stability Summary
Total crashes: [N]
Total hangs: [N]
Worst FPS observed: [N] fps at [checkpoint]
Performance degradation: stable / mild / severe
### Balance / Fatigue Summary
Fun curve: [engaged throughout / fatigue onset at T+N / repetitive from start]
Content exhaustion point: [never / at T+N / early]
Difficulty arc: [appropriate / too easy throughout / difficulty spike at T+N]
### Issues Found
| ID | Severity | Checkpoint | Description |
|----|----------|------------|-------------|
| SOAK-001 | S[1-4] | T+[N] | [description] |
---
## Verdict: PASS / PASS WITH CONCERNS / FAIL
**PASS**: No leaks detected, stability maintained, fun factor consistent
**PASS WITH CONCERNS**: Minor drift or fatigue noted; addressable in Polish
**FAIL**: Memory leak confirmed, stability breach, or severe fun fatigue
---
## Sign-Off
- **Tester**: [name] — [date]
- **QA Lead review**: [name] — [date]Present the protocol summary in conversation, then ask:
"May I write this soak test protocol to production/qa/soak-test-[date]-[duration].md?"
Write only after approval.
After writing:
"Protocol written. To run the soak:
production/qa/bugs//bug-triage sprint after the session to integrate any S1/S2 issuesIf the verdict is FAIL, run /smoke-check again after fixing the issues."
run a soak test automatically. The observations require a human observer.
doesn't need a 4h soak; a city-builder might. Use judgment and ask if unclear.
regression soaks after a specific fix, not the first pass
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.