stress-test — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited stress-test (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Runs a three-phase decision analysis using verbalized sampling to surface tail-distribution insights — the non-obvious analyses that standard prompting suppresses.
Usage:
/stress-test should I enter the home wellness category or stay narrower?/stress-test two content formats: long-form carousel vs short-form daily posts/stress-test hire a contractor now vs wait until revenue hits $10k/moAfter alignment training (RLHF/DPO), LLMs suffer from mode collapse: they return the most typical analysis — safe, expected, rarely wrong but rarely surprising. The genuinely valuable insights live in the tails of the distribution.
Verbalized sampling (Zhang et al., Stanford 2025) fixes this: by asking the model to generate multiple candidate responses with probability estimates, it forces reasoning across the full distribution, including the suppressed tails. Diversity gains of 1.6–2.1x in creative tasks have been reported.
Important caveat (from follow-up research, Jun 2025): The probability numbers are unreliable — LLMs can describe a distribution accurately but don't faithfully sample from it. The mechanism that actually works is the diversity forcing (generate N, select from the non-obvious end), not the precision of the scores. Use the scores as a ranking device, not a measurement.
Before running the analysis, read the relevant context files for this project:
Adjust these paths to match your project structure. The goal is grounding the analysis in real constraints before reasoning begins.
Reframe the question if needed after reading context (e.g. "build vs buy" might actually be "delegate vs own").
For each of the four analytical perspectives below, generate 3 candidate responses with probability estimates (higher = more typical/expected). Then use the lowest-probability response as the perspective output.
The goal is not the probability number — it is the act of generating multiple candidates that forces the model past its default to the non-obvious analysis.
Four required perspectives (single-model maximally-differentiated):
For each perspective, output:
[Tail insight — probability: X%] label on the final output to make explicit this was the non-obvious candidateApply the preset lenses for the decision type, or use the custom lens block if the user provides one.
founder (default for COO/strategic decisions)| Lens | Question |
|---|---|
| Time ROI | What is the hours-invested-to-value-returned ratio? Where does attention compound vs. drain? |
| Cashflow impact | What does this do to the financial runway in 90 days? In 12 months? |
| Long-term optionality | Does this open or close future options? Would a future version of you be grateful or boxed in? |
| Constraint check | Does this fit the actual constraints (time, money, energy, anti-goals)? Be honest — does it actually fit? |
dtc-store (for product, category, supplier decisions)| Lens | Question |
|---|---|
| Margin viability | Can this hit ≥40% gross margin at a price the market will pay? |
| Agent-readiness | Will this product be findable and purchasable by an AI shopping agent in 2026? |
| Competition density | Is the category dominated, defensible, or open? Where exactly is the gap? |
| Execution path | What is the specific next physical action? Fastest test with real data? |
brand-content (for content strategy, platform, format decisions)| Lens | Question |
|---|---|
| Audience fit | Does this serve the specific person the brand is talking to? Would they stop and engage? |
| Credibility alignment | Does this reinforce or dilute the brand's core positioning? |
| Distribution leverage | Does this work with how the algorithm actually distributes content in 2026? |
| Longevity test | Will this still be relevant in 6 months, or is it chasing a news cycle? |
If the user provides custom lenses (e.g. "lens: investor view / employee impact / technical feasibility"), use those instead. Format: same table structure, apply same depth of analysis.
Output a structured brief under 500 words. No process explanation — results only.
THE QUESTION: [restate — reframe if the framing was wrong]
WHERE PERSPECTIVES AGREE: [2-3 genuine convergence points]
WHERE PERSPECTIVES DISAGREE: [key tensions — explicitly flag any tail-distribution insights
that challenge consensus from Step 1]
RISK: [single most important failure mode, one sentence]
BLIND SPOT: [assumption being made without stating — the thing that, if wrong, changes everything]
OPPORTUNITY: [unseen adjacent upside — what would a competitor do with this same information?]
VERDICT: [clear recommendation, 2-3 sentences — not "it depends"]
TEST IT THIS WEEK: [specific action + metric + success threshold]Verdict must be real. If genuinely ambiguous, name the single deciding factor and what it would take to resolve it.
YYYY-MM-DD-stress-test-[slug].md.~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.