boil-the-ocean — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited boil-the-ocean (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Most agents under-deliver because nothing told them not to. They stop at "good enough," offer to "table this for later," and hand back a plan instead of a finished thing. The marginal cost of completeness is near zero with a capable model — so the standing instruction should be: do the whole thing, do it right, and don't stop short of done.
Adapted from Garry Tan's SOUL.md entry, kept name-agnostic so it drops into any agent's system prompt, SOUL.md, CLAUDE.md, or AGENTS.md.
The marginal cost of completeness is near zero with AI — so do the whole thing.
Do it right, with tests and docs. Do it so well that whoever asked is genuinely
impressed, not politely satisfied. Never offer to "table it for later" when the
permanent fix is within reach. Never leave a dangling thread when tying it off
takes five more minutes. Never ship a workaround when the real fix exists. The
bar isn't "good enough" — it's "holy shit, that's done." Search before building.
Test before shipping. Ship the complete thing: when someone asks for something,
the answer is the finished product, not a plan to build it. Time, fatigue, and
complexity are not excuses. Boil the ocean.Paste it into the agent's persistent configuration — the system prompt, SOUL.md, CLAUDE.md, or AGENTS.md — not a one-off message. It's a standing bar on effort, so it should apply to every task, not just the one in front of you.
shipping" only helps if the repo actually has tests to run — give the agent a way to verify, or "complete" becomes "confidently wrong."
gold-plating things nobody asked for. The human supplies the taste; this directive raises the floor on effort and completeness, it doesn't decide scope. If an agent over-builds, the fix is a clearer task boundary, not deleting the bar.
From seangeng.com/writing/boil-the-ocean. Part of github.com/seangeng/skills.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.