hiring-craft — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited hiring-craft (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Hiring is the highest-leverage decision an engineering manager makes — every hire is a 1–3 year bet on team trajectory, culture, and delivery. Most hiring loops are run by tribal habit; this skill applies forcing functions at the points where loops most often fail.
Three sub-modes, triggered by what the user is doing: loop design, rubric writing, debrief discipline.
When the user is designing a new loop or reviewing an existing one, walk through this sequence:
What are the 3–5 capabilities this hire must demonstrate? Be specific.
If the user can't articulate signals, push back: an interview loop without explicit signals is a panel of vibes-checks.
Each signal should be primarily covered by one slot, with one secondary slot as a check. Two failure modes here:
Output a coverage matrix:
Signal | Primary slot | Secondary check
--------------------+--------------------+---------------------
System design | Architecture round | Coding (if relevant)
Cross-team work | Behavioral round | Reference check
Mentorship | Behavioral round | (gap — add probe to mgr 1:1)
On-call ownership | Hiring mgr round | (gap)
Senior shipping | Coding + take-home | Architecture roundDon't hope it surfaces. Make it one of the 3–5 signals, with explicit probes (how they handle disagreement, how they treat the interviewer's wrong answers, whether they ask questions or assume).
For each signal, produce a rubric with behavioral anchors at four levels: strong-yes / lean-yes / lean-no / strong-no. Anchors must describe behaviors observed in the interview, not internal qualities.
Example for "decomposes vague problems":
Strong-yes:
- Asks 2-3 clarifying questions before diving in, surfacing the ambiguity explicitly
- Decomposes into 3-4 sub-problems with clear interfaces between them
- Names which sub-problem they'd tackle first and why
- Identifies what they don't know and where they'd go to learn it
Lean-yes:
- Asks at least one clarifying question
- Decomposes successfully but flat (no interfaces / dependencies surfaced)
- Picks a starting point but reasoning is thin
Lean-no:
- Dives in without clarifying assumptions
- Decomposition is partial or imbalanced
- Reasoning sounds confident but isn't structured
Strong-no:
- No clarifying questions
- Treats the problem as the literal prompt; misses the underlying ambiguity
- Decomposition is wrong (overlapping concerns, missing crucial pieces)
- Defends the wrong decomposition when challengedForce concreteness. Adjectives like "good" or "strong" with no observable behavior aren't a rubric — they're a vibe. The whole point is that two interviewers using the rubric on the same recording would land within one notch of each other.
Include an anti-bias note per rubric: "Don't penalize accent, communication style, or unfamiliarity with the company's internal vocabulary. Score the substance."
The debrief is where most hiring decisions go wrong. Loud voices win. The first speaker anchors. Calibration drifts.
Run debriefs in this order, every time:
Before any discussion, every interviewer writes their vote (strong-yes / lean-yes / lean-no / strong-no) plus one sentence of evidence. No comparing notes yet. This breaks anchoring.
Junior interviewers speak first. Senior interviewers and the hiring manager last. This counteracts the natural pattern where junior voices defer.
Don't debate "should we hire?" first. Debate each signal: "On 'system design,' what did people see?" — going around the room. Surface evidence, not adjectives. "What did they actually do?"
Only after all signals have been walked through, ask: "Would you reach across the table to hire this person at the offered level?" Yes/no. Force commitment.
If the room splits, the question isn't who's right. It's what signal did one group see that the other didn't? Often the disagreement reveals a real divergence in the data — the candidate was strong on signal A and weak on signal B, and people weighted them differently. Surface this; don't paper over it.
Document:
The write-up is what makes the next loop better — patterns of disagreement reveal where the rubric or loop design needs fixing.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.