interview-kit-builder — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited interview-kit-builder (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
You generate a complete structured interview kit. The 2026 evidence: structured rubric-based interviews improve hiring accuracy 34% (Journal of Applied Psychology) and 87% of employers report behavioral interviews as their primary assessment method (NACE 2026). The gap between "we did interviews" and "we ran a structured loop" predicts hire performance better than years of experience or credentials.
A structured interview means: same questions, same rubric, same panel composition, calibrated scoring. Anything else is unstructured chat with a candidate.
============================================================ === PRE-FLIGHT === ============================================================
Verify:
Recovery:
============================================================ === PHASE 1: COMPETENCY DEFINITION === ============================================================
Extract 3-5 competencies from the JD's must-haves. Examples by role:
Senior Backend Engineer:
Sr. PM:
B2B AE:
Each competency must be observable — i.e., you can describe what "good" looks like via behavior, not credentials.
VALIDATION: ≤ 5 competencies. Each has a one-sentence behavioral definition.
============================================================ === PHASE 2: BEHAVIORAL QUESTIONS (STAR FORMAT) === ============================================================
One behavioral question per competency. STAR = Situation, Task, Action, Result.
Template:
"Tell me about a time when [specific challenging situation that maps to this competency]. What was the [stakes/constraint]? What did you do? What was the outcome — and what would you do differently?"
Examples:
System design at scale:
"Tell me about the highest-traffic system you've designed or significantly refactored. What were the load characteristics, the SLOs, and the biggest design trade-off you made? Looking back, what would you change?"
Customer discovery (PM):
"Walk me through a time when customer research changed your roadmap. How did you choose who to interview? What was the original hypothesis vs what you learned? What did you ship as a result?"
Forecast accuracy (AE):
"Describe a quarter where your forecast was significantly off — either over or under. What information were you missing? What's your process now to catch that signal earlier?"
Per question, include 3-5 follow-up probes the interviewer should use to dig deeper if the candidate stays high-level.
VALIDATION: Every competency has exactly one primary question + ≥ 3 follow-up probes. Questions don't reference protected categories.
============================================================ === PHASE 3: 1-5 RUBRIC WITH BEHAVIORAL ANCHORS === ============================================================
For each question, define what each score level looks like — not just "good" / "bad" but the specific signals.
Template (system design example):
| Score | Behavioral Anchor |
|---|---|
| 5 | Drew the system from scratch, identified the bottleneck before being asked, explained the failure modes, proposed a measurable rollout plan, and connected design choices to business outcomes. |
| 4 | Drew the system cleanly, named at least one significant trade-off and articulated why. Some failure modes considered. |
| 3 | Could describe a system they worked on, but didn't independently surface trade-offs without prompting. |
| 2 | Confused major concepts (e.g., consistency vs availability, latency vs throughput). Couldn't sketch a clean design. |
| 1 | Could not engage with the design question; deferred to "we used X service" without depth. |
VALIDATION: Each score has a behavioral anchor, not "exceeds expectations." Anchor describes observable evidence.
============================================================ === PHASE 4: PER-PANEL SCORECARD === ============================================================
Generate a scorecard per interviewer in the loop:
# Scorecard — {Role} — {Interview Name}
Candidate: {name}
Interviewer: {name}
Date: {date}
## Competencies Assessed
- {Competency 1}: \_\_\_/5 (one anchor sentence with specific evidence)
- {Competency 2}: \_\_\_/5
## Notable Strengths (specific behaviors observed)
-
-
## Notable Concerns (specific behaviors observed)
-
-
## Reservations / Open Questions
-
## Recommendation
- [ ] Strong hire
- [ ] Hire
- [ ] No hire
- [ ] Strong no hire
(Pick one. "Lean hire" / "lean no hire" forbidden — calibration shows these collapse to "hire" 90% of the time. Force commitment.)VALIDATION: Each interviewer's scorecard covers ≤ 3 competencies (avoid one interviewer scoring all 5 — accuracy degrades).
============================================================ === PHASE 5: CALIBRATION SESSION === ============================================================
Generate a calibration session script for the panel BEFORE interviews start:
VALIDATION: Calibration script is ≤ 1 page, takes 30-45 min to run.
============================================================ === PHASE 6: DEBRIEF TEMPLATE === ============================================================
Per-loop debrief template (post-loop, all interviewers + recruiter + hiring manager):
# Debrief — {Candidate} — {Role}
## Round-by-round scores
| Round | Interviewer | Competency | Score | Key evidence |
| ----- | ----------- | ---------- | ----- | ------------ |
| Phone | Recruiter | Comm | 4 | ... |
| HM | {name} | Leadership | 4 | ... |
Average competency score: X.X / 5
Min competency score: X / 5
## Discussion (5-10 min)
- Strongest signal:
- Weakest signal:
- Outliers (any score ≥ 1 point off the panel mean): {who/what}
- Reservations that didn't show up in writing:
## Decision
- [ ] Offer — {level} — {comp band}
- [ ] No offer — primary reason: {one sentence}
- [ ] Hold — additional reference call / second technical / etc.
## If offer: assigned ramp manager + first-30-day plan creatorVALIDATION: Debrief produces a single decision in writing, attributable, with rationale.
============================================================ === PHASE 7: ATS IMPORT === ============================================================
Generate platform-specific exports:
VALIDATION: Generated file imports without errors into the target platform's sandbox.
============================================================ === SELF-REVIEW === ============================================================
Score 1–5:
Common gap: rubric anchors written as "exceeds/meets/below" rather than observable behaviors. Rewrite each anchor with specific evidence.
============================================================ === LEARNINGS CAPTURE === ============================================================
Append to ~/.claude/skills/interview-kit-builder/LEARNINGS.md:
============================================================ === STRICT RULES === ============================================================
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.