data-resiliency-testing-and-failure-injection — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited data-resiliency-testing-and-failure-injection (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Use this skill when the goal is to prove that a data system recovers safely under failure, not only when everything goes right. It helps agents design controlled drills for retries, restarts, dependency outages, state recovery, duplicate prevention, backlog catch-up, and publish protection.
Do not treat resilience testing as random breakage. The point is to validate recovery behavior with explicit safety limits and evidence.
Prioritize:
Include:
Prefer:
Use controlled exercises such as:
Check:
The best resilience test is one the team can rerun after changes, not a one-time exercise that gets forgotten.
safe-backfill-and-replay-orchestrationkafka-resilience-and-schema-evolutionspark-serverless-reliability-and-state-managementmcp-data-observability-integrationreferences/data-resiliency-testing-patterns.md| Rationalization | Reality |
|---|---|
| "If the job retries, we are resilient enough." | Retry alone does not prove replay safety, duplicate prevention, or publish protection. |
| "We can test recovery during a real incident." | Real incidents are the worst time to discover the recovery path is unclear or unsafe. |
| "Failure injection is too risky for data systems." | Uncontrolled failure is riskier than bounded, reviewable drills in safe environments. |
| "The scheduler health page already proves resilience." | Scheduler status does not prove data correctness, backlog catch-up, or downstream safety. |
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.