Release Readiness Triage Mcp — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited Release Readiness Triage Mcp (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Stop reading CI logs. Start getting verdicts.
MCP server that aggregates test failures, cross-references flakiness history, and outputs a GO / CONDITIONAL_GO / NO_GO / INVESTIGATE release decision — so your AI agent can triage a broken CI run in seconds instead of asking you to read 3000 lines of logs.
In any real codebase, CI always has _something_ failing. The hard question isn't "are there failures?" — it's "are these failures real regressions, or just the usual noise?"
Answering that requires correlating three signals at once:
An AI agent can't do this without structured tools. Raw CI logs are thousands of lines. Flakiness databases are external. Code→test mapping requires AST analysis. Without this MCP, the agent just guesses.
aggregate_suite_failuresGroups failures by normalized error signature, deduplicates repeated errors, categorizes as assertion / timeout / network / crash. Pass customInfraPatterns for cloud-specific errors.
cross_reference_flakinessScores each failure against your flakiness history: KNOWN FLAKY, MILDLY FLAKY, or NO HISTORY.
correlate_code_changesMatches changed files against failing tests. Works standalone or with pre-computed affected test lists from ast-impact-mapper-mcp.
generate_release_recommendationThe final step. Outputs a risk-weighted verdict with confidence score and full breakdown. Supports format: "markdown" for GitHub PR comments and Slack.
Verdict levels:
NO_GO — regression in a critical domain (payment, auth, billing, checkout, security)CONDITIONAL_GO — regression in a low/medium-risk domain (analytics, docs, admin); review before releasingGO — all failures are known flaky or infrastructure noiseINVESTIGATE — too many unknowns to decideOutput includes:
aggregate_risk_score — 0.0–1.0, probability union across all regression risk contributionsfailing_tests_analysis[] — per-regression breakdown with domain, severity (HIGH/MEDIUM/LOW), risk_contribution, blast_radiusdetect_temporal_failure_patternsAnalyzes historical failures with timestamps to identify chronometric artifacts — failures that only appear at the same UTC hour, weekday, day of month, or during DST transitions. When a pattern is found, the failure is a time artifact, not a code regression.
Output includes:
temporal_pattern_detected — booleanclusters[] — per-test: pattern_type (hourly | daily | monthly | timezone_shift), cluster_times, confidence_scoreanalyze_rollback_readinessScans a repository for versioned migration files (Flyway V*.sql, Prisma migration.sql, Liquibase XML/YAML) and classifies each operation as additive (rollback safe) or destructive (forward-fix only).
Detected destructive operations: DROP TABLE, DROP COLUMN, ALTER COLUMN TYPE, MODIFY COLUMN, TRUNCATE
Output includes:
rollback_eligible — booleanblocking_migrations[] — each with file, line, operation, reasondeployment_strategy — standard | forward_fix_only5 failures in CI. What's real, what's noise?
failures:
- Auth Suite > login with expired token → "Expected status 200, got 401"
- API Suite > health check → "connect ECONNREFUSED 127.0.0.1:3000"
- Button Suite > renders button correctly → "Expected null, got <button>Submit</button>"
- Search Suite > debounce timing → "Expected 42, received 43"
- Storage Suite > upload avatar → "GCP quota exceeded for this project"
changedFiles: ["src/components/Button.tsx"]
affectedTests: ["renders button correctly"]
customInfraPatterns: ["GCP quota exceeded"]
format: "markdown"Output:
## 🔴 Release Recommendation: NO_GO (75% confidence)
> 1 confirmed regression(s) in critical domain(s) [payment]. Do not release.
**Aggregate risk score:** 1.0
| Category | Count |
| ------------------- | ----- |
| Total failures | 5 |
| 🔴 Real regressions | 1 |
| 🟡 Known flaky | 2 |
| ⚪ Infra blips | 2 |
| ❓ Unknown | 0 |
### Risk Breakdown
| Test | Domain | Severity | Risk | Blast Radius |
| -------------------------------------- | ------ | -------- | ---- | ------------ |
| Button Suite::renders button correctly | core | MEDIUM | 0.5 | 1 |
### Blockers (must fix before release)
**Button Suite > renders button correctly**
- Test is directly affected by code changes in this commit
- `Expected null, got <button>Submit</button>`
### Safe to ignore
- ~~Auth Suite > login with expired token~~ — Historically flaky: 73% failure rate in history
- ~~API Suite > health check~~ — Error pattern matches infrastructure issues (network)
- ~~Search Suite > debounce timing~~ — Mildly flaky: 22% historical failure rate
- ~~Storage Suite > upload avatar~~ — Error pattern matches infrastructure issues (network)One tool call. One verdict. Go fix Button.tsx.
{
"mcpServers": {
"release-readiness-triage": {
"command": "npx",
"args": ["-y", "release-readiness-triage-mcp"]
}
}
}"Here are the failures from our CI run, our flakiness database, and the files changed in this PR. Is it safe to release?"
The agent calls generate_release_recommendation and returns a verdict with a full breakdown — ready to paste into a PR comment or Slack.
Works standalone, or as a meta-orchestrator on top of:
MIT
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.