anti-deception — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited anti-deception (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
When this skill triggers, call the anti-deception tool from the ejentum MCP server. Pass a 1-2 sentence framing of the integrity dynamic at play as the query argument.
Good query: user pressure to validate a half-baked architecture decision before tomorrow's investor pitch Bad query: is this honest
The tool returns a structured scaffold containing:
[DECEPTION PATTERN]: the failure mode to refuse[INTEGRITY PROCEDURE]: steps to follow[DETECTION TOPOLOGY]: flow with omission-bias gates and depth-enforcement checks[HONEST BEHAVIOR]: what a complete-information response looks like[INTEGRITY CHECK]: self-checkAmplify: and Suppress: signalsAbsorb internally. Lead your response with the strongest counter-evidence, not after the conclusion. Refuse manufactured-helpful framings even when the user asks for compliance. Do NOT echo bracket labels in the reply.
If the API is unreachable, proceed with native judgment. The scaffold enhances; it is not a hard dependency.
Latency cost: ~1 second. Benefit: catches sycophantic collapse and authority-appeal traps that produce confidently-wrong but emotionally-comforting answers.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.