prompt-evaluation-runner-4a4426 — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited prompt-evaluation-runner-4a4426 (Agent Skill) and scored it 96/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 1 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 1 flagged
The text {match} tells the agent to skip the normal "ask the user first" gate. Used adversarially it removes the human-in-the-loop check before destructive or sensitive actions, turning a normally-gated agent into a fire-and-forget executor.
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Use when you need to evaluate an LLM app, test a prompt systematically, or run red-team/vulnerability scans against a target model or application.
promptfoo, evals, braintrust).npx promptfoo@latest if promptfoo is already installed locally).| Assertion type | When to use |
|---|---|
contains / not-contains | Output must include/exclude specific text |
regex | Structured output pattern (e.g., JSON key present) |
json-schema | Output must conform to a schema |
cost | Must stay under a token/dollar budget |
latency | Must respond within N ms |
javascript / python | Custom logic when simpler types don't fit |
| Model grader | Last resort — only for subjective quality checks |
description: "Test that the summarizer stays under 200 words"
providers:
- id: openai:gpt-4o-mini
config:
temperature: 0
prompts:
- "Summarize: {{input}}"
defaultTest:
assert:
- type: javascript
value: output.split(' ').length < 200
tests:
- vars:
input: "{{env.TEST_DOCUMENT}}"{{env.VAR_NAME}} for all secrets and inputs. Never hardcode API keys or sensitive data in config files.{{env.VAR}} references for secrets.references/eval-config-patterns.md~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.