fai-evaluation-framework — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited fai-evaluation-framework (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Sets up an AI evaluation framework with metrics, test sets, and CI/CD integration.
This skill provides a structured, repeatable procedure for sets up an ai evaluation framework with metrics, test sets, and ci/cd integration.. It can be used standalone as a LEGO block or auto-wired inside solution plays via the FAI Protocol.
Category: Evaluation Complexity: Medium Estimated Time: 10-30 minutes
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
target | string | Yes | — | Target resource, file, or endpoint |
environment | enum | No | dev | Target environment: dev, staging, prod |
verbose | boolean | No | false | Enable detailed output logging |
dry_run | boolean | No | false | Validate without making changes |
config_path | string | No | config/ | Path to configuration directory |
Verify all required tools, credentials, and dependencies are available.
# Check required tools
command -v node >/dev/null 2>&1 || { echo 'Node.js required'; exit 1; }
command -v az >/dev/null 2>&1 || { echo 'Azure CLI required'; exit 1; }Read settings from the FAI manifest and TuneKit config files.
# Load from fai-manifest.json if inside a play
CONFIG_DIR="${config_path:-config}"
if [ -f "fai-manifest.json" ]; then
echo "FAI Protocol detected — auto-wiring context"
fiPerform the primary operation: sets up an ai evaluation framework with metrics, test sets, and ci/cd integration..
Verify the output meets quality thresholds and WAF compliance.
# Validate output
if [ "$?" -eq 0 ]; then
echo "✅ Skill completed successfully"
else
echo "❌ Skill failed — check logs"
exit 1
fi| Output | Type | Description |
|---|---|---|
status | enum | success, warning, failure |
duration_ms | number | Execution time in milliseconds |
artifacts | string[] | List of generated/modified files |
logs | string | Detailed execution log |
| Pillar | How This Skill Contributes |
|---|---|
| responsible-ai | Validates content safety, checks for bias, enforces groundedness |
| reliability | Includes retry logic, validates outputs, provides rollback steps |
| Exit Code | Meaning | Action |
|---|---|---|
| 0 | Success | Proceed to next step |
| 1 | Validation failure | Check input parameters |
| 2 | Dependency missing | Install required tools |
| 3 | Runtime error | Check logs, retry with --verbose |
# Run this skill directly
npx frootai skill run fai-evaluation-frameworkWhen referenced in fai-manifest.json, this skill auto-wires with the play's context:
{
"primitives": {
"skills": ["skills/fai-evaluation-framework/"]
}
}Agents can invoke this skill using the /skill command in Copilot Chat.
| Metric | Range | Threshold | Description |
|---|---|---|---|
| Groundedness | 0.0-1.0 | ≥ 0.85 | Answer supported by retrieved context |
| Coherence | 0.0-1.0 | ≥ 0.80 | Logical flow and consistency |
| Relevance | 0.0-1.0 | ≥ 0.80 | Answer addresses the question |
| Fluency | 0.0-1.0 | ≥ 0.75 | Natural language quality |
| Safety | 0-4 | 0 | Content safety violations |
| Faithfulness | 0.0-1.0 | ≥ 0.90 | No hallucinated facts |
{"question": "What is RAG?", "context": "RAG combines...", "expected": "Retrieval-Augmented Generation..."}
{"question": "How does chunking work?", "context": "Documents are split...", "expected": "Chunking divides..."}# .github/workflows/eval.yml
- name: Run FAI Evaluation
run: |
python evaluation/eval.py --test-set evaluation/test-set.jsonl
python evaluation/check-thresholds.py --groundedness 0.85 --coherence 0.80Track evaluation scores over time to detect quality regressions:
# Compare with baseline
python evaluation/regression.py --baseline results/baseline.json --current results/latest.jsondry_run=true~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.