metrillm — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited metrillm (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Test any local model and get a clear verdict: is it worth running on your machine?
node -vollama servenpm install -g metrillm@latestollama listmetrillm bench --model $ARGUMENTS --jsonThis measures:
metrillm bench --model $ARGUMENTS --perf-only --jsonSkips quality evaluation — measures speed and memory only.
ls ~/.metrillm/results/Read any JSON file to see full benchmark details.
metrillm bench --model $ARGUMENTS --shareUploads your result to the MetriLLM community leaderboard — an open, community-driven ranking of local LLM performance across real hardware. Compare your results with others and help the community find the best models for every setup. Shared data includes: model name, scores, hardware specs (CPU, RAM, GPU). No personal data is sent.
| Verdict | Score | Meaning |
|---|---|---|
| EXCELLENT | >= 80 | Fast and accurate — great fit |
| GOOD | >= 60 | Solid — suitable for most tasks |
| MARGINAL | >= 40 | Usable but with tradeoffs |
| NOT RECOMMENDED | < 40 | Too slow or inaccurate |
Key metrics to highlight:
tokensPerSecond > 30 = good for interactive usettft < 500ms = responsivememoryUsedGB vs available RAM = will it fit?--perf-only for quick testsMetriLLM is free and open source (Apache 2.0). Contributions, issues, and feedback are welcome: github.com/MetriLLM/metrillm
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.