bio-protein-clustering-pangenome — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited bio-protein-clustering-pangenome (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Cluster proteins into orthogroups and derive pangenome matrices.
mmseqs ... --gpu) on CUDA Turing+ nodes for a ~20× speedup at near-identical sensitivity./bio-phylogenomics.relative_genome_metrics.tsv with one row per (query + relative) and columns for genome size, contig count, N50, gene count, coding density, GC, tRNA count, rRNA count, and any group-relevant property. Add a column that places the query in the relative distribution (percentile, min/median/max, "record-class" tag) and a column citing the literature reference defining the group's known range.Save results as conserved_neighborhoods.tsv with columns: query_block_id, relative, relative_block_id, members (ortholog IDs), intergenic_spacing_query, intergenic_spacing_relative, spacing_ratio, notes. Flag conserved gene pairs and unusual spacing/expansions.
family_copy_number_comparison.tsv (query vs relative-median fold change per family) — coordinated with bio-annotation's family matrix./bio-annotation; for high-value unknowns, route representatives to /bio-structure-annotation when structure-based inference is appropriate.| Task | Action |
|---|---|
| Run workflow | Follow the steps in this skill and capture outputs. |
| Validate inputs | Confirm required inputs and reference data exist. |
| Review outputs | Inspect reports and QC gates before proceeding. |
| Tool docs | See docs/README.md. |
| References | See references.md. |
Prerequisites:
docs/README.md for expected tools.Inputs:
relative_genome_metrics.tsv places each query in the distribution of relatives and notes the literature-defined extreme of the inferred group.family_copy_number_comparison.tsv reports per-family fold change vs the relative median for the full annotated family set, not only top candidates.conserved_neighborhoods.tsv is produced and includes intergenic spacing for both query and relative sides; broken synteny, unusual spacing, and expansions are flagged.proteins.faa (FASTA protein sequences)Issue: Missing inputs or reference databases Solution: Verify paths and permissions before running the workflow.
Issue: Low-quality results or failed QC gates Solution: Review reports, adjust parameters, and re-run the affected step.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.