Mcp Docpilot Server — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited Mcp Docpilot Server (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
An MCP server that exposes document retrieval as tools any LLM provider can call. It puts a single, stable interface in front of a vector index (built from DocPilot's ingestion pipeline) so a model never has to know how the documents are stored or which embedding backend is in use - it just calls docpilot_search.
The server provides the tools and data access; the connected model does the generation. That split is what makes it provider-agnostic: Claude Desktop, or any client that speaks MCP, gets the same retrieval tools.
| Tool | What it does |
|---|---|
docpilot_search | Semantic search over the corpus; returns ranked chunks with source and score |
docpilot_list_sources | Lists indexed source documents with per-source chunk counts |
Both tools are read-only.
docs/*.md ──ingest.py──> chunk + embed ──> ChromaDB (persistent)
│
server.py exposes ──┤── docpilot_search
MCP tools over └── docpilot_list_sources
stdio or HTTP
│
Claude Desktop / any MCP client ──┘ (model calls the tools)python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Build the index from the docs folder (swap in your own .txt/.md files)
python ingest.py ./docsEmbeddings use ChromaDB's local default model, so it runs with no API key. To point it at a hosted embedding provider instead, set a ChromaDB embedding function in ingest.py and server.py - the rest of the pipeline is unchanged.
stdio (local clients like Claude Desktop):
python server.pyStreamable HTTP (remote server):
DOCPILOT_TRANSPORT=http python server.py
# serves MCP at http://localhost:8000/mcpThe SDK's HTTP transport supersedes the older SSE transport; point HTTP-based MCP clients at the /mcp endpoint.
Add this to claude_desktop_config.json:
{
"mcpServers": {
"docpilot": {
"command": "python",
"args": ["/absolute/path/to/mcp-docpilot-server/server.py"],
"env": {
"DOCPILOT_CHROMA_PATH": "/absolute/path/to/mcp-docpilot-server/chroma"
}
}
}
}Or, for an HTTP server:
claude mcp add --transport http docpilot http://localhost:8000/mcp| Env var | Default | Meaning |
|---|---|---|
DOCPILOT_CHROMA_PATH | ./chroma | Persistent ChromaDB store |
DOCPILOT_COLLECTION | docpilot | Collection name |
DOCPILOT_TRANSPORT | stdio | stdio or http |
DOCPILOT_CHUNK_SIZE | 800 | Characters per chunk (ingest) |
DOCPILOT_CHUNK_OVERLAP | 100 | Overlap between chunks (ingest) |
DOCPILOT_EMBEDDINGS | default | default (local ONNX model) or hash (offline, for CI/tests) |
Switching the embedding backend changes the vector space, so re-ingest into a fresh store when you change it (rm -rf chroma && python ingest.py ./docs). All backend selection lives in embeddings.py - that one file is the seam for the embedding lifecycle.
pytest -qThe test ingests a tiny corpus and confirms retrieval ranks the expected document first. CI runs it on every push (.github/workflows/ci.yml).
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.