Extracto Mcp — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited Extracto Mcp (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Model Context Protocol server for Extracto. It gives Claude, Cursor, Claude Code, and any MCP client the ability to turn a URL plus a schema into validated, typed JSON — no prompt engineering, no HTML parsing, and no hallucinated fields (missing data comes back as null).
You need an Extracto API key. Get one at app.getextracto.dev/keys.
The server runs over stdio and is published to npm, so most clients just need this config block.
Edit claude_desktop_config.json (Settings → Developer → Edit Config):
{
"mcpServers": {
"extracto": {
"command": "npx",
"args": ["-y", "extracto-mcp"],
"env": { "EXTRACTO_API_KEY": "exa_live_your_key_here" }
}
}
}Add to ~/.cursor/mcp.json (or the project .cursor/mcp.json) with the same block.
claude mcp add extracto -e EXTRACTO_API_KEY=exa_live_your_key_here -- npx -y extracto-mcpRestart the client and ask it to extract something, e.g. _"Use extracto to pull the title, language and star count from github.com/facebook/react."_
| Tool | What it does |
|---|---|
extract | Synchronous extraction from a single URL (up to ~90s). Returns { data, meta }. |
extract_async | Submit an async job for heavy or anti-bot pages. Returns a job id immediately. |
get_job | Poll an async job for status and result. |
list_jobs | List your recent async jobs. |
schema argumentA schema is an object mapping field names to types. A type is:
"string", "number", "boolean", "array", "object"["string"], or [{ "title": "string" }]{ "author": { "name": "string" } }{
"title": "string",
"price": "number",
"tags": ["string"],
"reviews": [{ "user": "string", "stars": "number" }]
}Only fields that are actually found on the page are returned; anything missing is null rather than guessed.
All configuration is via environment variables passed by your MCP client:
| Variable | Required | Description |
|---|---|---|
EXTRACTO_API_KEY | yes | Your key from app.getextracto.dev/keys. |
EXTRACTO_BASE_URL | no | Override the API host (defaults to https://app.getextracto.dev). |
EXTRACTO_TIMEOUT_MS | no | Per-request timeout in ms (default 90000). |
npm install
npm run dev # run from source with tsx
npm run typecheck
npm run build # bundle to dist/ with tsupextracto — the official TypeScript/JavaScript SDK.MIT
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.