Block-based PDF extraction MCP server optimized for LLM consumption
SaferSkills independently audited Pdf4Vllm Mcp (MCP Server) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
PDF reading MCP server optimized for vision LLMs.
<!-- mcp-name: io.github.PyJudge/pdf4vllm -->
<details> <summary><b>한국어</b></summary>
| 방식 | 문제점 |
|---|---|
| 텍스트 추출 | 인코딩 깨짐 → 쓰레기 출력, 이미지-텍스트 순서 뒤섞임 |
| 이미지 변환 | 토큰 폭발 (특히 페이지 많을 때) |
pdf4vllm은 PDF가 지저분하다고 가정합니다.
pip install pdf4vllm-mcp
# 또는
uvx pdf4vllm-mcpgit clone https://github.com/PyJudge/pdf4vllm-mcp.git
cd pdf4vllm-mcp
python scripts/install_mcp.py또는 직접 설정 (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"pdf4vllm": {
"command": "/python/경로",
"args": ["/pdf4vllm-mcp/경로/src/server.py"]
}
}
}| 도구 | 설명 |
|---|---|
list_pdfs | PDF 파일 찾기 (glob 패턴 name_pattern 지원) |
read_pdf | PDF 내용 블록으로 추출 |
grep_pdf | PDF 내 텍스트 검색 (pdfgrep 설치 필요) |
| 모드 | 설명 |
|---|---|
auto (기본) | 텍스트 추출 시도 → 손상 감지 시 이미지로 전환 |
text_only | 텍스트/표만 추출, 이미지 없음 |
image_only | 페이지를 이미지로만 렌더링 |
</details>
| Approach | Issue |
|---|---|
| Text extraction | Encoding corruption → garbage output, mixed text-image ordering |
| Image conversion | Token explosion (especially with many pages) |
pdf4vllm assumes PDFs are messy.
PDF Input
↓
Corruption Detection (pdfminer.six + pattern analysis)
↓
┌─────────────┬─────────────┐
│ Corrupted │ Clean │
│ → Image │ → Text + │
│ only │ Tables + │
│ │ Images │
└─────────────┴─────────────┘
↓
Ordered Blocks (JSON)pip install pdf4vllm-mcp
# or run without installing
uvx pdf4vllm-mcpgit clone https://github.com/PyJudge/pdf4vllm-mcp.git
cd pdf4vllm-mcp
python scripts/install_mcp.pyOr manually edit ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"pdf4vllm": {
"command": "/path/to/python",
"args": ["/path/to/pdf4vllm-mcp/src/server.py"]
}
}
}Create .mcp.json in your project:
{
"mcpServers": {
"pdf4vllm": {
"command": "uvx",
"args": ["pdf4vllm-mcp"]
}
}
}| Tool | Description |
|---|---|
list_pdfs | Find PDF files with glob filtering (name_pattern) |
read_pdf | Extract PDF content as ordered blocks |
grep_pdf | Search text in PDFs using pdfgrep (requires pdfgrep installed) |
| Mode | Description |
|---|---|
auto (default) | Try text extraction → switch to image if corrupted |
text_only | Text/tables only, no images |
image_only | Render pages as images only |
{
"pages": [
{
"page_number": 1,
"content_blocks": [
{"type": "text", "content": "..."},
{"type": "table", "content": "| A | B |"},
{"type": "image", "content": "[IMAGE_0]"}
]
}
]
}When text is corrupted:
{
"page_number": 2,
"content_blocks": [],
"text_corrupted": true,
"page_image": "[IMAGE_1]"
}config.json or environment variables:
{
"max_pages_per_request": 10,
"max_image_dimension": 842,
"page_image_dpi": 100
}export PDF_MAX_PAGES=20
export PDF_PAGE_IMAGE_DPI=150pip install pdf4vllm-mcp[test]
python test_server.py
# → http://localhost:8000MIT
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.