qdrant-scaling-qps — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited qdrant-scaling-qps (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Throughput scaling means handling more parallel queries per second. This is different from latency - throughput and latency are opposite tuning directions and cannot be optimized simultaneously on the same node.
High throughput favors fewer, larger segments so each query touches less overhead.
default_segment_number: 2) Maximizing throughputalways_ram=true to reduce disk IO Quantizationoptimizer_cpu_budget to limit indexing CPUs (e.g. 2 on an 8-CPU node reserves 6 for queries)If a single node is saturated on CPU after applying the tuning above, scale horizontally with read replicas.
replication_factor: 2+ and route reads to replicas Distributed deploymentSee also Horizontal Scaling for general horizontal scaling guidance.
If it is not possible to keep all vectors in RAM, disk I/O can become the bottleneck for throughput. In this case:
io_uring on Linux (kernel 5.11+) io_uring articlecpu_count - 1, which is optimal for RAM-based search but may be too low for disk-based search. See configuration reference~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.