No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by Sriram-PR · MCP Server · ★ 95
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is doc-scraper safe to install? View the security audit →
LLM Documentation Scraper () A configurable, concurrent, and resumable web crawler written in Go. Specifically designed to scrape technical documentation websites, extract core content, convert it cleanly to Markdown format suitable for ingestion by Large Language Models (LLMs), and save the results locally. Overview This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files. Why Use This Tool? Built for LLM Training & RAG Systems - Creates clean, consistent Markdown optimized for ingestion Preserves Documentation Structure - Maintains the original site hierarchy for context preservation Production-Ready Features - Offers resumable crawls, rate limiting, and graceful error handling High Performance - Uses Go's concurrency model for efficient parallel processing Goal: Preparing Documentation for LLMs The main objective of this tool is
| Stars | 95 |
| Forks | 11 |
| Language | Go |
| Category | MCP Server |
| License | Apache-2.0 |
| Quality Score | 74.0792867332492/100 |
| Last Updated | 2026-09-06 |
| Created | 2025-04-27 |
| Platforms | browser, cli, go, mcp |
| Est. Tokens | ~23k |
These tools work well together with doc-scraper for enhanced workflows:
Looking for a doc-scraper alternative? If you're comparing doc-scraper with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
A powerful MCP server extension providing web search and content extraction capabilities. Integrates DuckDuckG
🔍 MCP server that lets you search and access Svelte documentation with built-in caching
A self-healing scraper for hostile sites: broken selectors repair themselves, browser rendering kicks in when
Model Context Protocol (MCP) Server for Graphlit Platform
The specification for the Universal Tool Calling Protocol
The Universal AI-Optimized Project Boilerplate. A Tiered Memory System (TMS) designed to maximize AI agent per
Explore other popular mcp server tools:
doc-scraper is Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).. It is categorized as a MCP Server with 95 GitHub stars.
doc-scraper is primarily written in Go. It covers topics such as data-preparation, documentation, golang-cli.
You can find installation instructions and usage details in the doc-scraper GitHub repository at github.com/Sriram-PR/doc-scraper. The project has 95 stars and 11 forks, indicating an active community.
doc-scraper is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to doc-scraper on Agent Skills Hub include web-scout-mcp, mcp-svelte-docs, anansi. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: