No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by dukesun99 · Agent Tool · ★ 85
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is Corpus2Skill safe to install? View the security audit →
Corpus2Skill Compile bounded, topically structured document corpora into navigable skill hierarchies for LLM agents. Serving uses hierarchy navigation and document lookup, without a vector-search service. This is the official implementation of the paper "Corpus2Skill: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG" (Sun, Wei, and Hsieh, 2026), accepted to Findings of EMNLP 2026. Title history: Earlier versions of the same arXiv record were titled "Don’t Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG". Use the current title above; these are versions of one paper. Corpus2Skill converts a collection of documents into a structured tree of Anthropic Skills. At query time, the LLM agent navigates this hierarchy (reading SKILL.md / INDEX.md files, drilling into sub-topics) and fetches full documents on demand — without embeddings, vector stores, or BM25 at serve time. 📦 Initial release (v0.1) — This is the initial public release accompanying the Findings of EMNLP 2026 paper. The compile–serve–eval pipeline and a WixQA quick start are included.
| Stars | 85 |
| Forks | 7 |
| Language | Python |
| Category | Agent Tool |
| License | MIT |
| Quality Score | 57.4270145703541/100 |
| Open Issues | 1 |
| Last Updated | 2026-09-06 |
| Created | 2026-04-17 |
| Platforms | python |
| Est. Tokens | ~5k |
These tools work well together with Corpus2Skill for enhanced workflows:
Looking for a Corpus2Skill alternative? If you're comparing Corpus2Skill with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Self-hosted RAG platform for AI document search across GitHub, Notion, Google Drive, local files, and web sour
open source assistant hybrid using small models (2b - 5b) and gemini , with image and agentic tool capabilit
Architecture-first skill lifecycle for AI agents. BinEval binary scoring with threshold-blind, cross-family-ca
MCP server that lets Claude Code and other AI agents read and search large PDFs, one file or a whole folder: a
Enhance LLM agents with rich tool APIs
Talk with your notes in Claude. RAG over your Apple Notes using Model Context Protocol.
Explore other popular agent tool tools:
Corpus2Skill is Official Findings of EMNLP 2026 implementation of Corpus2Skill: compile a document corpus into a navigable skill hierarchy that LLM agents explore at query time, with document lookup instead of a serv. It is categorized as a Agent Tool with 85 GitHub stars.
Corpus2Skill is primarily written in Python. It covers topics such as agent-skills, agentic-rag, emnlp-2026.
You can find installation instructions and usage details in the Corpus2Skill GitHub repository at github.com/dukesun99/Corpus2Skill. The project has 85 stars and 7 forks, indicating an active community.
Corpus2Skill is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to Corpus2Skill on Agent Skills Hub include OpenDocuments, J.A.R.V.I.S.2.0, skill-conductor. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: