Flagged: reads sensitive env vars. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by mclenhard · MCP Server · ★ 132
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is mcp-evals safe to install? View the security audit →
MCP Evals A Node.js package and GitHub Action for evaluating MCP (Model Context Protocol) tool implementations using LLM-based scoring, with built-in observability support. This helps ensure your MCP server's tools are working correctly, performing well, and are fully observable with integrated monitoring and metrics. Installation As a Node.js Package As a GitHub Action Add the following to your workflow file: Usage -- Evals Create Your Evaluation File You can create evaluation configurations in either TypeScript or YAML format. Option A: TypeScript Configuration Create a file (e.g., ) that exports your evaluation configuration: typescrip
| Stars | 132 |
| Forks | 14 |
| Language | TypeScript |
| Category | MCP Server |
| License | MIT |
| Quality Score | 78.7880420727528/100 |
| Open Issues | 7 |
| Last Updated | 2025-06-23 |
| Created | 2025-04-23 |
| Platforms | mcp, node |
| Est. Tokens | ~14k |
Looking for a mcp-evals alternative? If you're comparing mcp-evals with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
Prefect MCP server
AgentEval is the comprehensive .NET toolkit for AI agent evaluation—tool usage validation, RAG quality metrics
An opinionated list of awesome Pydantic-AI frameworks, libraries, software and resources.
Run workflows, delegate to swarms, and verify outputs before you apply them.
Explore other popular mcp server tools:
mcp-evals is A Node.js package and GitHub Action for evaluating MCP (Model Context Protocol) tool implementations using LLM-based scoring. This helps ensure your MCP server's tools are working correctly and perfor. It is categorized as a MCP Server with 132 GitHub stars.
mcp-evals is primarily written in TypeScript. It covers topics such as ai, evals, mcp.
You can find installation instructions and usage details in the mcp-evals GitHub repository at github.com/mclenhard/mcp-evals. The project has 132 stars and 14 forks, indicating an active community.
mcp-evals is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to mcp-evals on Agent Skills Hub include skill-optimizer, my-pi, prefect-mcp-server. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: