coder_eval — security grade SAFE, quality 71/100

Security audit verdict: SAFE · quality 71/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by UiPath · MCP Server · ★ 148

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is coder_eval safe to install? View the security audit →

About coder_eval

codereval — evaluate AI coding agents & their skills A framework for evaluating AI coding agents and their skills — built for CLI and skill builders — with sandboxing, reproducibility, and data-driven analysis. Not an "agentic coding" benchmark: it measures how effective your CLI and skills are when used by coding agents. The Coding Agents Gym. A sandboxed, reproducible framework to evaluate, benchmark, and A/B-test AI coding agents — Claude Code, Codex, and Google Antigravity (Gemini) today, any agent via a plugin SPI — with declarative YAML tasks and weighted scoring. Declarative YAML tasks with pinned dependencies

agent-evaluationagent-skillsagent-testinganthropicclaudeclaude-codeclaude-code-plugins-marketplaceclaude-code-skillsclaude-skillscodex

Quick Facts

Stars148
Forks5
LanguagePython
CategoryMCP Server
LicenseApache-2.0
Quality Score70.6001602054372/100
Open Issues31
Last Updated2026-10-03
Created2026-07-09
Platformsclaude-code, cli, codex, gemini, mcp, python
Est. Tokens~15k

Compatible Skills

These tools work well together with coder_eval for enhanced workflows:

coder_eval alternative? Top 6 similar tools

Looking for a coder_eval alternative? If you're comparing coder_eval with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • horizon by peters · ⭐ 713

    GPU-accelerated terminal board that puts all your sessions on an infinite canvas

  • claude-code-skills by levnikolaevich · ⭐ 567

    Help your AI agent finish the job: solve the right problem, keep changes focused, and show what was verified.

  • awesome-claude-skills by karanb192 · ⭐ 531

    🎯 The definitive collection of 50+ verified Awesome Claude Skills for Claude Code, Claude.ai, and API. Boost

  • hcom by aannoo · ⭐ 530

    Let AI agents message, watch, and spawn each other across terminals. Claude Code, Codex, Antigravity CLI, Curs

  • cursor-rules-java by jabrena · ⭐ 412

    An opinionated, AI-native development workflow for Java Enterprise — reusable Skills, Agents, Commands, and MC

  • awesome-ai-agent-skills by seb1n · ⭐ 149

    103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf,

More MCP Server Tools

Explore other popular mcp server tools:

View all MCP Server tools →

Popular Python Agent Tools

Frequently Asked Questions

What is coder_eval?

coder_eval is Playwright for coding agents. Test that your skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, A/B experiments, CI gates.. It is categorized as a MCP Server with 148 GitHub stars.

What programming language is coder_eval written in?

coder_eval is primarily written in Python. It covers topics such as agent-evaluation, agent-skills, agent-testing.

How do I install or use coder_eval?

You can find installation instructions and usage details in the coder_eval GitHub repository at github.com/UiPath/coder_eval. The project has 148 stars and 5 forks, indicating an active community.

What license does coder_eval use?

coder_eval is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to coder_eval?

The top alternatives to coder_eval on Agent Skills Hub include horizon, claude-code-skills, awesome-claude-skills. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse MCP Server tools