No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by UiPath · MCP Server · ★ 116
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is coder_eval safe to install? View the security audit →
codereval — evaluate AI coding agents & their skills A framework for evaluating AI coding agents and their skills — built for CLI and skill builders — with sandboxing, reproducibility, and data-driven analysis. Not an "agentic coding" benchmark: it measures how effective your CLI and skills are when used by coding agents. The Coding Agents Gym. A sandboxed, reproducible framework to evaluate, benchmark, and A/B-test AI coding agents — Claude Code, Codex, and Google Antigravity (Gemini) today, any agent via a plugin SPI — with declarative YAML tasks and weighted scoring. Declarative YAML tasks with pinned dependencies
| Stars | 116 |
| Forks | 2 |
| Language | Python |
| Category | MCP Server |
| License | Apache-2.0 |
| Quality Score | 70.6001602054372/100 |
| Open Issues | 23 |
| Last Updated | 2026-08-19 |
| Created | 2026-07-09 |
| Platforms | claude-code, cli, codex, gemini, mcp, python |
| Est. Tokens | ~15k |
These tools work well together with coder_eval for enhanced workflows:
Looking for a coder_eval alternative? If you're comparing coder_eval with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Standalone engineering skills for Claude Code and Codex: review, audit, optimization, testing, product discove
🎯 The definitive collection of 50+ verified Awesome Claude Skills for Claude Code, Claude.ai, and API. Boost
An opinionated, AI-native development workflow for Java Enterprise — reusable Skills, Agents, Commands, and MC
Curated AI coding agent skills and AGENTS.md playbooks for Codex, Claude Code, Cursor, OpenClaw, and other SKI
Terraform Skill for Claude Code and Codex. LLMs hallucinate a lot with Terraform - TerraShark fixes this. It e
HeyClaude (formerly Claude Pro Directory) is a searchable collection of pre-built AI skills, agents, MCP serve
Explore other popular mcp server tools:
coder_eval is Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.. It is categorized as a MCP Server with 116 GitHub stars.
coder_eval is primarily written in Python. It covers topics such as agent-evaluation, agent-skills, agent-testing.
You can find installation instructions and usage details in the coder_eval GitHub repository at github.com/UiPath/coder_eval. The project has 116 stars and 2 forks, indicating an active community.
coder_eval is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to coder_eval on Agent Skills Hub include claude-code-skills, awesome-claude-skills, cursor-rules-java. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.