by claw-eval · Codex Skill · ★ 742
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is claw-eval safe to install? View the security audit →
Claw-Eval Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents. 300 human-verified tasks Completion · Safety · Robustness. Leaderboard Browse the full leaderboard and individual task cases at claw-eval.github.io. Evaluation Logic (Updated March 2026): Primary Metric: Pass^3. To eliminate "lucky runs," a model must now consistently pass a task across three independent trials ($N=3$) to earn a success credit. Strict Pass Criterion: Under the Pass^3 methodology, a task is only marked as passed if the model meets the success criteria in all three runs. Reproducibility: We are committed to end-to-end reproducibility. Our co
| Stars | 742 |
| Forks | 70 |
| Language | Python |
| Category | Codex Skill |
| License | MIT |
| Quality Score | 70.7067432558119/100 |
| Open Issues | 8 |
| Last Updated | 2026-08-07 |
| Created | 2026-03-11 |
| Platforms | python |
| Est. Tokens | ~15k |
These tools work well together with claw-eval for enhanced workflows:
Looking for a claw-eval alternative? If you're comparing claw-eval with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
A persistent, unified memory layer for all your AI agents (e.g. Claude Code, Codex), backed by Markdown and Mi
A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.
autonomous systems engineering cli agent for any cloud environment: AWS, GCP, Cloudflare, etc
Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, backgro
Byterover Cipher is an opensource memory layer specifically designed for coding agents. Compatible with Cursor
DeepBot is a system-level AI assistant built for both personal productivity and enterprise workflows — one-cli
Explore other popular codex skill tools:
claw-eval is Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.. It is categorized as a Codex Skill with 742 GitHub stars.
claw-eval is primarily written in Python. It covers topics such as agent, harness, llm.
You can find installation instructions and usage details in the claw-eval GitHub repository at github.com/claw-eval/claw-eval. The project has 742 stars and 70 forks, indicating an active community.
claw-eval is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to claw-eval on Agent Skills Hub include memsearch, awesome-azure-openai-llm, clanker. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.