claw-eval — Codex Skill by claw-eval

by claw-eval · Codex Skill · ★ 742

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is claw-eval safe to install? View the security audit →

About claw-eval

Claw-Eval Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents. 300 human-verified tasks Completion · Safety · Robustness. Leaderboard Browse the full leaderboard and individual task cases at claw-eval.github.io. Evaluation Logic (Updated March 2026): Primary Metric: Pass^3. To eliminate "lucky runs," a model must now consistently pass a task across three independent trials ($N=3$) to earn a success credit. Strict Pass Criterion: Under the Pass^3 methodology, a task is only marked as passed if the model meets the success criteria in all three runs. Reproducibility: We are committed to end-to-end reproducibility. Our co

agentharnessllmopenclaw

Quick Facts

Stars742
Forks70
LanguagePython
CategoryCodex Skill
LicenseMIT
Quality Score70.7067432558119/100
Open Issues8
Last Updated2026-08-07
Created2026-03-11
Platformspython
Est. Tokens~15k

Compatible Skills

These tools work well together with claw-eval for enhanced workflows:

claw-eval alternative? Top 6 similar tools

Looking for a claw-eval alternative? If you're comparing claw-eval with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • memsearch by zilliztech · ⭐ 2.4k

    A persistent, unified memory layer for all your AI agents (e.g. Claude Code, Codex), backed by Markdown and Mi

  • awesome-azure-openai-llm by kimtth · ⭐ 402

    A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.

  • clanker by bgdnvk · ⭐ 371

    autonomous systems engineering cli agent for any cloud environment: AWS, GCP, Cloudflare, etc

  • omnicoreagent by omnirexflora-labs · ⭐ 246

    Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, backgro

  • cipher by campfirein · ⭐ 3.6k

    Byterover Cipher is an opensource memory layer specifically designed for coding agents. Compatible with Cursor

  • deepbot by kevinluosl · ⭐ 2.4k

    DeepBot is a system-level AI assistant built for both personal productivity and enterprise workflows — one-cli

More Codex Skill Tools

Explore other popular codex skill tools:

View all Codex Skill tools →

Popular Python Agent Tools

Frequently Asked Questions

What is claw-eval?

claw-eval is Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.. It is categorized as a Codex Skill with 742 GitHub stars.

What programming language is claw-eval written in?

claw-eval is primarily written in Python. It covers topics such as agent, harness, llm.

How do I install or use claw-eval?

You can find installation instructions and usage details in the claw-eval GitHub repository at github.com/claw-eval/claw-eval. The project has 742 stars and 70 forks, indicating an active community.

What license does claw-eval use?

claw-eval is released under the MIT license, making it free to use and modify according to the license terms.

What are the best alternatives to claw-eval?

The top alternatives to claw-eval on Agent Skills Hub include memsearch, awesome-azure-openai-llm, clanker. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

View on GitHub → Browse Codex Skill tools