OpenClawProBench — security grade SAFE, quality 69/100

Security audit verdict: SAFE · quality 69/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by suyoumo · Codex Skill · ★ 340

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is OpenClawProBench safe to install? View the security audit →

About OpenClawProBench

OpenClawProBench Transparent live-first benchmark harness for evaluating model capability inside the OpenClaw runtime. 102 active scenarios, 162 catalog scenarios, deterministic grading, and OpenClaw-native coverage. OpenClawProBench focuses on real OpenClaw execution with deterministic grading, structured reports, and benchmark-profile selection. The default ranking path is the profile; broader active coverage remains available through , , , and . The current worktree inventory reports active scenarios and total catalog scenarios ( incubating) via and . Leaderboard Browse the public leaderboard and benchmark cases at suyoumo.github.io/bench. [](https://suyoumo.gi

agentbenchmarkevaluationharnessleaderboardllmopenclaw

Quick Facts

Stars340
Forks26
LanguagePython
CategoryCodex Skill
LicenseApache-2.0
Quality Score69.238284222667/100
Last Updated2026-04-11
Created2025-03-02
Platformspython
Est. Tokens~104k

Compatible Skills

These tools work well together with OpenClawProBench for enhanced workflows:

  • claw-eval — semantic(0.49)+complementary+rare_topics+same_lang+similar_pop+shared_platform (67%)
  • tau2-bench — semantic(0.24)+complementary+rare_topics+same_lang+similar_pop+shared_platform (63%)
  • MCPBench — semantic(0.31)+complementary+rare_topics+same_lang+similar_pop+shared_platform (60%)
  • ollama-benchmark — semantic(0.28)+complementary+rare_topics+same_lang+similar_pop+shared_platform (59%)
  • WildClawBench — semantic(0.38)+complementary+same_lang+similar_pop+shared_platform (58%)

OpenClawProBench alternative? Top 6 similar tools

Looking for a OpenClawProBench alternative? If you're comparing OpenClawProBench with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • claw-eval by claw-eval · ⭐ 750

    Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

  • Awesome-LLM-Eval by onejune2018 · ⭐ 654

    Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, ma

  • OpenJudge by agentscope-ai · ⭐ 778

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • bigcodebench by bigcode-project · ⭐ 519

    [ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI

  • awesome-azure-openai-llm by kimtth · ⭐ 402

    A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.

  • clanker by bgdnvk · ⭐ 371

    autonomous systems engineering cli agent for any cloud environment: AWS, GCP, Cloudflare, etc

More Codex Skill Tools

Explore other popular codex skill tools:

View all Codex Skill tools →

Popular Python Agent Tools

Frequently Asked Questions

What is OpenClawProBench?

OpenClawProBench is OpenClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.. It is categorized as a Codex Skill with 340 GitHub stars.

What programming language is OpenClawProBench written in?

OpenClawProBench is primarily written in Python. It covers topics such as agent, benchmark, evaluation.

How do I install or use OpenClawProBench?

You can find installation instructions and usage details in the OpenClawProBench GitHub repository at github.com/suyoumo/OpenClawProBench. The project has 340 stars and 26 forks, indicating an active community.

What license does OpenClawProBench use?

OpenClawProBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to OpenClawProBench?

The top alternatives to OpenClawProBench on Agent Skills Hub include claw-eval, Awesome-LLM-Eval, OpenJudge. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

View on GitHub → Browse Codex Skill tools