No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by suyoumo · Codex Skill · ★ 823
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is ClawProBench safe to install? View the security audit →
ClawProBench Transparent live-first benchmark harness for evaluating model capability inside the OpenClaw runtime. 102 active scenarios, 162 catalog scenarios, deterministic grading, and OpenClaw-native coverage. ClawProBench focuses on real OpenClaw execution with deterministic grading, structured reports, and benchmark-profile selection. The default ranking path is the profile; broader active coverage remains available through , , , and . The current worktree inventory reports active scenarios and total catalog scenarios ( incubating) via and . Leaderboard Browse the public leaderboard and benchmark cases at suyoumo.github.io/bench. [](https://suyoumo.github.io
| Stars | 823 |
| Forks | 54 |
| Language | Rust |
| Category | Codex Skill |
| License | Apache-2.0 |
| Quality Score | 68.4435401153524/100 |
| Last Updated | 2026-08-25 |
| Created | 2025-03-02 |
| Platforms | rust |
| Est. Tokens | ~23k |
Looking for a ClawProBench alternative? If you're comparing ClawProBench with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, ma
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, eva
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.
Explore other popular codex skill tools:
ClawProBench is ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.. It is categorized as a Codex Skill with 823 GitHub stars.
ClawProBench is primarily written in Rust. It covers topics such as agent, benchmark, evaluation.
You can find installation instructions and usage details in the ClawProBench GitHub repository at github.com/suyoumo/ClawProBench. The project has 823 stars and 54 forks, indicating an active community.
ClawProBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to ClawProBench on Agent Skills Hub include claw-eval, Awesome-LLM-Eval, trpc-agent-go. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: