ClawProBench — security grade SAFE, quality 68/100

Security audit verdict: SAFE · quality 68/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by suyoumo · Codex Skill · ★ 823

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is ClawProBench safe to install? View the security audit →

About ClawProBench

ClawProBench Transparent live-first benchmark harness for evaluating model capability inside the OpenClaw runtime. 102 active scenarios, 162 catalog scenarios, deterministic grading, and OpenClaw-native coverage. ClawProBench focuses on real OpenClaw execution with deterministic grading, structured reports, and benchmark-profile selection. The default ranking path is the profile; broader active coverage remains available through , , , and . The current worktree inventory reports active scenarios and total catalog scenarios ( incubating) via and . Leaderboard Browse the public leaderboard and benchmark cases at suyoumo.github.io/bench. [](https://suyoumo.github.io

agentbenchmarkevaluationharnessleaderboardllmopenclaw

Quick Facts

Stars823
Forks54
LanguageRust
CategoryCodex Skill
LicenseApache-2.0
Quality Score68.4435401153524/100
Last Updated2026-08-25
Created2025-03-02
Platformsrust
Est. Tokens~23k

ClawProBench alternative? Top 6 similar tools

Looking for a ClawProBench alternative? If you're comparing ClawProBench with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • claw-eval by claw-eval · ⭐ 769

    Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

  • Awesome-LLM-Eval by onejune2018 · ⭐ 654

    Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, ma

  • trpc-agent-go by trpc-group · ⭐ 1.8k

    A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, eva

  • OpenJudge by agentscope-ai · ⭐ 829

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • bigcodebench by bigcode-project · ⭐ 519

    [ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI

  • awesome-azure-openai-llm by kimtth · ⭐ 402

    A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.

More Codex Skill Tools

Explore other popular codex skill tools:

View all Codex Skill tools →

Popular Rust Agent Tools

Frequently Asked Questions

What is ClawProBench?

ClawProBench is ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.. It is categorized as a Codex Skill with 823 GitHub stars.

What programming language is ClawProBench written in?

ClawProBench is primarily written in Rust. It covers topics such as agent, benchmark, evaluation.

How do I install or use ClawProBench?

You can find installation instructions and usage details in the ClawProBench GitHub repository at github.com/suyoumo/ClawProBench. The project has 823 stars and 54 forks, indicating an active community.

What license does ClawProBench use?

ClawProBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to ClawProBench?

The top alternatives to ClawProBench on Agent Skills Hub include claw-eval, Awesome-LLM-Eval, trpc-agent-go. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Codex Skill tools