by InternLM · Codex Skill · ★ 452
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
WildClawBench []() []() Hard, practical, end-to-end evaluation for AI agents — in the wild. WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end-to-end, without hand-holding? We drop agents into a live OpenClaw environment — the same open-source personal AI assistant that real users rely on daily — and throw 60 original tasks at them: clipping goal highlights from a football match, negotiating meeting times over multi-round emails, hunting down contradictions in search results, writing inference scripts for undocumented codebases, catching privacy leaks before they happen. Useful things. Hard things. Hard enough that every frontier model we tested scores below 0.55 (top overall 0.52). That makes scores mean something. Why WildClawBench? Most agent benchmarks test isolated capabilities — calling a function, parsing JSON, following a sing
| Stars | 452 |
| Forks | 45 |
| Language | Python |
| Category | Codex Skill |
| License | MIT |
| Quality Score | 71.2414881444168/100 |
| Open Issues | 5 |
| Last Updated | 2026-06-25 |
| Created | 2026-03-23 |
| Platforms | python |
| Est. Tokens | ~18k |
These tools work well together with WildClawBench for enhanced workflows:
Looking for a WildClawBench alternative? If you're comparing WildClawBench with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
A simple yet powerful agent framework for personal assistants, designed to enable intelligent interaction, mul
A set of tools that gives agents powerful capabilities.
A curated list of OpenClaw resources, tools, skills, tutorials & articles. OpenClaw (formerly Moltbot / Clawdb
Model-agnostic plug-n-play LangChain/LangGraph agents powered entirely by MCP tools over HTTP/SSE.
Enterprise-ready MCP Gateway & Registry that centralizes AI development tools with secure OAuth authentication
AI Agent Orchestrator with Skills System - Give AI Agents superpowers: memory search, code graph queries, agen
Explore other popular codex skill tools:
WildClawBench is An in-the-wild benchmark for AI agents in the OpenClaw Environment.. It is categorized as a Codex Skill with 452 GitHub stars.
WildClawBench is primarily written in Python. It covers topics such as agentic-ai, agentic-evaluation, agents.
You can find installation instructions and usage details in the WildClawBench GitHub repository at github.com/InternLM/WildClawBench. The project has 452 stars and 45 forks, indicating an active community.
WildClawBench is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to WildClawBench on Agent Skills Hub include NagaAgent, tools, awesome-openclaw. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.