No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by agentscope-ai · Codex Skill · ★ 98
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is PawBench safe to install? View the security audit →
🐾 PawBench English · 简体中文 A Model × Harness co-evaluation benchmark for agentic AI. 150 agent tasks · 9 models · 3 harnesses · task slices · diagnostic traces The same model can behave very differently once it is placed inside a real agent runtime. A failure may come from model reasoning, missing tools, weak skill discovery, poor workspace awareness, brittle web access, or a completion check that is too loose. A single final pass rate cannot separate these causes. PawBench is built around one c
| Stars | 98 |
| Forks | 12 |
| Language | Python |
| Category | Codex Skill |
| License | Apache-2.0 |
| Quality Score | 66.7888531271658/100 |
| Open Issues | 6 |
| Last Updated | 2026-08-03 |
| Created | 2026-05-15 |
| Platforms | python |
| Est. Tokens | ~16k |
These tools work well together with PawBench for enhanced workflows:
Looking for a PawBench alternative? If you're comparing PawBench with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
High-value AI skills repository for Codex, Claude Code, OpenClaw, agents, prompts, and automation workflows.
A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.
autonomous systems engineering cli agent for any cloud environment: AWS, GCP, Cloudflare, etc
The Runtime Security Layer for OpenClaw/Hermes-agent, the essential safety harness for PII & sensitive credent
Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, backgro
A banchmark list for evaluation of large language models.
Explore other popular codex skill tools:
PawBench is A benchmark for evaluating LLM × harness performance.. It is categorized as a Codex Skill with 98 GitHub stars.
PawBench is primarily written in Python. It covers topics such as agent, benchmark, harness.
You can find installation instructions and usage details in the PawBench GitHub repository at github.com/agentscope-ai/PawBench. The project has 98 stars and 12 forks, indicating an active community.
PawBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to PawBench on Agent Skills Hub include Commonly-used-high-value-skills, awesome-azure-openai-llm, clanker. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: