PawBench — security grade SAFE, quality 67/100

Security audit verdict: SAFE · quality 67/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by agentscope-ai · Codex Skill · ★ 98

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is PawBench safe to install? View the security audit →

About PawBench

🐾 PawBench English · 简体中文 A Model × Harness co-evaluation benchmark for agentic AI. 150 agent tasks · 9 models · 3 harnesses · task slices · diagnostic traces The same model can behave very differently once it is placed inside a real agent runtime. A failure may come from model reasoning, missing tools, weak skill discovery, poor workspace awareness, brittle web access, or a completion check that is too loose. A single final pass rate cannot separate these causes. PawBench is built around one c

agentbenchmarkharnesshermesllmopenclawqwenpaw

Quick Facts

Stars98
Forks12
LanguagePython
CategoryCodex Skill
LicenseApache-2.0
Quality Score66.7888531271658/100
Open Issues6
Last Updated2026-08-03
Created2026-05-15
Platformspython
Est. Tokens~16k

Compatible Skills

These tools work well together with PawBench for enhanced workflows:

  • OpenClawProBench — semantic(0.72)+same_lang+similar_pop+shared_platform (60%)
  • a11y-llm-eval — semantic(0.23)+complementary+same_lang+similar_pop+shared_platform (58%)
  • ollama-benchmark — semantic(0.37)+complementary+same_lang+similar_pop+shared_platform (58%)
  • llm-d-benchmark — semantic(0.37)+complementary+same_lang+similar_pop+shared_platform (58%)

PawBench alternative? Top 6 similar tools

Looking for a PawBench alternative? If you're comparing PawBench with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • Commonly-used-high-value-skills by seaworld008 · ⭐ 70

    High-value AI skills repository for Codex, Claude Code, OpenClaw, agents, prompts, and automation workflows.

  • awesome-azure-openai-llm by kimtth · ⭐ 402

    A curated collection of resources for 🌌 Azure OpenAI, 🦙 LLMs (+RAG, Agents). Monthly Updates.

  • clanker by bgdnvk · ⭐ 380

    autonomous systems engineering cli agent for any cloud environment: AWS, GCP, Cloudflare, etc

  • clawshell by clawshell · ⭐ 300

    The Runtime Security Layer for OpenClaw/Hermes-agent, the essential safety harness for PII & sensitive credent

  • omnicoreagent by omnirexflora-labs · ⭐ 246

    Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, backgro

  • LLM-Agent-Benchmark-List by zhangxjohn · ⭐ 169

    A banchmark list for evaluation of large language models.

More Codex Skill Tools

Explore other popular codex skill tools:

View all Codex Skill tools →

Popular Python Agent Tools

Frequently Asked Questions

What is PawBench?

PawBench is A benchmark for evaluating LLM × harness performance.. It is categorized as a Codex Skill with 98 GitHub stars.

What programming language is PawBench written in?

PawBench is primarily written in Python. It covers topics such as agent, benchmark, harness.

How do I install or use PawBench?

You can find installation instructions and usage details in the PawBench GitHub repository at github.com/agentscope-ai/PawBench. The project has 98 stars and 12 forks, indicating an active community.

What license does PawBench use?

PawBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to PawBench?

The top alternatives to PawBench on Agent Skills Hub include Commonly-used-high-value-skills, awesome-azure-openai-llm, clanker. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Codex Skill tools