No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by sierra-research · LLM Plugin · ★ 2.1k
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is tau2-bench safe to install? View the security audit →
$\tau$-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains 🚀 τ³-bench is here! From text-only to multimodal, knowledge-aware agent evaluation. Voice full-duplex · Knowledge ret
| Stars | 2,062 |
| Forks | 524 |
| Language | Python |
| Category | LLM Plugin |
| License | MIT |
| Quality Score | 73.8759201914111/100 |
| Open Issues | 206 |
| Last Updated | 2026-09-17 |
| Created | 2025-06-09 |
| Platforms | python |
| Est. Tokens | ~15k |
These tools work well together with tau2-bench for enhanced workflows:
Looking for a tau2-bench alternative? If you're comparing tau2-bench with other llm plugin tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Sca
Universal memory layer for AI Agents. It provides scalable, extensible, and interoperable memory storage and r
Klavis AI: MCP integration platforms that let AI agents use tools reliably at any scale
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NP
Adversary simulation and Red teaming platform with AI
The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI
Explore other popular llm plugin tools:
tau2-bench is τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. It is categorized as a LLM Plugin with 2.1k GitHub stars.
tau2-bench is primarily written in Python. It covers topics such as ai, benchmark, conversational-agents.
You can find installation instructions and usage details in the tau2-bench GitHub repository at github.com/sierra-research/tau2-bench. The project has 2.1k stars and 524 forks, indicating an active community.
tau2-bench is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to tau2-bench on Agent Skills Hub include AI-Infra-Guard, MemMachine, klavis. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: