tau2-bench — security grade SAFE, quality 74/100

Security audit verdict: SAFE · quality 74/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by sierra-research · LLM Plugin · ★ 2.1k

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is tau2-bench safe to install? View the security audit →

About tau2-bench

$\tau$-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains 🚀 τ³-bench is here! From text-only to multimodal, knowledge-aware agent evaluation. Voice full-duplex · Knowledge ret

aibenchmarkconversational-agentslanguage-model-agentllm

Quick Facts

Stars2,062
Forks524
LanguagePython
CategoryLLM Plugin
LicenseMIT
Quality Score73.8759201914111/100
Open Issues206
Last Updated2026-09-17
Created2025-06-09
Platformspython
Est. Tokens~15k

Compatible Skills

These tools work well together with tau2-bench for enhanced workflows:

  • mcpmark — semantic(0.42)+complementary+rare_topics+same_lang+similar_pop+shared_platform (64%)
  • MCPBench — semantic(0.27)+complementary+rare_topics+same_lang+similar_pop+shared_platform (59%)
  • MLLM-Tool — semantic(0.35)+complementary+same_lang+similar_pop+shared_platform (57%)
  • Toucan — semantic(0.23)+complementary+same_lang+similar_pop+shared_platform (53%)
  • mcp-bench — semantic(0.20)+complementary+same_lang+similar_pop+shared_platform (52%)

tau2-bench alternative? Top 6 similar tools

Looking for a tau2-bench alternative? If you're comparing tau2-bench with other llm plugin tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • AI-Infra-Guard by Tencent · ⭐ 6.4k

    A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Sca

  • MemMachine by MemMachine · ⭐ 3.2k

    Universal memory layer for AI Agents. It provides scalable, extensible, and interoperable memory storage and r

  • klavis by Klavis-AI · ⭐ 5.8k

    Klavis AI: MCP integration platforms that let AI agents use tools reliably at any scale

  • lemonade by lemonade-sdk · ⭐ 5.8k

    Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NP

  • Viper by FunnyWolf · ⭐ 5.3k

    Adversary simulation and Red teaming platform with AI

  • code-graph-rag by vitali87 · ⭐ 5.2k

    The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI

More LLM Plugin Tools

Explore other popular llm plugin tools:

View all LLM Plugin tools →

Popular Python Agent Tools

Frequently Asked Questions

What is tau2-bench?

tau2-bench is τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. It is categorized as a LLM Plugin with 2.1k GitHub stars.

What programming language is tau2-bench written in?

tau2-bench is primarily written in Python. It covers topics such as ai, benchmark, conversational-agents.

How do I install or use tau2-bench?

You can find installation instructions and usage details in the tau2-bench GitHub repository at github.com/sierra-research/tau2-bench. The project has 2.1k stars and 524 forks, indicating an active community.

What license does tau2-bench use?

tau2-bench is released under the MIT license, making it free to use and modify according to the license terms.

What are the best alternatives to tau2-bench?

The top alternatives to tau2-bench on Agent Skills Hub include AI-Infra-Guard, MemMachine, klavis. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse LLM Plugin tools