AgentBench — security grade SAFE, quality 67/100

Security audit verdict: SAFE · quality 67/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by THUDM · Agent Tool · ★ 3.7k

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is AgentBench safe to install? View the security audit →

About AgentBench

AgentBench 🌐 Leaderboard (new) 📃 Paper 👋 Join our Slack for Q & A or collaboration on next version of AgentBench! 🔥[2025.10.10] Introducing AgentBench FC (Function Calling) based on AgentRL The current repository contains the function-calling version of AgentBench, integrated with AgentRL, an end-to-end multitask and mutliturn LLM Agent RL framework. If you wish to use the older version, you can revert to v0.1 and v0.2. Comparing to the original AgentBench, this version uses a function-calling style prompt, and adds fully-containerized deployment support for the following tasks: (AF) (DB) (KG) (OS) (WS) Quick Start We support a

chatgptgpt-4llmllm-agent

Quick Facts

Stars3,742
Forks280
LanguagePython
CategoryAgent Tool
LicenseApache-2.0
Quality Score67.438242037414/100
Open Issues77
Last Updated2026-02-08
Created2023-07-28
Platformspython
Est. Tokens~1970k

AgentBench alternative? Top 6 similar tools

Looking for a AgentBench alternative? If you're comparing AgentBench with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • langroid by langroid · ⭐ 4.1k

    Harness LLMs with Multi-Agent Programming

  • awesome-generative-ai by filipecalegario · ⭐ 3.5k

    A curated list of Generative AI tools, works, models, and references

  • gptme by gptme · ⭐ 4.4k

    Your agent in your terminal, equipped with local tools: writes code, uses the terminal, browses the web. Make

  • DecryptPrompt by DSXiangLi · ⭐ 3.4k

    总结Prompt&LLM论文,开源数据&模型,AIGC应用

  • aide by nicepkg · ⭐ 2.7k

    Conquer Any Code in VSCode: One-Click Comments, Conversions, UI-to-Code, and AI Batch Processing of Files! 在 V

  • awesome-llm-powered-agent by hyp1231 · ⭐ 2.3k

    Awesome things about LLM-powered agents. Papers / Repos / Blogs / ...

More Agent Tool Tools

Explore other popular agent tool tools:

View all Agent Tool tools →

Popular Python Agent Tools

Frequently Asked Questions

What is AgentBench?

AgentBench is A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24). It is categorized as a Agent Tool with 3.7k GitHub stars.

What programming language is AgentBench written in?

AgentBench is primarily written in Python. It covers topics such as chatgpt, gpt-4, llm.

How do I install or use AgentBench?

You can find installation instructions and usage details in the AgentBench GitHub repository at github.com/THUDM/AgentBench. The project has 3.7k stars and 280 forks, indicating an active community.

What license does AgentBench use?

AgentBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to AgentBench?

The top alternatives to AgentBench on Agent Skills Hub include langroid, awesome-generative-ai, gptme. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Agent Tool tools