LLM-Agent-Benchmark-List — security grade SAFE, quality 54/100

Security audit verdict: SAFE · quality 54/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by zhangxjohn · Agent Tool · ★ 169

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is LLM-Agent-Benchmark-List safe to install? View the security audit →

About LLM-Agent-Benchmark-List

LLM-Agent-Benchmark-List 🤗We greatly appreciate any contributions via PRs, issues, emails, or other methods. ⏳ Continuous update... :book: Introduction In the swiftly evolving landscape of artificial intelligence, Large Language Models (LLMs) have emerged as a pivotal cornerstone, revolutionizing how we interact with and harness the power of natural language processing. However, as LLMs gain widespread application in both research and industry sectors, the imperative shifts towards evaluating their efficacy rather than perpetuating a cycle of unbridled performance iterations. This paradigm shift raises critical questions: i) what to evaluate? ii) where to evaluate? iii)How to evaluate? Diverse research endeavors have proposed varying interpretations and methodologies in response to these queries. The aim of this work is to methodically review and organize benchmarks that are both LLMs and agent-powered, thereby providing a streamlined resource for those journeying towards Artificial General Intelligence (AGI). :dizzy: List Survey [2023/07] A Survey on Evaluation of Large Language Models. Yupeng Chang ( Jilin University) et al. arXiv.

agentbenchmarklarge-language-modelsllmsurvey

Quick Facts

Stars169
Forks12
CategoryAgent Tool
LicenseApache-2.0
Quality Score53.6349017474777/100
Open Issues3
Last Updated2026-08-19
Created2024-01-29
Est. Tokens~19k

LLM-Agent-Benchmark-List alternative? Top 6 similar tools

Looking for a LLM-Agent-Benchmark-List alternative? If you're comparing LLM-Agent-Benchmark-List with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • bigcodebench by bigcode-project · ⭐ 519

    [ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI

  • LLM-Tool-Survey by quchangle1 · ⭐ 489

    This is the repository for the Tool Learning survey.

  • OpenRCA by microsoft · ⭐ 401

    [ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?

  • awesome-lifelong-llm-agent by qianlima-lab · ⭐ 329

    TPAMI 2026 | This repository collects awesome survey, resource, and paper for lifelong learning LLM agents

  • LLM-Brained-GUI-Agents-Survey by vyokky · ⭐ 231

    GitHub page for "Large Language Model-Brained GUI Agents: A Survey"

  • Awesome-LLM-Papers-Comprehensive-Topics by shure-dev · ⭐ 223

    Awesome LLM Papers and repos on very comprehensive topics.

More Agent Tool Tools

Explore other popular agent tool tools:

View all Agent Tool tools →

Frequently Asked Questions

What is LLM-Agent-Benchmark-List?

LLM-Agent-Benchmark-List is A banchmark list for evaluation of large language models.. It is categorized as a Agent Tool with 169 GitHub stars.

How do I install or use LLM-Agent-Benchmark-List?

You can find installation instructions and usage details in the LLM-Agent-Benchmark-List GitHub repository at github.com/zhangxjohn/LLM-Agent-Benchmark-List. The project has 169 stars and 12 forks, indicating an active community.

What license does LLM-Agent-Benchmark-List use?

LLM-Agent-Benchmark-List is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to LLM-Agent-Benchmark-List?

The top alternatives to LLM-Agent-Benchmark-List on Agent Skills Hub include bigcodebench, LLM-Tool-Survey, OpenRCA. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Agent Tool tools