Awesome-LLM-Eval — security grade SAFE, quality 59/100

Security audit verdict: SAFE · quality 59/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by onejune2018 · Agent Tool · ★ 654

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is Awesome-LLM-Eval safe to install? View the security audit →

About Awesome-LLM-Eval

Awesome LLM Eval English | 中文 Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI. The is the official project of our survey: Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap. NOTE: As we cannot update the arXiv paper in real time, please refer to this repo for the latest updates and the paper may be updated later. We also welcome any pull request or issues to help us improve this work. Your contributions will be acknowledged in acknowledgements. If you find our survey useful, please kindly cite our paper: Table of Contents News [Tools]

awsome-listawsome-listsbenchmarkbertchatglmchatgptdatasetevaluationgpt3large-language-model

Quick Facts

Stars654
Forks83
CategoryAgent Tool
LicenseMIT
Quality Score58.6727772105924/100
Open Issues46
Last Updated2025-11-24
Created2023-04-26
Platformsaws
Est. Tokens~1132k

Awesome-LLM-Eval alternative? Top 6 similar tools

Looking for a Awesome-LLM-Eval alternative? If you're comparing Awesome-LLM-Eval with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • spacy-llm by explosion · ⭐ 1.4k

    🦙 Integrating LLMs into structured NLP pipelines

  • cai by ad-si · ⭐ 202

    User friendly CLI tool for AI tasks. Stop thinking about LLMs and prompts, start getting results!

  • sre by SmythOS · ⭐ 1.3k

    The SmythOS Runtime Environment (SRE) is an open-source, cloud-native runtime for agentic AI. Secure, modular,

  • tools by strands-agents · ⭐ 1.3k

    A set of tools that gives agents powerful capabilities.

  • samples by strands-agents · ⭐ 852

    Agent samples built using the Strands Agents SDK.

  • ClawProBench by suyoumo · ⭐ 823

    ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with determ

More Agent Tool Tools

Explore other popular agent tool tools:

View all Agent Tool tools →

Frequently Asked Questions

What is Awesome-LLM-Eval?

Awesome-LLM-Eval is Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.. It is categorized as a Agent Tool with 654 GitHub stars.

How do I install or use Awesome-LLM-Eval?

You can find installation instructions and usage details in the Awesome-LLM-Eval GitHub repository at github.com/onejune2018/Awesome-LLM-Eval. The project has 654 stars and 83 forks, indicating an active community.

What license does Awesome-LLM-Eval use?

Awesome-LLM-Eval is released under the MIT license, making it free to use and modify according to the license terms.

What are the best alternatives to Awesome-LLM-Eval?

The top alternatives to Awesome-LLM-Eval on Agent Skills Hub include spacy-llm, cai, sre. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Agent Tool tools