llm-d-benchmark — security grade SAFE, quality 66/100

Security audit verdict: SAFE · quality 66/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by llm-d · Agent Tool · ★ 71

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is llm-d-benchmark safe to install? View the security audit →

About llm-d-benchmark

llm-d-benchmark [![.github/workflows/ci-nightly-benchmark-ocp-standalone.yaml](https://github.com/llm-d/llm-d-benchmark/actions/workflows/

incubating

Quick Facts

Stars71
Forks148
LanguagePython
CategoryAgent Tool
LicenseApache-2.0
Quality Score65.6510863499555/100
Open Issues47
Last Updated2026-10-02
Created2025-05-13
Platformspython
Est. Tokens~25k

More Agent Tool Tools

Explore other popular agent tool tools:

View all Agent Tool tools →

Popular Python Agent Tools

Frequently Asked Questions

What is llm-d-benchmark?

llm-d-benchmark is llm-d benchmark scripts and tooling. It is categorized as a Agent Tool with 71 GitHub stars.

What programming language is llm-d-benchmark written in?

llm-d-benchmark is primarily written in Python. It covers topics such as incubating.

How do I install or use llm-d-benchmark?

You can find installation instructions and usage details in the llm-d-benchmark GitHub repository at github.com/llm-d/llm-d-benchmark. The project has 71 stars and 148 forks, indicating an active community.

What license does llm-d-benchmark use?

llm-d-benchmark is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Agent Tool tools