Flagged: sudo usage. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by xingyaoww · Agent Tool · ★ 141
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is mint-bench safe to install? View the security audit →
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback Official Repo for paper MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback by Xingyao Wang\, Zihan Wang\, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng and Heng Ji. MINT benchmark aims to evaluate LLMs' ability to solve tasks with multi-turn interactions by (1) using tools and (2) leveraging natural language feedback. :trophy: Please visit our website for the leaderboard. :warning: WARNING: Evaluation of LLMs requires executing untrusted model-generated code. Users are strongly encouraged to sandbox the code execution so that it does not perform destructive actions on their host or network. We highly recommend using the provided docker image for isolated execution. :rocket: Quick Start Environment Setup You can choose to use docker (recommended) or local setup as follows. Docker Setup (Recommended) You only need to ensure that you have docker installed on your local computer following [the official guide](https://docs.docker.co
| Stars | 141 |
| Forks | 11 |
| Language | Python |
| Category | Agent Tool |
| License | Apache-2.0 |
| Quality Score | 67.1695507218813/100 |
| Last Updated | 2024-06-04 |
| Created | 2023-09-18 |
| Platforms | python |
| Est. Tokens | ~4821k |
Looking for a mint-bench alternative? If you're comparing mint-bench with other agent tool tools, these 3 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Lightweight registry to discover, install, and manage all public Claude plugins and agent skills for your favo
A Claude Code skill that turns PDFs, docs, and codebases into Obsidian study vaults
Power rename/refactor tool for CLI and agent use
Explore other popular agent tool tools:
mint-bench is Official Repo for ICLR 2024 paper MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback by Xingyao Wang*, Zihan Wang*, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng and Hen. It is categorized as a Agent Tool with 141 GitHub stars.
mint-bench is primarily written in Python.
You can find installation instructions and usage details in the mint-bench GitHub repository at github.com/xingyaoww/mint-bench. The project has 141 stars and 11 forks, indicating an active community.
mint-bench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to mint-bench on Agent Skills Hub include claude-plugins, tutor-skills, repren. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: