No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by IBM · Agent Tool · ★ 68
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is vakra safe to install? View the security audit →
🔷 VAKRA: A Benchmark for Evaluating Multi-Hop, Multi-Source Tool-Calling in AI Agents VAKRA (eValuating API and Knowledge Retrieval Agents using multi-hop, multi-source dialogues) is a tool-grounded, executable benchmark designed to evaluate how well AI agents reason end-to-end in enterprise-like settings. Rather than testing isolated skills, VAKRA measures compositional reasoning across APIs and documents, using full execution traces to assess whether agents can reliably complete multi-step workflows, not just individual steps. VAKRA provides an executable environment where agents interact with over 8,000 locally hosted APIs backed by real databases spanning 62 domains, along with domain-aligned document collections. Resources: Leaderboard · Dataset · Blog Quick links: Requirements · Quick Start · Exploring Available Tools · Running Your Agent · Submit to Leaderboard What VAKRA Provides An executable benchmark environment with 8,000+ locally hosted APIs backed by real databases across 62 domains Domain-aligned document collections for retrieval-augmented, cross-source reasoning Tasks that require 3-7 step reasoning chains acros
| Stars | 68 |
| Forks | 8 |
| Language | Python |
| Category | Agent Tool |
| Quality Score | 59.841054885001/100 |
| Open Issues | 8 |
| Last Updated | 2026-09-22 |
| Created | 2026-02-25 |
| Platforms | python |
| Est. Tokens | ~21k |
These tools work well together with vakra for enhanced workflows:
Explore other popular agent tool tools:
vakra is A Benchmark for Evaluating Multi-Hop, Multi-Source Tool-Calling in AI Agents. It is categorized as a Agent Tool with 68 GitHub stars.
vakra is primarily written in Python.
You can find installation instructions and usage details in the vakra GitHub repository at github.com/IBM/vakra. The project has 68 stars and 8 forks, indicating an active community.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: