No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by aiverify-foundation · Agent Tool · ★ 344
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is moonshot safe to install? View the security audit →
Version 0.7.6 A simple and modular tool to evaluate any LLM-based AI systems. 🎯 Motivation Developed by the AI Verify Foundation, Moonshot is a tool to bring Benchmarking and Red-Teaming together to help AI developers, compliance teams evaluate LLM-based Apps and LLMs. 🚀 Why Moonshot In the rapidly evolving landscape of Generative AI, ensuring safety, reliability, and performance of LLM applications is paramount. Moonshot addresses this critical need by providing a unified platform for: Benchmark Tests: Systematically test LLM Apps or LLMs across critical trust & safety risks using a wide array of open-source benchmark dataset and metrics, including guided workflows to implement IMDA's Starter Kit for LLM-based App Testing. Red Team Attacks: Proactively identify vulnerabilities and potential misuse scenarios in your LLM applications through streamlined adversarial prompting. 🔑 Key Features User-friendly Interfaces: Interact with Moonshot via an intuitive Web UI for visual insights, and an interactive Command Line Interface (CLI) for quick operations. Comprehensiv
| Stars | 344 |
| Forks | 67 |
| Language | Python |
| Category | Agent Tool |
| License | Apache-2.0 |
| Quality Score | 68.7063525758279/100 |
| Open Issues | 1 |
| Last Updated | 2026-06-10 |
| Created | 2023-12-14 |
| Platforms | python |
| Est. Tokens | ~15282k |
These tools work well together with moonshot for enhanced workflows:
Looking for a moonshot alternative? If you're comparing moonshot with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
A security scanner for your LLM agentic workflows
Security toolkit for AI agents. Scan your machine for dangerous skills and MCP configs, monitor for supply cha
Bayesian Optimization as a Coverage Tool for Evaluating LLMs. Accurate evaluation (benchmarking) that's 10 tim
Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)
Security toolkit for AI agents. Scan your machine for dangerous skills and MCP configs, monitor for supply cha
A curated list of awesome LLM Red Teaming training, resources, and tools.
Explore other popular agent tool tools:
moonshot is Moonshot - A simple and modular tool to evaluate and red-team any LLM application.. It is categorized as a Agent Tool with 344 GitHub stars.
moonshot is primarily written in Python. It covers topics such as benchmarking, evaluation-framework, llm.
You can find installation instructions and usage details in the moonshot GitHub repository at github.com/aiverify-foundation/moonshot. The project has 344 stars and 67 forks, indicating an active community.
moonshot is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to moonshot on Agent Skills Hub include agentic-radar, agentseal, bocoel. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: