No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by NVIDIA · Codex Skill · ★ 533
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is SkillEvaluator safe to install? View the security audit →
SkillEvaluator SkillEvaluator is an open-source, multi-tier framework for evaluating AI agent artifacts, starting with agent skills: deterministic quality gates, semantic overlap detection, synthetic eval dataset generation, and live agent evaluation. Agent skills are folders of instructions and supporting files that extend AI agents, as defined by the Agent Skills specification. SkillEvaluator is part of the NVIDIA Verified Skills pipeline. Three-tier overview Tiers are independent entry points; nothing requires running earlier ones first. No API key for deterministic checks; the extra plus external Semgrep, SkillSpector, and Gitleaks for
| Stars | 533 |
| Forks | 59 |
| Language | Python |
| Category | Codex Skill |
| License | Apache-2.0 |
| Quality Score | 70.2814615592248/100 |
| Open Issues | 26 |
| Last Updated | 2026-10-02 |
| Created | 2026-06-24 |
| Platforms | claude-code, codex, python |
| Est. Tokens | ~15k |
These tools work well together with SkillEvaluator for enhanced workflows:
Looking for a SkillEvaluator alternative? If you're comparing SkillEvaluator with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Offline security scanner for AI-agent repos, skills, plugins, and MCP servers.
Universal Claude Code workflow plugin with agents, skills, hooks, and commands
Professional context and harness engineering for Claude Code and OpenAI Codex. Build production-grade software
Supercharge AI coding agents with portable skills. Install, translate & share skills across Claude Code, Curso
An IOS Simulator Skill for ClaudeCode. Use it to optimise Claude's ability to build, run and interact with you
Lightweight Agent Workstation for Codex + Claude code — with dots, task scheduler, git worktree & remote contr
Explore other popular codex skill tools:
SkillEvaluator is Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect a. It is categorized as a Codex Skill with 533 GitHub stars.
SkillEvaluator is primarily written in Python. It covers topics such as agent-evaluation, agent-security, agent-skills.
You can find installation instructions and usage details in the SkillEvaluator GitHub repository at github.com/NVIDIA/SkillEvaluator. The project has 533 stars and 59 forks, indicating an active community.
SkillEvaluator is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to SkillEvaluator on Agent Skills Hub include repo-forensics, claude-workflow-v2, pilot-shell. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: