by frontier-harness-eval · Codex Skill · ★ 169
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
Public results and task definitions for FrontierHarness Eval
| Stars | 169 |
| Forks | 10 |
| Language | JavaScript |
| Category | Codex Skill |
| Quality Score | 52.8591681145957/100 |
| Open Issues | 9 |
| Last Updated | 2026-09-08 |
| Created | 2026-08-31 |
| Platforms | claude-code, codex, node |
| Est. Tokens | ~13k |
These tools work well together with eval for enhanced workflows:
Looking for a eval alternative? If you're comparing eval with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
🥇 The strongest free web search plugin for DeepSeek Harness, and the search bridge for every model without na
Deepseek Harness、Openclaw知识图谱记忆插件。2026年4月受邀发布在清华大学讨论会。Knowledge Graph + Memory;Knowledge Graph Context Engine
Run Claude Code, Codex, Antigravity, Cursor Agent and OpenCode as one runtime — persistent sessions, multi-age
One dashboard for all your local AI coding agents. Switch providers, manage sessions, and orchestrate tasks ac
Fast, cross-platform, real-time token usage tracker and cost monitor for Claude Code / Codex CLI / Antigravity
Remote control for AI coding agents. Rust + GPUI + QUIC/UDP. Available on iOS/Android
Explore other popular codex skill tools:
eval is Public results and task definitions for FrontierHarness Eval. It is categorized as a Codex Skill with 169 GitHub stars.
eval is primarily written in JavaScript. It covers topics such as claude-code, codex, deepseek-harness.
You can find installation instructions and usage details in the eval GitHub repository at github.com/frontier-harness-eval/eval. The project has 169 stars and 10 forks, indicating an active community.
The top alternatives to eval on Agent Skills Hub include modsearch, graph-memory, claw-orchestrator. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: