No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by edonadei · MCP Server · ★ 50
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is caliper safe to install? View the security audit →
Caliper Reliability testing for agent skills. Define what success looks like, run your skill k times, and get a pass@k score you can track and compare. Agent skills are hard to test. A skill that works on your machine, on this prompt, today, might fail tomorrow after a model update or a one-line prompt edit. Caliper makes reliability measurable: define what success looks like, run the skill repeatedly, and get a pass@k score you can track over time. Use Caliper to answer questions like: Did my prompt edit actually improve the skill? Is the skill doing the work, or would the base agent pass without it? Does it still pass the workflows it passed last week? Which backend — Claude Code or Codex — runs this skill more reliably? Caliper runs each task with and without the skill, and shows you the difference:
| Stars | 50 |
| Forks | 8 |
| Language | Python |
| Category | MCP Server |
| License | MIT |
| Quality Score | 65.9946625043157/100 |
| Open Issues | 8 |
| Last Updated | 2026-08-30 |
| Created | 2026-05-13 |
| Platforms | claude-code, cli, codex, mcp, python |
| Est. Tokens | ~17k |
Looking for a caliper alternative? If you're comparing caliper with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Local-first coordination for human and agent work: durable work, decisions, dispatches, evidence, and prompt-f
Unified CLI for running AI coding agents in isolated containers. Includes built-in local metrics collection, H
AI agent 通用任务治理框架:对齐目标与事实,规划和调度能力,守住授权与风险边界,治理任务执行到真实验收与交付。Governance framework for evidence-driven planning,
A collection of skills for AI agents (Kiro, Cursor, Windsurf, Claude Code, and others). Each skill is a reusab
Code from anywhere — Telegram bridge for AI coding agents (Claude Code, Codex, OpenCode, Pi, Gemini CLI, Amp).
Battle-tested skill library for AI agents. Save 98% of API costs with ready-to-use code for crypto, PDFs, sear
Explore other popular mcp server tools:
caliper is Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.. It is categorized as a MCP Server with 50 GitHub stars.
caliper is primarily written in Python. It covers topics such as ai-agents, claude-code, cli.
You can find installation instructions and usage details in the caliper GitHub repository at github.com/edonadei/caliper. The project has 50 stars and 8 forks, indicating an active community.
caliper is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to caliper on Agent Skills Hub include maestro, vibepod-cli, odai. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: