caliper — security grade SAFE, quality 66/100

Security audit verdict: SAFE · quality 66/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by edonadei · MCP Server · ★ 50

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is caliper safe to install? View the security audit →

About caliper

Caliper Reliability testing for agent skills. Define what success looks like, run your skill k times, and get a pass@k score you can track and compare. Agent skills are hard to test. A skill that works on your machine, on this prompt, today, might fail tomorrow after a model update or a one-line prompt edit. Caliper makes reliability measurable: define what success looks like, run the skill repeatedly, and get a pass@k score you can track over time. Use Caliper to answer questions like: Did my prompt edit actually improve the skill? Is the skill doing the work, or would the base agent pass without it? Does it still pass the workflows it passed last week? Which backend — Claude Code or Codex — runs this skill more reliably? Caliper runs each task with and without the skill, and shows you the difference:

ai-agentsclaude-codeclicodexdsh-plugindsh-plugin-marketdsh-pluginsevalsevaluationhermes

Quick Facts

Stars50
Forks8
LanguagePython
CategoryMCP Server
LicenseMIT
Quality Score65.9946625043157/100
Open Issues8
Last Updated2026-08-30
Created2026-05-13
Platformsclaude-code, cli, codex, mcp, python
Est. Tokens~17k

caliper alternative? Top 6 similar tools

Looking for a caliper alternative? If you're comparing caliper with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • maestro by ReinaMacCredy · ⭐ 231

    Local-first coordination for human and agent work: durable work, decisions, dispatches, evidence, and prompt-f

  • vibepod-cli by VibePod · ⭐ 155

    Unified CLI for running AI coding agents in isolated containers. Includes built-in local metrics collection, H

  • odai by orziz · ⭐ 111

    AI agent 通用任务治理框架:对齐目标与事实,规划和调度能力,守住授权与风险边界,治理任务执行到真实验收与交付。Governance framework for evidence-driven planning,

  • skills by fabricioctelles · ⭐ 77

    A collection of skills for AI agents (Kiro, Cursor, Windsurf, Claude Code, and others). Each skill is a reusab

  • untether by littlebearapps · ⭐ 62

    Code from anywhere — Telegram bridge for AI coding agents (Claude Code, Codex, OpenCode, Pi, Gemini CLI, Amp).

  • open-skills by besoeasy · ⭐ 132

    Battle-tested skill library for AI agents. Save 98% of API costs with ready-to-use code for crypto, PDFs, sear

More MCP Server Tools

Explore other popular mcp server tools:

View all MCP Server tools →

Popular Python Agent Tools

Frequently Asked Questions

What is caliper?

caliper is Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.. It is categorized as a MCP Server with 50 GitHub stars.

What programming language is caliper written in?

caliper is primarily written in Python. It covers topics such as ai-agents, claude-code, cli.

How do I install or use caliper?

You can find installation instructions and usage details in the caliper GitHub repository at github.com/edonadei/caliper. The project has 50 stars and 8 forks, indicating an active community.

What license does caliper use?

caliper is released under the MIT license, making it free to use and modify according to the license terms.

What are the best alternatives to caliper?

The top alternatives to caliper on Agent Skills Hub include maestro, vibepod-cli, odai. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse MCP Server tools