No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by ivanfioravanti · LLM Plugin · ★ 96
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is llm_context_benchmarks safe to install? View the security audit →
LLM Context Benchmarks Benchmark prompt-processing and generation throughput across context sizes (0.5k–128k tokens) for many inference engines: Ollama (API & CLI), MLX, MLX Distributed, MLX-VLM, llama.cpp, LM Studio, Exo, Apple Foundation Models Serve, vMLX, oMLX, Paroquant, and any OpenAI-compatible endpoint. Optimized for Apple Silicon but works anywhere Python runs. Installation Engine-specific setup: (Optional) pre-commit hooks for Black + isort: Running Benchmarks bash List engines uv run benchmark --list-engines Generate test files (only needed once) uv run generate-context-files prideandprejudice.txt Run a benchmark (engine + model) uv run benchmark mlx mlx-community/Qwen3-4B-Instruct-2507-4bit uv run
| Stars | 96 |
| Forks | 11 |
| Language | Python |
| Category | LLM Plugin |
| License | Apache-2.0 |
| Quality Score | 70.7881896020751/100 |
| Last Updated | 2026-09-19 |
| Created | 2025-08-06 |
| Platforms | cli, python |
| Est. Tokens | ~15k |
Looking for a llm_context_benchmarks alternative? If you're comparing llm_context_benchmarks with other llm plugin tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Govern & Secure your AI
DEEPPOWERS is a Fully Homomorphic Encryption (FHE) framework built for MCP (Model Context Protocol), aiming to
Web app for teams of 20+ members. In-built connections to major LLMs via API. Share chats, prompts, and agents
Automatic LLM-based video generation using the manim library. Usage of a code-writer and code-reviewer feedbac
Python client library for improving your LLM app accuracy
Real-time behavioral enforcement for Claude Code. Monitors AI actions, detects violations, and interrupts misb
Explore other popular llm plugin tools:
llm_context_benchmarks is 📊 LLM Context Benchmarks - A comprehensive benchmarking tool for testing LLMs with varying context sizes using Ollama. Features dual benchmark modes (API/CLI), automatic hardware detection (optimized. It is categorized as a LLM Plugin with 96 GitHub stars.
llm_context_benchmarks is primarily written in Python. It covers topics such as ai, benchmarking, llms.
You can find installation instructions and usage details in the llm_context_benchmarks GitHub repository at github.com/ivanfioravanti/llm_context_benchmarks. The project has 96 stars and 11 forks, indicating an active community.
llm_context_benchmarks is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to llm_context_benchmarks on Agent Skills Hub include nexus, deeppowers, weam. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: