llm_context_benchmarks — security grade SAFE, quality 71/100

Security audit verdict: SAFE · quality 71/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by ivanfioravanti · LLM Plugin · ★ 96

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is llm_context_benchmarks safe to install? View the security audit →

About llm_context_benchmarks

LLM Context Benchmarks Benchmark prompt-processing and generation throughput across context sizes (0.5k–128k tokens) for many inference engines: Ollama (API & CLI), MLX, MLX Distributed, MLX-VLM, llama.cpp, LM Studio, Exo, Apple Foundation Models Serve, vMLX, oMLX, Paroquant, and any OpenAI-compatible endpoint. Optimized for Apple Silicon but works anywhere Python runs. Installation Engine-specific setup: (Optional) pre-commit hooks for Black + isort: Running Benchmarks bash List engines uv run benchmark --list-engines Generate test files (only needed once) uv run generate-context-files prideandprejudice.txt Run a benchmark (engine + model) uv run benchmark mlx mlx-community/Qwen3-4B-Instruct-2507-4bit uv run

aibenchmarkingllms

Quick Facts

Stars96
Forks11
LanguagePython
CategoryLLM Plugin
LicenseApache-2.0
Quality Score70.7881896020751/100
Last Updated2026-09-19
Created2025-08-06
Platformscli, python
Est. Tokens~15k

llm_context_benchmarks alternative? Top 6 similar tools

Looking for a llm_context_benchmarks alternative? If you're comparing llm_context_benchmarks with other llm plugin tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • nexus by Nexus-Router · ⭐ 436

    Govern & Secure your AI

  • deeppowers by deeppowers · ⭐ 252

    DEEPPOWERS is a Fully Homomorphic Encryption (FHE) framework built for MCP (Model Context Protocol), aiming to

  • weam by weam-ai · ⭐ 220

    Web app for teams of 20+ members. In-built connections to major LLMs via API. Share chats, prompts, and agents

  • manim-generator by makefinks · ⭐ 117

    Automatic LLM-based video generation using the manim library. Usage of a code-writer and code-reviewer feedbac

  • log10 by log10-io · ⭐ 97

    Python client library for improving your LLM app accuracy

  • claude-code-tamagotchi by Ido-Levi · ⭐ 433

    Real-time behavioral enforcement for Claude Code. Monitors AI actions, detects violations, and interrupts misb

More LLM Plugin Tools

Explore other popular llm plugin tools:

View all LLM Plugin tools →

Popular Python Agent Tools

Frequently Asked Questions

What is llm_context_benchmarks?

llm_context_benchmarks is 📊 LLM Context Benchmarks - A comprehensive benchmarking tool for testing LLMs with varying context sizes using Ollama. Features dual benchmark modes (API/CLI), automatic hardware detection (optimized. It is categorized as a LLM Plugin with 96 GitHub stars.

What programming language is llm_context_benchmarks written in?

llm_context_benchmarks is primarily written in Python. It covers topics such as ai, benchmarking, llms.

How do I install or use llm_context_benchmarks?

You can find installation instructions and usage details in the llm_context_benchmarks GitHub repository at github.com/ivanfioravanti/llm_context_benchmarks. The project has 96 stars and 11 forks, indicating an active community.

What license does llm_context_benchmarks use?

llm_context_benchmarks is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to llm_context_benchmarks?

The top alternatives to llm_context_benchmarks on Agent Skills Hub include nexus, deeppowers, weam. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse LLM Plugin tools