No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by waybarrios · MCP Server · ★ 1.6k
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is vllm-mlx safe to install? View the security audit →
vLLM-MLX vLLM-like inference for Apple Silicon - GPU-accelerated Text, Image, Video & Audio on Mac Overview vllm-mlx brings native Apple Silicon GPU acceleration to vLLM by integrating: MLX: Apple's ML framework with unified memory and Metal kernels mlx-lm: Optimized LLM inference with KV cache and quantization mlx-vlm: Vision-language models for multimodal inference mlx-audio: Speech-to-Text and Text-to-Speech with native voices mlx-embeddings: Text embeddings for semantic search and RAG Features Multimodal - Text, Image, Video & Audio in one platform Native GPU acceleration on Apple Silicon (M1, M2, M3, M4) Native TTS voices - Spanish, French, Chinese, Japanese + 5 more languages OpenAI API compatible - drop-in replacement for OpenAI client Anthropic Messages API - native /v1/me
| Stars | 1,587 |
| Forks | 222 |
| Language | Python |
| Category | MCP Server |
| License | Apache-2.0 |
| Quality Score | 70.9665687613498/100 |
| Open Issues | 121 |
| Last Updated | 2026-09-19 |
| Created | 2025-12-06 |
| Platforms | claude-code, mcp, python |
| Est. Tokens | ~16k |
These tools work well together with vllm-mlx for enhanced workflows:
Looking for a vllm-mlx alternative? If you're comparing vllm-mlx with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters inc
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpee
⚡️ Open-source AI Gateway — Use any SDK to call 100+ LLMs. Built-in failover, load balancing, cost control & e
AI Skills, MCP Tools, and CLI for Unity Engine. Full AI develop and test loop. Use cli for quick setup. Effici
LightAgent: Lightweight Python framework for OpenAI-compatible agents with tools, memory, guardrails, tracing,
Explore other popular mcp server tools:
vllm-mlx is High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.. It is categorized as a MCP Server with 1.6k GitHub stars.
vllm-mlx is primarily written in Python. It covers topics such as anthropic, anthropic-api, apple-silicon.
You can find installation instructions and usage details in the vllm-mlx GitHub repository at github.com/waybarrios/vllm-mlx. The project has 1.6k stars and 222 forks, indicating an active community.
vllm-mlx is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to vllm-mlx on Agent Skills Hub include mlx-serve, claude-code-local, smg. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: