No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by Anionex · Codex Skill · ★ 1.2k
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is agent-vision-toolkit safe to install? View the security audit →
agent-vision-toolkit What it thinks is what it sees — give any text-only coding agent eyes: image Q&A, OCR, screenshot understanding, visual grounding, and image-to-SVG, as a vision toolkit plus a skill, with optional drop-in integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. 🌐 中文 | English If your coding agent runs on a text-only model like DeepSeek V4, it can't look at images — screenshots, mockups, diagrams, and error dialogs are all dead ends. This repository gives it eyes in two layers: The toolkit — four CLIs, plus a skill that teaches your agent when to reach for each one. Works in any agent with a shell. Seamless integration (optional upgrade) — a transparent local proxy and single-file native extensions, so pasted images and built-in image tools work too, with no tool call and no extra prompting. All code has been verified in real Codex + DeepSeek sessions, and the same pipeline has been live-verified end-to-end in Claude Code, Pi, Oh My Pi, and OpenCode. Use cases include but are not limited to: image Q&A, screenshot analysis, Computer Use GUI operation, and multi-step image reasoning.
| Stars | 1,213 |
| Forks | 48 |
| Language | Python |
| Category | Codex Skill |
| License | MIT |
| Quality Score | 73.7626382783291/100 |
| Open Issues | 9 |
| Last Updated | 2026-10-02 |
| Created | 2026-08-01 |
| Platforms | claude-code, codex, python |
| Est. Tokens | ~17k |
Looking for a agent-vision-toolkit alternative? If you're comparing agent-vision-toolkit with other codex skill tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an
🥇 The strongest free web search plugin for DeepSeek Harness, and the search bridge for every model without na
A persistent, unified memory layer for all your AI agents (e.g. Claude Code, Codex, DSH), backed by Markdown a
⚡️ Open-source AI Gateway — Use any SDK to call 100+ LLMs. Built-in failover, load balancing, cost control & e
Make AI coding agents architecture-aware: baseline-first, evidence-verified, drift-checked, and safe across lo
Deepseek Harness、Openclaw知识图谱记忆插件。2026年4月受邀发布在清华大学讨论会。Knowledge Graph + Memory;Knowledge Graph Context Engine
Explore other popular codex skill tools:
agent-vision-toolkit is 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI. It is categorized as a Codex Skill with 1.2k GitHub stars.
agent-vision-toolkit is primarily written in Python. It covers topics such as agent, agent-skills, claude-code.
You can find installation instructions and usage details in the agent-vision-toolkit GitHub repository at github.com/Anionex/agent-vision-toolkit. The project has 1.2k stars and 48 forks, indicating an active community.
agent-vision-toolkit is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to agent-vision-toolkit on Agent Skills Hub include modlens, modsearch, memsearch. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: