No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by mbzuai-oryx · Agent Tool · ★ 105
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is VideoGLaMM safe to install? View the security audit →
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos [CVPR 2025🔥] Shehan Munasinghe , Hanan Gani , Wenqi Zhu , Jiale Cao, Eric Xing, Fahad Shahbaz Khan. Salman Khan, Mohamed bin Zayed University of Artificial Intelligence, Tianjin University, Linköping University, Australian National University, Carnegie Mellon University 📢 Latest Updates Feb-2025: Video-GLaMM is accepted at CVPR 2025! 🎊🎊 Overview VideoGLaMM is a large video multimodal video model capable of pixel-level visual grounding. The model responds to natural language queries from the user and intertwines spatio-temporal object masks in its generated textual responses to provide a detailed understanding of video content. V
| Stars | 105 |
| Forks | 4 |
| Language | Python |
| Category | Agent Tool |
| Quality Score | 64.4828452138918/100 |
| Open Issues | 9 |
| Last Updated | 2026-09-05 |
| Created | 2024-10-31 |
| Platforms | python |
| Est. Tokens | ~14k |
Looking for a VideoGLaMM alternative? If you're comparing VideoGLaMM with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
NEWTON: Agentic Planning for Physically Grounded Video Generation
Auto-Use Computer Use — drives your OS, browser, scours the web, writes your code. One agent, end to end.
基于多模态视觉感知与 LLM Agent 的 macOS 微信自动化框架 | Visual RPA for WeChat
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Q
LLM Agent that leverages cheminformatics tools to provide informed responses.
Code repo for the paper: Attacking Vision-Language Computer Agents via Pop-ups
Explore other popular agent tool tools:
VideoGLaMM is [CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos. It is categorized as a Agent Tool with 105 GitHub stars.
VideoGLaMM is primarily written in Python. It covers topics such as cvpr2025, foundation-models, llm-agent.
You can find installation instructions and usage details in the VideoGLaMM GitHub repository at github.com/mbzuai-oryx/VideoGLaMM. The project has 105 stars and 4 forks, indicating an active community.
The top alternatives to VideoGLaMM on Agent Skills Hub include NEWTON, Auto-Use, wechat-mac-rpa. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: