bigcodebench — security grade SAFE, quality 65/100

Security audit verdict: SAFE · quality 65/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by bigcode-project · Agent Tool · ★ 519

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is bigcodebench safe to install? View the security audit →

About bigcodebench

BigCodeBench 💥 Impact • 📰 News • 🔥 Quick Start • 🚀

agentagentsbenchmarkchatgptclaude-3code-generationdeepseekfunction-callinggeminigpt-4

Quick Facts

Stars519
Forks74
LanguagePython
CategoryAgent Tool
LicenseApache-2.0
Quality Score64.5815919159466/100
Open Issues30
Last Updated2026-01-03
Created2024-04-29
Platformsclaude-code, gemini, python
Est. Tokens~443k

Compatible Skills

These tools work well together with bigcodebench for enhanced workflows:

  • gptme — semantic(0.18)+complementary+rare_topics+same_lang+similar_pop+shared_platform (56%)
  • cursor-agent — semantic(0.16)+complementary+same_lang+similar_pop+shared_platform (56%)
  • agent-studio — semantic(0.16)+complementary+rare_topics+same_lang+similar_pop+shared_platform (55%)
  • PairCoder — semantic(0.38)+rare_topics+same_lang+similar_pop+shared_platform (52%)
  • mcp-gemini-search — semantic(0.17)+complementary+same_lang+similar_pop+shared_platform (51%)

bigcodebench alternative? Top 6 similar tools

Looking for a bigcodebench alternative? If you're comparing bigcodebench with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • DemoGPT by melih-unsal · ⭐ 1.9k

    🤖 Everything you need to create an LLM Agent—tools, prompts, frameworks, and models—all in one place.

  • kani by zhudotexe · ⭐ 607

    kani (カニ) is a highly hackable microframework for tool-calling language models. (NLP-OSS @ EMNLP 2023)

  • LLM-Tool-Survey by quchangle1 · ⭐ 489

    This is the repository for the Tool Learning survey.

  • awesome-llm-powered-agent by hyp1231 · ⭐ 2.3k

    Awesome things about LLM-powered agents. Papers / Repos / Blogs / ...

  • py-gpt by szczyglis-dev · ⭐ 1.9k

    Desktop AI Assistant powered by GPT-6, GPT-5, Gemini, Claude, Ollama, Grok, DeepSeek, Perplexity, chat, agents

  • awesome-gpt-prompt-engineering by snwfdhmp · ⭐ 1.6k

    A curated list of awesome resources, tools, and other shiny things for LLM prompt engineering.

More Agent Tool Tools

Explore other popular agent tool tools:

View all Agent Tool tools →

Popular Python Agent Tools

Frequently Asked Questions

What is bigcodebench?

bigcodebench is [ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI. It is categorized as a Agent Tool with 519 GitHub stars.

What programming language is bigcodebench written in?

bigcodebench is primarily written in Python. It covers topics such as agent, agents, benchmark.

How do I install or use bigcodebench?

You can find installation instructions and usage details in the bigcodebench GitHub repository at github.com/bigcode-project/bigcodebench. The project has 519 stars and 74 forks, indicating an active community.

What license does bigcodebench use?

bigcodebench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

What are the best alternatives to bigcodebench?

The top alternatives to bigcodebench on Agent Skills Hub include DemoGPT, kani, LLM-Tool-Survey. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Agent Tool tools