llm-srbench — security grade SAFE, quality 69/100

Security audit verdict: SAFE · quality 69/100

No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →

by deep-symbolic-mathematics · Agent Tool · ★ 117

Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h

🔒 Is llm-srbench safe to install? View the security audit →

About llm-srbench

: Benchmark for Scientific Equation Discovery or Symbolic Regression with LLMs This is the official repository for the paper "LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models" (ICML 2025 Oral) Overview In this paper, we introduce LLM-SRBench, a comprehensive benchmark with $239$ challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorized forms, and LSR-Synth, which introduces synthetic, discovery-driven problems requiring data-driven reasoning. Updates 9 June, 2025: 🌟 LLM-SRBench is selected for Oral presentation (top 1%) at ICML 2025 1 May, 2025: 🌟 LLM-SRBench is accepted for Spotlight poster at ICML 2025 16 Apr, 2025: 🌟 LLM-SRBench data and evaluation code i

ai4codeai4mathai4sciencelarge-language-modelsllm-agentscientific-discovery

Quick Facts

Stars117
Forks10
LanguagePython
CategoryAgent Tool
Quality Score69.28106042939/100
Open Issues4
Last Updated2025-07-31
Created2025-01-30
Platformspython
Est. Tokens~84k

llm-srbench alternative? Top 6 similar tools

Looking for a llm-srbench alternative? If you're comparing llm-srbench with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.

  • promptdesk by promptdesk · ⭐ 100

    Promptdesk is a tool designed for effectively creating, organizing, and evaluating prompts and large language

  • awesome-llms-fine-tuning by Curated-Awesome-Lists · ⭐ 525

    Explore a comprehensive collection of resources, tutorials, papers, tools, and best practices for fine-tuning

  • Awesome-Scientific-Skills by InternScience · ⭐ 507

    An open, curated collection of Agent Skills for scientific research — clone it, use it, extend it!

  • de-anthropocentric-research-engine by yogsoth-ai · ⭐ 499

    900+ pure-markdown skills for autonomous AI research, organized as 9 freely-composable packages over a 4-layer

  • edsl by expectedparrot · ⭐ 497

    Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market

  • LLM-Tool-Survey by quchangle1 · ⭐ 489

    This is the repository for the Tool Learning survey.

More Agent Tool Tools

Explore other popular agent tool tools:

View all Agent Tool tools →

Popular Python Agent Tools

Frequently Asked Questions

What is llm-srbench?

llm-srbench is [ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models. It is categorized as a Agent Tool with 117 GitHub stars.

What programming language is llm-srbench written in?

llm-srbench is primarily written in Python. It covers topics such as ai4code, ai4math, ai4science.

How do I install or use llm-srbench?

You can find installation instructions and usage details in the llm-srbench GitHub repository at github.com/deep-symbolic-mathematics/llm-srbench. The project has 117 stars and 10 forks, indicating an active community.

What are the best alternatives to llm-srbench?

The top alternatives to llm-srbench on Agent Skills Hub include promptdesk, awesome-llms-fine-tuning, Awesome-Scientific-Skills. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.

How this security grade is produced

Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.

The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.

Sources & who's responsible:

View on GitHub → Browse Agent Tool tools