No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by deep-symbolic-mathematics · Agent Tool · ★ 117
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is llm-srbench safe to install? View the security audit →
: Benchmark for Scientific Equation Discovery or Symbolic Regression with LLMs This is the official repository for the paper "LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models" (ICML 2025 Oral) Overview In this paper, we introduce LLM-SRBench, a comprehensive benchmark with $239$ challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorized forms, and LSR-Synth, which introduces synthetic, discovery-driven problems requiring data-driven reasoning. Updates 9 June, 2025: 🌟 LLM-SRBench is selected for Oral presentation (top 1%) at ICML 2025 1 May, 2025: 🌟 LLM-SRBench is accepted for Spotlight poster at ICML 2025 16 Apr, 2025: 🌟 LLM-SRBench data and evaluation code i
| Stars | 117 |
| Forks | 10 |
| Language | Python |
| Category | Agent Tool |
| Quality Score | 69.28106042939/100 |
| Open Issues | 4 |
| Last Updated | 2025-07-31 |
| Created | 2025-01-30 |
| Platforms | python |
| Est. Tokens | ~84k |
Looking for a llm-srbench alternative? If you're comparing llm-srbench with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Promptdesk is a tool designed for effectively creating, organizing, and evaluating prompts and large language
Explore a comprehensive collection of resources, tutorials, papers, tools, and best practices for fine-tuning
An open, curated collection of Agent Skills for scientific research — clone it, use it, extend it!
900+ pure-markdown skills for autonomous AI research, organized as 9 freely-composable packages over a 4-layer
Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market
This is the repository for the Tool Learning survey.
Explore other popular agent tool tools:
llm-srbench is [ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models. It is categorized as a Agent Tool with 117 GitHub stars.
llm-srbench is primarily written in Python. It covers topics such as ai4code, ai4math, ai4science.
You can find installation instructions and usage details in the llm-srbench GitHub repository at github.com/deep-symbolic-mathematics/llm-srbench. The project has 117 stars and 10 forks, indicating an active community.
The top alternatives to llm-srbench on Agent Skills Hub include promptdesk, awesome-llms-fine-tuning, Awesome-Scientific-Skills. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: