No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by HealthRex · Agent Tool · ★ 51
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is PhysicianBench safe to install? View the security audit →
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments Overview PhysicianBench is a benchmark for evaluating LLM agents on physician tasks grounded in real clinical workflows. It comprises 100 long-horizon tasks (670 sub-checkpoints) adapted from real primary-care/specialty consultations across 21 specialties, executed in an EHR environment with real patient records accessed via standard FHIR APIs. Solving each task requires retrieving data across encounters, reasoning over heterogeneous clinical information, executing consequential clinical actions, and producing clinical documentation. Main Results Overall model performance on PhysicianBench, ranked by pass@1 success rate. Trajectory Example !
| Stars | 51 |
| Forks | 4 |
| Language | Python |
| Category | Agent Tool |
| License | Apache-2.0 |
| Quality Score | 72.2209338015781/100 |
| Open Issues | 1 |
| Last Updated | 2026-08-13 |
| Created | 2026-04-28 |
| Platforms | python |
| Est. Tokens | ~14k |
These tools work well together with PhysicianBench for enhanced workflows:
Looking for a PhysicianBench alternative? If you're comparing PhysicianBench with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Python SDK for healthcare AI — typed, validated FHIR tools for agents, real-time EHR connectivity, production
[EMNLP'24] EHRAgent: Code Empowers Large Language Models for Complex Tabular Reasoning on Electronic Health Re
[npj Digital Medicine 2026] HealthFlow: Automating electronic health record analysis via a strategically self-
The evaluation benchmark on MCP servers
[ICLR 2025] A trinity of environments, tools, and benchmarks for general virtual agents
This project collects GPU benchmarks from various cloud providers and compares them to fixed per token costs.
Explore other popular agent tool tools:
PhysicianBench is The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".. It is categorized as a Agent Tool with 51 GitHub stars.
PhysicianBench is primarily written in Python. It covers topics such as benchmark, ehr, healthcare-ai.
You can find installation instructions and usage details in the PhysicianBench GitHub repository at github.com/HealthRex/PhysicianBench. The project has 51 stars and 4 forks, indicating an active community.
PhysicianBench is released under the Apache-2.0 license, making it free to use and modify according to the license terms.
The top alternatives to PhysicianBench on Agent Skills Hub include HealthChain, EhrAgent, HealthFlow. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: