No red flags found in any of the 11 categories — no credential harvesting, no data exfiltration, no curl-pipe-shell installer. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by Re-Align · Agent Tool · ★ 90
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is just-eval safe to install? View the security audit →
Just-Eval: A fine-grained evaluation of LLM Alignment This is part of the Re-Align project by AI2 Mosaic. Please find more information on our website: https://allenai.github.io/re-align/. Just-Eval-Instruct Dataset 💾 Check out our data on 🤗 Hugging Face: re-align/just-eval-instruct 📊 Check here for the leaderboard: https://allenai.github.io/re-align/justeval.html#leaderboard Data distribution Installation or Setup OpenAI API Key Scoring with Multiple Aspects One-click Helpfulness, Clarity, Factuality, Depth, and Engagement scoremulti
| Stars | 90 |
| Forks | 7 |
| Language | Python |
| Category | Agent Tool |
| License | MIT |
| Quality Score | 71.620128375141/100 |
| Open Issues | 2 |
| Last Updated | 2024-01-29 |
| Created | 2023-11-19 |
| Platforms | python |
| Est. Tokens | ~1201k |
Looking for a just-eval alternative? If you're comparing just-eval with other agent tool tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
AI agent simulation framework
Build, Improve Performance, and Productionize your LLM Application with an Integrated Framework
A list of LLMs Tools & Projects
A-RAG: Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces. State-of-the-art RAG fram
Bayesian Optimization as a Coverage Tool for Evaluating LLMs. Accurate evaluation (benchmarking) that's 10 tim
Explore other popular agent tool tools:
just-eval is A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.. It is categorized as a Agent Tool with 90 GitHub stars.
just-eval is primarily written in Python. It covers topics such as evaluation, gpt4, llm.
You can find installation instructions and usage details in the just-eval GitHub repository at github.com/Re-Align/just-eval. The project has 90 stars and 7 forks, indicating an active community.
just-eval is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to just-eval on Agent Skills Hub include skill-optimizer, synkro, palico-ai. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.
Grades come from a rule-based scan built on the SlowMist agent-security taxonomy, covering 11 red-flag categories including credential harvesting, data exfiltration, and curl | sh installers. It is a first-layer scan, not a manual audit — we say so rather than overstate it.
The scale of the problem is documented independently: Liu et al. (2026), in a study of 31,132 agent skills, report that 26.1% contain security vulnerabilities. Our own full-catalog census is published as a citable open dataset.
Sources & who's responsible: