eval
๐ฏSkillfrom alirezarezvani/claude-skills
`/hub:eval` โ rank all AgentHub agent results for a session using metric mode (run an eval command in each agent's worktree with `scripts/result_ranker.py --session ... --eval-cmd ... --metric ... --direction ...`), LLM judge mode (compare `git diff {base_branch}...{agent_branch}` plus each agent's `.agenthub/board/results/agent-{i}-result.md` on correctness / simplicity / quality), or hybrid (metric first, LLM tie-break within 10%). Updates session state via `session_manager.py --update ... --state evaluating` and points to `/hub:merge` for the winner.
Same repository
alirezarezvani/claude-skills(336 items)
Installation
npx vibeindex add alirezarezvani/claude-skills --skill evalnpx skills add alirezarezvani/claude-skills --skill eval~/.claude/skills/eval/SKILL.mdSKILL.md
More from this repository10
Plugin
Plugin
A product development skill package providing frameworks for product strategy, roadmap planning, user research analysis, and feature prioritization for product teams.
Production-ready skill packages for Claude AI combining best practices, analysis tools, and strategic frameworks for marketing, executive leadership, product development, and engineering teams.
Plugin
Plugin
AWS Architect plugin from a Claude Skills Library providing production-ready skill packages for Claude AI and Claude Code. Covers marketing teams, executive leadership, product development, and web/mobile engineering with reusable expertise bundles.
Plugin
A fullstack engineer skill from the Claude Skills Library, providing production-ready expertise bundles combining best practices, analysis tools, and strategic frameworks for web and mobile engineering.
220+ Claude Code skills & agent plugins for Claude Code, Codex, Gemini CLI, Cursor, and 8 more coding agents โ engineering, marketing, product, compliance, C-level advisory.