langsmith-evaluator
๐ฏSkillfrom langchain-ai/skills-benchmarks
A benchmark framework by LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Runs tasks with configurable treatments in Docker sandboxes and supports evaluation via LangSmith.
Same repository
langchain-ai/skills-benchmarks(21 items)
Installation
npx vibeindex add langchain-ai/skills-benchmarks --skill langsmith-evaluatornpx skills add langchain-ai/skills-benchmarks --skill langsmith-evaluator~/.claude/skills/langsmith-evaluator/SKILL.mdSKILL.md
More from this repository10
A benchmarking framework from LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with treatment-based test variations.
A benchmark suite by LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and repetitions.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. It supports multiple treatments, repetitions, and parallel test execution with Docker-sandboxed validation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and validation scripts.
Part of LangChain's skill benchmarks project that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed test tasks with configurable treatments.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and automated validation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Supports multiple treatments, repetitions, and parallel test execution in Docker-sandboxed environments.
A benchmark skill from LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, with support for multiple treatments and configurable model selection.
A benchmarking framework from LangChain that measures how skill documentation design affects Claude Code adherence to recommended patterns, using configurable treatments, Docker-sandboxed execution, and automated validation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. It runs tasks in Docker-sandboxed environments with configurable treatments and repetitions to evaluate skill effectiveness.