π―Skills21
A benchmarking framework from LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with treatment-based test variations.
A benchmark suite by LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and repetitions.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. It supports multiple treatments, repetitions, and parallel test execution with Docker-sandboxed validation.
Part of LangChain's skill benchmarks project that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed test tasks with configurable treatments.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and automated validation.
A benchmark skill from LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, with support for multiple treatments and configurable model selection.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and validation scripts.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Supports multiple treatments, repetitions, and parallel test execution in Docker-sandboxed environments.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. It runs tasks in Docker-sandboxed environments with configurable treatments and repetitions to evaluate skill effectiveness.
A benchmarking framework from LangChain that measures how skill documentation design affects Claude Code adherence to recommended patterns, using configurable treatments, Docker-sandboxed execution, and automated validation.
A benchmarking skill from LangChain that measures how skill documentation design affects Claude Code adherence to recommended LangGraph patterns, with Docker-based sandboxed execution and validation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, running tasks in Docker containers with configurable treatments and parallel test execution.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, running tasks in Docker with configurable treatments and parallel execution.
Part of Claude Code Skill Benchmarks by LangChain, a framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Provides Docker-sandboxed task execution with configurable treatments, repetition counts, and parallel workers for systematic evaluation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, with support for configurable task treatments and parallel test execution in Docker sandboxes.
A benchmark framework by LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Runs tasks with configurable treatments in Docker sandboxes and supports evaluation via LangSmith.
A benchmark framework by LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments, validation scripts, and parallel test execution.
A benchmarking framework that measures how skill documentation design affects Claude Code adherence to recommended patterns, supporting multiple tasks, treatments, and repetition-based statistical analysis.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, with Docker-based sandboxed execution and configurable treatments.
A benchmark suite that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Uses Docker-sandboxed task execution with configurable treatments, repetitions, and parallel workers for reproducible evaluation.
Benchmark framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Supports configurable tasks, treatments, and Docker-sandboxed validation with repeatable, parallel test execution.
