llm-serving-auto-benchmark
๐ฏSkillfrom bbuf/sglang-auto-driven-skills
Part of Lifeskills, a curated collection of non-coding skills for AI agents focused on business-critical communication, strategy, negotiation, and influence with decision-ready outputs.
Same repository
bbuf/sglang-auto-driven-skills(13 items)
Installation
npx vibeindex add bbuf/sglang-auto-driven-skills --skill llm-serving-auto-benchmarknpx skills add bbuf/sglang-auto-driven-skills --skill llm-serving-auto-benchmark~/.claude/skills/llm-serving-auto-benchmark/SKILL.mdSKILL.md
More from this repository10
A skill for triaging SGLang production serving incidents using a replay-first approach, helping diagnose queue growth, timeouts, wrong outputs, crashes, and distributed stalls by preserving evidence and reproducing the request path before patching.
An agent skill for AI infrastructure engineers that provides operational playbooks for torch profiler analysis, LLM serving benchmarks, and SGLang optimization. Includes skills for splitting prefill/decode profiler evidence and turning traces into kernel fusion opportunities.
Agent-ready operational playbooks for AI infrastructure engineers, covering LLM serving benchmarks across SGLang, vLLM, and TensorRT-LLM, torch-profiler trace triage, kernel optimization opportunities, SGLang patch review, and production incident replay.
Agent-ready playbooks for AI infrastructure engineers, covering LLM serving benchmarks, capacity planning, torch-profiler analysis, compute simulation, and SGLang/vLLM optimization. Includes 58 model PR histories and production incident triage skills.
An agent-ready playbook for LLM serving benchmarks, capacity planning, torch-profiler triage, SGLang/vLLM optimization, and production incident analysis. Provides structured workflows for AI infrastructure performance tuning and human code review.
An agent-ready playbook for LLM serving optimization, providing operational memory for benchmarking SGLang, vLLM, and TensorRT-LLM, analyzing serving capacity from logs, profiling at kernel level, and handling production incidents.
Agent-ready playbooks for AI infrastructure engineers, covering LLM serving benchmarks, capacity planning, torch profiler analysis, pipeline inspection, compute simulation, and SGLang/vLLM optimization.
Agent-ready playbooks for AI infrastructure engineers covering LLM serving benchmarks, capacity planning, torch-profiler analysis, pipeline inspection, compute simulation, and SGLang/vLLM optimization with production incident triage.
A collection of agent-ready playbooks for AI infrastructure engineers, covering LLM serving benchmarks, capacity planning, profiler triage, compute simulation, and SGLang/vLLM optimization with real maintainer discussion patterns.
An agent-ready playbook for AI infrastructure engineers that provides forward-pass, layer-level, and kernel-level timing analysis from torch profiler traces, part of a broader skill set covering LLM serving benchmarks, capacity planning, and SGLang/vLLM optimization.