On the leaderboard
| Rank | Repository | Stars |
|---|---|---|
| 897 | sgl-project/sglang | 35,510 |
Top repositories by stars
- sgl-project/sglang(on leaderboard)
SGLang is a high-performance serving framework for large language models and multimodal models.
Python35,510 - sgl-project/mini-sglang
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Python4,962 - sgl-project/SpecForge
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Python1,151 - sgl-project/sglang-omni
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
Python1,103 - sgl-project/sgl-learning-materials
Materials for learning SGLang
886 - sgl-project/sglang-jax
JAX backend for SGL
Python349 - sgl-project/genai-bench
Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.
Python327 - sgl-project/rbg
A workload for deploying LLM inference services on Kubernetes
Go291 - sgl-project/sgl-kernel-npu
SGLang kernel library for NPU
C++178 - sgl-project/sgl-project.github.io
This is the documentation repository for SGLang. It is auto-generated from https://github.com/sgl-project/sglang
Jupyter Notebook135 - sgl-project/DeepGEMM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Cuda40 - sgl-project/sgl-kernel-xpu
SGLang kernel library for Intel XPU
C++32 - sgl-project/whl
SGLang Kernel Wheel Index
HTML26 - sgl-project/sgl-flash-attn
Fast and memory-efficient exact attention
Python23 - Python17
- sgl-project/sgl-whl
SGLang wheels for multiple platforms
11 - sgl-project/sgl-cookbook
Cookbook of SGLang - Recipe
JavaScript10 - Python3
- Python3
- sgl-project/fast-hadamard-transform
Fast Hadamard transform in CUDA, with a PyTorch interface
C2 - sgl-project/ci-data-diffusion
SGLang diffusion CI ground-truth baselines, benchmark comparisons, and repro scripts (split from sgl-project/ci-data to shed accumulated trace history)
Python1 - Python1
- sgl-project/FlashMLA
FlashMLA: Efficient Multi-head Latent Attention Kernels
C++1 - sgl-project/DeepEP
DeepEP: an efficient expert-parallel communication library
Cuda0 - sgl-project/sgl-test-files
The test files for SGLang.
0