Portfolio concentration
99%
Top three share
Shows whether the organization is driven by one breakout repo or several visible projects.
Breadth
20 repos
Visible snapshot
12 repositories updated in the last 90 days.
Leading language
Python
Portfolio mix
Python (8), Unknown (6), Cuda (2)
Average size
1.5K
Stars per repository
Useful for distinguishing one flagship-heavy publisher from a repeatable portfolio.
99%
of the visible star count comes from this organization's top three repositories.
1.5K
stars per repository in this same snapshot.
Python
is the most common language here, with 12 repositories updated in the last 90 days.
Why this rank
This organization stands out because one flagship repo drives 65% of its visible star count.
Organization pages work best when you separate portfolio breadth from flagship concentration. In kvcache.ai's case, the visible top three repositories account for about 99% of total stars in this snapshot, which helps explain whether the organization is known for one breakout project or for a broader repeatable portfolio.
The dominant language mix here is Python (8), Unknown (6), Cuda (2). That makes this page useful not just for popularity checks, but also for seeing what technical shape an organization's public ecosystem actually has.
| # | Repository | Language | Stars |
|---|---|---|---|
| 1 | kvcache-ai/ktransformers A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations | Python | 19.6K |
| 2 | kvcache-ai/Mooncake Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. | C++ | 6.8K |
| 3 | kvcache-ai/AgentENV AgentENV (AENV) is a distributed platform for running agent environments at scale. | Rust | 3.6K |
| 4 | kvcache-ai/TrEnv-X | Go | 97 |
| 5 | kvcache-ai/kvcache-blog | JavaScript | 29 |
| 6 | kvcache-ai/vllm A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 19 |
| 7 | kvcache-ai/sglang SGLang is a fast serving framework for large language models and vision language models. | Python | 14 |
| 8 | kvcache-ai/custom_flashinfer FlashInfer: Kernel Library for LLM Serving | Cuda | 9 |
| 9 | kvcache-ai/AFD-Ledger AFD-Ledger is an analytical toolkit for provisioning and evaluating attention–FFN disaggregated LLM inference deployments. | Python | 6 |
| 10 | kvcache-ai/DeepEP_fault_tolerance DeepEP: an efficient expert-parallel communication library that supports fault tolerance | Cuda | 4 |
| 11 | kvcache-ai/accelerate 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support | Python | 3 |
| 12 | kvcache-ai/linux Linux kernel source tree for PVM | 2 | |
| 13 | kvcache-ai/transformers 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training. | Python | 2 |
| 14 | kvcache-ai/sglang_awq SGLang is a fast serving framework for large language models and vision language models. | Python | 2 |
| 15 | kvcache-ai/Model-Optimizer A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed. | 1 | |
| 16 | kvcache-ai/gpustack GPU cluster manager for optimized AI model deployment | 1 | |
| 17 | kvcache-ai/overlaybd Overlaybd: a block based remote image format. The storage backend of containerd/accelerated-container-image. | 0 | |
| 18 | kvcache-ai/kvcache-blog-private | 0 | |
| 19 | kvcache-ai/evalscope A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking. | Python | 0 |
| 20 | kvcache-ai/sglang-npu SGLang is a fast serving framework for large language models and vision language models. | 0 |
Total stars are useful as a discovery signal, but they do not tell you whether a team maintains every repository equally. Pair this page with release cadence, maintainer activity, and the flagship concentration shown above before making adoption decisions.
For broader background on GitStar's ranking logic and editorial guidance, see Methodology & Editorial Standards.