Llm-Serving
Llm-Serving
- LServe and SampleAttention: What Sparse Attention Actually Changes in Prefill and Decode
Structured sparsity aligned with GPU attention kernels: LServe’s unified serving stack for prefill and decode, and SampleAttention’s empirical patterns plus CRA as a runtime quality floor.