A daily log of LLM / ML-systems efficiency papers — inference serving, KV cache, speculative decoding, quantization, and whatever else makes models cheaper to run. Every paper is read in full text, not from the abstract.
18 papers read · 3 notes · 5 active days
vector-index 1 · cp-parallel-training 1 · rag-serving 1 · kernel-gen 1 · serving-sched 1 · compression-equivalence 1 · kv-tiering 1 · quantization-eval 1 · comm-overlap 1 · kv-compression 1 · tts-effectiveness 1 · moe-expert-prune 1
1 paper read · 1 note
- 📄 How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
- 💬 How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
6 papers read · 2 notes
- 📄 Higher-order pruning of experts in mixture-of-experts language models (HOPE)
- 📄 Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models
- 📄 A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality
- 📄 When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
- 📄 RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
- 📄 The Embedder's Dilemma: LLMs Are Better, but at What Cost?
- 💬 Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models
- 💬 The Embedder's Dilemma: LLMs Are Better, but at What Cost?
2 papers read
- 📄 Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
- 📄 Kinetics: Rethinking Test-Time Scaling Law
8 papers read
- 📄 Efficient Long-Context Language Model Training by Core Attention Disaggregation (DistCA)
- 📄 TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
- 📄 A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
- 📄 Simple Is Better: Multiplication May Be All You Need for LLM Request Scheduling (LMetric)
- 📄 Certifying Compressed Language Models: An Audit and a Statistical Toolkit
- 📄 A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving
- 📄 Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
- 📄 TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
1 paper read


