Skip to content
@z-lab

Z Lab

Efficient AI. PI: Zhijian Liu

Popular repositories Loading

  1. dflash dflash Public

    DFlash: Block Diffusion for Flash Speculative Decoding

    Python 6.1k 434

  2. paroquant paroquant Public

    [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

    Python 341 33

  3. sparselora sparselora Public

    [ICML 2025] SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity

    Python 79 7

  4. flashdrive flashdrive Public

    Flash Vision-Language-Action Inference for Autonomous Driving

    Python 79 8

  5. flash-colreduce flash-colreduce Public

    Fast, memory-efficient attention column reduction (e.g., sum, mean, max)

    Python 50 3

  6. flashvla flashvla Public

    FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLAs Inference

    Python 48 3

Repositories

Showing 10 of 11 repositories
  • flashdrive Public

    Flash Vision-Language-Action Inference for Autonomous Driving

    z-lab/flashdrive's past year of commit activity
    Python 79 MIT 8 0 0 Updated Sep 15, 2026
  • flashvla Public

    FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLAs Inference

    z-lab/flashvla's past year of commit activity
    Python 48 Apache-2.0 3 1 0 Updated Sep 15, 2026
  • vllm-fork Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    z-lab/vllm-fork's past year of commit activity
    Python 0 Apache-2.0 22,945 0 0 Updated Sep 1, 2026
  • llama.cpp-fork Public Forked from ggml-org/llama.cpp

    LLM inference in C/C++

    z-lab/llama.cpp-fork's past year of commit activity
    C++ 2 MIT 24,261 0 5 Updated Aug 26, 2026
  • paroquant Public

    [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

    z-lab/paroquant's past year of commit activity
    Python 341 MIT 33 12 2 Updated Aug 19, 2026
  • omlx-fork Public Forked from jundot/omlx

    LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

    z-lab/omlx-fork's past year of commit activity
    Python 2 Apache-2.0 1,956 0 2 Updated Aug 19, 2026
  • sglang-fork Public Forked from sgl-project/sglang

    SGLang is a high-performance serving framework for large language models and multimodal models.

    z-lab/sglang-fork's past year of commit activity
    Python 0 Apache-2.0 9,212 0 0 Updated Aug 18, 2026
  • dflash Public

    DFlash: Block Diffusion for Flash Speculative Decoding

    z-lab/dflash's past year of commit activity
    Python 6,122 MIT 434 92 14 Updated Aug 18, 2026
  • dflash-mlx-fork Public Forked from jundot/dflash-mlx

    Lossless DFlash speculative decoding for MLX on Apple Silicon

    z-lab/dflash-mlx-fork's past year of commit activity
    Python 0 Apache-2.0 68 0 0 Updated Aug 18, 2026
  • sparselora Public

    [ICML 2025] SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity

    z-lab/sparselora's past year of commit activity
    Python 79 MIT 7 2 0 Updated Mar 10, 2026

Top languages

Python C++

Most used topics

Loading…