Skip to content
View waynehacking8's full-sized avatar
:electron:
:electron:

Block or report waynehacking8

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
waynehacking8/README.md

Wei Cheng (Wayne) Chiu

LLM Infrastructure · GPU Inference · Solutions Architecture

Upstream contributions across the LLM-inference stack, paired with production work that turns deployment requirements into reliable, on-premises GenAI systems. Focused on model serving, GPU infrastructure, and solution architecture across Linux, Kubernetes, PyTorch, CUDA, TensorRT-LLM, vLLM, and Triton.

Portfolio · PR wall · LinkedIn · CV

Focus

  • Upstream inference work: correctness and performance across FlashInfer, vLLM, PyTorch, Dynamo, and the surrounding serving stack. The complete live record is on prs.wayne.is-a.dev.
  • GPU infrastructure: model serving, distributed communication, quantization, and kernel paths across vLLM, TensorRT-LLM, SGLang, Triton, Dynamo, and FlashInfer.
  • Solution architecture: deployment constraints, preflight validation, troubleshooting, acceptance, and production handover for on-premises LLM systems.

Representative merged work

Project Change
FlashInfer Fixed an SM120/121 multi-CTA radix top-k stream hang and NVFP4 attention correction layout
vLLM Made tokenizers survive pickling and fixed streaming message metadata
PyTorch Restored the dropped root argument in nccl.broadcast
Dynamo Cancelled in-flight KV-router recovery when a worker is removed and resolved aggregate planner workers by DGD component type

Selected projects

Project Evidence
trtllm-triton-serving TensorRT-LLM vs vLLM on H100; 12 controlled studies
tensor-core-from-scratch 10 CUDA matmul kernels from naive to Tensor Cores
inference-kernel-cookbook Flash Attention, KV cache, and paged attention from scratch
nccl-collectives-bench NCCL bandwidth, latency, NVLS, and TP-decode limits
nim-agent-blueprint NVIDIA NIM agentic RAG with evaluation and observability
llm-security-lab Reproducible LLM attacks and defenses

Technical writing

Pinned Loading

  1. coralline-codex coralline-codex Public

    Coralline-inspired status line companion for OpenAI Codex

    Shell 12 1