Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels Blog Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels TL;DR In this blog post, we present various novel kernel fusion techniques for common normalization…Jacky (Junqing) Zhou, Hongtao Yu, Jackie (Jiaqi) Xu, Menglu Yu, Ethan Che, Han Xu, Darren Liu, Peng Chen (Dev Infra), Daohang Shi, Max LeungJuly 10, 2026
From Minutes to Seconds: LLM-Guided Autotuning for Helion Kernels Blog From Minutes to Seconds: LLM-Guided Autotuning for Helion Kernels TL;DR Helion, PyTorch’s domain-specific language (DSL) for performance portable machine learning kernels, heavily relies on…Jongsok Choi, Ethan Che, Jason Ansel, Oguz UlgenJune 18, 2026
Portable vLLM Model Inference Kernels in Helion Blog Portable vLLM Model Inference Kernels in Helion TL;DR Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated…Sean Chen (Red Hat) and Yanan Cao (PyTorch, Meta Platforms)June 10, 2026