tile-ai/tilelangPublic

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

AI summary: Domain-specific language for developing high-performance hardware kernels easily.

Stars
8.3K
+53 today
Forks
838
Watchers
45
Open issues
214
Open PRs
156
Contributors
~211
Commits
2K
Branches
7

PythonOtherCreated Oct 3, 2024Last push 1d agoLatest release v0.1.15+259 stars this week+259 this month

Quick answers

What is tilelang?
Domain-specific language for developing high-performance hardware kernels easily.
What does tilelang do?
TileLang is a concise domain-specific language built to streamline the development of high-performance kernels for GPUs, CPUs, and accelerators. It utilizes a Pythonic syntax to make low-level programming more accessible and productive. The language is built on top of the TVM compiler infrastructure to ensure robust optimization. It excels at tasks like GEMM, FlashAttention, and LinearAttention implementations. Developers can achieve state-of-the-art hardware performance without sacrificing development speed.
Who is tilelang for?
Performance engineers, AI researchers, and systems programmers optimizing machine learning operations. It requires understanding of hardware architecture and compiler technologies.
How popular is tilelang on GitHub?
tile-ai/tilelang has 8,341 stars and 838 forks on GitHub, and gained 259 stars in the last 7 days.
What license does tilelang use?
tile-ai/tilelang is released under the Other license.

Star history

since Oct 1, 2026
02.5K5K7.5KOct 2026Oct 2026Oct 2026Oct 2026
8.3K stars as of Oct 4, 2026. Measured daily since Oct 1, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-05: 4 commits2025-10-06: 3 commits2025-10-07: 3 commits2025-10-08: 0 commits2025-10-09: 9 commits2025-10-10: 7 commits2025-10-11: 7 commits2025-10-12: 3 commits2025-10-13: 3 commits2025-10-14: 9 commits2025-10-15: 6 commits2025-10-16: 6 commits2025-10-17: 7 commits2025-10-18: 0 commits2025-10-19: 8 commits2025-10-20: 9 commits2025-10-21: 9 commits2025-10-22: 9 commits2025-10-23: 4 commits2025-10-24: 1 commit2025-10-25: 1 commit2025-10-26: 0 commits2025-10-27: 7 commits2025-10-28: 4 commits2025-10-29: 8 commits2025-10-30: 1 commit2025-10-31: 4 commits2025-11-01: 1 commit2025-11-02: 2 commits2025-11-03: 6 commits2025-11-04: 5 commits2025-11-05: 6 commits2025-11-06: 5 commits2025-11-07: 3 commits2025-11-08: 1 commit2025-11-09: 2 commits2025-11-10: 6 commits2025-11-11: 5 commits2025-11-12: 7 commits2025-11-13: 6 commits2025-11-14: 3 commits2025-11-15: 3 commits2025-11-16: 1 commit2025-11-17: 5 commits2025-11-18: 8 commits2025-11-19: 4 commits2025-11-20: 5 commits2025-11-21: 4 commits2025-11-22: 2 commits2025-11-23: 1 commit2025-11-24: 6 commits2025-11-25: 7 commits2025-11-26: 7 commits2025-11-27: 1 commit2025-11-28: 4 commits2025-11-29: 0 commits2025-11-30: 1 commit2025-12-01: 4 commits2025-12-02: 6 commits2025-12-03: 2 commits2025-12-04: 0 commits2025-12-05: 1 commit2025-12-06: 4 commits2025-12-07: 6 commits2025-12-08: 2 commits2025-12-09: 0 commits2025-12-10: 4 commits2025-12-11: 4 commits2025-12-12: 6 commits2025-12-13: 0 commits2025-12-14: 2 commits2025-12-15: 11 commits2025-12-16: 4 commits2025-12-17: 11 commits2025-12-18: 3 commits2025-12-19: 6 commits2025-12-20: 1 commit2025-12-21: 2 commits2025-12-22: 8 commits2025-12-23: 6 commits2025-12-24: 14 commits2025-12-25: 7 commits2025-12-26: 6 commits2025-12-27: 3 commits2025-12-28: 3 commits2025-12-29: 11 commits2025-12-30: 3 commits2025-12-31: 6 commits2026-01-01: 4 commits2026-01-02: 0 commits2026-01-03: 1 commit2026-01-04: 7 commits2026-01-05: 1 commit2026-01-06: 9 commits2026-01-07: 8 commits2026-01-08: 4 commits2026-01-09: 8 commits2026-01-10: 2 commits2026-01-11: 1 commit2026-01-12: 4 commits2026-01-13: 5 commits2026-01-14: 3 commits2026-01-15: 3 commits2026-01-16: 5 commits2026-01-17: 0 commits2026-01-18: 5 commits2026-01-19: 7 commits2026-01-20: 4 commits2026-01-21: 2 commits2026-01-22: 4 commits2026-01-23: 7 commits2026-01-24: 0 commits2026-01-25: 1 commit2026-01-26: 5 commits2026-01-27: 8 commits2026-01-28: 7 commits2026-01-29: 4 commits2026-01-30: 4 commits2026-01-31: 0 commits2026-02-01: 2 commits2026-02-02: 7 commits2026-02-03: 4 commits2026-02-04: 5 commits2026-02-05: 6 commits2026-02-06: 4 commits2026-02-07: 3 commits2026-02-08: 3 commits2026-02-09: 10 commits2026-02-10: 2 commits2026-02-11: 4 commits2026-02-12: 2 commits2026-02-13: 0 commits2026-02-14: 3 commits2026-02-15: 1 commit2026-02-16: 4 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 3 commits2026-02-20: 1 commit2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 2 commits2026-02-24: 5 commits2026-02-25: 6 commits2026-02-26: 0 commits2026-02-27: 2 commits2026-02-28: 2 commits2026-03-01: 0 commits2026-03-02: 2 commits2026-03-03: 2 commits2026-03-04: 1 commit2026-03-05: 2 commits2026-03-06: 4 commits2026-03-07: 0 commits2026-03-08: 1 commit2026-03-09: 1 commit2026-03-10: 3 commits2026-03-11: 3 commits2026-03-12: 5 commits2026-03-13: 2 commits2026-03-14: 0 commits2026-03-15: 1 commit2026-03-16: 0 commits2026-03-17: 1 commit2026-03-18: 3 commits2026-03-19: 2 commits2026-03-20: 4 commits2026-03-21: 0 commits2026-03-22: 4 commits2026-03-23: 4 commits2026-03-24: 4 commits2026-03-25: 2 commits2026-03-26: 4 commits2026-03-27: 2 commits2026-03-28: 0 commits2026-03-29: 1 commit2026-03-30: 4 commits2026-03-31: 4 commits2026-04-01: 1 commit2026-04-02: 2 commits2026-04-03: 1 commit2026-04-04: 1 commit2026-04-05: 2 commits2026-04-06: 0 commits2026-04-07: 6 commits2026-04-08: 3 commits2026-04-09: 0 commits2026-04-10: 2 commits2026-04-11: 3 commits2026-04-12: 0 commits2026-04-13: 5 commits2026-04-14: 4 commits2026-04-15: 5 commits2026-04-16: 1 commit2026-04-17: 6 commits2026-04-18: 2 commits2026-04-19: 0 commits2026-04-20: 2 commits2026-04-21: 7 commits2026-04-22: 7 commits2026-04-23: 3 commits2026-04-24: 2 commits2026-04-25: 4 commits2026-04-26: 0 commits2026-04-27: 5 commits2026-04-28: 5 commits2026-04-29: 2 commits2026-04-30: 4 commits2026-05-01: 1 commit2026-05-02: 0 commits2026-05-03: 2 commits2026-05-04: 1 commit2026-05-05: 1 commit2026-05-06: 9 commits2026-05-07: 5 commits2026-05-08: 2 commits2026-05-09: 1 commit2026-05-10: 2 commits2026-05-11: 2 commits2026-05-12: 4 commits2026-05-13: 2 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 1 commit2026-05-18: 5 commits2026-05-19: 6 commits2026-05-20: 9 commits2026-05-21: 12 commits2026-05-22: 5 commits2026-05-23: 1 commit2026-05-24: 3 commits2026-05-25: 10 commits2026-05-26: 10 commits2026-05-27: 4 commits2026-05-28: 9 commits2026-05-29: 4 commits2026-05-30: 0 commits2026-05-31: 1 commit2026-06-01: 2 commits2026-06-02: 3 commits2026-06-03: 1 commit2026-06-04: 5 commits2026-06-05: 3 commits2026-06-06: 1 commit2026-06-07: 1 commit2026-06-08: 4 commits2026-06-09: 4 commits2026-06-10: 3 commits2026-06-11: 2 commits2026-06-12: 5 commits2026-06-13: 0 commits2026-06-14: 2 commits2026-06-15: 5 commits2026-06-16: 3 commits2026-06-17: 3 commits2026-06-18: 4 commits2026-06-19: 1 commit2026-06-20: 0 commits2026-06-21: 1 commit2026-06-22: 6 commits2026-06-23: 6 commits2026-06-24: 6 commits2026-06-25: 2 commits2026-06-26: 5 commits2026-06-27: 0 commits2026-06-28: 1 commit2026-06-29: 2 commits2026-06-30: 3 commits2026-07-01: 4 commits2026-07-02: 6 commits2026-07-03: 5 commits2026-07-04: 0 commits2026-07-05: 1 commit2026-07-06: 2 commits2026-07-07: 7 commits2026-07-08: 4 commits2026-07-09: 0 commits2026-07-10: 1 commit2026-07-11: 1 commit2026-07-12: 3 commits2026-07-13: 8 commits2026-07-14: 5 commits2026-07-15: 3 commits2026-07-16: 6 commits2026-07-17: 6 commits2026-07-18: 1 commit2026-07-19: 0 commits2026-07-20: 11 commits2026-07-21: 10 commits2026-07-22: 16 commits2026-07-23: 5 commits2026-07-24: 6 commits2026-07-25: 2 commits2026-07-26: 0 commits2026-07-27: 16 commits2026-07-28: 8 commits2026-07-29: 11 commits2026-07-30: 11 commits2026-07-31: 1 commit2026-08-01: 1 commit2026-08-02: 2 commits2026-08-03: 9 commits2026-08-04: 13 commits2026-08-05: 5 commits2026-08-06: 7 commits2026-08-07: 3 commits2026-08-08: 3 commits2026-08-09: 1 commit2026-08-10: 6 commits2026-08-11: 8 commits2026-08-12: 1 commit2026-08-13: 6 commits2026-08-14: 1 commit2026-08-15: 1 commit2026-08-16: 2 commits2026-08-17: 6 commits2026-08-18: 2 commits2026-08-19: 2 commits2026-08-20: 5 commits2026-08-21: 2 commits2026-08-22: 0 commits2026-08-23: 3 commits2026-08-24: 4 commits2026-08-25: 7 commits2026-08-26: 5 commits2026-08-27: 11 commits2026-08-28: 3 commits2026-08-29: 2 commits2026-08-30: 2 commits2026-08-31: 8 commits2026-09-01: 3 commits2026-09-02: 14 commits2026-09-03: 9 commits2026-09-04: 6 commits2026-09-05: 1 commit2026-09-06: 2 commits2026-09-07: 2 commits2026-09-08: 2 commits2026-09-09: 2 commits2026-09-10: 1 commit2026-09-11: 1 commit2026-09-12: 4 commits2026-09-13: 0 commits2026-09-14: 2 commits2026-09-15: 1 commit2026-09-16: 5 commits2026-09-17: 3 commits2026-09-18: 2 commits2026-09-19: 2 commits2026-09-20: 2 commits2026-09-21: 3 commits2026-09-22: 4 commits2026-09-23: 2 commits2026-09-24: 4 commits2026-09-25: 1 commit2026-09-26: 1 commit2026-09-27: 1 commit2026-09-28: 2 commits2026-09-29: 2 commits2026-09-30: 3 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits
1,324 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Very active

    1,324 commits in 52 weeks

  • Community-driven

    ~211 contributors

  • Continuous integration

    Automated checks passing

  • Repeat trending

    5 trending appearances

What tilelang does

TileLang is a concise domain-specific language built to streamline the development of high-performance kernels for GPUs, CPUs, and accelerators. It utilizes a Pythonic syntax to make low-level programming more accessible and productive. The language is built on top of the TVM compiler infrastructure to ensure robust optimization. It excels at tasks like GEMM, FlashAttention, and LinearAttention implementations. Developers can achieve state-of-the-art hardware performance without sacrificing development speed.

Performance engineers, AI researchers, and systems programmers optimizing machine learning operations. It requires understanding of hardware architecture and compiler technologies.

  • Pythonic syntax: Employs a familiar, readable syntax to simplify the complex task of kernel development.
  • TVM backend integration: Leverages the powerful TVM compiler infrastructure for deep, low-level optimizations.
  • Broad hardware support: Targets GPUs, CPUs, and specialized NPUs/accelerators from a single codebase.
  • High-performance primitives: Built specifically for demanding operations like Dequant GEMM and FlashAttention.
  • Developer productivity: Focuses on rapid development cycles while maintaining state-of-the-art execution speeds.

Where teams use it

Custom hardware kernels

Develop highly optimized mathematical kernels for specific machine learning operations efficiently.

Cross-hardware compilation

Write a kernel once and compile it to run optimally across various CPUs, GPUs, and NPUs.

LLM optimization

Implement custom FlashAttention or GEMM routines to accelerate Large Language Model inference.

Performance engineering

Quickly prototype and test low-level hardware optimizations using a productive Pythonic interface.

README

main branch

TileLang logo

Tile Language (tile-lang) is a concise domain-specific language designed to streamline the development of high-performance GPU/CPU/NPU kernels (e.g., GEMM, Dequant GEMM, FlashAttention, LinearAttention). By employing a Pythonic syntax with an underlying compiler infrastructure on top of TVM, tile-lang allows developers to focus on productivity without sacrificing the low-level optimizations necessary for state-of-the-art performance.

TileLang tiled matrix multiplication example

Latest News

  • 2026-09-30 — Ascend 950 backend: TileLang now officially supports Huawei Ascend 950 NPUs with native code generation, automatic scheduling and synchronization, SIMD/SIMT vector programming, etc. Explore the Ascend examples for GEMM, FlashAttention, and more.
  • 2026-08-04 — TileLang LSP open sourced: published a Language Server Protocol implementation for TileLang with inlay hints for buffer shapes, dtypes, scopes, and inferred layouts, plus hover details and precise diagnostics.
  • 2026-08-03 — TileLang v0.1.13: shipped the multi-backend language dialect, source locations in compiler diagnostics, new CUDA and Metal hardware paths, and a broad set of correctness fixes. This release removes several legacy APIs; read the compatibility notes before upgrading.
  • 2026-07-30 — SM120 NVF4 block-scaled MMA: added an optimized Blackwell path for T.mma_gemm_blockscaled and a corresponding SM120 example.
  • 2026-07-28 — Metal 4 cooperative-tensor GEMM: added cooperative-tensor T.gemm support for Apple M5, while retaining the simdgroup fallback for unsupported shapes and systems.
Earlier news (2025–2026)

2026

  • 2026-07-28 — Source-aware compiler diagnostics: carried Python source locations into TIRX and surfaced them in compiler errors.
  • 2026-07-24 — Multi-backend language dialect: reorganized the language layer around shared semantics with static CUDA, ROCm, and Metal dialects.
  • 2026-07-23 — Block-causal attention for dLLM: added fixed-length and variable-length block-causal attention examples for diffusion language models.
  • 2026-07-22 — IR Lower Trace: introduced a debugging tool for inspecting IR changes across every compiler pass and the final code-generation step.
  • 2026-07-22 — DeepSeek V3.2 sparse MLA backward: selected the launch width adaptively from the head-block size.
  • 2026-07-21 — DeepSeek V3.2 top-k optimization: improved the sparse-attention top-k selector's memory access pattern, delivering approximately 1.9× higher performance in the reported benchmark.
  • 2026-07-16 — Compiler pass timing: added profiling for compiler passes with a configurable reporting threshold.
  • 2026-07-12 — IKET profiler integration: added CUDA timeline instrumentation and profiling support.
  • 2026-07-08 — TileLang v0.1.12: added the LLVM backend, tile scheduler, backend registry, pass visualizer, and expanded Blackwell support.
  • 2026-07-06 — Pass Visualizer: introduced a structure-tree browser for inspecting compiler transformations.
  • 2026-06-26 — Cross-host CUDA binary cache: enabled compiled CUDA binaries to be reused across compatible hosts.
  • 2026-06-24 — Tile scheduler: introduced persistent tile-scheduling primitives for kernel authors.
  • 2026-06-24 — Backend CodeGen registry: moved device and host CodeGen dispatch behind a backend registry.
  • 2026-06-18 — LLVM backend: added CPU lowering and execution through LLVM.
  • 2026-06-18 — Arbitrary-layout TMA lowering: enabled TMA transfers for swizzled shared-memory layouts.
  • 2026-06-16 — Pass Diff: added compiler-pass IR comparison for debugging lowering changes.
  • 2026-06-08 — TileLang v0.1.11: expanded scan, pipeline, backend, CUDA, ROCm, and Metal functionality.
  • 2026-05-25 — Scan operators: introduced tile-level scan primitives.
  • 2026-05-25 — TileLang v0.1.10: broadened AMD and Blackwell support, added initial Metal GEMM, improved Windows packaging, and expanded autotuning.
  • 2026-05-24 — CDNA4 MXFP4: added FP4 E2M1 matrix-core support for AMD gfx950.
  • 2026-05-22 — Metal simdgroup GEMM: added the first Metal T.gemm path using simdgroup_matrix MMA.
  • 2026-05-20 — Cluster copies: introduced T.copy_cluster for TMA multicast and SM-to-SM cluster transfers.
  • 2026-05-20 — TMA gather/scatter: added tile::gather4 and tile::scatter4 support.
  • 2026-05-20 — Native SM75 MMA GEMM: added FP16, INT8, and INT4 tensor-core paths for Turing GPUs.
  • 2026-05-20 — TIRX migration: moved TileLang IR usage to TVM's TIRX representation.
  • 2026-05-11 — Parallel autotuning: added pipelined compilation, grouped compilation, and multi-GPU benchmarking.
  • 2026-05-07 — DeepSeek V4 operators: added TileLang examples for DeepSeek V4 workloads.
  • 2026-05-06 — Windows support: added complete Windows build and runtime support with cross-platform fixes.
  • 2026-04-28 — MXFP8 grouped GEMM: added block-scaled grouped GEMM examples with transposed-B support on Blackwell.
  • 2026-04-25 — HISA sparse-attention indexer: added hierarchical sparse-attention indexing examples.
  • 2026-04-24 — Blackwell MXFP8 block-scaled GEMM: added MXFP8 block-scaled matrix multiplication on SM100.
  • 2026-04-22 — TileLang v0.1.9: delivered CuTe DSL GEMM V2, Metal code generation improvements, and build-without-host-toolchain support.
  • 2026-04-22 — RDNA3/RDNA3.5 WMMA: added WMMA lowering for AMD gfx11 GPUs.
  • 2026-04-20 — INT4 T.gemm: added INT4 matrix multiplication to the CUDA GEMM path.
  • 2026-04-17 — CUDA source kernels: introduced T.CUDASourceCodeKernel for embedding custom CUDA source.
  • 2026-04-15 — AutoDD frozen regions: added __freeze__ annotations to preserve selected code during automatic delta debugging.
  • 2026-03-27 — TMA stores: added store support to T.tma_copy.
  • 2026-03-24 — Two-SM Blackwell kernels: added two-SM TMA, TMEM, and TCGEN5 MMA support.
  • 2026-03-23 — AMD RDNA4: upgraded the ROCm path and added RDNA4 GPU support.
  • 2026-03-22 — FlashAttention on SM100: added Blackwell FlashAttention examples.
  • 2026-03-18 — Producer-consumer warp specialization: added automatic warp-specialized pipelines and the T.tma_copy API.
  • 2026-03-12 — Eager-mode autotuning: enabled the autotuner with eager JIT kernels.
  • 2026-03-10 — CPU T.gemm: added matrix multiplication support for the CPU target.
  • 2026-03-05 — IR dump configuration: added a TileLang pass configuration for dumping intermediate IR.
  • 2026-02-28 — CUDA cluster primitives: added cluster launch, query, synchronization, and barrier operations.
  • 2026-02-28 — TCGEN5 MMA tensor-shared path: added the tensor-memory/shared-memory Blackwell GEMM path.
  • 2026-02-24 — Host-toolchain-free builds: enabled installation without a host C/C++ toolchain when supported artifacts are available.
  • 2026-02-23 — CuTe DSL GEMM V2: added SM90 and SM100 GEMM V2 support to the CuTe DSL backend.
  • 2026-02-16 — TileLang v0.1.8: shipped dynamic pipeline improvements, logging documentation, richer layout representations, and AMD fixes.
  • 2026-02-14 — Cross-CUDA release wheels: unified multiple CUDA versions behind a single wheel.
  • 2026-02-14 — Hierarchical reductions: added hierarchical and warp-level reduction intrinsics.
  • 2026-02-09 — CUDA runtime stubs: added lazy-loading CUDART and NVRTC stubs for CUDA 11, 12, and 13 compatible wheels.
  • 2026-02-08 — Layout visualization improvements: improved rendering and inspection of TileLang layouts.
  • 2026-02-02 — TileLang Puzzles: published ten progressively harder exercises for learning TileLang interactively.

2025

See all releases for complete changelogs and compatibility notes.

Platform and Backend Support

TileLang is evolving into a multi-backend compiler (TileLang-X) built around a modular backend abstraction. See the backend architecture for the design, or ask a coding agent to use the backend integration skill when porting TileLang to a new backend.

The currently supported backends are listed below. Primary identifies TileLang's core backend, while Supported and Experimental backends are implemented in the main repository. Ecosystem adapters live in separate repositories, are not included in TileLang release wheels, and may follow independent compatibility schedules. Prebuilt wheels are available for Linux x86-64/AArch64, Windows x86-64, and macOS arm64.

TileLang uses Target objects to represent compilation targets. The default auto target detects CUDA, HIP, Metal, and Ascend devices; select an explicit target when compiling for another backend or architecture. See the target guide for target syntax, architecture options, and backend-specific notes, or the corresponding adapter repository for installation and tested-device details.

Backend Target Platforms and hardware Support level Notes
NVIDIA CUDA cuda Linux x86-64/AArch64, Windows x86-64; code paths from SM70 through SM120 Primary Release wheels and CI coverage; TMA, WGMMA, and TMEM features require the corresponding GPU architecture.
AMD ROCm/HIP hip Linux; CDNA and RDNA GPUs, including gfx942/gfx950 paths Supported Included in Linux wheels; a ROCm runtime is required. CI runs on a self-hosted gfx942 (MI300X) runner; gfx950 is not yet covered.
Huawei Ascend 950 ascend Linux; Ascend 950 NPUs Supported Build from source with USE_ASCEND=ON; requires CANN and torch_npu. See the Ascend guide.
Apple Metal metal macOS on Apple silicon Supported Release wheels and CI coverage; Metal 4 cooperative tensors are available on supported M5 systems.
LLVM CPU llvm Host CPUs Experimental Build from source with USE_LLVM=ON; LLVM 15 or newer is required.
NVIDIA CuTe DSL cutedsl NVIDIA GPUs Experimental Requires nvidia-cutlass-dsl.
WebGPU webgpu WebGPU runtimes Experimental Code generation and runtime integration are still evolving.
Huawei Ascend A2/A3 ascendc / pto / npuir Ascend A2 and A3 Ecosystem Developed in tilelang-ascend and the MLIR-based tilelang-mlir-ascend.
MetaX MACA maca MetaX C500 and C600 Ecosystem Developed in tilelang-metax; requires the MACA software stack.
Moore Threads MUSA musa S5000, S4000, and M1000 Ecosystem Developed in tilelang-musa and released independently.
HYGON hcu Linux; BW1000, BW1100, BW150 and K100_AI Ecosystem Developed in tilelang-hygon; requires the DTK software stack.
Sunrise-AI TANG tang Sunrise S2 and S3 Ecosystem Developed in tilelang-sunrise. The TANG software stack is required.

Installation

Install the latest stable release from PyPI:

pip install tilelang

Verify the installation:

python -c "import tilelang; print(tilelang.__version__)"

Nightly wheels provide recent features and fixes before the next stable release:

pip install tilelang --find-links https://tile-ai.github.io/whl/nightly

On AMD GPUs the same Linux wheels work out of the box: install a ROCm build of PyTorch first (e.g. pip install torch --index-url https://download.pytorch.org/whl/rocm7.0), then pip install tilelang. A host ROCm installation is required at runtime; see the ROCm notes in the installation guide.

Nightly builds may be less stable than official releases. For source builds, editable installs, Docker, ROCm setup, pip-provided CUDA toolchains, or a custom TVM checkout, follow the complete installation guide.

Quick Start

The following example defines, compiles, runs, and verifies an FP16 GEMM kernel with FP32 accumulation and a fused ReLU epilogue. It uses PyTorch CUDA tensors; PyTorch uses the same cuda device name on ROCm systems. TileLang selects the target automatically from the current environment.

import torch
import tilelang
import tilelang.language as T


@tilelang.jit
def matmul_relu(A, B, block_M: int = 128, block_N: int = 128, block_K: int = 32):
    M, N, K = T.const("M, N, K")
    A: T.Tensor((M, K), T.float16)
    B: T.Tensor((K, N), T.float16)
    C = T.empty((M, N), T.float16)

    with T.Kernel(T.ceildiv(N, block_N), T.ceildiv(M, block_M), threads=128) as (bx, by):
        A_shared = T.alloc_shared((block_M, block_K), T.float16)
        B_shared = T.alloc_shared((block_K, block_N), T.float16)
        C_local = T.alloc_fragment((block_M, block_N), T.float32)

        T.clear(C_local)
        for k in T.Pipelined(T.ceildiv(K, block_K), num_stages=3):
            T.copy(A[by * block_M, k * block_K], A_shared)
            T.copy(B[k * block_K, bx * block_N], B_shared)
            T.gemm(A_shared, B_shared, C_local)

        for i, j in T.Parallel(block_M, block_N):
            C_local[i, j] = T.max(C_local[i, j], 0)

        T.copy(C_local, C[by * block_M, bx * block_N])

    return C


M = N = K = 1024
a = torch.randn((M, K), device="cuda", dtype=torch.float16)
b = torch.randn((K, N), device="cuda", dtype=torch.float16)
c = matmul_relu(a, b)
torch.testing.assert_close(c, torch.relu(a @ b), rtol=1e-2, atol=1e-2)
print("GEMM + ReLU passed.")

@tilelang.jit specializes the kernel for the input shape and compile-time arguments on first use. T.Pipelined stages global-to-shared transfers, T.gemm maps the tile operation to the target backend, and T.Parallel expresses the elementwise ReLU epilogue. Continue with the language basics, then explore the GEMM examples for layouts, autotuning, and architecture-specific optimizations.

Examples

Browse the complete examples directory for additional operators, tests, and architecture-specific implementations.

Benchmark Summary

TileLang achieves exceptional performance across a variety of computational patterns. Comprehensive benchmark scripts and settings are available at tilelang-benchmark. Below are selected results showcasing its capabilities:

  • MLA Decoding Performance on H100

    mla decode performance bs64 on H100
    mla decode performance bs128 on H100
  • Flash Attention Performance on H100

    operator performance on H100
  • Matmul Performance on GPUs (RTX 4090, A100, H100, MI300X)

    gemm fp16 performance on Gpus
  • Dequantize Matmul Performance on A100

    dequantize gemv performance on A100

Join the Discussion

Welcome to join our Discord community for discussions, support, and collaboration!

Join our Discord

Acknowledgments

We would like to express our gratitude to the TVM community for their invaluable contributions. The initial version of this project was mainly developed by LeiWang1999, chengyupku and nox-410 with supervision from Prof. Zhi Yang at Peking University. Part of this work was carried out during an internship at Microsoft Research, where Dr. Lingxiao Ma, Dr. Yuqing Xia, Dr. Jilong Xue, and Dr. Fan Yang offered valuable advice and support. We deeply appreciate their mentorship and contributions.

View on GitHub

Recent activity

commits and pull requests

Discussions

all 25

Releases and announcements

23 total
  1. v0.1.15v0.1.15Sep 30, 202625 downloads

    ## Highlights - **Native Huawei Ascend 950 support** (#3308): an end-to-end NPU backend with native code generation, automatic Cube/Vector scheduling and synchronization, and mixed SIMD/SIMT programming. - **Automatic CUDA warp specialization** (#3059, #3185): an opt-in role-based scheduler that assigns TMA loads, MMA computation, TMA stores, and worker operations to specialized warp groups. - **Unified block-scaled GEMM** (#3237, #3257, #3284): common `T.gemm_blockscaled` semantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support. - **More expressive Python frontend** (#3230): compile-time iteration over Python iterables, `enumerate`, `zip`, comprehensions, and generator expressions. ## Ascend 950 - Add `tilelang.ascend.language` and `target="ascend"` for Huawei Ascend 950 (`dav-3510`). - Combine Cube GEMM and Vector computation in one kernel, with `T.SimdVF` and `T.SimtVF` regions. - Support explicit UB/L1/L0 storage, tiled copies, cross-core transfers, and MXFP8/MXFP4 block-scaled GEMM. - Add automatic scheduling, pipelining, multi-buffering, layout inference, and synchronization insertion. - Integrate Bis

  2. v0.1.14v0.1.14Sep 2, 202638 downloads

    ## Highlights - **Reducer v2** (#2940, #3093, #3043, #3044, #3079, #3100): `T.alloc_reducer` reworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates. - **Warp specialization schedules** (#2892): new scheduling and materialization mechanism for warp-specialized kernels. - **Layout inference cost models** (#2960, #3055, #3061): new IO-aware cost model for free-mode layout selection; register-count restored as the default, with an environment override to switch models. - **Unified backend resolution policy** (#2318) plus backend split-up (#2855, #2870, #2850): backend selection is now resolved through a single policy, and builtin ops / Python op proxies are split per backend (CUDA/ROCm/Metal). - **Compilation speed**: up to ~4x faster cold parallel/AOT compilation (#2809); Z3 solvers materialized lazily (#3105) and analyzer contexts isolated per kernel compilation (#2890). - **TMA rework**: TMA copy lowering unified on CuTe algebra (#3106); TMA layouts made region-awar

  3. v0.1.13v0.1.13Aug 3, 202633 downloads

    # TileLang v0.1.13 This release contains **138 commits** (79 bug fixes plus features, refactors, and examples) accumulated since v0.1.12 (2026-07-08 → 2026-08-02). The headline work is a **multi-backend language-dialect refactor** that replaces the runtime-activated language facade with static per-backend re-exports, alongside **two major new hardware paths**: SM120 NVF4 block-scale MMA for Blackwell and Metal 4 (M5) cooperative-tensor GEMM. On top of that, a large batch of correctness fixes landed across reductions/scans, atomics, TMA/copy lowering, loop-step preservation, and FP encoding edge cases. > **Breaking changes**: this release removes several legacy APIs and packages. See [Backend, API & Refactors](#backend-api--refactors) before upgrading. --- ## Highlights - **[CUDA] SM120 (Blackwell) NVF4 block-scale MMA support** (#2364) — `T.mma_gemm_blockscaled` now routes the packed-scale SM120 path internally through package-pingpong lowering, with an optimized non-persistent example (8192³ measured at ~1527 TFLOPS on SM120). The public `micro_pipeline` strategy knob was removed from the API. - **[Metal] M5 cooperative tensor `T.gemm`** (#2252) — TileLang-owned

  4. v0.1.12v0.1.12Jul 8, 2026142 downloads

    # TileLang v0.1.11 → v0.1.12 Changes Summary of the main changes between `v0.1.11` and `v0.1.12` (91 commits). ## New Features - **LLVM backend support** (#2409), with follow-up fixes for auto backend resolution (#2519) and module export (#2467) - **Tile scheduler** introduced (#2441) - **Backend registry architecture**: host and device CodeGen are now dispatched through a backend registry (#2442, #2446), target detection/normalization is registration-based, and ExecutionBackend was merged into the backend module (#2323); kernel launch is materialized per backend (#2387) - **Developer tooling**: `pass_visualizer` structure-tree pass browser (#2449) and pass-diff display for debugging (#2375) - New CUDA intrinsics exposed (#2473), `st.bulk` shared-memory zero fill on SM100+ (#2403), stmatrix m16n8 on Blackwell (#2417), SM75 MMA dispatchers for FP16 accumulation and UINT8 (#2392) ## CUDA / Codegen Improvements - **Optimized fp8↔half/bf16 casts**: vectorized and scalar cast codegen (#2511, #2475), plus a fix for vectorized fp16↔bf16 cast compilation (#2407) - **TMA lowering for arbitrary/swizzled SMEM layouts** (#2380) and GMMA/UMMA lowering for sliced (arbitrary-l

  5. v0.1.11v0.1.11Jun 8, 202669 downloads

    ## What's Changed * Fix atomic_load access_ptr lowering for dynamic indices by @VitalyAnkh in https://github.com/tile-ai/tilelang/pull/2157 * [Example] Add CLC-pipelined 2-CTA GEMM example for sm100 by @ighoshsubho in https://github.com/tile-ai/tilelang/pull/2169 * [Feature] Add thread_extent parameter to `T.tma_copy` for flexible TMA copy by @Rachmanino in https://github.com/tile-ai/tilelang/pull/2205 * Optimize disk cache source loading by @sepcnt in https://github.com/tile-ai/tilelang/pull/2176 * [CuTeDSL] Lower handle_add_byte_offset in Python codegen by @JayceSu98 in https://github.com/tile-ai/tilelang/pull/2261 * [FFI][Host] Refactor packed API binder to use FFI asserts by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2263 * [TileOP] Add scan operators by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2262 * [Feature] Add CUDA __ffs intrinsic for bit manipulation by @Rachmanino in https://github.com/tile-ai/tilelang/pull/2264 * [Bugfix] Fix cached source restore and Metal codegen fallback by @LeiWang1999 in https://github.com/tile-ai/tilelang/pull/2266 * [CuTeDSL] Represent tfloat32 storage as Float32 by @JayceSu98 in https://github.com/tile-ai/

Commits per week

last 52 weeks
500Week of 2025-10-05: 33 commitsWeek of 2025-10-12: 34 commitsWeek of 2025-10-19: 41 commitsWeek of 2025-10-26: 25 commitsWeek of 2025-11-02: 28 commitsWeek of 2025-11-09: 32 commitsWeek of 2025-11-16: 29 commitsWeek of 2025-11-23: 26 commitsWeek of 2025-11-30: 18 commitsWeek of 2025-12-07: 22 commitsWeek of 2025-12-14: 38 commitsWeek of 2025-12-21: 46 commitsWeek of 2025-12-28: 28 commitsWeek of 2026-01-04: 39 commitsWeek of 2026-01-11: 21 commitsWeek of 2026-01-18: 29 commitsWeek of 2026-01-25: 29 commitsWeek of 2026-02-01: 31 commitsWeek of 2026-02-08: 24 commitsWeek of 2026-02-15: 9 commitsWeek of 2026-02-22: 17 commitsWeek of 2026-03-01: 11 commitsWeek of 2026-03-08: 15 commitsWeek of 2026-03-15: 11 commitsWeek of 2026-03-22: 20 commitsWeek of 2026-03-29: 14 commitsWeek of 2026-04-05: 16 commitsWeek of 2026-04-12: 23 commitsWeek of 2026-04-19: 25 commitsWeek of 2026-04-26: 17 commitsWeek of 2026-05-03: 21 commitsWeek of 2026-05-10: 10 commitsWeek of 2026-05-17: 39 commitsWeek of 2026-05-24: 40 commitsWeek of 2026-05-31: 16 commitsWeek of 2026-06-07: 19 commitsWeek of 2026-06-14: 18 commitsWeek of 2026-06-21: 26 commitsWeek of 2026-06-28: 21 commitsWeek of 2026-07-05: 16 commitsWeek of 2026-07-12: 32 commitsWeek of 2026-07-19: 50 commitsWeek of 2026-07-26: 48 commitsWeek of 2026-08-02: 42 commitsWeek of 2026-08-09: 24 commitsWeek of 2026-08-16: 19 commitsWeek of 2026-08-23: 35 commitsWeek of 2026-08-30: 43 commitsWeek of 2026-09-06: 14 commitsWeek of 2026-09-13: 15 commitsWeek of 2026-09-20: 17 commitsWeek of 2026-09-27: 8 commitsOct 5, 2025Sep 27, 2026
1.3K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 21 commitsSun 1:00 — 11 commitsSun 2:00 — 11 commitsSun 3:00 — 4 commitsSun 4:00 — 0 commitsSun 5:00 — 1 commitsSun 6:00 — 1 commitsSun 7:00 — 0 commitsSun 8:00 — 2 commitsSun 9:00 — 0 commitsSun 10:00 — 3 commitsSun 11:00 — 1 commitsSun 12:00 — 7 commitsSun 13:00 — 7 commitsSun 14:00 — 6 commitsSun 15:00 — 9 commitsSun 16:00 — 8 commitsSun 17:00 — 16 commitsSun 18:00 — 7 commitsSun 19:00 — 5 commitsSun 20:00 — 12 commitsSun 21:00 — 11 commitsSun 22:00 — 11 commitsSun 23:00 — 11 commitsMon 0:00 — 18 commitsMon 1:00 — 9 commitsMon 2:00 — 8 commitsMon 3:00 — 4 commitsMon 4:00 — 1 commitsMon 5:00 — 0 commitsMon 6:00 — 2 commitsMon 7:00 — 1 commitsMon 8:00 — 1 commitsMon 9:00 — 0 commitsMon 10:00 — 5 commitsMon 11:00 — 9 commitsMon 12:00 — 19 commitsMon 13:00 — 48 commitsMon 14:00 — 26 commitsMon 15:00 — 27 commitsMon 16:00 — 36 commitsMon 17:00 — 36 commitsMon 18:00 — 15 commitsMon 19:00 — 24 commitsMon 20:00 — 24 commitsMon 21:00 — 14 commitsMon 22:00 — 16 commitsMon 23:00 — 9 commitsTue 0:00 — 14 commitsTue 1:00 — 19 commitsTue 2:00 — 11 commitsTue 3:00 — 7 commitsTue 4:00 — 1 commitsTue 5:00 — 1 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 1 commitsTue 10:00 — 3 commitsTue 11:00 — 22 commitsTue 12:00 — 33 commitsTue 13:00 — 34 commitsTue 14:00 — 41 commitsTue 15:00 — 38 commitsTue 16:00 — 24 commitsTue 17:00 — 24 commitsTue 18:00 — 17 commitsTue 19:00 — 15 commitsTue 20:00 — 16 commitsTue 21:00 — 12 commitsTue 22:00 — 13 commitsTue 23:00 — 13 commitsWed 0:00 — 17 commitsWed 1:00 — 8 commitsWed 2:00 — 11 commitsWed 3:00 — 3 commitsWed 4:00 — 2 commitsWed 5:00 — 1 commitsWed 6:00 — 0 commitsWed 7:00 — 3 commitsWed 8:00 — 0 commitsWed 9:00 — 1 commitsWed 10:00 — 6 commitsWed 11:00 — 20 commitsWed 12:00 — 17 commitsWed 13:00 — 37 commitsWed 14:00 — 40 commitsWed 15:00 — 40 commitsWed 16:00 — 36 commitsWed 17:00 — 31 commitsWed 18:00 — 15 commitsWed 19:00 — 13 commitsWed 20:00 — 21 commitsWed 21:00 — 10 commitsWed 22:00 — 18 commitsWed 23:00 — 19 commitsThu 0:00 — 14 commitsThu 1:00 — 9 commitsThu 2:00 — 10 commitsThu 3:00 — 6 commitsThu 4:00 — 1 commitsThu 5:00 — 0 commitsThu 6:00 — 2 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 4 commitsThu 10:00 — 3 commitsThu 11:00 — 21 commitsThu 12:00 — 22 commitsThu 13:00 — 32 commitsThu 14:00 — 22 commitsThu 15:00 — 26 commitsThu 16:00 — 25 commitsThu 17:00 — 30 commitsThu 18:00 — 22 commitsThu 19:00 — 14 commitsThu 20:00 — 21 commitsThu 21:00 — 11 commitsThu 22:00 — 10 commitsThu 23:00 — 19 commitsFri 0:00 — 9 commitsFri 1:00 — 9 commitsFri 2:00 — 8 commitsFri 3:00 — 5 commitsFri 4:00 — 2 commitsFri 5:00 — 1 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 1 commitsFri 9:00 — 0 commitsFri 10:00 — 4 commitsFri 11:00 — 17 commitsFri 12:00 — 18 commitsFri 13:00 — 27 commitsFri 14:00 — 27 commitsFri 15:00 — 23 commitsFri 16:00 — 25 commitsFri 17:00 — 38 commitsFri 18:00 — 11 commitsFri 19:00 — 7 commitsFri 20:00 — 9 commitsFri 21:00 — 12 commitsFri 22:00 — 6 commitsFri 23:00 — 16 commitsSat 0:00 — 9 commitsSat 1:00 — 8 commitsSat 2:00 — 4 commitsSat 3:00 — 5 commitsSat 4:00 — 1 commitsSat 5:00 — 1 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 6 commitsSat 10:00 — 3 commitsSat 11:00 — 3 commitsSat 12:00 — 4 commitsSat 13:00 — 12 commitsSat 14:00 — 10 commitsSat 15:00 — 8 commitsSat 16:00 — 10 commitsSat 17:00 — 12 commitsSat 18:00 — 11 commitsSat 19:00 — 9 commitsSat 20:00 — 8 commitsSat 21:00 — 6 commitsSat 22:00 — 6 commitsSat 23:00 — 10 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Oct 4, 2026weekly#11+800
Oct 3, 2026weekly#12+732
Oct 2, 2026daily#7+157
Oct 2, 2026weekly#12+481
Oct 1, 2026daily#7+157