Skip to main content

Triton

non-cncf
First Research Pass

Content Lifecycle & Evolution State

1
Initial Seed
Metadata Loaded
2
First Research
Automated Gathering
3
Tech Writing
Content Structured
4
Approval
Editorial Verified
First Research Pass Completed: Initial facts, key features, and documentation links have been gathered. Detailed technical writing and synthesis are currently in queue.

Overview

Triton is a language and compiler for writing highly efficient custom Deep-Learning primitives. It aims to provide an open-source environment for writing fast code at higher productivity than CUDA, but with higher flexibility than other existing DSLs.

Key Features

  • High Productivity: Provides an environment to write fast code more productively than CUDA.
  • High Flexibility: Offers greater flexibility compared to other domain-specific languages (DSLs) for deep learning.
  • MLIR-based Backend: Features a compiler backend rewritten to use MLIR.
  • Back-to-back Matmuls: Supports kernels that contain back-to-back matmuls, enabling operations like flash attention.

Use Cases

Writing highly efficient custom Deep-Learning primitives and optimizing deep learning workloads, such as flash attention and complex matmul operations.

Getting Started

Check out the official documentation at https://triton-lang.org/ for installation instructions and tutorials, or explore the Triton puzzles to practice running the Triton interpreter without a GPU.

Recent Updates

Version 3.8.0 includes Proton profiling improvements like CUDA graph profiling, new examples and tutorials, low-precision matmul additions, storage shape queries, performance improvements, standalone CUDA backend (can run without PyTorch), and various bug fixes.

Interesting Facts

Triton's foundations were described in a MAPL 2019 publication: 'Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations'.