The same function, a different execution plan同一个函数,不同的执行计划

The previous installments followed individual operators through dispatch, autograd and CUDA submission. Now consider (x + bias).relu() * 2. Eager execution handles these operations one at a time. Capturing the region as a graph lets a compiler consider them together: can the intermediate tensors disappear, and can one generated loop compute the result?前几期沿着单个算子走过了调度、自动微分与 CUDA 提交。现在看 (x + bias).relu() * 2:eager 模式逐个处理这些操作。把这一段捕获为图后,编译器便能一起分析:能否去掉中间张量,让一个生成的循环直接计算结果?

import torch

def block(x, bias):
    return (x + bias).relu() * 2

x = torch.tensor([-2., -1., 1., 2.], requires_grad=True)
bias = torch.ones(4, requires_grad=True)
compiled = torch.compile(block, fullgraph=True)
y = compiled(x, bias)
torch.testing.assert_close(y, block(x, bias))
y.sum().backward()
torch.testing.assert_close(x.grad, torch.tensor([0., 0., 2., 2.]))
torch.testing.assert_close(bias.grad, x.grad)
print("output:", y.tolist())
print("x.grad:", x.grad.tolist())
print("bias.grad:", bias.grad.tolist())
output: [0.0, 0.0, 4.0, 6.0]
x.grad: [0.0, 0.0, 2.0, 2.0]
bias.grad: [0.0, 0.0, 2.0, 2.0]

Baseline: PyTorch 2.10.0, commit 449b17684101, Linux CPU, default Inductor backend unless a custom backend is explicitly selected. The Python programs were executed with 2.10.0+cpu; forward values, gradients and capture counts below are observed results. CUDA code generation is discussed from the implementation, not presented as a GPU measurement.基线:PyTorch 2.10.0,提交 449b17684101,Linux CPU;除明确指定自定义后端外,使用默认 Inductor 后端。Python 程序已在 2.10.0+cpu 上执行,下文的前向值、梯度和捕获次数均为实际结果。CUDA 代码生成依据实现讲解,不作为 GPU 实测。

torch.compile(block) returns a wrapper. Compilation is generally triggered when that wrapper first sees inputs, not when the wrapper is created. fullgraph=True requires one captured graph for the function or raises an error; it does not promise one generated kernel. A graph may contain multiple kernels and external library calls.torch.compile(block) 返回包装后的可调用对象。编译通常在包装器首次收到输入时触发,而不是创建包装器时触发。fullgraph=True 要求把函数捕获为一张图,否则报错;它不承诺只生成一个内核。一张图中仍可能包含多个内核和外部库调用。

torch.compile API: graph capture, dynamic shapes and backend selectiontorch.compile API:图捕获、动态形状与后端选择

Separate the compilation path from the reuse path区分编译路径与复用路径

The diagram follows a forward invocation with static shape specialization. On a cache miss, the left branch prepares executable code. On a cache hit, guards allow the call to use an existing variant directly. The AOTAutograd box includes differentiation-graph preparation; backward machine-code compilation can be deferred until the first backward call. This is a control-flow illustration, not a time-proportional trace.图示跟踪静态形状特化下的一次前向调用。缓存未命中时,左侧路径准备可执行代码;命中时,守卫允许调用直接使用已有版本。AOTAutograd 方框包含求导图的准备过程,反向机器码的编译可能推迟到首次反向调用。这是控制流程示意,不是按耗时比例绘制的轨迹。

x, bias: float32[4] · static
GuardsCheck assumptions检查特化条件
DynamoPython → FX
AOTAutogradFW / BW求导图
InductorCodegen代码生成
Variants缓存版本{4}
ExecuteKernel call执行内核
Output输出[4]

First call: capture and compile a variant for length 4, then execute it.首次调用:为长度 4 捕获并编译一个版本,然后执行。

Compare four calls对照四次调用

Observe FX graphs and the conditions for reuse观察 FX 图与复用条件

Dynamo analyzes Python bytecode and extracts supported tensor operations into an FX GraphModule. It specializes under assumptions about inputs and surrounding state. Tensor guards can include type, dtype, device, gradient requirements, size and stride; Python constants and other state can matter too. Equal shapes alone are therefore not a complete cache key. Changing tensor values does not by itself require recompilation for this pointwise block.Dynamo 分析 Python 字节码,将支持的张量操作提取为 FX GraphModule。特化依赖输入与外围状态的一组假设。张量守卫可包含类型、dtype、设备、梯度需求、尺寸和步长;Python 常量等状态也可能参与判断。因此,仅形状相同并不构成完整的缓存键。对于这里的逐元素计算,单纯改变张量中的数值不要求重新编译。

eval_frame.py: Dynamo optimization entryeval_frame.py:Dynamo 优化入口

guards.py: fields used in tensor guard checksguards.py:张量守卫检查的字段

A backend receives the FX graph and example inputs, and returns a callable. Returning gm.forward below makes the capture visible without generating optimized kernels. The counter measures backend invocations, not GPU launches, and this experiment deliberately resets compiler state before each policy.后端接收 FX 图和示例输入,并返回可调用对象。下面返回 gm.forward,目的是观察图捕获,不生成优化内核。计数器统计的是后端调用次数,不是 GPU 启动次数;实验在每种策略前显式重置编译器状态。

import torch

def block(x, bias):
    return (x + bias).relu() * 2

for dynamic in (False, True):
    torch.compiler.reset()
    graphs = []
    def inspect_backend(gm, example_inputs):
        graphs.append(gm)
        return gm.forward
    probe = torch.compile(block, backend=inspect_backend,
                          fullgraph=True, dynamic=dynamic)
    counts = []
    for size in (4, 4, 8, 4):
        x, bias = torch.randn(size), torch.randn(size)
        torch.testing.assert_close(probe(x, bias), block(x, bias))
        counts.append(len(graphs))
    print("dynamic:", dynamic, "captures:", counts)
    print("ops:", [getattr(node.target, "__name__", str(node.target))
                   for node in graphs[0].graph.nodes
                   if node.op in ("call_function", "call_method")])
dynamic: False captures: [1, 1, 2, 2]
ops: ['add', 'relu', 'mul']
dynamic: True captures: [1, 1, 1, 1]
ops: ['add', 'relu', 'mul']

For these one-dimensional inputs of lengths 4 and 8, symbolic size handling avoids the second capture. It does not make every input valid: layout, dtype, control-flow constraints and special sizes can still require specialization. The default dynamic=None is a separate policy that may generalize after a size-related recompile. Do not turn an observation from this small graph into a promise of zero recompilations for a model.对于这些长度为 4 和 8 的一维输入,符号尺寸处理避免了第二次捕获,但并非任意输入都能复用:布局、dtype、控制流约束和特殊尺寸仍可能要求特化。默认的 dynamic=None 是另一种策略,可能在尺寸变化引起重新编译后尝试泛化。这个小图的结果不代表整个模型都不会重新编译。

The backend handoff is explicit in OutputGraph._call_user_compiler. FX describes dependencies and operations; its Python representation is not the final machine code.后端交接在 OutputGraph._call_user_compiler 中是显式调用。FX 描述操作及其依赖,图的 Python 表示并不是最终机器码。

compiled_fn = compiler_fn(gm, example_inputs)
assert callable(compiled_fn)

output_graph.py: pass an FX graph to a backend and receive a callableoutput_graph.py:将 FX 图传给后端并接收可调用对象

torch.fx: Graph, Node and GraphModuletorch.fx:Graph、Node 与 GraphModule

AOTAutograd prepares the backward computation tooAOTAutograd 也准备反向计算

For training, AOTAutograd traces differentiation and partitions a joint forward/backward representation into forward and backward graphs. The partition determines which values cross the boundary as saved tensors and which computations may be recomputed. “Ahead of time” here concerns preparing differentiation for the captured region; it does not mean the whole Python application has been compiled offline.训练时,AOTAutograd 跟踪求导过程,将前向/反向联合表示划分为前向图和反向图。划分决定哪些值作为保存的张量跨越边界,哪些计算可能重新执行。这里的“提前”指为捕获区域准备求导,不表示整个 Python 应用已离线编译。

In the first program, summation supplies an output cotangent of 1. The positive-input mask gives gradients [0, 0, 2, 2]; PyTorch’s ReLU uses derivative 0 at the zero boundary. Because the two inputs have identical shapes, no broadcast-gradient reduction is needed. With a broadcast bias, its gradient would also need reduction over the expanded dimensions.第一个程序中的求和为输出端提供全 1 的余切向量。正输入掩码给出梯度 [0, 0, 2, 2];PyTorch 的 ReLU 在零点采用导数 0。两个输入形状相同,因此不需要广播梯度归约。如果 bias 通过广播参与计算,它的梯度还需要沿展开的维度归约。

aot_autograd.py: joint tracing, graph partitioning and compiler callbacksaot_autograd.py:联合跟踪、图划分与编译回调

Inductor’s compile_fx connects separate forward, backward and inference compiler callbacks to AOTAutograd. Compiled backward regions still participate in autograd execution; ordinary torch.compile does not imply that every surrounding backward-engine action or hook has become one graph. This is also distinct from the separately configurable Compiled Autograd feature.Inductor 的 compile_fx 将前向、反向和推理编译回调分别连接到 AOTAutograd。编译后的反向区域仍参与 autograd 执行;普通 torch.compile 不表示外围反向引擎的所有动作或 hook 都变成了一张图。这也区别于单独配置的 Compiled Autograd 功能。

compile_fx.py: fw_compiler, bw_compiler and inference_compilercompile_fx.py:前向、反向与推理编译回调

runtime_wrappers.py: deferred backward compilationruntime_wrappers.py:延迟反向编译

Fusion removes intermediate tensor traffic融合减少中间张量的读写

Inductor lowers the graph into its internal representation and schedules compatible computations together. For this pointwise chain, one loop can load x[i] and bias[i], perform add → ReLU → multiply locally, then store y[i]. The two intermediate values still exist in the computation, but need not be materialized as full tensors.Inductor 将图转换为自身的低层中间表示,并把兼容计算组合调度。对于这个逐元素计算链,一个循环即可加载 x[i] 和 bias[i],在局部完成加法 → ReLU → 乘法,最后存储 y[i]。两个中间值仍存在于计算中,但不必物化为完整张量。

Three array passes三次遍历数组

add → t1
-1023
relu → t2
0023
mul → y
0046
7N = 28

Logical float reads + writes; t1 and t2 are materialized arrays.逻辑 float 读写次数;t1、t2 是物化数组。

One fused pass一次融合遍历

x[i] + bias[i]Load two inputs加载两个输入
max(s, 0) * 2Loop-local values循环内的局部值
store → y
0046
3N = 12

Same output; no full t1 or t2 buffer in this model.输出相同;此模型不需要完整的 t1、t2 缓冲区。

These counts assume float32, materialized intermediates and one read per input use. They describe logical array accesses, not measured DRAM traffic: caches, vectorization and register allocation alter hardware behavior. The ratio 7/3 is not a promised speedup. Fusion also has limits—dependencies, reductions, register pressure and external library boundaries affect what is legal and worthwhile.这些计数假设使用 float32、中间结果已物化、每次使用输入都读取一次。它们描述逻辑数组访问,不是实测 DRAM 流量:缓存、向量化与寄存器分配会改变硬件行为。7/3 不是承诺的加速比。融合也有限制,依赖、归约、寄存器压力与外部库边界都会影响是否合法、是否值得融合。

scheduler.py: fusion of compatible scheduler nodesscheduler.py:融合兼容的调度节点

#include <algorithm>
#include <cmath>
#include <iostream>
#include <stdexcept>
#include <string>
#include <vector>

int main(int argc, char** argv) {
    const int n = argc > 1 ? std::stoi(argv[1]) : 4;
    if (n < 1 || n > 100000) throw std::invalid_argument("invalid size");
    const float seed[] = {-2.f, -1.f, 1.f, 2.f};
    std::vector<float> x(n), bias(n, 1.f), t1(n), t2(n), separate(n), fused(n);
    for (int i = 0; i < n; ++i) x[i] = seed[i % 4];
    std::size_t separate_accesses = 0, fused_accesses = 0;
    for (int i = 0; i < n; ++i) {
        t1[i] = x[i] + bias[i]; separate_accesses += 3;
    }
    for (int i = 0; i < n; ++i) {
        t2[i] = std::max(t1[i], 0.f); separate_accesses += 2;
    }
    for (int i = 0; i < n; ++i) {
        separate[i] = t2[i] * 2.f; separate_accesses += 2;
    }
    for (int i = 0; i < n; ++i) {
        fused[i] = std::max(x[i] + bias[i], 0.f) * 2.f;
        fused_accesses += 3;
        if (std::abs(fused[i] - separate[i]) > 1e-6f)
            throw std::runtime_error("output mismatch");
    }
    std::cout << "output:";
    for (int i = 0; i < std::min(n, 8); ++i) std::cout << ' ' << fused[i];
    std::cout << "\nlogical accesses: " << separate_accesses
              << " -> " << fused_accesses << '\n';
}

Output checked locally and with GCC 14.2 on Godbolt:以下输出已在本地及 Godbolt 的 GCC 14.2 中核对:

output: 0 0 4 6
logical accesses: 28 -> 12

This standalone C++ program makes the dataflow and access-count model executable. It is not Inductor’s generated kernel and does not measure runtime memory transactions; a C++ compiler may itself optimize the loops. Counters cover expression evaluation only, excluding initialization, validation and printing.这个独立 C++ 程序让数据流和访问计数模型可以执行。它不是 Inductor 生成的内核,也不测量运行时内存事务;C++ 编译器自身还可能优化这些循环。计数器只覆盖表达式求值,不包含初始化、校验与打印。

Inspect the generated code, not just the FX graph检查生成代码,而不只看 FX 图

Use the following standalone inference program to keep the generated region small. Save it as compile_block.py and run it with the logging command below. The CPU run used for this article generated a fused cpp_fused_add_mul_relu_0 kernel with vector loads, local arithmetic and a final store. Names and vector widths depend on the build and target CPU.下面的独立推理程序让生成区域保持简洁。保存为 compile_block.py,再用后面的日志命令运行。本文的 CPU 运行生成了融合的 cpp_fused_add_mul_relu_0 内核,其中包含向量加载、局部算术和最终存储。名称与向量宽度取决于构建和目标 CPU。

import torch

def block(x, bias):
    return (x + bias).relu() * 2

compiled = torch.compile(block, fullgraph=True)
x = torch.tensor([-2., -1., 1., 2.])
bias = torch.ones(4)
with torch.inference_mode():
    y = compiled(x, bias)
    torch.testing.assert_close(y, block(x, bias))
    print(y.tolist())
TORCH_LOGS="graph_code,aot_graphs,output_code" python compile_block.py

An arithmetic fragment from the generated CPU kernel is shown below. The two loaded vectors are tmp0 and tmp1; tmp5 broadcasts the value 2. Vector load/store and wrapper code are omitted, so this excerpt is not standalone.下面摘出生成的 CPU 内核中的算术部分。tmp0 与 tmp1 是加载的两个向量,tmp5 是广播后的数值 2。这里省略了向量加载、存储和包装代码,片段不能单独运行。

auto tmp2 = tmp0 + tmp1;
auto tmp3 = at::vec::clamp_min(tmp2, decltype(tmp2)(0));
auto tmp6 = tmp3 * tmp5;

GraphLowering.codegen() invokes the scheduler and generates wrapper code; compile_to_module() makes the result loadable. For supported CUDA computations, Inductor commonly emits Triton kernels, while selected operations may use external libraries. It does not turn every Python statement into handwritten CUDA C++, and the best implementation of a matrix multiply need not be an elementwise fused loop.GraphLowering.codegen() 调用调度器并生成包装代码,compile_to_module() 将结果变为可加载模块。对于支持的 CUDA 计算,Inductor 常生成 Triton 内核,某些操作则可能调用外部库。它并非把每条 Python 语句变成手写 CUDA C++;矩阵乘法的合适实现也未必是逐元素融合循环。

graph.py: scheduling, wrapper generation and module compilationgraph.py:调度、包装代码生成与模块编译

codegen/cpp.py: CPU scalar and vector code generationcodegen/cpp.py:CPU 标量与向量代码生成

codegen/triton.py: Triton kernel generationcodegen/triton.py:Triton 内核生成

A graph break is not a guard failure图中断不是守卫失败

A graph break ends capture of a region so Python can execute outside that graph; capture may resume afterward. A guard failure means an existing specialization does not apply to a call. If no suitable cached variant exists, another capture/compilation may be needed. One concerns capture boundaries, the other reuse conditions. They can occur in the same program but diagnose different problems.图中断结束一个区域的捕获,让 Python 在图外执行,之后可能继续捕获。守卫失败表示已有特化不适用于本次调用;若没有合适的缓存版本,可能需要再次捕获/编译。前者关乎捕获边界,后者关乎复用条件;同一程序中可以同时出现,但对应不同问题。

import torch

@torch.compiler.disable
def observe(value):
    print("observed length:", value.shape[0])
    return value

def with_boundary(x):
    y = x + 1
    y = observe(y)
    return y.relu() * 2

regions = []
def record_backend(gm, example_inputs):
    regions.append(gm)
    return gm.forward

x = torch.tensor([-2., -1., 1., 2.])
result = torch.compile(with_boundary, backend=record_backend)(x)
print("output:", result.tolist(), "regions:", len(regions))
try:
    torch.compile(with_boundary, backend="eager", fullgraph=True)(x)
except torch._dynamo.exc.Unsupported:
    print("fullgraph: disabled function cannot be included")
observed length: 4
output: [0.0, 0.0, 4.0, 6.0] regions: 2
fullgraph: disabled function cannot be included

Here the boundary is intentional: torch.compiler.disable keeps the observer in ordinary Python. It demonstrates a graph break without relying on whether a particular print or scalar-extraction pattern happens to be supported in a release. The private exception class is used only to catch this diagnostic in the pinned version. In production, investigate unexpected breaks with TORCH_LOGS="graph_breaks,recompiles,guards", and avoid disabling guards to hide recompilations.这里的边界是有意设置的:torch.compiler.disable 让观察函数保留在普通 Python 中执行。这样演示图中断,无需依赖某个版本是否恰好支持特定打印或标量提取写法。私有异常类型仅用于在固定版本下捕获此诊断。实际项目中,可用 TORCH_LOGS="graph_breaks,recompiles,guards" 查找意外中断,不要通过关闭守卫掩盖重新编译。

torch.compiler.disable: exclude a function from compilationtorch.compiler.disable:将函数排除在编译区域之外

Separate first-call cost from steady-state execution将首次调用成本与稳态执行分开测量

The first invocation can include capture, compiler work, code loading and execution. A warm invocation still pays guard and wrapper costs. A new process may reuse an on-disk compiler artifact, so “first call” does not automatically mean a completely cold compilation. Report cache conditions and count shape variants rather than averaging compilation into an unexplained latency number.首次调用可能包含图捕获、编译器工作、代码加载与执行。预热后的调用仍有守卫和包装器开销。新进程也可能复用磁盘上的编译产物,因此“首次调用”不自动等于完全冷编译。报告缓存条件并统计形状版本,不要把编译混进一个未说明边界的平均延迟。

This CPU-only benchmark fixes the thread count and inference context for both paths, verifies results, then retains per-block measurements normalized by the number of runs. It reports the first compiled invocation separately. These block averages are not individual request-latency percentiles. To measure CUDA, use the synchronization boundaries from part 5.这个纯 CPU 基准测试为两条路径固定线程数和推理上下文,先验证结果,再保留按运行次数归一化的分块测量值,并单独报告首次编译调用。分块平均值不是单次请求延迟的分位数。若测量 CUDA,需要采用第五期介绍的同步边界。

import json
import platform
from time import perf_counter
import torch
from torch.utils.benchmark import Timer

def block(x, bias):
    return (x + bias).relu() * 2

torch.set_num_threads(1)
torch.manual_seed(0)
x = torch.randn(262144)
bias = torch.randn_like(x)
compiled = torch.compile(block, fullgraph=True)
with torch.inference_mode():
    begin = perf_counter()
    actual = compiled(x, bias)
    first_call = perf_counter() - begin
    torch.testing.assert_close(actual, block(x, bias))
    for _ in range(5):
        compiled(x, bias)
    measurements = {}
    for name, fn in (("eager", block), ("compiled", compiled)):
        sample = Timer(stmt="fn(x, bias)",
                       globals={"fn": fn, "x": x, "bias": bias},
                       num_threads=1).blocked_autorange(min_run_time=0.2)
        measurements[name] = {
            "median_seconds": sample.median,
            "runs_per_block": sample.number_per_run,
            "seconds_per_call_by_block": [
                t / sample.number_per_run for t in sample.raw_times
            ],
        }
print(json.dumps({
    "torch": torch.__version__, "machine": platform.machine(),
    "device": "cpu", "dtype": str(x.dtype), "shape": list(x.shape),
    "threads": 1, "inference_mode": True,
    "cache_policy": "existing disk caches may be reused",
    "first_compiled_call_seconds": first_call,
    "steady_state": measurements,
}, indent=2))

torch.utils.benchmark: Timer, blocked_autorange and Measurementtorch.utils.benchmark:Timer、blocked_autorange 与 Measurement

A short workload can lose to compilation and guard overhead, and a workload already dominated by an optimized library call may have little fusible work. For training, include the first and steady-state backward calls in the chosen measurement boundary. Inspect correctness, guards and generated code before attributing a change to faster arithmetic.短任务可能被编译和守卫开销抵消收益;已经主要耗时于优化库调用的任务,也可能没有多少可融合的工作。训练场景还应在选定边界内测量首次和稳态反向调用。在把性能变化归因于算术加速前,先检查正确性、守卫和生成代码。

Locate a failure by the stage that introduces it按引入问题的阶段定位故障

Backend后端What it exercises覆盖的阶段
backend="eager"Dynamo capture, then execute the captured graph with PyTorch.Dynamo 捕获,然后用 PyTorch 执行捕获的图。
backend="aot_eager"Add AOTAutograd graph preparation; execute without Inductor kernel generation.加入 AOTAutograd 图准备,但不进行 Inductor 内核生成。
backend="inductor"Add lowering, scheduling and generated or external kernels.加入低层表示转换、调度,以及生成的内核或外部内核。

Use these backends on a minimal reproducer with matching inputs, gradient mode and tolerances. A failure appearing only with Inductor narrows the investigation; it does not alone identify the faulty pass. The source trail is now connected: tensor metadata informs guards, operator graphs feed compilation, autograd supplies differentiation, and the resulting kernels still execute through the CPU or CUDA runtime.在最小复现中保持输入、梯度模式和误差容限一致,再依次使用这些后端。只有 Inductor 才出现的失败可以缩小排查范围,但不能单凭这一点确定哪个优化过程有错。至此,源码路径已经连起来:张量元信息参与守卫,算子图进入编译,autograd 提供求导,而最终内核仍通过 CPU 或 CUDA 运行时执行。

backends/debugging.py: eager and aot_eager debugging backendsbackends/debugging.py:eager 与 aot_eager 调试后端