The same numbers, three different contracts数值相同,承诺不同
Take two tensors, x = [1, 2] and b = [10, 20], and calculate x + 2 * b. The result is [21, 42]. Yet torch.add(x, b, alpha=2), x.add_(b, alpha=2) and torch.add(x, b, alpha=2, out=out) promise different things about the objects that hold those numbers. A compiler, a binding and the dispatcher need more information than “this is addition”.取两个张量 x = [1, 2]、b = [10, 20],计算 x + 2 * b,结果都是 [21, 42]。但 torch.add(x, b, alpha=2)、x.add_(b, alpha=2) 与 torch.add(x, b, alpha=2, out=out) 对承载数值的对象作出了不同承诺。编译器、语言绑定和调度器需要知道的,不止“这是加法”。
PyTorch records this part of the contract in an operator schema. For built-in ATen operators, a schema string lives in the func field of native_functions.yaml, alongside generation metadata. Together, these declarations serve an interface-description role—IDL, or Interface Definition Language. The YAML is the container; the schema inside it is the language describing an operator.PyTorch 用 operator schema(算子接口模式)记录这部分契约。对 ATen 内置算子,schema 字符串写在 native_functions.yaml 的 func 字段中,旁边还有代码生成所需的元信息。这套声明承担接口描述的职责,也就是这里所说的 IDL(Interface Definition Language)。YAML 是容器,内部的 schema 才是描述算子的语言。
This installment follows PyTorch 2.10.0, commit 449b17684101. The three Python programs below were executed with 2.10.0+cpu; the generation path was checked against that source revision. It connects part 3’s operator registration with part 6’s compiled execution: first establish what an operator promises, then follow how its interfaces are built. This IDL describes local tensor operations; it does not specify an RPC wire format.本期沿用 PyTorch 2.10.0,提交 449b17684101。下文三个 Python 程序已在 2.10.0+cpu 中执行,生成路径按这一版本源码核对。内容连接第三期的算子注册与第六期的编译执行:先看一个算子承诺什么,再追踪接口怎样建立。这里的 IDL 描述本地张量操作,不规定 RPC 通信格式。
Read a schema as a sentence把一行 schema 读成一句话
Start with the functional tensor overload. The registered name adds the ATen namespace to the declaration. Read the line from left to right: identify the operation, bind its inputs, apply argument rules, then describe its return.先看接收张量的函数式重载。注册后的名称在声明前加上 ATen 命名空间。按从左到右的顺序读:确定操作,绑定输入,说明参数规则,再描述返回值。
- aten::add.Tensor
- Namespace, operator name and overload name.
Tensorhere names the overload; it is not a device.命名空间、算子名与重载名。这里的Tensor是重载名称,不表示设备。 - Tensor self, Tensor other
- Two tensor arguments;
selfis the first argument even for the function API.两个张量参数;即使使用函数形式,self仍是第一个参数的名称。 - *, Scalar alpha=1
- The following alpha argument is keyword-only at the Python boundary, with default 1. Scalar accepts numeric scalar values.Python 接口中,此后的 alpha 只能按关键字传入,默认值为 1。Scalar 接收数值标量。
- → Tensor
- A tensor return with no declared alias to either input. The input arguments are not marked writable.返回张量,没有声明与输入的别名关系;两个输入参数也没有可写标记。
The overload name is part of the operator’s identity: aten::add.Tensor and aten::add.Scalar are distinct schemas. Python’s torch.add can expose multiple overloads under one callable. Conversely, variants: function, method can expose the same declaration as a function and a tensor method. Python spellings, dispatcher identities and C++ function signatures are related, but they are not interchangeable names.重载名属于算子标识的一部分:aten::add.Tensor 与 aten::add.Scalar 对应不同 schema。Python 的 torch.add 可以在一个可调用对象下提供多个重载;反过来,variants: function, method 又能把同一声明暴露为函数和张量方法。Python 写法、调度器中的算子标识、C++ 函数签名彼此关联,却不能直接当成同一个名字。
A few symbols recur throughout ATen: Tensor? permits None (omitting the argument also requires a default), Tensor[] is a tensor list, and SymInt supports a symbolic integer. A plain Tensor does not encode a fixed shape, dtype or device. Binding the declared argument types therefore does not establish that broadcasting, an in-place write or a particular backend will succeed.ATen 中还常见几个符号:Tensor? 允许传入 None(能否省略参数还取决于默认值),Tensor[] 表示张量列表,SymInt 支持符号整数。单独一个 Tensor 并不写明固定形状、dtype 或设备。因此,参数符合声明类型,不代表广播、原地写入或某个后端的计算一定成功。
ATen native-function guide: schema syntax and argument conventionsATen 原生函数指南:schema 语法与参数约定
Aliasing is part of the interface别名与修改,也是接口的一部分
aten::add.Tensor(Tensor self, Tensor other, *, Scalar alpha=1) -> Tensor
aten::add_.Tensor(Tensor(a!) self, Tensor other, *, Scalar alpha=1) -> Tensor(a!)
aten::add.out(Tensor self, Tensor other, *, Scalar alpha=1, Tensor(a!) out) -> Tensor(a!)
The letter a names an alias set inside this schema; it is not the variable name x, a memory address or a global identifier shared by every operator. The ! marks a writable tensor. In add_, the input self and the return carry the same annotation, expressing their alias relationship. In add.out, that relationship connects out and the return. The underscore is a naming convention; the annotation makes the effect explicit for machinery reading the schema.a 为当前 schema 内的别名集合命名:它不是变量 x 的名字、内存地址,也不是所有算子共享的全局编号。! 标记张量可以被写入。add_ 的输入 self 与返回值带有相同标注,表达二者的别名关系;add.out 中,这个关系连接 out 与返回值。下划线是命名惯例,标注则把副作用明确交给读取 schema 的程序。
Notice that add.out declares a required out argument, with no default. A public Python signature may offer out=None because the binding groups ordinary and out variants. That convenience belongs to the binding layer; it does not make the native out argument optional.留意 add.out 中的 out 没有默认值,是必需参数。公共 Python 签名可以提供 out=None,因为绑定层会把普通版本与 out 版本组织到一起。这个便利来自绑定层,不代表底层 out 参数变成了可选项。
Aliasing and mutation are separate facts. Tensor(a) can describe sharing without declaring a write, as needed for views. The unannotated inputs of functional add do not forbid callers from passing tensors that already alias each other; they state that this operator does not declare writes to those inputs or return an alias of them. For this add family, we can check the concrete return objects as well as the values.共享存储与修改存储是两件事。Tensor(a) 可以描述共享而不声明写入,例如视图所需的关系。函数式 add 的输入没有标注,不是禁止调用者传入本就互相别名的张量,而是这个操作没有声明写入这些输入,也没有声明返回它们的别名。对这组 add 操作,可以同时检查实际数值与返回对象。
import torch
x = torch.tensor([1., 2.])
b = torch.tensor([10., 20.])
y = torch.add(x, b, alpha=2)
assert x.tolist() == [1., 2.]
assert y.tolist() == [21., 42.]
assert y.data_ptr() != x.data_ptr()
z = x.clone()
inplace_result = z.add_(b, alpha=2)
assert inplace_result is z
out = torch.empty_like(x)
out_result = torch.add(x, b, alpha=2, out=out)
assert out_result is out
assert x.tolist() == [1., 2.]
torch.testing.assert_close(y, z)
torch.testing.assert_close(y, out)
print("functional:", y.tolist(), "x:", x.tolist())
print("in-place returns z:", inplace_result is z)
print("out returns out:", out_result is out)
functional: [21.0, 42.0] x: [1.0, 2.0]
in-place returns z: True
out returns out: True
Here x, b and out are separate dense CPU tensors without gradient recording, and out already has the required shape and dtype. These choices isolate the storage contract. They do not make every use of out allocation-free, or every in-place update legal under autograd. The pointer check establishes different starting addresses for these particular nonempty tensors; it is not a general test for overlapping views.这里的 x、b、out 是互相独立、不记录梯度的稠密 CPU 张量,out 已具备所需形状与 dtype,用来单独观察存储约定。这些条件不能推广成“任何 out 调用都不分配内存”或“任何原地操作都符合 autograd 约束”。指针检查只说明本例非空张量的起始地址不同,不能用作任意视图是否重叠的通用判据。
native_functions.yaml: add.Tensor, add_.Tensor and add.outnative_functions.yaml:add.Tensor、add_.Tensor 与 add.out
How one declaration reaches several interfaces一份声明,怎样抵达多层接口
Writing these contracts independently in Python bindings, C++ APIs and dispatcher registrations would create several places that must agree. ATen instead feeds declarations into generators. This happens while building PyTorch. Installing a wheel gives the reader the compiled result; an ordinary torch.add call does not reopen the YAML and regenerate its bindings.如果 Python 绑定、C++ API 与调度注册各自手写一遍这些契约,就会出现多处需要保持一致的定义。ATen 把声明送入生成器,这一步发生在构建 PyTorch 时。安装 wheel 得到的是编译后的产物;普通的 torch.add 调用不会重新打开 YAML 生成绑定。
Codegen.cmake: invoke torchgen during the buildCodegen.cmake:构建时调用 torchgen
- Declarations声明
native_functions.yamlSchema + generation options.schema 与生成选项。 - Parse and validate解析与校验
torchgen/model.pyFunctionSchema + NativeFunction.接口模型与生成模型。 - Generate sources生成源码
torchgen/gen.pyATen interfaces and registrations.ATen 接口与注册代码。
ATen/ops/add_ops.hRegisterSchema.cpp
RegisterCPU*.cpppython_torch_functions*.cppGenerated C++/CUDA sources are then compiled into the library by the build toolchain. Python bindings and autograd glue have their own generators and supporting inputs, including compatibility signatures and derivative rules. The parser separates the schema from surrounding options. FunctionSchema represents the callable contract; NativeFunction adds information used for generation. From this common model, different generators construct the signatures and wrappers required by their own layer. A keyword-only argument in a Python API need not appear as a keyword-only mechanism in generated C++.生成的 C++/CUDA 源码随后由构建工具链编译进库。Python 绑定与求导连接代码各有生成器,分别读取兼容签名、导数规则等辅助输入。解析器把 schema 与外围选项分开处理。FunctionSchema 表示可调用接口的契约,NativeFunction 补充生成所需的信息。不同生成器读取共同模型,再构造各自层次需要的签名与包装代码。Python API 中的“仅限关键字参数”,不意味着生成的 C++ 也有同样的语法机制。
torchgen/model.py: FunctionSchema and NativeFunction modelstorchgen/model.py:FunctionSchema 与 NativeFunction 模型
torchgen/gen.py: parsing and ATen output generationtorchgen/gen.py:解析与 ATen 产物生成
What the add declaration actually generatesadd 的声明实际交给生成器什么
The following excerpt retains only the fields needed to follow ordinary dense add in this version. The source has additional sparse, nested and other registrations. The first two entries delegate structured generation to add.out; this is a relationship used to generate wrappers, not evidence that Python literally calls torch.add(..., out=...) behind the scenes.下面只摘录这个版本中理解普通稠密 add 所需的字段;原文件还有稀疏、嵌套等注册。前两个条目把结构化生成委托给 add.out。这是生成包装代码时使用的关系,不意味着 Python 内部字面执行了一次 torch.add(..., out=...)。
- func: add.Tensor(Tensor self, Tensor other, *, Scalar alpha=1) -> Tensor
variants: function, method
structured_delegate: add.out
- func: add_.Tensor(Tensor(a!) self, Tensor other, *, Scalar alpha=1) -> Tensor(a!)
variants: method
structured_delegate: add.out
- func: add.out(Tensor self, Tensor other, *, Scalar alpha=1, Tensor(a!) out) -> Tensor(a!)
structured: True
structured_inherits: TensorIteratorBase
ufunc_inner_loop:
Generic: add (AllAndComplex, BFloat16, Half, ComplexHalf)
ScalarOnly: add (Bool)
variants controls the function/method API forms. structured_delegate lets the functional and in-place variants share structured machinery with the out variant, while keeping distinct output behavior: prepare a fresh result, operate on self, or use the provided out. structured_inherits names the base used by the structured implementation; it is not a tensor dtype or a dispatch key.variants 控制函数/方法两类 API 形式。structured_delegate 让函数式与原地版本复用 out 版本的结构化机制,同时保留各自输出行为:准备新结果、操作 self、使用已有 out。structured_inherits 指定结构化实现的基类,不是张量 dtype 或调度键。
For add, ufunc_inner_loop is another generation input. It identifies the scalar operation and supported type groups; the model and ufunc generators use it to supply CPU/CUDA metadata and emit backend code. Therefore, an explicit CPU: or CUDA: line is not required in this particular dispatch block. Other operators may instead name backend implementations directly with dispatch. Read the generation fields together rather than treating one field as a complete runtime table.对 add,ufunc_inner_loop 也是生成输入,指定标量操作及支持的类型组;模型与 ufunc 生成器据此补充 CPU/CUDA 元信息并生成后端代码。因此,这个条目的 dispatch 块不必显式写出 CPU: 或 CUDA:。其他算子可以用 dispatch 直接指定后端实现。应结合生成字段一起读,而不能把单个字段当作完整运行时调度表。
NativeFunction parsing: add generated ufunc backend metadataNativeFunction 解析:补充 ufunc 后端元信息
The arithmetic still has an authored definition. The scalar add body in aten/src/ATen/native/ufunc/add.h computes self + alpha * other. Structured metadata logic in BinaryOps.cpp prepares the iterator and checks constraints such as alpha’s type. Generators connect and specialize these definitions; the name add alone is insufficient to invent broadcasting, type promotion or an algorithm.数值运算仍有明确写出的定义。aten/src/ATen/native/ufunc/add.h 中的标量 add 计算 self + alpha * other;BinaryOps.cpp 中的结构化元信息逻辑准备迭代器,并检查 alpha 类型等约束。生成器把这些定义连接和特化起来,仅凭 add 这个名字无法创造广播、类型提升或计算算法。
add.h: the authored scalar arithmeticadd.h:明确写出的标量运算
BinaryOps.cpp: structured add metadata and implementation hooksBinaryOps.cpp:结构化 add 元信息与实现入口
torchgen/dest/ufunc.py: CPU and CUDA ufunc generationtorchgen/dest/ufunc.py:CPU 与 CUDA 的 ufunc 生成
A contract also has boundaries接口契约的边界
A schema can describe a mutable tensor without explaining a derivative, and it can accept two Tensor arguments whose shapes cannot be broadcast. Different pieces of the system answer these different questions. For this article’s add example, that division is concrete:schema 可以描述一个可修改张量,却不解释它的导数;也可以接收两个 Tensor 参数,而它们的形状并不能广播。系统中的不同部分回答不同问题。落实到本文的 add,可以这样区分:
| Question问题 | Definition定义位置 | Information提供的信息 |
|---|---|---|
| Arguments and effects参数与副作用 | func schema | Names, types, defaults, aliases and writes.名称、类型、默认值、别名与写入。 |
| Output metadata and legality输出元信息与合法性 | Structured meta / TensorIterator | Broadcasting, result dtype and operation-specific checks.广播、结果 dtype 与操作特定的检查。 |
| Numerical values数值结果 | ufunc/add.h + generated backend code | The scalar formula, iteration and backend execution.标量公式、迭代与后端执行。 |
| Backward rule反向求导规则 | tools/autograd/derivatives.yaml | Explicit derivative formulas used by autograd generation.供求导代码生成使用的明确导数公式。 |
The distinction matters to transformations. Reordering an operation that reads x across an add_ that writes x can change the answer. Alias and mutation information is one input to deciding which transformations preserve behavior; it does not alone prove that every reordering or fusion is safe. Similarly, derivatives.yaml supplies formulas rather than asking torchgen to differentiate the arbitrary C++ body of every kernel.这种分工会直接影响变换是否合法:把读取 x 的操作跨过一次写入 x 的 add_,可能改变结果。别名与修改信息帮助判断哪些变换保持行为,但它们本身不能证明任意重排或融合都安全。同样,derivatives.yaml 提供导数公式,并不是让 torchgen 自动对每个内核的任意 C++ 函数体求导。
A composite implementation built from differentiable operators can obtain gradients through those operations; a separately registered custom kernel may need an explicit autograd rule. This is why “the schema exists”, “a CPU implementation exists” and “backward is supported” are three separate checks. For the same reason, the introspection below prints schemas, not guarantees about every permitted shape or available device.由可求导算子组成的组合实现可以通过内部操作获得梯度;单独注册的自定义内核则可能需要明确的求导规则。因此,“存在 schema”“存在 CPU 实现”“支持 backward”需要分别检查。同理,下面的内省程序打印的是接口定义,不是对所有形状或所有设备的支持保证。
import torch
for op in (torch.ops.aten.add.Tensor,
torch.ops.aten.add_.Tensor,
torch.ops.aten.add.out):
print(op._schema)
try:
torch.add(torch.zeros(2), torch.zeros(3))
except RuntimeError:
print("Tensor arguments accepted; broadcasting failed")
_schema is a private diagnostic attribute used here against the pinned version, not a stable application API. The three printed schemas match the earlier block. The final call fails because lengths 2 and 3 cannot broadcast, even though both arguments are tensors._schema 是这里按固定版本使用的私有诊断属性,不是稳定的应用 API。打印出的三行对应前文的 schema。最后一次调用失败,是因为长度 2 与 3 无法广播,尽管二者都属于 Tensor 类型。
derivatives.yaml: authored autograd formulas, including add.Tensorderivatives.yaml:包括 add.Tensor 在内的求导公式
Define a small contract yourself亲手定义一个小型契约
An extension can register a schema without editing PyTorch’s built-in YAML. In a fresh Python process, the following CPU example defines feng_idl::scaled_add, supplies a numerical implementation, and registers a fake implementation for metadata propagation. Keep the Library object alive. The calculation uses separate inputs and returns a new tensor, matching its non-mutating schema.扩展也能注册 schema,无需修改 PyTorch 内置 YAML。在新的 Python 进程中,下面的 CPU 示例定义 feng_idl::scaled_add,提供数值实现,再注册用于传播元信息的 fake 实现。需要保留 Library 对象。运算使用独立输入并返回新张量,与不修改输入的 schema 对应。
import torch
lib = torch.library.Library("feng_idl", "DEF")
lib.define("scaled_add(Tensor x, Tensor y, *, Scalar alpha=1) -> Tensor")
def scaled_add_cpu(x, y, *, alpha=1):
return torch.add(x, y, alpha=alpha)
lib.impl("scaled_add", scaled_add_cpu, "CPU")
@torch.library.register_fake("feng_idl::scaled_add")
def scaled_add_fake(x, y, *, alpha=1):
return torch.add(x, y, alpha=alpha)
x = torch.tensor([1., 2.])
y = torch.tensor([10., 20.])
with torch.no_grad():
result = torch.ops.feng_idl.scaled_add(x, y, alpha=2)
assert result.tolist() == [21., 42.]
assert x.tolist() == [1., 2.]
print(result.tolist())
print(torch.library.opcheck(
torch.ops.feng_idl.scaled_add.default,
(x, y), {"alpha": 2},
))
[21.0, 42.0]
{'test_schema': 'SUCCESS', 'test_autograd_registration': 'SUCCESS',
'test_faketensor': 'SUCCESS', 'test_aot_dispatch_dynamic': 'SUCCESS'}
Although the two bodies have identical source text, their execution contexts differ. The CPU implementation computes actual values. Under fake execution, torch.add propagates output metadata, including broadcasting and dtype promotion, without computing tensor data. The fake body must use operations that themselves support fake execution. It supplies neither CUDA numerical support nor a derivative registration for this custom operator.两段函数体虽然写得相同,执行上下文却不同:CPU 实现计算真实数值;fake 执行下,torch.add 传播广播形状、dtype 提升等输出元信息,不计算张量数据。fake 函数体需要使用自身也支持 fake 执行的操作。它不会因此提供这个自定义算子的 CUDA 数值实现,也不会替它注册导数。
The example deliberately uses inputs without gradients and makes its real call under no_grad. opcheck tests registration consistency on supplied samples; it is not a numerical proof or a gradient check. Here a successful autograd-registration check does not demonstrate a working derivative, because these inputs do not require gradients. To support training, add an appropriate register_autograd rule and test it with gradient-requiring inputs and numerical gradient checks.示例刻意使用不需要梯度的输入,并把实际调用放在 no_grad 中。opcheck 针对提供的样本检查注册一致性,不是数值正确性证明,也不是梯度检验。这里即使 autograd 注册检查成功,也不能说明导数已经可用,因为输入并不需要梯度。若要支持训练,需补充恰当的 register_autograd 规则,再用需要梯度的输入和数值梯度检查验证。
torch.library: define, impl, register_fake, register_autograd and opchecktorch.library:define、impl、register_fake、register_autograd 与 opcheck
Two kinds of code generation, two different inputs两种代码生成,两种输入
Return to the opening call. The schema establishes how to call add and which tensors it may affect; build-time generation connects that declaration to interfaces and registered implementations. When torch.compile later wraps a Python function containing add, Dynamo captures a computation graph. A backend such as Inductor consumes that graph and may fuse compatible operations into new code; guards check the assumptions used to specialize the program. Operator behavior and metadata inform this process. These two stages solve different problems and can both be present in the same program.回到开头那次调用:schema 约定如何调用 add、它可能影响哪些张量,构建期生成把这份声明连接到接口与已注册实现。之后,torch.compile 包装包含 add 的 Python 函数,由 Dynamo 捕获计算图,Inductor 等后端读取图并可能将兼容运算融合成新代码,守卫则检查程序特化所依据的条件。算子行为与元信息参与这一过程。这两个阶段解决不同问题,也可以同时存在于一个程序中。
| Stage阶段 | Input输入 | Output输出 |
|---|---|---|
| torchgen / ATen buildtorchgen / ATen 构建 | Operator declarations and generation rules.算子声明与生成规则。 | Interfaces, wrappers and registrations; for some operators, generated backend code.接口、包装与注册;部分算子还包含生成的后端代码。 |
| torch.compile / Inductortorch.compile / Inductor | A Python function/model, then captured graphs and specialization assumptions.Python 函数/模型,随后是捕获的图与特化条件。 | Executable code for that graph, potentially with fused operations.面向该图的可执行代码,可能融合多个运算。 |
When reading a new operator, follow this order: find its full name and overload, read its arguments and alias annotations, inspect its generation options, then locate its metadata logic, numerical implementation and derivative rule. This keeps a single line of schema connected to the behavior that makes the line true.阅读一个新算子时,可以按这个顺序追踪:找到完整名称与重载,读参数和别名标注,检查生成选项,再定位元信息逻辑、数值实现与导数规则。这样,一行 schema 就能始终连接到兑现它的实际行为。
Python binding generator: signatures and generated wrappersPython 绑定生成器:签名与生成包装
For runtime selection, continue with operator registration and dispatch keys. For graph capture and fusion, return to torch.compile and generated kernels. The contract explains what an operator exposes; those installments explain how a call finds an implementation and how a program may get a different execution plan.要看运行时选择,接着读算子注册与调度键;要看图捕获与融合,回到torch.compile 与生成内核。接口契约解释算子对外提供什么,这两期则解释调用怎样找到实现、程序怎样获得不同的执行计划。