TileFoundry Spec — Parser¶
The parser turns Python source decorated with TileFoundry's IR decorators
into a core_ir.Module. This document covers the user-facing DSL
syntax (§1), the DSL namespace surface (§2), the parser architecture
(§3), the shared machinery both IRs use (§4), and the per-IR parser
bodies (§5 / §6). Validation and rejection rules are in §7.
1. DSL syntax¶
This section is the language reference for TileFoundry DSL source.
Each subsection introduces one construct using a grammar production
followed by a short description and an example. Productions use
the convention lhs ::= rhs; literal terminals are quoted; * is
zero-or-more, ? is optional.
1.1 Decorators¶
@tilefoundry.module(entry="<name>") decorates a class and evaluates to a
core_ir.Module: the decorated name binds to the Module itself
(not the class). It collects the class body's @func / @prim_func results,
child Modules, and plain Python functions, in definition order, into the
result's functions / modules / methods (full contract in
§2.7); entry is an optional argument naming
which collected function is the default step. A member MAY call a sibling
defined above it —
the call lowers to a Call targeting that sibling function; forward
references (a callee declared below the caller) are unresolved (see
§3.3). Functions are reached by name on the result (see
core-ir §1.1).
@tilefoundry.func and @tilefoundry.prim_func decorate functions and
evaluate to the parsed IR directly: @func to a hir.Function
(hir.md §1.1), @prim_func to a
tir.PrimFunction. The decorated name binds to that IR node, not to
the original Python function. Passing a standalone function to compile /
jit lifts it into an implicit single-function Module whose entry is
that function. @func parses with dispatch token "hir", @prim_func
with "tir".
A standalone @func MAY declare its own execution context —
@func(target=..., topologies=(...)). Declaring either makes that function
its own execution domain, so the decorated name binds to the implicit
single-function Module carrying the declaration rather than to the
hir.Function. A plain @func inside a @module class body never does
this: the class declares the domain, and the member stays a hir.Function
so it remains callable by its siblings and specializable through
.specialize.
# example
@tilefoundry.module(entry="f")
class M:
@tilefoundry.func
def g(...): ...
@tilefoundry.func
def f(...):
return g(...) # sibling g is declared above → resolves to g's Function
Specialization decorators. A function specializes its body per input
shape through Function.specialize. The base function is defined with
@tilefoundry.func; each variant is added by decorating a def
with @base.specialize(pattern):
# example
S = DimVar("S", 1, 9) # envelope [1, 9) = 1..8
@tilefoundry.func
def f(x: Tensor[(S,), "f32"]) -> Tensor[(S,), "f32"]:
pass # prototype base
@f.specialize(DimVarRangePat("S", 1, 5))
def small_s(x: Tensor[(S,), "f32"]) -> Tensor[(S,), "f32"]:
return small_impl(x) # variant [1, 5) = 1..4
@f.specialize(DimVarRangePat("S", 5, 9))
def _(x: Tensor[(S,), "f32"]) -> Tensor[(S,), "f32"]:
return large_impl(x) # variant [5, 9), unlabelled
@tilefoundry.funcevaluates to the basehir.Function.func()has nospecializations=parameter; specialization is reachable only through.specialize.- The prototype base body is
pass: it declares the signature and dispatch envelope only and parses toFunction.body is None. The implementations live in the variants (hir.md §1.1). Apassbody is legal only for a function that receives variants; apassbody with no variants, or a real body combined with variants, is rejected. base.specialize(pattern)returns a decorator. It parses the decorateddefinto a varianthir.Function(samenameas the base,specializations=(pattern,)), registers it onbase.variants, and returns the variant.- The decorated identifier is the variant's display label and nothing more.
It MUST NOT become the variant's
name, and MUST NOT take part in equality, hashing, dispatch or TIR identity (hir.md §1.1): the variants of one base share that base's name, and which one runs is decided by the pattern alone.def _states no label and remains legal —_is reusable across variants because the base is the persistent handle — and anything reporting an unlabelled variant names it by its canonical specialization signature instead. A label is for a reader who has to be told which of two implementations ran. patternMUST be a singleDimVarRangePat(see core-ir §3); otherPatternsubclasses are rejected for v0. The referencedDimVarand its(lo, hi)envelope live on the base parameter's shape. Each variant's range MUST fall within that envelope, and the full variant set MUST partition it — disjoint and complete (see hir.md §1.1). Two variants with the same canonical signature are rejected.- A
DimVarshape entry MAY be written inline (DimVar("S", lo, hi)AST node in the shape tuple) or as a named alias (S = DimVar("S", lo, hi)thenTensor[(S,), ...]). Both resolve to the sameDimVarinstance (type-level cache keyed by(name, lo, hi)). A secondDimVarwith the samenamebut conflicting(lo, hi)is a hard parse-time error. - Variants accumulate on the base only during authoring. Once the base
enters a
Module(see core-ir §1) it is sealed and.specializeraises. A variant lives only insidebase.variants; it is never a separateModule.functionsentry.
1.2 DSL namespace¶
import-form ::= 'from tilefoundry.dsl.tf import *'
| 'from tilefoundry.dsl.T import *'
| 'from tilefoundry.dsl import' ('tf' | 'T')
namespace-callee ::= ('tf' | 'T') '.' op-name
tilefoundry.dsl.tf (HIR) and tilefoundry.dsl.T (TIR) are the only
DSL-facing entries to the Op catalogue. The mechanism that backs
this surface is described in §2.
1.3 Op call¶
op-call ::= callee '(' arg-list ')'
callee ::= op-name ; bare-name path
| namespace-callee ; namespace-attribute path
op-name ::= identifier
| identifier '_' ; trailing-underscore = effect form (TIR only)
A bare-name callee MUST resolve to an _op_schema-bearing surface
value (an Op class for real-Op schemas, or an alias builder
function carrying _op_schema for surface-alias schemas) through
the function's closure (typically established by an import-form) or,
failing that, through dispatch.resolve_callable
(§4.2/§4.3) — the
path a trailing-underscore op-name always takes, since nothing binds
a literal foo_ name in the closure. The namespace-callee form
resolves on the namespace package directly (§2). When op-name is
registered with multiple schemas ("overloads"), the parser uses
first-match dispatch (§4.3). The trailing-underscore selector is
gated to the TIR token; using it in HIR is a verify error.
1.4 Tensor[...] and ConstTensor[...] annotations¶
Tensor and ConstTensor are parser-owned DSL authoring annotation sugar,
imported from tilefoundry.dsl (from tilefoundry.dsl import Tensor). It
is distinct from the IR-level tilefoundry.ir.types.TensorType (the
runtime carrier on Expr.type); the parser resolves a Tensor[...]
annotation into a TensorType at parameter / return-binding time.
tensor-annot ::= ('Tensor' | 'ConstTensor') '[' shape ',' dtype (',' layout)? (',' storage)? ']'
shape ::= '(' (dim (',' dim)*)? ')' ; '()' is rank-0
dim ::= integer-literal | dim-Expr ; dim-Expr per types §4
dtype ::= '"f32"' | '"f16"' | '"bf16"' | … ; see types §3
layout ::= layout-sugar ; see §1.5
| 'ShardLayout(' … ')' ; verbose, see shard §7
storage ::= '"host"' | '"gmem"' | '"smem"' | '"rmem"' | '"tmem"'
Tensor[...] and ConstTensor[...] resolve to the same ordinary TensorType;
the latter sets Var.is_const=True on a function parameter. is_const marks
external residency semantics only and does not embed a payload. Tensor[...] is
the carrier of optional layout sugar; it does not
own the sugar (which lives at §1.5). A rank-0 (scalar) tensor is
written Tensor[(), "bf16"]; the form Tensor["bf16"] (without
shape) is rejected.
Tensor[(4096, 2048), "bf16"]
Tensor[(4096, 2048), "bf16", (4096 @ gpu.cta, 2048)]
Tensor[(4096, 2048), "bf16", (4096 @ gpu.cta, 2048), "smem"]
Tensor[(), "bf16"] # scalar
- constraints:
- The dtype slot MUST use a canonical quoted
DType.namefrom types §3. - The parser MUST normalize that string to the corresponding process-lifetime
descriptor before constructing
TensorType. - An unknown dtype string MUST be rejected; it MUST NOT fall back to another descriptor.
1.5 Layout sugar¶
layout-sugar ::= axis-tuple ; implicit strides, no value-state
| '(' axis-tuple ',' stride-tuple ')' ; explicit strides, no value-state
| '(' axis-tuple ',' value-state ')' ; implicit strides + value-state
| '(' axis-tuple ',' stride-tuple ',' value-state ')' ; explicit strides + value-state
axis-tuple ::= '(' axis-spec (',' axis-spec)* ','? ')'
axis-spec ::= axis-extent ; a layout dim, not split (axis placement only)
| static-extent '@' mesh-axis ; Split(axis_index) on mesh-axis
| static-extent '@' '(' mesh-axis (',' mesh-axis)* ')' ; sequential decomposition
axis-extent ::= static-extent | dim-ref ; a bare axis may be dynamic
static-extent ::= integer-literal | static-dim-ref ; split axes & mesh dims: a static int
dim-ref ::= identifier ; closure-resolved DimVar (or int); bare axes only
static-dim-ref ::= identifier ; closure/global name bound to a static int (bool rejected)
stride-tuple ::= '(' integer-literal (',' integer-literal)* ','? ')'
value-state ::= '{' partial-spec (',' partial-spec)* ','? '}' ; a set; only the last outer item
partial-spec ::= mesh-axis '@' 'P(' '"' reduction '"' ')' ; Partial(reduction) on mesh-axis
mesh-axis ::= identifier '.' identifier ; e.g. gpu.cta
reduction ::= 'sum' | 'max' | 'min' | …
A bare axis extent (axis-extent) may be a static integer-literal
or a closure-resolved dim-ref — an identifier bound to a DimVar
(or an int) in the function's closure, e.g. a dynamic seq_len. A bare
axis is Broadcast (carries no mesh binding). A split extent
(static-extent on the left of @) MUST resolve to a static int: it
participates in mesh-extent canonicalisation (factorisation), which a
dynamic extent cannot. A dynamic split extent is rejected.
Closure/global int resolution applies to mesh-shape dims too; a bool or
dynamic value in a static-extent is rejected with a must be a static int
diagnostic.
A dynamic bare axis is admissible only where Reshard materialises strides
in the shared-engine form — a same-storage reshard off a plain (non
per-instance) source, or a low→high storage move. There the deferred stride
materialisation (§Stride materialization)
keeps static inner strides as plain ints and only the strides of axes
above the dynamic axis become symbolic dim-exprs. In a per-instance
materialisation (a high→low storage move, or a per-instance source), the
per-shard local layout shape must be a static int — a register / shared buffer
cannot be sized by a non-split dynamic axis — so a dynamic non-split bare
axis is rejected with a deliberate error (only a launch-provided CTA Split,
whose per-shard extent is a static 1, may consume a dynamic axis).
The axis-tuple carries only axis placement (Split inlined as
size @ mesh.axis; a bare size is a non-split layout dim). The optional
{...} value-state set carries the mesh-axis Partial states
(mesh.axis @ P("reduction")). It is a Python set literal recognized at
the AST level (its element order carries no meaning) and MUST be the last
item of the outer tuple; it is never mixed into the axis-tuple. A mesh
axis named in no Split and no Partial is Broadcast (the default) —
Broadcast is never written. There is no _ @ B(...) / _ @ P(...) form.
Outer-tuple discrimination: a bare axis-tuple is implicit-strides with no
value-state; an outer length-2 tuple whose second item is a stride-tuple
is explicit strides; length-2 whose second item is a value-state set is
implicit-strides + value-state; length-3 (axis-tuple, stride-tuple,
value-state) is explicit strides + value-state.
dim @ (m.a, m.b) expands to one Split axis per mesh axis
(each with extent = mesh extent), followed by a bare remainder axis
of size dim / ∏(mesh_extents). The remainder axis is always
appended last; the mesh-axis order in the tuple determines the tensor
axis order.
Canonicalization (single-mesh-axis form). Surface sugar
N @ m.a where N > mesh_extent(a) MUST be expanded at parse time
into the factorised pair (mesh_extent(a) @ m.a, N // mesh_extent(a))
before the ShardLayout is constructed. The first element becomes
a Split axis with local_shape = 1; the second becomes a bare
residual axis (non-Split). N // mesh_extent(a) MUST divide N
exactly; otherwise the sugar is rejected. The factorisation is
opaque to the user: input N @ m.a and input
(mesh_extent(a) @ m.a, N // mesh_extent(a)) produce the same IR.
Stride materialization (parser surface)¶
The first sugar form
('(' axis-spec ... ')') emits Layout(shape=canonical,
strides=None) — the layout strides are deferred to Reshard
typeinfer, which fills them in based on the storage-level direction
(see hir.md §1.3). The
verbose form ('(' axis-tuple ',' stride-tuple ')') emits a
concrete strides tuple; typeinfer respects it verbatim. The
parser does NOT inspect storage or do any physical-materialization
logic — that responsibility lives entirely in Reshard typeinfer.
Spec: shard.md §7.1.1, hir.md §1.3.
Layout sugar is accepted anywhere the expected surface value is
a ShardLayout. Dispatch is annotation-driven (see §4.4): the
parser consults the position's expected ParamDef.annotation (or
the Tensor[...] layout slot) to decide whether to invoke the
sugar parser. Omitted mesh axes default to Broadcast. Sugar
forms that would lose mesh / layout information fall through to
the verbose ShardLayout(...) constructor.
# Tensor[...] annotation slot
Tensor[(4096, 2048), "bf16", (4096 @ gpu.cta, 2048)]
# value-state set: a Partial on a mesh axis (implicit strides)
Tensor[(4, 64), "f32", ((4 @ trd.l, 64), {trd.t @ P("sum")}), "smem"]
# Op attribute slot whose ParamDef.annotation is ShardLayout
reshard(x, layout=((2048 @ gpu.cta, 64), {gpu.warp @ P("sum")}))
1.6 with Mesh(...) as m¶
The with Mesh(...) as m grammar is shared by both dialects; the
binding name m is visible only inside suite. The two dialects differ
in what it lowers to:
- HIR treats it as an active mesh context — a parser-lexical
alias for the constructed
Mesh, so layout sugar (§1.5) may bind axes with… @ m.axisand tensors authored under it reuse the oneMesh. It is not a tensor-binding scope and emits no IR node. Ordinary values assigned insidesuitefollow normal function-body visibility (not confined);returninsidesuitereturns from the enclosing@func(no mesh-region result), and a@funcMUST NOT be defined insidesuite. A tensor's mesh/layout lives on itsTensorType.layout, not on the block it is written in;reshardis the explicit boundary, and op typeinfer (hir §1.3) decides whether values combine. - TIR lowers it to an explicit
MeshScopeStmt (§6) carrying theMeshand the bindingVar.
1.7 for i in tile(...) / for i in range(...) (HIR-only)¶
for-loop ::= 'for' identifier 'in' ('tile' | 'range') '(' loop-args ')' ':' suite
loop-args ::= extent-Expr # tile: extent / range: stop
| extent-Expr ',' step-Expr # tile(extent, step)
| start-Expr ',' stop-Expr # range(start, stop)
| start-Expr ',' stop-Expr ',' step-Expr # range(start, stop, step)
tile(...) and range(...) share one loop domain (start, extent,
step) and lower to the same GridRegionExpr (hir §1.2) —
range is not a separate construct and is not statically unrolled. The
only difference is the loop-variable binding:
range(...)bindsito a scalar induction var (i: i64); use it asx[i]or write the window manually (i : i + step). Args follow Pythonrange:range(stop)(start0, step1),range(start, stop)(step1),range(start, stop, step).tile(extent)also binds a scalari: i64(start0, step1).tile(extent, step)bindsito a parser-sideRangeSlice(start = iv * step,stop = start + step) sox[:, i]lifts to aSliceover the current window.RangeSliceis parser-only and does not reach IR;startis0.
start-Expr / extent-Expr (the stop endpoint, not a length — the
domain is half-open [start, extent)) / step-Expr MAY be any ShapeDim
(types §4), including a dim expression such as C // N; the
value is carried verbatim into GridRegionExpr.start / .extent / .step
and resolved at evaluate time (hir §1.2).
A tensor subscript x[slice0, …] inside a loop body lifts to a
hir.tensor.Slice Op call. An ast.Assign whose single Name target is
bound in outer scope is a loop-carried rebinding (see §5 for the carry-out
lift). A nested for ... in tile/range(...) is allowed and lifts to a
nested GridRegionExpr; the carry scan recurses into nested loops, so an
outer-scope name rebound only inside a nested loop is still carried across the
outer loop (and the nested loop carries it too).
1.8 Hard schedule constraints¶
constraint-annotation ::= 'where' '(' constraint-field (',' constraint-field)* ')'
constraint-field ::= 'layout' '=' layout-constraint
| 'mesh' '=' mesh-expression
| 'storage' '=' storage-expression
layout-constraint ::= layout-axis-tuple
| '(' layout-axis-tuple ',' binding-set ')'
binding-set ::= '{' binding (',' binding)* '}'
binding ::= topology '@' 'B()'
| topology '@' 'P(' string-literal ')'
where(...) is keyword-only and non-empty. A layout axis is _, an integer
or symbolic extent, or extent @ topology. The split form binds an existing
Split attribute to the physical layout position. _ is a private
constraint wildcard and is distinct from Layout's None launch-provided
extent. B() and P(...) reuse the existing Broadcast and Partial
ShardAttr values; a topology may be bound at most once in one layout
constraint. mesh resolves to a Mesh, and storage resolves through the
current storage-kind registry. CTA capability checks do not occur in this
parser surface.
1.9 Compile-time values¶
A compile-time value is a Python number the parser can reach without building
any IR: a numeric literal, a name captured from the enclosing scope, an attribute
of a captured object, and arithmetic over those (+ - * / // % **, unary -).
compile-time-expr ::= number-literal | identifier | compile-time-expr '.' identifier
| compile-time-expr binary-arith-op compile-time-expr
| '-' compile-time-expr
- A compile-time value MAY appear anywhere a number is required: an attribute
argument, a shape or extent, a subscript index, or an op input, where it
becomes a rank-0 unmaterialized
Constant. - A body-local assignment whose right-hand side is a compile-time value binds
the value, not an
Expr; the name is then usable in every position above. A tuple target binds one number per name when every element is a compile-time number. - An op input that is a Python float carries no precision of its own: the parser MUST give it the float dtype of the operands it is used with, and MUST reject the call when those name more than one float dtype. A Python integer keeps its own dtype, so combining one with a float tensor MUST still be rejected (hir §1.3).
- Evaluation MUST NOT call anything reached from a speculative position: a value that is not statically reachable is parsed as IR instead.
A compile-time list holds Expr elements and never reaches the IR:
compile-time-list ::= '[' expr (',' expr)* ']'
| '[' expr 'for' identifier 'in' compile-time-sequence ']'
- A comprehension MUST declare exactly one
forclause, noifguard, and a plain-name target. Its sequence MUST be either the builtinrange(...)over compile-time integers or a compile-time tuple / list; any other call MUST be rejected rather than evaluated. - Subscripting the bound name with a compile-time integer selects one element. Python's negative indexing applies.
- The comprehension form is the fixed-length unrolled spelling; it does not
affect
for(§1.7), which always builds aGridRegionExpr.
A tensor subscript resolves per axis: an ast.Slice keeps the axis, a
compile-time integer drops it (as in torch), and a negative integer counts
back from the axis extent, which MUST therefore be static.
2. DSL namespace surface¶
2.1 Model¶
The two namespaces are real Python modules; resolution is module
__getattr__ over the OpSchema registry. There is no
DslNamespace class.
# tilefoundry/ir/core/op_registry.py
def _register_schema(schema: OpSchema) -> None: ...
def get_schemas(dialect: str, name: str) -> list[OpSchema]: ...
def iter_schema_names(dialect: str) -> Iterable[str]: ...
# tilefoundry/ir/core/op_schema.py — a frozen dataclass
class OpSchema:
dialect: str # "tf" or "T"
name: str
op_cls: type[Op]
signature: tuple[ParamDef, ...]
# tilefoundry/dsl/tf/__init__.py (TIR symmetric in tilefoundry/dsl/T)
def __getattr__(name: str) -> type[Op] | Callable: ...
def __dir__() -> list[str]: ...
2.2 Class diagram¶
classDiagram
class op_registry {
get_schemas(dialect, name) list~OpSchema~
iter_schema_names(dialect)
_register_schema(schema)
}
class OpSchema { dialect; name; op_cls; signature }
class tilefoundry_dsl_tf { __getattr__(name); __dir__() }
class tilefoundry_dsl_T { __getattr__(name); __dir__() }
class Op { <<abstract>> }
op_registry "1" --> "*" OpSchema : indexes
tilefoundry_dsl_tf ..> op_registry : queries dialect="tf"
tilefoundry_dsl_T ..> op_registry : queries dialect="T"
OpSchema --> Op : op_class (None for alias schemas)
2.3 Resolution algorithm¶
The CPython attribute-access path runs module.__getattr__(name)
for any name not found on the module's own namespace. Both
namespaces implement it identically:
def __getattr__(name: str) -> type[Op] | Callable: ... # resolve a dialect name to its Op class / alias builder
def __dir__() -> list[str]: ... # list the dialect's registered schema names
__getattr__ looks up op_registry.get_schemas(_DIALECT, name) and raises
AttributeError on a miss; for a single real-Op schema it returns the Op
class, for a surface-alias schema (op_class is None) the alias builder fn, and
for more than one schema an overload resolver.
Both forms (Op class for real-Op schemas, alias builder fn for
alias schemas) carry an _op_schema attribute, so the parser's
bare-name resolver looks the schema up with a single
getattr(val, "_op_schema", None) regardless of which form was
bound.
from tilefoundry.dsl.tf import * invokes __dir__ and then
__getattr__ for every returned name; from tilefoundry.dsl import tf
binds the module object itself, leaving each later tf.<name>
attribute access to __getattr__.
2.4 .pyi stub regeneration¶
The dynamic __getattr__ surface is invisible to static analysers
and editors. To restore IDE completion / type inference,
tilefoundry.dsl._stub_gen emits per-namespace .pyi stubs derived
from the OpSchema registry:
tilefoundry/dsl/tf/__init__.pyi # generated, gitignored
tilefoundry/dsl/T/__init__.pyi # generated, gitignored
The CLI is python -m tilefoundry.dsl regen. The generator walks
every OpSchema registered for the dialect and emits one
def <name>(<param>: <type>[, ...]) -> Expr: ... signature per
schema. Multi-schema overloads emit @typing.overload stubs in
registration order, followed by a final non-overload signature
that matches the runtime resolver.
Conventions:
kind="input"ParamDefs render asExprregardless of their declaredannotation(operands are always Exprs at the DSL surface).kind="attribute"ParamDefs render theirannotationverbatim (int/str/ShardLayout/ …). Referenced types are auto-imported in the generated header so the.pyiis self-contained.- A
DTypeattribute renders asLiteral[<canonical names>] | DType, with theLiteralmembers derived from the closed descriptor set in types §3. The string form is the canonical DSL authoring path and the parser normalizes it to the corresponding descriptor at the call boundary. A descriptor value remains accepted as the IR-canonical attribute value in direct Python expressions. - Any other string-valued enum attribute, such as
ReduceKind, renders asLiteral[<member strings>] | <EnumType>. ItsLiteralmembers derive from the enum, and the parser normalizes a string to the corresponding enum member at the call boundary.
Stubs are not part of the runtime resolution path; the parser still
goes through §2.3. They exist solely so editors can show typed
completions for tf.<name>(...).
2.5 Invariants¶
- Dialect isolation.
tilefoundry.dsl.tfMUST surface only schemas withdialect="tf";tilefoundry.dsl.TMUST surface only schemas withdialect="T". §4.6's strict per-dialect resolution depends on this. - Late-registration visibility. An Op registered after the
namespace module is first imported is visible on the next
__getattr__call; the namespace MUST NOT cache resolutions in a way that would hide it. - Implementation independence. The DSL surface MUST NOT depend
on the
tilefoundry.ir.<dialect>.<category>directory layout. DSL source addresses Ops only through(dialect, name). - Single-schema identity. For a single-schema name
n,getattr(tilefoundry.dsl.tf, n) is get_schemas("tf", n)[0].op_cls. No wrapper class is interposed.
2.6 Platform sub-namespaces¶
tilefoundry.dsl.T exposes platform-specific instruction and atom
surfaces under a fixed set of platform sub-namespaces (e.g.
T.cuda). dsl.T.__getattr__ resolves a platform name before the
OpSchema registry lookup (§2.3): a name in the platform set returns the
platform namespace object; every other name falls through to the
registry.
- A platform sub-namespace is not an
Opand carries no_op_schema. It surfaces platform-specific descriptors (instruction specs, atoms) only, never catalogue Ops. §2.5's dialect-isolation invariant is unaffected — the platform set is disjoint from registered Op names. - The platform set is fixed; a name outside it MUST resolve as an
ordinary
dialect="T"Op name, preserving late-registration visibility (§2.5). T.cuda.mmais the CUDA MMA surface:T.cuda.mma.<NAME>is anMmaOpSpecandT.cuda.mma.atom(op=...)anMmaAtom(tir §2.3). The folder name (cuda) matchescodegen/cuda/and the runtime tree.
In a @prim_func body a chain rooted at a platform sub-namespace is a
compile-time static binding, not a runtime value: op = T.cuda.mma.<NAME>
and atom = T.cuda.mma.atom(op=op) bind Python descriptor objects in
the parser environment and emit no LetStmt. A subsequent atom.A/B/C
attribute access resolves statically against the bound descriptor.
The .pyi stub generator (§2.4) emits the platform sub-namespace
surface so editors complete T.cuda.mma.<NAME> and .atom(...).
2.7 @module authoring surface¶
@module(entry="<name>", target=...) collects a class body into a Module
(core-ir §1). The decorated name binds to the
resulting Module. target declares the hardware the domain runs on; the
ordered Topology hierarchy is declared by a topologies assignment in the
class body instead (see below).
A file of @module classes is an ordinary importable Python module: every name
its class bodies read — the shape configuration above all — MUST resolve within
that file, and the file MUST NOT require execution with a namespace injected into
its globals. The Modules it defines are module-level values, so import reaches
them, a linter sees them, and a CLI selector addresses them by name. A file that
has to be executed with a namespace injected is reachable by none of those.
A @module class body is evaluated once where it stands, so a body at file scope
states one shape. A model asked about more than one structural configuration —
shapes differing in a submodule count or a per-layer tuple, not only in a tensor
axis — MAY instead place its class bodies in a function of that same file that
takes the configuration as a parameter, and publish that function; each call
states the same source at the shape its caller names. This is not injection: the
file is still an ordinary import and the configurations are values its own
package publishes. @func MUST resolve the names in its signature and body
against the locals of every scope enclosing it, innermost first, so a parameter
of that function is as visible to a nested class body as a module-level import
is.
- Every non-dunder class member MUST be one of four kinds: the
topologiesdeclaration; an@func/@prim_funcresult (anhir.Function/tir.PrimFunction); a childModule— or a tuple/list of them, how a factory attaches N identical instances under one attribute (each already named by the factory, e.g.renamed(f"layer{i}")); or a plain Python function (an orchestration method, e.g.forward/init_caches). Any other member — a stray attribute, an undecorated method that is neither a DSL function nor a plain orchestration function, … — MUST be rejected. A specialization variant (a@base.specializedef) and a per-weight converter (a@base.converter(name)def, runtime §1.1.2) are not standalone members — both live on their base function and are skipped when collecting. - A nested
classstatement is a legal member when it is itself decorated with@module(...): by the time the outer class body finishes running, the inner decorator has already replaced the name with aModuleinstance, so it is collected as an ordinary childModule. An undecorated nested class is rejected (it is none of the three kinds). @func/@prim_funcresults are collected in definition order intoModule.functions; the class body MUST contain at least one. A childModuleis collected intoModule.modules, renamed to the attribute it is attached under — torch / HuggingFace checkpoint-naming semantics: assigning a child toself.self_attnnames itself_attnin the tree, independent of the child's own class name (core-ir §1). A plain Python function is collected intoModule.methodsby its own name. A duplicate function name, or duplicate child module name, across the class body MUST be rejected.- A class body MUST declare at least one function, child
Module, or plain method; only an empty body MUST be rejected. A methods-only Module is therefore valid — it composes the children it is given. entryis optional. Supplied, it MUST name exactly one collected function and an unknown name MUST be rejected. Omitted, the Module has no default step.- A method's name is free, but
forwardis the one a bare<module>(...)delegates to; any other name is reached only by naming it. A class-body__call__MUST be rejected: Python resolves a dunder on the type, so one attached to the built Module instance would never run. - A member MAY call a sibling defined above it (the call resolves to the sibling function / launches a sibling device kernel); a forward reference to a sibling defined below stays unresolved and MUST fail.
- A
topologies = (Topology(...), ...)assignment declares the domain's complete ordered hierarchy and MUST precede the functions that name one of its levels. Omitting it inherits the owning class's hierarchy;()declares an explicitly topology-free domain. The value MUST be a tuple ofTopology; a value that is not, or an entry the parser cannot resolve statically, MUST be rejected rather than read as an empty hierarchy. A deferred extent (Topology("cta", None), target §4) is a declaration in its own right and MUST survive parsing, not be dropped as unresolvable. - The printer emits this surface: shared meshes at module level (before the
class) so the class body stays declaration-and-function-only, then
@module(entry="<entry>")— or a bare@modulefor a Module that declares neither an entry nor a target — then thetopologiesdeclaration when the Module makes one, then the functions and nested Modules.
Design rationale¶
entry is a function-name forward reference rather than a function object
because a class decorator's arguments are evaluated before the class body runs,
so the entry function does not yet exist when @module(entry=...) is called.
topologies is a class-body assignment rather than a decorator argument for
the mirror-image reason. A function body MAY name a level of its domain
(with Mesh(topology="cta", ...)), and that name has to resolve while the
body is parsed — which happens as the class body runs, before the decorator is
applied. A class-body assignment is already bound at that point, so the
declaration is readable exactly when it is needed; a decorator argument is
not. target stays a decorator argument because nothing consumes it until
after the Module exists.
3. Parser architecture¶
3.1 Model¶
The implementation lives under tilefoundry/parser/. The current
function-level entry points are independent calls; there is no
ModuleContext / FunctionDecl data class.
# tilefoundry/parser/hir_parser.py
def parse_func (fn, *, topologies=()) -> hir.Function: ...
def parse_func_source (src: str) -> core_ir.Module | hir.Function: ...
def parse_module_source(src: str) -> core_ir.Module: ...
def parse_script (src: str) -> core_ir.Module | hir.Function: ...
class _HirBodyVisitor(BaseExprVisitor): ...
# tilefoundry/parser/tir_parser.py
def parse_prim_func (fn) -> tir.PrimFunction: ...
class _TirBodyVisitor(BaseExprVisitor): ...
# tilefoundry/parser/base.py
def extract_ast(fn) -> ast.FunctionDef: ...
class BaseExprVisitor: ...
# tilefoundry/parser/symtab.py
class LexicalEnv:
def push_frame(self) -> None: ...
def pop_frame (self) -> None: ...
def define (self, name, value) -> None: ...
def lookup (self, name) -> object: ...
def innermost_mesh(self) -> Mesh | None: ...
# tilefoundry/parser/dispatch.py
def resolve_op (name) -> type | None: ...
def resolve_stmt (name) -> type | None: ...
def resolve_schema(name, dialect: str = "tf") -> OpSchema | None: ...
def resolve_callable(name, token: Literal["hir", "tir"]) -> tuple[str, type]: ...
Each function-level parser collects a closure dict from the live
Python function (_collect_closure(fn) -> dict[str, Any]), reads
the AST via extract_ast(fn), and walks the body with the
dialect's BaseExprVisitor subclass.
3.2 Class diagram¶
classDiagram
class parse_func
class parse_prim_func
class parse_module_source
class BaseExprVisitor { <<abstract>> }
class _HirBodyVisitor
class _TirBodyVisitor
class LexicalEnv
class dispatch { resolve_op; resolve_stmt; resolve_schema; resolve_callable }
parse_func ..> _HirBodyVisitor : drives
parse_module_source ..> parse_func : per @func (HIR only)
parse_prim_func ..> _TirBodyVisitor : drives
BaseExprVisitor <|-- _HirBodyVisitor
BaseExprVisitor <|-- _TirBodyVisitor
_HirBodyVisitor o-- LexicalEnv
_TirBodyVisitor o-- LexicalEnv
_HirBodyVisitor ..> dispatch : callee lookup
_TirBodyVisitor ..> dispatch : callee lookup
3.3 Description¶
parse_func / parse_prim_func consume a live Python function
(fn); topologies supplies the parse-time namespace a body may name,
and is not retained on the resulting Function. parse_func_source /
parse_module_source / parse_script accept Python source text.
Both paths return a core_ir.Module for module-level input, but they do not
cover the same member kinds. The runtime @tilefoundry.module decorator builds
the core_ir.Module from the class's already-parsed @func / @prim_func
methods. The source-text parse_module_source / parse_script build it from
the class body, including every @func it owns and each nested @module
class. The source path is HIR-only: it MUST reject a @prim_func, whether
as a class member or as the bare top-level function, and the error MUST name
that boundary rather than report only the absence of an @func. Reading TIR
back from source text is not part of the round trip, because the Python printer
emits no mixed HIR/TIR module
(inspection §2.2). A source that declares
only a bare @func returns that hir.Function, or the implicit
single-function core_ir.Module when the decorator declares execution
context — mirroring what executing the same source would bind.
The closure dict supplies same-module callee lookup. Names defined
in the user's Python module (other @func / @prim_func
functions, mesh / topology objects, Op classes imported from
tilefoundry.dsl.tf / T) are visible through the closure. The
closure also includes the @func / @prim_func bindings present in
the definition frame when the decorator runs — for a
@tilefoundry.module class body, that is the sibling methods declared
above the one being parsed. Each such binding is the sibling's
hir.Function / tir.PrimFunction IR node (the decorator evaluates to
the IR directly, §1.1), so a sibling callee resolves to that Function
and becomes the Call target directly. This is what makes
callee-before-caller sibling calls work; a forward reference is simply
absent from the closure and fails as an unresolved callee. The merge is
additive: it never shadows the function's own globals / freevars.
LexicalEnv is a frame stack used by both body visitors for
parser-time bindings (Mesh axes, RangeSlice from tile, SSA
aliasing). Frame push / pop matches the Python-source scope
(with Mesh(...), for i in tile(...)).
dispatch.resolve_callable(name, token) performs strict
per-dialect Op resolution against op_registry; both body visitors'
ast.Name callee resolution (BaseExprVisitor._resolve_call_target,
shared by HIR and TIR) delegates to it once the closure path misses, and
the TIR top-level-statement dispatch (_call_as_top_level_stmt) delegates
to it for a bare-name effect Stmt / intrinsic before falling back to
call_to_op_call. It never resolves an arbitrary Python name — only a
name already cataloged as an Op / Stmt / intrinsic under the body's own
dialect — so an undefined name still raises unknown Op name.
There is no parser-side intermediate IR; function bodies translate
directly into core_ir nodes plus dialect-specific subclasses.
4. Shared parsing machinery¶
4.1 Lexical environment¶
Both parsers use the same lexical-env stack. define(name, expr_node)
binds a Python name to an Expr object. Subsequent uses of that name
reuse the same Expr, which is how HIR's SSA-as-DAG sharing falls
out for free.
4.2 Closure-then-registry callee resolution¶
Bare-name callees resolve through the lexical env + the function's
closure first — the common case, covering every name reached via a
star-import or an explicit tf.<name> / T.<name> binding. When that
misses, resolution falls through to dispatch.resolve_callable(name,
token) (§4.6): dialect-strict dispatch against the Op / Stmt / intrinsic
catalogue, not an arbitrary-name lookup, so a name that is neither bound
in the closure nor cataloged under the body's own dialect still raises
unknown Op name.
The closure binding for a name from tilefoundry.dsl.tf /
tilefoundry.dsl.T is whatever its module __getattr__ returns:
- a real-Op class for single-schema names whose schema has an
op_class(e.g.tf.matmul→MatMul); - the alias's builder function for surface-alias schemas (e.g.
tf.add→_add_alias); the function carries_op_schemaso the parser still recovers the schema by attribute lookup.
Both forms expose _op_schema, so the parser's
_resolve_call_target returns an OpSchema uniformly. Namespace-
attribute callees (tf.add / T.copy) skip the closure binding and
go directly through dispatch.resolve_schema(name, dialect), which
honours alias prepend order — an alias schema (if any) wins over a
legacy real-Op schema sharing the same name.
4.3 OpSchema and overload resolution¶
A registered Op has one or more OpSchema entries indexed by
(dialect, name). Each schema lists the Op's ParamDef descriptors
(see core-ir §2.3). When the parser sees a callee:
- Look up the schema list via
op_registry.get_schemas(dialect, name). - Filter by arity.
ParamDef.is_required(i.e.default is MISSING) setsn_min;optionaldoes NOT lowern_min. - For the surviving candidates, walk each input ParamDef and run
pattern.match(arg_type).pattern is Noneaccepts any. - Return the first schema whose every input pattern matches. Registration order is the tiebreaker; there is no "best match" search.
4.4 Annotation-driven sugar dispatch¶
At each attribute slot of a call, the parser consults the matched
schema's ParamDef.annotation:
annotation=ShardLayout→parse_shard_layout_sugarannotation=Layout→parse_layout_sugar
Sugar dispatch is annotation-driven, not name-driven: an attribute
called shape will not be parsed as layout sugar unless its
ParamDef declares a layout annotation. Ops without a registered
schema fall through to a small legacy heuristic
(attr_name == "layout" ⇒ ShardLayout sugar) until they migrate.
Attribute string normalization is also annotation-driven:
annotation=DTyperesolves a canonical string by descriptornameand rejects any other string with aVerifyError.- A string-valued Enum annotation resolves by Enum value and rejects any other
string with a
VerifyError.
4.5 Tensor[...] annotation surface¶
Tensor[shape, dtype, layout, storage] is recognised by
try_parse_sugar_tensor_type, which is the entry point shared by
both @func and @prim_func parsers when reading parameter and
return-type annotations. The shape and dtype slots are required;
the layout and storage slots are optional. The same routine reads
the layout sugar described in §1.4.
The dtype slot MUST resolve to the canonical descriptor named by its string.
Unknown names MUST be rejected and MUST NOT silently select DType.f32.
4.6 Per-dialect strict resolution¶
parser.dispatch.resolve_callable(name, token) does NOT fall back
across dialects. An HIR-only Op (e.g. rope) raises unknown TIR
callable in a TIR body, and a TIR-only Op (e.g. copy) raises
unknown HIR callable in an HIR body. The trailing-underscore
selector is gated to the TIR token only.
5. HIR parser¶
The HIR parser walks an @func body. The body is a sequence of
Python statements that the parser folds into a single Expr tree.
| Python | HIR action |
|---|---|
x = expr |
define(x, expr_node); no IR node. Subsequent x reuses the same Expr (SSA-as-DAG). |
x + y |
Call(Binary(kind=ADD), (x, y)); the parser maps Python AST BinOp / Compare / BoolOp directly to a Binary instance with the matching BinaryKind. UnaryOp USub / Not maps similarly to Unary(kind=NEG) / Unary(kind=NOT). AST @ (matmul) routes to MatMul (a real Op, not kinded). |
foo(a, b) |
Call(target_op, args) where target_op is constructed by the resolved schema's builder (§4.2 / §4.3). For surface aliases (e.g. add / cmp_eq / neg), the alias's builder returns the kinded target Op (Binary(kind=...) / Unary(kind=...)); for real Ops, the default builder is the Op class itself. |
for i in tile(...) |
GridRegionExpr (see §1.7 and below). |
with Mesh(...) as m |
Push m onto the parser-lexical stack; pop on exit. No IR node. |
return expr |
Sets Function.body. A return without a value is rejected. |
return (a, b) / return a, b |
A literal tuple return (both spellings are the same AST) folds to a core Tuple body (core-ir §2.2); Function.return_type is the TupleType of the element types. Callers destructure via the existing tuple-unpack rule (o, s = f(...)). |
pass |
Accepted only as the entire body: sets Function.body = None, declaring a dispatch prototype whose implementations are registered via .specialize |
(hir §1.1). A pass mixed with any other statement is rejected. |
for / if / while over arbitrary ranges, conditionals, and other
Stmt forms are TIR-only. They are rejected by the HIR parser.
A pass body yields Function.body is None, declaring a dispatch
prototype that awaits variants. Immediately after @func def f: pass
and before any @f.specialize(...), the base is transiently
body is None, variants == () — a valid unsealed authoring state. The
sealed (verified) invariant is body is None ⟺ variants != ()
(hir.md §1.1); the verifier rejects a
body is None function with no variants, a variant whose body is pass
(a variant MUST carry a real body), and a real body combined with
variants. The @base.specialize(...) parse rejects a pass-bodied
variant directly.
5.1 GridRegionExpr carry-out lifting¶
Inside a for i in tile(...) body, an ast.Assign whose single
Name target is already bound in outer scope is a loop-carried
rebinding. The parser:
- Allocates a fresh phi
Varper carry name (same type, same name). - Records the phi in
GridRegionExpr.carried_args. - Snapshots the final RHS bound to that name as a
yield_value. - After the loop, rebinds the carry name in the outer frame to the
GridRegionExpr(single carry) or projects each carry value out of itsTupleTyperesult (multi-carry).
Only = assignments are accepted; += is rejected. return and a
nested with inside a loop body are rejected. A nested for ... in
tile/range(...) IS allowed and lifts to a nested GridRegionExpr; the
carry scan recurses into it, so an outer-scope name rebound only inside the
nested loop is carried across both loops.
5.2 Constraint attachment¶
The HIR parser attaches one ScheduleConstraintMetadata record to one
concrete tensor Expr. Tensor parameters, tensor-valued intermediate SSA
values, and bound tensor-valued TupleGetItem values are valid subjects.
Whole tuples, shape scalars, unit values, direct subscripts, and unresolved
names are rejected. Inline and standalone annotations for the same Expr are
duplicates, not merged declarations. Diagnostics identify the subject and
retain the authored source location. Constraint metadata does not alter the
tensor type or introduce an HIR node.
6. TIR parser¶
The TIR parser walks a @prim_func body. The body is a sequence of
imperative statements that fold into a Sequential of Stmts.
| Python | TIR action |
|---|---|
x = expr |
LetStmt(var=x, value=expr, body=<sequential rest>). The remaining body of the function is nested as body. |
a = Tensor(...) |
LetStmt(var=a, value=Call(tir.memory.AllocTensor, (), attrs=<TensorType fields>), body=<rest>). See tir §2.3. |
foo(a, b) (effect Op) |
Evaluate(target_op, args) Stmt. |
foo(a, b) (value Op) |
Call Expr embedded in the right-hand side of a LetStmt or another Stmt's Expr field. |
for i in range(...) |
For(induction_var=i, start, stop, step, body). |
if/elif/else |
If(cond, then_body, else_body). |
while |
While(cond, body). |
with Mesh(...) as m |
MeshScope(mesh, binding=m, body). |
return |
Return() Stmt. A return value is rejected. |
TIR has no SSA-as-DAG sharing rule; every binding is an explicit
LetStmt. for i in tile(...) is HIR-only and is rejected here.
7. Validation and rejection¶
- Any
astnode not in §4 / §5 is rejected — the unsupported forms includetry/withover non-Mesh contexts /lambda/ list / dict / set comprehensions /yield/async. - Cross-dialect callees fall through to unknown callable (§4.6).
- A bare-name callee that the lexical env / closure does not
resolve to an
_op_schema-bearing surface value (Opsubclass or alias builder function) is unknown Op name. Tensor[...]with the wrong number of slots, an unknown dtype, a non-injective layout, or anast.Sliceshape element is rejected.for tileis HIR-only; emitting it inside a@prim_funcis rejected.with Mesh(...) as mis accepted in both dialects — in a@prim_funcit lowers to aMeshScopeStmt (§6), unlike the no-IR-node HIR sugar (§1.6).- Layout sugar that would lose mesh information falls through to the verbose form (§1.4); if neither is acceptable, the type is rejected.