TileFoundry Spec — Code Organization¶
Implementation guide (not an architecture spec). Defines the Python source tree layout.
1. Directory skeleton¶
This spec describes the stable layers only — whether a single
.pyfile exists is decided by the naming rules in §2 and is not enumerated here (so adding a new Op does not require a spec edit). The current file list is whatevergit ls-files src/tilefoundryreports.
The top-level package src/tilefoundry/ is divided as follows. Each
directory has an owning spec — that spec is the single source of
truth for the directory's structure and invariants.
| Directory | Owning spec | Contents |
|---|---|---|
ir/core/ |
core-ir | Shared node algebra: Module / Expr / Var / Constant / Tuple / Op / Call / Stmt (base class) / OpSchema / ParamDef / @register_op / @register_alias / op_registry / errors. |
ir/types/ |
types | Type-system root: TensorType / TupleType / UnitType / CallableType / DType / StorageKind / resolve_storage / dim.* (with their typeinfer). |
ir/types/shard/ |
shard | Shard / layout sublayer: Topology / Mesh / Layout / ComposedLayout / ShardLayout / ShardAttr (Split / Broadcast / Dynamic / Partial). The physical nesting reflects the spec's "sublayer" relationship. |
ir/visitor.py |
visitor-mutator | ExprVisitor / ExprMutator / StmtVisitor / StmtMutator / StmtExprMutator. |
ir/hir/ |
hir | HIR Op layer; one subdirectory per category (math/ / tensor/ / nn/ / shape/ / sharding/). One real Op per .py (§2 rule 1); surface-alias schemas have no per-name file and live in each category's aliases.py (§2 rule 5). |
ir/tir/ |
tir | TIR layer: stmt.py re-exports the Stmt base from ir/core/stmt.py; stmts.py hosts the TIR Stmt subclasses (LetStmt / Evaluate / Sequential / MeshScope / …); prim_function.py; effect Ops and TIR-owned Expr Ops by category (memory/ / nn/ / …); launch.py owns Launch and its authored launch-attribute descriptors; arith.py / reduce.py for tag-dispatched Binary / Unary / Reduce; intrinsic.py for the @intrinsic decorator. Target-specific nodes nest under ir/tir/<target>/<category>/ (e.g. ir/tir/cuda/nn/mma.py) per §2 Rule 1c. |
parser/ |
parser | DSL → IR parsing: base.py (shared visitor base + dispatch), hir_parser.py (@func body), tir_parser.py (@prim_func body), layout sugar / range-slice / dispatch modules. Not under ir/ — the parser is a producer of IR, not an IR sublayer. |
analysis/ |
analysis | Fact layer over typed HIR: poly.py (the polyhedral model — extract / TileGraph and the facts measured over a time relation), and one module per analysis family. The compact public surface lives in analysis/__init__.py; per-target atom catalogues and Facts projections live with their owning Target. |
schedule/ |
schedule | The public Schedule boundary in schedule/__init__.py -- the schedule() operation, immutable options, result, and plan base. One directory per algorithm family: pipeline/ for asynchronous overlap within a cooperating unit, partition/ for spatial division across a device. Each owns its private program view, projected Facts, closed problem, solve, and typed plan export; concrete Target packages register which families serve which exact levels. |
passes/ |
passes | Pass framework (pass_base.py / pass_manager.py) plus concrete transforms (transforms/<pass_name>.py, §2 rule 6). |
target/ |
target | Compilation target capability descriptors and architecture/device facts: Architecture / Device / Target / CudaTarget / CpuTarget / resolve_target, plus the deferred loaders that let each backend's Facts projections and scheduling algorithms register themselves on import. |
target/hardware/ |
target | The installed hardware database and its generic machinery: the authored Architecture / Device documents, the envelope and evidence-leaf loader, HardwareSpecRegistry, and the exact-key schema reader. It fixes the envelope only; the fact namespace below facts belongs to the target package named by a document's schema. |
target/<backend>/spec.py |
target | One backend's typed hardware schemas and the documents it installs, registered into the shared registry as an import side effect. This is where a fact path, its unit, and its cross-field invariants are validated, and where a document becomes an immutable Architecture / Device value. |
registry.py |
analysis | The shared exact-key AlgorithmRegistry: algorithms bound to a (Target concrete type, selector) pair, with separate instances per dispatching stage rather than one differently shaped system per stage. Generic — it names no concrete target and applies no subclass or default fallback. |
analysis/api.py |
analysis | The public composed Analyze operation: preflight, dependency closure, ordering, single execution per member, Metadata-ownership enforcement, and semantic result assembly. |
analysis/registry.py |
analysis | Analysis registration: an algorithm's selector, declared dependencies, and the Metadata types it owns. |
analysis/errors.py |
analysis | AnalysisError, the one diagnostic the whole analysis layer raises, so catching an analysis failure catches every analysis failure rather than the subset the caller happened to import. |
analysis/walk.py |
analysis | How every family reads the authored program: SSA-DAG traversal order, the values of a function, the reachable call graph, a type's tensor leaves and byte size per storage level, and the execution count its meshes imply. Names no Target — two families cannot disagree about the program they measured. |
analysis/preflight.py |
analysis | The gate every analysis runs behind: authored-type re-derivation and rejection of a program no analysis can measure. Established once per public call rather than per family. |
analysis/facts.py |
analysis | The narrow Facts aggregates the families declare — the memory hierarchy graph, the throughput rates, and the parallel capacity. It is the record of how much hardware each measurement rests on, and names no backend. |
analysis/metadata.py |
analysis | The typed records the families leave on the IR, split by what each number depends on rather than by convenience. |
analysis/compute_cost.py |
analysis | The compute-cost family: logical flops per DType and bytes per storage level, from the authored program alone. |
analysis/memory.py |
analysis | The memory family: value lifetimes, per-level peaks, and the capacity comparisons against a target's hierarchy — failing on an over-full addressable level and advising on an over-full cache. |
analysis/roofline.py |
analysis | The roofline family: the recorded work divided by the target's published rates, per Call and aggregated per Function. Adds no count of its own. |
analysis/timeline.py |
analysis | The timeline family: execution-unit fusion, wave decomposition against a parallel capacity, and the placement the scheduling model solves for. |
schedule/api.py |
schedule | The public schedule() operation and ScheduleResult: resolve the Target and the requested level from the Module, dispatch once on the exact pair, and verify the returned Plan. Generic -- it names no concrete target. |
schedule/plan.py |
schedule | SchedulePlan, the extensible semantic base every algorithm's result derives from, and PlanVerificationError. It fixes three operations and no shape: there is no shared schema, version, deserializer, or renderer registry. |
schedule/registry.py |
schedule | The Schedule instance of the shared exact-key AlgorithmRegistry, keyed on (concrete Target type, topology name), and the registration decorator. |
schedule/errors.py |
schedule | ScheduleError, the one diagnostic the schedule layer raises for a request it cannot serve or a solve that failed. Distinct from PlanVerificationError, which says a plan was produced and does not hold together. |
target/<backend>/schedule.py |
target | One backend's scheduling algorithms, registered for the exact levels it schedules as an import side effect. This is where that backend's private problem construction, solve, and Plan export are composed. |
schedule/partition/ |
schedule | The spatial partition family: program extraction, PartitionFacts, the closed candidate problem, the CP-SAT solve, and the PartitionSchedulePlan export. Every hardware number enters through the Facts, so no module below the family entry holds a Target, and nothing in it rewrites the program it decided about. |
visitor_registry/op_cost.py |
analysis | Each operation's per-instance flops and bytes, registered into the shared cost-evaluator registry. Owned here rather than by any target package, because the work an operation asks for follows from its own semantics and operand types on every backend. |
schedule/facts.py |
schedule | AtomFact, the one instruction fact every algorithm family reads the same way. Everything else a family needs from a target is declared by that family, so no aggregate here becomes a vocabulary another family has to satisfy. |
inspection/analysis_report.py |
inspection | Text and JSON renderings of an analysis result and the records it left on the IR, both built from one report structure so the two cannot disagree. Analyses do not format their own output. |
target/<backend>/facts.py |
target | One backend's Facts projections: what that target tells each analysis family, registered under exact (Target concrete type, Facts type) pairs. Restates the installed documents in the shape a family declared and measures nothing. |
target/facts.py |
target | TargetFactsRegistry and the Target.as_facts projection boundary: conversions from a concrete Target to the immutable aggregate a requesting algorithm declares, keyed by the exact (Target concrete type, Facts type) pair. Generic — it names no concrete target, so a backend adds a registration rather than a branch here. Distinct from the hardware-specification and algorithm registries. |
codegen/ |
codegen | Code generation: the emitter registry, the linkable / linked products and the linker, and per-target emitters under <target>/ (mirroring ir/tir/ file layout — tir/<category>/<name>.py emitter, §2 rule 2). Not under ir/ — codegen is a consumer of IR; templates/ holds boilerplate only (kernel shells / host stubs). |
runtime/ |
runtime | Runtime support (per-target headers, function templates, launch helpers). |
inspection/ |
inspection | IR visualisation: DOT, Python printer, web viewer. |
dsl/ |
parser (authoring namespace) | User-facing import surface: tf/ (HIR namespace) / T/ (TIR namespace) / _stub_gen.py / __main__.py. The tf/__init__.pyi and T/__init__.pyi stubs are produced by python -m tilefoundry.dsl regen and are gitignored. |
compile.py |
architecture | tilefoundry.lower / build / compile top-level public verbs. |
script.py |
parser | @func / @prim_func / @module decorator entry points. |
utils/ |
code-organization | Shared leaf machinery: a module here MUST import nothing from ir/, parser/, passes/, codegen/, runtime/ or cli/, and MUST name no layer. It is depended on and depends on nothing, which is what lets a consumer outside the package — a pre-commit hook under an interpreter with nothing installed — load one of these modules by path and get the same implementation the package uses. A helper that needs to know a layer belongs in that layer; this is not a home for anything that did not fit. |
Stage boundary. The pipeline picture in
architecture §1 places
parser/, schedule/, and codegen/ outside ir/ (front-end producer,
decision service over typed HIR, and back-end consumer); the physical directory
layout reflects that boundary directly.
Reading notes:
ir/holds the IR proper and its sublayers only.ir/types/is the root of the type system;ir/types/shard/is its shard / layout sublayer (architecture §3). The physical nesting reflects the spec's conceptual "sublayer".- The placement of
shard/undertypes/is a filing decision, not a consumer restriction:Topology/Mesh/Layout/ShardLayoutare consumed directly byparser,tir, andcodegen. The hierarchy expresses "role in the type system", not "who may import it". codegen/andparser/sit outsideir/. By the architecture §1 pipeline they are the front-end producer and back-end consumer of IR, not IR sublayers.schedule/sits outsideir/because it defines an operation over typed HIR, not a new IR layer.schedule/__init__.pycontains only the public operation and its shared value structures; the construction stages are imported from their own modules, and each algorithm family's candidate graph, solver model, and decoded solution stay private to that family.analysis/sits outsideir/for the same reason: it derives facts about typed HIR rather than defining an IR layer. It reads the IR and theTarget, and neverschedule/— the dependency between the two runs one way (architecture §5).codegen/<target>/consumes only TIR. The subtree mirrorsir/tir/:prim_functionlives intir/, Stmt emitters intir/stmts/, andmemory//nn//arith//reduce//tensor/each have their own subdirectory. There is nocodegen/<target>/hir/.- Authored launch attributes belong to
ir/tir/launch.py; launch-geometry derivation (grid / block extents) is an internalcodegen/cuda/emit.pyhelper (_derive_launch_config), consumed within codegen itself rather than carried past it as a runtime-owned metadata type. The two launch contracts are distinct even though both are consumed across the codegen boundary.
2. File naming and content rules¶
Rule 1 — one real Op = one file. A real Op class lives in
ir/<hir|tir>/<category>/<op_name>.py. The file name is the
snake_case of the Op class CamelCase (MatMul → matmul.py,
RMSNorm → rms_norm.py). TIR effect Ops and TIR-owned Expr Ops
follow the same rule.
Rule 1a — surface-alias schemas have no per-name file. A surface
alias (core-ir §2.3)
has no IR class — its builder routes to a kinded target Op. All
aliases for a category live together in aliases.py (e.g. the 19
HIR math sugar names add / sub / cmp_eq / neg / … all
register in ir/hir/math/aliases.py).
Rule 1b — tag-dispatched IR classes. Binary / Unary /
Reduce and other Op classes that fold many surface names through a
kind attribute live in one file per IR class
(ir/hir/math/binary.py / ir/hir/math/unary.py /
ir/tir/arith.py / ir/tir/reduce.py). This does not contradict
Rule 1: "one Op = one file" means one IR class per file; aliases
are not IR classes, so they go through Rule 1a.
Rule 1c — target-specific IR nodes nest under the dialect. IR is
dialect-first: its primary organizing axis is the dialect, and most
nodes are target-neutral. A node or descriptor that is specific to one
compilation target nests as ir/{dialect}/{target}/{category}/<name>.py;
target-neutral abstractions stay at ir/{dialect}/{category}/. For
example the whole MMA surface is target-owned — the Mma op, the
MmaOpSpec / MmaAtom descriptors, the CUDA SM80 instruction spec, and its
fragment layouts all live under ir/tir/cuda/nn/ (mma.py + mma_atom.py),
and the HIR per-shape Mma_SM80_* / Wgmma_SM90_* ops under
ir/hir/cuda/nn/mma.py — because an MMA instruction fixes a concrete hardware
op. (codegen/ and runtime/ are target-first instead — their primary
axis is the target — so each tree is organized by its own primary axis.)
Rule 2 — one (node, target) codegen = one file. Each handler
lives at codegen/<target>/tir/<category>/<name>.py. Stmt emitters,
Expr-Op emitters, and tag-dispatched (arith, reduce) emitters
each get their own file. Codegen consumes TIR only.
Rule 3 — what an IR-class file contains:
- HIR Op file (
ir/hir/<cat>/<name>.py): Op class +@register_typeinfer(Op)+@register_cost_evaluator(Op)(if any). - TIR effect Op file (
ir/tir/<cat>/<name>.py): Op class +@register_typeinfer(Op)(returningUnitType) +@register_verify_stmt(Op). The verify rule keys on the Op class even though the invocation is anEvaluate(op, args)Stmt — see visitor-registry §5. - TIR-owned Expr Op file
(
ir/tir/memory/{alloc_tensor,ptr_of,memory_span,tensor_view}.py, …): Op class +@register_typeinfer(Op)+@register_cost_evaluator(Op)(if any). Call-position constraints (e.g.AllocTensormay only appear asLetStmt.value) are checked by the enclosing Stmt's@register_verify_stmt; there is no separate Op-level verify decorator for these. <category>/aliases.pyfile (Rule 1a): a list of@register_alias(...)declarations whose builders construct the target Op instance. Alias files hold no IR class and participate in neither typeinfer nor verify.
Rule 4 — what a target codegen file contains: the
@register_codegen_<target> for that (op / stmt) pair, and nothing
else.
Rule 5 — <category>/__init__.py re-export rules:
- Real Op submodules are re-exported via
from .<file> import <Cls>(maintained by hand; new Ops add their import here). aliases.pyis imported only for its@register_aliasside-effects; nothing is re-exported from it (aliases have no class to expose).- User imports go through the
tilefoundry.dsl.tf/tilefoundry.dsl.Tnamespaces'__getattr__, not through the per-category__init__.py(parser §2).
Rule 6 — one pass = one file. A pass class lives in
passes/transforms/<pass_name>.py (snake_case file name = pass
class CamelCase in snake form: HirToTirPass → hir_to_tir.py).
Internal visitors / mutators stay in the same file.
Rule 7 — what template files contain.
codegen/<target>/templates/*.j2 carry boilerplate assembly only
(kernel shells, host stubs, fixed-shape per-target wrappers).
Stmt-level and Op-level emitters are not template-driven; they live
in codegen/<target>/tir/.../*.py as Python walkers (see
codegen).
3. Multi-agent parallelism guarantee¶
The lock granularity is a single (node, target) pair. The naming
rules in §2 imply that two agents working on different
(node, target) pairs touch disjoint files; cross-cutting changes
(shard/ fields, kernel templates, pass framework) confine
themselves to the owning directory. Representative scenarios:
| Scenario | Files affected |
|---|---|
Agent A adds Conv2D, Agent B adds RMSNorm |
ir/hir/nn/conv2d.py + ir/hir/nn/rms_norm.py — no conflict |
Agent A edits MatMul.typeinfer, Agent B adds tir.cuda.nn.Mma CUDA codegen |
ir/hir/nn/matmul.py + codegen/cuda/tir/nn/mma.py — no conflict |
Agent A adds a new target cpu |
a fresh codegen/cpu/ subtree — no conflict |
Anything that fits the "one Op = one file" / "one (node, target) = one file" / "one pass = one file" rules above shares the same property by construction.
4. DSL package layout¶
The author-facing surface is delivered as a namespace package:
src/tilefoundry/dsl/
__init__.py # exports `tf` and `T` sub-namespaces
__main__.py # `python -m tilefoundry.dsl regen` CLI
_stub_gen.py # `.pyi` generator (run by regen)
py.typed # PEP 561 marker
tf/
__init__.py # module-level __getattr__ → OpSchema lookup
__init__.pyi # AUTO-GENERATED, gitignored
T/
__init__.py # same pattern for TIR
__init__.pyi # AUTO-GENERATED, gitignored
The two sub-packages (tf and T) follow the same pattern:
__getattr__(name) looks name up in the OpSchema registry for the
corresponding dialect and returns either the Op class (for real-Op
schemas) or the alias builder fn (for surface-alias schemas).
Unknown names raise AttributeError.
4.1 Built-in op-class location convention¶
For @register_op to auto-derive dialect + category, an Op
class MUST live under:
with cls.__module__ matching tilefoundry.ir.<hir|tir>.<category>.*.
The 4th segment of the dotted path is the category. Outside that
path, the decorator requires explicit dialect= and category=
kwargs.
cls.__name__.lower() is the default canonical Op name. When the
canonical name diverges from the class lowercase (e.g. RMSNorm →
rms_norm), pass name="..." explicitly.
4.2 .pyi stub regeneration¶
The .pyi stubs reflect the registered schemas only. After adding
a new @register_op / @register_alias, regenerate stubs via:
The CLI imports tilefoundry.ir (forcing every built-in schema to
register) and writes tf/__init__.pyi / T/__init__.pyi. Stubs
are gitignored — IDEs that need them locally SHOULD run regen on
package install.
5. DSL import surface¶
The author-facing exports route through tilefoundry.dsl:
# canonical authoring imports
from tilefoundry import func, prim_func
from tilefoundry.dsl import tf, T, Tensor
Tensoris the parser-owned DSL authoring-surface annotation sugar; it is owned bytilefoundry.dsl(defined undertilefoundry.dsl._tensor, re-exported astilefoundry.dsl.Tensor). It is not the IR tensor type — the IR type carrier istilefoundry.ir.types.TensorType. See parser §1.4 for the annotation grammar.DTypeis not re-exported. dtype values use string form in DSL source (Tensor[(8,), "bf16"],zeros((1, 64), "bf16", ...)); the parser converts strings toDType.<name>at attribute-binding time when the receivingParamDefdeclaresannotation=DType.- For users who prefer bare Op names (
add(...)/relu(...)),from tilefoundry.dsl.tf import *binds every registered HIR name into the call site's lexical scope. Without that import the parser requires the namespace formtf.add(...).
The tilefoundry.dsl.{tf, T} modules expose __all__ via their lazy
__getattr__, so a star-import sees every name registered against
the corresponding dialect, including custom Ops registered after
the DSL package first loaded.