* Documentation

Compiler Architecture

The five-pass pipeline: Lexer, Parser, Resolver, TypeChecker, Codegen — and why it is split this way.

Turf’s compiler pipeline is a standard deterministic front-to-back — nothing about compilation itself is probabilistic. What Turf adds on top is a diagnostics layer (--smart, see SLM Integration) that treats the compiler’s own internal state as ground truth for AI-assisted suggestions, rather than pattern-matching formatted error strings.

Pipeline stages

source.tr
   │
   ▼
 Lexer  ──tokens──▶  Parser  ──AST──▶  Resolver (2 passes)  ──annotated AST──▶  TypeChecker  ──type-annotated AST──▶  Codegen  ──▶  native code
PassInputOutput
Lexerraw source bytesa stream of Token values, each carrying a SourceLocation
Parsertoken streaman AST — a tree of ExprAST subclass instances (~42 concrete node kinds)
ResolverASTthe same AST, mutated in place: every name reference gets a resolved symbol/candidate-list annotation written directly onto its node
TypeCheckerresolved ASTthe same AST again, plus a side table mapping every node to an inferred type
Codegentype-annotated ASTLLVM IR → a native object file → a linked binary (or, with --emit-llvm, human-readable IR text)

This is a genuine architectural constraint, not just a description after the fact: Codegen only ever runs on an AST that has already passed both the Resolver and the TypeChecker cleanly. If the Resolver fails, the process returns before the TypeChecker is even constructed; if the TypeChecker fails, the process returns before a single codegen() call runs. Codegen is written to trust this — it reads resolved call targets and inferred types with no re-validation, on the assumption that if it’s running at all, those fields are already correct.

Why five separate passes instead of one combined walk?

Forward references. A name can be used in the source before it’s declared — two top-level functions calling each other, a struct field referencing a type declared later in the file, a call to a function defined further down. A single combined parse-and-check walk has no way to answer “does bar exist” while still checking foo’s body, if bar is textually declared after foo. Splitting name registration (hoisting, done once up front) from name use (checked afterward, with every name already registered) is the general solution — the Resolver itself does this in two internal passes (below). The same logic extends further down the pipeline: the TypeChecker needs every name already resolved before it can ask “what type does this resolve to,” and Codegen needs every type already checked before it can safely choose which LLVM operation to emit for e.g. a + b (integer add vs. float add vs. string concat).

The double-parse wrinkle

The compiler doesn’t literally run each pass once. It first runs a pre-pass: a throwaway Resolver/TypeChecker over the whole file, purely so LLVM function prototypes (and fully built struct types) exist before any function body is compiled — this is what makes main() calling a helper() defined later in the file work at the Codegen layer, independent of the Resolver’s own two-pass hoisting at the name-resolution layer. Concretely, the pre-pass parses every top-level declaration, runs an informational (non-user-facing) Resolver/TypeChecker pass, fully codegens every struct declaration first (struct layouts must be complete before any GEP referencing them is built), emits every function prototype with no body, then fully codegens every function body.

The file is then closed, reset, and re-parsed from scratch for the real, authoritative pass: a fresh AST is built, the mandatory Resolver/TypeChecker run over it, and only non-function top-level statements are codegenned again (function bodies were already emitted in the pre-pass). The pre-pass’s own diagnostics are discarded — the mandatory main-pass TypeChecker re-encounters the identical error and produces the real, user-facing diagnostic. This two-parse structure is an implementation detail, orthogonal to the five-pass pipeline conceptually: from the Resolver/TypeChecker/Codegen’s own point of view, each still only ever sees “an AST,” built the normal way, once per real invocation.

The Resolver: two-pass name resolution

Within the Resolver’s single pipeline stage, name resolution itself happens in two internal passes: a hoisting pass that registers every top-level declaration’s name into scope before checking any body, followed by a resolution pass that walks every body and annotates each name reference with what it resolved to (or a candidate list, for overloaded calls). This is what allows the forward-reference patterns above to work without special-casing them at each call site.

The TypeChecker: visitor-pattern type inference

A visitor-pattern pass over the already-resolved AST. It infers a type for every expression node and writes it into a side table rather than mutating the AST directly, then checks that every operation (assignment, call, binary op, cast, recover, generic instantiation) is well-typed under Turf’s type rules. Generic type parameters are resolved back to their concrete instantiation type at each use site here — this is why Box<int>.get() is understood to return int, not a bare T (see Generics).

Codegen: walking the type-checked AST

Each function is lowered to LLVM IR by walking its type-annotated AST body-first, using the type information already computed by the TypeChecker to choose the correct LLVM instruction at each node (e.g. fadd vs. add for +, depending on whether the TypeChecker resolved the operands to double or int). This is also where the compile-time-provable paths for state machine transitions are recognized and compiled to a plain store instead of a runtime switch.

The CFG static-analysis subsystem

Independent of ordinary type checking, a control-flow graph is built for every function body and a fixed sequence of dataflow analyses runs over it on every compile — definite assignment, liveness (dead-store/unused-variable), reachability, cyclomatic complexity, infinite-loop/infinite-recursion detection, and a malloc/free points-to-lite pass for use-after-free/double-free. See Tooling for the full list of what each one catches.

Generic type erasure

Every instantiation of a generic struct shares the exact same underlying memory layout: a field typed with the struct’s own type parameter erases to a uniform pointer-sized slot regardless of the concrete type, and a generic method’s body is type-checked and codegenned exactly once, shared across every instantiation — there is no per-instantiation re-checking or re-codegenning the way a monomorphizing implementation (Rust, C++ templates) would do it. This keeps generics cheap and the compiler simple, and it’s the direct cause of the one documented stdlib limitation around Map<K, V>/Set<T> string-key comparison (see Standard Library).

Backend

LLVM handles target-independent optimization and native code generation — Turf reuses LLVM’s own optimization passes rather than implementing custom ones, so backend performance tracks upstream LLVM directly. --emit-llvm writes the generated IR as human-readable text instead of proceeding to object-code emission and linking.