* Documentation

Compiler Internals

A deep dive into the Lexer, error recovery via setjmp/longjmp, and the file map for each pass.

This page goes one level deeper than Compiler Architecture — the mechanics of specific passes, grounded directly in the current source layout.

The Lexer

include/Lexer.h, src/Lexer.cpp. Converts the raw source byte stream into a sequence of Token values — a flat set of negative-int constants for multi-character tokens/keywords, plus raw ASCII values for single-character operators.

Scanning is a straightforward hand-written switch over the current character — maximal-munch tokenization with no separate tokenizer-generator or regex table. Each branch greedily consumes the longest token it can: the : branch peeks for a second : to decide between TOK_COLON and TOK_SCOPE (::); the . branch peeks up to two characters ahead to disambiguate . / .. (range) / ... (variadic).

Identifiers are checked against a keyword table after being scanned as an ordinary identifier run — and builtin/stdlib function names (print, malloc, abs, …) are deliberately not in that table. They’re ordinary identifiers, resolved later by the Resolver against lib/std/*.tr, not lexer-level keywords. This matters architecturally: the lexer has zero knowledge of the standard library at all. Every “builtin-looking” name is exactly as unprivileged as a user’s own function name until the Resolver decides otherwise.

Source locations and span tracking

struct SourceLocation {
  int Line;
  int Col;
  std::string Filename;
  int EndLine = -1;
  int EndCol = -1;
};

Line/Col mark where the token starts; EndLine/EndCol mark one column past where it ends, populated for every token kind uniformly. -1 means “no end position recorded,” distinct from “same column as the start” — this is still true for genuine multi-token spans (a whole declaration, a qualified A::B::C path, a failing sub-expression). This single/multi-token distinction is exactly what lets the diagnostic renderer underline a token’s full width (^ followed by ~ for the rest of the span) instead of a single caret whenever an end position is available, and fall back to a bare ^ when it isn’t — see Tooling.

Filename is stamped at token-scan time from a global that’s repointed at each imported module’s own path while that module is being parsed, then restored afterward. This is what lets a diagnostic raised by a later pass — Resolver, TypeChecker, or Codegen, all of which run after every file involved has already finished parsing — still recover the correct source file and line for a location that originated inside an imported module.

Error recovery: setjmp/longjmp, not exceptions

Diagnostics are collected, not thrown as C++ exceptions (LLVM is commonly built with exceptions disabled). Raising a TurfError records the diagnostic, then longjmp()s back to the nearest active setjmp() recovery point — used pervasively by the parser and the pre-pass/main-pass driver loops for per-statement recovery: on a parse error, the compiler jumps back to a known-good point (typically the start of the next statement or function) and keeps going, so a single-invocation compile can report multiple independent syntax errors instead of stopping at the first one.

A single DiagnosticEngine is the choke point every diagnostic — errors and warnings alike — funnels through, and it caps at one primary error per source line to avoid cascading noise from one bad statement (see Tooling).

File map by pass

StagePrimary files
Lexerinclude/Lexer.h, src/Lexer.cpp
Parsersrc/ParserCore.cpp, ParserDecl.cpp, ParserStmt.cpp, ParserLiteral.cpp, ParserModule.cpp; AST node definitions in include/AST.h, ASTExpr.h, ASTStmt.h, ASTDecl.h, ASTAccess.h
Resolversrc/ResolverCore.cpp, ResolverDecl.cpp, ResolverExpr.cpp, ResolverStmt.cpp, ResolverCall.cpp; symbol tables in include/ResolverSymbolTable.h, ResolverScopeTable.h, ResolverModuleTable.h
TypeCheckerinclude/TypeChecker.h; src/TypeCheckerCore.cpp, TypeCheckerCall.cpp, TypeCheckerDecl.cpp, TypeCheckerExpr.cpp, TypeCheckerStmt.cpp; the visitor interface itself in include/ASTVisitor.h
Codegensrc/CodegenCore.cpp, CodegenFunc.cpp, CodegenExpr.cpp, CodegenStruct.cpp, CodegenVar.cpp, CodegenArray.cpp, CodegenControlFlow.cpp
CFG subsysteminclude/CFG.h, src/CFG.cpp; include/Analysis.h, src/Analysis.cpp; src/AliasAnalysis.cpp; include/ValueRange.h, src/ValueRange.cpp; include/CallGraph.h, src/CallGraph.cpp
Type systeminclude/Types.h, src/Types.cpp
Driversrc/main.cpp
--smart pipelineinclude/SLMEngine.h/src/SLMEngine.cpp (the llama.cpp binding), include/SLMIntegration.h/src/SLMIntegration.cpp (pipeline wiring), include/Sandbox.h/src/Sandbox.cpp (speculative re-validation)

The AST

Roughly 42 concrete node kinds, all deriving from a common ExprAST base, split across AST.h (shared infrastructure), ASTExpr.h (expressions — literals, binary/unary ops, calls, casts), ASTStmt.h (statements — if/while/for/compare/recover), ASTDecl.h (top-level declarations — functions, structs, enums, machine, operator), and ASTAccess.h (member/index/namespace access). ASTVisitor.h defines the visitor interface the TypeChecker implements; ASTPrinter.h backs --dump-ast and friends; ASTDiff.h supports the --smart pipeline’s sandboxed before/after comparison.

Built-in functions

A small, fixed set of true compiler builtins (as opposed to stdlib functions, which are ordinary Turf source resolved through the module system) live in Builtins.cpp and are recognized directly by the compiler rather than resolved as calls into lib/std/.