* Documentation

SLM Integration

Grammar-constrained sampling, computed confidence, and sandboxed validation — how the local model in --smart is actually kept honest.

This page covers the second half of the --smart pipeline — the model-serving layer itself, once context assembly and classification have already produced a prompt. The whole point of this design is that the model’s output is never trusted on its own; it’s constrained at generation time and verified after.

4. Grammar-constrained output, at the token level

A GBNF grammar is loaded into the sampler chain via llama.cpp’s grammar-masking sampler, fixing the JSON schema exactly: six required fields in a fixed order, a fix_mode field restricted to the literal tokens MODIFY or NONE, and a category field restricted to the exact set of category names the compiler understands.

This isn’t a validation step run after the fact — it’s enforced during sampling, so it is not merely unlikely for the model to emit a malformed response, it is not sampleable in the first place. The sampler chain also applies a repetition penalty ahead of greedy decoding, specifically to counter the tendency of small models to loop under purely deterministic sampling.

5. Confidence is computed, never asked for

The model is never prompted to self-report a probability — small models anchor to whatever example values appear nearby in the prompt and echo them back regardless of actual certainty, making a self-reported number meaningless. Instead, confidence is computed after parsing, from measurable properties of the actual response:

6. The suggested fix is never trusted — it’s tested

Before anything is shown to you, the fix is applied to an isolated snapshot of the compiler state — a real sandboxed copy of the source, symbol table, and AST, not the live one — and speculatively re-parsed and re-validated. Only a fix that demonstrably resolves the original error, without introducing new ones, is surfaced. Anything else is silently discarded, and you see the compiler’s own original diagnostic, unmodified.

This sandbox mechanism lives in include/Sandbox.h/src/Sandbox.cpp, and the pipeline wiring that ties context assembly, the model call, and sandbox validation together lives in include/SLMIntegration.h/src/SLMIntegration.cpp (see Compiler Internals).

Model and runtime

The model runs locally via llama.cpp, vendored as a submodule under external/llama.cpp and built automatically as a static library on first configure if not already built (see Tooling). There is no network call and no external service dependency — --smart works entirely offline, and no source code ever leaves the machine running the compiler.

What this is not

This is not the @slm { } in-source syntax (“Prompt-Oriented Programming”) described in the Roadmap — that’s a separate, not-yet-implemented feature where a Turf program itself would contain a natural-language block the compiler turns into real AST nodes. --smart only ever assists with fixing an existing compiler-detected error; it never generates new program logic from a prompt.