The pipeline is a chain of typed components.
Each stage takes a typed input, applies a transformation, and emits a typed output. The output is durable. The next stage reads it. Nothing is recomputed from raw artifacts downstream.
The calibrator produces a typed importance matrix.
per_tensor_calibrate reads the model and a corpus and produces per-tensor statistics: kurtosis, effective rank, outlier density, percentile bounds. The output is a typed component (tensor_stats) keyed by tensor name, model hash, and (in multimodal mode) modality. The calibrator's fitness is offline-Frobenius — the calibrator itself is not the ground truth, but its output is the basis for the GA's search.
For multimodal models, the calibrator runs once per modality: text, image, audio, mm projector. The unified output is a single tensor_stats table with the modality stamp on every record.
The GA searches per-tensor alpha and clip against the kernel.
awq-evolve reads the tensor_stats and runs an island-model GA. The fitness is the Frobenius distance between the original weight and the kernel's dequant output, evaluated through the exact code path the runtime will use. The GA's search is over continuous reconstruction knobs (per-tensor alpha and clip), not discrete bit-widths — no prior quantization work searches this space against the kernel directly.
The output is a per-tensor policy: alpha, clip, and a verdict (which quant type to use for this tensor). The policy is a typed component (calibration_policy).
The writer emits the GGUF with the policy in the metadata.
llama-quantize --tessera-mode reads the policy, packs each tensor according to its verdict, and writes the GGUF. The policy is embedded in the metadata as a small string (the tensor-family list, not the U/V payloads). The full policy and the GA archive go in a sidecar JSON. SHA-256 in both, for audit.
For a multimodal model, the writer absorbs three additional components: vision tower, audio tower, mm projector. The singular-GGUF is the constitutional commitment; the multimodal pipeline does not split the artifact across files.
The A/B harness re-dequants through the kernel and reports the same t_l².
tessera-ab-harness re-runs the kernel dequant on the shipped artifact and compares the output to the original weight. The reported t_l² is the same number the GA optimized against. A green receipt is one where the kernel's dequant matches the calibrator's prediction within tolerance. A red receipt is a bug — either in the calibrator, the GA, or the kernel.
The receipt is durable. The artifact is the proof. The proof is the claim.
The fitness is measured against the kernel, not the calibrator.
Most quantization pipelines search a proxy (offline Frobenius, a separate cost model, a benchmark) and ship the artifact. Tessera searches against the actual kernel dequant output. The fitness signal and the runtime signal are the same signal. The artifact is the proof of the search; the search is the proof of the artifact.
The fitness and the runtime signal are the same signal. — Not a proxy. The kernel.