Stage 1 · Calibration

The calibrator produces a typed importance matrix.

per_tensor_calibrate reads the model and a corpus and produces per-tensor statistics: kurtosis, effective rank, outlier density, percentile bounds. The output is a typed component (tensor_stats) keyed by tensor name, model hash, and (in multimodal mode) modality. The calibrator's fitness is offline-Frobenius — the calibrator itself is not the ground truth, but its output is the basis for the GA's search.

For multimodal models, the calibrator runs once per modality: text, image, audio, mm projector. The unified output is a single tensor_stats table with the modality stamp on every record.

Stage 2 · AWQ policy search

The GA searches per-tensor alpha and clip against the kernel.

awq-evolve reads the tensor_stats and runs an island-model GA. The fitness is the Frobenius distance between the original weight and the kernel's dequant output, evaluated through the exact code path the runtime will use. The GA's search is over continuous reconstruction knobs (per-tensor alpha and clip), not discrete bit-widths — no prior quantization work searches this space against the kernel directly.

The output is a per-tensor policy: alpha, clip, and a verdict (which quant type to use for this tensor). The policy is a typed component (calibration_policy).

Stage 3 · Quantization

The writer emits the GGUF with the policy in the metadata.

llama-quantize --tessera-mode reads the policy, packs each tensor according to its verdict, and writes the GGUF. The policy is embedded in the metadata as a small string (the tensor-family list, not the U/V payloads). The full policy and the GA archive go in a sidecar JSON. SHA-256 in both, for audit.

Writer packs per-tensor verdict into singular GGUF GGUF metadata tensor_family_list sha256: policy • GA archive → sidecar.json vision tower mm projector multimodal → singular GGUF verification t_l² via kernel dequant same number the GA optimized
Fig. 1 — Writer output. One file, one policy, one sidecar. Multimodal towers are absorbed, not split.

For a multimodal model, the writer absorbs three additional components: vision tower, audio tower, mm projector. The singular-GGUF is the constitutional commitment; the multimodal pipeline does not split the artifact across files.

Stage 4 · Verification

The A/B harness re-dequants through the kernel and reports the same t_l².

tessera-ab-harness re-runs the kernel dequant on the shipped artifact and compares the output to the original weight. The reported t_l² is the same number the GA optimized against. A green receipt is one where the kernel's dequant matches the calibrator's prediction within tolerance. A red receipt is a bug — either in the calibrator, the GA, or the kernel.

The receipt is durable. The artifact is the proof. The proof is the claim.

Why this is different

The fitness is measured against the kernel, not the calibrator.

Most quantization pipelines search a proxy (offline Frobenius, a separate cost model, a benchmark) and ship the artifact. Tessera searches against the actual kernel dequant output. The fitness signal and the runtime signal are the same signal. The artifact is the proof of the search; the search is the proof of the artifact.

The fitness and the runtime signal are the same signal. — Not a proxy. The kernel.