Calibrate and quantize models with receipts the kernel agrees with.
Most quantization pipelines produce bit-exact artifacts and bit-divergent runtime behavior. The calibrator optimizes a proxy. The kernel dispatches a different code path. The user sees a green benchmark and a hallucinated answer. Tessera closes the gap: fitness is measured against the actual kernel dequant output, the policy ships with the artifact, and every release is auditable end-to-end.
The proxy is the lie.
A typical quantization pipeline is three layers deep. The calibrator produces an importance matrix on a corpus. The optimizer uses that matrix to pick a bit allocation per tensor. The kernel reads the bit allocation and dispatches one of N code paths. Each layer is tested in isolation. The pipeline as a whole is not.
The calibrator's loss is a proxy — typically a Frobenius distance to the original weight, evaluated in fp32 offline. The kernel sees a different weight, dispatched through a different code path, with a different rounding regime. The artifact is bit-exact; the runtime is not. The benchmark reports a green number; the user observes something else.
This is a receipt gap. The artifact carries no record of what the kernel actually computed. A reproduction attempt produces a different number, and there is no audit trail to explain why.
The benchmark reports a green number; the user observes something else. — The receipt gap, in one line
Measure fitness at the kernel, not the calibrator.
If the run is the artifact, the fitness is the distance between the artifact and the run. The calibrator is a means; the kernel is the ground truth. Tessera's fitness is a Frobenius distance measured after the weight has been dequantized through the exact code path the runtime will use — not a proxy, not an offline estimate, the actual output. The policy is shipped with the artifact. The receipt carries the kernel's hash. A reviewer can re-run the dequant and recover the same number.
Three steps. End to end.
-
01
Calibrate
Run
per_tensor_calibrateon a calibration corpus. Per-tensor statistics: kurtosis, effective rank, outlier density. The calibrator's output is a typed component, not a number. -
02
Quantize
Run the GA-prep walk with the calibration policy as the fitness signal. The artifact is a GGUF; the policy is in the metadata. The kernel's dequant output is compared against the original, the receipt is appended, the artifact is written.
-
03
Verify
Run
tessera-ab-harness. The harness re-dequants the artifact through the kernel and reports the same t_l². The receipt is the proof. The proof is the claim.
$ tessera-ab-harness --metric t_l2
aggregate t_l2 0.0016 validated 202/202 tensors · kernel T640 sha256:7f3a…c9e1
attn.q.00 0.0017 · ffn.down.11 0.0014 · norm 0.0006
receipt → tinyllama-1.1b.tessera-t640.receipt.json
A claim is not a claim until it has a receipt.
Every Tessera release ships with the receipts that produced it. The receipts include the calibration corpus, the per-tensor statistics, the GA-prep trace, and the kernel's dequant output for each tensor. The receipt chain is the link from the source model to the shipped artifact. Every link is preserved. Every link is verifiable.
The tools, in order of how you would use them.
| Tool | Purpose | Stage |
|---|---|---|
l1-l6 spine |
Six-layer runtime-aware telemetry (v3.1): FP16 reference capture, activation differentials, KL attribution, Hessian sensitivity, tail-weighted loss. | Measurement |
per_tensor_calibrate |
Per-tensor importance statistics on a calibration corpus. | Calibration |
awq-evolve |
GA-driven AWQ policy search; per-tensor alpha and clip. | Calibration |
l5_orchestrator |
Iterative requant loop; family-aware bit allocation. | Quantization |
unified_calibrate |
Co-calibrate trunk + drafters (DFlash, DSPark, MTP) in one pass. | Multimodal |
multimodal_calibrate |
Modality-aware per-tensor stats for vision / audio / mm projector. | Multimodal |
clip-capture |
Multimodal activation capture against the actual C++ mtmd clip graph. | Multimodal |
embedding_budget |
Size-envelope producer for shared embedding tensors across components. | Quantization |
tessera-ab-harness |
A/B comparison; re-dequants through the kernel and compares the t_l². | Verification |
llama-quantize --tessera-mode |
The C++ quantizer that writes the artifact. The full pipeline runs end to end; opt back to stock K-quants with --tessera-mode=off. |
Writer |
| Tessera Studio | The desktop app built on this engine. The same pipeline, surfaced as a productivity agent. | Surface |
Models with receipts, or yours without them.
Tessera treats the Granite T640 family as first-class: granite-4.1-30b, guardian-4.1-8b, and 4.0-h-small are fully supported and calibrated, with vision, speech-plus, docling, and embeddings covered in the pipeline. Nemotron 3.5 runs too. If you bring your own GGUF, the receipt chain starts the moment you calibrate it.
Who built it
A single architect, sustained.
Tessera is independently developed by Julian Torres. The architecture is a sustained argument that the LLM deployment stack's abstractions are inadequate, and that the right response is a different one. The argument is not a vision document; it is a calibrator, a quantizer, an artifact format, a fitness signal against the kernel, a C++ pipeline, a macOS app, and a constitutional model for receipts.