The engine

Calibrate and quantize models with receipts the kernel agrees with.

Most quantization pipelines produce bit-exact artifacts and bit-divergent runtime behavior. The calibrator optimizes a proxy. The kernel dispatches a different code path. The user sees a green benchmark and a hallucinated answer. Tessera closes the gap: fitness is measured against the actual kernel dequant output, the policy ships with the artifact, and every release is auditable end-to-end.

The problem

The proxy is the lie.

A typical quantization pipeline is three layers deep. The calibrator produces an importance matrix on a corpus. The optimizer uses that matrix to pick a bit allocation per tensor. The kernel reads the bit allocation and dispatches one of N code paths. Each layer is tested in isolation. The pipeline as a whole is not.

The calibrator's loss is a proxy — typically a Frobenius distance to the original weight, evaluated in fp32 offline. The kernel sees a different weight, dispatched through a different code path, with a different rounding regime. The artifact is bit-exact; the runtime is not. The benchmark reports a green number; the user observes something else.

This is a receipt gap. The artifact carries no record of what the kernel actually computed. A reproduction attempt produces a different number, and there is no audit trail to explain why.

The benchmark reports a green number; the user observes something else. — The receipt gap, in one line
The insight

Measure fitness at the kernel, not the calibrator.

If the run is the artifact, the fitness is the distance between the artifact and the run. The calibrator is a means; the kernel is the ground truth. Tessera's fitness is a Frobenius distance measured after the weight has been dequantized through the exact code path the runtime will use — not a proxy, not an offline estimate, the actual output. The policy is shipped with the artifact. The receipt carries the kernel's hash. A reviewer can re-run the dequant and recover the same number.

The plan

Three steps. End to end.

  1. 01

    Calibrate

    Run per_tensor_calibrate on a calibration corpus. Per-tensor statistics: kurtosis, effective rank, outlier density. The calibrator's output is a typed component, not a number.

  2. 02

    Quantize

    Run the GA-prep walk with the calibration policy as the fitness signal. The artifact is a GGUF; the policy is in the metadata. The kernel's dequant output is compared against the original, the receipt is appended, the artifact is written.

  3. 03

    Verify

    Run tessera-ab-harness. The harness re-dequants the artifact through the kernel and reports the same t_l². The receipt is the proof. The proof is the claim.

The receipt — what verify prints
$ tessera-ab-harness --metric t_l2
aggregate t_l2 0.0016  validated  202/202 tensors · kernel T640 sha256:7f3a…c9e1
attn.q.00  0.0017  ·  ffn.down.11  0.0014  ·  norm  0.0006
receipt → tinyllama-1.1b.tessera-t640.receipt.json
The evidence

A claim is not a claim until it has a receipt.

Every Tessera release ships with the receipts that produced it. The receipts include the calibration corpus, the per-tensor statistics, the GA-prep trace, and the kernel's dequant output for each tensor. The receipt chain is the link from the source model to the shipped artifact. Every link is preserved. Every link is verifiable.

See the latest evidence →

What is in the box

The tools, in order of how you would use them.

ToolPurposeStage
l1-l6 spine Six-layer runtime-aware telemetry (v3.1): FP16 reference capture, activation differentials, KL attribution, Hessian sensitivity, tail-weighted loss. Measurement
per_tensor_calibrate Per-tensor importance statistics on a calibration corpus. Calibration
awq-evolve GA-driven AWQ policy search; per-tensor alpha and clip. Calibration
l5_orchestrator Iterative requant loop; family-aware bit allocation. Quantization
unified_calibrate Co-calibrate trunk + drafters (DFlash, DSPark, MTP) in one pass. Multimodal
multimodal_calibrate Modality-aware per-tensor stats for vision / audio / mm projector. Multimodal
clip-capture Multimodal activation capture against the actual C++ mtmd clip graph. Multimodal
embedding_budget Size-envelope producer for shared embedding tensors across components. Quantization
tessera-ab-harness A/B comparison; re-dequants through the kernel and compares the t_l². Verification
llama-quantize --tessera-mode The C++ quantizer that writes the artifact. The full pipeline runs end to end; opt back to stock K-quants with --tessera-mode=off. Writer
Tessera Studio The desktop app built on this engine. The same pipeline, surfaced as a productivity agent. Surface
Tessera pipeline: calibrate to receipt Four stages: per_tensor_calibrate, awq-evolve GA, llama-quantize writer, tessera-ab-harness verifier. Each emits a typed component; the kernel dequant is the ground truth. per_tensor_calibrate tensor_stats · 202 tensors kurtosis / rank / outlier awq-evolve GA · fitness = kernel t_l2 alpha / clip per tensor llama-quantize --tessera-mode · writer GGUF + policy + sidecar ab-harness re-dequant · same t_l2 receipt is the proof kernel ggml-quants.c T640 · sha256:7f3a…c9e1 · measured after dequant, not before
Figure 1. The pipeline as a chain of typed components. Nothing downstream is recomputed from raw artifacts; every stage reads the previous stage's durable output. The fitness and the runtime signal are the same signal — the kernel's dequant.
Tessera Studio graph: four composable nodes — per_tensor calibrate, awq-evolve GA, quantize writer, harness receipt
Fig. 2 — Studio is the same pipeline as a graph. Each stage is a node; the graph is the artifact. The workflow model and SwiftUI editor are shipped; the agent loop — approval tiers, audit log, time-limited undo — has landed in the ux-fatigue waves.
What runs on it

Models with receipts, or yours without them.

Tessera treats the Granite T640 family as first-class: granite-4.1-30b, guardian-4.1-8b, and 4.0-h-small are fully supported and calibrated, with vision, speech-plus, docling, and embeddings covered in the pipeline. Nemotron 3.5 runs too. If you bring your own GGUF, the receipt chain starts the moment you calibrate it.

Who built it

A single architect, sustained.

Tessera is independently developed by Julian Torres. The architecture is a sustained argument that the LLM deployment stack's abstractions are inadequate, and that the right response is a different one. The argument is not a vision document; it is a calibrator, a quantizer, an artifact format, a fitness signal against the kernel, a C++ pipeline, a macOS app, and a constitutional model for receipts.