A five-minute first pass.
Build the pipeline, calibrate a tiny model, and read the receipt. The whole loop is a single llama-quantize invocation.
Or grab a binary.
Prebuilt binaries track the latest GitHub release. Same pipeline, no build. Checksums live beside each asset. If the release has no assets yet, you will land on the releases page — install from source below.
macOS
Tessera for Mac
curl -L -o llama-latest-bin-macos-arm64.tar.gz https://github.com/Tribunus-dev/tessera/releases/latest/download/llama-latest-bin-macos-arm64.tar.gz && tar -xzf llama-latest-bin-macos-arm64.tar.gz
Linux
Tessera for Linux
curl -L -o llama-latest-bin-ubuntu-x64.tar.gz https://github.com/Tribunus-dev/tessera/releases/latest/download/llama-latest-bin-ubuntu-x64.tar.gz && tar -xzf llama-latest-bin-ubuntu-x64.tar.gz
.tar.gz containing build/bin/ + LICENSE built with -DGGML_NATIVE=OFF for portability. Check sha256sum against the release notes.
What you need before you start — from source.
- A C++17 or later toolchain (clang or gcc).
- CMake 3.20+.
- Python 3.10+ for the calibration pipeline.
- A small model to calibrate against (a 1.1B parameter Llama is enough for the smoke test).
- A small calibration corpus. The Tessera CC0 procedural set is shipped baked into the binary; you do not need to download a corpus for the first run.
Build the pipeline.
git clone https://github.com/Tribunus-dev/tessera
cd tessera
cmake -B build -DGGML_METAL=ON -DGGML_OPENBLAS=ON
cmake --build build --target llama-quantize -j
The output binary is build/bin/llama-quantize. The Tessera C++ pipeline is in the same binary as the upstream quantizer. The default mode is Tessera; opt back to stock K-quants with --tessera-mode=off.
Calibrate, then quantize.
Pass a model to llama-quantize with a Tessera type and let the pipeline run end to end:
build/bin/llama-quantize \
--tessera-mode=on \
--tessera-policy=auto \
--imatrix \
--calib-corpus=tessera-cc0 \
models/tinyllama-1.1b.gguf \
models/tinyllama-1.1b.tessera-t640.gguf \
TESSERA_T640
The pipeline runs three stages: calibration (the built-in corpus), AWQ policy search (the GA), and quantization (the writer). Each stage writes its output to the sidecar. The final artifact is the GGUF; the receipt is in the sidecar.
Read the receipt.
Run the A/B harness to re-dequantize the artifact through the kernel and compare against the original:
python3 tools/tessera/tessera-ab-harness.py \
--baseline models/tinyllama-1.1b.gguf \
--candidate models/tinyllama-1.1b.tessera-t640.gguf \
--metric t_l2
The harness reports the per-tensor t_l² and the aggregate. The t_l² is the Frobenius distance between the original weight and the kernel's dequant output, normalized by the max weight magnitude. A green receipt is one where the t_l² matches what the policy predicted.
tessera-ab-harness v1 kernel T640 sha256:7f3a…c9e1
✓ aggregate t_l² 0.0016 validated 202/202 tensors within tolerance
receipt → tinyllama-1.1b.tessera-t640.receipt.json
Full per-tensor table at /evidence/
TESSERA_T640 failed: need --tessera-mode=on means you ran stock llama-quantize. Re-run with --tessera-mode=on --tessera-policy=auto. See Colophon → Build for the assembler step.
Where to go from here.
- The architecture — what the pipeline is doing under the hood.
- The evidence — receipts from the most recent runs.
- The changelog — what shipped, in order.