How Tessera differs. From llama.cpp, MLX, and vLLM.
Tessera is not a quantizer wrapper or a faster serving engine. It is a calibrated, audited quantization system with a local-first consumer product, a data flywheel, and a privacy contract. None of llama.cpp, MLX, or vLLM have all four.
What each one does. Honestly.
A fair comparison starts with what each project is. All three are good projects. Tessera is a different project.
llama.cpp
CPU-first C/C++ inference via ggml, the de facto standard for local LLM. MIT licensed, wide model coverage via GGUF conversion, the substrate most other local projects build on. Tessera is a fork of llama.cpp: same C++ substrate, but extended with the kernel-fidelity loop, the per-tensor policy and signed receipt, ANE prefill, the hybrid drafters, Tessera Studio, and the flywheel. If you already use llama.cpp and don't need any of those, you don't need Tessera.
MLX
Apple's array framework for ML on Apple Silicon, NumPy-like API, unified memory, tight Metal integration. Designed for research and small-to-medium models. MLX is a general-purpose ML framework — not a serving stack and not a quantization system. Tessera could theoretically consume an MLX-trained model, but the quantization is the differentiator, not the framework. If you are prototyping on Apple Silicon with a general framework, MLX is the right tool. If you want calibrated quantization, the kernel-fidelity loop, and a consumer product, Tessera is.
vLLM
High-throughput LLM serving for production. PagedAttention, continuous batching, prefix caching, GPU-first (CUDA, ROCm, TPU, Neuron). The standard for cloud-scale LLM serving. vLLM is for thousands of QPS in a data center; Tessera is for one user on a Mac. Different deployment target, different design constraints. If you are serving production traffic at scale, vLLM is the right tool. If you want local AI on your own hardware, Tessera is.
What Tessera is doing that none of the three are.
| Axis | llama.cpp | MLX | vLLM | Tessera |
|---|---|---|---|---|
| Calibrated & audited quantization | No (vanilla K-quants, no per-tensor policy) | No (general framework, no quantization system) | No (serving engine, no quantization) | Yes — kernel-fidelity loop, per-tensor GA, schema-versioned evidence, signed receipt on every artifact |
| Local-first consumer product | No (CLI / library) | No (framework) | No (serving engine) | Yes — Tessera Studio for Mac and iOS, the three destinations, the eight tools, AION, the agent loop |
| Data flywheel | No | No | No | Yes — distributed training improves the model over time, every user benefits |
| Privacy contract | Implicit (you can run locally) | Implicit (you can run locally) | No (cloud serving) | Yes — sanitization pipeline, public audit trail, raw data never leaves device |
The four axes are the spine. The prose above is the unpacking. The reader walks away with "I get it, Tessera is doing something none of these are doing."
The two architectural mechanisms that make it work.
The four-axis table above answers what Tessera is doing. The next two sections answer how. These are the architectural mechanisms the other three don't have; the rest of the comparison falls out of them.
Heterogeneous dynamic execution across CPU, GPU, and ANE
Apple Silicon has three compute engines: the M-series CPU, the GPU, and the Neural Engine. The right engine depends on the operation: ANE for matrix prefill, GPU for general compute and the drafter, CPU for control flow and serial bookkeeping. Tessera dispatches dynamically at runtime, not statically at compile time. The ANE prefill hands off to the GPU via IOSurface async, so the data flows without a sync barrier. Full dispatch policy on the architecture page.
llama.cpp uses CPU and Metal GPU but does not touch the ANE. MLX is GPU-focused. vLLM does not run on Apple Silicon in the same way (CUDA-first). Heterogeneous dynamic execution is unique to Tessera in this comparison.
Quantization-on-transport to go around the memory wall
The bottleneck in modern local inference is not compute, it is memory bandwidth — loading weights from HBM into the compute units. Standard quantization reduces the data size but at a fixed accuracy cost. Tessera does quantization as data moves between compute stages, not as a one-time pre-processing step. The calibration kernel — the one that grades the dequant at calibration time — is the same kernel that does the runtime dequant, so the quantization is kernel-aware and effectively lossless for the model's purposes. Idle compute during transport is the resource that makes this free. Full treatment on the architecture page.
None of llama.cpp, MLX, or vLLM do this. Their quantization is a static pre-processing step with a fixed accuracy cost. Quantization-on-transport is the architectural reason the kernel-fidelity loop and the runtime performance can coexist.
The positive case.
Use Tessera if you want local AI on your Mac that gets better the more people use it, and you want the receipts to prove the model is honest.
Use llama.cpp if you want a fast, minimal local inference engine and you are not building on a flywheel or kernel-aware calibration.
Use MLX if you are prototyping on Apple Silicon with a general framework, training from scratch, or doing research that does not need a quantized inference stack.
Use vLLM if you are serving thousands of QPS in a data center and your deployment is GPU-bound.
If you are unsure which one fits, the how-it-works page explains the deal in plain language. If you are ready to install, start here.