Concept · 01

Quantization

Definition. Making a model smaller by reducing the precision of its numbers. A 16-bit weight takes two bytes; an 8-bit weight takes one; a 4-bit weight takes half of one. The model fits in less memory, loads faster, runs faster — and the question is how much accuracy you give up.

Analogy. Like compressing a high-resolution photo. The file gets smaller, and at some point you stop noticing the difference. The trick is choosing which details to drop so the eye does not catch them.

See the full treatment on the architecture page.

Concept · 02

Memory wall

Definition. In modern ML inference, the bottleneck is usually memory bandwidth — moving weights from memory into the compute units — not the compute itself. The "memory wall" is when the cost of data movement dominates the cost of actual computation.

Analogy. Imagine a kitchen where the chef can cook a dish in 30 seconds, but it takes 5 minutes to bring the ingredients from the pantry. The bottleneck is not cooking — it is fetching. The memory wall is when the cook spends all their time waiting for ingredients, not cooking. Tessera's quantization-on-transport is like packing the ingredients into a smaller bag at the pantry door, so the cook carries less per trip.

See the dedicated section on the architecture page, with the diagrams.

Concept · 03

Kernel-fidelity loop

Definition. A calibration system that grades candidate quantization recipes against the output of the actual dequant kernel that will run in production, not against an offline proxy. The kernel that grades is the kernel that runs, so the artifact is shaped to match the kernel that signed off on it.

Analogy. Imagine grading student essays. The traditional way is to give the essay to a substitute teacher and trust their judgment. The kernel-fidelity loop is giving the essay to the actual teacher who will be using the grade. The grade reflects the teacher's standards, not someone else's. So when the teacher reads the essay later, the grade holds.

See the full treatment on the architecture page.

Concept · 04

Quantization-on-transport

Definition. Performing quantization as data moves between compute stages, using idle compute during transport, instead of as a one-time pre-processing step. Because the calibration kernel and the runtime kernel are the same, the quantization is kernel-aware and effectively lossless for the model's purposes.

Analogy. Like packing a suitcase while walking to the airport. Instead of packing at home and carrying the whole suitcase, you pack as you walk, using time you would otherwise spend idle. The end result is the same, but you never carry the full weight.

See the full treatment on the architecture page.

Concept · 05

Heterogeneous dynamic execution

Definition. Dispatching work across multiple compute engines — CPU, GPU, and ANE on Apple Silicon — based on what each operation needs. The dispatch is dynamic (decided at runtime, not at compile time) and heterogeneous (the engines are different kinds of hardware with different strengths).

Analogy. A construction site with different crews: the electricians do wiring, the plumbers do pipes, the carpenters do framing. A good foreman sends the right crew for the right job as conditions change. A bad foreman sends everyone to do everything. Tessera is the foreman.

See the dispatch policy on the architecture page.

Concept · 06

Hybrid drafters

Definition. Small, fast models that propose tokens for a larger model to verify. The drafter writes quickly, with errors; the verifier checks each token and accepts or rejects. Acceptance rate determines how much faster the system is than the verifier alone. The "hybrid" part: Tessera runs both a self-attention drafter and a Markov head, and combines their proposals.

Analogy. Like a court stenographer. The stenographer writes down words as the speaker talks — fast, but with errors. The judge (the verifier) checks each word and corrects as needed. The stenographer gets faster over time as they learn the speaker's patterns.

See the architecture page for the dispatch table.

Concept · 07

Apple Neural Engine (ANE)

Definition. Apple's matrix-math accelerator, separate from the GPU and CPU. The ANE is not as flexible as the GPU, but it is much faster for the specific operations it is good at — and uses much less power. Tessera uses the ANE for the verifier's prefill on long prompts; the drafter and the generation steps stay on the GPU.

Analogy. A specialized tool in a workshop. A general drill can do most things; a precision drill press does one thing much better. The ANE is the drill press. The engine chooses the right tool for each step.

Concept · 08

Multi-modal

Definition. Handling multiple types of input — text, vision, audio, and speech — in the same model. Tessera's base model (Gemma 4 12B unified) is a single transformer that ingests all four; the speech-to-speech layer is Qwen 3 TTS, which speaks the model's outputs back to you. Each modality has its own calibration regime; the model carries per-modality activation scales so the same kernel can dequant the right way for each.

Analogy. Most recipes are optimized for one type of cuisine. Multi-modal is like designing a kitchen that is equally good at Italian, Japanese, and Mexican — different tools for different dishes, but all sharing the same workspace.

Concept · 09

Signed receipt

Definition. A structured, signed record attached to every Tessera artifact. The receipt carries the upstream model SHA, the calibration corpus snapshot fingerprint, the per-tensor policy, the eval results, and a cryptographic signature over all of it. Anyone with the receipt can verify what the model is, where it came from, and what it was tested against.

Analogy. Like a notarized document. The notary records what was used, what was done, and what the result was, and signs it. Anyone can verify the document is real and has not been tampered with. The notary's seal does not tell you the document is good, but it tells you the document is real, and the contents are what they claim to be.

See the evidence page for the schema and the live receipts.

The long tail of technical terms — AWQ, MAP-Elites, importance matrix, regime routing, per-tensor GA, IOSurface, CoreML — lives on the architecture page. Each section there links back to the concept card if the reader needs the analogy first.

Where to go next

Pick the question that brought you here.