BrainBank

LLM Inference Engineering Quick Reference Guide

8/29/2026, 11:24:10 AM · updated 8/29/2026, 11:36:36 AM

AI-translated on 8/29/2026, 11:40:11 AM · by Qwen3.6 35B (fast, default)

#knowledge#quantization#llm-inference#local-llm#apple-silicon#inference-engineering#model-serving#kv-cache#speculative-decoding#runtime

This article systematically maps out the core conceptual framework of LLM inference engineering, helping readers quickly identify terminology across each layer — models, runtimes, inference engines, and servers — and master the underlying principles and practical applications of optimization techniques such as KV Cache, speculative decoding, and quantization.

Its goal is not to teach you a specific software, but to enable you to quickly judge when you see these terms going forward: Ollama、oMLX、MLX、MLX-LM、llama.cpp、vLLM、KV Cache、Prefix Cache、MTP、Speculative Decoding、Serving、Runtime、Engine

"What is it? Which layer does it belong to? What am I actually doing right now? What should I do next?"


LLM Inference Engineering

Reasoning Engineering Long-term Reference Guide

Applicable Scenarios: Apple Silicon / Mac Studio / Qwen / Agentic Coding / Claude Code / Local AI


1-Minute Version: Remember This Diagram First

                         AI Application
                             │
                             ▼
                    Agent / Claude Code
                             │
                             ▼
                    API / Model Router
                             │
                             ▼
               ┌───────────────────────────┐
               │   INFERENCE ENGINEERING   │
               │       推理工程             │
               └────────────┬─────────────┘
                            │
              ┌─────────────┼─────────────┐
              ▼             ▼             ▼
         Modeling        Optimization     Reliability
           Serving        Inference        Stability
              │             │               │
              ▼             ▼               ▼
         oMLX/vLLM      KV Cache        Monitoring
         Ollama          Prefix Cache    Benchmark
         llama.cpp       MTP            24/72h test
                         Spec Decode
                         Quantization
                           ▼
                   Inference Runtime
                            │
                     ┌──────┴───────┐
                     ▼              ▼
                 MLX-LM      llama.cpp
                       │
                       ▼
                         MLX
                        │
                        ▼
                  Apple Silicon / GPU
                          │
                          ▼
                           Qwen

In One Sentence:

The model determines "what to think", Runtime handles "how to compute", the Inference Engine/Serving Layer determines "how to efficiently serve requests", and Inference Engineering is responsible for "making the entire system as good as possible".


2. The Four Concepts Most Often Confused

When in doubt about terminology going forward, return to these four concepts first.

ConceptChinese TranslationCore Question
Model模型Model(模型)What is the AI computing?
Runtime推理运行时Runtime(推理运行时)How do we run the model?
Inference Engine / Server推理引擎/服务Inference Engine / Serving(推理引擎/服务)How to efficiently serve requests?
Inference Engineering推理工程Inference Engineering(推理工程)How to make the whole system as good as possible?

3. Model — Model Layer

Examples:

  • Qwen3.8-27B
  • Qwen 35B
  • Llama
  • DeepSeek
  • Mistral

These ARE:

The models themselves.

The model determines:

  • Parameter scale
  • Architecture
  • Reasoning capability
  • Context length
  • Tokenizer
  • Model quality

Example:

Qwen3.8-27B
       │
       ▼
   Model
       │
       ▼
Requires Runtime to execute

4. Compute Framework — Computing Framework

This is a layer even lower than Runtime.

Typical:

TechnologyMain Hardware
  • MLX | Apple Silicon
  • PyTorch | NVIDIA / AMD / CPU / Apple
  • CUDA | NVIDIA
  • Metal | Apple
  • ROCm | AMD

MLX

Remember:

MLX = Machine learning computing framework for Apple Silicon

Not a replacement for Ollama. Instead, it is more like being:

MLX 
 │
 ├── Tensor Operations
 ├── GPU computation
 ├── Memory
 └── Apple Silicon optimization

5 Inference Runtime -- Inference Runtime

This IS:

The execution environment that actually runs the model.

It solves:

Model Weight
    ↓
Loading
    ↓
Computation
    ↓Token Generation
        ↓
Generation

Examples:

  • llama.cpp
  • MLX-LM
  • ONNX Runtime
  • TensorRT, etc.

6. Inference Engine —— Inference Engine

Inference Engines focus more on:

How to efficiently run large volumes of inference requests. Main problems it addresses:

  • Scheduling
  • Batching
  • Continuous Batching
  • KV Cache
  • Prefix Cache
  • Streaming
  • Concurrency
  • Memory management
  • API serving

Typical examples:

  • vLLM
  • SGLang
  • TensorRT-LLM
  • oMLX

7. Inference Server —— Inference Server

It addresses:

"How does an application access a model?" For example:

Claude Code
      │
      │ HTTP API
      ▼
Inference Server
      │
      ▼
Inference Engine
      │
      ▼
Model

Typical APIs:

POST /v1/chat/completions

Or:

Anthropic Messages API

Therefore:

The Server is the entry point through which a model exposes its services.


8. Where does Ollama actually belong?

Remember:

Ollama is an "integrated local LLM platform." It spans multiple layers:

              Ollama
                  │
       ┌──────────┼──────────┐
       ▼           ▼           ▼
 Model Manager  Runtime    Server
       │           │           │
   下载模型    执行推理     API

Its greatest advantage:

Simplicity. Suitable for:

  • Local chat
  • Quick testing
  • Development
  • Simple APIs
  • Local model management

But:

Ollama ≠ complete Inference Engineering.


9. Where does oMLX actually belong?

This is very important.

oMLX is an LLM inference server / serving system for Apple Silicon. It is built on the MLX/MLX-LM ecosystem. Roughly:

Apple Silicon
       │
       ▼
     MLX
       │
       ▼
   MLX-LM
       │
       ▼
    oMLX
       │
 ┌────┼───────────────┐
 ▼     ▼                ▼
API  KV Cache      Scheduling
     Prefix Cache  Batching
     SSD Cache

Therefore:

MLX is the foundational compute framework. MLX-LM is the LLM inference layer. oMLX is a more complete LLM inference serving system. For: Mac Studio + Qwen + Claude Code + Agentic Coding oMLX is particularly worth paying attention to.


10. What exactly is llama.cpp?

The simplest way to remember:

llama.cpp = a high-performance, cross-platform LLM inference runtime that also provides server capabilities. Therefore it spans two categories:

llama.cpp
     │
     ├── Inference Runtime
     │
     └── llama-server
             │
             └── API Server

This is why you will see different articles describe llama.cpp as:

  • Runtime
  • Inference Engine
  • Inference Server

None of these are entirely wrong.


11. What is vLLM?

Remember:

vLLM = a high-performance LLM inference engine / serving system. Key focus:

  • GPU utilization
  • Continuous batching
  • KV Cache
  • High concurrency
  • Throughput
  • Scheduling

More oriented toward:

Servers / Data Centers Rather than: Local, personal use on a Mac


12. Final Glossary

Future reference — just look here.

EnglishChineseSimple Explanation
Model模型Qwen/Llama etc.
Model Weights模型权重Model parameters
Runtime运行时Execute the model
Backend后端Underlying compute implementation
Compute Framework计算框架MLX/PyTorch
Inference Engine推理引擎Efficiently execute inference
Inference Server推理服务器Provide APIs to applications
Model Serving模型服务Provide model as a service
API EndpointAPI 接口The location where applications call the model
Context上下文The information the model currently sees
Context Window上下文窗口The maximum context the model can process
KV CacheKV 缓存Save Attention intermediate states
Prefix Cache前缀缓存Reuse repeated prompts
Quantization量化Reduce model precision/memory
MTP多 Token 预测Predict multiple next tokens at once
Speculative Decoding投机解码Small models help speed up large models
Batching批处理Batch processing

Process multiple requests at once Continuous Batching Continuous Batching Dynamic request add/exit Scheduling Scheduling Decide who runs when Routing Routing Decide which model to use Streaming Streamed Output Tokens returned one by one Throughput Throughput How many tokens generated per unit time Latency Latency How long a request takes TTFT First Token Time First Token arrival time TPOT Per Token Time Average time between tokens


13. Prefill vs Decode

This is essential for future performance optimization.

             Prompt
                │
                ▼
            PREFILL
                │
                ▼
           KV Cache
                │
                ▼
             DECODE
                │
        ┌───────┼───────┐
        ▼        ▼        ▼
     Token   Token   Token
        1        2        3

Prefill

Processing input:

"How much content did the model read?" Typically affects: TTFT


Decode

Generating output:

"How many tokens can the model generate per second?" Typically focuses on: tokens/sec / TPOT


14. KV Cache

This is one of the most core concepts in Inference Engineering. Without Cache:

Prompt
 ↓
Recompute every time
 ↓
Generate

With KV Cache:

Prompt
 ↓
Compute
 ↓
KV Cache
 ↓
Reuse on next request
 ↓
Continue Generate

Especially for Agents:

System Prompt
+
Tool Definitions
+
Project Context
+
Previous Messages

Large amounts of content will repeat. Therefore, KV Cache is extremely important.


15. Prefix Cache

Prefix Cache can be simply understood as:

If the beginning of two requests is exactly the same, do not compute it again. For example:

Request 1:
[System][Tools][Project]
                ↓
              Cache

Request 2:
[System][Tools][Project][New Task]
                ↓
          Reuse Cache

For:

Claude Code / Coding Agent It is especially valuable.


16. Speculative Decoding

Basic idea:

             Large model
             Qwen 35B
                  ▲
                  │ verify
                  │
             Small model
             Qwen 3B/7B
                  │
                  ▼
            Guess multiple Tokens

Small model:

"I guess the next is A B C D." Large model: "Let me verify." If many are guessed correctly: The large model can generate faster.


17. MTP

MTP = Multi-Token Prediction Core idea:

The model attempts to predict multiple future tokens at once. It shares a similar goal with Speculative Decoding: Improve generation speed. But the implementation mechanisms differ. In the future, when you see:

MTP
Speculative Decoding
Draft Model
Draft Tokens

You should think of:

Inference Acceleration


18. Quantization

Quantization:

FP16
 ↓
INT8
 ↓
INT4

Purpose:

  • Reduce memory
  • Increase speed
  • Allow larger models
  • Improve local deployment capabilities However:

A balance between speed, memory, and quality is required.


19. The Most Important Performance Metrics

In future benchmarks, do not just look at:

"How many tok/s?" At least look at: Metric | Significance TTFT | How long the user waits to see the first token Prefill tok/s | How fast the prompt is processed Decode tok/s | How fast it generates TPOT | Interval between tokens Throughput | Overall throughput Concurrency | Concurrent capability Memory | Memory usage Cache Hit Rate | Cache utilization efficiency Error Rate | API errors Crash Rate | Crashes Uptime | Stable running time


20. The Complete Workflow of Agentic Coding

This is the Process Map you should reference most from now on.

                   Claude Code
                        │
                        ▼
                 User Request
                        │
                        ▼
                 Context Build
                        │
           ┌────────────┴────────────┐
           ▼                          ▼
    System Prompt              Project Files
           │                          │
           └────────────┬────────────┘
                        ▼
                 Prefix Cache
                        │
                        ▼
                    Router
                        │
               ┌────────┴────────┐
               ▼                  ▼
           Qwen 27B          Qwen 35B
               │                  │
               └────────┬────────┘
                        ▼
                 Inference Server
                        │
                        ▼
                    Runtime
                        │
                        ▼
                   Hardware
                        │
                        ▼
                  Token Output
                        │
                        ▼
                   Tool Call
                        │
                        ▼
                 Execute Tool
                        │
                        ▼
                   New Result
                        │
                        ▼
                 Update Context
                        │
                        ▼
                    KV Cache
                        │
                        ▼
                 LLM Reasoning Again
                        │
                        ▼
                 ...

21. Understanding Your Local AI Stack

Your current core environment:

                  Mac Studio
                        │
                        ▼
                Apple Silicon
                        │
               ┌────────┴────────┐
               ▼                  ▼
             MLX            llama.cpp
               │
           MLX-LM
               │
               ▼
            oMLX
               │
               ▼
        Qwen3.8-27B
               │
               ▼
        Claude Code
               │
               ▼
        Agentic Coding

Also:

Ollama
   │
   └── another inference stack

Therefore what you truly need to compare is not:

"Ollama vs MLX" but rather: "How do different Inference Stacks perform under my Agentic Coding workload?"


22. Your Four Test Routes

Suggested long-term retention:

Stack A — Ollama

Claude Code
      ↓
Ollama
      ↓
Ollama Runtime
      ↓
Qwen
      ↓
Apple Silicon

Stack B — oMLX

Claude Code
      ↓
oMLX API
      ↓
oMLX
      ↓
MLX / MLX-LM
      ↓
Qwen
      ↓
Apple Silicon

Stack C — llama.cpp

Claude Code
      ↓
llama-server
      ↓
llama.cpp
      ↓
Qwen / GGUF
      ↓
Apple Silicon

Stack D — MLX-LM

Application
      ↓
MLX-LM
      ↓
MLX
      ↓
Apple Silicon

It is better suited for:

Research, experimentation, low-level tuning rather than directly as a complete Agent Server.


23. When things go wrong, first locate which layer is the issue

This is the most useful diagram for future troubleshooting.

Claude Code
      │
     X ← API failed?
      │
Inference Server
      │
     X ← Scheduling / batching?
      │
Inference Engine
      │
     X ← KV / cache?
      │
Runtime
      │
     X ← MLX / llama.cpp?
      │
Backend
      │
     X ← Metal / GPU?
      │
Hardware

If API failed

Check:

Server / API layer

If model crashes

Check:

Runtime / memory / model compatibility

If it gets slower and slower

Check:

Context / KV Cache / Memory pressure

If tok/s is low

Check:

Runtime / backend / quantization / hardware utilization

If it crashes after a long time

Check:

Memory / Cache / fragmentation / leak / server reliability


24. Troubleshooting Checklist

When you encounter:

API Failed

□ Is the Server running?
□ Is the Port correct?
□ Is the API endpoint correct?
□ Does the OpenAI/Anthropic API format match?
□ Is Streaming working correctly?
□ Is Tool calling compatible?

Crash

□ RAM exhausted?
□ Swap spiking?
□ Context too long?
□ KV Cache too large?
□ Model correctly quantized?
□ Runtime supports this model?
□ Server has crash logs?

Performance getting slower and slower

□ Context getting longer?
□ KV Cache growing?
□ Prefix Cache hit rate?
□ Memory pressure?
□ Swap usage?
□ Thermal throttling?
□ Batch size reasonable?

25. The Complete Process Map of Inference Engineering

In the future, when you don't know "where exactly should I start", follow this map:

                     ① Define Workload
                          │
                          ▼
                 ② Select Model
                          │
                          ▼
                ③ Select Runtime
                          │
                          ▼
              ④ Select Inference Engine
                          │
                          ▼
                  ⑤ Establish API
                          │
                          ▼
                  ⑥ Integrate Agent
                          │
                          ▼
                ⑦ Measure Baseline
                          │
                          ▼
               ┌──────────┴──────────┐
               ▼                      ▼
         Performance Too Slow      Unstable
               │                      │
               ▼                      ▼
        Cache / Decode          Memory
        Quantization            Context
        Batching                API
        Scheduling              Runtime
               │                      │
               └──────────┬──────────┘
                          ▼
                   ⑧ Optimize
                          │
                          ▼
                   ⑨ Benchmark
                          │
                          ▼
                   ⑩ Stress Test
                          │
                          ▼
                   ⑪ 24/72 Hour Run
                          │
                          ▼
                   ⑫ Production

26. When you see a new tool in the future, how do you determine what it is?

Ask it five questions:

Q1

Is it responsible for executing the model? If yes:

Runtime


Q2

Does it handle batching / scheduling / cache / concurrency?

If yes:

Inference Engine


Q3

Does it provide an API for applications to call?

If yes:

Inference Server


Q4

Does it handle downloading, managing, and switching models?

If yes:

Model Manager


Q5

Is it optimizing the entire system?

If yes:

Inference Engineering


27. Final Classification Quick Reference Table

Tool/Concept Most Accurate Classification How to Remember Qwen Model Model MLX Compute Framework / Backend Computing Foundation MLX-LM LLM Runtime / Library Running LLMs on MLX llama.cpp Runtime + Server Efficient Local Inference Ollama Runtime + Model Manager + Server Most Convenient oMLX Apple Silicon Inference Server/Engine High-Performance Mac Inference vLLM Inference Engine + Server High-Throughput Serving SGLang Inference Engine + Server High Performance / Agent TensorRT-LLM NVIDIA Inference Engine NVIDIA Ultimate Optimization KV Cache Optimization Avoid Redundant Computation Prefix Cache Optimization Reuse Prompt Prefixes MTP Acceleration Multi-Token Prediction Speculative Decoding Acceleration Small Model Assist Large Model Quantization Optimization Reduce Memory / Improve Efficiency Batching Serving Optimization Process Multiple Requests Together Routing System Design Decide Which Model to Use Scheduling System Design Decide Who Goes First Benchmark Engineering Test Performance Stress Test Engineering Test Stability Inference Engineering Complete Engineering System All Connected / Tied Together


28. Finally, Remember This "Brain Map"

                     │    MODEL     │
                     │    Qwen      │
                     └──────┬──────┘
                            │
                            ▼                   ┌────────────────┐                   │ Compute         │
                   │ MLX / CUDA      │                   └───────┬────────┘                         │                               ▼                  ┌──────────────────┐                  │ Runtime           │                  │ MLX-LM            │
                  │ llama.cpp         │                  └────────┬─────────┘                    │                               ▼                 ┌─────────────────────┐                │ Inference Engine     │                │ oMLX / vLLM / SGLang│                └──────────┬──────────┘                     ▲                           │                               ▼                   ┌────────────────────┐                   │ Inference Server    │                   │ API / Streaming     │                   └─────────┬──────────┘                          │                           ▼                   Claude Code                         │
                           ▼                    AI Agent
                           │
                           ▼                 ┌────────────────────┐                 │ Optimization        │                 │ KV Cache             │                 │ Prefix Cache         │                 │ Quantization         │                 │ MTP                  │                 │ Speculative Decode   │                 │ Batching              │                 └─────────┬──────────┘                           │
                           ▼                 ★ INFERENCE ENGINEERING ★
                ```

## The Bottom Line
> **Ollama, oMLX, llama.cpp, MLX-LM, and vLLM are "tools and components for building an inference system"; KV Cache, Prefix Cache, Quantization, MTP, and Speculative Decoding are "inference optimization techniques"; and combining Model, Runtime, Engine, Server, Cache, Scheduling, Hardware, Agent, Benchmark, and Reliability into a cohesive stack while continuously optimizing for your specific workload—that is what truly defines Inference Engineering.**
For your current setup of **Mac Studio + Apple Silicon + Qwen3.8-27B + Claude Code + Agentic Coding**, the most valuable mindset to establish is:
> **Never ask "Which tool is best?"**
Instead, ask:
> **"For my workload, which inference stack strikes the best balance between speed, context handling, KV cache efficiency, memory usage, API stability, agent compatibility, and 72-hour reliability?"**
This is the dividing line between simply "knowing how to run a local LLM" and actually "doing Inference Engineering".

Learning map

LLM Inference Engineering Learning Roadmap

Stage 1: Foundational Knowledge Layer

  • Understand the concept of a Model (e.g., Qwen, Llama) — what determines what an AI "thinks"
  • Master core parameters: parameter count, architecture, context length, tokenizer, and model quality
  • Understand the role of Compute Frameworks and their mapping to hardware

Stage 2: Runtime & Engine

  • Understand Inference Runtime (e.g., MLX-LM, llama.cpp) — the execution environment for running model computations
  • Learn the concept of an Inference Engine — how vLLM and oMLX efficiently serve batched inference requests
  • Master the concepts of the Serving/Server API layer (OpenAI-compatible interface, streaming output)

Stage 3: Optimization Techniques Layer

  • KV Cache principles and the reuse mechanism of Prefix Cache
  • Quantization (impact of INT8/INT4 quantization on performance)
  • Speculative Decoding (leveraging small models to assist large models in accelerating generation)
  • MTP (Multi-Token Prediction) and Continuous Batching
  • Differences between the Prefill and Decode phases

Stage 4: Performance Evaluation & Tuning

  • Key Metrics System: TTFT, TPOT, Decode tok/s, Throughput, Cache Hit Rate
  • Monitoring & Observability — logs, metrics, tracing
  • Building Benchmark and Stress Test workflows

Stage 5: System Design & Reliability

  • Agent Inference Workflows and Routing / Scheduling design
  • Long-term operational stability (24–72h stress testing)
  • Layered Troubleshooting Methodology — pinpointing whether an issue lies in the API, runtime, engine, or hardware layer

Get hands-on — step by step

  1. Install a local LLM Runtime environment Choose an inference stack and install it (MLX/oMLX recommended for Mac Studio, llama.cpp recommended for generic systems)

  2. Download and load an open-source model Obtain weight files for models such as Qwen3.8-27B via Hugging Face or locally

  3. Start an inference API Server Run ollama serve, oMLX, or llama-server to expose a /v1/chat/completions endpoint

  4. Write a basic inference calling script Send the first prompt to the local server via the OpenAI-compatible API format and verify the output

  5. Enable KV Cache monitoring Configure cache metrics for oMLX/llama.cpp and observe the Cache Hit Rate and Prefix Cache hit situations

  6. Experiment with the impact of different quantization precisions on performance Load the same model using FP16, INT8, and INT4 respectively, and compare Token/s, memory usage, and output quality

  7. Enable Streaming mode Set the API call parameter stream=true and observe how TTFT (Time To First Token) changes

  8. Design an Agent inference workflow Simulate a Claude Code-like multi-turn interaction—incorporating context construction, Tool Call, and KV Cache reuse

  9. Build a performance benchmark suite Record TTFT, Decode tok/s, Throughput, and concurrency capacity across different scenarios

  10. Conduct a long-duration stress test Run the inference service over a 24–72 hour cycle and monitor memory growth, cache fragmentation, and crash rate

Top 3 sources

  1. 1
    Hugging Face Transformers 官方文档

    大规模预训练模型推理与微调的基础文档,涵盖核心概念和 API 使用指南。

    https://huggingface.co/docs/transformers

  2. 2
    llama.cpp 项目文档

    本地 LLM 推理的核心运行时工具库,支持跨平台 GPU/CPU 高效加载与推理模型。

    https://github.com/ggerganov/llama.cpp

  3. 3
    MLX 官方开发者指南 (Apple Silicon)

    Apple Silicon 机器学习计算框架的官方文档,涵盖张量操作、GPU 计算和 MLX-LM 推理入门。

    https://ml-explore.github.io/mlx/build/html/overview.html

Links are AI-suggested — worth a quick sanity check before diving in.