LLM Inference Engineering Quick Reference Guide
8/29/2026, 11:24:10 AM · updated 8/29/2026, 11:36:36 AM
AI-translated on 8/29/2026, 11:40:11 AM · by Qwen3.6 35B (fast, default)
This article systematically maps out the core conceptual framework of LLM inference engineering, helping readers quickly identify terminology across each layer — models, runtimes, inference engines, and servers — and master the underlying principles and practical applications of optimization techniques such as KV Cache, speculative decoding, and quantization.
Its goal is not to teach you a specific software, but to enable you to quickly judge when you see these terms going forward: Ollama、oMLX、MLX、MLX-LM、llama.cpp、vLLM、KV Cache、Prefix Cache、MTP、Speculative Decoding、Serving、Runtime、Engine
"What is it? Which layer does it belong to? What am I actually doing right now? What should I do next?"
LLM Inference Engineering
Reasoning Engineering Long-term Reference Guide
Applicable Scenarios: Apple Silicon / Mac Studio / Qwen / Agentic Coding / Claude Code / Local AI
1-Minute Version: Remember This Diagram First
AI Application
│
▼
Agent / Claude Code
│
▼
API / Model Router
│
▼
┌───────────────────────────┐
│ INFERENCE ENGINEERING │
│ 推理工程 │
└────────────┬─────────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Modeling Optimization Reliability
Serving Inference Stability
│ │ │
▼ ▼ ▼
oMLX/vLLM KV Cache Monitoring
Ollama Prefix Cache Benchmark
llama.cpp MTP 24/72h test
Spec Decode
Quantization
▼
Inference Runtime
│
┌──────┴───────┐
▼ ▼
MLX-LM llama.cpp
│
▼
MLX
│
▼
Apple Silicon / GPU
│
▼
Qwen
In One Sentence:
The model determines "what to think", Runtime handles "how to compute", the Inference Engine/Serving Layer determines "how to efficiently serve requests", and Inference Engineering is responsible for "making the entire system as good as possible".
2. The Four Concepts Most Often Confused
When in doubt about terminology going forward, return to these four concepts first.
| Concept | Chinese Translation | Core Question |
|---|---|---|
| Model模型 | Model(模型) | What is the AI computing? |
| Runtime推理运行时 | Runtime(推理运行时) | How do we run the model? |
| Inference Engine / Server推理引擎/服务 | Inference Engine / Serving(推理引擎/服务) | How to efficiently serve requests? |
| Inference Engineering推理工程 | Inference Engineering(推理工程) | How to make the whole system as good as possible? |
3. Model — Model Layer
Examples:
- Qwen3.8-27B
- Qwen 35B
- Llama
- DeepSeek
- Mistral
These ARE:
The models themselves.
The model determines:
- Parameter scale
- Architecture
- Reasoning capability
- Context length
- Tokenizer
- Model quality
Example:
Qwen3.8-27B
│
▼
Model
│
▼
Requires Runtime to execute
4. Compute Framework — Computing Framework
This is a layer even lower than Runtime.
Typical:
| Technology | Main Hardware |
|---|
- MLX | Apple Silicon
- PyTorch | NVIDIA / AMD / CPU / Apple
- CUDA | NVIDIA
- Metal | Apple
- ROCm | AMD
MLX
Remember:
MLX = Machine learning computing framework for Apple Silicon
Not a replacement for Ollama. Instead, it is more like being:
MLX
│
├── Tensor Operations
├── GPU computation
├── Memory
└── Apple Silicon optimization
5 Inference Runtime -- Inference Runtime
This IS:
The execution environment that actually runs the model.
It solves:
Model Weight
↓
Loading
↓
Computation
↓Token Generation
↓
Generation
Examples:
- llama.cpp
- MLX-LM
- ONNX Runtime
- TensorRT, etc.
6. Inference Engine —— Inference Engine
Inference Engines focus more on:
How to efficiently run large volumes of inference requests. Main problems it addresses:
- Scheduling
- Batching
- Continuous Batching
- KV Cache
- Prefix Cache
- Streaming
- Concurrency
- Memory management
- API serving
Typical examples:
- vLLM
- SGLang
- TensorRT-LLM
- oMLX
7. Inference Server —— Inference Server
It addresses:
"How does an application access a model?" For example:
Claude Code
│
│ HTTP API
▼
Inference Server
│
▼
Inference Engine
│
▼
Model
Typical APIs:
POST /v1/chat/completions
Or:
Anthropic Messages API
Therefore:
The Server is the entry point through which a model exposes its services.
8. Where does Ollama actually belong?
Remember:
Ollama is an "integrated local LLM platform." It spans multiple layers:
Ollama
│
┌──────────┼──────────┐
▼ ▼ ▼
Model Manager Runtime Server
│ │ │
下载模型 执行推理 API
Its greatest advantage:
Simplicity. Suitable for:
- Local chat
- Quick testing
- Development
- Simple APIs
- Local model management
But:
Ollama ≠ complete Inference Engineering.
9. Where does oMLX actually belong?
This is very important.
oMLX is an LLM inference server / serving system for Apple Silicon. It is built on the MLX/MLX-LM ecosystem. Roughly:
Apple Silicon
│
▼
MLX
│
▼
MLX-LM
│
▼
oMLX
│
┌────┼───────────────┐
▼ ▼ ▼
API KV Cache Scheduling
Prefix Cache Batching
SSD Cache
Therefore:
MLX is the foundational compute framework. MLX-LM is the LLM inference layer. oMLX is a more complete LLM inference serving system. For: Mac Studio + Qwen + Claude Code + Agentic Coding oMLX is particularly worth paying attention to.
10. What exactly is llama.cpp?
The simplest way to remember:
llama.cpp = a high-performance, cross-platform LLM inference runtime that also provides server capabilities. Therefore it spans two categories:
llama.cpp
│
├── Inference Runtime
│
└── llama-server
│
└── API Server
This is why you will see different articles describe llama.cpp as:
- Runtime
- Inference Engine
- Inference Server
None of these are entirely wrong.
11. What is vLLM?
Remember:
vLLM = a high-performance LLM inference engine / serving system. Key focus:
- GPU utilization
- Continuous batching
- KV Cache
- High concurrency
- Throughput
- Scheduling
More oriented toward:
Servers / Data Centers Rather than: Local, personal use on a Mac
12. Final Glossary
Future reference — just look here.
| English | Chinese | Simple Explanation |
|---|---|---|
| Model | 模型 | Qwen/Llama etc. |
| Model Weights | 模型权重 | Model parameters |
| Runtime | 运行时 | Execute the model |
| Backend | 后端 | Underlying compute implementation |
| Compute Framework | 计算框架 | MLX/PyTorch |
| Inference Engine | 推理引擎 | Efficiently execute inference |
| Inference Server | 推理服务器 | Provide APIs to applications |
| Model Serving | 模型服务 | Provide model as a service |
| API Endpoint | API 接口 | The location where applications call the model |
| Context | 上下文 | The information the model currently sees |
| Context Window | 上下文窗口 | The maximum context the model can process |
| KV Cache | KV 缓存 | Save Attention intermediate states |
| Prefix Cache | 前缀缓存 | Reuse repeated prompts |
| Quantization | 量化 | Reduce model precision/memory |
| MTP | 多 Token 预测 | Predict multiple next tokens at once |
| Speculative Decoding | 投机解码 | Small models help speed up large models |
| Batching | 批处理 | Batch processing |
Process multiple requests at once Continuous Batching Continuous Batching Dynamic request add/exit Scheduling Scheduling Decide who runs when Routing Routing Decide which model to use Streaming Streamed Output Tokens returned one by one Throughput Throughput How many tokens generated per unit time Latency Latency How long a request takes TTFT First Token Time First Token arrival time TPOT Per Token Time Average time between tokens
13. Prefill vs Decode
This is essential for future performance optimization.
Prompt
│
▼
PREFILL
│
▼
KV Cache
│
▼
DECODE
│
┌───────┼───────┐
▼ ▼ ▼
Token Token Token
1 2 3
Prefill
Processing input:
"How much content did the model read?" Typically affects: TTFT
Decode
Generating output:
"How many tokens can the model generate per second?" Typically focuses on: tokens/sec / TPOT
14. KV Cache
This is one of the most core concepts in Inference Engineering. Without Cache:
Prompt
↓
Recompute every time
↓
Generate
With KV Cache:
Prompt
↓
Compute
↓
KV Cache
↓
Reuse on next request
↓
Continue Generate
Especially for Agents:
System Prompt
+
Tool Definitions
+
Project Context
+
Previous Messages
Large amounts of content will repeat. Therefore, KV Cache is extremely important.
15. Prefix Cache
Prefix Cache can be simply understood as:
If the beginning of two requests is exactly the same, do not compute it again. For example:
Request 1:
[System][Tools][Project]
↓
Cache
Request 2:
[System][Tools][Project][New Task]
↓
Reuse Cache
For:
Claude Code / Coding Agent It is especially valuable.
16. Speculative Decoding
Basic idea:
Large model
Qwen 35B
▲
│ verify
│
Small model
Qwen 3B/7B
│
▼
Guess multiple Tokens
Small model:
"I guess the next is A B C D." Large model: "Let me verify." If many are guessed correctly: The large model can generate faster.
17. MTP
MTP = Multi-Token Prediction Core idea:
The model attempts to predict multiple future tokens at once. It shares a similar goal with Speculative Decoding: Improve generation speed. But the implementation mechanisms differ. In the future, when you see:
MTP
Speculative Decoding
Draft Model
Draft Tokens
You should think of:
Inference Acceleration
18. Quantization
Quantization:
FP16
↓
INT8
↓
INT4
Purpose:
- Reduce memory
- Increase speed
- Allow larger models
- Improve local deployment capabilities However:
A balance between speed, memory, and quality is required.
19. The Most Important Performance Metrics
In future benchmarks, do not just look at:
"How many tok/s?" At least look at: Metric | Significance TTFT | How long the user waits to see the first token Prefill tok/s | How fast the prompt is processed Decode tok/s | How fast it generates TPOT | Interval between tokens Throughput | Overall throughput Concurrency | Concurrent capability Memory | Memory usage Cache Hit Rate | Cache utilization efficiency Error Rate | API errors Crash Rate | Crashes Uptime | Stable running time
20. The Complete Workflow of Agentic Coding
This is the Process Map you should reference most from now on.
Claude Code
│
▼
User Request
│
▼
Context Build
│
┌────────────┴────────────┐
▼ ▼
System Prompt Project Files
│ │
└────────────┬────────────┘
▼
Prefix Cache
│
▼
Router
│
┌────────┴────────┐
▼ ▼
Qwen 27B Qwen 35B
│ │
└────────┬────────┘
▼
Inference Server
│
▼
Runtime
│
▼
Hardware
│
▼
Token Output
│
▼
Tool Call
│
▼
Execute Tool
│
▼
New Result
│
▼
Update Context
│
▼
KV Cache
│
▼
LLM Reasoning Again
│
▼
...
21. Understanding Your Local AI Stack
Your current core environment:
Mac Studio
│
▼
Apple Silicon
│
┌────────┴────────┐
▼ ▼
MLX llama.cpp
│
MLX-LM
│
▼
oMLX
│
▼
Qwen3.8-27B
│
▼
Claude Code
│
▼
Agentic Coding
Also:
Ollama
│
└── another inference stack
Therefore what you truly need to compare is not:
"Ollama vs MLX" but rather: "How do different Inference Stacks perform under my Agentic Coding workload?"
22. Your Four Test Routes
Suggested long-term retention:
Stack A — Ollama
Claude Code
↓
Ollama
↓
Ollama Runtime
↓
Qwen
↓
Apple Silicon
Stack B — oMLX
Claude Code
↓
oMLX API
↓
oMLX
↓
MLX / MLX-LM
↓
Qwen
↓
Apple Silicon
Stack C — llama.cpp
Claude Code
↓
llama-server
↓
llama.cpp
↓
Qwen / GGUF
↓
Apple Silicon
Stack D — MLX-LM
Application
↓
MLX-LM
↓
MLX
↓
Apple Silicon
It is better suited for:
Research, experimentation, low-level tuning rather than directly as a complete Agent Server.
23. When things go wrong, first locate which layer is the issue
This is the most useful diagram for future troubleshooting.
Claude Code
│
X ← API failed?
│
Inference Server
│
X ← Scheduling / batching?
│
Inference Engine
│
X ← KV / cache?
│
Runtime
│
X ← MLX / llama.cpp?
│
Backend
│
X ← Metal / GPU?
│
Hardware
If API failed
Check:
Server / API layer
If model crashes
Check:
Runtime / memory / model compatibility
If it gets slower and slower
Check:
Context / KV Cache / Memory pressure
If tok/s is low
Check:
Runtime / backend / quantization / hardware utilization
If it crashes after a long time
Check:
Memory / Cache / fragmentation / leak / server reliability
24. Troubleshooting Checklist
When you encounter:
API Failed
□ Is the Server running?
□ Is the Port correct?
□ Is the API endpoint correct?
□ Does the OpenAI/Anthropic API format match?
□ Is Streaming working correctly?
□ Is Tool calling compatible?
Crash
□ RAM exhausted?
□ Swap spiking?
□ Context too long?
□ KV Cache too large?
□ Model correctly quantized?
□ Runtime supports this model?
□ Server has crash logs?
Performance getting slower and slower
□ Context getting longer?
□ KV Cache growing?
□ Prefix Cache hit rate?
□ Memory pressure?
□ Swap usage?
□ Thermal throttling?
□ Batch size reasonable?
25. The Complete Process Map of Inference Engineering
In the future, when you don't know "where exactly should I start", follow this map:
① Define Workload
│
▼
② Select Model
│
▼
③ Select Runtime
│
▼
④ Select Inference Engine
│
▼
⑤ Establish API
│
▼
⑥ Integrate Agent
│
▼
⑦ Measure Baseline
│
▼
┌──────────┴──────────┐
▼ ▼
Performance Too Slow Unstable
│ │
▼ ▼
Cache / Decode Memory
Quantization Context
Batching API
Scheduling Runtime
│ │
└──────────┬──────────┘
▼
⑧ Optimize
│
▼
⑨ Benchmark
│
▼
⑩ Stress Test
│
▼
⑪ 24/72 Hour Run
│
▼
⑫ Production
26. When you see a new tool in the future, how do you determine what it is?
Ask it five questions:
Q1
Is it responsible for executing the model? If yes:
Runtime
Q2
Does it handle batching / scheduling / cache / concurrency?
If yes:
Inference Engine
Q3
Does it provide an API for applications to call?
If yes:
Inference Server
Q4
Does it handle downloading, managing, and switching models?
If yes:
Model Manager
Q5
Is it optimizing the entire system?
If yes:
Inference Engineering
27. Final Classification Quick Reference Table
Tool/Concept Most Accurate Classification How to Remember Qwen Model Model MLX Compute Framework / Backend Computing Foundation MLX-LM LLM Runtime / Library Running LLMs on MLX llama.cpp Runtime + Server Efficient Local Inference Ollama Runtime + Model Manager + Server Most Convenient oMLX Apple Silicon Inference Server/Engine High-Performance Mac Inference vLLM Inference Engine + Server High-Throughput Serving SGLang Inference Engine + Server High Performance / Agent TensorRT-LLM NVIDIA Inference Engine NVIDIA Ultimate Optimization KV Cache Optimization Avoid Redundant Computation Prefix Cache Optimization Reuse Prompt Prefixes MTP Acceleration Multi-Token Prediction Speculative Decoding Acceleration Small Model Assist Large Model Quantization Optimization Reduce Memory / Improve Efficiency Batching Serving Optimization Process Multiple Requests Together Routing System Design Decide Which Model to Use Scheduling System Design Decide Who Goes First Benchmark Engineering Test Performance Stress Test Engineering Test Stability Inference Engineering Complete Engineering System All Connected / Tied Together
28. Finally, Remember This "Brain Map"
│ MODEL │
│ Qwen │
└──────┬──────┘
│
▼ ┌────────────────┐ │ Compute │
│ MLX / CUDA │ └───────┬────────┘ │ ▼ ┌──────────────────┐ │ Runtime │ │ MLX-LM │
│ llama.cpp │ └────────┬─────────┘ │ ▼ ┌─────────────────────┐ │ Inference Engine │ │ oMLX / vLLM / SGLang│ └──────────┬──────────┘ ▲ │ ▼ ┌────────────────────┐ │ Inference Server │ │ API / Streaming │ └─────────┬──────────┘ │ ▼ Claude Code │
▼ AI Agent
│
▼ ┌────────────────────┐ │ Optimization │ │ KV Cache │ │ Prefix Cache │ │ Quantization │ │ MTP │ │ Speculative Decode │ │ Batching │ └─────────┬──────────┘ │
▼ ★ INFERENCE ENGINEERING ★
```
## The Bottom Line
> **Ollama, oMLX, llama.cpp, MLX-LM, and vLLM are "tools and components for building an inference system"; KV Cache, Prefix Cache, Quantization, MTP, and Speculative Decoding are "inference optimization techniques"; and combining Model, Runtime, Engine, Server, Cache, Scheduling, Hardware, Agent, Benchmark, and Reliability into a cohesive stack while continuously optimizing for your specific workload—that is what truly defines Inference Engineering.**
For your current setup of **Mac Studio + Apple Silicon + Qwen3.8-27B + Claude Code + Agentic Coding**, the most valuable mindset to establish is:
> **Never ask "Which tool is best?"**
Instead, ask:
> **"For my workload, which inference stack strikes the best balance between speed, context handling, KV cache efficiency, memory usage, API stability, agent compatibility, and 72-hour reliability?"**
This is the dividing line between simply "knowing how to run a local LLM" and actually "doing Inference Engineering".
Learning map
LLM Inference Engineering Learning Roadmap
Stage 1: Foundational Knowledge Layer
- Understand the concept of a Model (e.g., Qwen, Llama) — what determines what an AI "thinks"
- Master core parameters: parameter count, architecture, context length, tokenizer, and model quality
- Understand the role of Compute Frameworks and their mapping to hardware
Stage 2: Runtime & Engine
- Understand Inference Runtime (e.g., MLX-LM, llama.cpp) — the execution environment for running model computations
- Learn the concept of an Inference Engine — how vLLM and oMLX efficiently serve batched inference requests
- Master the concepts of the Serving/Server API layer (OpenAI-compatible interface, streaming output)
Stage 3: Optimization Techniques Layer
- KV Cache principles and the reuse mechanism of Prefix Cache
- Quantization (impact of INT8/INT4 quantization on performance)
- Speculative Decoding (leveraging small models to assist large models in accelerating generation)
- MTP (Multi-Token Prediction) and Continuous Batching
- Differences between the Prefill and Decode phases
Stage 4: Performance Evaluation & Tuning
- Key Metrics System: TTFT, TPOT, Decode tok/s, Throughput, Cache Hit Rate
- Monitoring & Observability — logs, metrics, tracing
- Building Benchmark and Stress Test workflows
Stage 5: System Design & Reliability
- Agent Inference Workflows and Routing / Scheduling design
- Long-term operational stability (24–72h stress testing)
- Layered Troubleshooting Methodology — pinpointing whether an issue lies in the API, runtime, engine, or hardware layer
Get hands-on — step by step
-
Install a local LLM Runtime environment Choose an inference stack and install it (MLX/oMLX recommended for Mac Studio, llama.cpp recommended for generic systems)
-
Download and load an open-source model Obtain weight files for models such as Qwen3.8-27B via Hugging Face or locally
-
Start an inference API Server Run ollama serve, oMLX, or llama-server to expose a /v1/chat/completions endpoint
-
Write a basic inference calling script Send the first prompt to the local server via the OpenAI-compatible API format and verify the output
-
Enable KV Cache monitoring Configure cache metrics for oMLX/llama.cpp and observe the Cache Hit Rate and Prefix Cache hit situations
-
Experiment with the impact of different quantization precisions on performance Load the same model using FP16, INT8, and INT4 respectively, and compare Token/s, memory usage, and output quality
-
Enable Streaming mode Set the API call parameter stream=true and observe how TTFT (Time To First Token) changes
-
Design an Agent inference workflow Simulate a Claude Code-like multi-turn interaction—incorporating context construction, Tool Call, and KV Cache reuse
-
Build a performance benchmark suite Record TTFT, Decode tok/s, Throughput, and concurrency capacity across different scenarios
-
Conduct a long-duration stress test Run the inference service over a 24–72 hour cycle and monitor memory growth, cache fragmentation, and crash rate
Top 3 sources
- 1Hugging Face Transformers 官方文档
大规模预训练模型推理与微调的基础文档,涵盖核心概念和 API 使用指南。
https://huggingface.co/docs/transformers
- 2
- 3MLX 官方开发者指南 (Apple Silicon)
Apple Silicon 机器学习计算框架的官方文档,涵盖张量操作、GPU 计算和 MLX-LM 推理入门。
https://ml-explore.github.io/mlx/build/html/overview.html
Links are AI-suggested — worth a quick sanity check before diving in.