BrainBank

Combines 3 small local LLMs, with reasoning capability rivaling Anthropic Fable 5

9/8/2026, 10:08:20 PM · Source

AI-translated on 9/9/2026, 12:04:43 AM · by Qwen3.6 35B (fast, default)

#step-by-step#llm#local-ai#logit-fusion#pytorch#model-merging

This paper introduces a dynamic model fusion method based on logit layers that achieves reasoning performance on par with frontier models—without modifying model weights—by running three specialized small LLMs in parallel on local GPUs and blending their outputs through weighted averaging.

This paper introduces a dynamic model fusion method at the logit layer — by running three specialized small LLMs in parallel on a local GPU and weighting their logit outputs together before sampling, it synergizes different models' advantages in grammar, creativity, and logical reasoning. By avoiding modifications to model weights, this approach also addresses data privacy, inference efficiency, and local deployment costs, offering a new paradigm for building high-performance private AI inference systems on consumer hardware.

imageimage

1. The High-End Subscription Trap

Every time you send a prompt, you pay $0.12 to a centralized API just so a closed-source model can parse simple customer logs — and you constantly worry it might hallucinate. It's slow, expensive, almost absurd. The industry wants you to believe that only through massive expenditures can you access cutting-edge intelligence. But this isn't the case.

Most developers look at lightweight local models like Llama 3 8B or Gemma 2 9B and think they're merely "toys." They run a few standard prompts, see the model occasionally make errors in logical reasoning, and immediately return to expensive cloud APIs. But what if we don't have to choose between the high cost of centralized servers and the performance limits of local devices?

Here's something rarely told: solving highly complex reasoning tasks doesn't require a massive 400-billion parameter model entirely. By fusing the probability distributions generated at the sampler level by three specialized small models, we can achieve near-leading-edge inference precision rivaling multi-agent frameworks like Anthropic Fable 5.

Below, I demonstrate exactly how this works at the low level.

# img

imageimage

2. The Jazz Trio in the Soundproofed Basement

Most tutorials stop at simple prompt routing — sending different queries to different API boxes. Don't do that; this approach is essentially just a very slow, fragile workaround. But if we want to truly understand the underlying mechanism of token generation, we must go deeper than token probability distributions. Standard models output probability distributions as logit arrays. If we load three small models simultaneously into your GPU memory, they don't have to execute the same prompt sequentially. Instead, their forward passes run in parallel, and based on task requirements, vector weights are assigned to outputs from different models, with unified token masks applied.

When a model computes its next-token vectors, we intercept the logits before sampling occurs. We apply weighted averaging across their probability matrices — multiplying each model's output vector by specific task-dependent weights and applying a combined token mask. The three models vote on the next character at the hardware level; without modifying any underlying parameters, their grammar and logic strengths are integrated in our fusion mechanism.

# img

imageimage

imageimage

imageimage

imageimage

Section 3: Avoiding Destructive Parameter Merging

You've probably already seen static model merging on Hugging Face — using tools like mergekit with spherical linear interpolation (SLERP) or task vectors, permanently fusing different models' weights.

Many tutorials recommend this approach. Don't adopt it so quickly.

The hidden problem is that static weight fusion can be highly destructive — when you permanently fuse the parameter matrices of two neural networks, the mathematical structures that give each model its unique capabilities might be flattened or diluted. Eventually, models may lose their inherent advantages and exhibit subtle "cognitive decay" that standard benchmarks struggle to capture.

Runtime dynamic fusion bypasses this limitation entirely. We retain each model's specialized parameters intact in isolated GPU buffers, allowing every model to evaluate the current context at its native capability. For example: the structured code model handles grammar analysis, the creative model judges narrative style, and the logic model cross-verifies facts. A logit router serves as an arbiter, achieving output precision comparable to top-tier closed-source models without parameter degradation.

90% of people get stuck here: writing a complex Python wrapper that loses milliseconds querying a server pool sequentially. If you want to build a truly responsive local system, we can write a lightweight PyTorch custom sampler and perform the math at the logit layer entirely within local GPU memory.

Next, let's look at a concrete implementation: using Python to fuse probability distributions generated by three local models, enabling high-quality handling of structured technical tasks.

# Dynamic Logit-Level Ensemble Fusion Controller

Wait — before going further, let's examine the math in this loop. It introduces no added latency. On modern consumer GPUs, running three small models in parallel takes less time than invoking a single massive cloud model through a standard network round-trip. The result: absolute data privacy, zero subscription fees, and extremely low-latency local inference performance. 图片图片 图片图片

Original title: I Fused 3 Tiny Local LLMs on my Laptop and Matched the Reasoning of Anthropic Fable 5 Original link: https://pub.towardsai.net/i-fused-3-tiny-local-llms-on-my-laptop-and-matched-the-reasoning-of-anthropic-fable-5-4e62930b2bf0

Learning map

Phase I: Understanding Local Large Language Models and Efficiency Pain Points

  • Understand the advantages and limitations of running lightweight models (such as Llama 3 8B, Gemma 2 9B) on consumer-grade hardware.
  • Grasp why expensive cloud APIs are not the only option for handling all complex tasks.

Phase II: Understanding Logit-Level Dynamic Fusion Principles

  • Dive into the mathematical underpinnings of token generation and probability distributions (logits).
  • Compare static weight merging (e.g., mergekit, SLERP) with dynamic runtime fusion to avoid cognitive degradation.

Phase III: Building a Local Logit Routing and Parallel Inference System

  • Learn how to run forward passes from multiple small models in parallel within GPU VRAM.
  • Write a custom sampler to implement weighted averaging across the logit layer and real-time arbitration.

Get hands-on — step by step

  1. Set up the local runtime environment and install PyTorch along with the dependency libraries needed for multi-model parallel inference.
  2. Download three specialized lightweight local LLMs tailored for different tasks (e.g., grammar, creativity, logical reasoning).
  3. Write a Python script to load these three models into isolated GPU buffers and implement parallel forward passes.
  4. Intercept the raw logit arrays output by the models and implement weighted averaging and composite token masking calculations before sampling.
  5. Run the custom sampler to test dynamic ensemble fusion effects, comparing against single-model outputs to verify improvements in inference performance.

Top 3 sources

  1. 1
    PyTorch Documentation

    官方 PyTorch 文档,是实现自定义采样器和张量数学运算的核心参考指南。

    https://pytorch.org/docs/stable/index.html

  2. 2
    Hugging Face Transformers

    提供开源模型加载、分词器管理以及底层 Logit 访问的标准库文档。

    https://huggingface.co/docs/transformers/index

  3. 3
    Towards AI Publication

    本文出处,汇集了大量前沿 AI 技术解析与本地部署实战案例。

    https://pub.towardsai.net/

Links are AI-suggested — worth a quick sanity check before diving in.