๐ŸฅŠ Framework Showdown: Ollama vs LM Studio vs llama.cpp vs vLLM

This is Part 2 of the Self-Hosting LLMs series. Part 1 covers why self-hosting is worth it.


Youโ€™ve decided to run an LLM locally. Youโ€™ve read about the cost savings and privacy benefits. Now comes the practical question: which tool do you actually use?

The ecosystem has converged around four main options, and theyโ€™re not interchangeable โ€” each fills a different niche. Picking the wrong one wonโ€™t break anything, but picking the right one saves you hours of frustration.

This post gives you a hands-on tour of all four, with real benchmarks, configuration examples, and a decision framework at the end.


๐Ÿ—บ๏ธ 1. The Landscape at a Glance

Before diving deep, hereโ€™s how these four relate to each other:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                   Experience Layers                  โ”‚
โ”‚                                                     โ”‚
โ”‚     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”             โ”‚
โ”‚     โ”‚  Ollama   โ”‚          โ”‚ LM Studio โ”‚             โ”‚
โ”‚     โ”‚  (CLI)    โ”‚          โ”‚   (GUI)   โ”‚             โ”‚
โ”‚     โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜          โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜             โ”‚
โ”‚          โ”‚                      โ”‚                    โ”‚
โ”‚          โ–ผ                      โ–ผ                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚         llama.cpp / MLX              โ”‚  โ—€โ”€ Engine โ”‚
โ”‚  โ”‚     (C++ inference runtime)          โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ”‚                                                     โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚  โ”‚              vLLM                    โ”‚  โ—€โ”€ Server โ”‚
โ”‚  โ”‚     (Python production server)       โ”‚            โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Ollama and LM Studio are experience layers โ€” they wrap an engine (llama.cpp, and on Apple Silicon, MLX) and make it easy to use. llama.cpp is the engine itself โ€” maximum control, maximum portability. vLLM is a different beast entirely โ€” a production serving system designed for throughput, not convenience.

Understanding this hierarchy saves confusion: Ollama isnโ€™t competing with llama.cpp, itโ€™s built on llama.cpp.


๐Ÿฆ™ 2. Ollama โ€” The Docker of LLMs

What it is: A single Go binary that wraps llama.cpp (and MLX on Apple Silicon) with a model registry, automatic GPU detection, and an OpenAI-compatible REST API.

Best for: Solo developers who want to get running in under 2 minutes.

Install and Run

# Install (macOS / Linux)
curl -fsSL https://ollama.com/install.sh | sh

# Pull and run a model
ollama run qwen3:8b

Thatโ€™s it. Ollama detects your GPU (NVIDIA, AMD, or Apple Silicon), downloads the model, loads it, and gives you a chat prompt. No Python environment, no CUDA toolkit, no config files.

The Model Registry

Ollama maintains a curated library of 500+ pre-quantized models at ollama.com/library. Models follow a Docker-like naming convention:

ollama pull llama3.1:8b              # default quant (Q4_K_M)
ollama pull llama3.1:8b-q5_K_M       # specific quant
ollama pull qwen3:32b                # larger model
ollama pull deepseek-r1:14b          # reasoning model

This is one of Ollamaโ€™s biggest advantages โ€” you donโ€™t need to browse Hugging Face, figure out which GGUF file to download, or worry about tokenizer configs. It just works.

OpenAI-Compatible API

Every Ollama install exposes a REST API on port 11434 that mirrors the OpenAI chat completions format:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",
)

response = client.chat.completions.create(
    model="qwen3:8b",
    messages=[{"role": "user", "content": "Explain PagedAttention in 3 sentences."}],
)

print(response.choices[0].message.content)

This means any tool that supports the OpenAI API works with Ollama โ€” LangChain, LlamaIndex, Continue.dev, Open WebUI, and hundreds more. Same code, different URL.

Modelfiles โ€” Custom Model Variants

Ollamaโ€™s Modelfile system lets you create custom model configurations:

FROM qwen3:8b

SYSTEM """You are a senior Python developer. Always include type hints,
write clean code, and explain your reasoning step by step."""

PARAMETER temperature 0.3
PARAMETER num_ctx 8192
ollama create my-coder -f Modelfile
ollama run my-coder

This bakes a system prompt, temperature, and context length into a named model variant you can share or reuse.

GPU Management

Ollama handles GPU detection automatically:

  • NVIDIA: Uses CUDA, detects multi-GPU setups
  • AMD: ROCm support for RX 6000/7000 series
  • Apple Silicon: Uses Metal (and MLX since Ollama 0.19+)
  • Fallback: Automatically offloads to CPU if VRAM is insufficient

You can check how a model is loaded:

ollama ps
# NAME         ID         SIZE      PROCESSOR    UNTIL
# qwen3:8b     a6990ed6   5.2 GB    100% GPU     4 minutes from now

Limitations

  • Single-user optimized: Under concurrent load, throughput drops significantly (~41 tok/s vs vLLMโ€™s ~793 tok/s in benchmarks)
  • Limited quantization control: You get what the registry offers; custom GGUF files require a Modelfile workaround
  • No tensor parallelism: Canโ€™t split a model across multiple GPUs for inference (it can use multiple GPUs, but loads complete model copies)

๐Ÿ–ฅ๏ธ 3. LM Studio โ€” The GUI-First Experience

What it is: A desktop application (Electron-based) with a built-in Hugging Face model browser, chat interface, and local API server.

Best for: People who prefer clicking over typing, or who want to visually browse and compare models.

The Model Browser

LM Studioโ€™s standout feature is its Discover tab โ€” a visual model browser connected directly to Hugging Face. You search for a model, see:

  • Compatibility ratings based on your hardware
  • File sizes for each quantization level
  • Download counts and community ratings
  • VRAM requirements with your specific GPU highlighted

Then you click Download. No command line, no manual file management.

Chat Interface

The Chat tab feels like using ChatGPT, but everything runs locally. You can:

  • Adjust temperature, top-p, and context length with sliders
  • Compare outputs from different models side-by-side
  • Save and organize conversation histories
  • Upload documents for in-context processing (RAG)

Developer Mode โ€” Local API Server

LM Studio can start a local server on port 1234 thatโ€™s OpenAI API-compatible:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="lm-studio",
)

response = client.chat.completions.create(
    model="qwen3-8b",
    messages=[{"role": "user", "content": "Hello!"}],
)

It also ships lms, a CLI tool, and llmster, a headless daemon for running models without the GUI.

When LM Studio Shines

  • Model discovery: The visual browser is genuinely useful when youโ€™re exploring options
  • Quick A/B testing: Load two models, send the same prompt, compare results visually
  • Non-terminal users: If your workflow is GUI-native, LM Studio fits naturally
  • Document chat: Built-in RAG for chatting with local files

Limitations

  • Electron overhead: Heavier on system resources than Ollamaโ€™s single binary
  • Smaller model registry: Relies on Hugging Face directly โ€” more models available, but less curation than Ollamaโ€™s registry
  • macOS and Windows only: No official Linux support (though community workarounds exist)
  • Closed source: Unlike the other three tools, LM Studio itself is not open source

โš™๏ธ 4. llama.cpp โ€” The Engine Under the Hood

What it is: A C/C++ implementation of LLM inference, originally created by Georgi Gerganov. Itโ€™s the engine that powers Ollama, LM Studio, and much of the local LLM ecosystem.

Best for: Power users who want maximum control, embedded deployments, or need to run on unusual hardware.

Why llama.cpp Matters

llama.cpp is the reason local LLMs are practical. Before it existed (March 2023), running a language model required Python, PyTorch, and a CUDA-capable GPU. llama.cpp proved you could run Llama models in pure C++ on a MacBook CPU โ€” and the local LLM movement exploded from there.

Today the project has 118K+ GitHub stars, 450+ contributors, and ships new builds almost daily.

Direct Usage

# Build from source
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON  # or DGGML_METAL=ON for Mac
cmake --build build --config Release

# Run a model
./build/bin/llama-cli \
    -m models/qwen3-8b-q4_K_M.gguf \
    -p "Explain the difference between TCP and UDP" \
    -n 256

Server Mode

llama.cpp includes a built-in HTTP server with an impressive feature set:

./build/bin/llama-server \
    -m models/qwen3-8b-q4_K_M.gguf \
    --host 0.0.0.0 \
    --port 8080 \
    --n-gpu-layers 99 \
    --ctx-size 8192 \
    --parallel 4

Server mode features include:

  • OpenAI + Anthropic-compatible API routes
  • Parallel decoding for multiple concurrent requests
  • Function calling support
  • Speculative decoding for 2โ€“3x faster inference with a draft model
  • Grammar-constrained generation โ€” force outputs to follow a JSON schema or other grammar
  • Embeddings and reranking endpoints
  • Router mode for loading multiple models with LRU eviction
  • Built-in web UI with chat interface

Speculative Decoding

One of llama.cppโ€™s most powerful features โ€” use a small โ€œdraftโ€ model to propose tokens, then verify them with the main model:

./build/bin/llama-speculative \
    -m models/qwen3-8b-q4_K_M.gguf \
    -md models/qwen3-0.6b-q8_0.gguf \
    -p "Write a Python HTTP server" \
    --draft-max 8

This can achieve 2โ€“3x speedup on generation with minimal quality impact, because the small modelโ€™s correct predictions are kept and only wrong ones are regenerated.

Grammar-Constrained Generation

Force the model to output valid JSON, SQL, or any format defined by a GBNF grammar:

./build/bin/llama-cli \
    -m models/qwen3-8b-q4_K_M.gguf \
    --grammar-file grammars/json.gbnf \
    -p "List 3 programming languages with their year of creation"

This guarantees syntactically valid output โ€” something you canโ€™t easily get from cloud APIs without post-processing.

When to Use llama.cpp Directly

  • Embedded systems: Runs on Raspberry Pi, Android, iOS, WebAssembly
  • Custom builds: Need specific CUDA flags, quantization options, or experimental features
  • Speculative decoding: Ollama doesnโ€™t expose this (yet)
  • Grammar-constrained output: When you need guaranteed-valid structured output
  • Maximum performance: Direct llama.cpp can squeeze out more tok/s than the Ollama wrapper

Limitations

  • More setup: You need to build from source (or download pre-built binaries), manage GGUF files yourself, and configure options manually
  • No model registry: You download models from Hugging Face yourself
  • GGUF only: Doesnโ€™t support GPTQ or AWQ formats (use vLLM for those)

๐Ÿš€ 5. vLLM โ€” The Production Powerhouse

What it is: A Python-based LLM serving system built for maximum throughput. Uses PagedAttention, continuous batching, and tensor parallelism to serve many concurrent users efficiently.

Best for: Team deployments, production APIs, and any scenario where you need to serve multiple users simultaneously.

The Throughput Gap

Hereโ€™s why vLLM exists โ€” single-user performance is roughly the same across all tools (~30 tok/s on an RTX 4090 with a 24B model). But under concurrent load:

Tool Single user 10 concurrent users 50 concurrent users
Ollama ~30 tok/s ~41 tok/s total Queues requests
llama.cpp server ~30 tok/s ~120 tok/s total ~200 tok/s total
vLLM ~30 tok/s ~400 tok/s total ~793 tok/s total

vLLM achieves 16โ€“20x Ollamaโ€™s concurrent throughput. If youโ€™re building something that more than one person will use at the same time, this matters enormously.

How It Achieves This

PagedAttention โ€” Instead of pre-allocating contiguous memory for each requestโ€™s KV cache, vLLM borrows the concept of virtual memory paging from operating systems. Memory is divided into fixed-size blocks allocated on-demand, reducing waste by up to 4x and enabling much larger batch sizes.

Continuous batching โ€” Traditional serving waits for an entire batch to finish before starting a new one. vLLM uses rolling batches โ€” as soon as one sequence finishes, a new request immediately takes its slot. GPU utilization stays high even with variable-length outputs.

Tensor parallelism โ€” Split a model across 2, 4, or 8 GPUs transparently. A 70B model that doesnโ€™t fit on one GPU runs across two with near-linear scaling.

Getting Started

pip install vllm

# Start the server
vllm serve meta-llama/Llama-3.1-8B-Instruct \
    --dtype auto \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.9

vLLMโ€™s API is also OpenAI-compatible:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="vllm",
)

response = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Hello!"}],
)

Multi-GPU Deployment

# Tensor parallel across 2 GPUs
vllm serve meta-llama/Llama-3.1-70B-Instruct \
    --tensor-parallel-size 2 \
    --dtype auto \
    --quantization awq

Quantization Support

vLLM supports quantization formats optimized for GPU throughput:

  • GPTQ โ€” calibrated 4-bit quantization with Marlin kernels (~5x faster than naive INT4)
  • AWQ โ€” activation-aware quantization, good quality-speed balance
  • FP8 โ€” 8-bit floating point, minimal quality loss, supported on H100/RTX 4090

Note: vLLM does not support GGUF format. If you need GGUF, use Ollama or llama.cpp.

When to Use vLLM

  • Multi-user serving: Internal team tools, API endpoints, chatbots with concurrent users
  • Maximum throughput: Batch processing, document pipelines, high-volume inference
  • Multi-GPU setups: When your model needs more than one GPU
  • Production deployments: When uptime, monitoring, and scaling matter

Limitations

  • Slow cold start: Loading models takes minutes (compiles CUDA kernels on first run)
  • GPU required: No CPU fallback โ€” you need NVIDIA (or AMD ROCm) GPUs
  • Heavier setup: Python environment, more dependencies, more configuration
  • Overkill for personal use: If itโ€™s just you querying a model, Ollama is simpler and just as fast

๐Ÿ“Š 6. Head-to-Head Comparison

Feature Ollama LM Studio llama.cpp vLLM
Setup time 2 min 5 min 15-30 min 10-20 min
Interface CLI + API GUI + API CLI + API API
Model format GGUF (auto) GGUF (auto) GGUF GPTQ, AWQ, FP8
Model registry Built-in (500+) HuggingFace browser Manual download HuggingFace (auto)
GPU support NVIDIA, AMD, Apple NVIDIA, Apple NVIDIA, AMD, Apple, Vulkan NVIDIA, AMD
CPU inference Yes (fallback) Yes (fallback) Yes (primary feature) No
Multi-GPU Limited No Layer splitting Tensor parallelism
Concurrent users Poor (~41 tok/s) Poor Moderate (~200 tok/s) Excellent (~793 tok/s)
Speculative decoding No No Yes (2-3x speedup) Yes
Grammar constraints No No Yes Yes (guided decoding)
OpenAI API Yes Yes Yes Yes
Open source Yes (MIT) No Yes (MIT) Yes (Apache 2.0)
Best for Getting started Visual exploration Power users, embedded Production serving

๐ŸŽฏ 7. Decision Framework

Start Here

Are you a solo developer getting started with local LLMs?
  โ””โ”€โ†’ Ollama. Done. You can always switch later.

Do you prefer a GUI over terminal?
  โ””โ”€โ†’ LM Studio for browsing and chatting.
      Still use Ollama for API integration.

Are you building something multiple people will use?
  โ””โ”€โ†’ vLLM. The throughput difference is massive.

Do you need structured output, speculative decoding,
or deployment on unusual hardware?
  โ””โ”€โ†’ llama.cpp directly.

The Progression Most People Follow

  1. Start with Ollama โ€” get your first model running, experiment with different sizes
  2. Try LM Studio โ€” visually browse models, compare outputs side-by-side
  3. Hit a limit โ€” you need more speed, structured output, or multi-user support
  4. Graduate to llama.cpp or vLLM โ€” depending on whether you need control or throughput

Most developers stay at step 1 or 2 for a long time. Thatโ€™s fine. The experience layers exist precisely so you donโ€™t need to think about the engine.

Can You Mix Them?

Absolutely. A common setup:

  • Ollama on your laptop for development and quick queries
  • vLLM on a GPU server for your teamโ€™s shared inference endpoint
  • llama.cpp in a CI pipeline for structured test output with grammar constraints

The OpenAI-compatible API standard means your application code doesnโ€™t care which backend is running. Change the URL, keep the code.


๐Ÿ”ฎ Whatโ€™s Next

This series covers everything from motivation to model selection:

Next up: if youโ€™ve picked your framework and want to understand which model to run on it โ€” head to Part 4. If you want to understand the quantization settings you keep seeing (Q4_K_M, Q5_K_S, GPTQ), Part 3 breaks it all down.


Questions or corrections? Leave a comment below, or find me on GitHub at @scrowten.




    Enjoy Reading This Article?

    Here are some more articles you might like to read next:

  • ๐Ÿ—บ๏ธ Open Source Model Landscape: Benchmarks and Picking the Right One
  • ๐Ÿ”ญ GitSanity: arxiv-sanity, but for GitHub
  • ๐Ÿ–ฅ๏ธ Terminal System Monitors: Why I Keep Coming Back to htop
  • ๐Ÿค– Choosing Your Personal AI Assistant: Why I Landed on NanoClaw
  • ๐Ÿ  Why Run Your Own LLM? The Case for Self-Hosting in 2026