Running Models Locally

Beginner

Run large language models on your own hardware -- no cloud, no API keys, no limits.

Last updated: Sep 13, 2026

Why Run Locally?

Running models on your own machine gives you capabilities that cloud APIs cannot match.

Local data control

Prompts can stay on your device when the application and selected features make no external calls. Check logging, telemetry, integrations and cloud options separately.

Zero API Costs

After the one-time hardware investment, every token is free. Run as many queries as you want.

Offline Access

Works without internet. Use AI on planes, in secure environments, or anywhere connectivity is limited.

Full Customization

Choose any model, any quantization, any parameters. Fine-tune for your specific use case.

Deep Learning

Nothing teaches you how LLMs work like running and experimenting with them directly.

Total Control

No rate limits, no content filters you did not choose, no surprise API changes or deprecations.

Hardware Requirements

Select a model size and quantization level to see how much VRAM you need and which GPUs can handle it.

Start with storage arithmetic. A hypothetical 70B model needs at least 35 GB (32.6 GiB) for raw 4-bit weights or 70 GB (65.2 GiB) for 8-bit weights. Quantization metadata, KV cache and runtime buffers come on top. A 24 GB GPU therefore needs partial offload or a smaller model.

VRAM Calculator

The MoE Advantage for Local Inference

Mixture of Experts (MoE) models route each token through only a subset of "expert" layers. The key advantage is speed: fewer active parameters means faster generation. But all parameters still live in VRAM — MoE does not save memory.

Faster Generation Speed

Only selected experts compute for each token. This can reduce arithmetic relative to a dense model with the same total parameters, but does not by itself predict tokens per second.

Large-Model Intelligence

All potentially selected expert weights must remain accessible. Total parameters govern stored weights; active parameters govern only part of per-token work.

VRAM Is Still Based on Total Params

All expert weights must be loaded into memory. Mixtral 8x7B at Q4 needs ~26 GB VRAM — similar to a dense 30B model, not a 13B. MoE saves compute, not memory.

Active parameters govern only part of the per-token work. Stored weights include every expert. Change the two counts to see why a small active count does not imply a small model file. Routing changes with each token; inactive experts can be needed later.

Raw 4-bit weight storage
35.0 GB
Raw active weights per token
5.0 GB

MoE separates stored parameters from per-token computation. All experts contribute to storage even though only a subset is selected for a particular token. Offloading adds transfers or CPU work, so benchmark the actual placement.

MoE is a fundamental architecture shift, not just an optimization trick. Understanding how expert routing works helps you pick the right model for your hardware.

Deep dive into Mixture of Experts →

Popular Tools

The local inference ecosystem has matured rapidly. Here are the tools that matter, from beginner-friendly to production-grade.

Choose a runtime by platform, model format and device support. Ease-of-use or speed scores need a defined workload, so no universal rating is shown.

Ollama: local model runner with a CLI and HTTP API. Check model licensing, downloads, device support and whether any selected feature calls a cloud service.

Sources

llama.cpp: GGUF inference with explicit device and offload controls. The CLI help documents --gpu-layers, --cpu-moe and --n-cpu-moe. --override-kv changes model metadata, not expert prefetching.

Sources

LM Studio: desktop interface for local inference. Check the supported model formats and operating-system requirements in its documentation.

Sources

The Quantization Tradeoff

Quantization is the key technology that makes local inference practical. By reducing the precision of model weights, you can fit much larger models into limited VRAM.

A hypothetical 70B model requires 140 GB for raw FP16 weights, 70 GB for 8-bit weights, or 35 GB for raw 4-bit weights. Quantization metadata, mixed tensor types, KV cache and runtime buffers add memory. Actual quality and speed depend on the chosen quantizer and runtime.

Deep dive into Quantization →
🧮

VRAM Calculator →

Not sure if a model fits your GPU? Calculate VRAM requirements and estimated speed for any model and quantization level.

Getting Started

Follow these five steps to go from zero to running your first local model.

1

Pick a Tool

Start with Ollama or LM Studio -- they handle everything for you. Move to llama.cpp or vLLM when you need more control.

2

Check Your VRAM

Run nvidia-smi (NVIDIA) or check Activity Monitor (Mac). This determines what models you can run.

3

Choose a Model Size

Start with 7B models. They are fast, capable, and fit on most GPUs. Move to 13B or 70B as you need more capability.

4

Pick a Quantization Level

Q4 is the sweet spot for most users: good quality with reasonable VRAM use. Go Q8 if you have the memory, Q2 if you are tight.

5

Run It

Download the model and start chatting. With Ollama: ollama pull llama3.2 then ollama run llama3.2. That is it.

Quickstart Demo

Here is what it looks like to install Ollama and run your first model -- three commands and you are chatting.

Install a runtime using its instructions for your operating system. Start with a small model, inspect available memory, then measure the actual run. These are command examples, not a simulated terminal session.

ollama --help
ollama list
ollama ps

llama-cli --help
Ollama

Tips and Tricks

  • 1Context length directly impacts VRAM usage. A 7B model with 128K context needs significantly more memory than with 4K context. Start small and increase as needed.
  • 2GPU offloading lets you split a model between GPU and CPU. You get GPU speed for the layers that fit, with CPU handling the rest. Slower than full GPU, but runs larger models.
  • 3CPU, GPU and accelerator speed depend on the workload and backend. Apple Silicon can run GPU inference through Metal with unified memory; that is not the same as unusually fast CPU-only inference.
  • 4Size the exact model file, context and runtime reserve against available memory. A 70B 4-bit model cannot be held entirely on a 24 GB GPU. Running it requires offload or additional devices. Even 32 GB is below its raw 35 GB weight footprint.
  • 5Llama 3.2, Mistral, Phi-3, and Qwen 2.5 are excellent choices for local inference. Each excels at different tasks -- experiment to find your best fit.
  • 6Run models as an API server (Ollama and LM Studio both support this) to integrate local models into your own applications, scripts, and workflows.