Why Run Locally?
Running models on your own machine gives you capabilities that cloud APIs cannot match.
Local data control
Prompts can stay on your device when the application and selected features make no external calls. Check logging, telemetry, integrations and cloud options separately.
Zero API Costs
After the one-time hardware investment, every token is free. Run as many queries as you want.
Offline Access
Works without internet. Use AI on planes, in secure environments, or anywhere connectivity is limited.
Full Customization
Choose any model, any quantization, any parameters. Fine-tune for your specific use case.
Deep Learning
Nothing teaches you how LLMs work like running and experimenting with them directly.
Total Control
No rate limits, no content filters you did not choose, no surprise API changes or deprecations.
Hardware Requirements
Select a model size and quantization level to see how much VRAM you need and which GPUs can handle it.
Start with storage arithmetic. A hypothetical 70B model needs at least 35 GB (32.6 GiB) for raw 4-bit weights or 70 GB (65.2 GiB) for 8-bit weights. Quantization metadata, KV cache and runtime buffers come on top. A 24 GB GPU therefore needs partial offload or a smaller model.
VRAM CalculatorThe MoE Advantage for Local Inference
Mixture of Experts (MoE) models route each token through only a subset of "expert" layers. The key advantage is speed: fewer active parameters means faster generation. But all parameters still live in VRAM — MoE does not save memory.
Faster Generation Speed
Only selected experts compute for each token. This can reduce arithmetic relative to a dense model with the same total parameters, but does not by itself predict tokens per second.
Large-Model Intelligence
All potentially selected expert weights must remain accessible. Total parameters govern stored weights; active parameters govern only part of per-token work.
VRAM Is Still Based on Total Params
All expert weights must be loaded into memory. Mixtral 8x7B at Q4 needs ~26 GB VRAM — similar to a dense 30B model, not a 13B. MoE saves compute, not memory.
Active parameters govern only part of the per-token work. Stored weights include every expert. Change the two counts to see why a small active count does not imply a small model file. Routing changes with each token; inactive experts can be needed later.
MoE separates stored parameters from per-token computation. All experts contribute to storage even though only a subset is selected for a particular token. Offloading adds transfers or CPU work, so benchmark the actual placement.
MoE is a fundamental architecture shift, not just an optimization trick. Understanding how expert routing works helps you pick the right model for your hardware.
Deep dive into Mixture of Experts →Popular Tools
The local inference ecosystem has matured rapidly. Here are the tools that matter, from beginner-friendly to production-grade.
Choose a runtime by platform, model format and device support. Ease-of-use or speed scores need a defined workload, so no universal rating is shown.
Ollama: local model runner with a CLI and HTTP API. Check model licensing, downloads, device support and whether any selected feature calls a cloud service.
Sourcesllama.cpp: GGUF inference with explicit device and offload controls. The CLI help documents --gpu-layers, --cpu-moe and --n-cpu-moe. --override-kv changes model metadata, not expert prefetching.
SourcesLM Studio: desktop interface for local inference. Check the supported model formats and operating-system requirements in its documentation.
SourcesThe Quantization Tradeoff
Quantization is the key technology that makes local inference practical. By reducing the precision of model weights, you can fit much larger models into limited VRAM.
A hypothetical 70B model requires 140 GB for raw FP16 weights, 70 GB for 8-bit weights, or 35 GB for raw 4-bit weights. Quantization metadata, mixed tensor types, KV cache and runtime buffers add memory. Actual quality and speed depend on the chosen quantizer and runtime.
Deep dive into Quantization →VRAM Calculator →
Not sure if a model fits your GPU? Calculate VRAM requirements and estimated speed for any model and quantization level.
Getting Started
Follow these five steps to go from zero to running your first local model.
Pick a Tool
Start with Ollama or LM Studio -- they handle everything for you. Move to llama.cpp or vLLM when you need more control.
Check Your VRAM
Run nvidia-smi (NVIDIA) or check Activity Monitor (Mac). This determines what models you can run.
Choose a Model Size
Start with 7B models. They are fast, capable, and fit on most GPUs. Move to 13B or 70B as you need more capability.
Pick a Quantization Level
Q4 is the sweet spot for most users: good quality with reasonable VRAM use. Go Q8 if you have the memory, Q2 if you are tight.
Run It
Download the model and start chatting. With Ollama: ollama pull llama3.2 then ollama run llama3.2. That is it.
Quickstart Demo
Here is what it looks like to install Ollama and run your first model -- three commands and you are chatting.
Install a runtime using its instructions for your operating system. Start with a small model, inspect available memory, then measure the actual run. These are command examples, not a simulated terminal session.
ollama --help ollama list ollama ps llama-cli --helpOllama
Tips and Tricks
- 1Context length directly impacts VRAM usage. A 7B model with 128K context needs significantly more memory than with 4K context. Start small and increase as needed.
- 2GPU offloading lets you split a model between GPU and CPU. You get GPU speed for the layers that fit, with CPU handling the rest. Slower than full GPU, but runs larger models.
- 3CPU, GPU and accelerator speed depend on the workload and backend. Apple Silicon can run GPU inference through Metal with unified memory; that is not the same as unusually fast CPU-only inference.
- 4Size the exact model file, context and runtime reserve against available memory. A 70B 4-bit model cannot be held entirely on a 24 GB GPU. Running it requires offload or additional devices. Even 32 GB is below its raw 35 GB weight footprint.
- 5Llama 3.2, Mistral, Phi-3, and Qwen 2.5 are excellent choices for local inference. Each excels at different tasks -- experiment to find your best fit.
- 6Run models as an API server (Ollama and LM Studio both support this) to integrate local models into your own applications, scripts, and workflows.