What Is a Local LLM? How to Run AI Models Locally
Say you want to run an LLM on your own machine — a local LLM, in the community's vocabulary. The reasons are usually one of three: you're pasting things into a chat box that shouldn't leave your device, you want it to work on a plane or in a lab with no internet, or your API bill has started to look like a car payment. All three are legitimate. Then you download your first model, ask it something, and hit the wall: it's slower than the cloud, and the answers are… noticeably worse than what you get from a paid plan.
That gap — between what the benchmark promised and what your GPU delivered — is the actual subject of this guide. Because the honest answer about a local LLM — about any local AI model you can actually download — is that they work, they're getting better fast, and almost every beginner is silently running a degraded version of the model they think they downloaded. I'll get to that. First, the basics, then the hardware math, then the tools, then the part nobody tells you about.
What you're actually deploying
Every local LLM deployment has three layers, and "deployment" means all of them:
The weights — the model itself, which is just a large file. A 7-billion-parameter model at 4-bit precision is about 4 GB on disk. You download it once, like a movie.
The runtime — the software that loads those weights into memory and computes answers. Ollama, llama.cpp, LM Studio, vLLM are all runtimes (or wrappers around one). This layer decides how fast you go and how much memory gets used.
The interface — a terminal window counts, but most people eventually want a chat UI or an API endpoint so other apps can call the model.
Here's the cost framing most tutorials skip: downloading the weights removes the vendor's token fee, and that's the only line item it removes. The machine, the electricity, the storage, the maintenance, and your own time all remain — as one hardware-tracking site puts it, downloaded weights do not make inference free. Local deployment trades a recurring bill for a capital cost plus friction. Whether that trade wins depends entirely on your usage pattern, which is why the "when should you" question comes later in this piece with actual criteria.
What your computer can actually run
VRAM is the hard wall, so let's do the arithmetic once. Rough formula: parameter count × bits per parameter ÷ 8 ≈ weights in GB. A 27B model at 4-bit is about 13.5 GB of weights — decimal, weights only.
That "weights only" qualifier matters, because it's where beginner frustration is manufactured. On top of the weights you need quantization metadata, the non-quantized layers, runtime buffers, the KV cache (the model's working memory for your conversation — the longer the context, the bigger it gets, and we have a whole piece on the context window if you want that mechanism), and whatever your operating system is already using. A 27B model that "fits" in 16 GB of weights realistically wants a 24 GB-class card to breathe.
The practical tiers:
| Memory you have | What runs comfortably |
|---|---|
| 8 GB | 3B–7B models at 4-bit. Usable for chat and summarizing, rough for anything serious |
| 16 GB | 7B–13B at Q4/Q8; 20B-class MoE models (with a caveat below) |
| 24 GB | 27B at 4-bit — where "this is genuinely good" starts |
| 48 GB+ / multi-GPU | 70B-class, or big context windows on 27B |
Two exceptions to memorize. Apple Silicon: unified memory means the GPU shares system RAM, so a 48 GB Mac can hold models that no single consumer video card can. It works well — llama.cpp supports Metal and CPU-GPU hybrid inference — but budget memory for macOS itself and treat CPU offload as part of your tested setup, not invisible extra capacity. Mixture-of-Experts models: when OpenAI says GPT-OSS 20B "runs within 16 GB," the 20B is total parameters; only 3.6B are active per token, which explains the modest compute, but all experts still have to sit in memory. Active parameters explain speed, not storage. The 16 GB figure is also a provider claim for one specific quantization format — not a promise about every runtime and context length.
And one more reality check while we're doing arithmetic: model files are not the same as their memory needs. A weights file that looks like it fits can blow past your card once you raise the context length, because the KV cache grows with it. Ollama exposes the actual processor split via ollama ps — trust that output over any chart, including mine.
Three tool routes, by what you want
Beginner guides love listing six tools. In practice the choice collapses to one of three routes:
Route 1: LM Studio — you want a GUI. Install an app, search for a model, click download, chat. It also runs a local API server compatible with OpenAI's format, so your code can point at it later. If the terminal scares you, start here and stop feeling bad about it.
Route 2: Ollama — you want the ecosystem. One command installs it; ollama run llama3 downloads a model and drops you into a chat. It serves a local API on port 11434, speaks the OpenAI format, and has become the default backend that other local LLM tools integrate with — 9M+ installs a month at this point. Pair it with Open WebUI (one Docker command) and you get a ChatGPT-style interface over your own models. This is the route I'd recommend to most people reading this.
Route 3: vLLM — you're serving users, not just yourself. It's a production inference engine built for throughput: one benchmark shows Llama 70B under concurrent load at 793 tokens/second on vLLM versus 41 on Ollama. That number is meaningless for a single user chatting alone — the gap only matters once multiple requests hit the model simultaneously. If "local deployment" for you means a team server, vLLM (or SGLang) is the layer you'll grow into.
Underneath all of these sits llama.cpp, the C/C++ engine most of the ecosystem is built on. You don't need to touch it directly — knowing it exists just helps the rest of the stack make sense.
Quantization: the trick that makes it possible and makes it worse
Here's the part almost every tutorial mentions in one paragraph and moves on: the model you download is almost never the model that was benchmarked.
Providers publish scores from full-precision (BF16) reference implementations running on datacenter hardware. What lands on your disk is a quantized copy — weights compressed from 16-bit to 4-bit or 8-bit so they fit in your VRAM. Running a local LLM means living with this trade: quantization buys size and speed with precision, and "precision" here isn't an abstract term: each generated token is a probability distribution over the vocabulary, and compression nudges that distribution. Nudge it enough and the top candidate flips — the model starts choosing different words. One flipped token early in a long answer cascades, because everything after it is conditioned on it.
A hardware forum ran the experiments I wish every guide would run. Same model, same weights, changing one variable at a time. Just swapping the attention backend — nothing else — caused visible disagreement between runs at longer context lengths. Quantizing only the KV cache (weights untouched): 8-bit degraded and recovered; 4-bit produced a tool call that failed and never recovered. A five-way bakeoff of weight quantizations ended with a vendor-published FP4 package hitting roughly 50% next-token flips at 88k context — dead last — while both 4-bit entries botched a command-line task the full-precision model handled. Meanwhile a community INT8 build quietly beat the official FP8 one, because quantization quality depends on calibration and on which layers get excluded, not on the vendor's badge.
Two takeaways for you. First, the effect is distance-dependent: short chats and one-shot questions look nearly identical across quantizations; the damage shows up in long contexts, multi-step tool calling, and agent workflows. Second, don't trust KLD or "near-lossless" claims on a model card without the methodology — measurement setup changes the number dramatically.
For everyday use the community's settled default is Q4_K_M — the balance point validated across enormous amounts of usage. Q8 is near-indistinguishable from full precision when you have the memory for it. And a free debugging tip from that same forum thread: if your model loops endlessly in its "thinking" output, your temperature is probably set too low. I've seen that one waste an afternoon.
When local actually makes sense — and when it doesn't
The decision isn't ideological. Deploy locally when one of these is true:
The data genuinely can't leave. Contracts, medical records, internal documents, source code with restrictions. If the constraint is real, no cloud feature is worth violating it — and note there's a middle tier here: many companies land on private-cloud deployments before anyone needs a GPU under a desk.
You need it offline. Planes, field work, air-gapped machines, flaky hotel Wi-Fi. The model doesn't care; it's already on your disk.
Your volume makes the meter spin. If you're running millions of tokens a day through an API — batch processing, a busy internal tool — hardware amortizes. Compute your electricity against the token bill; the crossover exists and it's findable.
And the honest counter-list, when you shouldn't bother:
- You need top-tier output quality. The community's own dividing line, from people running this daily: below roughly the 27B class, results are "still very limited." If quality is the product, you want the big cloud models or big local hardware.
- You're running long-context agents. The exact workload where quantization damage surfaces. A quantized model that nails quick Q&A can fail a 60k-token tool-calling chain.
- You don't want a hobby. Drivers, memory settings, quant formats, context tuning. This is maintenance whether you enjoy it or not.
- Your usage is light. Twenty questions a day costs pennies over API. The local setup "wins" that comparison only if your time is free.
One more use case deserves the spotlight because it's the most practical of all: feeding your own documents to a model that stays on your machine. That's a RAG setup — retrieval over your files, generation by the local model — and it's the configuration where local deployment goes from "neat" to "I use this weekly."
What people actually run down there
Numbers from the community, not marketing. A teacher with a 3090 runs an entire assessment pipeline locally — generating quizzes for 500-student mock exams, the correction tooling, all of it — on a 27B-class model, and puts the dividing line bluntly: below that size, "very limited." Another user runs Qwen-class 27B at 4-bit with a 96k context squeezed into a 24 GB card, 43 tokens/second, driving a coding agent inside VS Code. On the other end, the person who started a now-famous "let's be honest about local LLMs" thread — laptop, 8 GB GPU — concluded it was mostly a hobby for them. That's fine too. Hobbies are allowed.
If you want to know what it takes to run an LLM locally on your exact hardware, skip the calculators and look at measured reports: there's a community site, vram.wiki, that collects real setups — GPU, model, quantization, context, tokens-per-second, honest limitations — with a strict rule of "unknown stays unknown." 150+ entries last I checked. Seeing five people run the model you're considering on the card you own is worth more than any spec sheet. And if multimodal input matters to you — dropping images or audio into the same model — check which model families carry that capability before you commit hardware; the small multimodal options are a different shortlist than the text-only one.
Quick answers
Is running a local LLM free? The models are (most popular ones carry Apache 2.0 licenses). The inference isn't: electricity, hardware wear, storage, and your setup time are all real costs. You're trading a token bill for a hardware bill.
Can a local LLM replace ChatGPT? For specific jobs — privacy-bound tasks, offline work, high-volume simple tasks — yes, and sometimes it's the better tool. As a general-purpose quality match for frontier cloud models, not yet, especially below the 27B class.
Is 8 GB of VRAM enough? For 3B–7B quantized models: yes, for light tasks. For the experience you've seen in demos: no. That's the tier where "hobby" is the honest word.
Can a Mac run a local LLM? Genuinely well, yes — it's a fine local LLM machine — Apple Silicon's unified memory and Metal support make Macs some of the best consumer hardware for this, provided you respect the memory macOS itself needs.
What's Ollama, in one sentence? The tool that made local models a one-command install — it downloads, runs, and serves models on your machine with an OpenAI-compatible API, and it's the default backend most local AI tools now expect.
The meta-advice, if you take one thing: benchmark claims are about a model you don't have, on hardware you don't own. Before you run an LLM locally, check the hardware requirements you actually measured, not the ones a model card advertised. Run your own tasks, on your own quant, at your own context length — that's the only score that describes the LLM in front of you.