~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / how-to-run-llm-locally $
Tutorials

How to Run an LLM Locally: Step by Step with Ollama in 2 Commands (No GPU)

LS Lucas Souza · · 9 min read
How to Run an LLM Locally: Step by Step with Ollama in 2 Commands (No GPU)

You can run a decent language model on your laptop. Without sending a single byte to the cloud, without an API key, without a bill at the end of the month. The model downloads, loads on your machine, and answers right there, on localhost.

The question is no longer "can you run an LLM locally?". You can. The question is: when does it make sense, what actually runs on your hardware, and how do you get it up without suffering through drivers, quantization, and blown VRAM.

In this guide you'll install Ollama, run your first local model in two commands, understand the hardware math (how much RAM/VRAM each model size needs), and learn to decide, case by case, when local beats the API and when it doesn't.

TL;DR

  • What it is: running an LLM locally means executing the model on your own machine, without calling an external API. The text never leaves your hardware.
  • Stack: Ollama (runner + server), open models like Llama, Qwen, Gemma, gpt-oss, DeepSeek.
  • Cost/Access: free and open source. The cost is the hardware you already have (or are going to buy).
  • Minimum: 8 GB of RAM, 10 GB of free disk, and a 64-bit CPU with AVX2, with no GPU required. A GPU speeds things up a lot, but it's not a prerequisite to get started.

If your question is the math rather than the setup (how much it costs to run local versus paying for an API in production, with a month-by-month break-even), see Local LLM vs API in 2026: the real 90-day spreadsheet. Here the focus is getting the model running on your machine.

The context: why "local LLM" is a topic again

Two years ago, running a local model was a hobby for people with an expensive graphics card and the patience to compile llama.cpp by hand. That changed.

First, open models closed the gap. A 7–8B Qwen or Llama today handles a good chunk of day-to-day work (summarization, extraction, classification, code drafts) at a level that not long ago only a frontier API delivered. Second, the tooling got dead simple to use: Ollama hides the annoying part (quantization, GPU acceleration, HTTP server) behind a single command.

And then there's the timing. Ollama v0.32 shipped on July 11, with a built-in interactive agent mode (chat, code, web search, and task delegation straight from the terminal), flash attention for older NVIDIA GPUs (compute capability 6.x), and vision offload on iGPUs. Translation: running local got faster precisely on modest hardware, which is where most people are.

If you want the cold math on when open source really pays off, with cost, privacy, and speed side by side, we already covered that in Are open source AI models worth it in 2026. Here the focus is hands-on: get it running and use it.

Prerequisites

Before you download anything:

  • [ ] RAM: 8 GB is the floor for a small model (1–3B). 16 GB opens the door to 7–8B models with room to spare.
  • [ ] Disk: 10 GB free to start. Each model weighs from ~2 GB (quantized 3B) to dozens of GB (70B+).
  • [ ] GPU (optional, but it changes everything): any NVIDIA/AMD card with decent VRAM, or a Mac with Apple Silicon (unified memory is great for this). Without a GPU, it runs on the CPU, just slower.
  • [ ] Terminal: you're going to live in it. No IDE setup.

A rule of thumb for sizing: in the original format (FP16), the model needs about 2 GB of VRAM per billion parameters. An 8B in FP16 is ~16 GB. That's where quantization comes in, and it's what makes local viable on your laptop.

Hands-on: from zero to first token

Step 1: install Ollama

On Linux, one command:

curl -fsSL https://ollama.com/install.sh | sh

On macOS and Windows, download the installer at ollama.com/download. It starts a local server listening on http://localhost:11434. Remember that port, it matters in a moment.

Step 2: run your first model

One command downloads it and opens the chat:

ollama run qwen3:8b

If the model isn't on your machine yet, Ollama pulls it on its own (about 5 GB) and opens an interactive prompt. Type your question and that's it: you're talking to an LLM that runs 100% on your machine. To exit, /bye.

If 8B is too heavy for your hardware, start smaller:

ollama run qwen3:1.7b

Step 3: understand quantization (the part that saves your RAM)

Quantizing means reducing the precision of the model's weights so it takes up less memory. The gold standard on consumer hardware is Q4_K_M: it compresses the weights to 4 bits and cuts ~75% of the VRAM compared to FP16, with a quality loss you barely notice in practice. That 16 GB 8B becomes ~5–6 GB.

The math by hardware tier looks like this (reference):

VRAM Runs well (Q4_K_M)
6–8 GB 7–8B models
16 GB 13B models
24 GB 32B models
48 GB+ 70B+ models

A 7–8B in Q4_K_M delivers 40+ tokens per second on an entry-level GPU, fast enough that you're not left waiting. And watch the context: the KV cache eats extra memory as the window grows (about 2.5 GB more on a 7B with 32k of context, ~10 GB at 128k). Giant context isn't free, even locally.

Step 4: plug it into your code

The terminal is the start. What matters for building a product is the API. Ollama exposes an OpenAI-compatible endpoint at /v1/chat/completions, meaning the same SDK you'd use for GPT, pointed at your localhost:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # ignored, but the SDK requires something
)

resp = client.chat.completions.create(
    model="qwen3:8b",
    messages=[{"role": "user", "content": "Summarize this text in 3 bullets: ..."}],
)
print(resp.choices[0].message.content)

Swap base_url and model, and the rest of the code stays the same. That's what makes local an architecture decision, not a toy: you plug it into the same place where you call the cloud today.

Useful commands for managing what's on your machine: ollama list (what you've downloaded), ollama ps (what's loaded in memory right now), and ollama rm <model> to free up disk.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

When local beats the API (and when it doesn't)

Local isn't always the answer. It's a tool with a clear trade-off.

Local wins when:

  • Privacy is a requirement. Sensitive customer data, PII, contracts, medical records. If it can't leave your infra, the discussion ends here: the text never becomes an outbound request.
  • High, predictable volume. Thousands of classifications/summaries per day. On the API that becomes a recurring bill; locally, the marginal cost per call is ~zero after the hardware.
  • Offline or low latency. No network, no round trip to another continent.
  • Free experimentation. Run evals, test prompts in a loop, without fear of the meter running.

The API wins when:

  • You need top-tier capability. Heavy reasoning, complex code, long-running agents. The best open model that runs on your laptop still loses to the frontier.
  • Sporadic load. A few calls a day don't pay for the cost (and the pain) of maintaining hardware.
  • You don't want to be the SRE. With local, you're the one who deploys, updates, monitors, and fixes.

The mature take is hybrid: local for what's high-volume, sensitive, or repetitive; API for what demands the capability ceiling. It's not religion, it's engineering.

Limitations and things to watch

  • Quality has a ceiling. A local 8B is not a frontier model. For tasks that demand deep reasoning, you'll feel the difference. Pick the right battle.
  • Aggressive quantization degrades. Q4 is the sweet spot. Below that (Q3, Q2) quality drops visibly. Don't push it just to fit in VRAM.
  • CPU-only is slow. It runs, but be ready to wait on larger models. For serious use, a GPU or Apple Silicon makes a brutal difference.
  • Long context costs memory. That 128k window the cloud gives you for free, locally you pay for in RAM/VRAM.
  • You become responsible for the infra. Model updates, drivers, monitoring: it's all yours. There's no "someone else's computer" here.

Quick FAQ

Do I need a GPU to run an LLM locally? No. Ollama's minimum is a 64-bit CPU with AVX2 and 8 GB of RAM. A GPU speeds things up a lot (the difference between seconds and tens of seconds per response), but you can start without one. A Mac with Apple Silicon is an excellent route because of unified memory.

Which model should I pick to start? A 7–8B quantized to Q4_K_M is the best balance for consumer hardware: it fits in ~6 GB and runs fast. Qwen3, Llama, and Gemma at that size cover most day-to-day cases.

Is local really free? The software is. The model is. But the hardware isn't, and neither is the electricity. For high volume, it still comes out much cheaper than an API. For sporadic use, it's probably not worth keeping a machine running.

Can I use Ollama in my existing app? Yes. It exposes an OpenAI-compatible endpoint at localhost:11434/v1. You swap the base_url of your current client and the rest of the code stays the same.

Conclusion

Running an LLM locally is no longer a feat. Install Ollama, run ollama run qwen3:8b, and in minutes you have a model answering without sending anything to the cloud. The secret isn't in the command, it's in the decision: understanding the VRAM math, picking the right quantization, and knowing when local beats the API and when it doesn't.

That's the real trick. Running the model is easy; turning it into a product that holds up in production, with the right hybrid architecture, evaluation, and cost under control, is where the real engineering lives. That's exactly the kind of decision we break down live in the Clã Beer and Code, the largest AI engineering community in Brazil, with mentorship, hands-on labs, and the crowd that puts agents in production. The next step, if you've already run your first local model, is to make it worth it: reread the real math on running open source and design your hybrid setup.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing