article thumbnail

Ollama

Run Powerful AI Models Right on Your Own Machine

18 min read
#ai, #ollama, #llm, #localai, #friday4

Somewhere in Palo Alto right now, roughly 9 million developers are quietly doing something that would have sounded like science fiction three years ago: running a genuine, GPT-class AI model on their own laptop, with no cloud account, no API key, and no bill. The tool making that possible — Ollama — just raised a $65 million Series B, sits inside 85% of the Fortune 500, and is free. If that combination sounds improbable, stick around; the backstory explains it.

For the last few years, "using AI" has meant sending your words to somebody else's computer — a giant data center owned by a handful of companies — and waiting for an answer to come back over the wire. That works, but it comes with strings attached: monthly bills, rate limits, an internet connection, and the uncomfortable fact that everything you type leaves your machine. Ollama flips that model on its head. It lets you download a real, capable large language model and run it entirely on your own computer — offline, private, and free to use as much as you like. If Docker made it trivial to run any app in a container, Ollama makes it trivial to run any open model on your desktop. Let's dig in.

The Backstory: Two Docker Engineers, Round Two

Ollama wasn't built by strangers to the "make hard infrastructure feel effortless" problem. Co-founders Jeffrey Morgan and Michael Chiang met as students at the University of Waterloo and built a startup called Kitematic — a friendly GUI for running Docker containers — that Docker itself acquired in 2015. Their work there became the core of Docker Desktop, the tool millions of developers open every day without a second thought. In 2023 they set out to do for local AI models what they'd already done for containers: hide the painful parts behind one clean command. It worked. By mid-2026 Ollama had grown to nearly 9 million monthly developers and closed a $65M Series B led by Theory Ventures (with Benchmark, 8VC, and Y Combinator among the backers) — a strong signal that "run it yourself" is becoming a first-class way to use AI, not just a hobbyist workaround.

What Is Ollama?

Ollama is a free, open-source (MIT-licensed) tool that downloads, manages, and runs open-weight large language models locally. It bundles everything that used to be painful — the model weights, the runtime, GPU acceleration, memory management — behind a single, friendly command. Where running a local model once meant compiling C++ and hand-wrangling gigabytes of weight files, Ollama reduces the whole experience to one line:

ollama run llama3.2

That's it. The first time, it downloads the model; every time after, it just starts chatting. Under the hood Ollama builds on the excellent open-source llama.cpp inference engine, but you never have to think about that — it hands you a clean CLI, a local REST API, and sensible defaults.

Why Run a Model Locally?

Cloud AI is convenient, but local models win on several fronts that matter more every year:

The trade-off is honest: a model running on your laptop won't match the very largest frontier models in the cloud. But the gap has narrowed dramatically, and for a huge range of everyday tasks — drafting, summarizing, coding help, classification, brainstorming — a good local model is more than enough.

Cross-Platform: It Runs Everywhere

Ollama offers native installers for all three major desktop platforms — no WSL gymnastics required on Windows.

| Platform | Install | |----------|---------| | Windows | Download the ~5MB installer from ollama.com/download/windows — no admin rights needed, install takes about a minute | | macOS | Download from ollama.com, or brew install ollama | | Linux | curl -fsSL https://ollama.com/install.sh \| sh |

After installing, confirm it's alive:

ollama --version

On Windows and macOS the installer runs Ollama as a background service; on Linux you can start the server manually with ollama serve if it isn't already running.

Your First Conversation

The run command pulls a model (if needed) and drops you into an interactive chat:

ollama run llama3.2
>>> Explain what a database index is, in two sentences.

Type your message, get a response, keep the conversation going. Exit with /bye. You can also pipe a one-shot prompt straight in — perfect for scripts:

echo "Summarize this in one line: the quick brown fox jumps over the lazy dog" | ollama run llama3.2

Managing Models

Ollama treats models a lot like Docker treats images. A handful of commands cover almost everything:

| Command | What it does | |---------|--------------| | ollama pull <model> | download a model without running it | | ollama run <model> | run (and pull if missing) a model interactively | | ollama list | show models you've downloaded | | ollama ps | show models currently loaded in memory | | ollama rm <model> | delete a model to reclaim disk space | | ollama show <model> | view a model's parameters and details | | ollama cp <src> <dst> | copy/rename a model |

Models are published in a central library at ollama.com/library, covering families like Llama, Mistral, Gemma, Phi, Qwen, and DeepSeek, among many others — including Meta's Llama 4, which switched to a mixture-of-experts design (the "Scout" and "Maverick" variants) so a model with hundreds of billions of total parameters only activates a fraction of them per token, keeping it fast enough to actually run at home. New and updated models arrive constantly — browse the library for what's current, since whatever tops the charts the week you read this will likely be old news a month later.

Tags, Sizes, and Quantization

Most models come in several sizes, measured in billions of parameters (e.g. 1B, 3B, 8B, 70B). Bigger models are smarter but need more memory and run slower. You choose a size with a tag:

ollama run llama3.2:1b      # tiny, fast, runs almost anywhere
ollama run llama3.2:3b      # a good all-rounder
ollama run llama3.1:70b     # very capable, needs serious hardware

You'll also see tags mentioning quantization — labels like q4_K_M. Quantization shrinks a model by storing its numbers at lower precision, trading a little quality for a lot less memory and disk. The default tag is usually a sensible 4-bit quantization that balances quality and footprint. A rough rule of thumb for what fits:

| Model size | Rough RAM/VRAM needed | Good for | |-----------|----------------------|----------| | 1B–3B | ~2–4 GB | laptops, quick tasks, even a Raspberry Pi 5 | | 7B–8B | ~6–8 GB | the everyday sweet spot on most machines | | 13B–14B | ~10–16 GB | stronger reasoning, needs a good GPU or lots of RAM | | 70B+ | 40 GB+ | workstations with high-end or multiple GPUs |

Ollama automatically uses your GPU when it can — NVIDIA (CUDA), Apple Silicon (Metal), and AMD (ROCm) are all supported — and falls back to CPU otherwise. CPU works; it's just slower.

The Local API: Ollama as a Server

Here's where Ollama becomes a building block instead of just a chat toy. Whenever it's running, it exposes a local REST API on port 11434 that any program on your machine can call. Send it a prompt with a plain HTTP request:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Write a haiku about databases.",
  "stream": false
}'

Even better, Ollama speaks an OpenAI-compatible API, so a lot of existing code and tools that were written for cloud AI work against your local model with almost no changes — just point them at http://localhost:11434/v1 and use any string as the API key:

curl http://localhost:11434/v1/chat/completions -d '{
  "model": "llama3.2",
  "messages": [{ "role": "user", "content": "Hello!" }]
}'

This one detail is why Ollama plugs so easily into editors, chat front-ends (like Open WebUI), scripting libraries, and frameworks — anything that already knows how to talk to an OpenAI-style endpoint can talk to your laptop instead.

Customizing with a Modelfile

Ollama borrows another idea from Docker: the Modelfile, a tiny recipe that customizes a base model into your own named variant. Set a system prompt to give it a persona, tweak parameters like temperature (creativity), and bake it into a reusable model:

FROM llama3.2

SYSTEM You are a terse senior DBA. Answer in at most three sentences, and prefer SQL examples.

PARAMETER temperature 0.3

Save that as Modelfile, then build and run your custom model:

ollama create dba-helper -f Modelfile
ollama run dba-helper

Now dba-helper is a first-class model you can call from the CLI or the API, with your personality and settings locked in. It's a clean way to turn a general model into a purpose-built assistant.

Beyond the Laptop: Editor Integrations and Cloud Bursting

Ollama's OpenAI-compatible local API turned out to be a gateway drug — it's now a supported local backend for coding agents like Claude Code and GitHub Copilot CLI, so you can point a familiar AI coding assistant at a model that never leaves your machine. And when a task genuinely needs more horsepower than your hardware can offer, Ollama's own hosted cloud models let you burst to bigger weights through the same CLI and API you already know — same commands, same Modelfiles, just more muscle when you ask for it. The local-first philosophy stays intact either way: you choose when data leaves your machine, not a vendor's default.

A Few Practical Tips

Why Ollama Belongs in Your Toolkit

The era where serious AI could only live in a hyperscaler's data center is ending. With Ollama, a capable model runs on the laptop in front of you — quietly, privately, and on your terms. Download it, ollama run something small, and start exploring. The future of AI isn't only in the cloud; a very useful slice of it now fits on your own hard drive.

Next time you reach for a cloud API to summarize a paragraph or draft a snippet, ask yourself: could my own machine have handled that? More and more often, the answer is yes.

Enjoyed this article? Share it with someone who'd love it too.

Most covered topics