How to Run an AI Model on Your Own Computer Offline (Ollama and LM Studio)

Run open AI models on your own computer without the internet using Ollama or LM Studio: hardware needs, setup, your first chat, and how to pick a model.

Published 5 min read

A laptop showing a neural network, a crossed-out Wi-Fi icon and a processor chip
In this article
  1. Why run AI locally?
  2. What hardware do you need?
  3. Option 1: Ollama from the command line
  4. Option 2: LM Studio with a graphical interface
  5. Which model should you pick?
  6. Tips for better performance
  7. When is running locally the right choice?

Yes, you can run an AI model on your own computer and chat with it offline, using free apps such as Ollama or LM Studio. You install the app, download an open model once, and from then on every conversation runs on your machine without sending your data to a server.

In short: make sure you have at least 8 GB of RAM (16 GB is better), install Ollama from its official site, type ollama run followed by a small model's name, and start chatting. Prefer a graphical app? Use LM Studio.

Why run AI locally?

Cloud tools like ChatGPT are easy and powerful, but running a model locally has real advantages:

  • Privacy: your text and files never leave your computer, which matters for sensitive documents.
  • Works offline: handy when travelling or on a weak connection.
  • No subscription: the apps and open models are free. Your only costs are your hardware and electricity.
  • Learning: you get a feel for how language models work and can connect them to your own programs through a local API.

The trade-off: local models are smaller than the big cloud models, so they're usually weaker on complex tasks, and their speed depends on your hardware.

What hardware do you need?

The most important factor is memory: system RAM, or video memory (VRAM) if you have a dedicated GPU. Models are described by their number of parameters, such as 3B or 8B (3 or 8 billion). More parameters generally means more capability and more memory.

Model size Rough memory needed Good for
1B to 4B 8 GB Everyday laptops, simple tasks
7B to 9B 16 GB Most day-to-day use
12B to 14B 16 to 32 GB Better answers on a strong machine
Larger 32 GB or more, ideally a strong GPU Advanced users

These are rough figures for quantized (compressed) models, which is what these apps usually download by default. Macs with Apple Silicon do well here because the CPU and GPU share the same memory. You'll also need free disk space: a single model is typically anywhere from 1 GB to 10 GB or more.

Option 1: Ollama from the command line

Ollama is free, open-source software for Windows, macOS and Linux that lets you run a model with a single command.

  1. Go to the official site, ollama.com, download the installer for your system and install it.
  2. Open a terminal (Terminal on macOS and Linux, PowerShell on Windows).
  3. Check the install:
ollama --version
  1. Run a small model. The first time, Ollama downloads it automatically and then opens a chat:
ollama run llama3.2
  1. Type your question after the >>> prompt. Type /bye to leave the chat.

A terminal runs ollama run, downloads the model and starts a chat on the local machine, with a crossed-out cloud icon

Useful Ollama commands

ollama pull llama3.2   # download a model without starting it
ollama list            # show installed models
ollama ps              # show models currently running
ollama rm llama3.2     # delete a model to free up space

Model names change as new models are released, so browse the model library on the Ollama website and copy the name from there.

Ask a quick question without opening a chat

You can put your question straight after the model name. Ollama prints the answer and returns you to the terminal:

ollama run llama3.2 "Explain what RAM is in two sentences"

This is handy when you want one quick answer, or want to use the model inside a simple script.

The local API

Ollama also serves a local API at http://localhost:11434, which you can call from your own apps. A simple example with curl on macOS or Linux:

curl http://localhost:11434/api/generate -d '{"model": "llama3.2", "prompt": "Why is the sky blue?"}'

The request never leaves your computer, because localhost means the machine itself. If APIs are new to you, read what is an API.

Option 2: LM Studio with a graphical interface

If you'd rather avoid the command line, LM Studio is a great choice. It's a desktop app for Windows, macOS and Linux with a chat window similar to ChatGPT.

  1. Download LM Studio from its official site, lmstudio.ai, and install it.
  2. Open the model search inside the app and type a model family such as Llama, Gemma or Qwen.
  3. The app lists the available versions and often indicates whether they'll fit in your machine's memory. Pick a small one to start.
  4. Once it's downloaded, load the model from the top of the chat window and start typing.

Open-source alternatives such as Jan work in a similar way.

Which model should you pick?

Several companies publish open model families, including:

  • Llama from Meta.
  • Gemma from Google.
  • Qwen from Alibaba.
  • Mistral from Mistral AI.
  • Phi from Microsoft.
  • DeepSeek.

Each family keeps releasing new versions and sizes. A simple rule: start with the smallest recent version from a well-known family, test it on questions from your real work, and if the answers are weak and your machine can handle it, move up a size. If you need a language other than English, test that first, because language support varies a lot between models.

Tips for better performance

  • Close heavy apps, like a browser with dozens of tabs, to free memory while the model runs.
  • Start small. A fast small model beats a big one that types a word per second.
  • Write clear prompts. Small models need more precise instructions than big ones. See how to write good prompts.
  • Remove models you don't use to save disk space.
  • Download only from official sources: the app's own site and its built-in model library. Avoid model files from unknown websites.

When is running locally the right choice?

Go local if privacy is a priority, you often work offline, or you want to learn and connect models to your own code. If you need the best possible quality for complex writing or large coding tasks, cloud tools are still ahead; you can compare them in ChatGPT vs Claude vs Gemini. Many people use both: a local model for private, quick tasks and a cloud tool for the hard ones.

Frequently asked questions

Are local models as good as ChatGPT or Claude?

Usually not. Models that fit on a personal computer are much smaller than the big cloud models, so they can struggle with complex tasks. They're still very good for summarising, rewording, general questions, and learning how LLMs work.

Do I need a graphics card (GPU) to run a local model?

No. Small models run on a normal CPU and RAM, just more slowly. A GPU with plenty of video memory, or a Mac with Apple Silicon, makes responses much faster.

Does Ollama send my chats to the internet?

You only need the internet to download the app and the model the first time. After that, processing happens on your machine, and you can disconnect and keep chatting.

Can I use local models in a commercial product?

It depends on each model's licence. Some allow commercial use with conditions and some are more restrictive, so read the licence on the model's official page before building a product on it.