How to run DeepSeek-R1 on your own computer (Windows and Mac)

Rewritten in August 2026
The original version of this article, from January 2025, covered a different way of installing DeepSeek. I’ve rewritten it with the current route, which is also much shorter: Ollama.
DeepSeek is a family of open-source AI models from the Chinese company of the same name. DeepSeek-R1 is its reasoning model: before answering, it writes out its chain of thought and then responds. That makes it good at math and programming, and noticeably slower at everything else.
In this guide you’ll install it locally, on Windows or Mac, without depending on any external service and without your data leaving your computer.
Which model your machine can handle
This is the first thing to decide, because it’s what makes the experience good or frustrating.
The full DeepSeek-R1 has 671 billion parameters and needs hundreds of gigabytes of video memory spread across several GPUs. It won’t fit on a laptop, and no trick changes that.
What does run on a normal machine are the distillations: smaller models (based on Qwen or Llama) trained to imitate R1’s reasoning. They’re what almost everyone uses and what Ollama installs.
| Model | Approximate memory | Typical machine |
|---|---|---|
deepseek-r1:1.5b |
~4 GB | Almost any recent laptop |
deepseek-r1:7b |
~5 GB | Mid-range GPU (RTX 4060), Mac with 16 GB |
deepseek-r1:14b |
~9-10 GB | Just fits a 12 GB GPU: lower the context or it runs out of memory |
deepseek-r1:32b |
~18-20 GB | RTX 4090 or Mac with 32 GB or more |
deepseek-r1:671b |
~376 GB | Several GPUs. Not here. |
As for speed, the 32B model on an RTX 4090 runs at around 40 tokens per second, and about 30 on a 3090. That’s fast enough to read along comfortably, but with a reasoning model you have to add the time it spends “thinking” before it starts answering.
On Apple Silicon Macs memory is shared, so read the memory column as total RAM: with 16 GB you’re fine up to the 7B model, and with 32 GB you can go up to 32B.
I’ve run it on a Mac with an M3 and 16 GB, and performance was very good. But that doesn’t generalize: the same setup will behave very differently depending on the computer, so take the table as a starting point and test.
If you don’t know where to start, try 7b. It fits on almost any machine, responds at a good pace, and if it falls short you can download the 14B model later; both can be installed at the same time.
Installation
1. Install Ollama
Ollama is the simplest way to run models locally: it downloads the models, manages memory and exposes an API. There are installers for Windows and Mac.
If you prefer the terminal, on Mac with Homebrew:
brew install ollama
And on Windows with winget:
winget install Ollama.Ollama
2. Download and run the model
One command is enough. The first time it downloads the model, and after that it reuses it:
ollama run deepseek-r1:7b
When the download finishes it drops you straight into a chat in the terminal, and that’s it, it’s running.
To exit, type /bye. To come back, run the same command again, this time without the download.
3. Check what you have installed
ollama list
And if you want to free up space later:
ollama rm deepseek-r1:7b
Using it with a graphical interface
The terminal is fine for trying it out, but for everyday use an interface is more comfortable. Two options work with what you’ve just installed:
- Open WebUI: a ChatGPT-style interface that connects to your local Ollama. It keeps your conversation history and lets you switch models from a dropdown.
- LM Studio: a desktop app that does it all, from downloading models to chatting and serving an API. If you don’t want to touch the terminal at all, you can start here instead of with Ollama.
Connecting it to your code
For programming, the most useful part is that Ollama exposes an API at http://localhost:11434, so you can call it from any script:
curl http://localhost:11434/api/generate -d '{
"model": "deepseek-r1:7b",
"prompt": "Explain what a debounce does in JavaScript",
"stream": false
}'
The response includes the model’s chain of thought inside <think> tags, before the final answer. If you’re going to show the result in an interface, you’ll usually want to separate the two.
Why run it locally?
- Privacy: nothing you type leaves your computer. If you work with code under a confidentiality agreement or with client data, this is the main reason.
- No subscription: once it’s downloaded there’s no message limit and no pay-per-use.
- Works offline: after the initial download you don’t need the internet.
- Open source: the weights are published under the MIT license, so you can use it in commercial projects and fine-tune it if you need to.
That said, a 7B distillation isn’t at the level of GPT-5 or Claude. For its size it’s quite capable, but if you expect the quality of the big paid models you’ll be disappointed. The fair comparison is with having nothing when you’re offline or can’t send your code to a third party.
For more on the model, its benchmark results and the license, see the DeepSeek-R1 repository on GitHub.

