1970 words · 10 min read

Table of contents

Introduction

Discussing AI with colleagues and friends (just like everyone these days), I shared my interest in learning more about model architecture in depth, and how in the process, local models caught my attention. I often got the question: how do you start testing local models and navigate the absurd amount of information on the topic? This article aims at addressing exactly that.

Disclaimer: If you spend a lot of time on X or YouTube, you will get exposed too much to influencers and sponsored content, telling you either that local models are now matching the reasoning depth of frontier models at every new generation (1, 2) or, on the contrary, that local models are dumb and frontier models are inches away from AGI. Stay away from those sirens and shape your own opinion.

Ecosystem is moving fast, so do local models. Here is my honest simple guide to step in to local model usage as a Software Engineer as of April 2026.

Inference engine setup

An inference engine is the software layer responsible for loading a model’s weights into memory and running the computations needed to generate tokens. It handles the heavy lifting: memory management, batching, hardware acceleration (CPU, GPU, Apple Silicon, etc.), and exposes an interface (typically a local HTTP API) so that other tools can talk to your model. Without it, a model is just a large file of numbers sitting on your disk.

As Ollama opens the world to both local LLMs and open-source frontier models with a simple, developer-friendly interface, this guide will focus on it. There are many other inference engines out there, such as LM Studio, Jan, GPT4All, LocalAI, and llama.cpp (the underlying engine that Ollama itself builds on). Each has its own trade-offs in terms of control, ease of use, and hardware support.

Setup Ollama

Let’s follow the official setup guide and Install using this command:

curl -fsSL https://ollama.com/install.sh | sh

Then launch Ollama. On macOS, you should see a llama icon appear in your menu bar, indicating the service is running. On Linux, the service starts automatically in the background and listens on http://localhost:11434 by default. You can verify everything is working by running:

ollama list

This will display the models you have pulled locally. An empty list is expected on a fresh install.

Ollama basics

Choose your model

The weight of a model and its quantization are going to be the two main parameters impacting your hardware’s ability to run it efficiently.

  • Model size (in billions of parameters) determines the baseline memory requirement. A 7B model will load faster and run on more modest hardware than a 70B model.
  • Quantization is a compression technique that reduces the precision of the model’s weights (e.g., from 32-bit floats down to 4-bit integers). This drastically cuts memory usage at the cost of a small quality degradation. Common formats include q4_K_M, q8_0, and nvfp4.

A good rule of thumb: your model should fit comfortably within your available unified memory (RAM + VRAM). If even one layer is offloaded to CPU, inference speed will drop significantly.

Some tools like llmfit can help you determine what is your best option based on your hardware specs.

I personally use these days primarily the qwen3.6 family (qwen3.6:27b-coding-nvfp4, qwen3.6:35b-a3b-nvfp4) on both my personal and professional laptops (M1 Pro 32Go, M5 32 Go).

Once you chose a model, just run :

# format: ollama pull <model_name_and_tag>
ollama pull qwen3.6:35b-a3b-nvfp4

Get ready, you are in for several gigabytes of downloading.

First steps, commands & GUI

After pulling the model, you can start to interact with it with the following command:

ollama run --verbose qwen3.6:35b-a3b-nvfp4

Disclaimer: The first message you send will be noticeably slow, as the model needs to be loaded into memory. Send a quick Hello! twice and you will see the speed roughly double on the second run once the weights are fully cached.

The --verbose option will allow you to see some stats. Here is what I get using the M5 Macbook Pro:

total duration: 22.948880125s
load duration: 46.027042ms
prompt eval count: 89 token(s)
prompt eval duration: 641.67725ms
prompt eval rate: 138.70 tokens/s
eval count: 953 token(s)
eval duration: 22.259480584s
eval rate: 42.81 tokens/s

If you prefer a graphical interface over the terminal, tools like Open WebUI or Enchanted (macOS-native) connect directly to your local Ollama instance and give you a ChatGPT-like experience in your browser or as a native app.

Performances & Optimizations

If inference feels too slow, here are the levers to pull in order:

  1. Switch to a smaller or more aggressively quantized model. This is the most impactful change. A q4 quantized 14B model will often outperform a q8 27B model on constrained hardware.

  2. Enable Flash Attention. Flash Attention is an optimized attention algorithm that reduces memory bandwidth usage and speeds up the attention computation significantly, especially for longer contexts. In Ollama, you can enable it by setting the following environment variable before launching:

    OLLAMA_FLASH_ATTENTION=1 ollama serve

    On macOS, you can persist this with: launchctl setenv OLLAMA_FLASH_ATTENTION 1

    Then restart Ollama. The speedup is most noticeable on longer prompts or when using extended context windows.

  3. Limit context size. Running with a very large context window (e.g., 128k tokens) is expensive. Unless you need it, keep the context to a reasonable size (8k–32k) to preserve speed.

Bonus: frontier model in Ollama

Creating a free account on ollama.com unlocks access to cloud models: large frontier-grade models (up to 480B parameters) running on Ollama’s datacenter hardware, without touching your local machine. The experience is seamless, cloud models behave exactly like local ones and use the same CLI commands. Ollama does not retain your data.

Available cloud models as of writing include qwen3-coder:480b-cloud, gpt-oss:120b-cloud, gpt-oss:20b-cloud, and deepseek-v3.1:671b-cloud.

From account creation to first run in 3 steps:

1. Create an account at ollama.com, then sign in from your terminal:

ollama signin

2. Pull a cloud model (no gigabytes to download, just a reference):

ollama pull gpt-oss:120b-cloud

3. Run it exactly like any local model:

ollama run gpt-oss:120b-cloud

That is it. The -cloud suffix in the tag name is the only difference from a local model. Your existing tools, API calls, and agentic setups work without any change.

Note The free tier covers smaller cloud models. Larger ones (e.g. qwen3-coder:480b-cloud) may require a paid plan. Check ollama.com/pricing for the current limits. ### 2. Agentic coding

Chatting with a model is great for exploration, but the real productivity gain for engineers is agentic coding: letting a model autonomously read your codebase, plan changes, write code, run tests, and iterate. This is where local models start to become genuinely useful in a daily workflow.

Ollama exposes an OpenAI-compatible API, which means most agentic coding tools that support custom endpoints will work out of the box. To use it with Claude Code (Anthropic’s CLI agent), use the following command:

ollama launch claude --model qwen3.6:35b-a3b-nvfp4

This tells Claude Code to use the Ollama API instead of Anthropic’s servers, with your local model as the backend.

Now you can run your flows using this local LLM 😎

Disclaimer: The first prompt is going to be slow, just as for the previous session, give it a real shot in a small codebase to start with.

I take the opportunity to suggest you to use lighter, safer and simpler agent than claude-code like Pi, Crush, OpenCode. I personaly use crush as :

  • It’s Go based (and with the amount of js vulns & supply-chain attacks these days…)
  • It’s pleasant aesthetically
  • It doesn’t hide the thought process of the LLM
  • Allows to switch fast from one model to another (frontiers, locals, multi-providers, etc)

Suggested workflow

A local model is not a replacement for a frontier model in an agentic setup; it is a cost-effective executor for well-defined, granular tasks. The key is to structure your workflow so that each tool is used where it shines.

Here is the approach I find most effective:

1. Ground the agent in your codebase. Before starting any task, have the agent analyze the codebase and produce an AGENTS.md file. This document summarizes the project structure, the key modules, the conventions used, and the entry points an agent needs to know to navigate the code autonomously. This is a one-time investment per project that pays dividends in every subsequent agentic session.

2. Plan with a frontier model. For any non-trivial feature or refactoring task, use a frontier model (Opus 4.x, Kimi 2.6, GLM5.1, Qwen3.6 Plus) to produce the high-level plan. Frontier models are better at reasoning about architectural trade-offs, understanding ambiguous requirements, and producing coherent multi-step strategies. This step is cheap (a single prompt) and sets the direction.

3. Break down the plan into detailed, atomic steps. Take the high-level plan and decompose it into small, well-scoped sub-tasks. Each step should be achievable in a single agentic pass with limited context. This decomposition can itself be done by a model, either the frontier one from the previous step or a mid-tier local model.

4. Execute the steps with your local model. Now hand each atomic step to your local model via the agentic coding interface. Because the task is well-defined and self-contained, a capable local model (27B–35B, well quantized) can execute it reliably without the broader reasoning capacity of a frontier model. You keep full control, nothing leaves your machine, and the cost is zero.

This pattern, plan big remotely and execute small locally, gives you the best of both worlds.

Bonus: Accessing your powerful machine from anywhere

You have a powerful machine at home and want to leverage it from your workstation, your phone, or a lightweight laptop? A personal VPN mesh is the cleanest solution.

Install Netbird (🇪🇺) or Tailscale (🇺🇸) on all your devices and log in through the same account. Here is the process to install Netbird on macOS:

brew install --cask netbirdio/tap/netbird-ui sudo netbird service install sudo netbird service start

Open the Netbird GUI and click Connect (macOS: right-click on the menu bar icon). Make sure you allow it to send and receive connections.

Repeat this setup on all the machines you want to connect, then visit https://app.netbird.io/peers. You should see all your connected devices listed there with their assigned private IP addresses.

Next, configure Ollama to broadcast on all network interfaces so it can be reached remotely. On macOS:

launchctl setenv OLLAMA_HOST "0.0.0.0:11434"

Restart Ollama. You can now point any agentic coding interface on your laptop or phone to http://<home-machine-ip>:11434 and run your full local model stack remotely, without any cloud provider in the loop.

Few words to leave you with

Local models are a genuinely exciting tool for engineers, but approaching them with the right expectations is everything.

Do not expect frontier model performance. Local models have improved dramatically, but a quantized 30B model is not Claude Sonnet. For complex reasoning, architecture decisions, or subtle debugging, frontier models are still ahead. The goal is not to replace them but to reduce your dependency on them for the right class of tasks.

Combine rather than replace. The most productive setups use local models as high-throughput, zero-cost executors for well-defined subtasks, while reserving frontier model calls for planning, evaluation, and anything that requires nuanced judgment. Think of it as delegating to a fast junior developer who executes reliably when given precise instructions.

Keep the pace. The local model ecosystem moves quickly. New model families appear monthly, quantization techniques improve, and inference engines keep getting faster. Stay curious: periodically pull a new model, try an alternate agentic interface, and revisit your workflow. What felt sluggish six months ago may now run smoothly on your current hardware.

With that in mind, I share Julien Chaumond’s opinion, certain models such as Qwen3.6 27B feel very close to hitting Sonnet or Opus performances for deterministic agentic coding.

The barrier to entry has never been lower. Start small, iterate, and build your own opinion.