On-screen text
#1
ollama
Run Llama, Mistral, Gemma and hundreds of other models with a single command. The fastest way to get a real model running on your own laptop.
ollama/ollama
Get up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
#2
llama.cpp
The engine that started local AI. Pure C++ inference that runs big models on ordinary hardware, from a MacBook to a Raspberry Pi.
ggml-org/llama.cpp
LLM inference in C/C++
#3
open-webui
A polished ChatGPT-style interface for your local models. Chat, documents, RAG and multiple users, all running fully offline on your machine.
open-webui/open-
webui
User-friendly AI Interface (Supports Ollama, OpenAI
API, ...)
#4
vllm
The serving engine for when it gets serious. High-
throughput inference that squeezes every token out of your
GPU, built for production loads.
vllm-project/vllm
A high-throughput and memory-efficient inference
and serving engine for LLMs
#5
jan
An offline ChatGPT alternative as a clean desktop app. Download a model, start chatting, no account and no cloud. Your conversations never leave your device.
janhq/jan
Jan is an open source alternative to ChatGPT that
runs 100% offline on your computer.
#6
exo
Turn the devices you already own into one AI cluster. It
splits big models across your MacBook, iPhone and iPad so
together they run what one alone cannot.
exo-explore/exo
Run frontier AI locally.
#7
LocalAI
A drop-in replacement for the OpenAI API that runs on
your own hardware. Point your existing apps at it and they work,
no code changes and no API bill.
mudler/LocalAI
LocalAI is the open-source AI engine. Run any
model - LLMs, vision, voice, image, video - on any
hardware. No...
Save this for
your next build
Your data stays yours.
@joshualevi.ai ->