Jan
Open-source desktop app for running LLMs locally with a clean GUI
Jan is an open-source desktop application for running large language models locally on macOS, Windows, and Linux. It provides a ChatGPT-style interface backed by local inference (via llama.cpp) with no internet connection required after model download. Supports Llama, Mistral, Phi, and Gemma families, and can connect to remote APIs (OpenAI, Anthropic, Groq) as an alternative front-end. Built by Menlo Research under Apache-2.0.
Pricing: Free
Jan is a desktop application that runs language models locally, using llama.cpp as its engine and MLX on Apple Silicon. It is aimed at people who want a ChatGPT-style interface without sending conversations to a provider, and it can also be pointed at remote APIs when a local model is not enough.
Hardware requirements are modest but real. The Windows documentation asks for 8GB of RAM (16GB recommended), 6GB of VRAM for NVIDIA, AMD or Intel Arc acceleration, 10GB of free disk, and a CPU with AVX2, which means Intel Haswell (2013) or AMD Excavator (2015) and later. macOS and Linux builds are available. Release v0.8.4 (23 July 2026) moved stored credentials out of browser localStorage and into the OS keyring.
Two things affect adoption. The licence is Apache 2.0 with an added line requesting attribution in user-facing materials, which is why GitHub does not classify it as a standard Apache-2.0 project; the wording is a request rather than an obligation, but it defeats automated licence detection. And Menlo Research, the company behind Jan, now presents itself publicly as a humanoid robotics company built around its Asimov platform, with Jan framed as the origin of that work rather than the main product. Development is still active, with commits through August 2026.
Jan Alternatives
Explore 35 products in the Frameworks & Stacks category. View all Jan alternatives.
Anycloud
CLI and Python SDK for running AI jobs, services and VMs across your own AWS, Azure, GCP, Lambda and Vast accounts
vLLM
High-throughput LLM inference engine with PagedAttention for efficient GPU memory usage
Ollama
Run large language models locally with a single command
Work on Jan? Feature it at the top of Frameworks & Stacks.
Is your product missing?