
PrivateGPT is an open-source API layer that turns local models into production AI applications. Rather than running models itself, it connects to any OpenAI-compatible inference server — Ollama, llama.cpp, vLLM, and others — and provides the higher-level building blocks you need to build private AI products: a standard messages API, document ingestion, retrieval-augmented generation (RAG) with citations, embeddings, custom tools, MCP connectors, and structured access to databases and CSVs.
The project began as a viral 2023 proof-of-concept that let you chat with your own documents fully offline, with no data leaving your machine. It crossed 50,000 GitHub stars and remains one of the most-watched AI repositories. Maintained by the Zylon team and released under the Apache-2.0 licence, PrivateGPT follows the Claude API as its reference model, so the API surface is familiar and complete — with streaming, async processing, token counting, extended thinking, and structured outputs. It ships a built-in workbench UI (at /ui) for testing and demos, while the API itself is the real product for developers.
Linux, macOS, and Windows (pip/uv, Homebrew, or Docker).
Apache-2.0 (open source)