Text Generation Web UI (commonly called “textgen” or “oobabooga”) is a free, open-source Gradio web interface for running large language models locally. First released in late 2022, it was the original go-to local LLM interface and still powers thousands of rigs today. You load an open-weight model and get a clean, ChatGPT-style chat window — plus character chat, a notebook mode for free-form generation, multimodal vision, image generation, and an OpenAI/Anthropic-compatible API server that makes your local model a drop-in replacement for cloud APIs.
The project supports a wide range of inference backends — llama.cpp, ik_llama.cpp, Hugging Face Transformers, ExLlamaV3, and TensorRT-LLM — and lets you switch between them without restarting. It is 100% offline and private, with zero telemetry or external remote requests. Released under the AGPL-3.0 licence with roughly 46,000 GitHub stars, textgen provides portable one-click builds (no setup) for GGUF models on Windows, Linux, and macOS, as well as a full one-click installer for the extended feature set.
Key Features
Multiple local backends: llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM — switch without restarting.
Chat and characters: instruct, chat, and chat-instruct modes with automatic Jinja2 prompt formatting.
OpenAI/Anthropic-compatible API: Chat, Completions, and Messages endpoints with tool-calling — a local drop-in for cloud APIs.
Tool-calling: models can call custom functions, including web search, page fetch, and math, plus MCP support.
Multimodal and attachments: attach images, PDFs, and .docx files to the conversation.
Image generation: a dedicated tab for diffusers models with quantization.
Extensions: built-in and community extensions for TTS, voice input, translation, and more.
100% offline and private: no telemetry, no external resources.
Use Cases
Running open-weight models locally behind a polished chat interface.
An offline and private chatbot or character-based conversation platform.
A local OpenAI-compatible API for your own apps and integrations.
Experimenting with multimodal models, image generation, and tool-calling on your own hardware.
Platforms
Windows, Linux, and macOS (portable builds or one-click installer; browser-based web UI).
Overview LM Studio is a desktop application for discovering, downloading, and running large language models locally. It offers a polished graphical interface where you can browse a built-in model catalogue, run models with fast inference on your own hardware, and chat or build against them with an OpenAI-compatible local API. Popular for its ease of […]
Overview Neuronto is a federated Agentic Resource Discovery (ARD) registry — a search index that answers the question “what tool can do this task?” on behalf of AI agents. Instead of hard-coding integrations one by one, an agent sends a single query and Neuronto searches its own catalogue plus every other public ARD registry at […]