Skip to content
LanceLRQPublic

About

llama.cpp model management panel: self-hosted Docker web UI for GGUF model lifecycle, config, downloads, monitoring, and OpenAI-compatible inference proxy / llama.cpp 模型管理面板:自托管 Docker Web 控制台,管理 GGUF 模型的启停、配置、下载与监控、API 中转

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

llamapad

License: MIT Docker Pulls Docker Image Size Docker Build Powered by llama.cpp

中文 | English

llamapad is a self-hosted management panel for llama.cpp. It manages a Dockerized llama.cpp service and model files in the browser.

llamapad 是一个自托管的 llama.cpp 模型管理面板,在浏览器里管理 Docker 化的 llama.cpp 服务与模型文件,部署本地大模型。

Note

This is a preview release, so features and config formats may still change before a stable version. If you run into a bug or have a feature idea, please open an issue. Thanks for trying it out!

Overview: charts for CPU, memory, GPU memory and inference metrics, plus the running model and the event log Repo archive: GGUF files grouped by quantization, showing which are downloaded and which are auxiliary models

Features

  • 🎛️ Model management - Model list with one-click start/stop (Docker + GPU acceleration); run several models at once, with automatic port shifting on clashes and model-based routing in the API relay
  • ⚡ MTP acceleration - Detects from GGUF metadata whether a weight carries MTP layers; turn on speculative decoding with one switch, or link a separate draft weight
  • 🏠 Models home - Running models, recently updated repos, and trending or searched GGUF repos on HuggingFace, one click away from a download
  • 📝 Parameter editing - Form-based editing in the panel, showing the merged final parameters; configs support YAML import/export and automatic snapshots you can commit to git
  • 🗂️ Namespaces - Group models into custom spaces, share GGUF files across spaces, delete safely with reference checks
  • 📥 Model downloads - HuggingFace (official and mirror) plus direct URLs, resumable with sha256 verification, proxy configurable in the panel; pasting a repo auto-groups files by quantization (Q4/Q8/…), and split files are grouped automatically
  • 🧙 Creation wizard - Pick a repo, choose files, save the config; done in one pass
  • 📁 File management - ComfyUI-style unified file browser; move/rename with reference checks, disk usage at a glance
  • 📊 Monitoring - Container CPU/memory, llama.cpp inference metrics (slots, token rates), GPU memory and temperature, host disk and network, live logs
  • 💬 Playground - Built-in chat page; /llama-proxy/* also reverse-proxies the inference API, so an SSH tunnel only needs to expose one port
  • 🔐 Auth & API - Login protection, plus a REST API you can call from scripts
  • 🌏 Bilingual UI - Chinese/English interface with a built-in documentation center

Quick Start

Install with one command (Linux + Docker required; NVIDIA Container Toolkit for GPU acceleration):

curl -fsSL https://raw.githubusercontent.com/LanceLRQ/llamapad/main/deploy/llamapad.sh | bash

The script checks your Docker setup, installs to /opt/llamapad by default, and walks you through picking a model directory (with free space listed per disk), a runtime user, GPU, port, and admin password (leave blank to generate one). It detects the docker.sock gid on its own, then pulls the image and starts the panel.

Once installed, run llamapad from any directory to open the management menu:

Command What it does
llamapad Interactive arrow-key menu
llamapad start / stop / restart / status Start, stop, restart, check status
llamapad logs -f Follow logs
llamapad config Change port, listen address, model directory, GPU, admin password, and more
llamapad build [--repo path] Build the image locally (shows up in the menu as "Build image" when a repo is found)
llamapad upgrade Upgrade the script and image
llamapad doctor Environment self-check
llamapad uninstall Uninstall

If you'd rather deploy the compose file by hand, see Deployment. To take over an existing manual deployment, just run the script in that directory (back up first; your data and models are left untouched).

Documentation

Full documentation lives in docs/guide/en/ (also available in Chinese), and can be read inside the panel from the sidebar. No required reading order; look up what you need.

Getting started

Document Contents
Quick Start Three steps to deploy, first login, launching your first model
Glossary Quick reference: GGUF, quantization, splits, namespaces, and other terms

Deployment

Document Contents
Deployment & Operations Directory layout, runtime user and permissions, build-time proxy, upgrades and backups
HTTPS Reverse Proxy Reference nginx configs for single-domain and subdomain setups

Usage

Document Contents
Model Management Create/edit/clone, parameter groups, running several models, readiness checks
Model Downloads HF and direct links, resumable downloads, verification, proxy configuration
Files & Namespaces Directory structure, namespace semantics, reference checks, the three-layer deletion model
Settings Reference All four settings groups, item by item

Operations & Troubleshooting

Document Contents
Monitoring & Logs Metric definitions, multi-GPU aggregation, retention tiers and source fallback
Config Format & Migration Fields of the exported YAML, hand editing, migrating from llama-launcher
Troubleshooting Known pitfalls, each verified on a real machine

API

Document Contents
Inference API Playground, the proxy endpoint, clients and SDK integration
Panel API Authentication, common task examples, full endpoint list

Chinese documentation: docs/guide/zh/.

See the changelog for what changed in each release.

Development

pnpm install       # pnpm is the package manager
pnpm run dev       # dev server (PANEL_DOCKER defaults to mock; no real docker.sock needed)
pnpm test          # tests (vitest)
pnpm run lint      # eslint
pnpm run build     # production build (next build, standalone output)

Contributing

Issues are welcome; please open one before sending a PR. When reporting a bug, attach the output of llamapad doctor if you can.

License

MIT — see LICENSE.


If llamapad is useful to you, please consider giving it a ⭐ on GitHub and Docker Hub.

About

llama.cpp model management panel: self-hosted Docker web UI for GGUF model lifecycle, config, downloads, monitoring, and OpenAI-compatible inference proxy / llama.cpp 模型管理面板:自托管 Docker Web 控制台,管理 GGUF 模型的启停、配置、下载与监控、API 中转

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages