A universal OpenCode plugin for dynamic model discovery with flexible configuration for OpenAI-compatible providers.
-
Updated
Aug 21, 2026 - TypeScript
A universal OpenCode plugin for dynamic model discovery with flexible configuration for OpenAI-compatible providers.
Self-hosted Qwen3.8-27B (FP8) inference with vLLM, KServe and Envoy AI Gateway on RTX 6000 PRo or 2× RTX 4080 Super
Terraform setup for deploying a private coding LLM on Vast.ai with vLLM, Qwen3 Coder, and OpenCode.
Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.
Enterprise-grade Sovereign AI Stack optimized for NVIDIA Blackwell (sm_120) & vLLM. Features 256K context window, 5.8k tok/s prefill, and integrated observability via Langfuse.
The control plane for self-hosted AI inference. Warm-state GPU routing, multi-runtime orchestration across Ollama, vLLM, llama.cpp, TGI and MLX . Single Go binary. Apache-2.0.
A platform-agnostic format + method for sizing local-LLM hardware from your real agent sessions, calibrated with capability-oracle runs.
Windows utility for using Codex Desktop GUI with Ollama and other local/self-hosted LLM backends via profiles.
Alternatively prompt, an LLM-based plugin for the IntelliJ product family. Answer questions and generate code with self-hosted models.
Air-gapped pre-deployment network change validation against a real containerlab digital twin, with sealed PCI/SOC2/NIST evidence. Zero egress.
Production-grade multimodal RAG assistant using open-source LLMs and vector databases.
convert instagram reels to markdown recipe using self hosted LLMs on VLLM
GLM-5.2 744B at 4-bit on Modal 4x H200 via vLLM, plus a static streaming chat UI.
Self-hosted vLLM inference stack with an OpenAI-compatible API, Docker Compose templates, Caddy reverse proxy, and NVIDIA GPU thermal guard.
A university-scale LLM serving platform in miniature: vLLM on cloud GPU, LiteLLM gateway with per-faculty governance, Prometheus/Grafana SLOs, k3d/ArgoCD GitOps. All numbers measured, all failures documented.
Self-hosted AI coding platform — provisions GPU compute on Vast.ai and deploys open-weight LLMs via llama.cpp as a bring-your-own-key backend for GitHub Copilot (also auto-configures Continue.dev/Cline).
AI-powered herbal remedies chatbot based on "The Little Handbook of Natural Remedies" by Michael Martin. RAG system for natural medicine research.
A custom AI agent that gives real-time weather-based outfit advice, built with a self-hosted vLLM model. Features persistent memory, a manual ReAct loop, and layered guardrails — with annotated source code explaining every design decision.
A curated collection of applications built using [LangChain](https://docs.langchain.com/) and related frameworks. This repository serves as a **monorepo** showcasing different use cases, integrations, and patterns for agentic AI development.
KV-cache-aware intelligent routing for self-hosted and hybrid LLM fleets. Route requests using model quality, latency, cost, policy, and live GPU state.
Add a description, image, and links to the self-hosted-llm topic page so that developers can more easily learn about it.
To associate your repository with the self-hosted-llm topic, visit your repo's landing page and select "manage topics."