Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
-
Updated
Sep 15, 2026 - C++
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
Crossplatform Swift SDK for Cactus hybrid inference and running LLMs locally in your app.
🧠 A comprehensive toolkit for benchmarking, optimizing, and deploying local Large Language Models. Includes performance testing tools, optimized configurations for CPU/GPU/hybrid setups, and detailed guides to maximize LLM performance on your hardware.
LiAgent OS is a local-first AI agent OS for building a private personal assistant with governed autonomy across local models and hybrid cloud services. It brings conversation, tool use, multi-agent orchestration, task scheduling, heartbeat execution, human approval, and auditability into one long-lived loop.
Official Python SDK for the TQNN Fault-Tolerant Inference Platform.
Execution infrastructure for local-first AI. Reason locally, execute globally.
AI powered financial document analysis platform that extracts insights from reports and generates structured summaries and answers using LLMs
Build a private AI assistant that runs locally with controlled autonomy, combining task management, tool use, and human oversight in one system.
Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery
This repository presents a sophisticated hybrid quantum-classical framework for language identification, specifically tailored for English (eng), French (fra), and Twi (twi). The system leverages classical bigram-based feature extraction with quantum-optimized weighting to achieve high accuracy and efficiency.
Hybrid contextual inference routing for Microsoft Scout - route tasks across cloud, on-device and org-hosted models with an enforced egress boundary.
On-device vs. cloud routing policy for React Native LLM apps. The RN sibling of local-first-llm, wrapping react-native-executorch.
Provider-agnostic, per-request on-device/cloud LLM routing for native Android (AICore + any cloud). Kotlin sibling of local-first-llm.
This repository implements an ultra-long-context inference framework in PaddlePaddle, featuring an accelerated sparse-quantized SQAttn attention kernel and hybrid CPU-GPU execution. It supports up to 1,024K-token inference and enables efficient large-scale attention computation for long-context language models.
Per-request on-device vs cloud LLM routing — an OpenAI-SDK-shaped router that asks: can this device handle this request, right now?
[ENG] A hybrid cloud/local LLM inference orchestrator via Model Context Protocol (MCP) and Ollama / [JPN] Model Context Protocol (MCP) と Ollama を活用したクラウド・ローカル LLM ハイブリッド推論オーケストレーター
To associate your repository with the hybrid-inference topic, visit your repo's landing page and select "manage topics."