Skip to content
#

hybrid-inference

Here are 17 public repositories matching this topic...

🧠 A comprehensive toolkit for benchmarking, optimizing, and deploying local Large Language Models. Includes performance testing tools, optimized configurations for CPU/GPU/hybrid setups, and detailed guides to maximize LLM performance on your hardware.

  • Updated Mar 27, 2025
  • Shell
Liagent_OS_V0.1.2

LiAgent OS is a local-first AI agent OS for building a private personal assistant with governed autonomy across local models and hybrid cloud services. It brings conversation, tool use, multi-agent orchestration, task scheduling, heartbeat execution, human approval, and auditability into one long-lived loop.

  • Updated Mar 9, 2026
  • Python

This repository presents a sophisticated hybrid quantum-classical framework for language identification, specifically tailored for English (eng), French (fra), and Twi (twi). The system leverages classical bigram-based feature extraction with quantum-optimized weighting to achieve high accuracy and efficiency.

  • Updated Jan 5, 2026
  • Jupyter Notebook

This repository implements an ultra-long-context inference framework in PaddlePaddle, featuring an accelerated sparse-quantized SQAttn attention kernel and hybrid CPU-GPU execution. It supports up to 1,024K-token inference and enables efficient large-scale attention computation for long-context language models.

  • Updated Oct 30, 2025
  • Python

Add this topic to your repo

To associate your repository with the hybrid-inference topic, visit your repo's landing page and select "manage topics."

Learn more