Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 14 additions & 28 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,40 +1,26 @@
# Devino
# devino

An independently maintained experimental toolkit for model conversion and local inference environments. Validated features can be adopted individually by consuming products. Environment advice is supplied by [intel-gpu-wsl-advisor](https://github.com/zixcel/intel-gpu-wsl-advisor).
Experiment with model conversion and local inference setup before adopting a verified feature in another tool.

# Project Overview
## What you can do

This repository provides an environment for leveraging OpenVINO and OneAPI to enable LLM inference on Intel devices. It includes scripts for setting up an Ubuntu-based environment, converting models to OpenVINO IR format, and deploying an OpenVINO model server.
- Review conversion and environment scripts.
- Compare explicitly configured local inference experiments.

## Repository Structure
## Current scope

- [setup](./setup): Provides scripts for setting up a Pytorch-XPU environment using Ubuntu 22 and Poetry.
installation.
- Model Directories: Named according to Hugging Face model IDs, each containing conversion and inference scripts based on the setup environment.
- [playground](./playground/): Contains sample scripts tested in an OpenVINO 2025 and Ubuntu 24 environment.
This is an experimental toolset. Hardware, model files and wheel sources are registered inputs; a script or documentation example is not production compatibility evidence.

## Key Features
Package distribution is not activated by this documentation. Use the checked-in source and the declared dependency versions; published availability must be verified separately.

- **OpenVINO IR Conversion**: Converts Hugging Face models to OpenVINO IR format for optimized inference.
- **OpenVINO Model Server (OVMC)**: Implements an OpenVINO model server for running converted models.
- **GPU Acceleration**: Provides performance improvements for inference using Intel GPUs.
## Getting started

## Getting Started
Start with the implementation and examples linked below. Review registered configuration and prerequisites before running a command that writes state or contacts a service.

1. **Verify Ubuntu Compatibility**: Check the appropriate WSL Ubuntu version using [intel-gpu-wsl-advisor](https://github.com/zixcel/intel-gpu-wsl-advisor).
- The advisor is optional; callers may verify the Windows driver, WSL kernel and runtime requirements directly.
2. **Setup Environment**: Use the scripts in `setup/` to install dependencies and configure Pytorch-XPU.
3. **Convert Models**: Run the provided conversion scripts to transform models into OpenVINO IR format.
4. **Deploy Model Server**: Install OpenVINO GenAI's OVMC server and execute converted models. by [Makefile](./Makefile)
## Documentation and source

## References
[Interface reference](docs/interface-reference.md)

- [OpenVINO Documentation (2025)](https://docs.openvino.ai/2025/index.html)
- [Hugging Face Optimum-Intel](https://huggingface.co/blog/deploy-with-openvino)
[Usage guide](docs/getting-started.md)

This repository is under active development, integrating new features for optimized inference and deployment on Intel hardware.


## Registered local inputs

Some historical setup recipes use a local PyTorch wheel. Supply the selected, verified wheel at `registration/torch.whl` before running such a recipe; registration data and downloaded model weights are excluded from Git. External model/runtime licenses and hardware compatibility must be verified for the selected experiment. The migration validates source and configuration without downloading models, starting servers, or changing the host.
[Contributing](CONTRIBUTING.md) · [Security reporting](SECURITY.md) · [License](LICENSE) · [Attribution notices](NOTICE)
22 changes: 22 additions & 0 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Using devino

Experiment with model conversion and local inference setup before adopting a verified feature in another tool.

## Before you start

This is an experimental toolset. Hardware, model files and wheel sources are registered inputs; a script or documentation example is not production compatibility evidence.

## First steps

Read the declared configuration and implementation before enabling an operation. Choose an explicitly registered target and review any write or external effect.

## How to assess the result

- Review conversion and environment scripts.
- Compare explicitly configured local inference experiments.

A passing source-level check establishes only what that check observes. Keep missing configuration, unavailable services and unverified deployment paths visible.

## Continue reading

[Repository overview](../README.md)
35 changes: 35 additions & 0 deletions docs/interface-reference.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# devino interface reference

Use the [usage guide](getting-started.md) for the first steps. This reference preserves the current interface details and operational limits. Run command examples from the repository root, after preparing the exact declared dependencies and registered configuration.

## Repository Structure

- [setup](../setup): Provides scripts for setting up a Pytorch-XPU environment using Ubuntu 22 and Poetry.
installation.
- Model Directories: Named according to Hugging Face model IDs, each containing conversion and inference scripts based on the setup environment.
- [playground](../playground): Contains sample scripts tested in an OpenVINO 2025 and Ubuntu 24 environment.

## Key Features

- **OpenVINO IR Conversion**: Converts Hugging Face models to OpenVINO IR format for optimized inference.
- **OpenVINO Model Server (OVMC)**: Implements an OpenVINO model server for running converted models.
- **GPU Acceleration**: Provides performance improvements for inference using Intel GPUs.

## Getting Started

1. **Verify Ubuntu Compatibility**: Check the appropriate WSL Ubuntu version using [intel-gpu-wsl-advisor](https://github.com/zixcel/intel-gpu-wsl-advisor).
- The advisor is optional; callers may verify the Windows driver, WSL kernel and runtime requirements directly.
2. **Setup Environment**: Use the scripts in `setup/` to install dependencies and configure Pytorch-XPU.
3. **Convert Models**: Run the provided conversion scripts to transform models into OpenVINO IR format.
4. **Deploy Model Server**: Install OpenVINO GenAI's OVMC server and execute converted models. using the checked-in experiment scripts

## References

- [OpenVINO Documentation (2025)](https://docs.openvino.ai/2025/index.html)
- [Hugging Face Optimum-Intel](https://huggingface.co/blog/deploy-with-openvino)

This repository is under active development, integrating new features for optimized inference and deployment on Intel hardware.

## Registered local inputs

Some historical setup recipes use a local PyTorch wheel. Supply the selected, verified wheel at `registration/torch.whl` before running such a recipe; registration data and downloaded model weights are excluded from Git. External model/runtime licenses and hardware compatibility must be verified for the selected experiment. The migration validates source and configuration without downloading models, starting servers, or changing the host.
Loading