Skip to content

feat: let the listener and audio images run speech on the device - #200

Merged
goldyfruit merged 1 commit into
devfrom
feat/local-speech
Oct 6, 2026
Merged

goldyfruit merged 1 commit into
devfrom
feat/local-speech

Conversation

@goldyfruit

@goldyfruit goldyfruit commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

ovos-installer is adding a local speech choice (OpenVoiceOS/ovos-installer#648): onnx-asr for recognition and phoonnx for the voice, on a Raspberry Pi 5 with 8 GB or an equivalent machine. The virtualenv method already supports it. The containers method needs three things from this repository:

  • Listener image: the alpha image installs ovos-stt-plugin-onnx-asr and onnx-asr[cpu,hub]. The hub extra brings huggingface_hub, without which onnx-asr cannot download a model. Like the [onnx] extra it needs ovos-plugin-manager>=2, so stable skips it with a message. The audio image already has phoonnx through ovos-audio[extras].
  • Model volumes: ovos_stt_models (listener, ~/.local/share/ovos_stt_plugin_onnxasr) and ovos_tts_models (audio, ~/.cache/phoonnx), in docker-compose.yml, .windows.yml and .macos.yml, so an image update does not download the models again. Both images create these directories, so a fresh volume starts out owned by ovos.
  • Per-service memory limits: LISTENER_MEMORY_LIMIT and AUDIO_MEMORY_LIMIT, read before STANDARD_MEMORY_LIMIT, including in the Raspberry Pi override. The Parakeet int8 recognition model peaks at about 1.2 GB on its own, above the 1G and 512M limits; the voice peaks at about 290 MB. Unset, every limit is what it is today.

contract.yml is regenerated: the two variables are listed as optional.

How it reaches installs

The installer pins this repository at v2.1.0. The installer side (containers run ovos-config autoconfigure --offline, LISTENER_MEMORY_LIMIT is raised for local speech, and the models are pre-fetched in the listener and audio containers) can only follow once these compose changes are in a release it can pin. Until then the installer keeps containers on the public servers.

Test plan

  • docker compose config (v2.40.3) for the default, Raspberry Pi, Windows and macOS combinations:
    • Unset: the listener and audio limits stay at 1G (512M on the Pi).
    • LISTENER_MEMORY_LIMIT=4G AUDIO_MEMORY_LIMIT=1G changes only those two services.
    • STANDARD_MEMORY_LIMIT=768M still reaches them through the nested default.
    • Both volumes are declared and mounted everywhere.
  • python3 scripts/contract.py and python3 scripts/test_contract.py.
  • The image builds in this PR's CI, in particular the onnx-asr install on alpha.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added persistent storage for on-device speech recognition and speech synthesis models across supported Compose configurations.
    • Added separate memory limit settings for listener and audio services, with platform-specific defaults.
    • Listener images now attempt to include the ONNX-based speech recognition plugin; builds continue if installation fails.

ovos-installer is gaining a local speech choice: onnx-asr for recognition and
phoonnx for the voice, on hardware that can run them. The virtualenv method
already does; the containers method needs three things from here.

The alpha listener image now installs ovos-stt-plugin-onnx-asr with onnx-asr's
cpu and hub extras (hub brings huggingface_hub, without which onnx-asr cannot
download a model). Like the [onnx] extra it needs ovos-plugin-manager>=2, so
stable skips it. The audio image already ships phoonnx through ovos-audio[extras].

The downloaded models get named volumes, ovos_stt_models and ovos_tts_models, so
an image update does not fetch up to a few gigabytes again. Both images create
the directories first, so the volumes start out owned by the ovos user.

onnx-asr's recognition model alone peaks around 1.2 GB, above the standard
limit, so the listener and the audio service read LISTENER_MEMORY_LIMIT and
AUDIO_MEMORY_LIMIT before STANDARD_MEMORY_LIMIT. Unset, nothing changes: 1G, or
512M with the Raspberry Pi override. contract.yml lists both as optional.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: eb2244d7-2b72-4429-a6df-4f538ce7fa14
📥 Commits

Reviewing files that changed from the base of the PR and between 4b19b0e and b9d1572.

📒 Files selected for processing (8)
  • audio/Dockerfile
  • compose/.env.example
  • compose/docker-compose.macos.yml
  • compose/docker-compose.raspberrypi.yml
  • compose/docker-compose.windows.yml
  • compose/docker-compose.yml
  • contract.yml
  • listener/Dockerfile

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The listener image now attempts to install the ONNX ASR plugin, and the audio image creates a Phoonnx cache directory. Compose configurations add persistent speech-model volumes and separate memory limits for the listener and audio services.

Changes

Speech service setup

Layer / File(s) Summary
Speech plugin and cache setup
listener/Dockerfile, audio/Dockerfile
The listener image attempts to install the ONNX ASR plugin and creates its data directory. The audio image creates a Phoonnx cache directory.
Persistent speech model volumes
compose/docker-compose.yml, compose/docker-compose.macos.yml, compose/docker-compose.windows.yml
The configurations define persistent volumes for STT and TTS models and mount them in the listener and audio services.
Service-specific memory limits
contract.yml, compose/.env.example, compose/docker-compose*.yml
The contract and example environment file add listener and audio memory-limit variables. Compose configurations use separate service limits, with platform-specific defaults and reservations.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to b9d15

No concrete merge-blocking issue is established. The image builds remain untested, including whether ONNX ASR installs successfully on each channel.

Security Architecture Review

Security architecture risk: 🔵 Low · up to b9d15

The changes preserve existing service privileges and network exposure while adding persistent speech caches and optional memory controls. No introduced security vulnerability was established, but recovery from interrupted downloads and compatibility after image rollback remain unverified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The directly changed state is the two speech services' model caches on the deployment host. Access through these mounts remains within their existing service authority; no new IAM role, secret mount, privileged mode, or cross-service model-volume mount is introduced. Effects of malicious or corrupted model content were not established because the consuming implementations were not inspected.

Trust Boundaries and Controls

  • observed — Both images retain non-root runtime users. Existing host networking, sound-device access, audio sockets, and configuration mounts predate this PR and are unchanged. The added dependency and durable model files extend trusted inputs within that existing boundary; the source does not establish model authenticity or parsing controls.

Resilience and Maintainability Implications

  • inferred — Separate caches contain normal cross-service writes, but persistent partial or incompatible state may outlive container replacement. No such failure was demonstrated; atomic publication, validation, cleanup, and recovery guarantees remain unverified in the external speech implementations.

Hardening Proposals

  • proposed — Before relying on persistent caches for rollout recovery, verify model integrity checks, atomic downloads, concurrent-writer behavior, and downgrade compatibility in the speech implementations. Establish an explicit cache-reset recovery procedure; if multiple isolated stacks are supported, use stack-specific volume identities.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: adding on-device speech support to the listener and audio images.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@goldyfruit goldyfruit self-assigned this Oct 6, 2026
@goldyfruit goldyfruit added the enhancement New feature or request label Oct 6, 2026
@goldyfruit goldyfruit added this to the Pac-Man milestone Oct 6, 2026
@goldyfruit
goldyfruit merged commit fb239f3 into dev Oct 6, 2026
9 checks passed
goldyfruit added a commit to OpenVoiceOS/ovos-installer that referenced this pull request Oct 6, 2026
Container installs now get the speech choice too, for the ovos and listener
profiles (a containers satellite runs hivemind-docker images, which carry no
speech plugins). They pass --offline or --online to ovos-config autoconfigure,
so public now means both public there as well, instead of autoconfigure's
hybrid mode. Local speech raises LISTENER_MEMORY_LIMIT and AUDIO_MEMORY_LIMIT,
then runs the same speech tasks as the virtualenv inside the cli, listener and
audio containers, each container fetching its own model into its own volume.

It only goes local when the compose files in use read LISTENER_MEMORY_LIMIT
and mount ovos_stt_models, and the pulled listener and audio images carry
onnx-asr and phoonnx (OpenVoiceOS/ovos-docker#200). Until the pinned release
has both, containers stay public and the installer says so, so this is safe
to merge before that release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant