Repository navigation
feat: let the listener and audio images run speech on the device - #200
Conversation
ovos-installer is gaining a local speech choice: onnx-asr for recognition and phoonnx for the voice, on hardware that can run them. The virtualenv method already does; the containers method needs three things from here. The alpha listener image now installs ovos-stt-plugin-onnx-asr with onnx-asr's cpu and hub extras (hub brings huggingface_hub, without which onnx-asr cannot download a model). Like the [onnx] extra it needs ovos-plugin-manager>=2, so stable skips it. The audio image already ships phoonnx through ovos-audio[extras]. The downloaded models get named volumes, ovos_stt_models and ovos_tts_models, so an image update does not fetch up to a few gigabytes again. Both images create the directories first, so the volumes start out owned by the ovos user. onnx-asr's recognition model alone peaks around 1.2 GB, above the standard limit, so the listener and the audio service read LISTENER_MEMORY_LIMIT and AUDIO_MEMORY_LIMIT before STANDARD_MEMORY_LIMIT. Unset, nothing changes: 1G, or 512M with the Raspberry Pi override. contract.yml lists both as optional. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (8)
Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe listener image now attempts to install the ONNX ASR plugin, and the audio image creates a Phoonnx cache directory. Compose configurations add persistent speech-model volumes and separate memory limits for the listener and audio services. ChangesSpeech service setup
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Feature Merge Risk: ⚪ Minimal · up to No concrete merge-blocking issue is established. The image builds remain untested, including whether ONNX ASR installs successfully on each channel. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The changes preserve existing service privileges and network exposure while adding persistent speech caches and optional memory controls. No introduced security vulnerability was established, but recovery from interrupted downloads and compatibility after image rollback remain unverified. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Container installs now get the speech choice too, for the ovos and listener profiles (a containers satellite runs hivemind-docker images, which carry no speech plugins). They pass --offline or --online to ovos-config autoconfigure, so public now means both public there as well, instead of autoconfigure's hybrid mode. Local speech raises LISTENER_MEMORY_LIMIT and AUDIO_MEMORY_LIMIT, then runs the same speech tasks as the virtualenv inside the cli, listener and audio containers, each container fetching its own model into its own volume. It only goes local when the compose files in use read LISTENER_MEMORY_LIMIT and mount ovos_stt_models, and the pulled listener and audio images carry onnx-asr and phoonnx (OpenVoiceOS/ovos-docker#200). Until the pinned release has both, containers stay public and the installer says so, so this is safe to merge before that release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Summary
ovos-installer is adding a local speech choice (OpenVoiceOS/ovos-installer#648): onnx-asr for recognition and phoonnx for the voice, on a Raspberry Pi 5 with 8 GB or an equivalent machine. The virtualenv method already supports it. The containers method needs three things from this repository:
ovos-stt-plugin-onnx-asrandonnx-asr[cpu,hub]. Thehubextra bringshuggingface_hub, without which onnx-asr cannot download a model. Like the[onnx]extra it needs ovos-plugin-manager>=2, so stable skips it with a message. The audio image already has phoonnx throughovos-audio[extras].ovos_stt_models(listener,~/.local/share/ovos_stt_plugin_onnxasr) andovos_tts_models(audio,~/.cache/phoonnx), indocker-compose.yml,.windows.ymland.macos.yml, so an image update does not download the models again. Both images create these directories, so a fresh volume starts out owned byovos.LISTENER_MEMORY_LIMITandAUDIO_MEMORY_LIMIT, read beforeSTANDARD_MEMORY_LIMIT, including in the Raspberry Pi override. The Parakeet int8 recognition model peaks at about 1.2 GB on its own, above the 1G and 512M limits; the voice peaks at about 290 MB. Unset, every limit is what it is today.contract.ymlis regenerated: the two variables are listed as optional.How it reaches installs
The installer pins this repository at
v2.1.0. The installer side (containers runovos-config autoconfigure --offline,LISTENER_MEMORY_LIMITis raised for local speech, and the models are pre-fetched in the listener and audio containers) can only follow once these compose changes are in a release it can pin. Until then the installer keeps containers on the public servers.Test plan
docker compose config(v2.40.3) for the default, Raspberry Pi, Windows and macOS combinations:LISTENER_MEMORY_LIMIT=4G AUDIO_MEMORY_LIMIT=1Gchanges only those two services.STANDARD_MEMORY_LIMIT=768Mstill reaches them through the nested default.python3 scripts/contract.pyandpython3 scripts/test_contract.py.🤖 Generated with Claude Code
Summary by CodeRabbit