You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Does #105 help? It makes the model weights disk-backed, so that they may swap to disk if RAM is insufficient. Otherwise, consider choosing a smaller model such as Kokoro_espeak_Q4.gguf, which uses merely ~1.4GB max for a small prompt.
I tried the q5 of parler mini, although I decided against using Parler after reading in an issue on this repo that it mumbles. Yes, mmap would immensely help, and could be acceptable with a few of GB of swap.
Yeah about that, Kokoro is the only practical model at this time. The others seem academic because they are too large or unreliable.
I've had discussions about this spread across PRs and issues. I'll update README.md to make this fact more visible.
With Parler and your settings, I'm getting 2.5455017 GiB RssAnon in the middle of dac_runner::run. This might be a problem because a comment says splitting this out from the primary graph so that we can better manage streaming (i.e. sentence chunks are better performed this way) but I don't see any streaming implemented.
Possibly this isn't an issue, but I could not figure out how to not have 5GB memory allocated up front, despite tweaking several numbers.