Skip to content

5GB memory allocated up front #103

Description

@aPaleBlueDot

Possibly this isn't an issue, but I could not figure out how to not have 5GB memory allocated up front, despite tweaking several numbers.

Activity

  1. danielzgtg commented on Aug 13, 2025

    @danielzgtg
    Collaborator

    Which model?

  2. danielzgtg commented on Aug 13, 2025

    @danielzgtg
    Collaborator

    Does #105 help? It makes the model weights disk-backed, so that they may swap to disk if RAM is insufficient. Otherwise, consider choosing a smaller model such as Kokoro_espeak_Q4.gguf, which uses merely ~1.4GB max for a small prompt.

  3. aPaleBlueDot commented on Aug 14, 2025

    @aPaleBlueDot
    ContributorAuthor

    I tried the q5 of parler mini, although I decided against using Parler after reading in an issue on this repo that it mumbles. Yes, mmap would immensely help, and could be acceptable with a few of GB of swap.

  4. danielzgtg commented on Aug 14, 2025

    @danielzgtg
    Collaborator

    Yeah about that, Kokoro is the only practical model at this time. The others seem academic because they are too large or unreliable.
    I've had discussions about this spread across PRs and issues. I'll update README.md to make this fact more visible.

  5. added a commit that references this issue on Aug 14, 2025
  6. aPaleBlueDot commented on Aug 14, 2025

    @aPaleBlueDot
    ContributorAuthor

    Some users caught Kokoro saying gibberish for out of training data, so I can't risk it.

  7. danielzgtg commented on Aug 14, 2025

    @danielzgtg
    Collaborator

    Which model are you using now?

  8. aPaleBlueDot commented on Aug 16, 2025

    @aPaleBlueDot
    ContributorAuthor

    Undecided, still searching.

  9. danielzgtg commented on Aug 19, 2025

    @danielzgtg
    Collaborator

    With Parler and your settings, I'm getting 2.5455017 GiB RssAnon in the middle of dac_runner::run. This might be a problem because a comment says splitting this out from the primary graph so that we can better manage streaming (i.e. sentence chunks are better performed this way) but I don't see any streaming implemented.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions