Skip to content

Adding CoreML Export Support - Looking for Feedback #14

Description

@DjKesu

Before I make the pr, just want to make sure that my process makes sense:

Started by exploring how to handle EdgeTAM's dynamic prompts (points, boxes, masks) in CoreML. The main challenge was that CoreML prefers static input shapes, but EdgeTAM needs to handle variable numbers of points and different prompt types.

Initially tried exporting the entire model as one piece, but quickly realized that wouldn't work well for mobile - the prompt handling was too dynamic and CoreML's tracing couldn't handle the conditional logic properly.

Ended up splitting the model into 3 parts:

  • Image encoder (~9.6MB) - handles the heavy lifting of processing the 1024x1024 input image
  • Prompt encoder (~2MB) - processes user inputs with fixed slots (up to 4 points, 1 box, 1 mask)
  • Mask decoder (~8MB) - generates the final masks

The tricky part was getting the prompt encoder right. Had to dig into the SAM2 code to understand how prompts are embedded, then create wrapper classes that make the dynamic parts static enough for CoreML.

Let me know if this makes sense! happy to create the pr soon :)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions