Skip to content

Thank you for creating this model. How do I run it on my phone? #5

Description

@Tylersuard

Hello. I am excited about your invention. You have it at 4-bit, I am wondering if you could squeeze even more speed and density out of it by going one-bit, like PrismML does.

Also, how can I run this on my iPhone? I am looking forward to trying it out.

Activity

  1. linyubupa commented on Sep 10, 2026

    @linyubupa
    Collaborator

    Thanks so much for the kind words! On 1-bit: we've actually built 1-bit models, and this framework runs them without any problem — but the damage to model quality was too large, so we never released them, and the speed / memory gains were nowhere near dramatic enough to make it worth the trade-off. On iPhone: the app code is being tidied up right now and we'll be releasing it within the next few days.

  2. dzy1997 commented on Sep 11, 2026

    @dzy1997

    Looking forward to the App Store app release!

  3. JeffMII commented on Sep 14, 2026

    @JeffMII

    I think ternary shows more promising results in quality. That may be worthwhile if not already explored.

  4. pszemraj commented on Sep 14, 2026

    @pszemraj

    +1 to code existing to run phone-class LLM on a phone. I don't know about you guys, but I don't use my MacBook as a phone

  5. Raphaeal19 commented on Sep 16, 2026

    @Raphaeal19

    This is a very interesting project, and something that i need desperately for my app (requires on device LLM runs). Is there a way i can contribute to this project?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions