Skip to content

Ask Ying-Ja Chen if they know about Multi-Process Service to lower the VRAM overhead, potentially having it run on any modern GPU architecture #27

Description

@keiran-rowell-unsw

Early indications is that typical query sizes will not fully saturate everything and it's mostly about start-up cost, or idle time if running in GPU-server mode.

Would need to bundle together user request batches to really get the scale.

Mentor shared that the NIMs already use MMSeqs-GPU in the background (including AF2!), but not prior experience with Multi-Process Service so that idea was likely idealistic

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions