• wewbull@feddit.uk
    link
    fedilink
    English
    arrow-up
    7
    ·
    21 hours ago

    That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

    It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.