Capacity planning guide

Muse Glimmer 30B System Requirements

LAST REVIEWEDSOURCE-BACKED

Direct answer

Meta publishes memory targets, not minimum requirements:24 GiB for K-Quant 17 GB,32 GiB for K-Quant Dynamic, and 64 GiB for BF16. Real fit also depends on selected companions, context allocation, the runtime build, offload, drivers, and operating-system overhead.

Published capacity tiers

Start with the official target, then budget the runtime

Wide table: swipe or use Shift + mouse wheel to inspect every column.

Published target tiers and exact main-artifact sizes. A target is not a guaranteed device result.
Official targetMain artifactExact bytesBinary sizeComponent note
24 GiBK-Quant 17 GB16,756,681,05615.606 GiBVision projector and DFlash drafter are separate files when selected.
32 GiBK-Quant Dynamic19,653,957,98418.304 GiBVision projector and DFlash drafter are separate files when selected.
64 GiBBF16 weights59,553,435,27255.463 GiBThe perception encoder is included; DFlash is not an optional BF16 companion here.

Separate GGUF artifacts

Add only the companions your workload needs

Official companionVision projector

mmproj-kquant.gguf

1,400,328,928 bytes · 1.304 GiB

Official companionDFlash drafter

dflash-kquant.gguf

1,631,205,312 bytes · 1.519 GiB

Configuration ceiling131,072 tokens

This is the model configuration limit, not proof that the full window fits or performs well on a particular machine.

Context memory boundary

Use an active and a conservative KV estimate

At the configured context ceiling, the architecture-derived active estimate is1.7 GiB. It counts all13 global layers across the selected context and caps the39 local layers at the2,048-token sliding window.

A conservative full-allocation boundary is6.5 GiB. It assumes all52 layers allocate the full context with16-bit KV, 2 KV heads, and head dimension128.

Planning reserve is a site assumption

This wiki adds 1.5 GiB as a small planning reserve. Meta does not publish it as a runtime requirement. Drivers, buffers, batch size, offload, cache dtype, and operating-system use can move the real peak in either direction.

Limitations

  • Published target capacity is not a minimum, certification, or guarantee for any GPU, Mac, or CPU-only system.
  • Exact file size is only one part of memory use; weights, KV cache, runtime buffers, and system usage coexist.
  • The configured context ceiling does not promise acceptable latency or a successful full-window allocation.
  • No first-party per-device peak-memory matrix was available at the last review date.

System requirement sources

  1. Official source
    Meta Research: Introducing Muse Glimmer
  2. Official source
    Official Muse Glimmer 30B model card
  3. Official source
    Official Muse Glimmer 30B GGUF model card
  4. Official source
    Official Muse Glimmer 30B GGUF file tree

LOCAL INDEX · SIX TASK PAGES

Search Muse Glimmer Wiki

Results are static internal links. Queries stay in this page and are cleared when the dialog closes.

↑ ↓ move · Tab follows native order · Enter opens the focused link