UPDATED AUG 10, 2026 · SOURCE-BACKED

Muse Glimmer 30B: Can You Run It Locally?

Start with the published artifact sizes, Meta's capacity targets, and a runtime-aware KV range. The estimator answers hardware fit without pretending a target tier is a tested machine.

Architecture
29.6B dense
Config limit
131K context
Inputs
Text + Image
Terms
Apache-2.0 + Usage Policy

Facts checked against the Meta introduction, official model card, and Usage Policy.

Hardware-fit estimator

Can I run it?

Local calculation
GiB

Use unified memory or VRAM available to the model.

Optional components

DFlash is a K-Quant companion and is disabled for BF16.

Result

Official target

Estimate, not a guarantee

Meets Meta's published capacity target for this variant. This is a target tier, not verified device compatibility.

Active allocation

21.63 GiB active / 26.43 GiB conservative

Capacity 24 GiB

Meta target 24 GiB

  • Weights15.61 GiB
  • Vision1.3 GiB
  • DFlash1.52 GiB
  • Active KV1.7 GiB
  • Planning reserve1.5 GiB
  • Headroom2.37 GiB
Active KV
1.7 GiB
Conservative KV
6.5 GiB
Active headroom
2.37 GiB
Conservative headroom
-2.43 GiB
How this range is planned

The site adds a 1.5 GiB planning reserve for runtime overhead. It is a site assumption, not an official requirement.

Active KV follows the model's global/local attention pattern. The conservative total assumes every layer allocates the full selected context, so actual runtimes can land between the two estimates.

Evidence: Meta published target tier

Official target means Meta's published capacity tier, not a verified device result.

  • Capacity arithmetic uses binary GiB (1 GiB = 1,073,741,824 bytes).
  • 1.5 GiB is reserved for runtime overhead as a site planning assumption, not an official requirement.
  • Active KV uses 13 global layers at the selected context and 39 local layers capped at 2,048 tokens.
  • KV arithmetic uses 16-bit KV (2 bytes per element) with 2 KV heads and head dimension 128; common F16/BF16 cache allocation is covered, but runtime cache type and allocation may vary.
  • Conservative KV assumes all 52 layers allocate the full selected context.

Published capacity tiers

Official targets

Official claim
Target, not minimum: these are Meta's published capacity tiers, not guaranteed device compatibility.
Memory targetVariantMeta quality notePlanning note
24 GiBK-Quant 17 GB1% average degradationMeta target for the smaller official quant; task-level quality can differ from the reported average.
32 GiBK-Quant Dynamic0.2% average degradationMeta target for the dynamic official quant; additional memory can absorb runtime-dependent KV allocation.
64 GiBBF16Reference weightsMeta target for the official BF16 weights; this capacity tier is not a verified device result.

Runtime watch

What is verified right now

Checked Aug 10, 2026

Official GGUF

Available

Meta publishes the K-Quant files and companion artifacts in the official repository.

Inspect official files

llama.cpp

Merged to master

Muse Glimmer support was merged at commit 62bf73d. Older builds may fail before model loading or lack the required support.

Review merge commit

Other integrations

Unknown

Version-specific compatibility for integrations not backed by a dated source has not been verified here. Availability is not inferred.

Method and limits

What the estimate can - and cannot - tell you

The calculator adds exact published file sizes, selected companion artifacts, a KV-cache estimate derived from the model configuration, and a small site planning reserve. It reports active and conservative totals because allocation behavior depends on the runtime.

It is a capacity screen for choosing what to investigate next. It does not measure speed, prove that a particular GPU or Mac works, or replace build-specific runtime documentation.

  1. 01

    Estimate, not a guarantee. No device is certified by this tool.

  2. 02

    Target, not minimum. Official target means Meta's capacity tier.

  3. 03

    Reserve is local policy. The 1.5 GiB runtime reserve is this site's planning assumption.

  4. 04

    KV varies. Runtime cache type and allocation can move the real total.

  5. 05

    Support is versioned. Unsourced integration availability remains unknown.

Key sources

  1. Official source
    Meta Research: Introducing Muse Glimmer
  2. Official source
    Official Muse Glimmer 30B model card
  3. Official source
    Official Muse Glimmer 30B GGUF model card
  4. Official source
    Official Muse Glimmer 30B GGUF file tree
  5. Official source
    Muse Glimmer Usage Policy
  6. Upstream source
    llama.cpp PR #26841: Muse Glimmer support
  7. Upstream source
    llama.cpp Muse Glimmer merge commit

LOCAL INDEX · SIX TASK PAGES

Search Muse Glimmer Wiki

Results are static internal links. Queries stay in this page and are cleared when the dialog closes.

↑ ↓ move · Tab follows native order · Enter opens the focused link