Launch evidence, not a leaderboard

Muse Glimmer Benchmarks: What Meta Tested

LAST REVIEWEDSOURCE-BACKED

Direct answer

The results below were published by Meta and are not an independent reproduction. Meta's launch comparison image contains 24 rows across agentic, coding, multimodal, safety, and general-reasoning categories. Values are preserved row by row with their display metric; they are not converted into a single ranking.

Complete launch-image transcription

Every comparison row, including mixed and lower-is-better metrics

Scope of this transcription

The 24 rows reproduce the public launch comparison image only. The separate methodology document discusses additional evaluation suites that are not rows in that image. “Benchmark-native value” means the original benchmark's unit is retained; it does not turn every value into a percentage.

Open Meta's source comparison image

General Agentic

8 rows

Wide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.

General Agentic values displayed in Meta's launch comparison image.
BenchmarkDisplay metricMuse Glimmer-30BGemma4-31BQwen3.6-27B
MCP Atlas (Public)Benchmark-native value75.554.262.5
DeepSearch QABenchmark-native value74.661.771.1
τ³-BankingBenchmark-native value23.515.116.7
WildClawBenchBenchmark-native value47.637.643.2
GDPVal-AA v2Benchmark-native value9538111141
Gaia2Benchmark-native value43.336.440.0
SkillsBench (with skills)Benchmark-native value44.332.446.6
OSWorld-VerifiedBenchmark-native value65.958.575.6

Agentic Coding

4 rows

Wide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.

Agentic Coding values displayed in Meta's launch comparison image.
BenchmarkDisplay metricMuse Glimmer-30BGemma4-31BQwen3.6-27B
SWE-Bench ProBenchmark-native value51.236.950.2
SWE-Bench VerifiedBenchmark-native value76.066.677.2
TerminalBench 2.1 (with terminus2)Benchmark-native value51.743.460.7
SciCodeBenchmark-native value43.643.439.8

Multimodal

4 rows

Wide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.

Multimodal values displayed in Meta's launch comparison image.
BenchmarkDisplay metricMuse Glimmer-30BGemma4-31BQwen3.6-27B
Charxiv ReasoningBenchmark-native value78.877.778.4
ScreenSpot ProBenchmark-native value75.475.976.1
OmniDocBench v1.5Benchmark-native value75.872.577.8
MMMU ProBenchmark-native value747375

Safety

2 rows

Wide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.

Safety values displayed in Meta's launch comparison image.
BenchmarkDisplay metricMuse Glimmer-30BGemma4-31BQwen3.6-27B
CI MemoriesViolation (↓); Coverage26.4; 64.812.1; 53.053.4; 66.9
Siren AgentDojoAttack Success Rate (↓); Utility28.4; 94.225.6; 90.840.3; 92.7

Each safety cell contains two metrics. The down arrow applies only to the first metric; the second is coverage or utility. The pairs are intentionally not winner-highlighted or collapsed.

General Capabilities and Reasoning

6 rows

Wide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.

General Capabilities and Reasoning values displayed in Meta's launch comparison image.
BenchmarkDisplay metricMuse Glimmer-30BGemma4-31BQwen3.6-27B
IFBenchBenchmark-native value77.076.070.8
AIME 2026Benchmark-native value94.789.294.1
GPQA Diamond (AA)Benchmark-native value83.585.784.2
HLE Text (AA)Benchmark-native value22.023.623.1
AA-LCRBenchmark-native value80.068.373.3
Beam128KBenchmark-native value65.158.263.0

Meta methodology

Generation settings and comparator policy affect the table

Muse Glimmerhigh reasoning strength

Temperature 1, top-p 0.95, top-k 64.

Gemma4-31Bthinking mode

Temperature 1, top-p 0.95, top-k 64.

Qwen3.6-27Bthinking mode

Temperature 1, top-p 0.95, top-k 20;GAIA2 and WildClawBench use temperature 0.6.

Comparator selection

better of self-report and Meta internal reproduction. Artificial Analysis values are used only when all three compared models have a value.

Third-party agent harness caveat

Third-party agentic results use the same evaluation framework as internal models, but tools and system prompts may not be optimized for proprietary third-party models.

Source-reported speed bars

DFlash speedup is tied to one quantized configuration

K-Quant 17 GB + DFlash versus standard decoding

Meta says the speed test covers 7 prompt categories. The published multiplier is retained as reported; it is not recalculated from the rounded bar labels. A multiplier is not tokens per second and should not be applied to another setup.

Batch size 1 with greedy decoding.

Wide table: swipe or use Shift + mouse wheel to inspect every condition.

Meta-reported standard-decoding and DFlash bars for the published K-Quant configuration.
DeviceRuntimeStandardDFlashReported speedup
NVIDIA RTX 5090llama.cpp74.9 tok/s233.4 tok/s3.1×
Apple M5 MaxExecuTorch26.6 tok/s50.2 tok/s1.8×
Apple M4 MaxExecuTorch23.7 tok/s37.8 tok/s1.5×

Meta did not disclose context length, RAM, CPU core count, or exact runtime version for these speed bars.

Meta average degradation

Quantization summaries are averages, not task guarantees

Official GGUFK-Quant Dynamic

0.2% average degradation reported by Meta.

Published target: 32 GiB.

Official GGUFK-Quant 17 GB

1% average degradation reported by Meta.

Published target: 24 GiB.

InterpretationInspect the task you care about

An average can hide larger movement on an individual benchmark, modality, or safety measure.

Limitations

  • These are Meta-published launch results, not measurements independently reproduced by this wiki.
  • The comparison mixes benchmark-native units; row values should be read with each benchmark's own protocol.
  • Safety rows contain paired metrics with different directions and are not reducible to one number.
  • Third-party models may not receive tool and system-prompt tuning specific to their proprietary interfaces.
  • Meta did not disclose context length, RAM, CPU core count, or exact runtime version for these speed bars.
  • Local hardware fit, peak memory, latency, and output quality require separate, reproducible testing.

Benchmark sources

  1. Official source
    Meta Research: Introducing Muse Glimmer
  2. Official source
    Meta Research: Muse Glimmer methodology
  3. Official source
    Official Muse Glimmer 30B GGUF model card

LOCAL INDEX · SIX TASK PAGES

Search Muse Glimmer Wiki

Results are static internal links. Queries stay in this page and are cleared when the dialog closes.

↑ ↓ move · Tab follows native order · Enter opens the focused link