Complete launch-image transcription
Every comparison row, including mixed and lower-is-better metrics
Scope of this transcription
The 24 rows reproduce the public launch comparison image only. The separate methodology document discusses additional evaluation suites that are not rows in that image. “Benchmark-native value” means the original benchmark's unit is retained; it does not turn every value into a percentage.
General Agentic
8 rowsWide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.
| Benchmark | Display metric | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B |
|---|---|---|---|---|
| MCP Atlas (Public) | Benchmark-native value | 75.5 | 54.2 | 62.5 |
| DeepSearch QA | Benchmark-native value | 74.6 | 61.7 | 71.1 |
| τ³-Banking | Benchmark-native value | 23.5 | 15.1 | 16.7 |
| WildClawBench | Benchmark-native value | 47.6 | 37.6 | 43.2 |
| GDPVal-AA v2 | Benchmark-native value | 953 | 811 | 1141 |
| Gaia2 | Benchmark-native value | 43.3 | 36.4 | 40.0 |
| SkillsBench (with skills) | Benchmark-native value | 44.3 | 32.4 | 46.6 |
| OSWorld-Verified | Benchmark-native value | 65.9 | 58.5 | 75.6 |
Agentic Coding
4 rowsWide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.
| Benchmark | Display metric | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B |
|---|---|---|---|---|
| SWE-Bench Pro | Benchmark-native value | 51.2 | 36.9 | 50.2 |
| SWE-Bench Verified | Benchmark-native value | 76.0 | 66.6 | 77.2 |
| TerminalBench 2.1 (with terminus2) | Benchmark-native value | 51.7 | 43.4 | 60.7 |
| SciCode | Benchmark-native value | 43.6 | 43.4 | 39.8 |
Multimodal
4 rowsWide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.
| Benchmark | Display metric | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B |
|---|---|---|---|---|
| Charxiv Reasoning | Benchmark-native value | 78.8 | 77.7 | 78.4 |
| ScreenSpot Pro | Benchmark-native value | 75.4 | 75.9 | 76.1 |
| OmniDocBench v1.5 | Benchmark-native value | 75.8 | 72.5 | 77.8 |
| MMMU Pro | Benchmark-native value | 74 | 73 | 75 |
Safety
2 rowsWide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.
| Benchmark | Display metric | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B |
|---|---|---|---|---|
| CI Memories | Violation (↓); Coverage | 26.4; 64.8 | 12.1; 53.0 | 53.4; 66.9 |
| Siren AgentDojo | Attack Success Rate (↓); Utility | 28.4; 94.2 | 25.6; 90.8 | 40.3; 92.7 |
Each safety cell contains two metrics. The down arrow applies only to the first metric; the second is coverage or utility. The pairs are intentionally not winner-highlighted or collapsed.
General Capabilities and Reasoning
6 rowsWide table: swipe or use Shift + mouse wheel. Values are not sorted by apparent magnitude.
| Benchmark | Display metric | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B |
|---|---|---|---|---|
| IFBench | Benchmark-native value | 77.0 | 76.0 | 70.8 |
| AIME 2026 | Benchmark-native value | 94.7 | 89.2 | 94.1 |
| GPQA Diamond (AA) | Benchmark-native value | 83.5 | 85.7 | 84.2 |
| HLE Text (AA) | Benchmark-native value | 22.0 | 23.6 | 23.1 |
| AA-LCR | Benchmark-native value | 80.0 | 68.3 | 73.3 |
| Beam128K | Benchmark-native value | 65.1 | 58.2 | 63.0 |
Meta methodology
Generation settings and comparator policy affect the table
Temperature 1, top-p 0.95, top-k 64.
Temperature 1, top-p 0.95, top-k 64.
Temperature 1, top-p 0.95, top-k 20;GAIA2 and WildClawBench use temperature 0.6.
Comparator selection
better of self-report and Meta internal reproduction. Artificial Analysis values are used only when all three compared models have a value.
Third-party agent harness caveat
Third-party agentic results use the same evaluation framework as internal models, but tools and system prompts may not be optimized for proprietary third-party models.
Source-reported speed bars
DFlash speedup is tied to one quantized configuration
K-Quant 17 GB + DFlash versus standard decoding
Meta says the speed test covers 7 prompt categories. The published multiplier is retained as reported; it is not recalculated from the rounded bar labels. A multiplier is not tokens per second and should not be applied to another setup.
Batch size 1 with greedy decoding.
Wide table: swipe or use Shift + mouse wheel to inspect every condition.
| Device | Runtime | Standard | DFlash | Reported speedup |
|---|---|---|---|---|
| NVIDIA RTX 5090 | llama.cpp | 74.9 tok/s | 233.4 tok/s | 3.1× |
| Apple M5 Max | ExecuTorch | 26.6 tok/s | 50.2 tok/s | 1.8× |
| Apple M4 Max | ExecuTorch | 23.7 tok/s | 37.8 tok/s | 1.5× |
Meta did not disclose context length, RAM, CPU core count, or exact runtime version for these speed bars.
Meta average degradation
Quantization summaries are averages, not task guarantees
0.2% average degradation reported by Meta.
Published target: 32 GiB.
1% average degradation reported by Meta.
Published target: 24 GiB.
An average can hide larger movement on an individual benchmark, modality, or safety measure.
Limitations
- These are Meta-published launch results, not measurements independently reproduced by this wiki.
- The comparison mixes benchmark-native units; row values should be read with each benchmark's own protocol.
- Safety rows contain paired metrics with different directions and are not reducible to one number.
- Third-party models may not receive tool and system-prompt tuning specific to their proprietary interfaces.
- Meta did not disclose context length, RAM, CPU core count, or exact runtime version for these speed bars.
- Local hardware fit, peak memory, latency, and output quality require separate, reproducible testing.
Put the claims in context
- Estimate hardware fit
Use artifact sizes and KV boundaries rather than benchmark scores to screen capacity.
- Compare official GGUF files
Put Meta’s average degradation claims next to exact artifact sizes and targets.
- Plan memory and context
Separate a published capacity target from a measured device result.
- Check runtime support
Verify the upstream merge before interpreting local speed.