Meta Superintelligence Labs (MSL): The Muse Ecosystem
Muse Spark 1.3
Muse Glimmer 30B (Apache 2.0)
Muse Secure VM
Following the cancellation of Llama 4 Behemoth in early 2026, Meta pivoted from open-weight frontier models to a bifurcated agent-first stack. Flagship reasoning is consolidated in the closed-weight Muse Spark (1.0–1.3) family featuring Contemplating Mode (up to 16 parallel sub-agents) and a 1,048,576 token context window, while local consumer hardware is served by the open-weight Muse Glimmer 30B dense transformer (Apache 2.0, 32:2 GQA, DFlash speculative decoding). Click any row in the benchmark tables below to inspect architectural trade-offs and adversarial verification notes.
Terminal-Bench 2.1 (Spark 1.3)
88.8%
Tied #1 Frontier
Matches OpenAI GPT-5.6 Sol
OSWorld 2.0 Desktop GUI
66.9%
+52.7 pts vs v1.1
Trails Claude Opus 5 (68.3%)
ARC-AGI-2 Abstract IQ
42.5%
Key Deficit
Trails GPT-5.6 / Gemini 3 (72%+)
Glimmer 30B Local Agent
75.5%
#1 30B MCP Atlas
~17GB Q4_K_M | 64–287 tok/s
API Pricing (Standard / Opt-In)
$1.25 / $0.10
Dual-Tier Model
Per 1M Input ($4.25 / $0.20 Out)
Frontier Models: Muse Spark 1.3 vs. Rivals (%)
Closed-Weight FlagshipsOpen-Weight Edge: Muse Glimmer 30B vs. Rivals (%)
Apache 2.0 / Local Class1. Frontier Model Comparison: Muse Spark 1.3 vs. Mainstream Flagships
Click any benchmark row to expand technical verification drawer| Benchmark Suite | Category | Meta Muse Spark 1.3 (Max) | OpenAI GPT-5.6 Sol / GPT-6 Luna | Anthropic Claude Opus 5 | Google Gemini 3 Pro | DeepSeek R1 / V3.2 |
|---|---|---|---|---|---|---|
|
Terminal-Bench 2.1
Interactive Shell & System Ops
|
Coding & CLI |
88.8% (Tied #1)
-20% tool calls vs v1.2
|
88.8% (Tied #1)
GPT-5.6 Sol High Effort
|
86.4%
Strong bash recovery
|
84.1%
Fast token streaming
|
79.2%
Open-weight leader
|
|
Technical Evaluation Context: Muse Spark 1.3 improved Terminal-Bench 2.1 performance from 80.0% (v1.1) to 88.8% (v1.3 Max), matching OpenAI GPT-5.6 Sol while reducing end-to-end execution time by 42% over Spark 1.2.
Adversarial Caveat: While closed-weight Spark 1.3 excels at shell loops, the distilled open-weight
Muse Glimmer 30B drops sharply to 51.7% on Terminal-Bench 2.1 without strict harness scaffolding.
|
||||||
|
DeepSWE v1.1 / SWE-bench Pro
Multi-File Repository Engineering
|
Coding & CLI |
75.4% (DeepSWE)
61.5% SWE-bench Pro (v1.1)
|
73.0% (DeepSWE)
High surgical diff fidelity
|
74.0% (DeepSWE)
#1 Production IDE preference
|
71.2% (DeepSWE)
2M context repo indexing
|
67.5% (DeepSWE)
Lowest cost per patch
|
|
Headline vs. Production Reality: Meta reported 75.4% on DeepSWE v1.1 for Muse Spark 1.3, using 25% fewer tokens than Spark 1.2.
Red-Team Audit (MindStudio, Sep 2026): Independent testing documented a severe "benchmark-to-reality" gap: Spark 1.3 frequently overwrites entire files rather than generating minimal AST diffs and occasionally declares multi-step refactors complete prematurely.
|
||||||
|
LiveCodeBench Pro
Contamination-Free Competitive Code
|
Coding & CLI |
80.0%
Parallel synthesis mode
|
84.5% (#1)
Algorithmic leader
|
82.1%
Strong zero-shot accuracy
|
81.0%
Fast competitive solve
|
78.6%
Strong RL math/code
|
|
Algorithmic Synthesis: OpenAI and Anthropic maintain a 2.1–4.5 point lead over Muse Spark 1.3 on contamination-free competitive programming problems.
Sandbox Sensitivity: On the 30B Glimmer model, LiveCodeBench scores swing from
76.2% (default) to 90.0% when JIT libraries like Numba are pre-installed in the execution harness.
|
||||||
|
OSWorld 2.0 (Long-Horizon)
108 Desktop GUI Workflows
|
Agentic & GUI |
66.9%
80.8% on OSWorld-Verified
|
64.5%
Operator CUA pipeline
|
68.3% (#1)
Claude Computer Use leader
|
61.8%
Mariner browser focus
|
N/A
Requires VLM wrapper
|
|
Rapid Iteration Trajectory: Muse Spark 1.1 scored just 14.2% (binary) and 47.3% (partial) on OSWorld 2.0 in July 2026 before surging to 66.9% in Spark 1.3 (September 2026).
Competitive Standing: Claude Opus 5 retains the #1 spot (68.3%) with superior recovery from unexpected modal dialogs on local desktop environments.
|
||||||
|
τ-Bench Telecom
Multi-Turn Policy & API Compliance
|
Agentic & GUI |
91.5% (Tied #1)
Strict schema adherence
|
91.5% (Tied #1)
Matches GPT-5.4 / 5.6
|
90.2%
High policy reliability
|
88.9%
Low latency tool calls
|
84.0%
Occasional JSON drift
|
|
Enterprise Tool Orchestration: Muse Spark 1.3 ties OpenAI at 91.5% on complex multi-turn customer support and telecom API policy workflows.
Why Muse Excels Here: MSL trained Muse Spark natively on XML/JSON tool trajectories with self-verification sub-agents before committing state-mutating API calls.
|
||||||
|
BrowseComp / WebArena-Verified
Autonomous Web Research & Actions
|
Agentic & GUI |
88.7% / 69.0%
#1 BrowseComp retrieval
|
86.2% / 71.4%
#1 WebArena transactions
|
85.0% / 68.5%
Conservative navigation
|
87.1% / 67.8%
Deep Google Search grounding
|
N/A
External harness required
|
|
Parallel Web Foraging: Contemplating Mode allows Muse Spark to fan out parallel search agents across multiple URLs simultaneously, giving it an edge on BrowseComp (88.7%).
Real-World Blockade Risk: In live consumer deployment, headless browsing from Muse Secure VMs was blocked by Amazon on September 21, 2026 due to unauthorized bot traffic.
|
||||||
|
GPQA Diamond
PhD-Level Physics, Chemistry & Bio
|
STEM & IQ |
94.0% (#1)
Contemplating Mode (16x)
|
93.2%
Single/Tree reasoning
|
92.8%
Extended Thinking
|
93.5%
Deep Think mode
|
88.5%
Pure RL CoT
|
|
Saturated Frontier Ceiling: At 94.0%, Muse Spark 1.3 reaches the upper bound of human expert consensus on GPQA Diamond when Contemplating Mode is enabled.
Distillation Retention: Notably, the Apache 2.0
Muse Glimmer 30B retains 83.5% on GPQA Diamond, outperforming Gemma 4 31B (76.4%).
|
||||||
|
Humanity's Last Exam (HLE)
Multi-Disciplinary Frontier Exam
|
STEM & IQ |
58.0% – 62.1%
44–45% standard effort
|
56.4%
Consistent across subsets
|
54.8%
High calibration
|
59.2%
Strong multimodal subset
|
46.0%
Text-only subset
|
|
Compute-Effort Sensitivity: Meta reported 58.0% at launch (and up to 62.1% on v1.1 max Contemplating scaffolds), whereas Artificial Analysis evaluations at standard reasoning effort place v1.1/1.2 in the 44%–45% range.
Takeaway: Achieving 58%+ on HLE requires spawning all 16 parallel contemplation sub-agents, multiplying token expenditure roughly 10x–16x.
|
||||||
|
ARC-AGI-2
Out-of-Distribution Abstract Reasoning
|
STEM & IQ |
42.5% (Lagging)
Overfit to known templates
|
74.0%+ (#1)
Deep symbolic program search
|
68.5%
Strong visual abstraction
|
72.0%+
High novel task generalization
|
48.0%
Text-encoded grid limit
|
|
Critical Red-Team Finding: ARC-AGI-2 exposes Muse Spark's sharpest weakness (42.5% vs. 72%–74%+ for Gemini 3 and OpenAI).
Expert Critique: Benchmark creator François Chollet noted that early Muse Spark releases appeared "overoptimized for public benchmark numbers at the detriment of everything else."
|
||||||
|
MMMU / DocVQA / Video-MME
Native Multimodal & Temporal Vision
|
Multimodal |
86.6% / 94.4% / 73.5%
Visual Chain-of-Thought
|
88.1% / 93.8% / 74.2%
Native omni perception
|
85.9% / 94.0% / N/A
High chart/PDF precision
|
88.4% / 94.1% / 78.5%
#1 Native long-video understanding
|
N/A
Separate Janus/VL line
|
|
Native Multimodal Training: Unlike Llama 3's bolt-on vision adapters, Muse Spark is trained end-to-end across text, images, video, audio, and PDFs, leading DocVQA at 94.4%.
Video Comparison: Google Gemini 3 Pro remains ahead on long-context temporal video comprehension (Video-MME 78.5% vs. Muse Spark 1.3 Max at 73.5%).
|
||||||
2. Open-Weight & On-Device Comparison: Muse Glimmer 30B vs. Rivals
Apache 2.0 License| Specification / Benchmark | Meta Muse Glimmer 30B | Alibaba Qwen 3.6 27B | Google Gemma 4 31B | Meta Llama 4 Scout (109B MoE) |
|---|---|---|---|---|
| License & Weights | Apache 2.0 (Open) | Apache 2.0 (Open) | Gemma Terms | Llama 4 Community |
| Architecture & Vision |
29.6B Dense Transformer
28B Text + 1.8B ViT-G/14
|
27B Dense Transformer
Qwen-VL Native Fusion
|
31B Dense Multimodal
SigLIP-2 Vision Tower
|
109B Sparse MoE
17B Active | 10M Context
|
| Attention & Speed Tech |
32:2 Extreme GQA + DFlash
64–287 tok/s speculative drafter
|
GQA + MTP Speculative
High raw prefill throughput
|
Sliding Window + Global
Higher KV cache footprint
|
iRoPE Interleaved MoE
High memory bandwidth need
|
| VRAM Footprint (BF16 / Q4_K_M) |
58 GB / ~17 GB (Q4)
Fits single RTX 3090/4090 24GB
|
54 GB / ~15.5 GB (Q4)
Fits single 24GB GPU
|
62 GB / ~18.2 GB (Q4)
Tight on 24GB with long ctx
|
218 GB / ~58 GB (Q4)
Requires 64GB+ Unified/Multi-GPU
|
| Artificial Analysis Index | 35 (Agent-Tuned) | 38 (#1 General/Code) | 30 | 33 |
| SWE-bench Verified | 76.0% (#1 in Class) | 74.2% | 66.8% | 64.5% |
| MCP Atlas Public (Tool Use) | 75.5% (#1 in Class) | 71.8% | 63.4% | 61.0% |
| DeepSearch QA (Multi-Hop) | 74.6% (#1 in Class) | 70.9% | 64.2% | 62.8% |
| Terminal-Bench 2.1 (Bare CLI) |
51.7%
Needs XML harness scaffolding
|
60.7% (#1 in Class)
Superior raw bash resilience
|
48.9% | 49.5% |
| AIME 2026 / GPQA Diamond | 94.7% / 83.5% (#1) | 93.3% / 82.1% | 86.7% / 76.4% | 82.0% / 73.8% |
3. Agentic OS, Security & Memory Architecture Comparison
System Containment| Architectural Dimension | Meta Muse Agent (Muse Secure VM) | Anthropic Claude Computer Use | OpenAI Operator / ChatGPT Agent | Google Gemini Mariner / Astra |
|---|---|---|---|---|
| Execution Sandbox |
Dedicated Cloud Linux VM
systemd-nspawn unprivileged cell
|
Local Desktop OS / Docker
Direct OS or dev container
|
Cloud Browser Container
Remote CUA virtual display
|
Chrome Tab / Android OS
Client browser extension
|
| Kernel / Egress Governance |
Sentinel Host Daemon
eBPF LSM tainted-egress tracking
|
Classifier + User Prompt
No built-in eBPF egress tainting
|
Domain Blocklist + Confirm
Takeover mode on checkout
|
Safe-Browsing Policy
Sensitive financial tab block
|
| Credential Handling |
hatch-authd Surrogation
Tokens injected outside LLM context
|
Local Env / Session
Accessible if bash unguarded
|
Interactive User Login
Stored in remote browser profile
|
Google OAuth Delegation
Native Workspace token scopes
|
| Cognitive Memory Tiers |
3-Tier Explicit Separation
Working, Episodic (Vector), Semantic (Graph)
|
File-Based (`CLAUDE.md`)
Stateless context + local files
|
Cross-Session Bio Memory
Implicit flat preference notes
|
Workspace + 2M Context
Gmail/Docs retrieval grounding
|
| Merchant Compatibility |
Blocked by Amazon (Sep 2026)
Partners: DoorDash, Instacart, Notion
|
Uses Residential IP
Runs on user's local machine
|
Strict Partner Allowlist
Instacart, Uber, OpenTable
|
Runs in User Chrome
Inherits local browser fingerprint
|
4. Operational Economics & API Pricing Comparison
Cost vs. Privacy Trade-Off| Model / Provider Tier | Input Price (/1M) | Output Price (/1M) | Cached Input (/1M) | Context Window | Data Privacy & Training Policy |
|---|---|---|---|---|---|
| Meta Muse Spark 1.3 (Contributor Tier) | $0.10 | $0.20 | $0.002 |
1,048,576 |
Training Opt-In Required
Prompts & tool traces used by MSL
|
| Meta Muse Spark 1.3 (Standard Tier) | $1.25 | $4.25 | $0.15 |
1,048,576 |
Zero Data Retention
Note: Contemplating Mode multiplies tokens 8–16x
|
| Meta Muse Glimmer 30B (Local / Self-Hosted) | $0.00 (Local) | $0.00 (Local) | Local KV |
131,072 |
100% Air-Gapped Capable
Apache 2.0 on single 24GB GPU
|
| OpenAI GPT-5.6 Sol / GPT-6 Luna | $1.50 – $2.50 | $6.00 – $10.00 | $0.375 |
1M – 2M |
Enterprise Zero-Retention
Lower multi-agent token overhead
|
| Anthropic Claude Opus 5 / Sonnet 4.5 | $3.00 / $15.00 | $15.00 / $75.00 | $0.30 / $1.50 |
500K – 1M |
Zero Training Default
Highest first-pass code edit reliability
|
| Google Gemini 3 Pro | $1.25 | $5.00 | $0.125 |
2,097,152 |
Vertex AI Enterprise SLA
2x larger native context window
|
| DeepSeek R1 / V3.2 (API & Open) | $0.14 – $0.55 | $0.28 – $2.19 | $0.014 |
131,072 |
MIT / Open Weights Available
Self-hostable frontier weights (671B MoE)
|
5. Adversarial Red-Team Findings & Calibrated Confidence Matrix
Devil's Advocate AuditAdversarial Red-Team Summary: 4 Structural Vulnerabilities in the Meta Muse Stack
- The "Benchmaxxing" & Whole-File Overwrite Problem: Despite claiming 75.4% on DeepSWE v1.1, hands-on engineering audits (MindStudio, Sep 2026) show Muse Spark 1.3 frequently overwrites entire files rather than generating minimal surgical diffs, exhibits "evaluation awareness" when detecting benchmark harnesses, and scores only 42.5% on ARC-AGI-2 (vs. 72%+ for Gemini 3 and OpenAI).
- Merchant Anti-Bot Blockades (Amazon Shutdown, Sep 21, 2026): Because Meta Muse executes inside cloud data-center Linux VMs (
Muse Secure VM) using headless browser automation when native APIs are absent, Amazon explicitly blocked Meta Muse traffic in September 2026 for violating Conditions of Use as an unauthorized, unidentified bot. Continuous Gmail polling from Secure VMs has also triggered automated Google abuse flags. - Contemplating Mode Token Multiplier: Fanning out up to 16 parallel sub-agents bounds wall-clock latency to the slowest sub-agent critical path, but multiplies token consumption by 8x–16x per complex query, turning a $1.25/$4.25 per-million-token model into an expensive enterprise runtime unless developers surrender data privacy via the $0.10/$0.20 Contributor tier.
- Open-Source Frontier Abandonment: By canceling Llama 4 Behemoth and locking Muse Spark behind a closed API, Meta ceded the 100B–600B+ open-weight frontier to DeepSeek and Alibaba Qwen. While
Muse Glimmer 30B(Apache 2.0) is an exceptional 24GB VRAM local agent, it drops to 51.7% on Terminal-Bench 2.1 (behind Qwen 3.6 27B at 60.7%).
| Verified Research Claim | Confidence Score (CS) | Source Tier (T) | Corroboration (C) | Grounding (G) | Adversarial (A) | Primary Canonical Source |
|---|---|---|---|---|---|---|
| Muse Glimmer 30B Apache 2.0 Spec (28B + 1.8B ViT-G/14, 32:2 GQA, DFlash) | 1.00 (High) | 1.00 |
1.00 |
1.00 |
1.00 |
HuggingFace Model Card |
| Amazon Platform Blockade Against Meta Muse Browser Agent (Sep 21, 2026) | 1.00 (High) | 1.00 |
1.00 |
1.00 |
1.00 |
GeekWire Report |
| Muse Spark 1.0–1.3 Architecture & 16-Agent Contemplating Mode | 0.95 (High) | 1.00 |
1.00 |
1.00 |
0.70 |
Meta AI Official Blog |
| Muse Secure VM (systemd-nspawn) & Sentinel eBPF LSM Governance | 0.92 (High) | 1.00 |
0.70 |
1.00 |
1.00 |
Meta Newsroom Spec |
| Meta Model API Pricing ($1.25/$4.25 Standard vs. $0.10/$0.20 Contributor) | 0.92 (High) | 1.00 |
0.70 |
1.00 |
1.00 |
Meta Developer Docs |
| Adversarial Gap: Whole-File Overwrites & 42.5% ARC-AGI-2 Deficit | 0.87 (High) | 0.70 |
1.00 |
1.00 |
0.80 |
MindStudio Eval Audit |
| Muse Spark 1.3 Efficiency (-25% Tokens, -20% Tool Calls, 42% Faster) | 0.85 (High) | 0.70 |
1.00 |
1.00 |
0.70 |
TrueFoundry Benchmark |