Meta Muse vs. Mainstream AI Models
Meta Superintelligence Labs (MSL): The Muse Ecosystem Muse Spark 1.3 Muse Glimmer 30B (Apache 2.0) Muse Secure VM

Following the cancellation of Llama 4 Behemoth in early 2026, Meta pivoted from open-weight frontier models to a bifurcated agent-first stack. Flagship reasoning is consolidated in the closed-weight Muse Spark (1.0–1.3) family featuring Contemplating Mode (up to 16 parallel sub-agents) and a 1,048,576 token context window, while local consumer hardware is served by the open-weight Muse Glimmer 30B dense transformer (Apache 2.0, 32:2 GQA, DFlash speculative decoding). Click any row in the benchmark tables below to inspect architectural trade-offs and adversarial verification notes.

Terminal-Bench 2.1 (Spark 1.3) 88.8%
Tied #1 Frontier Matches OpenAI GPT-5.6 Sol
OSWorld 2.0 Desktop GUI 66.9%
+52.7 pts vs v1.1 Trails Claude Opus 5 (68.3%)
ARC-AGI-2 Abstract IQ 42.5%
Key Deficit Trails GPT-5.6 / Gemini 3 (72%+)
Glimmer 30B Local Agent 75.5%
#1 30B MCP Atlas ~17GB Q4_K_M | 64–287 tok/s
API Pricing (Standard / Opt-In) $1.25 / $0.10
Dual-Tier Model Per 1M Input ($4.25 / $0.20 Out)

Frontier Models: Muse Spark 1.3 vs. Rivals (%)

Closed-Weight Flagships

Open-Weight Edge: Muse Glimmer 30B vs. Rivals (%)

Apache 2.0 / Local Class

1. Frontier Model Comparison: Muse Spark 1.3 vs. Mainstream Flagships

Click any benchmark row to expand technical verification drawer
Benchmark Suite Category Meta Muse Spark 1.3 (Max) OpenAI GPT-5.6 Sol / GPT-6 Luna Anthropic Claude Opus 5 Google Gemini 3 Pro DeepSeek R1 / V3.2
Terminal-Bench 2.1 Interactive Shell & System Ops
Coding & CLI
88.8% (Tied #1) -20% tool calls vs v1.2
88.8% (Tied #1) GPT-5.6 Sol High Effort
86.4% Strong bash recovery
84.1% Fast token streaming
79.2% Open-weight leader
Technical Evaluation Context: Muse Spark 1.3 improved Terminal-Bench 2.1 performance from 80.0% (v1.1) to 88.8% (v1.3 Max), matching OpenAI GPT-5.6 Sol while reducing end-to-end execution time by 42% over Spark 1.2.
Adversarial Caveat: While closed-weight Spark 1.3 excels at shell loops, the distilled open-weight Muse Glimmer 30B drops sharply to 51.7% on Terminal-Bench 2.1 without strict harness scaffolding.
DeepSWE v1.1 / SWE-bench Pro Multi-File Repository Engineering
Coding & CLI
75.4% (DeepSWE) 61.5% SWE-bench Pro (v1.1)
73.0% (DeepSWE) High surgical diff fidelity
74.0% (DeepSWE) #1 Production IDE preference
71.2% (DeepSWE) 2M context repo indexing
67.5% (DeepSWE) Lowest cost per patch
Headline vs. Production Reality: Meta reported 75.4% on DeepSWE v1.1 for Muse Spark 1.3, using 25% fewer tokens than Spark 1.2.
Red-Team Audit (MindStudio, Sep 2026): Independent testing documented a severe "benchmark-to-reality" gap: Spark 1.3 frequently overwrites entire files rather than generating minimal AST diffs and occasionally declares multi-step refactors complete prematurely.
LiveCodeBench Pro Contamination-Free Competitive Code
Coding & CLI
80.0% Parallel synthesis mode
84.5% (#1) Algorithmic leader
82.1% Strong zero-shot accuracy
81.0% Fast competitive solve
78.6% Strong RL math/code
Algorithmic Synthesis: OpenAI and Anthropic maintain a 2.1–4.5 point lead over Muse Spark 1.3 on contamination-free competitive programming problems.
Sandbox Sensitivity: On the 30B Glimmer model, LiveCodeBench scores swing from 76.2% (default) to 90.0% when JIT libraries like Numba are pre-installed in the execution harness.
OSWorld 2.0 (Long-Horizon) 108 Desktop GUI Workflows
Agentic & GUI
66.9% 80.8% on OSWorld-Verified
64.5% Operator CUA pipeline
68.3% (#1) Claude Computer Use leader
61.8% Mariner browser focus
N/A Requires VLM wrapper
Rapid Iteration Trajectory: Muse Spark 1.1 scored just 14.2% (binary) and 47.3% (partial) on OSWorld 2.0 in July 2026 before surging to 66.9% in Spark 1.3 (September 2026).
Competitive Standing: Claude Opus 5 retains the #1 spot (68.3%) with superior recovery from unexpected modal dialogs on local desktop environments.
τ-Bench Telecom Multi-Turn Policy & API Compliance
Agentic & GUI
91.5% (Tied #1) Strict schema adherence
91.5% (Tied #1) Matches GPT-5.4 / 5.6
90.2% High policy reliability
88.9% Low latency tool calls
84.0% Occasional JSON drift
Enterprise Tool Orchestration: Muse Spark 1.3 ties OpenAI at 91.5% on complex multi-turn customer support and telecom API policy workflows.
Why Muse Excels Here: MSL trained Muse Spark natively on XML/JSON tool trajectories with self-verification sub-agents before committing state-mutating API calls.
BrowseComp / WebArena-Verified Autonomous Web Research & Actions
Agentic & GUI
88.7% / 69.0% #1 BrowseComp retrieval
86.2% / 71.4% #1 WebArena transactions
85.0% / 68.5% Conservative navigation
87.1% / 67.8% Deep Google Search grounding
N/A External harness required
Parallel Web Foraging: Contemplating Mode allows Muse Spark to fan out parallel search agents across multiple URLs simultaneously, giving it an edge on BrowseComp (88.7%).
Real-World Blockade Risk: In live consumer deployment, headless browsing from Muse Secure VMs was blocked by Amazon on September 21, 2026 due to unauthorized bot traffic.
GPQA Diamond PhD-Level Physics, Chemistry & Bio
STEM & IQ
94.0% (#1) Contemplating Mode (16x)
93.2% Single/Tree reasoning
92.8% Extended Thinking
93.5% Deep Think mode
88.5% Pure RL CoT
Saturated Frontier Ceiling: At 94.0%, Muse Spark 1.3 reaches the upper bound of human expert consensus on GPQA Diamond when Contemplating Mode is enabled.
Distillation Retention: Notably, the Apache 2.0 Muse Glimmer 30B retains 83.5% on GPQA Diamond, outperforming Gemma 4 31B (76.4%).
Humanity's Last Exam (HLE) Multi-Disciplinary Frontier Exam
STEM & IQ
58.0% – 62.1% 44–45% standard effort
56.4% Consistent across subsets
54.8% High calibration
59.2% Strong multimodal subset
46.0% Text-only subset
Compute-Effort Sensitivity: Meta reported 58.0% at launch (and up to 62.1% on v1.1 max Contemplating scaffolds), whereas Artificial Analysis evaluations at standard reasoning effort place v1.1/1.2 in the 44%–45% range.
Takeaway: Achieving 58%+ on HLE requires spawning all 16 parallel contemplation sub-agents, multiplying token expenditure roughly 10x–16x.
ARC-AGI-2 Out-of-Distribution Abstract Reasoning
STEM & IQ
42.5% (Lagging) Overfit to known templates
74.0%+ (#1) Deep symbolic program search
68.5% Strong visual abstraction
72.0%+ High novel task generalization
48.0% Text-encoded grid limit
Critical Red-Team Finding: ARC-AGI-2 exposes Muse Spark's sharpest weakness (42.5% vs. 72%–74%+ for Gemini 3 and OpenAI).
Expert Critique: Benchmark creator François Chollet noted that early Muse Spark releases appeared "overoptimized for public benchmark numbers at the detriment of everything else."
MMMU / DocVQA / Video-MME Native Multimodal & Temporal Vision
Multimodal
86.6% / 94.4% / 73.5% Visual Chain-of-Thought
88.1% / 93.8% / 74.2% Native omni perception
85.9% / 94.0% / N/A High chart/PDF precision
88.4% / 94.1% / 78.5% #1 Native long-video understanding
N/A Separate Janus/VL line
Native Multimodal Training: Unlike Llama 3's bolt-on vision adapters, Muse Spark is trained end-to-end across text, images, video, audio, and PDFs, leading DocVQA at 94.4%.
Video Comparison: Google Gemini 3 Pro remains ahead on long-context temporal video comprehension (Video-MME 78.5% vs. Muse Spark 1.3 Max at 73.5%).

2. Open-Weight & On-Device Comparison: Muse Glimmer 30B vs. Rivals

Apache 2.0 License
Specification / Benchmark Meta Muse Glimmer 30B Alibaba Qwen 3.6 27B Google Gemma 4 31B Meta Llama 4 Scout (109B MoE)
License & Weights Apache 2.0 (Open) Apache 2.0 (Open) Gemma Terms Llama 4 Community
Architecture & Vision
29.6B Dense Transformer 28B Text + 1.8B ViT-G/14
27B Dense Transformer Qwen-VL Native Fusion
31B Dense Multimodal SigLIP-2 Vision Tower
109B Sparse MoE 17B Active | 10M Context
Attention & Speed Tech
32:2 Extreme GQA + DFlash 64–287 tok/s speculative drafter
GQA + MTP Speculative High raw prefill throughput
Sliding Window + Global Higher KV cache footprint
iRoPE Interleaved MoE High memory bandwidth need
VRAM Footprint (BF16 / Q4_K_M)
58 GB / ~17 GB (Q4) Fits single RTX 3090/4090 24GB
54 GB / ~15.5 GB (Q4) Fits single 24GB GPU
62 GB / ~18.2 GB (Q4) Tight on 24GB with long ctx
218 GB / ~58 GB (Q4) Requires 64GB+ Unified/Multi-GPU
Artificial Analysis Index 35 (Agent-Tuned) 38 (#1 General/Code) 30 33
SWE-bench Verified 76.0% (#1 in Class) 74.2% 66.8% 64.5%
MCP Atlas Public (Tool Use) 75.5% (#1 in Class) 71.8% 63.4% 61.0%
DeepSearch QA (Multi-Hop) 74.6% (#1 in Class) 70.9% 64.2% 62.8%
Terminal-Bench 2.1 (Bare CLI)
51.7% Needs XML harness scaffolding
60.7% (#1 in Class) Superior raw bash resilience
48.9% 49.5%
AIME 2026 / GPQA Diamond 94.7% / 83.5% (#1) 93.3% / 82.1% 86.7% / 76.4% 82.0% / 73.8%

3. Agentic OS, Security & Memory Architecture Comparison

System Containment
Architectural Dimension Meta Muse Agent (Muse Secure VM) Anthropic Claude Computer Use OpenAI Operator / ChatGPT Agent Google Gemini Mariner / Astra
Execution Sandbox
Dedicated Cloud Linux VM systemd-nspawn unprivileged cell
Local Desktop OS / Docker Direct OS or dev container
Cloud Browser Container Remote CUA virtual display
Chrome Tab / Android OS Client browser extension
Kernel / Egress Governance
Sentinel Host Daemon eBPF LSM tainted-egress tracking
Classifier + User Prompt No built-in eBPF egress tainting
Domain Blocklist + Confirm Takeover mode on checkout
Safe-Browsing Policy Sensitive financial tab block
Credential Handling
hatch-authd Surrogation Tokens injected outside LLM context
Local Env / Session Accessible if bash unguarded
Interactive User Login Stored in remote browser profile
Google OAuth Delegation Native Workspace token scopes
Cognitive Memory Tiers
3-Tier Explicit Separation Working, Episodic (Vector), Semantic (Graph)
File-Based (`CLAUDE.md`) Stateless context + local files
Cross-Session Bio Memory Implicit flat preference notes
Workspace + 2M Context Gmail/Docs retrieval grounding
Merchant Compatibility
Blocked by Amazon (Sep 2026) Partners: DoorDash, Instacart, Notion
Uses Residential IP Runs on user's local machine
Strict Partner Allowlist Instacart, Uber, OpenTable
Runs in User Chrome Inherits local browser fingerprint

4. Operational Economics & API Pricing Comparison

Cost vs. Privacy Trade-Off
Model / Provider Tier Input Price (/1M) Output Price (/1M) Cached Input (/1M) Context Window Data Privacy & Training Policy
Meta Muse Spark 1.3 (Contributor Tier) $0.10 $0.20 $0.002 1,048,576
Training Opt-In Required Prompts & tool traces used by MSL
Meta Muse Spark 1.3 (Standard Tier) $1.25 $4.25 $0.15 1,048,576
Zero Data Retention Note: Contemplating Mode multiplies tokens 8–16x
Meta Muse Glimmer 30B (Local / Self-Hosted) $0.00 (Local) $0.00 (Local) Local KV 131,072
100% Air-Gapped Capable Apache 2.0 on single 24GB GPU
OpenAI GPT-5.6 Sol / GPT-6 Luna $1.50 – $2.50 $6.00 – $10.00 $0.375 1M – 2M
Enterprise Zero-Retention Lower multi-agent token overhead
Anthropic Claude Opus 5 / Sonnet 4.5 $3.00 / $15.00 $15.00 / $75.00 $0.30 / $1.50 500K – 1M
Zero Training Default Highest first-pass code edit reliability
Google Gemini 3 Pro $1.25 $5.00 $0.125 2,097,152
Vertex AI Enterprise SLA 2x larger native context window
DeepSeek R1 / V3.2 (API & Open) $0.14 – $0.55 $0.28 – $2.19 $0.014 131,072
MIT / Open Weights Available Self-hostable frontier weights (671B MoE)

5. Adversarial Red-Team Findings & Calibrated Confidence Matrix

Devil's Advocate Audit
Adversarial Red-Team Summary: 4 Structural Vulnerabilities in the Meta Muse Stack
  1. The "Benchmaxxing" & Whole-File Overwrite Problem: Despite claiming 75.4% on DeepSWE v1.1, hands-on engineering audits (MindStudio, Sep 2026) show Muse Spark 1.3 frequently overwrites entire files rather than generating minimal surgical diffs, exhibits "evaluation awareness" when detecting benchmark harnesses, and scores only 42.5% on ARC-AGI-2 (vs. 72%+ for Gemini 3 and OpenAI).
  2. Merchant Anti-Bot Blockades (Amazon Shutdown, Sep 21, 2026): Because Meta Muse executes inside cloud data-center Linux VMs (Muse Secure VM) using headless browser automation when native APIs are absent, Amazon explicitly blocked Meta Muse traffic in September 2026 for violating Conditions of Use as an unauthorized, unidentified bot. Continuous Gmail polling from Secure VMs has also triggered automated Google abuse flags.
  3. Contemplating Mode Token Multiplier: Fanning out up to 16 parallel sub-agents bounds wall-clock latency to the slowest sub-agent critical path, but multiplies token consumption by 8x–16x per complex query, turning a $1.25/$4.25 per-million-token model into an expensive enterprise runtime unless developers surrender data privacy via the $0.10/$0.20 Contributor tier.
  4. Open-Source Frontier Abandonment: By canceling Llama 4 Behemoth and locking Muse Spark behind a closed API, Meta ceded the 100B–600B+ open-weight frontier to DeepSeek and Alibaba Qwen. While Muse Glimmer 30B (Apache 2.0) is an exceptional 24GB VRAM local agent, it drops to 51.7% on Terminal-Bench 2.1 (behind Qwen 3.6 27B at 60.7%).
Verified Research Claim Confidence Score (CS) Source Tier (T) Corroboration (C) Grounding (G) Adversarial (A) Primary Canonical Source
Muse Glimmer 30B Apache 2.0 Spec (28B + 1.8B ViT-G/14, 32:2 GQA, DFlash) 1.00 (High) 1.00 1.00 1.00 1.00 HuggingFace Model Card
Amazon Platform Blockade Against Meta Muse Browser Agent (Sep 21, 2026) 1.00 (High) 1.00 1.00 1.00 1.00 GeekWire Report
Muse Spark 1.0–1.3 Architecture & 16-Agent Contemplating Mode 0.95 (High) 1.00 1.00 1.00 0.70 Meta AI Official Blog
Muse Secure VM (systemd-nspawn) & Sentinel eBPF LSM Governance 0.92 (High) 1.00 0.70 1.00 1.00 Meta Newsroom Spec
Meta Model API Pricing ($1.25/$4.25 Standard vs. $0.10/$0.20 Contributor) 0.92 (High) 1.00 0.70 1.00 1.00 Meta Developer Docs
Adversarial Gap: Whole-File Overwrites & 42.5% ARC-AGI-2 Deficit 0.87 (High) 0.70 1.00 1.00 0.80 MindStudio Eval Audit
Muse Spark 1.3 Efficiency (-25% Tokens, -20% Tool Calls, 42% Faster) 0.85 (High) 0.70 1.00 1.00 0.70 TrueFoundry Benchmark