AI Infrastructure

Meta Muse Glimmer: Local Agent Model Guide

Choose Muse Glimmer weights, quantization, runtimes, tool calling, and hardware for a practical local multimodal agent.

Abstract multimodal local model connected to a GPU and laptop through a quantized inference pipeline
AgentPedia conceptual illustration of local multimodal model deployment. It is not an official Meta runtime diagram. View image source.

Meta Muse Glimmer is best understood as a local-agent stack, not just a 30B checkpoint. Meta combines a multimodal dense model, native ATEM tool calling, long-context support, quantized artifacts, a speculative-decoding companion, and integrations for llama.cpp, ExecuTorch, Transformers, vLLM, MLX, Ollama, and other runtimes.

What Muse Glimmer is

Meta’s model page describes Muse Glimmer as an open model for always-on local agents. The model card reports approximately 29.6 billion parameters, a dense decoder-only transformer, and an approximately 1.8B ViT-G/14 perception encoder. The listed license is Apache 2.0 for the model artifacts.

If you are comparing local deployment sizes, AgentPedia’s Liquid LFM2.5–2.6B on-device guide covers a much smaller class of model, while the Shieldstral multimodal moderation guide shows a different image-and-text workload.

PropertyWhat the primary sources reportPractical meaning
ParametersAbout 29.6B, marketed as 30BLarge enough for serious local agent behavior, too large for casual laptop deployment
InputText and imagesMultimodal prompts are supported when the runtime includes the perception path
OutputTextAudio is not a native input/output mode
Context128K default in Meta docs; model card reports 131,072+Pin the runtime’s actual context behavior instead of assuming every path is identical
Tool formatATEMThe server or client must translate the model-native format correctly
LicenseApache 2.0 for listed release artifactsOpen weights do not imply open training data or open training code

The model card says video is not explicitly optimized as a first-class modality. A video input is handled as individual frames, which is not the same thing as native temporal video understanding. That distinction matters for anyone designing a camera or screen agent.

Tool calling and persistent state

Muse Glimmer’s native prompting guide uses an ATEM format with an assistant tool-call block and an invocation payload. The important constraint is easy to miss: the model-native format supports one tool call per turn, not parallel tool calls. The agent loop should execute the tool, return its result, and then ask the model to choose the next action.

A serving layer can expose an OpenAI-compatible API while translating ATEM internally. For example, the official vLLM recipe uses Muse-specific tool and reasoning parsers:

vllm serve meta-models/Muse-Glimmer-30B \
  --enable-auto-tool-choice \
  --tool-call-parser muse_glimmer \
  --reasoning-parser muse_glimmer

Pin the vLLM version and confirm the parser names against the current recipe before using this in production. A compatible HTTP surface does not mean every client understands the same reasoning controls.

The other common misunderstanding is persistent memory. Meta’s product language discusses always-on agents and persistent state, but the implementation evidence is more precise: persistence across restarts comes from the surrounding harness. You still need a memory store, a session identifier, serialization, retrieval policy, and recovery behavior. The weights do not remember a previous process by themselves.

Hardware and quantization

The full BF16 artifact is a server-class deployment. Unsloth reports roughly 58 GB, while Meta describes the full model as requiring more than 55 GB. Meta’s recommended quantized artifacts are more practical:

ArtifactTarget classMeta-reported degradation
K-Quant-Dynamic32 GB0.2% average across 15 benchmarks
K-Quant-17GB24 GB1.0% average across 15 benchmarks
GGUF text buildAbout 16.8–19.7 GB depending on variantMeasure on your workload
Vision projectorAbout 1.4 GB in the GGUF distributionRequired for image input
DFlash drafterAbout 1.6 GB in the GGUF distributionOptional speculative decoding companion

Those numbers are artifact sizes or vendor target classes, not the total memory budget for a useful agent. Context length, image processing, KV cache, runtime allocations, tool results, and concurrency all add overhead. A 24 GB system may load the 17 GB path and still fail when the context or vision path grows.

Meta reports the following local speed examples for K-Quant-17GB plus DFlash at batch size one and greedy decoding:

Hardware/runtimeWithout DFlashWith DFlashReported change
RTX 5090 / llama.cpp74.9 tok/s233.4 tok/s3.1×
M4 Max / ExecuTorch23.7 tok/s37.8 tok/s1.5×
M5 Max / ExecuTorch26.6 tok/s50.2 tok/s1.8×

These are Meta-reported measurements under narrow conditions. They do not predict end-to-end agent latency, image latency, tool-call pauses, or reasoning quality. DFlash proposes tokens and Muse Glimmer verifies them; the output distribution is intended to remain the target model’s distribution, but speed depends on acceptance rate, hardware, draft length, and workload.

Deployment paths

Choose the runtime according to the job rather than downloading every artifact.

RuntimeBest fitImportant constraint
TransformersDirect experiments and custom PythonYou manage processor, chat template, tools, and memory yourself
llama.cppLocal GGUF server or CLIImage input needs llama-mtmd-cli, the multimodal projector, and --jinja; the model card requires a recent build
ExecuTorchDocumented Apple Silicon or NVIDIA PTE pathsNo CPU variant, no continuous batching, and execution is serialized
vLLMOpenAI-compatible servingUse Muse-specific tool/reasoning parsers; quantization is hardware-dependent
MLX / Ollama / LM StudioConvenient local ecosystem pathsVerify current integration and artifact support before promising multimodal tools

For llama.cpp, the GGUF card specifies a minimum build around b10353 or newer at the time of research. A text-only GGUF is not enough for image input. Download the matching multimodal projector and use the multimodal executable. This is a common reason a successful text launch turns into a failed vision launch.

ExecuTorch provides pre-exported artifacts for Apple Metal and NVIDIA sm80+ptx, with text/text-image and solo/DFlash variants. The repository is large, so selective downloads are important. The documented runner also has no continuous batching, no cross-session prefix sharing, and no persisted session state. Those are runtime constraints, not proof that the model itself lacks the concepts.

For vLLM, BF16, FP8, and NVFP4 recipes are listed, but NVFP4 is a Blackwell-specific path rather than a generic consumer GPU option. Do not compare it directly with a K-Quant GGUF and call the result a model-level speed or quality difference.

The broader deployment tradeoff is also useful to compare with the GPT-5.6 family guide, which covers hosted model tiers rather than local weights.

No official Meta, Hugging Face, or vLLM source checked for this guide documents a Muse Glimmer Tinker integration. Treat Tinker support as not documented by the sources checked, not as proof that an integration is impossible.

How to read the benchmarks

Meta’s published table includes strong agentic and multimodal numbers, including SWE-Bench Verified 76.0, SWE-Bench Pro 51.2, TerminalBench 2.1 51.7, MCP Atlas 75.5, OSWorld-Verified 65.9, and CharXiv Reasoning 78.8. These should be labeled Meta-reported.

They are not bare next-token scores. Agent benchmarks include tools, scaffolds, sandboxes, judges, sampling settings, and sometimes modified benchmark implementations. Meta’s methodology reports, for example, four runs for MCP Atlas and DeepSearch QA, 89 tasks for TerminalBench 2.1, 361 tasks for OSWorld-Verified, and multiple attempts for multimodal evaluations. Some comparisons use Artificial Analysis results or internally reproduced competitors.

A particularly important caveat is OmniDocBench: Meta’s methodology describes an internal implementation with a modified two-component score and simpler matching than the official version. A number from that setup should not be placed beside an official leaderboard result without labeling the difference.

A reproducibility checklist

Before comparing Muse Glimmer with another local model, record:

  • exact model revision and artifact;
  • quantization format and runtime version;
  • context length and KV-cache policy;
  • sampler settings, including temperature, top_p, and top_k;
  • reasoning strength and chat template;
  • tool definitions and agent scaffold;
  • hardware, driver, and concurrency;
  • judge model and number of attempts;
  • whether the result is a bare-model or agent-system score.

For a local deployment test, start with one text-only task, one image task, and one sequential tool loop. Measure load time, steady-state tokens per second, first-token latency, tool-call latency, memory headroom, and failure behavior. Then repeat with the intended context size. A model that looks comfortable at 2K tokens may become unusable at 64K.

FAQ

Can Muse Glimmer run on a normal laptop?

Usually not comfortably. The 17 GB and 32 GB targets assume high-end hardware, and context, vision, and runtime overhead require additional memory.

Does the model have persistent memory?

No. Persistent state is supplied by the agent harness, memory store, and restart logic around the model.

Does it support parallel tool calls?

The native prompting guide says no: use one tool call per turn and return its result before the next decision.

Which runtime should I choose?

Use llama.cpp for a local GGUF path, vLLM for an OpenAI-compatible server, Transformers for direct experimentation, and ExecuTorch for its documented Apple Silicon or NVIDIA artifacts.

Sources and links

Primary

Model artifacts

Runtime references