# Meta Muse Glimmer: Local Agent Model Guide

> Choose Muse Glimmer weights, quantization, runtimes, tool calling, and hardware for a practical local multimodal agent.

- **Published**: 2026-08-11
- **Category**: AI Infrastructure
- **URL**: https://agentpedia.codes/blog/meta-muse-glimmer-local-agent-model-guide

---

Meta Muse Glimmer is best understood as a **local-agent stack**, not just a 30B checkpoint. Meta combines a multimodal dense model, native ATEM tool calling, long-context support, quantized artifacts, a speculative-decoding companion, and integrations for llama.cpp, ExecuTorch, Transformers, vLLM, MLX, Ollama, and other runtimes.

> **Note callout**

**Practical verdict:** Muse Glimmer is a serious option for developers with roughly 24-32 GB of usable memory and a willingness to tune the runtime. "Local" here means a high-end Mac, GPU, or unified-memory system--not an ordinary laptop. Persistent memory still belongs to your agent harness, database, and restart logic.

## What Muse Glimmer is

Meta's model page describes Muse Glimmer as an open model for always-on local agents. The model card reports approximately **29.6 billion parameters**, a dense decoder-only transformer, and an approximately 1.8B ViT-G/14 perception encoder. The listed license is Apache 2.0 for the model artifacts.

If you are comparing local deployment sizes, AgentPedia's [Liquid LFM2.5-2.6B on-device guide](/blog/liquid-lfm2-5-2-6b-on-device-agent-guide) covers a much smaller class of model, while the [Shieldstral multimodal moderation guide](/blog/mistral-shieldstral-3b-multimodal-moderation-guide) shows a different image-and-text workload.

| Property | What the primary sources report | Practical meaning |
| --- | --- | --- |
| Parameters | About 29.6B, marketed as 30B | Large enough for serious local agent behavior, too large for casual laptop deployment |
| Input | Text and images | Multimodal prompts are supported when the runtime includes the perception path |
| Output | Text | Audio is not a native input/output mode |
| Context | 128K default in Meta docs; model card reports 131,072+ | Pin the runtime's actual context behavior instead of assuming every path is identical |
| Tool format | ATEM | The server or client must translate the model-native format correctly |
| License | Apache 2.0 for listed release artifacts | Open weights do not imply open training data or open training code |

The model card says video is not explicitly optimized as a first-class modality. A video input is handled as individual frames, which is not the same thing as native temporal video understanding. That distinction matters for anyone designing a camera or screen agent.

## Tool calling and persistent state

Muse Glimmer's native prompting guide uses an **ATEM** format with an assistant tool-call block and an invocation payload. The important constraint is easy to miss: the model-native format supports **one tool call per turn**, not parallel tool calls. The agent loop should execute the tool, return its result, and then ask the model to choose the next action.

A serving layer can expose an OpenAI-compatible API while translating ATEM internally. For example, the official vLLM recipe uses Muse-specific tool and reasoning parsers:

```bash
vllm serve meta-models/Muse-Glimmer-30B \
  --enable-auto-tool-choice \
  --tool-call-parser muse_glimmer \
  --reasoning-parser muse_glimmer
```

Pin the vLLM version and confirm the parser names against the current recipe before using this in production. A compatible HTTP surface does not mean every client understands the same reasoning controls.

The other common misunderstanding is persistent memory. Meta's product language discusses always-on agents and persistent state, but the implementation evidence is more precise: persistence across restarts comes from the surrounding harness. You still need a memory store, a session identifier, serialization, retrieval policy, and recovery behavior. The weights do not remember a previous process by themselves.

## Hardware and quantization

The full BF16 artifact is a server-class deployment. Unsloth reports roughly 58 GB, while Meta describes the full model as requiring more than 55 GB. Meta's recommended quantized artifacts are more practical:

| Artifact | Target class | Meta-reported degradation |
| --- | --- | ---: |
| K-Quant-Dynamic | 32 GB | 0.2% average across 15 benchmarks |
| K-Quant-17GB | 24 GB | 1.0% average across 15 benchmarks |
| GGUF text build | About 16.8-19.7 GB depending on variant | Measure on your workload |
| Vision projector | About 1.4 GB in the GGUF distribution | Required for image input |
| DFlash drafter | About 1.6 GB in the GGUF distribution | Optional speculative decoding companion |

Those numbers are artifact sizes or vendor target classes, not the total memory budget for a useful agent. Context length, image processing, KV cache, runtime allocations, tool results, and concurrency all add overhead. A 24 GB system may load the 17 GB path and still fail when the context or vision path grows.

Meta reports the following local speed examples for K-Quant-17GB plus DFlash at batch size one and greedy decoding:

| Hardware/runtime | Without DFlash | With DFlash | Reported change |
| --- | ---: | ---: | ---: |
| RTX 5090 / llama.cpp | 74.9 tok/s | 233.4 tok/s | 3.1x |
| M4 Max / ExecuTorch | 23.7 tok/s | 37.8 tok/s | 1.5x |
| M5 Max / ExecuTorch | 26.6 tok/s | 50.2 tok/s | 1.8x |

These are Meta-reported measurements under narrow conditions. They do not predict end-to-end agent latency, image latency, tool-call pauses, or reasoning quality. DFlash proposes tokens and Muse Glimmer verifies them; the output distribution is intended to remain the target model's distribution, but speed depends on acceptance rate, hardware, draft length, and workload.

## Deployment paths

Choose the runtime according to the job rather than downloading every artifact.

| Runtime | Best fit | Important constraint |
| --- | --- | --- |
| Transformers | Direct experiments and custom Python | You manage processor, chat template, tools, and memory yourself |
| llama.cpp | Local GGUF server or CLI | Image input needs `llama-mtmd-cli`, the multimodal projector, and `--jinja`; the model card requires a recent build |
| ExecuTorch | Documented Apple Silicon or NVIDIA PTE paths | No CPU variant, no continuous batching, and execution is serialized |
| vLLM | OpenAI-compatible serving | Use Muse-specific tool/reasoning parsers; quantization is hardware-dependent |
| MLX / Ollama / LM Studio | Convenient local ecosystem paths | Verify current integration and artifact support before promising multimodal tools |

For llama.cpp, the GGUF card specifies a minimum build around `b10353` or newer at the time of research. A text-only GGUF is not enough for image input. Download the matching multimodal projector and use the multimodal executable. This is a common reason a successful text launch turns into a failed vision launch.

ExecuTorch provides pre-exported artifacts for Apple Metal and NVIDIA `sm80+ptx`, with text/text-image and solo/DFlash variants. The repository is large, so selective downloads are important. The documented runner also has no continuous batching, no cross-session prefix sharing, and no persisted session state. Those are runtime constraints, not proof that the model itself lacks the concepts.

For vLLM, BF16, FP8, and NVFP4 recipes are listed, but NVFP4 is a Blackwell-specific path rather than a generic consumer GPU option. Do not compare it directly with a K-Quant GGUF and call the result a model-level speed or quality difference.

The broader deployment tradeoff is also useful to compare with the [GPT-5.6 family guide](/blog/gpt-5-6-sol-terra-luna-explained), which covers hosted model tiers rather than local weights.

No official Meta, Hugging Face, or vLLM source checked for this guide documents a Muse Glimmer Tinker integration. Treat Tinker support as **not documented by the sources checked**, not as proof that an integration is impossible.

## How to read the benchmarks

Meta's published table includes strong agentic and multimodal numbers, including SWE-Bench Verified 76.0, SWE-Bench Pro 51.2, TerminalBench 2.1 51.7, MCP Atlas 75.5, OSWorld-Verified 65.9, and CharXiv Reasoning 78.8. These should be labeled **Meta-reported**.

They are not bare next-token scores. Agent benchmarks include tools, scaffolds, sandboxes, judges, sampling settings, and sometimes modified benchmark implementations. Meta's methodology reports, for example, four runs for MCP Atlas and DeepSearch QA, 89 tasks for TerminalBench 2.1, 361 tasks for OSWorld-Verified, and multiple attempts for multimodal evaluations. Some comparisons use Artificial Analysis results or internally reproduced competitors.

A particularly important caveat is OmniDocBench: Meta's methodology describes an internal implementation with a modified two-component score and simpler matching than the official version. A number from that setup should not be placed beside an official leaderboard result without labeling the difference.

## A reproducibility checklist

Before comparing Muse Glimmer with another local model, record:

- exact model revision and artifact;
- quantization format and runtime version;
- context length and KV-cache policy;
- sampler settings, including `temperature`, `top_p`, and `top_k`;
- reasoning strength and chat template;
- tool definitions and agent scaffold;
- hardware, driver, and concurrency;
- judge model and number of attempts;
- whether the result is a bare-model or agent-system score.

For a local deployment test, start with one text-only task, one image task, and one sequential tool loop. Measure load time, steady-state tokens per second, first-token latency, tool-call latency, memory headroom, and failure behavior. Then repeat with the intended context size. A model that looks comfortable at 2K tokens may become unusable at 64K.

> **Warning callout**

Do not describe a quantized checkpoint as "lossless" because a vendor reports a small average degradation. Quantization can change edge-case behavior, and agent loops amplify small errors through tool selection and retries.

## FAQ

### Can Muse Glimmer run on a normal laptop?

Usually not comfortably. The 17 GB and 32 GB targets assume high-end hardware, and context, vision, and runtime overhead require additional memory.

### Does the model have persistent memory?

No. Persistent state is supplied by the agent harness, memory store, and restart logic around the model.

### Does it support parallel tool calls?

The native prompting guide says no: use one tool call per turn and return its result before the next decision.

### Which runtime should I choose?

Use llama.cpp for a local GGUF path, vLLM for an OpenAI-compatible server, Transformers for direct experimentation, and ExecuTorch for its documented Apple Silicon or NVIDIA artifacts.

## Sources and links

### Primary

- [Meta Muse Glimmer model page](https://developer.meta.com/ai/models/muse-glimmer/)
- [Meta research launch post](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)
- [Meta model overview](https://dev.meta.ai/docs/muse-glimmer.md)
- [Meta prompting and ATEM tools](https://dev.meta.ai/docs/muse-glimmer/prompting.md)
- [Meta quantization guide](https://dev.meta.ai/docs/muse-glimmer/quantization.md)
- [Meta speculative decoding guide](https://dev.meta.ai/docs/muse-glimmer/spec-decode.md)
- [Meta evaluation methodology](https://research.meta.ai/static/muse-glimmer-methodology)

### Model artifacts

- [Muse Glimmer BF16 model](https://huggingface.co/meta-models/Muse-Glimmer-30B)
- [Muse Glimmer GGUF](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF)
- [Muse Glimmer ExecuTorch artifacts](https://huggingface.co/meta-models/Muse-Glimmer-30B-ExecuTorch-PTE)
- [Muse Glimmer collection](https://huggingface.co/collections/meta-models/muse-glimmer)

### Runtime references

- [vLLM Muse Glimmer recipe](https://recipes.vllm.ai/meta-models/Muse-Glimmer-30B)
- [Unsloth local guide](https://unsloth.ai/docs/models/muse-glimmer)
- [llama.cpp](https://github.com/ggml-org/llama.cpp)
- [Apache 2.0 license](https://www.apache.org/licenses/LICENSE-2.0)


---

- [All articles](https://agentpedia.codes/blog)