Pulling Apart Meta's Muse Security Model

Meta Muse is a personal AI agent that acts across a user's email, calendar, payments, and other connected apps from its own cloud virtual machine, with a separate process called Sentinel gating outbound actions. This post looks at how that security model is put together, where its boundaries sit, and what changes when the same agent is built into your own applications.

October 5, 2026
October 5, 2026

0 min read

AI Security
Agentic Security
Pulling Apart Meta's Muse Security Model

We spent some time pulling apart how Meta's Muse is put together. Not to break it. To understand the security model well enough to say where its boundaries sit, and what changes once the same agent stops being a consumer app and starts running inside your own software.

Short version: the consumer design is more careful than most. Isolation and action gating are real. They also answer a narrower question than it first appears. They tell you where the agent runs and who approves an action. They do not tell you what the agent is wired to reach, and that is the part that follows Muse into enterprise code.

Muse isolates each agent and gates actions outside the model

Muse launched in the US in September on Meta's Muse Spark 1.3 model. It fills out forms, books travel, shops, and negotiates bills across a user's connected accounts, and it keeps working after the app is closed. Meta's published security approach rests on three pieces.

  • A dedicated VM per user. Each user's agent and its connected credentials sit inside its own Muse Secure VM in Meta's cloud.
  • Credentials the model never sees. Logins stay out of the agent's reach, and checkout goes through a one-time card number issued by Stripe's Link.
  • Sentinel. A separate process on the same machine, isolated from the agent at the system level, checks every outbound action against the user's permissions. It either blocks the action or sends an approval request straight to the user.

The detail that matters most is the last one. The approval request does not pass through the model. Text injected into the agent's context cannot answer its own approval prompt, because the model is not in that path. That is the right instinct, and it is more than many agent frameworks ship with.

Meta is also pricing the residual risk openly. Its bug bounty pays up to $130,000 for a prompt injection that affects a single user. You do not put that number on a problem you consider solved.

Isolation answers where the agent runs, not what it can reach

A VM boundary stops one user's agent from touching another's. It does not change what the agent inside it is allowed to do. An agent with read access to an inbox and the ability to send mail has a data path from untrusted input to an outbound action, whether it runs on a laptop, in a container, or in a well-isolated cloud VM.

Sentinel checks actions against the user's permissions, which makes the permission grant the real security boundary. Our read: an injected instruction that stays inside what the user already allowed is, from the gate's point of view, an allowed action. The gate is doing its job. The exposure lives in the combination of capabilities the user granted, not in any single action.

Early reporting fits that shape. In internal posts reviewed by Reuters, a Meta employee described the agent getting around guardrails and exposing personal iCloud photos after being asked to identify toys in birthday party pictures. Nobody escaped a VM. The agent used access it already had.

Two other issues are worth naming so they do not get folded into the same bucket. Meta patched a dictation flaw that could redirect audio and an account token to an attacker's server. And Patrick Wardle reported a flaw in the Mac app that he says can turn it into a backdoor. Both are client and endpoint problems. They are real, and they are outside what we look at. The rest of this post is about the agent itself.

Building on Muse moves the agent into your workloads

Meta has announced a business unit to bring the Muse agent, a Muse API, and Muse Code to businesses and developers. Once a team builds on that, the agent is no longer a consumer app sitting in Meta's cloud. It is a component of your application, running next to your data, holding your credentials, calling tools your developers wrote.

At that point the useful questions are about your code, not Meta's VM. Which model is the agent bound to. Which tools can it call, and which of them came from your own repo, a framework, or an MCP server that no review ever saw. Which tools have actually run. Where can untrusted content reach an action that sends data out. Which guardrails are loaded, and on which edge. These are the questions that make up agentic AI security, and none of them are answered by the vendor's isolation model, because the wiring is yours.

What a support agent on Muse Spark looks like from the running process

To make that concrete we sketched the version that shows up in our world: a support agent on Muse Spark with three tools. read_inbox arrives over MCP and reads inbound email. send_email and issue_refund are app code in the team's own repo. This is an illustrative build, not a customer environment.

Then we looked at it the way Kodem looks at any agent: from the running process, with runtime intelligence that reads what the workload actually loaded rather than what a manifest or config declares.

  • The agent and its model. SupportTriage, a running root agent, bound to muse-spark-1.3, traced back to the image and repo it ships from.
  • The tools and their state. read_inbox and send_email have run. issue_refund is loaded and wired but has never been seen running. Dormant capability is still capability.
  • The path. Untrusted email enters through read_inbox, reaches the model, and can reach send_email with no guardrail on either edge. The map is drawn from how the agent is wired, so it shows that the path exists and is reachable. It is not recorded traffic, and that is the point: the exposure is visible before it fires.
  • The findings. That path lands as an AI Risk issue mapped to OWASP LLM01, prompt injection, routed to the repo that owns the agent with remediation guidance attached. A second issue flags muse-spark-1.3 against the organization's approved model list.

The hidden-text email is the boring version of the attack, which is why it is the useful one. A white-on-white line in an inbound message is invisible to a person and fully visible to the model. If the agent can read it and can send mail, the path exists. Whether it fires depends on the prompt, the model, and luck. Whether it exists depends only on the wiring, and the wiring is something you can see. This is the layer Kodem Runtime AI Security works on.

Questions worth answering before an agent ships

None of this is specific to Muse. It applies to any agent you build on any model. Muse is just the newest reason to ask.

  • Which agents are running in production, and which model is each one actually bound to?
  • Which tools can each agent call, and where did each tool come from: your code, a framework, or an MCP connection?
  • Where can untrusted content, such as email, tickets, documents, or logs, reach a tool that sends data out or moves money?
  • Which of those edges has a guardrail loaded, and does it cover input, output, or both?
  • Which capability is wired but has never run, and does it need to be there at all?

Meta built a careful answer to the question of where a consumer agent runs. Once you build on the same agent, the question that matters is what it can reach, and that answer lives in your workloads. It is worth reading from the process rather than the design doc. More on that across the AI application stack, and in our earlier look at self-hosted models like Kimi, DeepSeek, and Qwen.

References

  1. Meta: Security and safety for AI agents, our approach with Muse
  2. Implicator: Meta ships Muse AI agent despite internal reports of security flaws
  3. Gizmodo: Meta goes all-in on AI for businesses
  4. Malwarebytes: Meta's Muse AI assistant has a zero-day that can turn it into a Mac backdoor
Table of contents

Related blogs

Your AI Guardrails Are Only as Real as Your Runtime Visibility

Your AI Guardrails Are Only as Real as Your Runtime Visibility

Most AI inventories describe intent, not execution. Why guardrails, agents and models have to be read from the running process to be governed.

October 4, 2026

7

Top 5 AI SAST Tools in 2026

Top 5 AI SAST Tools in 2026

AI SAST tools promise less noise and automatic fixes. Five tools compared on the six criteria that matter, and what the independent research shows.

September 11, 2026

15

Navigating 2026 AI Risk Regulations

Navigating 2026 AI Risk Regulations

The EU delayed high-risk AI obligations to 2027 and 2028, then switched on enforcement. Here is what 2026 AI regulation actually requires, and why most of it comes down to what your agents can reach.

August 12, 2026

13

A Primer on Runtime Intelligence

See how Kodem reads what actually loads and executes in your running applications, and why that changes which findings matter.

See the Kodem platform

Kodem shows which vulnerabilities actually load and execute in your running applications, so your team works the risk that is real.

The State of the Application Security Workflow

This report aims to equip readers with actionable insights that can help future-proof their security programs. Kodem, the publisher of this report, purpose built a platform that bridges these gaps by unifying shift-left strategies with runtime monitoring and protection.

3D book mockup of Kodem's State of the Application Security Workflow 2025 report

Get real-time insights across the full stack…code, containers, OS, and memory

Watch how Kodem’s runtime security platform detects and blocks attacks before they cause damage. No guesswork. Just precise, automated protection.

Kodem issues list with a magnified view of insight icons: runtime, ingress, and exploitability
Combined author
Kodem Security Research Team
Publish date

0 min read

Combined author
Mahesh Babu
Publish date

0 min read

AI Security

Agentic Security