Peter Yang put a number on Meta's free assistant: $800 a year off his cable and phone bills. Not a benchmark score. Not a feeling of being more productive. Money that stopped leaving his account.

That number is the reason Muse is worth a CEO's twenty minutes. It is also the reason to be careful with almost everything else being said about it.

Figure 1. Assistant Benchmark v0.2
Muse leads on fewer dimensions
SCORE MuseInstinctGrok Bot 9.18.47.3 DIMENSIONS SCORED, OUT OF 15 Muse: 7 of 15 Instinct: 11 of 15
108 assistants, 15 dimensions, 15 tasks, 180 runs. The winning score was produced on a narrower evaluation than the runner-up’s, which is the part the ranking does not show.

What Muse is

Muse is Meta's personal AI assistant. It is free. By the middle of September 2026 it sat on top of the independent Assistant Benchmark at 9.1, ahead of Instinct at 8.4 and Grok Bot at 7.3.

Read that number with the asterisk attached. Muse was scored on 7 of the benchmark's 15 dimensions. Instinct was scored on 11. Muse leads on a narrower evaluation, which is not the same thing as leading. Alexandr Wang, Meta's chief AI officer, endorsed the benchmark publicly, which tells you it is credible and also that the product in first place had every reason to amplify it.

The $800 carries an asterisk too. It is one user, self-reported, unaudited. Nobody has seen the before and after bills.

Peter Yang
Peter
Yang

"the best personal agent I've tried to date. It actually save me $800+ a year on my cable and phone bills, which is just insane value for a free AI agent."

X  ·  September 18, 2026

What people actually use it for

Not chat. Deedy Das, an investor at Menlo Ventures, listed his favourite jobs across Muse and Instinct, and every one of them is paperwork with a deadline on it.

Deedy Das
Deedy
Das

"1. Submit FOIA requests to request data from the US government 2. Creating spend-limited Privacy cards to spend on subscriptions without having them recur 3. End to end filed an entire visa form for a country"

X  ·  September 20, 2026

Freedom of information requests. Cards that cannot renew a subscription. A visa application filed start to finish. None of these need creativity. They need somebody to sit down and do the form, and the only reason they stay undone is that nobody wants to be that somebody.

Garry Tan flagged a detail that matters more to a technical buyer than a consumer. Muse and Grok Bot ship Tailscale support out of the box, so they can reach into a private network. The cloud containers for Codex and Claude Code do not. A free consumer assistant with more network reach than the two serious coding harnesses is an odd fact, and a useful one to hold onto when you decide where you would and would not point it.

The design read is the one that lasts

Claire Vo ran her own hands-on and gave the feedback straight to Alexandr Wang. She did not grade the model. She graded the structure.

Claire Vo
Claire
Vo

"ux is :chefskiss: designers cooked. browser use/audio to text is not best in class. love primitives of goals, ideas, library. love lineage of tasks."

X  ·  September 13, 2026

Goals, ideas, library. Lineage of tasks. Those are the parts to care about, because lineage is what makes an agent auditable. You can see what it did and in what order and off what instruction. That is the line between an assistant you can put near a real process and one you can only use on yourself. Her negative is worth the same weight: browser use and speech to text are not best in class on the product currently sitting in first place.

The claim to leave out of your board meeting

Bill D'Alessandro uses all three and ranked Muse first. Then he offered an explanation for why.

Bill D'Alessandro
Bill
D'Alessandro

"Meta's new Muse AI assistant is extremely good. will probably replace both OpenClaw and Instinct for me."

"I think I've figured out why. It's OpenClaw under the hood."

X  ·  September 20, 2026

The second sentence is his inference. There is no Meta statement behind it, no technical disclosure, and no evidence presented. It has been repeated since as if it were a fact about the product.

Nat Eliason landed on the same ranking the same week and made no such claim.

Nat Eliason
Nat
Eliason

"Muse -> personal assistant. GrokBot -> work assistant. All you need now."

X  ·  September 20, 2026

So the quality judgment has two independent operators behind it. The explanation has one person's hunch. If the hunch were true it would be the most interesting fact in the whole category, a hyperscaler shipping its flagship consumer assistant on a community agent runtime. That is precisely why it needs a source before you repeat it.

The lead may be timing, not position

Signull
Signüll

"every major company is arriving at roughly the same offering: persistent memory, email/calendar/messages, browser + computer use, background tasks, proactive notifications, voice, app/tool execution, ambient context, & some notion of a personal agent sitting above everything."

X  ·  September 21, 2026

If that is right, Muse is ahead because it shipped, not because it is different. Every item on that list will be on every product inside a year. Which means the thing you are actually choosing between is not the feature grid. It is who you hand the connections to.

The selection rule worth stealing

Put the rankings down. The useful observation in all of this is Peter Yang's, and it is not really about Meta.

The consumer agent that gets kept is the one that recovered a specific, verifiable amount of money on a job the user was already annoyed about.

Every word in that sentence is load-bearing. Specific, so you can say the number out loud. Verifiable, so somebody else can check it. Money rather than time saved, because time saved is a story and money is a line. And already annoyed about, because that is the only place where a person will forgive a bad first run instead of quietly never opening the thing again.

That is a selection rule, not a review. It works on your company too.

Most first agents get picked the other way around. The most visible process. The one the loudest person in the room complains about. The one a vendor demo made look easy. Six months later nobody can tell you what it produced, and the honest answer is that nobody was ever going to be able to.

Run the test before you build. Find the job where the number already exists. A renewal nobody negotiated. An invoice line nobody queried. A vendor tier nobody right-sized. A refund nobody chased. Something with a figure before, a figure after, and a person who has been irritated by it for a year.

Your first agent will not be the impressive one. It will be the one still running next quarter.

Pick the boring one.

Which agent would your team actually keep?

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We pick your first agent the way the ones that stick get picked: an already-outsourced job, a checkable answer, and a number you can verify afterwards.

Book the Diagnostic →
Sources
1Peter Yang, 2026-09-18: Meta's Muse is "the best personal agent I've tried to date. It actually save me $800+ a year on my cable and phone bills, which is just insane value for a free AI agent." Self-reported and unaudited.
2Assistant Benchmark v0.2, mid-September 2026: 108 assistants, 15 dimensions, 15 tasks, 180 runs. Muse 9.1, Instinct 8.4, Grok Bot 7.3. Muse was scored on 7 of the 15 dimensions against Instinct's 11, so it leads on a narrower evaluation. Alexandr Wang called the benchmark "surprisingly comprehensive."
3Claire Vo, 2026-09-13, reviewing Muse for Alexandr Wang: praised the primitives of "goals, ideas, library" and "lineage of tasks", flagged browser use and audio-to-text as "not best in class."
4Nat Eliason, 2026-09-20: "Muse -> personal assistant. GrokBot -> work assistant. All you need now" and "Muse > Instinct." Bill D'Alessandro the same day: "will probably replace both OpenClaw and Instinct for me."
5Bill D'Alessandro, 2026-09-20: "It's OpenClaw under the hood." This is his inference. No Meta statement, no technical evidence, and Nat Eliason reached the same ranking without making the claim. Do not repeat it as fact.
6Deedy Das, 2026-09-20, on favourite Muse and Instinct jobs: submitting FOIA requests, creating spend-limited Privacy cards, and filing an entire visa form end to end.
7Garry Tan, 2026-09-17: "every AI harness needs to support Tailscale / Muse and Grok Bot do it out of box / Codex and Claude Code cloud containers are missing this right now."
8Signüll, 2026-09-21, on convergence: every major company is arriving at roughly the same offering, which would make any current lead a matter of release timing rather than position.
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.