Peter Yang put a number on Meta's free assistant: $800 a year off his cable and phone bills. Not a benchmark score. Not a feeling of being more productive. Money that stopped leaving his account.
That number is the reason Muse is worth a CEO's twenty minutes. It is also the reason to be careful with almost everything else being said about it.
What Muse is
Muse is Meta's personal AI assistant. It is free. By the middle of September 2026 it sat on top of the independent Assistant Benchmark at 9.1, ahead of Instinct at 8.4 and Grok Bot at 7.3.
Read that number with the asterisk attached. Muse was scored on 7 of the benchmark's 15 dimensions. Instinct was scored on 11. Muse leads on a narrower evaluation, which is not the same thing as leading. Alexandr Wang, Meta's chief AI officer, endorsed the benchmark publicly, which tells you it is credible and also that the product in first place had every reason to amplify it.
The $800 carries an asterisk too. It is one user, self-reported, unaudited. Nobody has seen the before and after bills.
Yang
"the best personal agent I've tried to date. It actually save me $800+ a year on my cable and phone bills, which is just insane value for a free AI agent."
What people actually use it for
Not chat. Deedy Das, an investor at Menlo Ventures, listed his favourite jobs across Muse and Instinct, and every one of them is paperwork with a deadline on it.
Das
"1. Submit FOIA requests to request data from the US government 2. Creating spend-limited Privacy cards to spend on subscriptions without having them recur 3. End to end filed an entire visa form for a country"
Freedom of information requests. Cards that cannot renew a subscription. A visa application filed start to finish. None of these need creativity. They need somebody to sit down and do the form, and the only reason they stay undone is that nobody wants to be that somebody.
Garry Tan flagged a detail that matters more to a technical buyer than a consumer. Muse and Grok Bot ship Tailscale support out of the box, so they can reach into a private network. The cloud containers for Codex and Claude Code do not. A free consumer assistant with more network reach than the two serious coding harnesses is an odd fact, and a useful one to hold onto when you decide where you would and would not point it.
The design read is the one that lasts
Claire Vo ran her own hands-on and gave the feedback straight to Alexandr Wang. She did not grade the model. She graded the structure.
Vo
"ux is :chefskiss: designers cooked. browser use/audio to text is not best in class. love primitives of goals, ideas, library. love lineage of tasks."
Goals, ideas, library. Lineage of tasks. Those are the parts to care about, because lineage is what makes an agent auditable. You can see what it did and in what order and off what instruction. That is the line between an assistant you can put near a real process and one you can only use on yourself. Her negative is worth the same weight: browser use and speech to text are not best in class on the product currently sitting in first place.
The claim to leave out of your board meeting
Bill D'Alessandro uses all three and ranked Muse first. Then he offered an explanation for why.
D'Alessandro
"Meta's new Muse AI assistant is extremely good. will probably replace both OpenClaw and Instinct for me."
"I think I've figured out why. It's OpenClaw under the hood."
The second sentence is his inference. There is no Meta statement behind it, no technical disclosure, and no evidence presented. It has been repeated since as if it were a fact about the product.
Nat Eliason landed on the same ranking the same week and made no such claim.
Eliason
"Muse -> personal assistant. GrokBot -> work assistant. All you need now."
So the quality judgment has two independent operators behind it. The explanation has one person's hunch. If the hunch were true it would be the most interesting fact in the whole category, a hyperscaler shipping its flagship consumer assistant on a community agent runtime. That is precisely why it needs a source before you repeat it.
The lead may be timing, not position
"every major company is arriving at roughly the same offering: persistent memory, email/calendar/messages, browser + computer use, background tasks, proactive notifications, voice, app/tool execution, ambient context, & some notion of a personal agent sitting above everything."
If that is right, Muse is ahead because it shipped, not because it is different. Every item on that list will be on every product inside a year. Which means the thing you are actually choosing between is not the feature grid. It is who you hand the connections to.
The selection rule worth stealing
Put the rankings down. The useful observation in all of this is Peter Yang's, and it is not really about Meta.
The consumer agent that gets kept is the one that recovered a specific, verifiable amount of money on a job the user was already annoyed about.
Every word in that sentence is load-bearing. Specific, so you can say the number out loud. Verifiable, so somebody else can check it. Money rather than time saved, because time saved is a story and money is a line. And already annoyed about, because that is the only place where a person will forgive a bad first run instead of quietly never opening the thing again.
That is a selection rule, not a review. It works on your company too.
Most first agents get picked the other way around. The most visible process. The one the loudest person in the room complains about. The one a vendor demo made look easy. Six months later nobody can tell you what it produced, and the honest answer is that nobody was ever going to be able to.
Run the test before you build. Find the job where the number already exists. A renewal nobody negotiated. An invoice line nobody queried. A vendor tier nobody right-sized. A refund nobody chased. Something with a figure before, a figure after, and a person who has been irritated by it for a year.
Your first agent will not be the impressive one. It will be the one still running next quarter.
Pick the boring one.
Which agent would your team actually keep?
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We pick your first agent the way the ones that stick get picked: an already-outsourced job, a checkable answer, and a number you can verify afterwards.
Book the Diagnostic →