Two operators published their Grok Bot results in the same week. One found the bot had covered its own subscription. The other published a firing notice.
Fallen
"My Grok Bot just paid its own salary. Which is completely insane. I gave it one job... Win back churned customers. It: Found everyone who left in the last 6 months. Emailed them. Won several back. Collected their feedback."
Note what he did not say. There is no revenue figure in that post. "Won several back" is the whole claim, and he left it there. Overnight the same bot went back through the replies and wrote up why those customers had left, which is arguably the more valuable half.
The firing came from a plumbing company owner with zero engineers on staff, who had gone from download to automated dispatch and office chores in 24 hours. He posted that win. Then he posted the other one: "Dana Dispatcher, the first bot I ever fired (my fault). Slow, very slow, and would make mistakes."
Read the parenthetical again. My fault. That is the correct diagnosis, and it is the whole article.
What the wins have in common
One job. Written as a sentence. With an outcome you could check on Friday.
Liam's job description was five words: win back churned customers. Not "help with retention." Not "own the customer lifecycle." A verb, an object, and a population the bot could enumerate itself. Every measured outcome in the current use-case catalog has that shape. A sales play that used to take 90 minutes is now roughly 95% automated. Most of a custom CRM rebuilt by a small bot crew in about a day and a half. 200-plus recruiting applications triaged in one pass. A Salesforce end-of-quarter question answered from a phone in about ten seconds. A community ops fleet saving 20-plus hours a week.
Peter Yang published the cleanest single number. His concierge bot found a Tokyo round trip $2,700 cheaper than his preferred itinerary, by reasoning across his whole vacation document rather than matching one price alert. Scoped to an outcome, not to a tool.
Baker
"I think @bot is another 'Claude Code' moment for AI. I would estimate my personal AI usage is up something like 100x. And for everyone who reached out about how to build a 'podcast summarizer' it took me about 15 seconds in Grok Bot and is better than what I had before."
Fifteen seconds, because the job was already a sentence before he opened the app.
What the failures have in common
"Dana Dispatcher" is a job title. It is not a job description. Title-shaped scope is how you get slow and wrong: the bot infers the boundary of its own work on every run, so it reads more, does more, and produces something nobody specified well enough to reject quickly.
It is also the expensive failure mode, because these agents run continuously. A vague job does not sit idle waiting for clarity. It burns tokens finding things to do.
You cannot price that risk from published data either. No reproducible agentic task-success benchmark shipped with the launch, and escalation reliability is unquantified: no independent measure exists of how often the bot correctly stops and asks a human. The one efficiency figure in circulation is vendor-sourced and self-reported, that internal users report feeling 2 to 3 times more efficient. Feeling. Treat it accordingly.
The actual cost
There is no free tier. Access comes through Cursor Ultra at $200 a month, SuperGrok Heavy at $300, or Cursor Teams Premium at $120 per seat. Nate B Jones describes the on-ramp as going "from free usage for a while, maybe less than a day, all the way to $200 a month." He built 12 bots and says 2 of them alone justify the subscription.
The subscription is the floor, not the bill. Weekly usage is included; anything past it bills at token cost on top. An early-access user summarized one month on Hacker News: "I've used less tokens in the last 5 years prior to this month than I have this month. Always on perpetual agents use a LOT of tokens."
Underneath it sits a competitive floor that appeared within a week. Nous priced Hermes Cloud at "3 cents per day when idle and 29 cents for nonstop usage," posted directly under Musk's own promotion of Grok Bot. The bull case is the line making the rounds: Grok Bot costs $300 a month, a junior employee costs that in two days. The bear case, from @bitslix on August 18: "If AI companies keep building this future exclusively for people with enormous disposable incomes, they shouldn't be surprised when 'mass adoption' never actually becomes mass adoption."
The verdict
Buy now if two things are true. You already pay for SuperGrok Heavy or Cursor Ultra, which makes this effectively free on top of a bill you signed anyway. And you can write one job as a single sentence with a countable result, then name the person who checks its output daily for two weeks.
Wait if you are not on those tiers. At $200 to $300 with no way to try it in isolation, you are buying a capability you cannot scope before you pay. Miles Deutscher, after a week of use, landed exactly there.
Deutscher
"Pound for pound, it IS the most useful and powerful agentic software on the market right now. But here's the honest answer for most of you: it's probably not worth it yet."
His read on what you are paying for is the sharpest published: "The value isn't that Grok Bot does something fundamentally new. It's that it removes the setup headache and maintenance headache from tools like Hermes/OpenClaw." That is a real product. It is also a convenience premium rather than a capability premium, and convenience premiums get competed away. See: 3 cents a day.
If you wait, do not wait idle. Deutscher's alternative is the right one: "go get properly fluent with Claude Code first. You'll learn more, faster, for a fraction of the cost, and get most of what Grok Bot can build anyway." Spend the quarter writing job descriptions instead of buying seats. Pick three recurring jobs, write each as one sentence with a measurable outcome, run them by hand for two weeks, and log every point where a human stepped in. That log is the escalation policy nobody has published a number for. When you do buy, you hand the bot a job instead of a title.
The split between the bot that paid for itself and the bot that got fired was never the model. Same product, same week. One operator handed it a sentence. The other handed it a nameplate.
Scope the job first.
Find your one-sentence job.
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We map which recurring jobs in your business are scoped tightly enough for an agent today, and which would burn $300 a month producing work nobody asked for.
Book the Diagnostic →