ISA After Hours Augsburg
A phone propped against a dark green mug on a café table at night, its screen showing a blurred chat, a notebook and pen beside it, rain and city lights on the window.

AI at work · Tools

Grok Bot: what it is good at, who it fits, and what it costs

· About 35 min read

It works with the old internal system that will never have a proper integration, because it clicks through it the way a person would. You teach it by showing it once instead of describing it. That combination decides the whole question of fit: which jobs it does well, which teams those jobs belong to, and which kinds of work it should not be given at all.

In this article
  1. What it actually is
  2. What is actually new
  3. The jobs it is doing, and who they suit
  4. Is it for you?
  5. Why it clicks for so many people
  6. You show it, you do not explain it
  7. A team of them
  8. The benchmarks nobody published
  9. What it costs, in three real jobs
  10. Grok Bot or a developer agent
  11. Four briefs you can copy
  12. What to be careful about
  13. A teammate, not a system
  14. Where it breaks
  15. A first week, and the 30 days after it
  16. Bottom line
  17. FAQ
  18. Sources

Short version

Grok Bot gives you AI workers rather than an AI chat. Each one holds a role, signs into your tools, remembers how you like things done, and keeps working when your laptop is shut. Two properties decide who it fits. It needs no integration — it operates your systems through the screen, like a person, so it reaches the internal tools that will never have a programming interface. And you teach it by demonstrating a task once rather than describing it. Since 26 August 2026 it has been bundled into every paid Grok and Cursor plan. It suits small, repeatable, boring jobs; it does not suit work needing isolation, an audit trail or a named model.

Key facts

What this article is about

Most write-ups of a tool like this measure it against the most powerful alternative and mark it down. That misses the point. Convenience is not a weaker version of capability — it is a different requirement, and for a great deal of ordinary work it is the binding one. So this piece is organised around fit: the jobs Grok Bot does well, the kinds of team those jobs belong to, what a month of it costs, and the work it should not be given. Everything factual below comes from xAI's own documentation or from a named outlet; where the marketing and the documentation disagree, the article says so.

What it actually is

You are not opening a chat window. You are staffing a role. xAI's documentation defines a Bot as "a single persistent, named agent" — one AI teammate — that runs on a persistent cloud VM with a browser, filesystem and terminal. A VM, or virtual machine, is simply a computer that exists as software in a data centre rather than on your desk; the practical point is that it stays switched on when yours does not.

Two terms are worth defining here, because the whole product turns on the difference. An API (application programming interface) is a purpose-built connection that lets one piece of software talk to another directly — clean, fast, and only available if somebody built it. Computer use is the alternative: the software looks at a screen and moves a pointer, the way a person does. xAI's docs say a Bot "can use connectors/MCP where available, and computer use for apps and websites without a clean API". That second half is the sentence that explains the adoption.

ItemWhat it isWhat it means for you
Launch11 August 2026, described by xAI as an early beta (xAI launch post)Under a month old at the time of writing. Terms, limits and features are still moving.
Where it runsA persistent cloud computer with browser, filesystem and terminal (xAI docs)Work continues when your laptop is closed. It is not running on your machine.
IsolationOne computer per account, not per Bot; all Bots share files, browser sessions and logins (xAI docs)You cannot wall one Bot off from another. Anything you connect is available to all of them.
How it reaches your systemsConnectors ("Plugins") where they exist, otherwise computer use through the ordinary screen (xAI docs)Legacy and internal software becomes automatable without anyone building an integration.
How you teach it"Teach a task": it records up to ten minutes of visible screen interaction, no audio, and drafts a reusable skill (xAI docs)Ten minutes is the hard ceiling on a single demonstration. Long processes must be taught in pieces.
Scheduling limitsUp to 50 routines per Bot; the 20 most recent run records kept per routine (xAI docs)Enough for a real workload, but your run history is short — export anything you need for an audit.
Model choiceRouted automatically; no manual selection (Gadgety, Hebrew)You cannot say which model processed which data. For some compliance positions that is a hard stop.
Price, checked 6 September 2026No standalone subscription; bundled into eight Grok and Cursor plans from $20 to about $300 a month (CellCog, xAI)If you already pay for either product, you may already have it. Check before buying anything.
PlatformsmacOS, Windows, Linux and iOS in beta at launch; AY Automate reports Android has since shipped on Google Play (AY Automate, checked 6 September 2026)The mobile app is reported to carry every bot, chat and routine over from the desktop, so you can approve an action, or take over a session, from a phone.

Prices and availability checked 6 September 2026. xAI's own consumer pricing page returns an error to automated fetches, so plan prices here come from CellCog and are corroborated by VentureBeat, Gadgety and AY Automate — see Sources.

Every Grok Bot concept explained for normal people — Nate Herk's walkthrough (YouTube, not affiliated with xAI or with us). Plays from youtube-nocookie.com only after you click.

What is actually new

Five things, none of which is a research breakthrough, all of which change who can use automation at all.

The jobs it is doing, and who they suit

Fit is decided by the job, not by the product, so the use cases come before the mechanism. xAI documents eight starter roles and is unusually direct about the boundary: Bots suit work that "owns a repeatable outcome, not a loose category of questions", and the vendor says they are not for unsupervised production changes. Let's AI, reviewing it in Hebrew, arrives at the same place from the other end — it lists weekly reports, pre-meeting audits, CRM updates and draft emails as the sweet spot.

The jobWhat the bot doesRoute it usesWho it fits, and why
The weekly reportShown once how to pull the numbers, then produces it every week unaskedRoutine on a schedule (docs)A team lead whose numbers sit in a system with no usable export. The task is identical every week, which is exactly the repeatable outcome the vendor says a Bot is for (docs)
Keeping the CRM honestTurns call notes into updated records, which nobody has ever enjoyed doingConnector if one exists, otherwise the screen (docs)A small sales team with no operations person behind it. "Account health" is one of the eight roles xAI documents (docs)
Expense reconciliationWeekly reconciliation and drafting the chasing emails — a documented vendor use case (docs)Connector plus approval before sendingFinance in a company too small for a finance system. The drafting stops at an approval gate, and the vendor warns an approval does not reverse completed work
Watching somethingChecks a page or an account on a schedule and tells you when it changesComputer use on a routine (docs)Anyone monitoring a supplier portal or a competitor page nobody built an alert for. Watch the bill here: long accumulated context crosses into the higher token tier (CellCog)
Email triagePulls new messages and summarises what actually needs youConnectorAnyone whose inbox is the work queue. Let's AI names drafted emails and pre-meeting audits as the reviewed sweet spot (Let's AI)
Research inside a systemLogs into the CRM itself and works through the accounts like a person wouldComputer use (docs)Roles stuck inside industry software with a few hundred customers worldwide — the case where no other tool reaches the data at all (Gadgety)

Notice how unglamorous that list is. Not one of these is a task anybody enjoys, and not one of them is intellectually hard. They are the twenty minutes a day that never justified a project. That is almost always where automation actually pays, and it is a category that traditional automation tools were structurally bad at reaching, because a twenty-minute task never justified an integration either.

What this means for you: the fit signal is the same in every row — a boring, repeatable outcome, in a system nobody was going to integrate, where the person who does it could demonstrate it faster than they could write it down. If your candidate task has all three, this product is aimed at it. If it has none of them, no amount of setup will make it a good match.

Is it for you?

The honest version of the fit question is not "is this good", it is "which of these four statements is true of your work". Each row below is a condition the documentation or the reporting settles, not a matter of taste.

Suited if…Not suited if…
Your work sits in software that has no programming interface and never will — screen-driving is what the product is built on (xAI docs)You have to be able to say which model processed which data: routing is automatic, with no manual selection (Gadgety)
You could demonstrate the task in ten minutes but could never write it down — that is the exact shape of the "Teach a task" recording (xAI docs)You need one job's access walled off from another's: every Bot shares one cloud computer, and xAI says not to use separate Bots as a security boundary
The job owns a small, recurring, read-only outcome that nobody was ever going to build for you — the vendor's own definition of the fit (xAI docs)Someone will ask you for an audit trail: Action Recording and audit logs are Enterprise-only, and Action Recording is off by default
You already pay for a Grok or Cursor plan and want something working this week — it is bundled rather than sold separately (xAI, CellCog)You need to read and change the machinery rather than a drafted skill: you can edit the skill, but the routing beneath it is closed (xAI docs)

A lot of work sits in the left column, including plenty done by people who would describe themselves as technical — and a single organisation usually has both columns in it at once. Let's AI's review names the three conditions for looking elsewhere in the same terms: cost-conscious users, highly sensitive data and strict permission isolation.

Why it clicks for so many people: nothing needs to be integrated

This is the part that explains the adoption, and it is almost never the headline.

Automation tools have always asked the same question first: does your system have an interface we can connect to? If the answer was no — the internal system nobody has touched in a decade, the supplier portal, the industry software with four hundred customers worldwide, the ERP your company paid to have customised — you were out. Not "harder". Out.

Grok Bot does not ask. Its Bots work through the ordinary screen, clicking and typing the way a person does. Gadgety's launch coverage describes them as connecting to services, applications and websites through their regular user interface, bypassing the need for an API and making legacy systems reachable. xAI's own documentation says the same thing more soberly, and adds a caveat the marketing does not: "Prefer a connector when one is available: it is often more reliable than clicking through a website."

Read that vendor sentence twice. The company selling you screen-driving is telling you screen-driving is the fallback. That is unusually honest documentation, and it is also the single most useful line in the product for planning purposes: use the connector where one exists, and treat computer use as the thing that rescues the systems nothing else can reach.

Think about what that means for who can use this. The people whose work sits in systems that will never get a modern integration are not a niche — in most companies, they are the majority. Finance, operations, purchasing, HR, service. They have been told for a decade that automation was for other people's tools. This is the first version of the promise that arrives on their side of that line.

What this means for you: if a proper connector exists for the system you care about, use it and expect it to hold. If it does not, you now have an option you did not have — and you should expect it to need supervision, because the vendor has told you it is the less reliable path.

Two routes into a system, and what each one costs you in reliability
How a Bot reaches a system Bot cloud computer Connector (Plugin / MCP) structured — xAI: "prefer a connector" Computer use drives the screen — breaks when the page moves Your system CRM, ERP, portal, inbox Source: xAI documentation, "Computer and apps", checked 6 September 2026

You show it, you do not explain it

The second barrier that quietly disappears is the harder one.

Automating a task has always meant describing it precisely — in a script, a workflow builder, or a very careful set of instructions. That description was the bottleneck, and not because people are not clever. It is because most of any job consists of small decisions the person stopped noticing years ago. Ask someone to write down how they do the Monday report and you will get eight steps. Watch them do it and there are twenty-three.

Here is the mechanism, precisely. You open a Bot conversation with the computer view, choose Teach a task, say what result you want, and do the job once. xAI's documentation says the system records visible computer interaction for up to ten minutes and does not record audio, then produces a skill — the vendor's word for a reusable set of instructions covering steps, decision rules, expected output and safety boundaries. A routine is the separate object that decides when a skill runs. And the documentation is blunt that what comes out of a demonstration is a draft that still needs decision rules and failure handling added by you.

That last clause is the whole difference between a demo and a working automation. One recording teaches the machine what you did on a day when nothing went wrong. It does not teach it what to do when the page fails to load, when a field is empty, or when the number looks implausible. The vendor also notes the feature is rolling out gradually, so it may not be switched on for your account yet — in which case you write the skill from instructions instead.

What this means for you: budget an hour, not five minutes. The demonstration takes ten minutes at most; the review that turns the draft into something you would leave alone takes longer, and is the part that decides whether this works.

A team of them, not one of them

The design goes a step further than a single assistant. You can run several Bots at once, each with a specialism, and they can pass work and information between themselves in a shared conversation. Gadgety's launch write-up describes xAI's own internal setup: a "Chief of Staff" Bot overseeing specialists in research, support, HR and development, transferring tasks and information to each other independently. Chief of Staff is also one of the eight roles xAI documents as intended use cases, alongside sales outbound, talent scout, paid media, expense manager, product performance, bug reproduction and account health (xAI docs).

Technically the parallelism is real but bounded. xAI's documentation states that several Bots can use browser and desktop tools in parallel, but that one Bot can run only one computer-use task on its screen at a time, and that those screens are "separate work surfaces, not separate security boundaries". Each Bot can hold up to 50 routines, which is more than most people will ever fill.

That is a genuinely different mental model, and it is the one most likely to be oversold. Coordinating several agents adds management overhead of exactly the kind you would recognise from managing people: unclear ownership, work handed on with half the context, and nobody noticing when something stalls.

What this means for you: start with one Bot. A team of agents is something to grow into once you have watched a single one work for a month, not something to design on day one because the diagram looks impressive.

The benchmarks nobody published

This section is short because there is almost nothing in it, and that absence is itself the most useful fact in the article.

A benchmark is a standardised test that lets you compare one system against another on the same task. xAI's launch announcement contains none. Reading the announcement, what you get is capability claims — it signs into tools, it works while you sleep, it learns your preferences — and customer testimonials, including one claiming users became "2-3x more efficient". That is a vendor claim from an unnamed methodology, and it should be read as marketing, not measurement. VentureBeat's coverage makes the same observation from the outside, noting no published performance benchmarks and quoting early tester Matt Shumer's description instead: "an agent for everything, not just code."

Nor is there any independent measurement. We could not find, from any source reachable here, a third-party test of how often a Grok Bot routine completes without intervention, how it recovers from a failed step, or how stable a taught skill is across weeks. For a chat product that gap would be tolerable. For a product whose selling point is that it works unattended inside your systems, reliability is the specification, and nobody outside xAI has measured it.

What this means for you: you are the benchmark. Whatever you pilot, count two numbers yourself — how many runs finished without you, and how long checking took. Those are the only figures about this product that currently exist with a method behind them.

What it costs, in three real jobs

There is no standalone Grok Bot subscription; it is bundled. As of 26 August 2026 xAI states that "Grok Bot comes with its own usage, separate from your Grok and Cursor plans, so anything you hand off to a Bot won't count against your existing usage" — which is the detail that turns it from an upsell into something you may already be paying for.

The access ladder moved fast. CellCog's timeline records the entry point falling from $200 a month at launch on 11 August, to $60 on 21 August, to $20 on 26 August, by which point eight plans bundled it. Note a disagreement worth knowing about: xAI's own launch page today lists the full set of plans, while CellCog's dated timeline records that set arriving in three steps over two weeks. Vendor pages get quietly updated; the timeline is the better guide to how young this is.

Cheapest monthly plan that includes Grok Bot, by date (US dollars)
Entry price for Grok Bot access, August 2026 (USD per month) 0 100 200 $200 11 Aug Cursor Ultra $60 21 Aug Cursor Pro+ $20 26 Aug Cursor Pro Source: CellCog pricing timeline, checked 6 September 2026

Per-token rates are not pricing, so here is what three ordinary jobs actually cost. Each plan carries a weekly Grok Bot allowance; Let's AI reports the caps are measured in agent steps and tokens rather than messages, and that a trial credit runs seven days. Past the allowance, CellCog reports on-demand billing at $2 per million input tokens and $6 per million output, doubling to $4 and $12 above 200,000 prompt tokens. A token is roughly three-quarters of an English word, so a million tokens is about 750,000 words.

The jobRough size per runOverage cost if you exceeded your allowance
Weekly sales report: log in, pull four screens, write 600 wordsOur estimate: about 60,000 input tokens, 1,000 outputAbout $0.13 a run, or $0.55 a month at $2/$6 per million
Daily inbox triage: read 80 messages, summarise 10Our estimate: about 120,000 input tokens, 2,000 outputAbout $0.25 a run, roughly $5.50 a month at the same rates
Overnight monitoring: 40 page checks, long accumulated contextOur estimate: over 200,000 prompt tokens, so the higher tier appliesAbout $0.85 a run at $4/$12 per million, roughly $25 a month

Token counts are our own estimates for illustration, not vendor figures; the per-million rates and the 200,000-token threshold are CellCog's, checked 6 September 2026. The honest headline: for the jobs most people will try, the subscription is the cost and the tokens are noise — but the long-context tier is where an unattended monitoring Bot can surprise you, and xAI's documentation states that "a separate Grok Bot spend cap is not available today".

Grok Bot or a developer agent: which job goes to which

The obvious alternative to Grok Bot is not another chatbot, it is a developer agent — a command-line tool that will do more or less anything if somebody assembles the environment around it. The two are not competing for the same job, and knowing which category a task belongs to is most of the decision.

The dividing line is who owns the machinery. Grok Bot ships a finished environment you staff: Gadgety's launch coverage states that it routes models by task type with no option for manual selection of a specific model by the user, and the workspace is pre-built and largely closed. A developer agent hands you the parts instead — including xAI's own Grok Build, the command-line coding agent it launched on 18 May 2026, which is a different product for a different job.

QuestionGrok BotA developer agent (Claude Code, Codex, Grok Build)
Who sets it up?You, in an afternoon, from a desktop or phone app (xAI)Someone comfortable in a terminal, over days (eWeek)
Which model runs your job?Chosen for you, not overridable (Gadgety)You choose, per job
Reaching a system with no APINative — that is the point (xAI docs)Possible, but you build it
Can you read and edit the steps?You edit a drafted skill; the routing beneath it is closed (xAI docs)Everything is inspectable by design
Isolation between jobsNone — one shared computer per account (xAI docs)Whatever you configure
Audit trailAction Recording, Enterprise only and off by default (xAI docs)Your own logs, at your own cost

That is the trade, and it is a real one, made in both directions. Let's AI's Hebrew review framed it as choosing low friction over raw power, naming Claude Code, Codex, OpenClaw and Hermes as tools that do comparable things while demanding more setup. Owning an environment is a cost as well as a capability: for a recurring business task it buys nothing, and for something other people will depend on it is the only thing that makes the work reviewable.

What this means for you: route the job, not the person. A boring recurring task inside a system nobody will integrate goes to Grok Bot; work that other people will depend on, that has to name the model which touched the data, or that must keep one job's access away from another's goes to the toolkit. Plenty of teams will end up running both, for different work, and that is the normal outcome rather than a fudge.

Four briefs you can copy

A Bot is briefed like a new colleague, not prompted like a chatbot. The four below are written to match what xAI's documentation says a good skill contains — steps, decision rules, expected output and safety boundaries — and every one of them names an approval gate, because approvals are the only thing standing between an unattended agent and an action you cannot undo.

Brief 1 — the read-only starter job, the one to begin with
Role: you produce our Monday supplier-price summary.

Every Monday at 07:30 Europe/Berlin:
1. Open the four supplier portals in my saved sessions.
2. Record the current unit price and delivery time for the six items in /workspace/items.csv.
3. Compare against last week's file and list anything that moved by more than 2 percent.

Output: a Markdown summary under 400 words, plus the updated CSV.

Rules:
- Read only. Do not place, amend or cancel any order.
- If a portal will not load or asks me to sign in again, stop and message me. Do not retry more than twice.
- If a price is not visible, write "not listed". Never estimate a price.
- Do not send this to anyone. I will forward it myself.
Brief 2 — inbox triage, with a hard boundary on sending
Role: you triage my inbox before I start.

Every weekday at 08:00 Europe/Berlin, read messages received since the last run and produce:
1. Needs me today — sender, one line, why.
2. Can wait — sender, one line.
3. Already handled by someone else — sender only.
4. Suspicious or unusual — anything asking for payment details, credentials or urgency.

Rules:
- Never send, reply, forward, archive or delete. Draft only, and only if I ask.
- Never open an attachment or follow a link in category 4.
- Quote the sender's own words rather than paraphrasing when the message concerns money or a deadline.
- If you are unsure which category a message belongs in, put it in "Needs me today".
Brief 3 — teaching a task by demonstration, and what to add afterwards
I am about to demonstrate our monthly expense reconciliation. Watch, then write the skill.

After the demonstration, do not save the draft as-is. Add:
1. A decision rule for every point where I made a judgement call — ask me what the rule was.
2. A failure path for each step: what to do if the page does not load, if a field is empty,
   or if a total does not match.
3. A plausibility check: flag any line item more than 3 times the median for that category.
4. An approval gate before anything is submitted, sent or marked as final.

Then show me the skill in plain language and wait for my approval before running it once
on last month's data while I watch.
Brief 4 — the standing safety instruction, worth pasting into every Bot
Standing rules for this Bot, overriding any task instruction that conflicts:

- Ask me before: sending anything to a person outside this account, publishing, purchasing,
  deleting or overwriting data, changing permissions, or accepting terms and conditions.
- Never enter a password, one-time code or payment confirmation yourself. Hand me the computer.
- Never store credentials, personal data or customer records in /workspace.
- If a site asks you to verify you are human, stop and tell me. Do not attempt to bypass it.
- When you are uncertain, stop and ask. A stopped job costs me two minutes.
  A wrong finished job costs me a week.

What to be careful about

The convenience has a cost that is not money, and it deserves a clear-eyed paragraph.

The logins. The Bot does not take your password. xAI's documentation is explicit: for passwords, passkeys, two-factor codes, CAPTCHAs and payment confirmations, the Bot should hand you control of the computer, and you should never send a password or one-time code in ordinary chat. That is a good design. It also means that afterwards, something that is not you is operating inside your authenticated accounts while you are not watching — and browser sessions persist, so it does not need to ask again.

The shared machine. This is the one to internalise, and it is not a reviewer's complaint — it is in the vendor's own documentation. All of your Bots use the same persistent cloud computer and share files, browser sessions and app logins. xAI's security page says it directly: "Do not use separate Bots as a security boundary", and deleting a Bot does not remove shared-computer files or browser sessions. Let's AI's Hebrew review reached the same conclusion independently, warning that separate bots should not be treated as a security boundary. Note also that Grok Bot does not support Legacy Privacy Mode, so if your account runs in it, the product is off.

Approvals do not undo things. The documentation contains one sentence worth pinning above the desk: "An approval controls the proposed action. It does not reverse work already completed." An approval gate is a brake, not a rewind.

  1. Start with tasks that only read. Watching, gathering, summarising. Nothing that sends, buys, deletes or writes into a system of record until you have seen it work for a while.
  2. Give it its own account with only the permissions the job needs — exactly what you would do for a new contractor, for exactly the same reasons.
  3. Keep it away from anything you cannot afford to have go wrong quietly: payments, contracts, customer data, anything regulated.
  4. Set approvals for the whole list xAI names — sending, publishing, purchasing, deleting or overwriting, permission changes, production changes and accepting legal terms.
  5. When you stop using a Bot, do the cleanup by hand: pause its routines, sign out of the services on the shared computer, uninstall its connectors and remove sensitive files from the workspace. Deleting the Bot does not do any of this for you.
  6. Check its work for the first few weeks. Not because it is likely to misbehave, but because that is how you build the sense of when it has.

Questions for the people who own this in a German company — these are questions, not legal advice, and the answers belong to your data protection officer and your works council, not to us:

  • Does a Auftragsverarbeitungsvertrag (data processing agreement) exist covering this processor, and does it cover personal data reaching a shared cloud computer through a browser session?
  • Where is that computer, and can anyone tell you? xAI's documentation does not state data residency; the teams page directs you to the account team for commitments.
  • Under the EU AI Act's transparency expectations, do the people whose emails, records or performance data an agent reads know that it is doing so?
  • Does this change how work is monitored or performed? In Germany that is usually the works council's business before it is a pilot.
  • If you need an audit trail, note that Action Recording and audit logs are Enterprise-only, and Action Recording is off by default. Answer this before the pilot, not after it.

A teammate, not a system

The sharpest line written about this product so far comes from Beam's enterprise analysis, which calls it "a self-serve teammate, not a governed enterprise deployment" — and for anyone putting agents anywhere near production systems, that distinction is the whole story. Beam also notes how it arrives: through individual consumer subscriptions, activated on personal laptops before IT is involved.

It is worth spelling out what "governed" means here, because it is not a synonym for expensive.

  • Separation. The shared cloud computer does not give you one assistant walled off from another; the vendor says so in its own security documentation. If your compliance position depends on that separation existing, it does not.
  • Permissions and review. Beam's assessment is that there are no scoped permission controls and no built-in audit trail at the level most teams buy at.
  • Visibility. The controls that would give an administrator a record — Action Recording, audit logs, OpenTelemetry export, SCIM provisioning, network controls — are Enterprise-tier only, and Action Recording is off unless someone turns it on. Single sign-on through Okta or Microsoft Entra ID is available; the logging is the gap.
  • Choosing the engine. The system routes work to models automatically and you cannot override it (Gadgety). VentureBeat reported an early tester's view that the router "wasn't great" initially. If you have to be able to say which model processed which data, that is a hard stop.
  • Independent evidence. Reliability over long runs, recovery from failure and repeatability week after week have not been tested by anyone outside the vendor, and the launch post carries no benchmarks at all.

One thing that is easy to get wrong: Grok Bot is not confined to its own app. There is an iOS app reported to have full parity with the desktop (AY Automate), and routines can be triggered by events from connected accounts, including a Slack message or a GitHub notification. The channels are not the problem. The record of what happened in them is.

None of this is a reason for an individual not to try it on a small task. It is the reason a company should not treat it as infrastructure yet. Organisation-wide enablement is an Enterprise-only switch, and CellCog reports that from 3 September 2026 enterprise customers get two free weeks but pricing remains unpublished and requires a sales conversation — which is itself an admission of where the product is.

Where it breaks

The thing that makes it work is also its weak point. Driving a screen means depending on the screen staying the same — which is precisely why xAI's own documentation tells you to prefer a connector wherever one exists.

  • The site changes and the routine breaks. A moved button is nothing to a person and fatal to a recorded sequence.
  • Logins expire, and CAPTCHAs exist. Both stop an unattended Bot dead by design: the documentation says the Bot should pause rather than bypass a re-verification check.
  • A demonstration teaches whatever you happened to do. The vendor calls the result a draft that still needs decision rules and failure handling. If the day you recorded it contained an unusual case, you have taught it the exception as though it were the rule.
  • Ten minutes is the ceiling on one demonstration (xAI docs). A long process has to be broken into taught pieces, and the joins are where things go wrong.
  • Your run history is short. Only the 20 most recent run records per routine are retained. If you need evidence of what happened six weeks ago, export it.
  • It is a beta. Expect things to move, including terms — the access ladder changed three times in the first fifteen days.

None of this is disqualifying. It does mean that "set it up and forget it" is the wrong expectation, and that the right question after a month is not whether it worked but how often somebody had to go and fix it.

A first week, and the 30 days after it

Pick something you already do and fully understand. xAI's own guidance points the same way: it says Bots suit a role that owns a repeatable outcome rather than a loose category of questions, and that they are not for unsupervised production changes.

  1. Days 1 to 2: pick one recurring task you do yourself, that takes 10 to 30 minutes and only involves reading and summarising. Check first whether a connector exists for the system involved; if it does, use it.
  2. Day 3: set up one Bot. Paste in the standing safety brief above, then demonstrate the task once, end to end, including the parts you normally do without thinking.
  3. Day 4: review the drafted skill and add what the recording could not know — the decision rules, and what to do when a step fails.
  4. Day 5: let it run while you watch. Compare what it produced with what you would have produced.
  5. Days 6 to 7: let it run without you, then check it properly. This is the first honest test.
  6. Days 8 to 21: leave it on a schedule and record the two numbers in the table below every single run. Change nothing else.
  7. Days 22 to 30: decide. Only then consider a second Bot, or a task that writes something — and if it writes, put an approval gate on it.
What to measureHowWhat a pass looks like
Unattended completion rateRuns that finished without you, divided by total runsAbove 80% by week three, and rising rather than falling
Checking timeMinutes spent verifying the output, per runLess than a third of the time the task used to take
Silent errorsWrong outputs you only caught because you lookedZero. One is a reason to narrow the task, not to add a second Bot
Intervention causesNote why each failure happened: page changed, login expired, unusual caseA short list you can fix, not a scatter
Cost beyond the planAny on-demand token charges (rates)Effectively zero for a read-and-summarise job; watch it if you add monitoring

If checking takes as long as doing, you have not saved anything — and learning that in a month is a good outcome. One thing to hold on to while you do this: our piece on what happens when AI replaces effort instead of adding to it applies here more than anywhere. A Bot that quietly takes over a task nobody understands any more is a cost that arrives long after the saving does.

Bottom line

Grok Bot is not the most capable agent available, and that is not what it is competing on. It is the first one that meets ordinary people where their work actually is: inside systems nobody was ever going to integrate, described in a language nobody can write down but everybody can demonstrate. It is also under a month old, ships with no published benchmarks, and its own documentation tells you that separate Bots are not a security boundary.

So the question is which of your jobs it is for. If your work lives in tools that were never going to be automated, and the barrier has always been that you could show the task but not specify it, this is aimed squarely at that work — start with one Bot, one read-only job, and one month of counting. If the job instead has to name the model that touched the data, keep one job's access away from another's, or hand an auditor a trail, it belongs somewhere else, and that is a statement about the requirement rather than a rating of the product.

FAQ

What is Grok Bot, in one sentence?

xAI's product for assigning a job to an AI worker that signs into your tools, remembers how you like things done, and keeps running when your laptop is closed. It has been in early beta since 11 August 2026.

Why do people find it convenient?

Two reasons. It works with systems that have no integration, because it operates them through the screen like a person. And you teach it by demonstrating a task once instead of describing the steps.

Do I need to be technical?

No. There is no code and no flowchart. If you can do the task while it watches, you can set it up — though you will still have to add the decision rules the recording could not capture.

What does it cost?

There is no standalone subscription. It is bundled into Grok and Cursor plans, and xAI says its usage is separate from your existing Grok and Cursor allowances. CellCog's timeline records the entry price falling from $200 a month at launch to $20 by 26 August 2026. Checked 6 September 2026; terms have moved repeatedly, so verify before budgeting.

Does each bot really get its own computer?

No, and this is the most important correction to the marketing. xAI's documentation states that all of your Bots share one persistent cloud computer scoped to your account, including files, browser sessions and app logins. Each Bot gets its own screen, which the docs describe as a work surface, not a security boundary.

Is it safe to give it access to my accounts?

It never holds your password — xAI's docs say it should hand you control for passwords, two-factor codes, CAPTCHAs and payment confirmations. But it then acts inside your accounts unattended, the shared computer means one Bot's access is not walled off from another's, and deleting a Bot does not remove its files or sessions. Start read-only, give it the narrowest access that works, and keep it away from regulated data.

How long a task can I demonstrate?

Ten minutes. The Teach a task recording captures visible screen interaction for up to ten minutes and does not record audio. Longer processes have to be taught in pieces.

Can I choose which AI model does the work?

No. Gadgety's launch coverage states there is no option for manual selection of a specific model; routing is automatic. If your compliance position requires naming the model that processed a given piece of data, that is a hard stop.

Are there benchmarks showing how reliable it is?

None that we could find. xAI's launch post publishes capability claims and testimonials but no benchmark scores, and VentureBeat noted the same absence. No independent measurement of unattended completion rates exists from any source we reached.

Will it keep a record of what it did?

Partly. A routine keeps only its 20 most recent run records, and the administrator-visible Action Recording and audit logs are Enterprise-only, with Action Recording off by default. If you need evidence months later, export it yourself.

Can it spend money without me noticing?

Purchases are on xAI's list of actions that should require approval, and the documentation warns that an approval controls the proposed action but does not reverse completed work. On the bill side, xAI states that a separate Grok Bot spend cap is not available today, so account-level controls are what you have.

Which is better, the connector or the screen?

The connector, on the vendor's own advice: xAI's documentation says to prefer a connector when one is available, because it is often more reliable than clicking through a website. Screen-driving is what rescues the systems nothing else reaches, not the default.

Why would someone choose a different tool?

If you need to control which model handles a task, own the environment, or inspect and change the steps rather than accept a recorded routine, Grok Bot is deliberately closed in those areas. Let's AI names Claude Code, Codex, OpenClaw and Hermes as tools that give you that control in exchange for far more setup.

Should we roll it out across a team?

Not yet. Beam's analysis calls it a self-serve teammate rather than a governed enterprise deployment, and organisation-wide enablement is an Enterprise-only switch. One person, one Bot, one read-only task, one month — then a conversation with whoever owns data protection, and in Germany with the works council, before it touches business systems.

ISA After Hours · Augsburg

Tried this on something real and want to compare notes?

ISA After Hours is a community of Israeli and international tech professionals in Augsburg. We meet after work, run these tools on actual tasks, and are equally interested in the ones that did not work out.

Join ISA After Hours →

Sources