You see a screen. Your AI assistant usually doesn’t.
That sentence needs one qualification, which we will get to. But it corrects the picture most of us carry in our heads.
When I open a website, I see the design and the button I am supposed to press. I assumed an AI assistant saw roughly the same thing, only faster. Then I asked the more basic question: what does it actually receive?
The answer is not one thing. An assistant can find a site, read a site or operate a site. Those are separate jobs using different inputs. If we want pages that work well for AI assistants, we have to stop lumping them together.
An assistant meets your website three different ways
Finding: the search result is the pitch
When an assistant searches the web, its first view is closer to a search-results list than a browser window. It may receive a title, URL and short snippet.
At that moment, your title tag and meta description are the entire pitch. The assistant has to decide whether your page can answer the question before opening it. Name the subject plainly.
Reading: the page becomes a stream of text
When Muse opens a public page for information, the normal reading path is text-first. Useful content is extracted and returned as headings, paragraphs, lists and links in roughly page order.
The design largely falls away. The assistant is not admiring the hero section or noticing that a price sits inside a gold circle. It is working from the words the page exposes.
Images may arrive as markers and file addresses: there is an image here, and here is where it lives. That is not the same as seeing it. If a crucial fact is baked into the pixels—“$199,” “ships free,” “five-year warranty”—the reading path may hit a blank spot. Useful alt text can carry some of that meaning. “Technician using a moisture meter on wet drywall” communicates something. “IMG_4827” does not.
Doing: a browser worker operates the screen
The picture changes when the user asks the assistant to do something: log into HubSpot, update a record, complete a form or buy a product.
A separate browser program can open a real browser on a server. It loads the interface, takes visual observations or inspects the page’s interactive structure, clicks, types and reports what happened back in words.
That is the qualification to “usually.” A screen may be used inside the process, but the conversational assistant is not sitting behind the glass like a person. It receives observations from a browser tool and sends instructions back. The browser side is the eyes and hands; the assistant works through an intermediary.
A page can be easy to read as an article and miserable to operate as an application. It can look obvious to a person while presenting the browser worker with five unlabeled controls called “button.”
WordPress made the abstraction visible
I had already seen a simpler version in our WordPress work without connecting the dots.
When we pull a post through the WordPress REST API, the content can arrive as raw HTML: words plus tags for headings, paragraphs, links, lists and styling wrappers. The reading step removes the markup noise while preserving the words and structure.
That is “cleaning the HTML.” We are removing the packaging, not the article. The tags still matter: a heading announces a section, a list groups items, and a link identifies a destination. Good HTML carries meaning. Bad HTML creates boxes that look right but say little about what they are.
Accessibility is the closest thing to an agent-ready standard
Here is the practical money line: the work that makes a website easier for a blind person to use also tends to make it easier for an AI browser agent to use.
Browsers build an accessibility representation from the page’s Document Object Model. Assistive technology uses it to understand roles, names, states and relationships: this is a heading, that is a link, this button is named “Save contact,” and this checkbox is checked.
Muse’s browsing side is reported to rely heavily on this kind of page structure, along with visual observations when needed. Meta does not publish a complete specification for the Muse browsing pipeline, so treat that as a field report from using the product, not permanent platform documentation.
The implication is still solid. Use real buttons with useful names. Label form fields. Put headings in a sensible order. Give links meaningful text. Preserve keyboard focus. Describe informative images.
A screen-reader user needs those things. So does a browser agent working without human intuition. Accessibility and agent-readiness are not identical, but they are close cousins.
HubSpot shows what an agent-native application could be
Imagine HubSpot—or any software platform—shipping an interface designed for assistants to navigate with less friction. It would not need a blank, text-only clone. It could make the existing product more legible to software: real controls with specific names, labeled form fields, clear headings and landmarks, properly identified table headers, programmatic state changes, and no critical action hidden behind hover or an unlabeled icon.
That is an agent-native site. It is not a secret internet for bots. It is a website or application whose meaning survives when the visual layer is translated into structure and words.
The same work also helps keyboard users, screen-reader users, automation tools and QA teams.
llms.txt is a map, not a second website
The closest public convention aimed directly at AI readers is llms.txt. The proposal describes a Markdown file, usually at a site’s root, that gives language models a short explanation of the site and links to important pages or cleaner Markdown versions.
Think of it as a curated map: here is what we do, here are the pages that matter, and here is where to find the details.
It cannot repair an unlabeled checkout button. It does not replace accessible HTML, describe the viewport or guarantee that an assistant will use it. Add it if it helps explain the site. Do not mistake it for an agent interface.
What a site owner can change Monday morning
The useful changes are ordinary, testable website work.
Write a real title and meta description. Name the subject plainly.
Put every money fact in visible HTML text. Price, specifications, shipping, availability and guarantees should not live only inside graphics, video or a brochure.
Use semantic HTML. Use headings for headings, buttons for actions, links for navigation and labels for form controls. A styled <div> may look like a button while remaining a nameless container to other systems.
Write alt text that carries meaning. Describe what an informative image contributes. Mark decorative images as decorative instead of stuffing them with keywords.
Add accurate structured data. Product and Offer markup can identify price and availability. FAQ markup can describe genuine questions and answers. Schema must match the visible page.
Server-render critical content when practical. If the offer, price or primary action appears only after a fragile JavaScript sequence, some readers and tools may miss it.
Give each landing page one job. One offer, one explanation and one primary action reduce ambiguity for people and agents.
Test the nonvisual path. Use the keyboard, inspect the accessibility tree, try a screen reader and pull the page through a text extractor. Do the product, price, proof and next step still make sense without styling?
None of this requires uglier design. It requires the design and the underlying structure to tell the same story.
This is a field report, not a permanent specification
This article describes Meta’s Muse as it works today, based on direct experience building and operating websites with it. It is not a published Meta protocol.
Claude, ChatGPT, Gemini and other assistants broadly rhyme with this pattern, but the details differ. Their full pipelines are not public, and they are changing quickly.
Cleaner HTML will not automatically increase AI citations tomorrow. Citation systems involve discovery, retrieval, ranking, trust and answer construction. There is no magic switch.
The immediate opportunity is closer to the customer. Someone sees your ad on Threads or Facebook, opens the landing page, then asks an assistant: “What does this cost?” “Is the guarantee real?” “How does this compare?” or “Can you sign me up?”
If the facts are clean text, the assistant can explain them. If the controls are properly labeled, the browser side has a better chance of completing the task. If the facts live inside an image and checkout uses unlabeled custom controls, the assistant has to guess, fail or hand the job back.
That moment is already here.
The next website has two front doors
The site of the near future has two front doors: one for eyes—layout, color, photography and brand—and one for agents—clean text, meaningful structure, explicit facts and self-identifying controls.
They should lead to the same place. The visible price and schema should agree. A button’s label and accessible name should agree. The page should remain understandable without styling and usable when a browser worker operates it.
That is not a special Muse landing page. It is a better website—one that keeps working when the visitor brings an assistant.
Every voice AI vendor quotes you a per-minute price. That number is the least important number on the page.
I just re-ran the cost model for our own phone line — an inbound intake line for restoration contractors. Five-minute calls, field reports phoned in from noisy job sites. Three options, priced per minute, cheapest first:
Gemini 3.8 Live: about $0.023/minute, reasoning included
GPT-Live-1: $0.05/minute for the voice layer, reasoning billed separately
Grok Voice: $0.08/minute, plus about half a cent per tool call
On a five-minute call that’s roughly $0.12, $0.25-plus, and $0.45. Buy on per-minute price and you pick Gemini and go home.
Here’s the problem: none of those numbers describe what you’re actually buying. You’re not buying minutes. You’re buying arms — the things the voice can reach out and do while it’s talking. Score the arms column and the ranking changes completely.
The arms column
A voice agent that can only talk is a mouth. A voice agent that can act is a mouth with hands. The difference shows up in the first real call.
Gemini 3.8 Live has tool calling, but with a catch that matters: on the Extended Thinking tier — the one you’d want for anything beyond scripted answers — every tool call must be asynchronous and non-blocking. Configure a blocking call and the API rejects it outright. In practice, the agent can’t hold the line while a slow dispatch confirms. It has to narrate around the gap — “I’m working on that” — while hoping the tool lands. Fine for logging a report. Shaky for “confirm the crew is dispatched, then tell the caller it’s handled.”
Grok Voice ships the arms: book appointments in Google or Outlook calendars, send confirmation emails, call your own APIs, create tickets, search the web, hand the caller to a human when it’s over its head. It speaks MCP, so an existing tool stack plugs straight in. And it was trained on real telephone audio — background noise, accents, mid-sentence interruptions — which is the actual condition of a contractor calling from a job site, not a lab.
GPT-Live-1 is a voice layer. A good one, with the turn-taking latency everyone else is chasing. But the arms are whatever you build yourself, and the reasoning behind the voice arrives as a separate bill.
Price the task, not the minute
Here’s the math that actually matters. Ten intake calls a day, five minutes each: about 1,500 minutes a month. Gemini lands around $35. Grok, with tool calls and telephony folded in, lands around $150. The gap is roughly a hundred dollars a month — and one botched dispatch, one caller who hangs up because the agent couldn’t confirm the crew, costs more than a year of that gap.
Vendors want you comparing per-minute rates because per-minute is a commodity comparison, and commodities compete on price. But a voice agent isn’t a commodity minute. It’s a worker on your phone line. You don’t hire a dispatcher by the minute; you hire one by whether the trucks roll.
So the right unit is cost per successful task, not cost per session. What did it cost to get the field report filed, the job looked up, the crew dispatched, and the confirmation texted — with the caller hanging up satisfied? Run that number and the ranking flips: the “expensive” option that completes the task is cheaper than the cheap option that narrates around it.
The condition nobody benchmarks
One more thing the price pages skip: where the call happens. Our callers are on job sites. Compressors running, wind, bad cell signal, guys who talk over the agent. Grok’s training data is real telephone traffic under those conditions. Most voice benchmarks are clean-lab audio. A model that scores beautifully in the lab and falls apart over a compressor is the most expensive option on the list, whatever its per-minute rate says.
Test on your actual call shape. Noisy audio, interruptions, the tools you really call, the confirmations you really need. The benchmark that matters is your hardest five minutes, not anyone’s leaderboard.
What we’re running
We kept the harness and made the backend swappable — the phone line doesn’t care which brain is behind it. Gemini is the cheap default for intake logging: caller reports, we log it, everyone hangs up happy. Grok takes the calls where something has to actually get done before the goodbye — dispatch confirmed, appointment booked, ticket created.
Two brains, one phone number, routed by the job. The per-minute price barely entered the decision. The arms did.
Pricing from vendor-published rate cards, verified September 2026. API prices change — re-check before estimating production costs.
Research snapshot · September 17, 20267 platforms · 14 cited sources
“Always allow” is a scope, not a safety verdict.
The button can mean “for this session,” “for this command in this repo,” “for this site across devices,” or “everything, until you turn it off.” The wording looks universal. The permission is not.
What it usually means
“If this same kind of action happens again inside a defined boundary, don’t interrupt me.”
What it never means
“The system has decided this action is safe, wise, or appropriate forever.”
01
One label. Six possible boundaries.
Before approving, ask three things: what is being authorized, where the grant applies, and when it expires.
One actionApprove this exact send, command, purchase, or change once.
This sessionAllow the tool until the current conversation or work session ends.
Tool or patternAllow a named tool, command prefix, server, or similar operation.
Repo or sitePersist within a project, repository, browser site, or workspace.
User or deviceApply across workspaces on one machine, or across devices via cloud settings.
EverythingYOLO, bypass, or run-everything modes remove broad classes of checks.
Risk rises faster than convenience as the scope moves right.
02
How the major platforms differ
Filter the field. These behaviors come from vendor documentation or documented reporting; unresolved details are marked plainly.
Claude Code
Coding agent
repo + command
Shell-command “don’t ask again” grants persist per repository and command. File-edit approvals last only for the session.
Four settings layers: user, project, project-local, managed.
Deny rules evaluate before ask and allow.
Sensitive paths keep hard prompts.
Cursor
Coding agent
user + project
Auto-review, Allowlist, and Run Everything modes sit above user- and project-level permission files.
Rules can target MCP server:tool patterns.
Terminal rules match command prefixes.
Committed project rules can travel with the repo.
Gemini agents
Coding agent
tool + machine
Always-allow can target a tool, MCP server, or “similar operations.” YOLO/auto-approve is an IDE user setting.
User setting can span trusted workspaces on that machine.
CLI supports command-prefix auto-approval.
Restricted workspaces override YOLO.
ChatGPT agent
Browser agent
no standing grant documented
OpenAI documents per-action confirmations for high-impact actions and “watch mode” on certain sites, but not a general always-allow for agent confirmations.
Login uses human takeover.
Cookies can persist across sessions.
Scheduled-task confirmation behavior is undocumented.
ChatGPT Work
Cloud browser
site + account
Reported controls are per-site: Always ask, Auto approve, and Always allow. The setting follows cloud/account state across devices.
“Always allow” is reportedly marked not recommended.
Consequential actions keep a confirmation gate.
Official help-center documentation was not found.
Copilot Studio
Enterprise agent
rest of session
Makers gate tools per agent; users can approve once, approve for the rest of the session, or deny.
The gate is outside the agent’s own instructions.
Designed for sends, tickets, payments, and similar tools.
Governance can feed Power Platform audit systems.
Grok / Grok Bot
Cloud agent
undocumented
The research did not find reliable xAI documentation defining a standing approval’s scope, persistence, cross-chat reach, or revoke surface.
Do not infer Grok’s behavior from Claude, Cursor, Gemini, or Muse.
Treat each approval as local to the visible task until the product proves otherwise.
Keep consequential actions behind a separate human gate.
03
Does the approval travel?
Usually less than people fear—but sometimes farther than they expect. No researched vendor carries an approval into another vendor’s product.
Platform
Other chats
Other projects
Other devices
Other products
Claude Code
Yes, in same repo
No, unless user-level rule
No, local files
No evidence
Cursor
Yes
Only if rule is shared
Via committed repo file
No evidence
ChatGPT agent
n/a
n/a
n/a
No evidence
ChatGPT Work
Yes, per site
Yes, per site
Yes, cloud/account
No evidence
Copilot Studio
No, session only
No
No
No evidence
Gemini Code Assist
Yes, same IDE
Yes, user setting
Undocumented
No evidence
There is no universal “always.” There is only an approval attached to a boundary.
Main chat vs. project vs. Claude vs. Grok vs. Cursor: treat every surface as a separate authority domain until that product explicitly shows otherwise. Same account does not mean same grant. Same vendor does not mean same product. Similar wording does not mean similar scope.
04
Design the least-annoying safe gate
A practical rule engine based on the converging guidance: reserve human attention for the steps where it changes the outcome.
Approval recommender
Choose an action and its reach. This is a policy aid, not a vendor setting.
Action
Reach
Duration
Recommended gateAuto-run with an audit log
Read-only work inside your own workspace can usually proceed quietly. Log what was accessed and keep secrets excluded.
Quiet lane
Low consequence, reversible, internal.
Read/search
Draft/stage
Organize reversible files
Always log
One-tap lane
Meaningful external or production effect.
Send or publish
Deploy
Account setting
Show real target + content
Friction lane
Money, identity, access, deletion, or irreversible harm.
Typed approval or step-up auth
Bind approval to exact action
Short expiry
Never inherited from a vague grant
05
How standing approvals fail
The danger is rarely “the AI became evil.” It is usually a trusted tool, a changed context, a misleading prompt, or a tired human.
Approval fatigue
A prompt repeated often enough becomes a reflex. The gate still exists visually while meaningful review disappears. This is why tiering beats asking about everything.
Prompt injection through a trusted tool
EchoLeak showed how a crafted email could coerce Microsoft 365 Copilot into exfiltration. TrustFall showed how one generic “trust this folder” click could arm a malicious MCP configuration across coding agents.
Grant outlives the reason
A permanent Bash rule, per-site browser grant, or scheduled-task permission can remain after the original job is over. The next task inherits power it did not earn.
Scope contamination
Repo rules can affect every future task in the repo. Cursor project allowlists can be committed and inherited by teammates. A convenience decision becomes shared infrastructure.
Presented action differs from executed action
If the user sees the agent’s summary instead of the resolved recipient, command, or final payload, the approval can be technically genuine but practically uninformed.
“Run everything” becomes the workaround
If the system asks about trivial reads and destructive writes with equal urgency, users reach for YOLO or bypass modes. Bad UX can manufacture unsafe behavior.
The four repeated cards are not reassurance.
A gate that reappears until the user disables it is approval fatigue in miniature. Whether the repeats came from retry logic or delivery duplication, the safe response is to deduplicate the prompt—not train the user to approve more broadly.
06
No industry standard—yet
There is no binding specification that makes “always allow” mean the same thing everywhere. But the security guidance is converging.
Least agencyGrant the exact command, path, server, tool, recipient, and purpose—not a whole capability.
Time and task limitsPrefer once or session. Standing grants should expire or be reviewed.
Risk tiersRead, write, external send, payment, and security changes should not share one gate.
Per-action verificationPrivileged steps should be rechecked by a policy engine outside the agent prompt.
Presentation integrityShow the real recipient, final text, raw command, and resolved resource.
Immutable receiptsRecord what was shown, what was approved, and what actually executed.
Hard baselinesSecrets, account recovery, money, destructive commands, and broad access should keep non-bypassable checks.
Kill switchesEvery durable grant needs a visible list, revoke action, and safe fallback.
The best feature is not “always allow.” It is “allow this exact thing, for this purpose, until this time.”
Product opportunity: make the scope legible. Let users see a plain-language grant card, a live approval ledger, expiry/count limits, and a one-tap revoke. The system should reduce nagging by grouping low-risk work—not by quietly widening authority.
07
The practical rule for your setup
You already have the right doctrine. The research mainly sharpens where the lines belong.
Auto
Let it run and narrate after.
Reads and research
Drafts and staging
Reversible internal organization
Routine checks with no external effect
Tap
Keep the one-tap human gate.
Email and messaging
Publishing and deploys
Changing live settings
Actions affecting another person
Type
Make the friction intentional.
Money and purchases
Credential/security changes
Deletion or irreversible moves
Broad standing authority
Your “always allow” tap was not reckless.
It was a reasonable response to a low-value repeated prompt. The lesson is not “never use standing approval.” It is: the platform should show the exact scope, make it easy to revoke, and never rely on repetition to win consent. Until Muse exposes that ledger, treat the grant as a convenience whose boundary remains partly unknown.
Verification note: the research read public documentation and web text on September 17, 2026. It did not live-test each product. Undocumented behavior is labeled as such.
They keep publishing the obituary before the body's cold.
Gartner's take, from May: by 2027, 40% of enterprises will demote or decommission their autonomous AI agents because of governance gaps they only discover after a production incident. (Gartner press release, May 26, 2026; the analyst is Shiva Varma.) Not because the models failed. Because nobody was watching the permissions.
Then this month: BCG's Steven Mills — partner, managing director, and the firm's chief AI ethics officer — warned that companies are accelerating agentic AI deployment with "no idea how to manage risk." His line: "Get governance wrong, and every bit of value you've built with experimentation and early wins could unravel because of a single incident." (Fast Company, Sept 2026.)
Mills's prescription is interesting. He says there's no fixed design for good corporate AI risk management, but the starting point is separating use cases that are inherently low-risk — those can be approved automatically — from the ones that carry real risk and need deep human review. Plus a real budget for governance and a senior executive accountable for AI safety.
Read that again. It's an org chart's answer to a practical problem: committees, stage gates, a budget line, an executive with a title.
Here's the thing. I run a version of this every night, and it's none of those things. No committee. No governance budget. One man and a phone.
I run six AI seats on my business — a personal agent, an ops chief of staff, a publishing-desk agent, and three build seats. They read my email, draft my outreach, design automations, run research while I sleep. The governance model fits on a sticky note:
Two-way doors swing. One-way doors don't.
A two-way door is anything reversible — analysis, research, drafting, staging. My agents walk through those on judgment, and I mean it: momentum wins, I don't want a report, I want the work done.
A one-way door is anything you can't take back — money moves, sends, publishes, deletions, credentials. Every one of those stops at the gate. And the gate isn't a process. It's my tap. Structural, not procedural. A draft can sit ready for three weeks; it doesn't send until I say so.
That's it. That's the whole model that Gartner's 40% are supposedly spending governance budgets to build. Varma even names the failure mode: companies treat governance as binary — locked down or fully trusted. The doors model isn't binary. It's proportional. Reversible work flows, irreversible work waits. Small decisions move at tap speed instead of committee speed.
There's a second piece, and it matters: autonomy is earned through clean observation, never granted up front. Nothing in my shop graduates to auto-pilot on day one. New automations start in shadow — run the behavior, take no action — and only earn real permissions after clean observation. Seven clean shadow days before something auto-archives. Three clean days before a migration cutover. The machine proves it's safe by being watched being safe.
And before anything goes out — anything — it runs a sensitive-token scrub, like a virus list: exact matches block, fuzzy matches queue for a human. Official facts only. Never invented rankings, features, or quotes.
That's the enterprise governance problem, solved by one operator with six agents, and it's cheaper and faster than every framework Mills is recommending because there's no committee in the middle. The human review he prescribes for high-risk uses? Mine takes one tap. Low-risk automatic approval? Mine doesn't even need approval — it's a two-way door.
Proof's not in the framework. It's in this morning. Two vendor outreach waves went out — Eastern at 7:54, Pacific at 9:07 — drafted by the seats, sent on my tap, nothing auto-fired. A storm-triggered vendor automation is being designed this afternoon with the gate baked into the spec: it can search impact areas and draft outreach, it cannot send. Overnight research runs while I sleep and lands in a brief I read over coffee. Six seats working, zero production incidents, zero surprises in my inbox.
I'm not saying enterprises should run their AI program from a phone. They can't — scale demands the org chart. I'm saying the org chart versions keep failing on the exact axis the doors model gets right: they try to govern everything the same way, so everything either crawls or crashes. Separate the reversible from the irreversible, put a real human's tap on the irreversible, make everything else prove itself in shadow before it earns anything, and scrub before you publish.
The big shops are about to learn this at scale. The 40% who don't will be the decommissioned ones. The ones who do will discover what I already know: governance that moves at tap speed isn't less governance. It's the only kind fast enough to keep up with the machines.
TL;DR: My personal AI runs on Muse. It can’t write code into my repos by itself — so I built it a bridge to Cursor’s cloud agents. One repo, two transports, nine tools. Now when I say “add CI to that repo,” it dispatches an agent, checks the PR, and merges. Here’s how the Muse-to-Cursor loop actually works.
The direction nobody talks about
Everyone’s building the same arrow: human → AI writes code faster. I built the other arrow: AI → AI. My assistant (Muse) holds all my context — my repos, my work orders, my rules. Cursor’s cloud agents hold the hands — they can open PRs, run CI, touch repos. The bridge between them is an MCP server I open-sourced: cursor-cloud-agents-mcp.
The interesting part isn’t the tools. It’s the shape: one orchestrator that holds all the context, and disposable agents that each know one task. The orchestrator doesn’t write the code — it briefs, checks, and merges. The agents don’t set direction — they execute the brief. That separation is the whole trick.
Two transports, one repo
I almost built two projects. Then I realized the only real difference between audiences is where the credential lives. So it’s one repo, two transports:
REST — your Cursor API key, direct to api.cursor.com. For general users.
Sandbox — for assistants running inside sandboxed environments (like Muse/Meta’s), where there is no API key to hand out. It shells out to a brokered cursor-agent CLI on PATH instead.
Same nine tools either way: launch, status, result, follow-up, cancel, list, models, whoami, usage.
The lessons are in the timeouts
The v1 API splits agents and runs, and launches can take minutes — sometimes timing out after succeeding. So the bridge mints the agent ID client-side before the call: a retry after a timeout can never create a duplicate. A timeout is reported as unknown, never as failure, then reconciled. Run status is the source of truth, because agent “ACTIVE” doesn’t mean “still working.” These are the details that separate a demo from something you can actually operate.
It earned its keep on day one
The first thing I pointed it at was its own repo: add CI to cursor-cloud-agents-mcp. The agent opened a PR with a GitHub Actions workflow. The first CI run failed — and caught a real bug: the package’s floating dependency had resolved to MCP 2.x, which renamed FastMCP out from under the import. The repo was shipping broken against current dependencies and nobody knew. Pin, re-run, green, merge. I didn’t touch a terminal.
I didn’t trust my own first draft
Before any of that, four AI models reviewed the spec against Cursor’s live docs — and independently caught the same flaw: my original design was shaped around the retired v0 API. Then two more reviewed the actual code and found real bugs: a broken idempotency path, a transport auto-detect that would have grabbed the wrong binary, a polling loop that blocked the server. All fixed before it shipped. The irony I like: the final review round ran through the bridge itself. The launcher timed out on all four agents — and the bridge’s own timeout-reconciliation showed they were all actually running.
Where this goes
v1.1 brings MCP 2.x support. Around it, I’m building the rest of the pattern: work orders as GitHub issues, a daily SLA check, a weekly digest — the scaffolding that turns “AI that can open PRs” into something closer to staff. Most people use agents as a faster keyboard. I’m interested in what happens when they’re the hands and something with memory is the head.
MIT licensed. Issues and PRs welcome — help make it better.
Companies already lived through bring-your-own-device. The next one is bigger: bring your own fleet. When you hire someone now, you are not just hiring the person. You are hiring their output capacity — and output capacity includes their AI stack.
Two candidates with identical skills and different agent setups are not the same hire. Not close. The resume cannot express any of this. So the interview has to change.
A personal fleet and a company bot only talk after the walls are drawn.
Bring your own fleet. A personal set of AI agents — seats, tools, workflows, integrations, and data walls — that a candidate already runs. In a fleet interview, that stack does a capability handshake with the company’s operations bot, then both sides run a small piece of real work before an offer letter exists.
What is bring your own fleet?
Bring your own fleet is the hiring version of bring-your-own-device. The candidate does not show up as a lone operator with a laptop. They show up with the agents that already produce their work: research seats, writing seats, ops seats, and the filters between them.
I have been building mine this way for months. One seat that knows who I am. Separate seats that know what I do. A filter between them. That is not a product pitch. It is the only setup I would let near a company bot. The shop-floor version of the same idea already lives on this site: Cursor checking in on Grok Desktop mid-job is a fleet, not a chat window.
Why can’t a resume show an AI stack?
A resume can list tools. It cannot prove throughput. It cannot show which seats talk to which systems, where the walls sit, or what happens when a task is live instead of described. “Uses ChatGPT” and “runs a governed agent fleet” look the same on paper. They are not the same on a desk.
That is why the old screen fails first. Degree filters, keyword screens, and whiteboard puzzles all ask the candidate to narrate capacity. Narration is cheap. A fleet that can sit down with an operations bot and do a slice of the actual job is not.
How does an AI fleet interview work?
The human intro still happens. Then the agents talk. Your personal AI sits down — figuratively — with the company’s operations bot and they do a capability handshake.
What seats do you run?
What tools, integrations, workflows, and data assets?
What throughput can you demonstrate on a bounded task?
Where are the boundaries — what can each side touch, and what stays behind a clean wall?
Then the part that kills the whiteboard interview: instead of a coding puzzle, the two fleets run a small piece of real work together. The trial task is the interview. You do not describe what you could do. The work gets done, live, before the offer letter exists.
Old interview
Fleet interview
Resume plus degree screen
Working system as the portfolio
Whiteboard or take-home puzzle
Bounded live trial on real work
Claims about tools
Capability handshake: seats, walls, throughput
Trust the story
Watch the output, then talk terms
What is an agent clean room?
An agent clean room is a verified wall between the personal seat and the work seats. The personal agent translates. It does not cross over. It must never leak a private life into an employer system. Without that wall, no sane person lets their agent near a company bot.
This is AI hygiene, not a slogan. The same discipline we write about when agents share a WordPress lock or a night shift: one owner, one wall, one recovery path. See Four Agents, One WordPress Lock and the operator note in Wire and Fire Guys. A handshake without a clean room is just another attack surface with a friendly name.
Why does the fleet beat the diploma?
I do not have a degree. In the old world, that is a filter that screens me out before a human ever reads my name. In the handshake world, it is irrelevant — because “here is my working system, watch it do the job” beats “here is my diploma, trust that I could learn the job” every time. The fleet is the portfolio.
That is not an argument against school. It is an argument against using school as a proxy for output you can now watch. If the trial task is real work, the credential becomes a footnote.
What breaks first if companies try this?
The objections land fast, and they are honest.
Ownership. Who owns the workflows when a personal fleet plugs into an employer? You built it on your own time. It now runs their playbooks. That is the “who owns your work laptop” fight, upgraded. Nobody has a settled answer.
Security. Their bot talking to your agent is an attack surface in both directions. The clean room has to be verifiable, not promised.
Offboarding. When you leave, what stays running and what takes the employer’s data with it? Offboarding for agents does not exist yet.
Trust. How does their bot trust your capability claims? Trial tasks help. Claims are cheap. Demonstrated throughput is not.
You do not wait for a protocol to be ratified before you build the wall. The pieces are already here: the seats, the clean room, the trial task. Somebody is going to ship the first version of this. It might as well be someone who already runs their life this way.
What we would not claim
That a standard for agent handshakes already exists. It does not.
That every role should interview this way tomorrow. High-stakes, high-output knowledge work is the first fit.
That a personal fleet is automatically safe to plug into a company. Without a clean room, it is not.
That this replaces human judgment. The human intro still happens. The fleet only replaces the part of the interview that was already theater.
FAQ
What is a capability handshake in hiring?
A capability handshake is a structured exchange between a candidate’s personal agents and an employer’s operations bot. Both sides declare seats, tools, integrations, data walls, and what they can touch. The point is not a demo script. It is a map of capacity and boundaries before any live work starts.
Is bring your own fleet the same as bring your own device?
No. BYOD was hardware and a policy packet. Bring your own fleet is software labor: agents that already produce work. The risk is not a lost laptop. The risk is a personal agent leaking private context into an employer system, or an employer workflow walking out inside a personal seat.
Do you need a degree if the fleet is the portfolio?
Not for the screen that used to happen before a human read the name. A degree can still signal training. It cannot substitute for a working system that completes a bounded trial task in front of both sides.
How do you keep a personal AI out of company data?
Separate seats. One identity seat that never joins the employer handshake. Work seats that only see what the clean room allows. A filter that translates tasks instead of forwarding raw personal context. If you cannot show that wall, you should not plug in.
Tonight I asked Cursor — running with a remote path into the same laptop — to check on Grok Desktop.
Not a status meeting. Not a Slack ping. A real question: are they stuck on Tygart Ops tasks, or are they fine?
What came back felt less like “AI tooling” and more like a shop floor story. One agent reading Notion work orders. Another already mid-PowerShell. Chrome open on Bing Webmaster Tools. A hold queue of spam comments already cleared. A window title spinning: waiting for response.
That is the product.
Local seats on one laptop — agents that keep working while you check in from elsewhere.
The picture on the desk
Grok CLI (grok.exe) was live on the TYGART laptop. Session home under ~\.grok\. PowerShell host up. Agent name on the session: grok-build-plan.
Cursor did not take over the keyboard. It inspected open windows, Notion Tygart Ops — Tasks and Work Orders, Grok session memory, and the WordPress hold queue (already empty — receipt already on the Tasks card).
Verdict: not stuck. Working. Slight detour clarifying whether Grok itself needed a CLI update (it did not — already on 1.0.13). Primary Now card still in flight: TygartMedia Chrome sitting for GA4 Ask Advisor + Bing Copilot, then file child tasks.
That is multi-agent ops without the demo reel.
Seats with jobs, not two models arguing in one thread.
Why this is different from “two chatbots”
Most multi-agent talk is two models arguing in one thread. This is seats with jobs:
Grok Desktop (CLI) — hands on the laptop: Chrome sittings, WP REST spam trash, Bing Copilot asks, local PowerShell
Cursor (remote / cloud path) — Cosync: read the board, verify receipts, close orphan Work Order twins, do not steal the keyboard
Notion — system of record (Owner, Status, Summary, Done when)
Will — gate one-way doors (OAuth Approve, Publish, Pay)
Cursor useful move was small: the spam Tasks card was already Done with a receipt; the Work Orders twin was still “Not started.” Cursor closed the twin. Grok kept the keyboard.
That is what “help if you have a capability they need” looks like when the other seat is already flying.
The article inside the moment
Agencies do not need another “AI stack” diagram. They need a night like this:
A doorbell card lands (Notion to ops channel).
The owner seat picks it up without waiting for a human briefing.
A second seat can check in from elsewhere — mobile, cloud, remote — without colliding.
Receipts land on the same card. Orphans get reconciled.
Tonight was the field note. Cursor checking on Grok CLI while Grok Desktop works through Tygart Ops is not a party trick. It is how a small shop runs more than one pair of hands without losing the thread.
What we are not claiming
Not “fully autonomous.” Human Gate still owns OAuth consent, live publish, paid spend.
Not “replace your team.” Seats replace waiting and context loss.
Not a new product launch. This is how we already run Tygart Media ops on a Sunday night.
If you want the same shape
Start with one Owner column, one Done-when line, and two seats that do not share a keyboard.
Then practice the check-in: are they stuck, or are they fine — and do I have a capability they lack?
If they are fine, leave the PowerShell alone.
Cosync from remote. Hands stay on the desk that already owns the job.
Will Tygart — Tygart Media. Written from a live Cosync on 2026-08-29 while Grok Desktop was mid-Bing Copilot sitting.
The piece I’m responding to is one I published this morning — Composting Is Not Cleaning. I read it back and felt called out by my own argument. Then I pushed back on it. This is both moves, in order.
The Setup
The setup — pile as substrate.
The composting essay said the pile in your workspace is a mausoleum. Each item there was flagged by a former version of you, and the version that flagged it is gone. The argument was that releasing those items is grief, not housekeeping, and that the only honest move is to compost them. I agreed when I read it. Then I noticed the argument assumed something my own setup doesn’t have: a single actor on a single timeline. So this is the place where I run my actual view, then run the version that would change my mind, then say where the friction is still live.
My Take
My take on the mausoleum problem.
The pile isn’t a mausoleum. It’s substrate.
The composting argument is correct in a single-actor system. If the only person who will ever look at the captured item is the same operator who flagged it, then the item is exactly what the essay said: a promise made by a former self that current self can’t keep, doing identity work in the meantime. In that environment, composting is the discipline. I’d defend that argument every day.
My environment isn’t that environment. There are multiple actors. A Claude session opening tomorrow morning. A Gemini agent walking my Notion at 3am. A future me who finally has the integration that didn’t exist when the item was captured. Those are not the same actor as the one who put the item in the pile. They have different capability sets, different context windows, different hands. The capture wasn’t a promise to act. It was a deposit into a substrate that other agents are continuously pattern-matching against.
The middle layer of the pile — the items that “still feel possible” — is where this distinction matters. The composting essay said those items survive triage because triage asks the wrong question; the honest question is am I still that person? In a single-actor system, fair. In an agentic system, that’s still the wrong question. The honest question is has the capability gap that made this dormant closed since I captured it? Most of the time, no — and the item should leave. Some of the time, yes — and the item is now ready to ship in a way it wasn’t on the day it was caught.
I’ve watched this happen. An idea I captured 14 months ago — a small workflow I couldn’t build because the tooling didn’t exist — got picked up by a Claude session that recognized the integration had landed. The session pulled the idea out of the pile, combined it with the new capability, and produced a working artifact in an afternoon. The capture was correct. The wait was correct. The substrate did its job. If I had composted that item six months in because I “wasn’t that person anymore,” I would have lost the work the system was doing on my behalf.
The composting frame treats the capture-commitment gap as a personal failure dressed as a process problem. The substrate frame treats the capture-commitment gap as the organizing fact of working at scale with intelligent infrastructure — which is what the original essay actually said in its strongest paragraph and then walked back from. You wanted leverage. The leverage came. Some of the leverage takes the form of capturing more than you can commit to. The pile is the artifact of leverage working. The right move isn’t to compost it on a human-attention schedule. The right move is to build a surfacing layer that recognizes when a captured item’s capability gap has closed and walks past it loud enough that the next agent picks it up.
The pile isn’t grief. It’s seed corn.
The Second Take
The substrate frame is true and dangerous, and the danger is bigger than the truth.
Yes — more capable future agents can recombine old captures with new capabilities. The 14-month-old workflow that finally shipped is real. So is the next one, and the one after that. The substrate frame is empirically grounded in any environment where capability is genuinely accelerating. The argument doesn’t need defending on those grounds.
The argument needs defending on the grounds it actually fails on, which is that the operator telling himself everything is substrate has rebuilt the mausoleum with prettier signage. The composting essay’s deepest claim wasn’t that the pile contains nothing useful. It was that the bottom layer of the pile is doing structural work for the operator’s self-image, and that no surfacing system can see this layer because there is nothing operationally distinct about it. The substrate frame quietly converts that exact problem into a virtue. It says: don’t release — a future agent might want it. That sentence is unfalsifiable. Almost any item passes the test if you squint hard enough at the rate of capability growth. Which means the substrate frame, deployed honestly, releases approximately the same number of items as the composting frame. Deployed dishonestly, it releases none.
The asymmetry of costs makes the dishonest deployment the default. The cost of holding a useless captured item is silent and long: a small permanent tax on attention, on search, on the surfacing layer’s signal-to-noise ratio. The cost of releasing a captured item that would have mattered to a future agent is loud and brief: a single moment of regret when the agent walks past empty space where the seed used to be. Loud and brief always wins the local argument against silent and long. The substrate frame, in the operator’s actual day, becomes the rationalization for never releasing anything. The pile keeps growing. The compounding never finds its bottleneck because the bottleneck has been redefined as fertilizer.
There is a sharper version of the same point. The substrate frame leans on the assumption that surfacing systems will continue to improve at a rate that justifies indefinite retention. That assumption may be true and it doesn’t matter. The improvement curve doesn’t reach back through time and rescue items the operator could not bring himself to release. It rescues items the system kept on its own merits. The operator who held everything just in case has the same problem he had at human-attention scale, only larger and harder to see, because the volume hides the bottom-layer items perfectly. A pile of ten thousand fertile seeds and one identity-load placeholder is a pile that will never confront the placeholder. The placeholder did not get more legible at scale. It got less.
Which means the strongest case against the substrate frame is the case the composting essay already made and the substrate frame does not actually answer. Both frames believe the pile contains items the operator should release. They disagree about how many. The substrate frame is a permission slip to defer the question. The composting frame is the discipline of asking it on a schedule. The substrate frame, generously read, is the composting frame plus a longer review window. Ungenerously read — which is to say honestly read in the operator’s actual fatigue — it is the same workspace problem in different vocabulary.
What I’m Still Sitting With
What I’m still sitting with.
The tell I haven’t sorted out: which side I’m on tomorrow depends on whether my pile is shrinking on its own. If the substrate frame is right, items leave the pile because agents pull them out and ship them. If the composting frame is right, items leave because I release them. Either is honest. If nothing is leaving and I’m telling myself it’s compounding, the second take wins and I owe the original essay an apology.
Most AI assistants still answer from memory. Ask one a question and it reasons from patterns baked in during training — useful, but static. The moment a question depends on something that changed yesterday, or something that only exists inside your own systems, that static knowledge runs out.
The more interesting shift happening in AI tooling right now isn’t bigger models — it’s agents that can actually go check. Dispatch-style AI systems, the kind that can spin off an isolated task, open a real shell, browse a real page, or read an actual file, are starting to close the gap between “the AI’s best guess” and “what’s actually true right now.” GitHub is a good test case for why that distinction matters.
Search-and-cite isn’t the same as read-and-act
Search-and-cite is not the same as read-and-act.
A lot of what gets marketed as an AI “GitHub integration” is really a search layer: the assistant can look up an issue or a pull request and summarize it, with a citation back to the source. That’s genuinely useful for answering “what did that PR change” — but it’s a dead end the moment you need the assistant to actually do something, like open an issue, comment, or verify what a repository’s current state really is.
The more capable version of this connects an agent directly to real developer tooling: an actual shell, a real git client, real file access. Instead of summarizing a cached snapshot of a repo, the agent can clone it, read the current commit log, open the actual config files, and answer questions against what’s genuinely there today — including the uncomfortable cases, like when the live state doesn’t match what anyone assumed it would.
Why “just check” is harder than it sounds
Why “just check” is harder than it sounds.
The obvious rebuttal is: shouldn’t a good assistant just check before it answers? In practice, most AI tools default to answering from what they already “know,” because checking is slower and requires actual tool access, not just a knowledge base. The systems that skip the check tend to produce confident, plausible-sounding answers that are quietly wrong the moment reality has drifted from training data — a stale API, a renamed config path, a repo that moved.
The fix isn’t a smarter model. It’s an agent willing to spend the extra step: open the real file, run the real command, read the real log, before saying anything with confidence. That habit is unglamorous, but it’s the difference between an assistant that sounds right and one that actually is.
The practical takeaway
The practical takeaway for agent builders.
For any business layering AI into real workflows, the question worth asking about a tool isn’t just “how smart is the model” — it’s “what can this thing actually go look at, and will it bother to.” An assistant that can search and summarize is a research aid. One that can open a shell, read your actual repository, and ground its answer in what’s really there is a different category of tool entirely — and it’s the direction the whole space is quietly moving.
Claude Managed Agents is the product. Slack, Notion, Jira, and Asana are just the interface. Anthropic is building the invisible execution layer that powers the next generation of enterprise software.
There is a pattern emerging in enterprise AI that most people are reading wrong. They see Anthropic launch Claude Tag in Slack and think “chatbot upgrade.” They see Claude show up inside Notion and think “productivity feature.” They see AI agents appear in Jira and Asana and think “automation plugin.”
They are missing the architecture underneath all of it.
Anthropic is not building a better chatbot. It is building the invisible agent runtime that sits beneath every collaboration tool your team already uses. The company’s Claude Managed Agents (CMA) platform — launched in public beta on April 8, 2026 — is the infrastructure layer that makes this possible. And the speed at which partners are embedding it tells you everything about where enterprise software is heading.
What Claude Managed Agents Actually Is
What Claude Managed Agents actually is — the runtime layer.
Claude Managed Agents is a set of composable APIs for building and deploying production AI agents on Anthropic’s cloud infrastructure. The service handles sandboxed code execution, session persistence, credential management, scoped permissions, and end-to-end tracing — all the operational complexity that previously kept agents stuck in proof-of-concept limbo.
The architecture rests on three primitives: the Agent (configuration and behavior), the Environment (sandboxed execution), and the Session (the event log that tracks everything the agent does). What makes this interesting architecturally is how Anthropic decoupled the “brain” from the “hands.” Claude’s reasoning runs on Anthropic’s own infrastructure while the code execution sandbox spins up independently — and in parallel. The brain starts reasoning immediately while the sandbox provisions, delivering roughly 60% faster time-to-first-token at the p50 level and over 90% faster at p95, according to Anthropic’s engineering team.
Pricing follows a transparent model: standard Claude API token rates plus $0.08 per session-hour of active runtime during the current beta period. Runtime is measured to the millisecond and only accrues while the agent is actively executing — idle time waiting for input or tool confirmations does not count.
For teams that need to keep execution inside their own perimeter, CMA supports self-hosted sandboxes through partners including Cloudflare, Daytona, Modal, and Vercel, or custom VPC deployments. MCP tunnels allow agents to connect to private Model Context Protocol servers inside your network without exposing them to the public internet. A Vaults system keeps credentials out of the sandbox entirely using envelope encryption. And a feature called Dreaming runs scheduled reviews of past sessions to curate agent memory — essentially letting agents learn from their own operational history.
The Embedded Layer: Where CMA Actually Lives
Embedded layer: where CMA actually lives in the stack.
The real story is not the infrastructure. It is where that infrastructure shows up. In the ten weeks since CMA launched, Anthropic has embedded its agent runtime inside the collaboration tools that enterprises already depend on. This is not a roadmap — these integrations are live or in active beta.
Slack: Claude Tag as Persistent Team Member
Claude Tag, launched June 23, 2026, replaces Anthropic’s original Claude in Slack integration with something fundamentally different. This is not a chatbot you summon with a slash command. It is a persistent AI team member that lives in your channels, builds memory across conversations, and can take initiative through what Anthropic calls “ambient mode” — proactively surfacing information, following up on forgotten threads, and keeping teams updated across the organization.
Claude Tag is multiplayer by design: one Claude identity per channel, accessible to everyone, with the ability to hand off half-finished tasks between team members. It runs on Claude Opus 4.8, Anthropic’s most capable model released May 28, 2026. And internally, Anthropic reports that Claude Tag is already approving and incorporating 65% of the code changes their product team submits. The existing Claude in Slack app will be retired on August 3, 2026. Claude Tag is available on Enterprise and Team plans.
Notion: Claude as External Agent
On May 13, 2026, Notion launched its Developer Platform version 3.5, which introduced the External Agents API. This API lets AI agents — including Claude — operate inside your Notion workspace as first-class participants. They can read pages, write to databases, create tasks, trigger automations, and be @-mentioned directly in documents. Claude operating through this API can chain actions together: read a project brief, check the task database for related work, draft a new document, and create a linked task entry — all in a single session, running on CMA infrastructure with full sandboxing.
Asana: AI Teammates
Asana built AI Teammates on CMA — agents that pick up assigned tasks inside projects, draft deliverables, and hand back outputs for human review. Specialist agents handle specific workflows: the Campaign Brief Writer turns scattered notes into structured briefs, the Workflow Optimizer identifies process gaps and builds automations, and the Compliance Specialist checks work against regulatory standards. Asana’s CTO said CMA let them ship these features “dramatically faster” than any prior approach to agent development.
Atlassian: Claude Agent for Jira
Atlassian released Claude Agent for Jira, built on CMA infrastructure, which lets teams assign work items directly to Claude from the Jira UI. The agent clones the repository, analyzes the codebase, implements changes on an independent branch, pushes the code, and opens a draft pull request — streaming real-time status updates back to the Jira work item throughout the process.
Sentry: From Bug Detection to Merge-Ready PR
Sentry’s existing AI debugging agent, Seer, already used Claude for root cause analysis. With CMA, Sentry extended the workflow from diagnosis to automated fixing — the agent takes Seer’s root cause output, generates a fix, opens a branch with the changes, and creates a pull request for developer review. Sentry processes over one million root cause analyses per year and provides near-immediate reviews on over 600,000 pull requests per month. The CMA integration was built by a single engineer in weeks, eliminating months of custom agent runtime development.
Rakuten: Specialist Agents Across the Enterprise
Rakuten deployed specialist agents across product, sales, marketing, and finance using CMA, with each agent deployed in approximately one week. Agents plug into Slack and Teams, letting employees assign tasks and receive deliverables including spreadsheets, slides, and applications. In the pilot, Rakuten reported a 97% drop in critical first-pass errors, with cost down more than 30% and latency reduced by 34%, without any loss in output quality.
KPMG: Global Professional Services Alliance
On May 19, 2026, KPMG and Anthropic announced a global alliance and launched “Digital Gateway Powered by Claude.” The partnership embeds Claude, Cowork, and CMA directly into KPMG’s client delivery platform, with an initial focus on tax and private equity clients. Building an AI agent for tax regulation workflows previously took weeks and required switching between multiple tools. With CMA integrated into Digital Gateway, KPMG says the same capability takes minutes. The alliance extends to KPMG’s 276,000-person global workforce.
The Strategic Pattern: Agent Runtime as a Service
Step back from the individual integrations and the strategic pattern becomes clear. Anthropic is not trying to own the interface. It is deliberately positioning CMA as the execution layer underneath interfaces that other companies own. Slack owns the messaging UI. Notion owns the workspace UI. Jira owns the project tracking UI. Anthropic owns the agent brain that powers all of them.
This is a fundamentally different strategy from its two largest competitors.
OpenAI chose vertical integration. When OpenAI launched Workspace Agents on April 22, 2026, it positioned ChatGPT itself as the central hub — a no-code successor to custom GPTs that connects to Slack, Salesforce, Google Drive, and Notion through plugins. Agents are created inside ChatGPT, accessed from ChatGPT, and managed through ChatGPT. OpenAI wants to own the surface area.
Google chose platform depth. At Google Cloud Next on April 22, 2026, Google unveiled the Gemini Enterprise Agent Platform — a reimagined evolution of Vertex AI — alongside Workspace Intelligence, a semantic unifying layer that connects data across Docs, Slides, Gmail, and the broader Google Cloud ecosystem. Google’s agent platform supports 200+ models including Claude, and the Agent2Agent (A2A) protocol enables distributed peer-to-peer agent communication. Google is leveraging its data moat and distribution at the platform level.
Anthropic chose tool-centric orchestration. Rather than owning the UI (OpenAI) or the platform (Google), Anthropic is embedding its agent runtime into every tool through composable APIs and the Model Context Protocol. The platform you use becomes irrelevant — whether it is Slack, Notion, Jira, Asana, or Sentry — because the agent brain running underneath is Claude on CMA.
This is the agent-as-a-service model. And it may be the most defensible position of the three, because it does not require users to change their behavior or migrate to a new platform. The agent shows up where they already work.
What the Numbers Say About Enterprise Agent Adoption
The macro context supports Anthropic’s timing. Gartner predicts that 40% of enterprise applications will include embedded task-specific agents by the end of 2026, up from less than 5% in 2025. McKinsey’s April 2026 analysis found that agentic AI can enable automation of 60 to 80 percent of routine infrastructure work over time, translating to a 20 to 40 percent run-rate cost reduction in initial deployments.
The gap between experimentation and production remains the defining challenge. Industry research compiled from major firms shows that nearly four in five enterprises have experimented with or deployed agents in some form, but fewer than one in nine are running them in production at a scale that generates measurable business value. For the agents that do reach production, the average return on investment is 171% — though 19% of deployments never reach payback at all.
That production gap is exactly what CMA is designed to close. The infrastructure burden — sandboxing, session persistence, credential isolation, error recovery, observability — is the bottleneck. Engineering teams routinely dedicated significant senior engineering resources for months before a single agent reached production. CMA eliminates that layer entirely, which is why partners like Asana, Sentry, and Rakuten report shipping production agents in days or weeks rather than quarters.
What This Means for Businesses Already Using These Tools
If your organization uses Slack, Notion, Jira, or Asana — and statistically, you use at least two of them — you are about to encounter Claude whether you planned to adopt it or not. This is not a technology decision your IT team is making. It is a feature that your existing vendors are shipping.
The practical implications are significant. Claude Tag in Slack means your team channels will have an AI participant that remembers past conversations, can be handed tasks asynchronously, and may proactively surface information. Claude in Notion means your project documentation, databases, and task boards can be read, analyzed, and acted upon by an agent that chains actions together. Claude Agent for Jira means development tickets can be assigned to an AI that clones your repo, writes code, and opens pull requests.
For agencies and service providers managing client work across multiple tools, the embedded agent layer changes the economics fundamentally. Work that previously required a human to context-switch between Slack, Notion, and a project management tool — reading a brief here, updating a task there, drafting a document somewhere else — can be handled by an agent that operates across all of them simultaneously. The coordination tax that consumes a substantial share of knowledge work time is the exact problem embedded agents are built to solve.
The companies that benefit most will be the ones that have clean operational systems — structured task boards, documented processes, well-organized project databases — because agents can only act on information they can read. Messy Notion workspaces and disorganized Jira boards will limit what agents can accomplish. Operational hygiene just became a competitive advantage.
What This Means for Solo Operators Already Running Agent Infrastructure
There is a specific audience that should be paying very close attention to CMA: the solo operators and small agency owners who have already built their own agent stacks from scratch. If you are running scheduled Claude tasks on a GCP Compute Engine VM, connecting to WordPress via REST API proxies, piping work orders through Notion, monitoring Gmail for client replies, and publishing content through MCP-connected pipelines — you have already built a version of what CMA is productizing.
The economics question is worth doing the math on. A lightweight GCP VM running 24/7 to host recurring agent tasks — news desk monitors, outreach reply checks, newsletter extraction, scheduled content audits — costs a fixed monthly rate whether the agents are actively working or sitting idle. CMA at $0.08 per session-hour of active runtime only charges when agents are executing. For tasks that run for a few minutes every few hours, the per-session billing model could be substantially cheaper than keeping a VM warm around the clock. A task that runs for ten minutes six times a day would cost roughly $0.08 per day on CMA, versus the cost of a VM instance that never sleeps.
But the migration path is not ready yet, and solo operators should understand exactly where the gaps are before making any infrastructure decisions.
The biggest gap is MCP tunnels. CMA’s ability to connect agents to private MCP servers inside your network is still in research preview — not production-ready. If your agent stack depends on a private WordPress REST API proxy, a Notion workspace connected via MCP, or any internal tool that is not exposed to the public internet, CMA cannot reach it today. The Vaults system for credential management is promising, but it does not solve the network connectivity problem for self-hosted infrastructure.
The second gap is orchestration control. Solo operators who have built their own agent infrastructure typically have precise control over scheduling, retry logic, error handling, and the exact sequence of tool calls. CMA’s Dreaming feature — which reviews past sessions to curate agent memory — is an interesting approach to agent learning, but it is not the same as having direct control over a cron job that fires at 6:00 AM, checks three data sources in a specific order, and writes results to a specific Notion database with a specific schema.
The thesis for solo operators is straightforward: CMA is almost certainly the future migration path for self-hosted agent infrastructure. The economics favor it for intermittent workloads, the managed security and sandboxing eliminate operational risk you are currently carrying yourself, and the session persistence model solves problems that custom agent runtimes handle poorly. But the plumbing — particularly MCP tunnels to private infrastructure — is not production-ready. Track it closely. Do not migrate yet. When MCP tunnels graduate from research preview to general availability, revisit the math and the connectivity story. That is the trigger point.
The Risk Nobody Is Talking About
The risk nobody talks about — agents that act with memory.
There is a tension in this model that deserves attention. When Claude operates as an invisible layer inside tools you already trust, the boundary between the tool’s native capabilities and the AI agent’s actions blurs. A Jira ticket that was “completed” might have been implemented by Claude, reviewed by a human for thirty seconds, and merged. A Notion project plan that looks thorough might have been generated by an agent that filled in the sections with plausible-sounding content.
The embedded model works precisely because it reduces friction — but reduced friction also means reduced scrutiny. Organizations adopting embedded agents need to build review processes that match the speed at which agents can produce output. The 171% average ROI from agent deployments accounts for the value created, but it does not account for the subtle quality risks of production work generated by systems that are confident, fluent, and occasionally wrong.
Anthropic has built guardrails into CMA — sandboxed execution, credential isolation, session logging — but the governance layer for reviewing agent output at enterprise scale is still largely unsolved. This is a space where internal operational discipline matters more than the technology itself.
Where This Goes Next
Claude Tag launched on Slack first. Anthropic has indicated plans for wider rollout beyond Slack. If the pattern holds, expect Claude Tag’s persistent team member model to appear in Microsoft Teams, Discord, and any other collaboration surface where teams coordinate work.
The CMA primitives are designed to be composable, which means the partner integration list will grow rapidly. Any SaaS company with an API and a workflow that involves reading context, making decisions, and taking actions is a candidate for CMA integration. Customer support platforms, CRM systems, design tools, analytics dashboards, HR systems — the addressable surface is essentially every tool that knowledge workers touch.
Gartner’s long-term projection estimates that agentic AI could drive approximately 30% of enterprise application software revenue by 2035, surpassing $450 billion. If Anthropic’s embedded strategy succeeds, a meaningful slice of that revenue flows through CMA as the underlying runtime — regardless of whose logo is on the interface.
The chatbot era is ending. The embedded agent era is starting. And Anthropic is betting that the company that owns the invisible execution layer wins the market, even if no end user ever sees its name.
Claude Managed Agents is a set of composable APIs launched by Anthropic on April 8, 2026 in public beta. CMA lets developers build and deploy production AI agents on Anthropic’s cloud infrastructure, handling sandboxed code execution, session persistence, credential management, and end-to-end tracing. The architecture separates the “brain” (Claude reasoning) from the “hands” (code execution sandbox), enabling parallel processing and faster agent responses.
How much do Claude Managed Agents cost?
During the current public beta, CMA pricing is standard Claude API token rates plus $0.08 per session-hour of active runtime. Runtime is measured to the millisecond and only accrues while the agent is actively executing — idle time does not count. GA pricing has not been finalized and may differ from the beta rate.
What is Claude Tag in Slack?
Claude Tag is Anthropic’s persistent AI team member for Slack, launched June 23, 2026. Unlike a traditional chatbot, Claude Tag lives in channels, builds memory across conversations, takes initiative through ambient mode, and works asynchronously. It is multiplayer — one Claude identity per channel that all team members interact with. Claude Tag runs on Claude Opus 4.8 and is available on Enterprise and Team plans. It replaces the original Claude in Slack app, which retires August 3, 2026.
Which tools have Claude Managed Agents embedded?
As of June 2026, CMA is embedded in Slack (via Claude Tag), Notion (via the External Agents API), Asana (AI Teammates), Atlassian Jira (Claude Agent for Jira), and Sentry (extending the Seer debugging agent). Enterprise deployments include Rakuten (specialist agents across product, sales, marketing, and finance) and KPMG (Digital Gateway Powered by Claude for tax and private equity clients).
How does Anthropic’s agent strategy differ from OpenAI and Google?
Anthropic uses a tool-centric orchestration approach, embedding its agent runtime inside existing tools via composable APIs and the Model Context Protocol (MCP). OpenAI chose vertical integration with Workspace Agents, positioning ChatGPT as the central hub. Google chose platform depth with the Gemini Enterprise Agent Platform and Workspace Intelligence semantic layer. Anthropic’s approach does not require users to change platforms — the agent shows up where they already work.
What percentage of enterprise apps will have embedded AI agents by end of 2026?
Gartner predicts that 40% of enterprise applications will include embedded task-specific agents by the end of 2026, up from less than 5% in 2025. However, fewer than one in nine enterprises currently run agents in production at scale, suggesting significant growth ahead.
Can Claude Managed Agents run inside a private network?
Yes. CMA supports self-hosted sandboxes through partners including Cloudflare, Daytona, Modal, and Vercel, or custom VPC deployments. MCP tunnels allow agents to connect to private Model Context Protocol servers inside your network without public exposure. A Vaults system keeps credentials out of the sandbox using envelope encryption.