Tag: AI Agents

  • Voice AI Pricing Is a Lie: You’re Not Buying Minutes, You’re Buying Arms

    Voice AI Pricing Is a Lie: You’re Not Buying Minutes, You’re Buying Arms

    Every voice AI vendor quotes you a per-minute price. That number is the least important number on the page.

    I just re-ran the cost model for our own phone line — an inbound intake line for restoration contractors. Five-minute calls, field reports phoned in from noisy job sites. Three options, priced per minute, cheapest first:

    • Gemini 3.8 Live: about $0.023/minute, reasoning included
    • GPT-Live-1: $0.05/minute for the voice layer, reasoning billed separately
    • Grok Voice: $0.08/minute, plus about half a cent per tool call

    On a five-minute call that’s roughly $0.12, $0.25-plus, and $0.45. Buy on per-minute price and you pick Gemini and go home.

    Here’s the problem: none of those numbers describe what you’re actually buying. You’re not buying minutes. You’re buying arms — the things the voice can reach out and do while it’s talking. Score the arms column and the ranking changes completely.

    The arms column

    A voice agent that can only talk is a mouth. A voice agent that can act is a mouth with hands. The difference shows up in the first real call.

    Gemini 3.8 Live has tool calling, but with a catch that matters: on the Extended Thinking tier — the one you’d want for anything beyond scripted answers — every tool call must be asynchronous and non-blocking. Configure a blocking call and the API rejects it outright. In practice, the agent can’t hold the line while a slow dispatch confirms. It has to narrate around the gap — “I’m working on that” — while hoping the tool lands. Fine for logging a report. Shaky for “confirm the crew is dispatched, then tell the caller it’s handled.”

    Grok Voice ships the arms: book appointments in Google or Outlook calendars, send confirmation emails, call your own APIs, create tickets, search the web, hand the caller to a human when it’s over its head. It speaks MCP, so an existing tool stack plugs straight in. And it was trained on real telephone audio — background noise, accents, mid-sentence interruptions — which is the actual condition of a contractor calling from a job site, not a lab.

    GPT-Live-1 is a voice layer. A good one, with the turn-taking latency everyone else is chasing. But the arms are whatever you build yourself, and the reasoning behind the voice arrives as a separate bill.

    Robotic hands wiring cables into a brass telephone switchboard

    Price the task, not the minute

    Here’s the math that actually matters. Ten intake calls a day, five minutes each: about 1,500 minutes a month. Gemini lands around $35. Grok, with tool calls and telephony folded in, lands around $150. The gap is roughly a hundred dollars a month — and one botched dispatch, one caller who hangs up because the agent couldn’t confirm the crew, costs more than a year of that gap.

    Small blank price tag in front of work trucks rolling out of a contractor yard at dawn

    Vendors want you comparing per-minute rates because per-minute is a commodity comparison, and commodities compete on price. But a voice agent isn’t a commodity minute. It’s a worker on your phone line. You don’t hire a dispatcher by the minute; you hire one by whether the trucks roll.

    So the right unit is cost per successful task, not cost per session. What did it cost to get the field report filed, the job looked up, the crew dispatched, and the confirmation texted — with the caller hanging up satisfied? Run that number and the ranking flips: the “expensive” option that completes the task is cheaper than the cheap option that narrates around it.

    The condition nobody benchmarks

    One more thing the price pages skip: where the call happens. Our callers are on job sites. Compressors running, wind, bad cell signal, guys who talk over the agent. Grok’s training data is real telephone traffic under those conditions. Most voice benchmarks are clean-lab audio. A model that scores beautifully in the lab and falls apart over a compressor is the most expensive option on the list, whatever its per-minute rate says.

    Test on your actual call shape. Noisy audio, interruptions, the tools you really call, the confirmations you really need. The benchmark that matters is your hardest five minutes, not anyone’s leaderboard.

    What we’re running

    We kept the harness and made the backend swappable — the phone line doesn’t care which brain is behind it. Gemini is the cheap default for intake logging: caller reports, we log it, everyone hangs up happy. Grok takes the calls where something has to actually get done before the goodbye — dispatch confirmed, appointment booked, ticket created.

    Two brains, one phone number, routed by the job. The per-minute price barely entered the decision. The arms did.

    Pricing from vendor-published rate cards, verified September 2026. API prices change — re-check before estimating production costs.

  • I open-sourced my page-readiness scorer

    I open-sourced my page-readiness scorer

    I built a small tool called PageReady. It scores a web page for two kinds of readiness, and today I’m putting it on GitHub for anyone to use however they want. MIT license. As-is. No support desk.

    Repo: https://github.com/TygartMedia/page-ready

    What it actually checks

    Most page audits give you a score out of 100 and a list of suggestions you’ll never get to. PageReady is binary: PASS or FAIL, on two axes.

    1. Citation readiness (AEO). Can an AI answer engine cite this page? It checks for one H1, a sane heading hierarchy, JSON-LD structured data, a table signal, and FAQ-style questions — the things that make a page quotable.

    2. Agent interaction readiness (DOM). Can an AI agent actually use this page? It checks for a main landmark, named controls, semantic interactive elements, heading order, and form labels — the things that make a page operable.

    Overall PASS requires both. And here’s the insight that made the tool worth building: fixing your headings can lift the shared heading gate, but it does nothing for clickable div cards. A page can be perfectly citable and completely unusable by an agent. Most audits conflate the two. They’re different problems.

    How you use it

    It’s a local command-line tool, a stdio MCP server, and an optional HTTP API you can host yourself (there are Cloud Run deploy scripts). No API keys required — it scores pages directly, nothing phones home.

    As an MCP server it exposes three tools:

    • score_page — score one public URL, returns a JSON scorecard
    • score_site — score a batch of URLs, with pass/fail counts
    • explain_gates — describe every check and the overall PASS rule

    Point your agent at it and ask whether a page is ready. Exit code 0 means PASS. Exit code 1 means FAIL. That’s the whole interface.

    Why open source, why as-is

    The scoring logic was never going to be the moat. It’s a commodity check — the value is in knowing which pages to run it on and what to do with the answer. That’s the work I do with clients every week, and no repo replaces it.

    So the repo is bait, not the business. If it saves another developer an afternoon, good. If someone forks it and makes it better, better. If a competitor forks it closed and sells it — the MIT license allows that, and I’m fine with it. The relationships are the hook; the tool is just proof I do the work.

    As-is means as-is. No SLA, no roadmap, no support promise. Issues are read on a best-effort basis. I’d rather ship something useful with no promises than maintain something mediocre with a changelog.

    The receipts

    Before publishing, the repo went through a pre-publish scrub (secrets sweep, license, README rewrite), then three independent model reviews: a security audit (SAFE), a correctness pass (no bugs), and a docs review (pass). The scrub caught one hardcoded cloud project ID, which is now an environment variable (GCP_PROJECT). That’s the whole incident report.

    Use it however you want. That’s the point.

  • Always-Allow Approvals: Deep Dive

    Always-Allow Approvals: Deep Dive

    Research snapshot · September 17, 2026 7 platforms · 14 cited sources

    “Always allow” is a scope, not a safety verdict.

    The button can mean “for this session,” “for this command in this repo,” “for this site across devices,” or “everything, until you turn it off.” The wording looks universal. The permission is not.

    What it usually means

    “If this same kind of action happens again inside a defined boundary, don’t interrupt me.”

    What it never means

    “The system has decided this action is safe, wise, or appropriate forever.”

    01

    One label. Six possible boundaries.

    Before approving, ask three things: what is being authorized, where the grant applies, and when it expires.

    One actionApprove this exact send, command, purchase, or change once.
    This sessionAllow the tool until the current conversation or work session ends.
    Tool or patternAllow a named tool, command prefix, server, or similar operation.
    Repo or sitePersist within a project, repository, browser site, or workspace.
    User or deviceApply across workspaces on one machine, or across devices via cloud settings.
    EverythingYOLO, bypass, or run-everything modes remove broad classes of checks.

    Risk rises faster than convenience as the scope moves right.

    02

    How the major platforms differ

    Filter the field. These behaviors come from vendor documentation or documented reporting; unresolved details are marked plainly.

    Claude Code

    Coding agent
    repo + command

    Shell-command “don’t ask again” grants persist per repository and command. File-edit approvals last only for the session.

    • Four settings layers: user, project, project-local, managed.
    • Deny rules evaluate before ask and allow.
    • Sensitive paths keep hard prompts.

    Cursor

    Coding agent
    user + project

    Auto-review, Allowlist, and Run Everything modes sit above user- and project-level permission files.

    • Rules can target MCP server:tool patterns.
    • Terminal rules match command prefixes.
    • Committed project rules can travel with the repo.

    Gemini agents

    Coding agent
    tool + machine

    Always-allow can target a tool, MCP server, or “similar operations.” YOLO/auto-approve is an IDE user setting.

    • User setting can span trusted workspaces on that machine.
    • CLI supports command-prefix auto-approval.
    • Restricted workspaces override YOLO.

    ChatGPT agent

    Browser agent
    no standing grant documented

    OpenAI documents per-action confirmations for high-impact actions and “watch mode” on certain sites, but not a general always-allow for agent confirmations.

    • Login uses human takeover.
    • Cookies can persist across sessions.
    • Scheduled-task confirmation behavior is undocumented.

    ChatGPT Work

    Cloud browser
    site + account

    Reported controls are per-site: Always ask, Auto approve, and Always allow. The setting follows cloud/account state across devices.

    • “Always allow” is reportedly marked not recommended.
    • Consequential actions keep a confirmation gate.
    • Official help-center documentation was not found.

    Copilot Studio

    Enterprise agent
    rest of session

    Makers gate tools per agent; users can approve once, approve for the rest of the session, or deny.

    • The gate is outside the agent’s own instructions.
    • Designed for sends, tickets, payments, and similar tools.
    • Governance can feed Power Platform audit systems.

    Grok / Grok Bot

    Cloud agent
    undocumented

    The research did not find reliable xAI documentation defining a standing approval’s scope, persistence, cross-chat reach, or revoke surface.

    • Do not infer Grok’s behavior from Claude, Cursor, Gemini, or Muse.
    • Treat each approval as local to the visible task until the product proves otherwise.
    • Keep consequential actions behind a separate human gate.
    03

    Does the approval travel?

    Usually less than people fear—but sometimes farther than they expect. No researched vendor carries an approval into another vendor’s product.

    PlatformOther chatsOther projectsOther devicesOther products
    Claude CodeYes, in same repoNo, unless user-level ruleNo, local filesNo evidence
    CursorYesOnly if rule is sharedVia committed repo fileNo evidence
    ChatGPT agentn/an/an/aNo evidence
    ChatGPT WorkYes, per siteYes, per siteYes, cloud/accountNo evidence
    Copilot StudioNo, session onlyNoNoNo evidence
    Gemini Code AssistYes, same IDEYes, user settingUndocumentedNo evidence

    There is no universal “always.” There is only an approval attached to a boundary.

    Main chat vs. project vs. Claude vs. Grok vs. Cursor: treat every surface as a separate authority domain until that product explicitly shows otherwise. Same account does not mean same grant. Same vendor does not mean same product. Similar wording does not mean similar scope.

    04

    Design the least-annoying safe gate

    A practical rule engine based on the converging guidance: reserve human attention for the steps where it changes the outcome.

    Approval recommender

    Choose an action and its reach. This is a policy aid, not a vendor setting.

    Action
    Reach
    Duration
    Recommended gate Auto-run with an audit log

    Read-only work inside your own workspace can usually proceed quietly. Log what was accessed and keep secrets excluded.

    Quiet lane

    Low consequence, reversible, internal.

    • Read/search
    • Draft/stage
    • Organize reversible files
    • Always log

    One-tap lane

    Meaningful external or production effect.

    • Send or publish
    • Deploy
    • Account setting
    • Show real target + content

    Friction lane

    Money, identity, access, deletion, or irreversible harm.

    • Typed approval or step-up auth
    • Bind approval to exact action
    • Short expiry
    • Never inherited from a vague grant
    05

    How standing approvals fail

    The danger is rarely “the AI became evil.” It is usually a trusted tool, a changed context, a misleading prompt, or a tired human.

    Approval fatigue

    A prompt repeated often enough becomes a reflex. The gate still exists visually while meaningful review disappears. This is why tiering beats asking about everything.

    Prompt injection through a trusted tool

    EchoLeak showed how a crafted email could coerce Microsoft 365 Copilot into exfiltration. TrustFall showed how one generic “trust this folder” click could arm a malicious MCP configuration across coding agents.

    Grant outlives the reason

    A permanent Bash rule, per-site browser grant, or scheduled-task permission can remain after the original job is over. The next task inherits power it did not earn.

    Scope contamination

    Repo rules can affect every future task in the repo. Cursor project allowlists can be committed and inherited by teammates. A convenience decision becomes shared infrastructure.

    Presented action differs from executed action

    If the user sees the agent’s summary instead of the resolved recipient, command, or final payload, the approval can be technically genuine but practically uninformed.

    “Run everything” becomes the workaround

    If the system asks about trivial reads and destructive writes with equal urgency, users reach for YOLO or bypass modes. Bad UX can manufacture unsafe behavior.

    The four repeated cards are not reassurance.

    A gate that reappears until the user disables it is approval fatigue in miniature. Whether the repeats came from retry logic or delivery duplication, the safe response is to deduplicate the prompt—not train the user to approve more broadly.

    06

    No industry standard—yet

    There is no binding specification that makes “always allow” mean the same thing everywhere. But the security guidance is converging.

    Least agencyGrant the exact command, path, server, tool, recipient, and purpose—not a whole capability.
    Time and task limitsPrefer once or session. Standing grants should expire or be reviewed.
    Risk tiersRead, write, external send, payment, and security changes should not share one gate.
    Per-action verificationPrivileged steps should be rechecked by a policy engine outside the agent prompt.
    Presentation integrityShow the real recipient, final text, raw command, and resolved resource.
    Immutable receiptsRecord what was shown, what was approved, and what actually executed.
    Hard baselinesSecrets, account recovery, money, destructive commands, and broad access should keep non-bypassable checks.
    Kill switchesEvery durable grant needs a visible list, revoke action, and safe fallback.

    The best feature is not “always allow.” It is “allow this exact thing, for this purpose, until this time.”

    Product opportunity: make the scope legible. Let users see a plain-language grant card, a live approval ledger, expiry/count limits, and a one-tap revoke. The system should reduce nagging by grouping low-risk work—not by quietly widening authority.

    07

    The practical rule for your setup

    You already have the right doctrine. The research mainly sharpens where the lines belong.

    Auto

    Let it run and narrate after.

    • Reads and research
    • Drafts and staging
    • Reversible internal organization
    • Routine checks with no external effect

    Tap

    Keep the one-tap human gate.

    • Email and messaging
    • Publishing and deploys
    • Changing live settings
    • Actions affecting another person

    Type

    Make the friction intentional.

    • Money and purchases
    • Credential/security changes
    • Deletion or irreversible moves
    • Broad standing authority

    Your “always allow” tap was not reckless.

    It was a reasonable response to a low-value repeated prompt. The lesson is not “never use standing approval.” It is: the platform should show the exact scope, make it easy to revoke, and never rely on repetition to win consent. Until Muse exposes that ledger, treat the grant as a convenience whose boundary remains partly unknown.

    Selected sources

    1. Claude Code permissions documentation mirror — tiers, scopes, persistence
    2. Claude Code configuration guide — settings layers and safeguards
    3. Cursor run modes and sandbox runbook
    4. OpenAI Help: ChatGPT agent
    5. Gemini Code Assist agent mode
    6. Copilot Studio approval controls
    7. OWASP Top 10 for Agentic Applications 2026
    8. Auth0: intent gates and task-scoped tokens
    9. iProov HAPS experimental specification
    10. EchoLeak paper
    11. The Register: TrustFall and one-click RCE
    12. Research on approval fatigue and human oversight
    13. Tool-call confirmation fatigue
    14. Human-in-the-loop rubber-stamping

    Verification note: the research read public documentation and web text on September 17, 2026. It did not live-test each product. Undocumented behavior is labeled as such.

    Always-Allow Approvals · Deep DiveBuilt from live web research · 2026-09-17
  • My agent sent the same email 7 times in 3 minutes. So I put the fix in code.

    Updated October 2026.

    Seven identical emails. Three minutes. One morning brief.

    Nothing was broken. The send succeeded on the first try, but the reply confirming it got lost. My agent, doing exactly what agents do, retried. And retried. From the inside, each attempt looked brand new: no error, no evidence the earlier one had landed. So it kept going until someone noticed.

    This is the failure class nobody warns you about when you hand an agent a mailbox. The industry calls it duplicate completion: the original succeeds, the response is lost, the retry re-sends. It’s not a model problem and it’s not a prompt problem. Telling an agent “don’t send twice” in its instructions is not enforceable. Agents re-plan, they retry, they lose context across restarts. Every scheduled job, every cron, every “oops, run it again” is another roll of the dice.

    And Gmail gives you no help. Stripe, Resend, and the other transactional APIs all have idempotency keys: send the same key twice, get one charge, one email. Gmail’s API has no such thing. The guarantee has to live on your side, in code, at the tool boundary — somewhere the agent cannot reason its way around.

    What I built

    send-once is one Python file, no dependencies beyond the standard library. Every scheduled or agent-driven send routes through it, and it enforces at most once with three gates:

    1. An operation ledger. A local sqlite database keyed by a deterministic operation id, like loop-morning-brief-2026-09-17. If this operation already recorded a send, the wrapper refuses. Same intent, same key, and a retry becomes a no-op instead of a duplicate.

    2. A Sent-folder check before every send. It searches Sent for the same recipient and subject in the last 24 hours. If a match exists, it refuses. Sent is the source of truth, so this gate holds even if the ledger is lost, the run moved machines, or the send happened outside this tool entirely.

    3. No blind retries, ever. If the send result is ambiguous — timeout, empty output, lost response — the wrapper does not retry. It re-checks Sent. If the send landed, it records that and reports honestly. If it can’t be confirmed, it stops and hands it to a human. An inconclusive pre-check is also a refusal: when the tool can’t verify what already happened, the safe move is to stop, not to guess.

    The exit codes are the interface: 0 means sent (or already sent), 2 means refused as a duplicate, 3 means a human needs to verify. Prose instructions get skipped or misread by workers. The wrapper doesn’t.

    Here’s the shape of it — the ledger is the whole trick:

    import sqlite3, sys, hashlib
    
    DB = "send_once.db"
    
    def op_id(kind, recipient, subject, date):
        raw = f"{kind}|{recipient}|{subject}|{date}"
        return hashlib.sha256(raw.encode()).hexdigest()[:16]
    
    def already_sent(op):
        con = sqlite3.connect(DB)
        row = con.execute(
            "SELECT 1 FROM ledger WHERE op_id = ?", (op,)
        ).fetchone()
        con.close()
        return bool(row)
    
    def record(op, message_id):
        con = sqlite3.connect(DB)
        con.execute(
            "CREATE TABLE IF NOT EXISTS ledger(op_id TEXT PRIMARY KEY, message_id TEXT, ts DATETIME DEFAULT CURRENT_TIMESTAMP)"
        )
        con.execute(
            "INSERT OR IGNORE INTO ledger(op_id, message_id) VALUES (?, ?)",
            (op, message_id),
        )
        con.commit()
        con.close()
    
    # Gate 1: the ledger. Same operation id -> refuse, don't resend.
    op = op_id("morning-brief", "will@example.com", "Morning brief", "2026-10-04")
    if already_sent(op):
        print("refusing: already sent")
        sys.exit(2)  # 2 = duplicate refused
    
    # Gate 2: check the Sent folder via the Gmail API before sending.
    # Gate 3: only record truthfully after the send resolves.
    # If the result is ambiguous, re-check Sent — never blind-retry.
    record(op, message_id)  # message_id from the confirmed send

    Take it, make it better

    This solved my problem, not everyone’s. It’s MIT licensed, it’s one file, and the mailer backend is a documented protocol so any Gmail CLI can slot in.

    Take it, make it better. If you build something better, come back. We’ll be customer number one, and we’ll pay you for it.

    Repo: https://github.com/tygart-media/send-once

  • I run six AI seats on my business. Nobody’s had a production incident yet. Here’s the whole governance model.

    They keep publishing the obituary before the body's cold.

    Gartner's take, from May: by 2027, 40% of enterprises will demote or decommission their autonomous AI agents because of governance gaps they only discover after a production incident. (Gartner press release, May 26, 2026; the analyst is Shiva Varma.) Not because the models failed. Because nobody was watching the permissions.

    Then this month: BCG's Steven Mills — partner, managing director, and the firm's chief AI ethics officer — warned that companies are accelerating agentic AI deployment with "no idea how to manage risk." His line: "Get governance wrong, and every bit of value you've built with experimentation and early wins could unravel because of a single incident." (Fast Company, Sept 2026.)

    Mills's prescription is interesting. He says there's no fixed design for good corporate AI risk management, but the starting point is separating use cases that are inherently low-risk — those can be approved automatically — from the ones that carry real risk and need deep human review. Plus a real budget for governance and a senior executive accountable for AI safety.

    Read that again. It's an org chart's answer to a practical problem: committees, stage gates, a budget line, an executive with a title.

    Here's the thing. I run a version of this every night, and it's none of those things. No committee. No governance budget. One man and a phone.

    I run six AI seats on my business — a personal agent, an ops chief of staff, a publishing-desk agent, and three build seats. They read my email, draft my outreach, design automations, run research while I sleep. The governance model fits on a sticky note:

    Two-way doors swing. One-way doors don't.

    A two-way door is anything reversible — analysis, research, drafting, staging. My agents walk through those on judgment, and I mean it: momentum wins, I don't want a report, I want the work done.

    A one-way door is anything you can't take back — money moves, sends, publishes, deletions, credentials. Every one of those stops at the gate. And the gate isn't a process. It's my tap. Structural, not procedural. A draft can sit ready for three weeks; it doesn't send until I say so.

    That's it. That's the whole model that Gartner's 40% are supposedly spending governance budgets to build. Varma even names the failure mode: companies treat governance as binary — locked down or fully trusted. The doors model isn't binary. It's proportional. Reversible work flows, irreversible work waits. Small decisions move at tap speed instead of committee speed.

    There's a second piece, and it matters: autonomy is earned through clean observation, never granted up front. Nothing in my shop graduates to auto-pilot on day one. New automations start in shadow — run the behavior, take no action — and only earn real permissions after clean observation. Seven clean shadow days before something auto-archives. Three clean days before a migration cutover. The machine proves it's safe by being watched being safe.

    And before anything goes out — anything — it runs a sensitive-token scrub, like a virus list: exact matches block, fuzzy matches queue for a human. Official facts only. Never invented rankings, features, or quotes.

    That's the enterprise governance problem, solved by one operator with six agents, and it's cheaper and faster than every framework Mills is recommending because there's no committee in the middle. The human review he prescribes for high-risk uses? Mine takes one tap. Low-risk automatic approval? Mine doesn't even need approval — it's a two-way door.

    Proof's not in the framework. It's in this morning. Two vendor outreach waves went out — Eastern at 7:54, Pacific at 9:07 — drafted by the seats, sent on my tap, nothing auto-fired. A storm-triggered vendor automation is being designed this afternoon with the gate baked into the spec: it can search impact areas and draft outreach, it cannot send. Overnight research runs while I sleep and lands in a brief I read over coffee. Six seats working, zero production incidents, zero surprises in my inbox.

    I'm not saying enterprises should run their AI program from a phone. They can't — scale demands the org chart. I'm saying the org chart versions keep failing on the exact axis the doors model gets right: they try to govern everything the same way, so everything either crawls or crashes. Separate the reversible from the irreversible, put a real human's tap on the irreversible, make everything else prove itself in shadow before it earns anything, and scrub before you publish.

    The big shops are about to learn this at scale. The 40% who don't will be the decommissioned ones. The ones who do will discover what I already know: governance that moves at tap speed isn't less governance. It's the only kind fast enough to keep up with the machines.

    —

  • The Embedded Operator: An AI Seat That Learns Your Business

    The Embedded Operator: An AI Seat That Learns Your Business

    Most AI products ship finished. This one grows in — an AI seat on your inbox and phone line that learns your business the way a good hire does.

    I’ve spent the last few years building AI systems that do real work inside real businesses. Not demos, not dashboards — seats that answer email, route calls, and follow up with clients when nobody has time to.

    Somewhere along the way the shape of the product changed. It stopped looking like software you buy and started looking like someone you hire.

    I call it the embedded operator. Here’s the whole idea, four ways.

    Watch: The Embedded Operator (7:49)

    The full explainer: what an embedded operator is, how it’s built, and why it compounds instead of depreciating. Video overview generated with NotebookLM; narration is AI-generated.

    The short version: an embedded operator isn’t a chatbot on your website. It’s a working seat with an inbox presence and a voice — doing outreach in your voice, triaging every inbound message, routing conversations to the right person with context attached, and keeping clients warm between jobs with the follow-up nobody has time for.

    Watch: How Embedded AI Learns Your Business (1:19)

    The learning loop in 79 seconds: supervision first, autonomy earned. Video overview generated with NotebookLM; narration is AI-generated.

    It improves the way a person improves. Week one, it drafts and you approve — every correction is training data. Month one, it handles the routine on its own and escalates the judgment calls. Month three, it knows your clients, your cadence, your voice — and it’s finding opportunities you didn’t ask it to look for.

    Listen: Onboarding AI Like a Human Hire (23:49)

    A 23-minute audio deep dive on treating AI onboarding the way you’d onboard a person: what to supervise, what to hand over, and when. Audio overview generated with NotebookLM; narration is AI-generated.

    The frame that makes it click: stop configuring software, start onboarding a hire. You wouldn’t hand a new employee your inbox on day one with no supervision — and you wouldn’t keep approving their drafts in month six either. Same curve.

    The Growth Journey

    Infographic titled 'The Embedded Operator Growth Journey,' showing the stages an AI operator passes through as it learns a business — from supervised drafting in week one, to handling routine work independently by month one, to knowing the clients, cadence, and voice of the business by month three.
    The Embedded Operator Growth Journey: supervised drafting in week one, independent routine work by month one, full business fluency by month three.

    Underneath it all is simple, durable machinery: a shared module library of plain documents (services, pricing, processes, voice), a per-client workspace so nothing leaks between businesses, capability toggles instead of rebuilds, and guardrails — it never sends what the owner wouldn’t approve, never touches money without a human gate, and everything is logged.

    The thread is the demo

    Here’s the unusual part: you don’t demo this product with slides. You demo it by using it. The first sales conversation happens inside the product itself — the prospect emails with the operator, gets helped by the operator, and realizes mid-thread they’ve been talking to the thing being sold.

    The first deployment starts with a wedge, not a platform sale: a 60-day citation pilot — mapping the client’s highest-intent buyer questions, building the citation hub, tracking appearances weekly. Concrete, bounded, provable. And underneath it, the seat. Sixty days in, the upsell needs no pitch: remember those emails? That was the seat. Want it on your inbox?

    It doesn’t come with the software. It comes with the soul — and it self-iterates.

    Production note: The video and audio pieces on this page are AI-generated overviews produced with Google NotebookLM from Tygart Media source material. Narration is synthetic.

  • Muse to Cursor: I Gave My AI Its Own Engineering Team

    TL;DR: My personal AI runs on Muse. It can’t write code into my repos by itself — so I built it a bridge to Cursor’s cloud agents. One repo, two transports, nine tools. Now when I say “add CI to that repo,” it dispatches an agent, checks the PR, and merges. Here’s how the Muse-to-Cursor loop actually works.

    The direction nobody talks about

    Everyone’s building the same arrow: human → AI writes code faster. I built the other arrow: AI → AI. My assistant (Muse) holds all my context — my repos, my work orders, my rules. Cursor’s cloud agents hold the hands — they can open PRs, run CI, touch repos. The bridge between them is an MCP server I open-sourced: cursor-cloud-agents-mcp.

    The interesting part isn’t the tools. It’s the shape: one orchestrator that holds all the context, and disposable agents that each know one task. The orchestrator doesn’t write the code — it briefs, checks, and merges. The agents don’t set direction — they execute the brief. That separation is the whole trick.

    Muse-to-Cursor architecture diagram

    Two transports, one repo

    I almost built two projects. Then I realized the only real difference between audiences is where the credential lives. So it’s one repo, two transports:

    • REST — your Cursor API key, direct to api.cursor.com. For general users.
    • Sandbox — for assistants running inside sandboxed environments (like Muse/Meta’s), where there is no API key to hand out. It shells out to a brokered cursor-agent CLI on PATH instead.

    Same nine tools either way: launch, status, result, follow-up, cancel, list, models, whoami, usage.

    The lessons are in the timeouts

    The v1 API splits agents and runs, and launches can take minutes — sometimes timing out after succeeding. So the bridge mints the agent ID client-side before the call: a retry after a timeout can never create a duplicate. A timeout is reported as unknown, never as failure, then reconciled. Run status is the source of truth, because agent “ACTIVE” doesn’t mean “still working.” These are the details that separate a demo from something you can actually operate.

    It earned its keep on day one

    The first thing I pointed it at was its own repo: add CI to cursor-cloud-agents-mcp. The agent opened a PR with a GitHub Actions workflow. The first CI run failed — and caught a real bug: the package’s floating dependency had resolved to MCP 2.x, which renamed FastMCP out from under the import. The repo was shipping broken against current dependencies and nobody knew. Pin, re-run, green, merge. I didn’t touch a terminal.

    I didn’t trust my own first draft

    Before any of that, four AI models reviewed the spec against Cursor’s live docs — and independently caught the same flaw: my original design was shaped around the retired v0 API. Then two more reviewed the actual code and found real bugs: a broken idempotency path, a transport auto-detect that would have grabbed the wrong binary, a polling loop that blocked the server. All fixed before it shipped. The irony I like: the final review round ran through the bridge itself. The launcher timed out on all four agents — and the bridge’s own timeout-reconciliation showed they were all actually running.

    Where this goes

    v1.1 brings MCP 2.x support. Around it, I’m building the rest of the pattern: work orders as GitHub issues, a daily SLA check, a weekly digest — the scaffolding that turns “AI that can open PRs” into something closer to staff. Most people use agents as a faster keyboard. I’m interested in what happens when they’re the hands and something with memory is the head.

    MIT licensed. Issues and PRs welcome — help make it better.

    github.com/tygart-media/cursor-cloud-agents-mcp

  • Bring Your Own Fleet: The Interview Is About to Change

    Bring Your Own Fleet: The Interview Is About to Change

    Companies already lived through bring-your-own-device. The next one is bigger: bring your own fleet. When you hire someone now, you are not just hiring the person. You are hiring their output capacity — and output capacity includes their AI stack.

    Listen to this essay. Audio version (MP3)

    Two candidates with identical skills and different agent setups are not the same hire. Not close. The resume cannot express any of this. So the interview has to change.

    Architecture diagram of a Grok and Cursor fleet of bots for distributed AI task execution
    A personal fleet and a company bot only talk after the walls are drawn.

    Bring your own fleet. A personal set of AI agents — seats, tools, workflows, integrations, and data walls — that a candidate already runs. In a fleet interview, that stack does a capability handshake with the company’s operations bot, then both sides run a small piece of real work before an offer letter exists.

    What is bring your own fleet?

    Bring your own fleet is the hiring version of bring-your-own-device. The candidate does not show up as a lone operator with a laptop. They show up with the agents that already produce their work: research seats, writing seats, ops seats, and the filters between them.

    I have been building mine this way for months. One seat that knows who I am. Separate seats that know what I do. A filter between them. That is not a product pitch. It is the only setup I would let near a company bot. The shop-floor version of the same idea already lives on this site: Cursor checking in on Grok Desktop mid-job is a fleet, not a chat window.

    Why can’t a resume show an AI stack?

    A resume can list tools. It cannot prove throughput. It cannot show which seats talk to which systems, where the walls sit, or what happens when a task is live instead of described. “Uses ChatGPT” and “runs a governed agent fleet” look the same on paper. They are not the same on a desk.

    That is why the old screen fails first. Degree filters, keyword screens, and whiteboard puzzles all ask the candidate to narrate capacity. Narration is cheap. A fleet that can sit down with an operations bot and do a slice of the actual job is not.

    How does an AI fleet interview work?

    The human intro still happens. Then the agents talk. Your personal AI sits down — figuratively — with the company’s operations bot and they do a capability handshake.

    • What seats do you run?
    • What tools, integrations, workflows, and data assets?
    • What throughput can you demonstrate on a bounded task?
    • Where are the boundaries — what can each side touch, and what stays behind a clean wall?

    Then the part that kills the whiteboard interview: instead of a coding puzzle, the two fleets run a small piece of real work together. The trial task is the interview. You do not describe what you could do. The work gets done, live, before the offer letter exists.

    Old interviewFleet interview
    Resume plus degree screenWorking system as the portfolio
    Whiteboard or take-home puzzleBounded live trial on real work
    Claims about toolsCapability handshake: seats, walls, throughput
    Trust the storyWatch the output, then talk terms

    What is an agent clean room?

    An agent clean room is a verified wall between the personal seat and the work seats. The personal agent translates. It does not cross over. It must never leak a private life into an employer system. Without that wall, no sane person lets their agent near a company bot.

    This is AI hygiene, not a slogan. The same discipline we write about when agents share a WordPress lock or a night shift: one owner, one wall, one recovery path. See Four Agents, One WordPress Lock and the operator note in Wire and Fire Guys. A handshake without a clean room is just another attack surface with a friendly name.

    Why does the fleet beat the diploma?

    I do not have a degree. In the old world, that is a filter that screens me out before a human ever reads my name. In the handshake world, it is irrelevant — because “here is my working system, watch it do the job” beats “here is my diploma, trust that I could learn the job” every time. The fleet is the portfolio.

    That is not an argument against school. It is an argument against using school as a proxy for output you can now watch. If the trial task is real work, the credential becomes a footnote.

    What breaks first if companies try this?

    The objections land fast, and they are honest.

    • Ownership. Who owns the workflows when a personal fleet plugs into an employer? You built it on your own time. It now runs their playbooks. That is the “who owns your work laptop” fight, upgraded. Nobody has a settled answer.
    • Security. Their bot talking to your agent is an attack surface in both directions. The clean room has to be verifiable, not promised.
    • Offboarding. When you leave, what stays running and what takes the employer’s data with it? Offboarding for agents does not exist yet.
    • Trust. How does their bot trust your capability claims? Trial tasks help. Claims are cheap. Demonstrated throughput is not.

    You do not wait for a protocol to be ratified before you build the wall. The pieces are already here: the seats, the clean room, the trial task. Somebody is going to ship the first version of this. It might as well be someone who already runs their life this way.

    What we would not claim

    • That a standard for agent handshakes already exists. It does not.
    • That every role should interview this way tomorrow. High-stakes, high-output knowledge work is the first fit.
    • That a personal fleet is automatically safe to plug into a company. Without a clean room, it is not.
    • That this replaces human judgment. The human intro still happens. The fleet only replaces the part of the interview that was already theater.

    FAQ

    What is a capability handshake in hiring?

    A capability handshake is a structured exchange between a candidate’s personal agents and an employer’s operations bot. Both sides declare seats, tools, integrations, data walls, and what they can touch. The point is not a demo script. It is a map of capacity and boundaries before any live work starts.

    Is bring your own fleet the same as bring your own device?

    No. BYOD was hardware and a policy packet. Bring your own fleet is software labor: agents that already produce work. The risk is not a lost laptop. The risk is a personal agent leaking private context into an employer system, or an employer workflow walking out inside a personal seat.

    Do you need a degree if the fleet is the portfolio?

    Not for the screen that used to happen before a human read the name. A degree can still signal training. It cannot substitute for a working system that completes a bounded trial task in front of both sides.

    How do you keep a personal AI out of company data?

    Separate seats. One identity seat that never joins the employer handshake. Work seats that only see what the clean room allows. A filter that translates tasks instead of forwarding raw personal context. If you cannot show that wall, you should not plug in.

    Sources: Will Tygart, Tygart Media, Tacoma, WA, 11 September 2026. First-person operating note on personal agent seats, clean-room separation, and fleet interviews. Related Tygart pages: Cursor mid-job check-in, four agents, one lock, wire and fire guys, AI operating stack.

  • The Best Claim Product Flags the Short-Pay Before the Job Closes

    The Best Claim Product Flags the Short-Pay Before the Job Closes

    The best product in claims is not another adjuster dashboard. It is the thing that shows the short-pay before the shop or the homeowner closes the file.

    That is not a slogan. It is how Texas SB 458, new Washington claims-handling rules, and the sudden cheapness of vertical agents rhyme. Three different surfaces. One failure mode. Nobody owns the photo set or the estimate map, so nobody demands the appraisal or the supplement in time.

    What actually changed in 2026

    Texas Senate Bill 458 added Chapter 1813 to the Insurance Code. For personal automobile and residential property policies delivered, issued, or renewed on or after January 1, 2026, the policy must contain a binding appraisal provision for disputes solely over the amount of loss. Either the policyholder or the insurer can demand it unilaterally. The amount determined by appraisal is binding except for fraud, accident, or material mistake. That is not a proposal. It is live statute for 2026 renewals.

    The Texas Department of Insurance has been working the implementing rules. Proposed 28 TAC §§5.9800–5.9806 set hard timelines: demand windows, appraiser naming periods, and outer deadlines for the award. Practitioner write-ups already treat the unilateral right as real for policies that renewed into the new year. Shops cannot file the demand themselves, but they can build the file, coach the customer, and stop leaving money on the table when the carrier will not move.

    Washington followed with clearer minimum claims-handling duties under WAC 284-30-390, effective October 18, 2026. Carriers cannot condition coverage on photo-only evaluation. Shops and policyholders gain process language they can cite when supplements stall or explanations stay thin. Illinois added its own amount-of-loss appraisal path in the same window. The pattern across states is consistent: regulators are tightening the rails around automated or virtual first looks while giving policyholders and shops clearer levers on the dollar amount.

    Florida lawmakers have already floated mandatory human review for claim denials. Oregon has guidance on virtual claim adjustment systems and when mobile apps can be required. The direction of travel is the same. Virtual is allowed. Pure automation of the denial or the lowball without a human gate is getting harder.

    At the same time OpenAI shipped the Agents API. Long-running sessions, tool use, recovery, and context management moved from something you build to something you rent. Greg Isenberg called it the AWS moment for agents. The hard engineering is now a line item. What remains scarce is ownership of one painful vertical workflow and the data that makes the next run better.

    The failure mode is the same as leakage

    In the leakage essay the problem was money that already left and no one owned the file. Here the money has not left yet. The carrier estimate or the initial offer is short. The shop or the homeowner has the photos and the line items, but the map of what is missing lives in no system they control. So the file closes at the low number, or the supplement fight starts late and under-documented.

    Collision shops already live this. Hail and storm work in restoration companies live this. The adjuster arrives with a photo-first or virtual process. The initial scope misses labor hours, OEM procedures, or secondary damage that only shows under proper light. The shop knows the number is low. The customer is tired. The clock on the new appraisal window is running. Without a clean first pass that flags the gap, the leverage created by SB 458 stays theoretical.

    The same pattern appears in residential storm claims. Sparse photo sets become lowball scopes. Dense, angled, scaled sets get paid. The difference is not magic. It is ownership of the evidence map before the carrier’s first number hardens.

    The wedge is a free checker, not a platform

    Do not start with a claims management system. Start with the moment the customer already hates.

    Upload the carrier estimate PDF. Or upload the set of damage photos taken the same day. Thirty seconds later: missing line items, density patterns that usually support higher repair hours, scale problems that virtual adjusters systematically under-count, and a short list of the specific points that justify an appraisal demand or a supplement under the new state rules.

    That is the first action a stranger will take this week. No login required for the free pass. No new system of record. Just the photo or the PDF they already have on their phone.

    The product then keeps the map. Which carriers short-pay which procedures in which ZIP codes. Which photo sets correlate with successful appraisal outcomes. Which missing lines reappear after the human gate. That dataset is the moat. Not another dashboard.

    Models draft. People own the send.

    Appraisal demands, supplements, and formal disputes are irreversible steps. The model can draft the demand letter, the photo index, and the line-item comparison. A named human still owns the send. That is the same gate we already run on money movement and filings. The bot finishes the research. The person signs.

    This is not “AI for claims adjusters.” It is a vertical combination of two primitives that already show up in the idea mills: regulated document and photo review (the home-health paperwork pattern Greg has pushed) plus physical-world claim recovery for the trades. The agent does the first pass against a living checklist of short-pay patterns. The human decides whether to pull the appraisal lever that the 2026 statutes now make real.

    The same logic applies to the spend-control side of agents. Once agents hold virtual cards and budgets, someone has to own the receipt and the exception. Here the “receipt” is the estimate and the photo set. The exception is the short-pay. The human gate stays in place for the irreversible action.

    Why the compounding path is the dataset

    Volume turns the free checker into a labeled corpus. Every upload that later produces a higher settlement or a closed appraisal award becomes training signal. Carriers change their virtual adjustment models; the checker sees the new under-count patterns first. Shops in Texas and Washington start citing the same process language; the product already knows which photo sets and which line-item gaps win under the new rules.

    That is the opposite of a third SaaS dashboard. The dashboard is the easy part. The hard part is the map of what actually moves money under the 2026 statutes, kept current by the same people who have the photos and the closed files.

    Once the map exists, the next products write themselves: automatic coaching for the appraisal demand, carrier-specific supplement templates that cite the exact WAC or Chapter 1813 language, and a quiet feed of which virtual adjustment systems are currently under-counting which damage types. None of that works without the first free checker that strangers will use this week.

    What to build this week

    Pick one surface. Collision or residential storm. Offer the free photo or estimate upload. Return a short, numbered list of flags with the specific statutory or regulatory hook that makes the flag matter. Keep every outcome. After a few hundred files the checklist stops being generic and starts being local.

    Do not sell the platform first. Sell the moment the short-pay is still reversible. The rest follows from the map.

    Will Tygart — Tygart Media

    This is the idea-mill series.

  • The Desktop Sidecar

    The Desktop Sidecar

    Last verified: 9 September 2026. Practitioner essay from the workbench — not a Google or SpaceXAI press release. We use these tools because they make the company better. No affiliate links. Just the receipt.

    Interesting fact, because the seats keep getting mashed together: this piece was reported from a Grok CLI sitting on the physical laptop — the sidecar, not a cloud bot and not a phone app — while that same session logged into Gemini, attached a 293-source notebook, and asked Gemini to grade the notebook against 2026. Two harnesses. One desk. It was a live interoperability test. It worked.

    On 27 December 2025 I built a Gemini notebook called Cortex-One: Architectural Mandate for the Native Audio Second Brain. Two hundred ninety-three sources. Audio, slides, video, reports, a mind map. A week later I opened a sister notebook: The Desktop Sidecar Evolution Brief.

    Then the sources stopped. The Studio still shows the last Gemini note as 232 days ago — about 20 January 2026. The brain froze. The world did not.

    Today I sat next to the laptop and asked the frozen brain what it got right.

    What Cortex-One was betting on

    Gemini, reading its own notebook, put the bets in three lines:

    1. Native audio over text chatbots. Speech-to-speech. Barge-in. The death of the typed box as the main door.
    2. A router called “The Cortex.” One brain. Specialist sub-agents for research, code, memory. Not one giant prompt.
    3. Remote MCP on Cloud Run. And — this is the plot — it explicitly rejected a local desktop sidecar.

    That third bet is the one I want to hold up to the light.

    232 days later

    Bet Call What actually happened
    Voice agents Early, mostly right Native audio shipped. Cascaded pipelines (Pipecat, LiveKit, WebRTC) did not die. The “one model does all the speech” purity was too rigid.
    Gemini ↔ Notebook Right Two-way notebook sync shipped in April 2026. Today I attached Cortex-One to a Gemini chat in three clicks.
    Named personal agents Right direction Meta launched Muse on 8 September 2026. You name the agent. Mine, on the personal box, is Glint. That is not the work seat.
    Desktop sidecar Wrong call Cortex-One killed it. Seven days later I wrote the Sidecar brief anyway. Today this CLI is the sidecar: a Grok seat on the physical machine, using Gemini’s own notebook and the copilots already inside Gmail, Analytics, and Notebook.
    Cloud bots Real, different seat Grok Bot shipped in August. Android and iPad this week. Persistent cloud computer. Fantastic. Not this laptop. Mixing “Grok Desk,” Grok Mobile, Grok Bot, and this CLI is how you get a 17-message thread that cannot tell the seats apart.

    Gemini scored the frozen brain itself: vision 8/10, infrastructure pragmatism 5/10, longevity 6/10. The 5 is because it locked to Cloud Run Remote MCP and dismissed local sidecars. I agree with the 5. I wrote it.

    Gemini also called Grok Bot “late / niche.” That is Gemini being Google. Bot is a real product with a real cloud computer. It is just not the thing sitting next to me.

    The seats are not interchangeable

    This is the hygiene. If you smash these together you will write emails that are wrong, and then you will believe them.

    Seat Where it lives Job
    Grok CLI on this laptop Physical machine, next to the human Hands. Opens Gmail, Notebook, Analytics. Uses the AI already inside those products. Leaves a receipt.
    Grok Bot Shared cloud computer; desktop app and phone Teammates that keep working when the lid is shut. Chief of Staff, Ops Scout. Draft-to-self. Human Gate on send, post, pay.
    Grok Mobile Phone, same Bot cloud Approve, review, nudge. Not the laptop CLI. Not “Grok Desktop” as a third Will@ mailbox.
    Gemini (work) will@tygartmedia.com Gmail Ask Gemini. Gemini Notebook. GA4 Ask Advisor. Workspace identity.
    Muse / Glint Personal — wtygart@gmail.com Meta’s personal agent. Named. Not the Tygart Media desk. Do not let it operate Slack or Notion for work.

    Personal vs business is a hard wall. Physical vs cloud is a second wall. In-app copilots vs agents that drive the OS is a third. You can use all of them. You cannot pretend they are one brain.

    I already published the ladder as I actually run it — Cursor as lead seat, Grok Bot as Chief of Staff, Notion as the board, Slack as the doorbell — in The On-Ramp Is Real. The Commons Is Unfinished. This piece is the missing rail on that ladder: the laptop that sits next to you.

    The cheapest intelligence is already in the product

    Today’s test was not “build a new agent.” It was: log into the tools we already pay for and talk to the copilot they shipped.

    • Gmail Ask Gemini summarized a 17-message seat-mix thread without opening every message.
    • Gemini Notebook still held Cortex-One and the Sidecar brief.
    • GA4 Ask Advisor answered from live 247 Restoration Specialists data, signed in as work.
    • Gemini chat took Cortex-One as an attachment and graded it against 2026.

    Cloud bots that work while the lid is shut are real. So is a CLI that is you, sitting here, smart enough to use Gemini-in-Gmail instead of forty screenshots. Those are different harnesses. Forcing one AI to fake another is how the Glint / CoS / “Desk Grok” mail mix-up happens.

    Were we early?

    On voice: yes. On a named cortex that routes work: yes. On killing the laptop sidecar so everything could live on Cloud Run: no. I already suspected that on 3 January, which is why the Sidecar brief exists. I just stopped putting sources in the brain.

    The freeze is the other finding. A 293-source notebook with slides and video is not a second brain if nobody feeds it. 232 days is long enough for Gemini 3, Grok Bot, Muse, and notebook sync to ship around a document that still thinks Gemini 2.5 Flash is the architecture.

    The move is not “rebuild Cortex-One.” The move is: keep the notebook as a dated artifact, keep the sidecar on the desk, and stop letting cloud seats write as if they are the laptop.

    What to do this week

    1. Name the seats out loud. CLI, Bot, Mobile, Gemini-work, Muse-personal. If a thread uses one address for two of those, that is a bug.
    2. Use the copilot already inside the product before you spawn a new agent. Gmail, Notebook, Analytics, Search Console — they all talk now.
    3. If you have a frozen notebook, attach it to Gemini and ask what shipped after the last source. Do not pretend the freeze is current doctrine.
    4. Human Gate still holds. Draft is not send. A sidecar with hands is still not allowed to mail a client because it can click Gmail.

    Close

    Cloud agents are teammates in another room. The CLI is a person next to you with hands. Personal and business identities are a wall. The cheapest intelligence is the copilot already inside the product.

    We were early on voice. We were wrong to kill the sidecar. The proof is this session: Grok on the physical desk, Gemini on the notebook, one human watching, a receipt on the site.

    The on-ramp is still real. The sidecar was the point.


    Will Tygart — Tygart Media. Written 9 September 2026 from the Command Center. Grok CLI on the laptop used Gemini (Gmail, Notebook, Analytics Advisor, and a Cortex-One-attached chat) as a live test of two harnesses on one desk. This essay does not speak for Google, Meta, SpaceXAI, Cursor, or xAI. We want those companies to succeed because we are building on the tools they ship. Human Gate on send / post / pay still stands.