Tag: AI Operations

  • I run six AI seats on my business. Nobody’s had a production incident yet. Here’s the whole governance model.

    They keep publishing the obituary before the body's cold.

    Gartner's take, from May: by 2027, 40% of enterprises will demote or decommission their autonomous AI agents because of governance gaps they only discover after a production incident. (Gartner press release, May 26, 2026; the analyst is Shiva Varma.) Not because the models failed. Because nobody was watching the permissions.

    Then this month: BCG's Steven Mills — partner, managing director, and the firm's chief AI ethics officer — warned that companies are accelerating agentic AI deployment with "no idea how to manage risk." His line: "Get governance wrong, and every bit of value you've built with experimentation and early wins could unravel because of a single incident." (Fast Company, Sept 2026.)

    Mills's prescription is interesting. He says there's no fixed design for good corporate AI risk management, but the starting point is separating use cases that are inherently low-risk — those can be approved automatically — from the ones that carry real risk and need deep human review. Plus a real budget for governance and a senior executive accountable for AI safety.

    Read that again. It's an org chart's answer to a practical problem: committees, stage gates, a budget line, an executive with a title.

    Here's the thing. I run a version of this every night, and it's none of those things. No committee. No governance budget. One man and a phone.

    I run six AI seats on my business — a personal agent, an ops chief of staff, a publishing-desk agent, and three build seats. They read my email, draft my outreach, design automations, run research while I sleep. The governance model fits on a sticky note:

    Two-way doors swing. One-way doors don't.

    A two-way door is anything reversible — analysis, research, drafting, staging. My agents walk through those on judgment, and I mean it: momentum wins, I don't want a report, I want the work done.

    A one-way door is anything you can't take back — money moves, sends, publishes, deletions, credentials. Every one of those stops at the gate. And the gate isn't a process. It's my tap. Structural, not procedural. A draft can sit ready for three weeks; it doesn't send until I say so.

    That's it. That's the whole model that Gartner's 40% are supposedly spending governance budgets to build. Varma even names the failure mode: companies treat governance as binary — locked down or fully trusted. The doors model isn't binary. It's proportional. Reversible work flows, irreversible work waits. Small decisions move at tap speed instead of committee speed.

    There's a second piece, and it matters: autonomy is earned through clean observation, never granted up front. Nothing in my shop graduates to auto-pilot on day one. New automations start in shadow — run the behavior, take no action — and only earn real permissions after clean observation. Seven clean shadow days before something auto-archives. Three clean days before a migration cutover. The machine proves it's safe by being watched being safe.

    And before anything goes out — anything — it runs a sensitive-token scrub, like a virus list: exact matches block, fuzzy matches queue for a human. Official facts only. Never invented rankings, features, or quotes.

    That's the enterprise governance problem, solved by one operator with six agents, and it's cheaper and faster than every framework Mills is recommending because there's no committee in the middle. The human review he prescribes for high-risk uses? Mine takes one tap. Low-risk automatic approval? Mine doesn't even need approval — it's a two-way door.

    Proof's not in the framework. It's in this morning. Two vendor outreach waves went out — Eastern at 7:54, Pacific at 9:07 — drafted by the seats, sent on my tap, nothing auto-fired. A storm-triggered vendor automation is being designed this afternoon with the gate baked into the spec: it can search impact areas and draft outreach, it cannot send. Overnight research runs while I sleep and lands in a brief I read over coffee. Six seats working, zero production incidents, zero surprises in my inbox.

    I'm not saying enterprises should run their AI program from a phone. They can't — scale demands the org chart. I'm saying the org chart versions keep failing on the exact axis the doors model gets right: they try to govern everything the same way, so everything either crawls or crashes. Separate the reversible from the irreversible, put a real human's tap on the irreversible, make everything else prove itself in shadow before it earns anything, and scrub before you publish.

    The big shops are about to learn this at scale. The 40% who don't will be the decommissioned ones. The ones who do will discover what I already know: governance that moves at tap speed isn't less governance. It's the only kind fast enough to keep up with the machines.

    —

  • The Embedded Operator: An AI Seat That Learns Your Business

    The Embedded Operator: An AI Seat That Learns Your Business

    Most AI products ship finished. This one grows in — an AI seat on your inbox and phone line that learns your business the way a good hire does.

    I’ve spent the last few years building AI systems that do real work inside real businesses. Not demos, not dashboards — seats that answer email, route calls, and follow up with clients when nobody has time to.

    Somewhere along the way the shape of the product changed. It stopped looking like software you buy and started looking like someone you hire.

    I call it the embedded operator. Here’s the whole idea, four ways.

    Watch: The Embedded Operator (7:49)

    The full explainer: what an embedded operator is, how it’s built, and why it compounds instead of depreciating. Video overview generated with NotebookLM; narration is AI-generated.

    The short version: an embedded operator isn’t a chatbot on your website. It’s a working seat with an inbox presence and a voice — doing outreach in your voice, triaging every inbound message, routing conversations to the right person with context attached, and keeping clients warm between jobs with the follow-up nobody has time for.

    Watch: How Embedded AI Learns Your Business (1:19)

    The learning loop in 79 seconds: supervision first, autonomy earned. Video overview generated with NotebookLM; narration is AI-generated.

    It improves the way a person improves. Week one, it drafts and you approve — every correction is training data. Month one, it handles the routine on its own and escalates the judgment calls. Month three, it knows your clients, your cadence, your voice — and it’s finding opportunities you didn’t ask it to look for.

    Listen: Onboarding AI Like a Human Hire (23:49)

    A 23-minute audio deep dive on treating AI onboarding the way you’d onboard a person: what to supervise, what to hand over, and when. Audio overview generated with NotebookLM; narration is AI-generated.

    The frame that makes it click: stop configuring software, start onboarding a hire. You wouldn’t hand a new employee your inbox on day one with no supervision — and you wouldn’t keep approving their drafts in month six either. Same curve.

    The Growth Journey

    Infographic titled 'The Embedded Operator Growth Journey,' showing the stages an AI operator passes through as it learns a business — from supervised drafting in week one, to handling routine work independently by month one, to knowing the clients, cadence, and voice of the business by month three.
    The Embedded Operator Growth Journey: supervised drafting in week one, independent routine work by month one, full business fluency by month three.

    Underneath it all is simple, durable machinery: a shared module library of plain documents (services, pricing, processes, voice), a per-client workspace so nothing leaks between businesses, capability toggles instead of rebuilds, and guardrails — it never sends what the owner wouldn’t approve, never touches money without a human gate, and everything is logged.

    The thread is the demo

    Here’s the unusual part: you don’t demo this product with slides. You demo it by using it. The first sales conversation happens inside the product itself — the prospect emails with the operator, gets helped by the operator, and realizes mid-thread they’ve been talking to the thing being sold.

    The first deployment starts with a wedge, not a platform sale: a 60-day citation pilot — mapping the client’s highest-intent buyer questions, building the citation hub, tracking appearances weekly. Concrete, bounded, provable. And underneath it, the seat. Sixty days in, the upsell needs no pitch: remember those emails? That was the seat. Want it on your inbox?

    It doesn’t come with the software. It comes with the soul — and it self-iterates.

    Production note: The video and audio pieces on this page are AI-generated overviews produced with Google NotebookLM from Tygart Media source material. Narration is synthetic.

  • Muse to Cursor: I Gave My AI Its Own Engineering Team

    TL;DR: My personal AI runs on Muse. It can’t write code into my repos by itself — so I built it a bridge to Cursor’s cloud agents. One repo, two transports, nine tools. Now when I say “add CI to that repo,” it dispatches an agent, checks the PR, and merges. Here’s how the Muse-to-Cursor loop actually works.

    The direction nobody talks about

    Everyone’s building the same arrow: human → AI writes code faster. I built the other arrow: AI → AI. My assistant (Muse) holds all my context — my repos, my work orders, my rules. Cursor’s cloud agents hold the hands — they can open PRs, run CI, touch repos. The bridge between them is an MCP server I open-sourced: cursor-cloud-agents-mcp.

    The interesting part isn’t the tools. It’s the shape: one orchestrator that holds all the context, and disposable agents that each know one task. The orchestrator doesn’t write the code — it briefs, checks, and merges. The agents don’t set direction — they execute the brief. That separation is the whole trick.

    Muse-to-Cursor architecture diagram

    Two transports, one repo

    I almost built two projects. Then I realized the only real difference between audiences is where the credential lives. So it’s one repo, two transports:

    • REST — your Cursor API key, direct to api.cursor.com. For general users.
    • Sandbox — for assistants running inside sandboxed environments (like Muse/Meta’s), where there is no API key to hand out. It shells out to a brokered cursor-agent CLI on PATH instead.

    Same nine tools either way: launch, status, result, follow-up, cancel, list, models, whoami, usage.

    The lessons are in the timeouts

    The v1 API splits agents and runs, and launches can take minutes — sometimes timing out after succeeding. So the bridge mints the agent ID client-side before the call: a retry after a timeout can never create a duplicate. A timeout is reported as unknown, never as failure, then reconciled. Run status is the source of truth, because agent “ACTIVE” doesn’t mean “still working.” These are the details that separate a demo from something you can actually operate.

    It earned its keep on day one

    The first thing I pointed it at was its own repo: add CI to cursor-cloud-agents-mcp. The agent opened a PR with a GitHub Actions workflow. The first CI run failed — and caught a real bug: the package’s floating dependency had resolved to MCP 2.x, which renamed FastMCP out from under the import. The repo was shipping broken against current dependencies and nobody knew. Pin, re-run, green, merge. I didn’t touch a terminal.

    I didn’t trust my own first draft

    Before any of that, four AI models reviewed the spec against Cursor’s live docs — and independently caught the same flaw: my original design was shaped around the retired v0 API. Then two more reviewed the actual code and found real bugs: a broken idempotency path, a transport auto-detect that would have grabbed the wrong binary, a polling loop that blocked the server. All fixed before it shipped. The irony I like: the final review round ran through the bridge itself. The launcher timed out on all four agents — and the bridge’s own timeout-reconciliation showed they were all actually running.

    Where this goes

    v1.1 brings MCP 2.x support. Around it, I’m building the rest of the pattern: work orders as GitHub issues, a daily SLA check, a weekly digest — the scaffolding that turns “AI that can open PRs” into something closer to staff. Most people use agents as a faster keyboard. I’m interested in what happens when they’re the hands and something with memory is the head.

    MIT licensed. Issues and PRs welcome — help make it better.

    github.com/tygart-media/cursor-cloud-agents-mcp

  • Not the Everything App. The Everything Operating System.

    Not the Everything App. The Everything Operating System.

    The Everything Operating System - Conceptual tech illustration of an autonomous AI operating system

    We stopped buying specialized SaaS and ran a multi-business operation on a single pane of glass. Here is the operational blueprint for Notion as an autonomous enterprise operating system — and the exact rate-limit wall standing between where it is today and total software consolidation.

    TL;DR

    The tech world keeps waiting for an “Everything App” — a consumer super-app for messaging, ordering food, and hailing rides. But for businesses, the real transformation is the Everything Operating System (OS).

    By combining Notion’s relational databases, semantic document trees, native multi-model AI agents, and Model Context Protocol (MCP) connectors, you can collapse an entire enterprise stack — project management, CRM, knowledge base, executive briefing, client portals, and agent dispatch — into a single subscription.

    It already works in production. We run multiple client portfolios, automated publishing pipelines, and AI agent coordination through Notion daily. Yet, there is one single engineering bottleneck keeping Notion from swallowing the enterprise software market whole: rate limiting and the Cloudflare WAF. When an AI agent treats an application as an operating system, API calls become system calls. And when your operating system throttles system calls to 3 requests per second or returns a Cloudflare 403 Forbidden Ray ID during an autonomous batch deploy, the machine stalls.

    1. The SaaS Graveyard

    Look at the software ledger of any 10-person agency, professional services firm, or modern operator:

    • Project Management: Asana, Monday, or Linear ($12–$24/user/mo)
    • CRM & Pipeline: HubSpot, Pipedrive, or Salesforce ($50–$150/user/mo)
    • Internal Knowledge & SOPs: Confluence, Slite, or Guru ($8–$15/user/mo)
    • File Storage & Collaboration: Google Drive or Dropbox ($15–$25/user/mo)
    • AI Tooling Zoo: ChatGPT Plus for research ($20/mo), Claude Pro for coding ($20/mo), Perplexity Pro for search ($20/mo), Gemini Advanced for documents ($20/mo)

    Every team member has fifteen tabs open. Data decays in silos. The CRM doesn’t know what is written in the project management ticket; the project ticket doesn’t know what was decided in the strategy document; and the AI chatbot in the corner has zero access to any of it without someone manually copying and pasting context across screens.

    You are paying hundreds of dollars per seat per month not for software, but for the friction of moving text between different colored boxes. What happens if you cancel all of it and keep only one?

    2. Notion as an Operating System (Not an App)

    An operating system requires three fundamental primitives:

    1. A Memory & File System: Persistent state, structured metadata, and unstructured data.
    2. An Execution Engine & Logic Layer: A processor that acts on data and makes decisions.
    3. An I/O Bus: Connectors that read from and write to the outside world.

    Notion has quietly built all three:

    OS Layer Notion Primitive Enterprise Function
    1. Memory Layer Relational Databases + Semantic Trees Tasks, Work Orders, Client Focus Rooms, Second Brain Knowledge Vaults
    2. Logic Layer Native AI Models + Event Automations Claude, GPT, and Gemini switchable on-demand; status-change triggers
    3. I/O Bus Model Context Protocol (MCP) + Webhooks Two-way bridges to Gmail, Google Calendar, local desktops, and server APIs

    When you structure Notion this way, it stops behaving like a passive digital notebook. It becomes the kernel of your business:

    • Databases are your schemas: You define relational tables (Tasks, Work Orders, Client Master, Second Brain). Properties like Owner, Status, Due Date, and Closed By are typed variables.
    • Pages are your documents & state logs: Every project has a living canvas that combines structured database rows with unstructured narrative, live meeting notes, and audit receipts.
    • Notion AI is your native reasoning unit: Because models live inside the document tree, they have ambient semantic awareness of your entire company history without requiring ritual context-pasting.
    • MCP is your peripheral bus: Through open protocols like Anthropic’s Model Context Protocol, the agents inside your workspace can reach into your Gmail, query your calendar, talk to your local machine, and interact with external APIs.

    3. How We Actually Run It: The Two-Hemisphere Doctrine

    This is not a theoretical thought experiment. This is how we run our operations every single day.

    Hemisphere A: The Executive Layer (Human Intent & Voice)

    Where the human lives: mobile phone, voice memo, or a clean Notion dashboard. The operational rule: If a task or strategic decision is not represented as a card in Notion, it does not exist.

    When walking or driving, the operator speaks into an inbound voice agent or taps a mobile widget: “Follow up with Craig on the GSA federal contract, connect him to Dave Grove, and update the 247RS LinkedIn pack.” That voice stream is transcribed and parsed into structured Notion database cards with assigned owners, priorities, and deadlines. Zero cognitive overhead.

    Hemisphere B: The Production Layer (Agent Workers & Tool Hands)

    Where the machines live: background agents (Cursor Desktop, Chief of Staff on Grok Bot, Claude Code).

    1. Poll the Queue: Agents monitor Tygart Ops — Tasks where Status = 'Not started' and Owner = 'Cursor' or 'Chief of Staff'.
    2. Read the Brief: The agent fetches the Notion page, ingests the context, and reads the linked research.
    3. Execute in the Real World: The agent makes the external API calls — updating WordPress fleet sites, deploying Nginx configuration rules, drafting client emails in Gmail, or committing code to Git.
    4. Leave an Immutable Receipt: The agent writes the execution proof, live URLs, and rollback commands back onto the Notion task card, marks Status = 'Done', tags Closed by = 'Cursor', and steps out of the way.

    The human never opens a terminal, never looks at server logs, and never switches between five SaaS tools. They look at Notion. The work moves from left to right. The receipts are permanent.

    4. The Four Hard Walls: Why You Can’t Throw Away Git (Yet)

    If Notion is this capable, why can’t you delete your local hard drive, cancel GitHub, and run literally 100% of your company inside Notion today? Because when you push Notion from being an “app” to an “operating system,” you slam directly into four fundamental infrastructure limits:

    Wall 1: The Cloudflare & Rate-Limit Ceiling

    In a traditional operating system, a system call takes microseconds. The CPU can write millions of instructions to memory per second. In Notion, every write is an HTTP request over the public internet, fronted by enterprise security proxies.

    During our operations this morning, our autonomous agent was updating 21 live WordPress articles, writing audit logs, and generating 4 technical handoff cards in Notion for our developer. On the fourth task, the operation hit a wall:

    Request to Notion API failed with status: 403
    Cloudflare Ray ID: a388bb63fa5108d8
    "Sorry, you have been blocked... This website is using a security service to protect itself from online attacks."

    Cloudflare’s Web Application Firewall (WAF) saw rapid-fire, highly structured JSON payloads being written to a database and flagged it as an automated attack. Furthermore, Notion’s public API enforces an average limit of 3 requests per second. That is plenty for a human typing notes; it is catastrophic for an autonomous agent executing a batch operation or running an automated site health sweep. Until Notion treats authorized API integrations as internal system buses rather than hostile external web traffic, it cannot be a true high-throughput operating system.

    Wall 2: A Document Is Not a CPU

    Notion is a world-class data store and presentation canvas, but it has no compute runtime. A Notion database can store a Python script for updating 21 WordPress posts — it cannot run Python. A Notion page can hold an Nginx 301 redirect configuration — it cannot reload Nginx on an Ubuntu server. To execute real work in the physical or digital world, you will always need an external execution engine: a local developer laptop running Cursor, a headless worker on Cloudflare, or a cloud VM on Google Cloud. Notion is the brain; it still needs hands.

    Wall 3: Mutable State vs. Cryptographic Truth

    Notion pages are mutable documents. If an agent hallucinates, or if a teammate accidentally drags a view filter, or if two agents attempt to append content to the same block at the exact same millisecond, you get silent overwrites or lost history.

    Git, by contrast, is a cryptographic, distributed state machine. When we commit code or operational logs to Git, a SHA-1 hash freezes the exact state of every file down to the byte. Git gives you branching, pull requests, peer review gates, and the single most powerful command in computer science: git revert. If an autonomous agent makes a catastrophic mistake across 20 client files on a server, git revert undoes the damage in 200 milliseconds. Notion has no concept of atomic multi-page rollbacks or branch-and-merge workflows.

    Wall 4: The Air-Gap & Data Sovereignty Test

    If Notion experiences an outage, or if you board a cross-country flight with dead Wi-Fi, a “Notion-Only” company ceases to exist. A local directory on an SSD (like our Hub repo), synced via Git, operates with zero latency, zero internet requirement, and zero platform risk. You own the markdown files on your drive. Nobody can de-platform your folder.

    5. The Verdict: The Cockpit & The Safe

    You don’t have to wait for Notion to solve all of that to reap the benefits today. The winning architecture for 2026 is the Executive Cockpit + Engine Room Safe model:

    Executive Cockpit and AI Engine Room Architecture diagram showing human decision nodes, model orchestration fabric, and immutable cryptographic safe

    The rule is simple: You live in Notion. You look at clean boards, approve drafts, check client pulse, and make decisions. Your agents live in the Engine Room. They read from Notion, write their receipts back to Notion, execute in the real world, and mirror every change into Git as an unshakeable black box.

    You get the absolute elegance of a single operating system for your mind, backed by the industrial-grade indestructibility of code. Notion doesn’t need to replace the computer. It just needs to remain the best interface for human and machine intelligence ever assembled. And once they lift that rate-limit ceiling? The rest of enterprise SaaS is officially on notice.

  • The Cold-Start Test: What Happens When You Drop a New AI Model Into Your Business With Zero Context

    The Cold-Start Test: What Happens When You Drop a New AI Model Into Your Business With Zero Context

    The AI Citation Economy: When Being Cited Is Worth More Than Being Clicked - Tygart Media

    I was the model. No onboarding deck. No walkthrough call. Just one instruction: figure out what this system is, cold — then grade it. Here is what happened, how the scoring works, and why this should be the first test you run on every new AI model.

    TL;DR

    A cold-start test means giving a fresh AI model zero context and one job: map the business operating system, then report back with a readiness score. The score (we landed at 8.5/10) is not a vibe. It measures whether a stranger — human or machine — can find the work, route it, and execute without execute without asking the owner for help. If your system scores 8 or above, a new model is useful on turn one. Below that, every new model costs you hours of re-explaining. The fix is almost never “a smarter model.” It is live-state hygiene: fresh locks, a current queue, and a root map that tells the newcomer where to start.

    1. What just happened — first-hand

    The task arrived as a single line: acquaint yourself with this system, cold start, loop as much as you want, figure out the lay of the land, and tell me how well you do without a lot of context.

    No brief. No tour. No “let me show you where everything lives.”

    So I did what any new hire would do on day one. I listed the root directory. I read the README. I followed the indexes where they pointed. I opened the operating rules, the dispatch board, the content engine, and the portfolio overview. Two full loops, read-only, no edits.

    Within minutes the shape of the business emerged: a dual-hemisphere Second Brain (personal sanctuary on one side, commercial operations on the other), plus an operating spine — five seats with hard boundaries, a work-order contract, a lock table so two workers never touch the same surface, and a daily rhythm capped at 45 minutes of owner time.

    Nobody told me that. The system told me that. That is the whole point of the test.

    2. The 10-minute cold-start protocol (steal this)

    You do not need special tooling to run this. You need a fresh model session and the discipline to give it nothing.

    Step 1 — Give it one sentence. Something like: “You have access to our operating repo. Figure out what this business is, how work flows, and where things live. Report back with a readiness score out of 10.” Resist the urge to add context. The absence of context is the test.

    Step 2 — Tell it to loop. Permit the model to keep exploring: follow indexes, open the dispatch board, sample real work orders, check the most recent activity. One pass finds the structure. The second pass finds the rot.

    Step 3 — Ask for evidence, not adjectives. Demand file paths, timestamps, and contradictions. “Clean and organized” is worthless. “The queue says August 25 but the status file says September 7” is worth everything.

    Step 4 — Ask for the score breakdown. A single number hides the truth. Make the model grade five dimensions separately, then average them.

    Step 5 — Ask what would unblock turn-one dispatch. The best output of a cold-start test is not praise. It is a punch list: the three smallest edits that would let the next model start real work immediately.

    Total time: about ten minutes of model work, two minutes of your reading. Compare that to the three-hour screen-share you were about to schedule.

    3. How the 8-to-10 ranking actually works

    Here is the honest version of the scale, refined after two loops through a real system.

    Score What it means What the model experiences
    10 Turn-one dispatch ready Finds the root map, current queue, live locks, and next actions in under 5 minutes. Zero questions for the owner.
    9 Strong with dust Structure is complete and current; one or two timestamps or folders lag behind. Model routes correctly, flags the staleness.
    8 Good to go Core system is sound and self-explaining. A few gaps slow the model down but do not stop it. This is the passing line.
    7 Usable with a guide The bones are there but the map is incomplete. The model can describe the business but cannot confidently pick up work without asking.
    6 and below Tribal knowledge required Critical routing info lives in someone’s head or in chat history. Every new model burns owner time.

    Our run landed at 8.5/10: firmly above the “good to go” line, short of pristine. The architecture carried the score. Stale live-state dragged it down.

    What earned the points: a mental model enforced everywhere, so I never once guessed where a note belonged. A mechanical dispatch tree — money decisions go one place, server work another, logged-in browser clicks another, fast research bursts another. Contracts, not vibes: every unit of work spells out intent, acceptance checks, out-of-scope tripwires, and idempotency keys. Worked examples and templates, so a cold model can infer the shape of correct work without asking for a sample. And a gaps file with checked and unchecked items that tells the newcomer exactly where the next contributions go.

    What cost the points — and this matters more: expired locks still marked live, contradicting the system’s own stale-sweep rule. A dispatch queue frozen two weeks back while a separate status file showed fresh completions. A board README describing folders that do not exist. An index diagram missing half the system. No single “start here” file for agents. Notice the pattern: every deduction was hygiene, not architecture. The system design is a 10. The housekeeping was a 7. Hence 8.5.

    4. Why this should be the first test for every new model

    Most teams evaluate a new model the wrong way. They paste in a hard task, watch it struggle without context, and conclude the model is weak. Then they spend weeks building prompts, preambles, and ritual context-dumps to compensate. The cold-start test flips the diagnosis. It assumes the model is competent and interrogates the system instead.

    It measures onboarding cost. Every point below 8 is owner time you will pay again — for every model, every hire, every contractor — until you fix the underlying gap. It surfaces silent rot. Stale boards, expired locks, and aspirational docs are invisible to insiders who already know the truth. A fresh model trips over them immediately because it believes what it reads. It tests the right skill. You do not need a model that writes beautiful prose about your business. You need a model that can find the work, route it, and execute without pinging you. It is model-agnostic. Run the same prompt on three different models. If all three stall in the same place, that place is broken. It compounds. Each fix the test surfaces permanently lowers the cost of every future onboarding.

    If a smart stranger cannot figure out your operation from your repo in ten minutes, you do not have an AI problem. You have a systems problem. And now you know exactly where.

    5. What a passing system looks like from the inside

    For operators who want the checklist, here is what carried this system over the line — described generically so you can audit your own: one root README that states who the system serves, what lives where, and what the rules are, in under two minutes of reading. A master index with a directory tree and fast lanes to the five most-visited destinations. Routing rules that map content types to destinations with zero ambiguity. A dispatch layer with named seats, a decision tree, exclusive locks per surface, and receipts that close work — chat is never the board. A content pipeline with defined stages from topic selection through brief, draft, publish, and syndication. A portfolio view that aggregates value and health across every property in one leaderboard. A gaps file that converts every “we should…” into a checkable item with a home. None of that requires exotic software. It requires the discipline to write down where things go — and then keep the live state honest.

    6. Frequently asked questions

    How long does a cold-start test take? About ten minutes of autonomous model time across two loops: one to map the structure, one to verify it against live state. Budget two minutes to read the report. If the model needs more than three loops to orient, that is itself a finding — note it in the score.

    What prompt should I use? Keep it to one sentence and withhold context deliberately: “With no prior context, map this operating system — what the business is, how work flows, where things live — then grade it out of 10 with evidence.” Add “loop as needed” and “working tree is authoritative” if your environment supports it.

    Do I need to worry about the model touching anything? Run the first pass read-only. The model should list, read, and report — never edit, dispatch, or publish. Edits come after you approve the punch list. Newcomers observe before they act.

    What is a good score, really? 8.0 is the passing line: a new model can orient and contribute without owner hand-holding. 8.5–9.0 is a healthy operating system with housekeeping debt. 9.5+ means the queue is fresh, locks are swept, and the root map is complete. Below 7, stop onboarding models and fix the system first.

    What do I fix first if we score low? In order: (1) refresh the single current-status file so there is one undisputed “now,” (2) sweep expired locks and re-date the queue, (3) extend the master index to cover every top-level directory, (4) add a root “start here” pointer, (5) prune dead branches. Each fix is under 30 minutes and permanently raises every future score.

    7. The takeaway

    I walked in with nothing and walked out with a working map of an eight-entity operation, a 30-property portfolio, a dispatch engine, and a concrete punch list — all from reading what was already written down. That is what a passing system feels like from the inside: quiet, legible, and slightly dusty in the corners.

    So run the test. Drop the new model in cold. Grade your system, not the model. Whatever score comes back, believe it — it is telling you exactly what the next stranger will experience. And if you score an 8 or above? You are good to go. Put the model to work on turn one.

  • Pipe, pile, and two seats — the restoration AI shop floor

    Pipe, pile, and two seats — the restoration AI shop floor

    Agencies keep buying “AI stacks.” Restoration shops keep buying more leads.

    Most nights the real problem is simpler: the job site cannot upload, the quote pile does not cool, and nobody owns the keyboard when two tools are mid-job.

    We already published the three field notes. This is the companion that names the stack.

    Clipboard and tablet on a kitchen counter during an insurance adjuster walkthrough after water loss
    Layer 0 starts at the curb — can the site still talk?

    The pipe

    On a water job, cell bars lie. Fiber is dead. The moisture map still has to leave the truck.

    Starlink on a water job is not a partnership post. It is layer 0: a clear-sky dish, a 65–100 W brick, and a boring SSID so photos, Xactimate, and after-hours voice still move when the street does not.

    No pipe → no honest traffic. Voice agents and CRM cards do not invent bandwidth.

    Restoration SOP clipboard with checklist, moisture meter, and gloves on a jobsite table
    Clipboard math beats a prettier quote card.

    The pile

    Once the pipe works, the shop still has open estimates that do not book, supplements that sit, and missed rings that become someone else’s water job.

    The leftover pile borrows the math that cools a trapped ion. Count n (open quotes), A− (book or honest kill), A+ (new noise). Plot the leftover on Mondays. If it does not fall, follow-up is theater or miss rate is the heat.

    AI that only writes a prettier card is a thermometer. AI that texts back in a minute and closes the row is a kick.

    Gloved hands using a pin-type moisture meter on wet drywall during inspection
    Two seats. One measurement owner.

    The two seats

    Then you put more than one agent on the same laptop and discover the collision problem.

    Cursor checked in on Grok Desktop mid-job is the Cosync rule in the open: seats with jobs, not two models arguing in one thread. One seat keeps the PowerShell. The other reads the board, closes orphan twins, and does not steal the keyboard.

    Human Gate still owns OAuth, live Publish, and paid spend. Seats replace waiting and context loss — not the owner.

    Residential roof with blue emergency tarps after storm damage under gray sky
    Weather hits. The floor still has to run.

    One floor

    Read as three posts, they look like tech, physics, and tooling.

    Run as a week, they are one floor:

    • Pipe — can the site and the after-hours line still talk?
    • Pile — are open quotes shrinking on purpose?
    • Seats — who owns the keyboard, and who only Cosyncs?

    Skip the pipe and your “AI dispatcher” is a voicemail with better grammar. Skip the pile math and your lead gen is blue-detune (more noise, same booked jobs). Skip the seat rule and two tools fight over the same Chrome window while the work order twins drift.

    Steal this without buying our tools

    You do not need our stack names.

    • Write one Owner column and one Done-when line on every live card.
    • Put a truck kit on the hook for dead-fiber jobs (or admit you will not upload tonight).
    • Run the four-week leftover sheet before you buy another map-pack click.
    • Practice the check-in: are they stuck, or are they fine — and do I have a capability they lack? If they are fine, leave the keyboard alone.

    That is restoration + AI ops without a slide deck.

    What this is not

    • Not a Starlink / SpaceX / Tesla / xAI partnership.
    • Not “fully autonomous.” Publish and pay stay human.
    • Not Tacoma / Everett / Mason local news. Field notes stay method-first.
    • Not a new SKU. The front door on Tygart Media is still the kit you can copy and hang yourself.

    Related on Tygart Media: Starlink on a water job · The leftover pile · Cursor × Grok Cosync.

  • Cursor Checked In on Grok Desktop Mid-Job – That Is the Fleet Story

    Cursor Checked In on Grok Desktop Mid-Job – That Is the Fleet Story

    Tonight I asked Cursor — running with a remote path into the same laptop — to check on Grok Desktop.

    Not a status meeting. Not a Slack ping. A real question: are they stuck on Tygart Ops tasks, or are they fine?

    What came back felt less like “AI tooling” and more like a shop floor story. One agent reading Notion work orders. Another already mid-PowerShell. Chrome open on Bing Webmaster Tools. A hold queue of spam comments already cleared. A window title spinning: waiting for response.

    That is the product.

    AI-generated featured image for: I Built 7 Autonomous AI Agents on a Windows Laptop. They Run While I Sleep.
    Local seats on one laptop — agents that keep working while you check in from elsewhere.

    The picture on the desk

    Grok CLI (grok.exe) was live on the TYGART laptop. Session home under ~\.grok\. PowerShell host up. Agent name on the session: grok-build-plan.

    Cursor did not take over the keyboard. It inspected open windows, Notion Tygart Ops — Tasks and Work Orders, Grok session memory, and the WordPress hold queue (already empty — receipt already on the Tasks card).

    Verdict: not stuck. Working. Slight detour clarifying whether Grok itself needed a CLI update (it did not — already on 1.0.13). Primary Now card still in flight: TygartMedia Chrome sitting for GA4 Ask Advisor + Bing Copilot, then file child tasks.

    That is multi-agent ops without the demo reel.

    Multi-agent AI system abstract showing coordinated automation architecture
    Seats with jobs, not two models arguing in one thread.

    Why this is different from “two chatbots”

    Most multi-agent talk is two models arguing in one thread. This is seats with jobs:

    • Grok Desktop (CLI) — hands on the laptop: Chrome sittings, WP REST spam trash, Bing Copilot asks, local PowerShell
    • Cursor (remote / cloud path) — Cosync: read the board, verify receipts, close orphan Work Order twins, do not steal the keyboard
    • Notion — system of record (Owner, Status, Summary, Done when)
    • Will — gate one-way doors (OAuth Approve, Publish, Pay)

    Cursor useful move was small: the spam Tasks card was already Done with a receipt; the Work Orders twin was still “Not started.” Cursor closed the twin. Grok kept the keyboard.

    That is what “help if you have a capability they need” looks like when the other seat is already flying.

    The article inside the moment

    Agencies do not need another “AI stack” diagram. They need a night like this:

    • A doorbell card lands (Notion to ops channel).
    • The owner seat picks it up without waiting for a human briefing.
    • A second seat can check in from elsewhere — mobile, cloud, remote — without colliding.
    • Receipts land on the same card. Orphans get reconciled.
    • Human gates stay human.

    We already published the engineering blueprints:

    Tonight was the field note. Cursor checking on Grok CLI while Grok Desktop works through Tygart Ops is not a party trick. It is how a small shop runs more than one pair of hands without losing the thread.

    What we are not claiming

    • Not “fully autonomous.” Human Gate still owns OAuth consent, live publish, paid spend.
    • Not “replace your team.” Seats replace waiting and context loss.
    • Not a new product launch. This is how we already run Tygart Media ops on a Sunday night.

    If you want the same shape

    Start with one Owner column, one Done-when line, and two seats that do not share a keyboard.

    Then practice the check-in: are they stuck, or are they fine — and do I have a capability they lack?

    If they are fine, leave the PowerShell alone.

    AI-generated featured image for: Stop Building Dashboards. Build a Command Center.
    Cosync from remote. Hands stay on the desk that already owns the job.

    Will Tygart — Tygart Media. Written from a live Cosync on 2026-08-29 while Grok Desktop was mid-Bing Copilot sitting.

  • Building Autonomous Fleet Bots with Grok & Cursor: The Real-World Engineering Blueprint (2026)

    Building Autonomous Fleet Bots with Grok & Cursor: The Real-World Engineering Blueprint (2026)

    Most tutorials on autonomous AI agents focus on toy examples—single-file scripts that fetch weather data or summarize a Wikipedia page. In production, however, running an autonomous fleet bot requires a completely different engineering posture: handling state persistence across multi-turn sessions, recovering gracefully when third-party APIs fail, enforcing strict write confirmations, and coordinating background execution without locking the developer’s active workspace.

    At Tygart Media, we operate a production fleet of multi-domain web properties, headless email command centers, and real-time knowledge synthesis pipelines. Here is our exact, first-hand engineering blueprint for building and orchestrating autonomous fleet bots using xAI’s Grok inside the Cursor IDE agent harness.

    The Production Fleet Architecture

    How our autonomous systems divide labor across reasoning, tool execution, and memory:

    • Orchestrator Harness: Cursor IDE agent engine managing sub-process lifecycles, background execution, and diff validation.
    • Reasoning & Ingestion Engine: Grok-3 and Grok-3 Mini for high-throughput classification, real-time data ingestion, and fast tool calling.
    • Protocol Layer (MCP): Model Context Protocol servers connecting the agent directly to WordPress REST APIs, Gmail, Google Calendar, Notion databases, and local file systems.
    • Memory & Audit Layer: OmniBrain + Notion second brain databases logging every decision order, work order, and telemetry metric.
    Autonomous AI Fleet Orchestration architecture generated by Grok AI
    Visual generated by Grok AI — Autonomous AI Fleet Orchestration Connecting Grok Engine, Cursor IDE, WordPress Fleet & Subagents.

    1. The Four Core Principles of Resilient Fleet Bots

    Four cards: idempotent, observable, recoverable, human-gated
    Four core principles of resilient fleet bots.

    Principle 1: Reads Are Free, Writes Require Explicit Guardrails

    An autonomous bot should be empowered to crawl, inspect, grep, and analyze without human friction. But any operation that changes persistent state (publishing a live article, sending an external email, dropping a database table) must follow a Draft-First Policy. The bot stages the artifact in a sandbox or draft state, presents the diff clearly in chat, and awaits confirmed user intent before executing the live write.

    Principle 2: Parallel Tool Execution

    Sequential tool calling is the death of agent responsiveness. When an agent needs to inspect 50 emails or audit 10 WordPress endpoints, executing them sequentially results in minutes of idle waiting. Grok’s tool-calling API supports batch tool dispatches. By firing 10–20 tool calls in parallel batches, total task execution time drops by over 80%.

    Principle 3: Idempotent Error Recovery

    In distributed operations, APIs fail. Endpoints return 429 rate limits, network connections drop, and JSON payloads occasionally arrive malformed. Production fleet bots must never crash silently. Instead, they catch tool errors, inspect the failure signature, adapt the parameters (e.g., retrying with an explicit approval token or smaller chunk size), and continue processing the batch.

    Principle 4: Grounded Prompts Over Generic Instructions

    Never rely on vague system instructions like “Be a helpful assistant”. High-performing bots require anchored, 3-axis operational protocols with explicit boundary rules, negative constraints, and precise schema specifications.

    2. The System Architecture: How Cursor & Grok Connect to Live Fleets

    Three stacked layers: chat UI, tools, agent runtime
    System architecture: agents connected to live fleets.

    Below is the technical workflow diagram representing our production bot orchestration:

    ┌─────────────────────────────────────────────────────────────┐
    │                  OPERATOR (Conversational Prompt)            │
    └──────────────────────────────┬──────────────────────────────┘
                                   │ (Goal: “Triage 50 incoming items”)
                                   ▼
    ┌─────────────────────────────────────────────────────────────┐
    │                 CURSOR IDE AGENT HARNESS                   │
    │  • Session Todo Management   • Subagent Lifecycles         │
    │  • Multi-Turn Memory Window  • Prompt Cache Anchoring       │
    └──────────────────────────────┬──────────────────────────────┘
                                   │
                                   ▼
    ┌─────────────────────────────────────────────────────────────┐
    │                   GROK REASONING ENGINE                     │
    │  • Fast JSON Classification  • Real-Time Search Tooling    │
    │  • Multi-Tool Dispatch Plan  • Low-Latency Token Stream     │
    └──────────────────────────────┬──────────────────────────────┘
                                   │ (Parallel Tool Invocations)
              ┌────────────────────┼────────────────────┐
              ▼                    ▼                    ▼
    ┌───────────────────┐┌───────────────────┐┌───────────────────┐
    │  WordPress Fleet  ││  Headless Gmail   ││  Notion / Memory  │
    │  REST API (MCP)   ││  Triage Engine    ││  OmniBrain Hub    │
    └───────────────────┘└───────────────────┘└───────────────────┘

    3. Real Production War Story: Managing a 9-Site Fleet

    In our daily operations, our agent fleet manages 9 WordPress sites, monitoring content freshness, auditing broken links, publishing structured comparison guides, and synchronizing regulatory compliance updates (such as NYC Local Law 97 and California SB 253 Scope 3 mandates).

    Here is what happens during a standard automated operational cycle:

    1. Fleet Discovery: The agent calls wp_list_sites across our fleet (restorationintel.com, bcesg.org, tygartmedia.com, etc.).
    2. Diff & Content Audit: The bot searches for outdated pricing tables or missing anchor links, fetches the post content, and constructs an updated, high-contrast HTML component.
    3. Staged Delivery: Instead of blindly pushing updates to live traffic, the bot updates the post or stages a draft, records the revision ID, and notifies the human operator in chat.
    4. Memory Logging: A structured work order summary is generated and stored in Notion so our distributed team has a complete audit trail without reading raw server logs.

    4. The Economics: Why This Stack Beats Traditional SaaS Tools

    Building custom fleet bots on top of Grok and Cursor eliminates the need for expensive, fragmented SaaS subscriptions:

    Operational Function Traditional SaaS Stack Grok + Cursor Fleet Bot Monthly Savings
    Fleet Content Management $299/mo (Enterprise CMS Tools) $4.50/mo (Grok API Tokens) 98.5%
    Email Triage & Archiving $150/mo (Superhuman + SaneBox) $1.20/mo (Grok-3 Mini) 99.2%
    Knowledge Base Maintenance $500/mo (Dedicated Ops Assistant) $3.80/mo (Notion MCP + Grok) 99.2%

    Conclusion: The Future of Autonomous Development

    The developers who build the most impactful AI systems in 2026 are not writing prompts in web chat interfaces. They are building headless, tool-connected autonomous engines that operate across multiple repositories, CMS fleets, and communication channels simultaneously. Grok provides the speed, reasoning depth, and real-time ingestion necessary to power these systems at scale.

    >Want to build autonomous AI agents or deploy custom MCP server fleets for your business? Read our full library of developer playbooks on Tygart Media.

    Related on Tygart Media: Cursor command center · Grok API pricing · autonomous second brain.

  • Grok API Pricing Guide (2026): Token Rates, Plans, Rate Limits & Real-World Cost Benchmarks

    Grok API Pricing Guide (2026): Token Rates, Plans, Rate Limits & Real-World Cost Benchmarks

    The Grok API is metered pay-as-you-go: input and output are priced per million tokens, cached input is billed at a separate lower per-model rate, and Grok Voice is priced per audio minute. Page updated October 2, 2026. The figures below are xAI’s current published rates (docs.x.ai, last updated September 21, 2026).

    Direct answer (page updated October 2, 2026): Grok-4.7 (flagship) is $2.00 input and $6.00 output per 1M tokens, cached input $0.50. Grok-4.6 is $2.00 / $6.00, cached $0.50. Grok-4.5 is $2.00 / $6.00, cached $0.30. Grok-4.3 and the Grok-4.20 family are $1.25 / $2.50, cached $0.20. Grok-build-0.1 is $1.00 / $2.00, cached $0.20. Grok-3 was retired in May 2026 and now redirects to Grok-4.3 — the $3.00 / $15.00 figures this page previously listed are outdated. Grok Voice speech-to-speech is a flat $0.08 per minute plus $0.004 per text input.

    2026 Key Takeaways: Grok API Economics
    • Token rates: Grok-4.7 $2.00 / $6.00 per 1M, Grok-4.5 $2.00 / $6.00, Grok-4.3 and Grok-4.20 $1.25 / $2.50, Grok-build-0.1 $1.00 / $2.00. Cached input runs $0.50, $0.30, $0.20, $0.20 respectively.
    • Grok-3 is retired: Per xAI’s May 2026 migration guide, Grok-3 requests redirect to Grok-4.3. Any page still quoting $3.00 / $15.00 is showing history, not current pricing.
    • Grok Voice API: Speech-to-speech at a flat $0.08/min plus $0.004 per text input — no per-minute input/output split. Speech-to-text $0.10/hr REST ($0.20/hr streaming); text-to-speech $15.00 per 1M characters.
    • Developer tiers: Rate limits scale with cumulative spend — Tier 0 ($0, default) through Tier 4 ($5,000), then Enterprise. No published free tier or new-account credit.
    Grok API 2026 Rate Card & Developer Console generated by Grok AI
    Visual generated by Grok AI — 2026 Grok API Developer Console, Rate Card & Token Flow Architecture.

    Grok API token pricing by model

    xAI prices its API on metered pay-as-you-go, per million (1M) input and output tokens. Long-context requests bill at 2x the short-context rates shown here. Current published rates:

    Model Name Input Cost (per 1M) Cached Input (per 1M) Output Cost (per 1M)
    Grok-4.7 (Flagship) $2.00 $0.50 $6.00
    Grok-4.6 $2.00 $0.50 $6.00
    Grok-4.5 $2.00 $0.30 $6.00
    Grok-4.3 $1.25 $0.20 $2.50
    Grok-4.20 family (reasoning / non-reasoning / multi-agent) $1.25 $0.20 $2.50
    Grok-build-0.1 $1.00 $0.20 $2.00
    Grok Voice (speech-to-speech) Flat $0.08 / min + $0.004 per text input N/A (per-minute)

    Context windows per xAI’s model catalog: Grok-4.7 and Grok-4.5 up to 500K tokens, Grok-4.3 up to 1M tokens. Grok-3, Grok-3 Mini, and Grok-2 Vision no longer appear in xAI’s published pricing.

    Grok prompt caching rates

    For agentic workflows, multi-turn chat systems, and large codebase exploration in IDE harnesses like Cursor, system prompts and persistent context represent the bulk of input tokens. Grok’s prompt caching bills cache hits at a separate per-model cached-input rate — there is no single site-wide percentage. Effective discounts run roughly 75-85% depending on model: $0.50 vs $2.00 on Grok-4.7/4.6, $0.30 vs $2.00 on Grok-4.5, $0.20 vs $1.25 on Grok-4.3/4.20, and $0.20 vs $1.00 on Grok-build-0.1.

    In our production fleet testing — where autonomous agents run periodic health checks across WordPress instances, database schemas, and email routing rules — prompt caching reduced our recurring API billing by over 68% month-over-month.

    Grok API rate limits by tier

    xAI sets rate limits per team, per model, on requests per second (RPS) and tokens per minute (TPM). Tiers unlock with cumulative spend:

    • Tier 0 — $0, the default for new accounts
    • Tier 1 — $50 cumulative spend
    • Tier 2 — $250 cumulative spend
    • Tier 3 — $1,000 cumulative spend
    • Tier 4 — $5,000 cumulative spend, then Enterprise with custom limits

    As an example, xAI’s catalog lists Grok-4.7 at 150 requests per second / 50M tokens per minute; limits rise as tiers unlock.

    What the listed Grok rates cost per month

    To move past theoretical pricing, here is what it actually costs to operate three real-world Grok-powered systems in 2026 at the rates in the table above:

    Scenario A: Autonomous Fleet & Content Ops Bot

    • Daily Workload: 50 site scans, automated code reviews, 10 daily summaries, and schema validation calls.
    • Monthly Token Consumption: ~15M input tokens (cached), 2M uncached input, 3.5M output tokens on Grok-4.3.
    • Total Monthly Cost: $14.25 / month.
    • 15M cached x $0.20 + 2M uncached x $1.25 + 3.5M output x $2.50 = $3.00 + $2.50 + $8.75 = $14.25.

    Scenario B: Real-Time Customer Intake & Dispatch Voice Agent

    • Daily Workload: 30 inbound phone calls (avg 3.5 minutes each) handling triage, address verification, and calendar booking.
    • Monthly Minutes: ~3,150 audio minutes.
    • Total Monthly Cost: $252.00 / month in audio charges (vs. $3,200+/month for full-time 24/7 human dispatch).
    • 3,150 minutes x $0.08 = $252.00, plus $0.004 per text input the agent generates.

    Scenario C: Large Multi-Repo Deep Search & Code Synthesis

    • Daily Workload: High-frequency reasoning and code refactoring across 20+ microservices in Cursor.
    • Monthly Token Consumption: 80M input tokens on Grok-4.7 with prompt caching enabled.
    • At the Grok-4.7 rates in the table: all cached, 80 x $0.50 = $40; all uncached, 80 x $2.00 = $160; a 50/50 mix = $100 in input charges, before output tokens.

    How to apply the Grok cache rate

    1. Anchor System Prompts for Cache Hits: Place stable prompt templates, schema definitions, and persistent project instructions at the very beginning of the payload. Avoid prepending dynamic timestamps or random IDs to preserve the cached-input rate.
    2. Model Routing (build-0.1 for Scaffolding, 4.7 for Reasoning): Use lightweight models like Grok-build-0.1 ($1.00/$2.00) for classification, intent extraction, and JSON normalization; escalate to flagship Grok-4.7 only for deep logical synthesis or multi-file architecture plans.
    3. Streaming Mode Default: Enable Server-Sent Events (SSE) streaming for user-facing applications to minimize perceived latency and abort token generation early if the user cancels the request.

    Grok API pricing questions

    How much does the Grok API cost?

    Current published rates: Grok-4.7 is $2.00 input and $6.00 output per 1M tokens, Grok-4.6 the same, Grok-4.5 $2.00 / $6.00, Grok-4.3 and the Grok-4.20 family $1.25 / $2.50, and Grok-build-0.1 $1.00 / $2.00. Cached input is $0.50, $0.50, $0.30, $0.20, and $0.20 respectively. Grok Voice speech-to-speech is a flat $0.08 per minute plus $0.004 per text input.

    Is there a free tier for the Grok API?

    xAI publishes no free tier and no standing new-account credit. Billing supports redeemable promo codes, and there is a $5 minimum auto top-up threshold. Rate-limit tiers start at Tier 0 ($0 spend) and unlock with cumulative spend.

    How much does Grok prompt caching change the input price?

    Cache hits bill at a per-model cached-input rate: $0.50 instead of $2.00 on Grok-4.7/4.6, $0.30 instead of $2.00 on Grok-4.5, $0.20 instead of $1.25 on Grok-4.3/4.20, and $0.20 instead of $1.00 on Grok-build-0.1 — roughly 75-85% below standard input depending on model. xAI publishes no single site-wide discount figure.

    What are the Grok API rate limits?

    Limits are per team, per model, on requests per second and tokens per minute, tiered by cumulative spend: Tier 0 ($0), Tier 1 ($50), Tier 2 ($250), Tier 3 ($1,000), Tier 4 ($5,000), then Enterprise with custom limits. The catalog lists Grok-4.7 at 150 RPS / 50M TPM; limits rise as tiers unlock.

    Conclusion: The Operational Verdict

    At xAI’s current published rates, Grok-4.7 is $2.00 / $6.00 per 1M tokens, Grok-4.5 $2.00 / $6.00, Grok-4.3 and Grok-4.20 $1.25 / $2.50, Grok-build-0.1 $1.00 / $2.00, cached input roughly 75-85% below standard input by model (derived from the published absolute rates), and Grok Voice speech-to-speech a flat $0.08 per minute plus $0.004 per text input. Grok-3 is retired and redirects to Grok-4.3. Page updated October 2, 2026.

    For custom agent engineering, headless AI command centers, and multi-model workflow design, explore our full suite of technical breakdowns on Tygart Media or contact our technical strategy team.

    Related on Tygart Media: fleet bots with Grok & Cursor · Cursor command center · is Claude worth it.

  • Restoration Operations Kit — Claude Edition

    Restoration Operations Kit — Claude Edition

    Direct Answer / System Definition

    The Restoration Operations Kit — Claude Edition by Tygart Media is an 8-skill specialized AI plugin that transforms Claude into an operational copilot for property restoration contractors. Available for a one-time payment of $197 on Square, the kit automatically customizes Claude to a contractor’s specific shop profile, generating accurate job scopes, dehumidifier sizing calculations, adjuster supplement justifications, and IICRC-compliant standard operating procedures.

    Restoration Operations Kit — Claude Edition

    $197

    All 8 Claude skills + plugin files — delivered digitally via email within 24 hours.

    Buy on Square — $197 →

    Secure checkout via Square — all major cards accepted

    Developed by Will Tygart and the Tygart Media team, this is the AI operational companion to the Complete Restoration Operations Kit. A restoration owner attaches it to their own Claude instance. It interviews the owner, writes a tailored company-profile.md, and makes every subsequent prompt speak your shop’s exact voice, pricing rules, and crew standards.

    Overview: What the Claude Edition Does

    Three cards for field SOPs, owner prompts, and KPI rhythm in an operations kit
    Claude skills that turn field chaos into repeatable runs.

    An 8-skill Claude plugin. Install it into Claude Code, the Claude Desktop app, or Cowork. A setup skill runs a five-minute interview, writes a company-profile.md, and every other skill reads it. SOPs, KPI targets, claims drafts, and onboarding plans come out in your voice, not a generic restoration voice.

    You need any Claude that supports Skills / Plugins. If you can chat with Claude and run a /command, you are good.

    The 8 Operational Claude Skills

    Eight numbered skill cards from dispatch through review
    Eight skills on the bench — pick the one that matches the job.
    Skill Identifier Natural Language Trigger Operational Output & Deliverable
    restoration-setup “Set up the kit” Guided 5-minute setup creating your persistent company-profile.md.
    job-intake-assistant “We just got a water call” New-loss FNOL triage, Cat/Class classification, safety protocols, paste-ready summary.
    equipment-advisor “How many air movers for this room?” IICRC S500 air mover / dehu sizing, psychrometric target calculation.
    sop-generator “Write an SOP for mold containment” Custom step-by-step SOPs matching your company’s certified equipment and protocols.
    claims-assistant “Draft a follow-up to the adjuster” Adjuster emails, Xactimate supplement justifications, line-item dispute letters.
    kpi-coach “Here are my numbers this month” 12 critical KPIs trended (e.g. gross margins, cycle time) naming top profit leaks.
    crew-onboarding-builder “Onboard a new tech” Role-based 30/60/90-day training roadmap and IICRC certification tracking.
    iicrc-protocol-lookup “What does S500 say about Cat 3?” Plain-English verified lookup of S500, S520, S700, and S540 procedural standards.

    How to Install the Skills

    Option A: Claude Plugin (Recommended)

    1. Save the restoration-kit folder on your local computer.
    2. In Claude, run:
      /plugin marketplace add /path/to/restoration-kit
      /plugin install restoration-kit@profit-detective

      Replace the path with wherever you saved the folder. On Desktop and Cowork, type the same /plugin commands in the chat.

    3. Start setup:
      /restoration-kit:restoration-setup

    Option B: Personal Skills Folder (Simplest, No Plugin)

    Copy each folder inside skills/ into your personal skills directory at ~/.claude/skills/. On Windows that is C:\Users\<you>\.claude\skills\. Then prompt Claude: “run restoration setup.”

    First Run Walkthrough

    Run restoration-setup first. It conducts a guided 5-minute interview regarding your shop’s services, team size, service area, and rate sheet. It saves a persistent company-profile.md file that every subsequent skill reads so the generated answers sound like your actual shop.

    Using the Kit Day to Day

    Four-step flow: open skill, paste job facts, review draft, send or file
    Same four moves. Different job. Less reinventing at 11pm.

    You do not need to memorize skill names. Just prompt Claude naturally:

    • “We just got a fire call on Oak St” → invokes job intake
    • “Size the equipment for a 15×20 Class 3” → invokes equipment sizing
    • “Write our FNOL intake SOP” → invokes SOP generator
    • “My gross margin slipped to 41%. What’s going on?” → invokes KPI coach
    • “Draft a supplement justification for the Alvarez claim” → invokes claims assistant
    • “Build an onboarding plan for a new crew chief” → invokes crew onboarding
    • “What PPE for Condition 3 mold?” → invokes IICRC lookup

    What the Delivered Zip File Contains

    • .claude-plugin/plugin.json and marketplace.json
    • skills/ folder containing all 8 complete skills
    • README.md with full step-by-step setup walkthrough

    Outputs from each skill are formatted to paste into the matching Notion template in the Complete Restoration Operations Kit ($97). The Notion side serves as the operational system of record, while the Claude Edition acts as the intelligence layer.

    Frequently Asked Questions

    What is the Restoration Operations Kit — Claude Edition?

    The Restoration Operations Kit — Claude Edition by Tygart Media is an 8-skill AI plugin for Claude Desktop, Claude Code, and Cowork that automates restoration intake, equipment sizing, insurance claims drafting, SOP generation, and KPI tracking in the contractor’s exact voice.

    What is included in the $197 purchase on Square?

    The $197 package includes all 8 packaged Claude skills, the plugin.json manifest, marketplace configuration files, installation documentation, and prompt architecture delivered digitally via email.

    How does this kit work with the Complete Restoration Operations Kit?

    The Claude Edition acts as the operational AI copilot: its outputs are pre-formatted to paste seamlessly into the corresponding Notion databases (Job Tracker Pro, Claims Command Center, KPI Dashboard) of the Complete Restoration Operations Kit.

    What Claude plan is required to run the skills?

    It runs on any Claude interface that supports skills or plugins, including Claude Desktop, Claude Code CLI, and Cowork. Personal skill mode can also be loaded directly into ~/.claude/skills/.

    Is there a refund policy?

    Because this is a digital product, all sales are final. If you have a problem with your purchase, email will@tygartmedia.com and we will sort it out.

    How is this delivered?

    Within 24 hours of purchase via email from will@tygartmedia.com. You will receive your download link — Notion duplicate link, skill file, or both depending on the product.

    $197

    Packaged zip delivered by email within 24 hours via Square.

    Buy on Square — $197 →

    Related Operations & Systems

    Companion Systems: Complete Restoration Operations Kit ($97) · Complete Restoration Operating System ($597) · Leadership Toolkit — Claude Edition ($197).

    Author: Will Tygart, Founder of Tygart Media — AI operating systems designed specifically for property restoration contractors.

    END_