Most conversations about AI crawlability focus on one file: llms.txt. But if you look at what Anthropic, Vercel, and LangGraph actually ship – and what GEO crawler research found AI agents fetching most – the file that matters more is its companion: llms-full.txt.
Here’s the practical reality: llms.txt is the map. llms-full.txt is the territory. And in 2026, the agents that matter for citation traffic are fetching the territory.
The Full File Family You Probably Don’t Know About
The original llms.txt proposal – published by Jeremy Howard in September 2024 – defined one file. Implementers built the rest. The complete family as of mid-2026 is four files, but most sites only need two:
File
What’s in it
When to use
/llms.txt
Curated index – H1, summary, link sections
Always. The orientation layer.
/llms-full.txt
Full content of every linked page, concatenated as Markdown
When you want a model to deep-ingest your docs in a single fetch
/llms-ctx.txt
Pre-expanded context without URLs
FastHTML-style implementations
/llms-ctx-full.txt
Pre-expanded context with URLs preserved
Same, but URL-aware
The pattern that works – and the one Anthropic, Vercel, and LangGraph all run – is the index + export pair: llms.txt for orientation, llms-full.txt for deep ingestion.
Why llms-full.txt Gets Crawled More
Why llms-full.txt gets crawled more.
GEO researchers analyzing AI crawler behavior – including work cited by Profound – have noted that agents from Microsoft, OpenAI, and others tend to fetch llms-full.txt more frequently than llms.txt when both are present. The working explanation is structural: when a file contains the full content, it removes one retrieval step. An agent that fetches llms-full.txt gets everything it needs in a single HTTP request instead of fetching the index, parsing the links, then fetching each linked page individually. This is consistent with how developer documentation platforms like Mintlify describe the behavior of IDE agents operating under tight latency budgets.
For IDE agents (Cursor, Continue, Cline) and MCP integrations, this is even more pronounced. These tools are operating under tight context windows and latency budgets. A single fetch that returns a clean Markdown blob of your entire docs is structurally preferable to a multi-step crawl.
The implication: if you’ve shipped llms.txt but not llms-full.txt, you’ve done half the job.
How to Build llms-full.txt
How to build llms-full.txt.
The construction logic is simple: take every URL in your llms.txt, fetch each page, strip HTML to Markdown, and concatenate. In practice, most sites do this in their build pipeline.
Here’s the minimal Node.js pattern:
const fs = require('fs');
const fetch = require('node-fetch');
const TurndownService = require('turndown');
const turndown = new TurndownService();
async function buildLlmsFullTxt(llmsIndexPath, outputPath) {
const index = fs.readFileSync(llmsIndexPath, 'utf8');
const urlRegex = /\[.*?\]\((https?:\/\/[^\)]+)\)/g;
const urls = [...index.matchAll(urlRegex)].map(m => m[1]);
let output = '';
for (const url of urls) {
const res = await fetch(url);
const html = await res.text();
const markdown = turndown.turndown(html);
output += \n\n---\n# Source: \n\n;
}
fs.writeFileSync(outputPath, output);
console.log(Built llms-full.txt: pages, chars);
}
buildLlmsFullTxt('./public/llms.txt', './public/llms-full.txt');
One constraint to manage: keep llms-full.txt under roughly 200,000 tokens (about 150K words, around 700KB). That’s the threshold where most models can ingest the file in a single context window. If your docs are larger, segment by product or language the way Supabase does – llms-full-api.txt, llms-full-guides.txt – and list the segmented files in your main llms.txt.
The 2026 robots.txt Stack That Completes the Picture
The 2026 robots.txt stack that completes the picture.
Shipping llms.txt and llms-full.txt is the visibility layer. The access-control layer is robots.txt – and it changed significantly in Q2 2026.
The key development: Anthropic split its crawler into two separate user-agents. ClaudeBot is the training scraper (high bandwidth, no citation value – block it). Claude-Web is the live-retrieval agent that fetches pages to answer Claude.ai user queries in real time (allow it, because it drives citation traffic). Brands that blanket-block “all Anthropic crawlers” lose Claude citations entirely.
Meta also shipped two active training scrapers in March 2026 – FacebookBot and Meta-ExternalAgent – at GPTBot-level crawl volume. Most sites have no rules for them yet.
One important caveat on robots.txt enforcement: aggressive training scrapers often ignore the file or spoof their user-agents. The robots.txt rules signal intent and work for compliant bots; a WAF rule at the edge is the only deterministic block for non-compliant crawlers.
The Honest State of the Technology
The SERanking study of 300,000 domains (November 2025) found no measurable correlation between having llms.txt and being cited by ChatGPT, Claude, Gemini, or Perplexity. Google’s John Mueller compared the file to the deprecated keywords meta tag – something site owners declare but that search systems derive from the content itself.
None of that means you shouldn’t ship both files. The cost is low, the optionality is real, and the IDE-agent ecosystem (Cursor, Continue, Cline) does actively use llms.txt. But the robots.txt work is the lever that moves outcomes today. The llms.txt + llms-full.txt pair is infrastructure investment – you want to be correct when major LLM providers start honoring it, and building the build pipeline now costs far less than retrofitting it later.
The practical sequence for a site that hasn’t done this yet:
Update robots.txt first. Add the Q2 2026 user-agent rules above. This takes twenty minutes and immediately affects how training scrapers treat your content.
Ship llms.txt. Curated index, 20-50 priority pages, one-sentence description per link, sections in priority order.
Build llms-full.txt. Concatenated Markdown of every linked page, under 200K tokens. Run it in your build pipeline so it stays current.
Verify both files are served correctly.curl -I https://yoursite.com/llms.txt should return 200 with Content-Type: text/plain. A 404 on either file is the most common implementation error.
Add an access-log check. Once per month, grep your logs for requests to /llms.txt and /llms-full.txt by user-agent. You want to see live-retrieval agents (Claude-Web, OAI-SearchBot, PerplexityBot) in the results – not just training scrapers.
The goal isn’t to optimize for a standard that isn’t fully adopted yet. It’s to build the infrastructure correctly now, while the field is still forming, so that adoption changes work in your favor rather than requiring catch-up.
What is the difference between llms.txt and llms-full.txt?
llms.txt is a curated index — an H1, a summary, and link sections that orient an AI agent to your site. llms-full.txt is the full content of every linked page concatenated as Markdown, so an agent can deep-ingest your documentation in a single fetch. The index is the map; the full file is the territory.
Why do AI agents crawl llms-full.txt more often than llms.txt?
Fetching llms-full.txt removes a retrieval step: the agent gets everything in one HTTP request instead of fetching the index, parsing links, and fetching each page individually. For IDE agents like Cursor, Continue, and Cline operating under tight latency and context budgets, a single clean Markdown blob is structurally preferable to a multi-step crawl.
How big should llms-full.txt be?
Keep it under roughly 200,000 tokens (about 150K words, around 700KB) so most models can ingest it in a single context window. If your docs are larger, segment by product or language — for example llms-full-api.txt and llms-full-guides.txt — and list the segmented files in your main llms.txt.
Does having llms.txt actually improve AI citations?
Not measurably on its own. A November 2025 SERanking study of 300,000 domains found no correlation between having llms.txt and being cited by ChatGPT, Claude, Gemini, or Perplexity, and Google’s John Mueller compared it to the deprecated keywords meta tag. The lever that moves outcomes today is robots.txt configuration; llms.txt and llms-full.txt are low-cost infrastructure for when adoption grows.
Which AI crawlers should I allow in robots.txt in 2026?
Allow live-retrieval agents that drive citation traffic — Claude-Web, OAI-SearchBot, ChatGPT-User, anthropic-ai, and PerplexityBot. Block high-bandwidth training scrapers with no referral value such as GPTBot, CCBot, ClaudeBot, FacebookBot, and Meta-ExternalAgent, and opt out of Google-Extended to skip Gemini training while keeping Search indexing intact.
Most “GEO” advice is recycled SEO with the word “AI” pasted on top. This guide is different. It describes what actually happens when Microsoft Copilot, Bing’s AI answers, and Google’s AI Overviews build a response and decide whose page to cite — based on running content sites that get cited tens of thousands of times a month. The short version: AI engines do not cite the page that ranks #1 for a head term. They cite the page that most directly answers the specific sub-question the model is grounding on. That distinction changes everything about what you should write.
How grounding actually works (the part nobody explains)
How grounding actually works.
When you ask Copilot or Bing’s AI a question, the model does not answer from memory. It runs a retrieval step called grounding: it rewrites your question into one or more search queries, fetches a handful of live web results, reads them, and composes an answer with inline citations pointing back at the pages it used. Google’s AI Overviews work the same way with a technique it calls “query fan-out” — one user question becomes many narrower synthetic queries.
Two things follow directly from this mechanism:
The model is not searching for your keyword. It is searching for the answer to a decomposed sub-question. A user who asks “what’s the best way to instantly index a new page” triggers grounding queries like “IndexNow API endpoint”, “submit URL to Bing programmatically”, and “IndexNow key file location”. The page that wins is the one that answers those narrow strings, not the one optimized for “indexing tips”.
Citations are extracted at the passage level, not the page level. The model lifts the specific sentence or table that answers the sub-question. If your answer is buried under 600 words of preamble, it loses to a page that states the fact in the first line under a matching heading.
This is why a niche, specific page routinely out-cites a high-authority generalist. The generalist ranks; the specialist gets quoted.
Why operational and comparison pages win over head terms
Across real citation data, the pages that get pulled into AI answers cluster into three shapes. None of them are “ultimate guide to X”.
1. Operational pages with real commands, configs, and error messages
When someone asks an AI assistant “how do I fix [specific error]” or “what’s the exact command to do X”, the model needs a page that contains the literal command, the literal config, or the literal error string. Generic advice cannot be cited because there is nothing concrete to quote. A page that says:
curl "https://www.bing.com/indexnow?url=https://example.com/new-page/&key=YOUR_KEY"
# 200 = received (not "indexed"), 422 = URL/key mismatch, 429 = too many submits
…is citation gold, because the model can extract that block verbatim and the user can act on it. The error-code annotations matter: questions about failures (“IndexNow 422”, “why am I getting 429”) are high-intent and low-competition, and a page that names the exact codes owns them.
2. Comparison pages (“X vs Y”)
“Which is better, X or Y” is one of the most common shapes of AI query, and comparison content is structurally easy to cite because it maps cleanly to a decision. If you maintain honest, current head-to-head pages, you become the default source the model reaches for when a user is choosing between tools. This is exactly why we keep dedicated comparison pages like Claude Code vs Cursor and Claude Code vs Codex — they answer a decision the model is constantly being asked to make, and a table of differences is trivially quotable.
3. Fresh, dated pages on fast-moving topics
For anything that changes — pricing, model versions, API limits, feature availability — grounding strongly favors recency. The model would rather cite a page dated this month than an “authoritative” page from two years ago that might be wrong. A visible “Last verified” date and a real publish/update timestamp are not decoration; they are a relevance signal the retrieval layer reads.
The losing move is chasing broad head terms. “Best AI coding assistant” is saturated, generic, and rarely the literal grounding query. The winning move is to own the long, specific, operational and comparison strings that the fan-out actually generates.
IndexNow: how to get cited the same day you publish
IndexNow — cited the same day you publish.
Grounding can only cite pages the engine knows about. The bottleneck for new content is crawl latency — and IndexNow collapses it. IndexNow is an open protocol (backed by Microsoft Bing and Yandex) that lets you push a URL to the index the instant you publish, instead of waiting for a crawler to wander by.
Setup is two steps:
Host a key file. Generate a key of 8-128 hex characters and place it at your site root as a UTF-8 text file named {key}.txt containing exactly that key. Example: https://example.com/daa44a2c....txt. This proves you own the host.
A 200 means the endpoint received your URL (not that it is indexed yet). Submitting to api.indexnow.org shares the ping with all participating engines, so you do not need to hit Bing and Yandex separately. Most WordPress SEO plugins (Rank Math, Yoast, SEOPress) have IndexNow built in — turn it on and it fires automatically on every publish and update. The practical payoff: pages can enter Bing’s crawl queue within hours, which means they are eligible to be grounded and cited the same day, not next week.
One caveat worth stating plainly: IndexNow accelerates indexing, which is a precondition for citation. It does not force a citation. You still need the page to be the best answer to the sub-question. But for fresh, time-sensitive content, same-day indexing is often the difference between getting cited while the topic is hot and showing up after the conversation has moved on.
How to actually measure your AI citations
For a long time AI citations were invisible — you could see referral clicks in analytics but not the citations themselves (most AI answers are zero-click). That changed. As of February 2026, Bing Webmaster Tools ships an AI Performance report (public preview) that shows when your pages are cited across Microsoft Copilot, Bing’s AI answers, and partner surfaces. It is the first direct, free window into AI citation behavior, and you should be reading it weekly.
The four metrics that matter:
Total citations — how many times your site was cited as a source in AI answers over the period.
Average cited pages — the daily average count of unique URLs from your site that got referenced. This tells you whether citations are concentrated on one page or spread across the site.
Grounding queries — sample query phrases the AI used to retrieve and cite you. This is the single most actionable field in the report. It is a literal list of the sub-questions you are winning, which tells you exactly which operational/comparison angles to expand next.
Page-level citation activity — citations by URL, so you can see which pages are doing the work.
Two limitations to keep in mind so you read the data honestly: the report does not show click data (you see citations, not visits from them), and it aggregates Copilot with Bing summaries, so you cannot isolate one surface from the other. For Google’s AI Overviews there is still no equivalent citation dashboard — the closest proxy is watching impressions and referral patterns in GA4 and Search Console, plus spot-checking your target queries by hand.
The workflow that works: pull the grounding-queries list, find the patterns, and feed them straight back into your content plan. If you are getting cited for “claude mcp setup” variants, that is a signal to deepen pages like the Claude MCP setup guide and adjacent operational walkthroughs, not to chase a new head term.
A repeatable checklist for citation-optimized pages
Checklist for citation-optimized pages.
Everything above reduces to a build pattern. For any page you want AI engines to cite:
Lead with the answer. Put a short, factual, quotable answer in the first 1-2 sentences under each heading. Assume the model reads only that passage.
Use question-shaped headings. H2s and H3s that mirror real queries (“How does IndexNow work?”, “How do I measure AI citations?”) match the grounding query and give the extractor a clean anchor.
Be specific and operational. Real commands, real config, real numbers, real error codes and fixes. Concrete text is extractable; vague advice is not.
Add a visible FAQ near the end. Plain question/answer pairs are the single most citation-friendly format, because each pair is a self-contained answer to a discrete sub-question. You do not need JSON-LD schema for this to work — visible Q&A text is what the model reads.
Date it and keep it current. A “Last verified” line plus genuine updates on fast-moving topics buys you the recency edge in grounding.
Push it with IndexNow so it is indexable the same day, then watch the AI Performance report to see which sub-questions it wins.
If you want the larger system this fits into — the full toolchain for operating as an AI-first publisher, from MCP servers to publishing pipelines — start with the AI operator’s stack.
FAQ
Do AI engines cite the page that ranks #1 on Google?
Not reliably. AI engines run their own grounding retrieval and cite the page that most directly answers the specific decomposed sub-question, which is often a niche, operational page rather than the head-term winner. Ranking helps your page be discoverable, but the citation goes to whichever passage best answers the exact grounding query.
What is grounding in AI search?
Grounding is the retrieval step where an AI assistant rewrites your question into search queries, fetches live web pages, reads them, and builds an answer with inline citations to those pages. It is why current, specific pages can get cited even by a model whose training data predates them.
Does IndexNow guarantee my page will be cited by AI?
No. IndexNow guarantees fast indexing, which is a precondition for being cited. The page still has to be the best, most specific answer to the sub-question the model is grounding on. Think of IndexNow as removing the crawl-latency excuse, not as buying a citation.
How do I measure how often AI cites my site?
Use the AI Performance report in Bing Webmaster Tools (public preview since February 2026). It shows total citations, average cited pages per day, sample grounding queries, and citation counts by URL across Microsoft Copilot and Bing AI answers. It does not yet show click-through from those citations, and there is no equivalent dashboard for Google AI Overviews.
Do I need JSON-LD or schema markup to get cited?
No. Citation extraction works on visible, well-structured text — question-shaped headings, short factual answers, and a plain visible FAQ. Schema can help search features generally, but it is not required for AI grounding to read and quote your page.
What kind of pages get cited most?
Three shapes dominate: operational pages with real commands, configs, and error fixes; comparison pages that resolve a “X vs Y” decision; and fresh, dated pages on fast-moving topics like pricing and model versions. Broad head-term content tends to get skipped because it rarely matches the literal grounding query and offers nothing concrete to quote.
All fall, Microsoft has been selling one idea: the future is the AI PC — a Copilot+ machine with a dedicated neural chip (an NPU), Recall, Click to Do, a thousand dollars and up, and your old laptop need not apply.
I had a $400 budget laptop on my desk — an AMD Ryzen 5 7520U, 16 GB of RAM, no NPU — and a hunch that the whole framing was backwards. The AI-first laptop was never about the chip. It’s about architecture.
A few hours later, that $400 laptop had a private AI brain, voice control, and a control panel I run from my phone. On the things that actually matter for operating a machine, it does more than the Copilot+ PC it’s supposedly too cheap to be. Here’s the exact build.
The thesis: AI-first is architecture, not a chip
AI-first is architecture, not a chip.
The trick is to stop asking your laptop to be the supercomputer. Split the job:
The brain lives in the cloud. The heavy reasoning runs on a frontier model (I use Claude) with effectively unlimited horsepower. No NPU on Earth competes with that.
The body lives on your laptop. Your machine becomes the always-on hands: it holds your private data, runs small models locally for anything sensitive, and executes the actions the brain decides on.
An NPU optimizes a handful of on-device Windows features. Architecture gives you an actual operator. Guess which one you feel every day.
Step 0 — Make it always-on
An operator rig is a little server, and servers don’t nap. My laptop kept sleeping and killing background jobs, so the first move was to take that off the table (while plugged in):
Screen never blanks, never sleeps, and it keeps running with the lid closed — while still sleeping on battery as a safety. Now it’s a real always-on host.
Step 1 — A private AI brain that lives on the laptop
A private AI brain that lives on the laptop.
The local engine is Ollama; the chat interface is open-webui (running in Docker). If you want the multi-agent version of this idea, I’ve also written up building a free AI agent army with Ollama and Claude. The only thing standing between me and a private, offline ChatGPT was one wrong setting — open-webui was pointed at a dead address. The fix was to aim it at the host:
The proof: a 3-billion-parameter model (Llama 3.2) introduced itself in about 10 seconds at ~12 tokens/second — on the CPU, no NPU, no discrete GPU. Fast enough for real Q&A, drafting, and summaries. Seven models sit ready on disk, and the whole thing is reachable from my phone over a private network.
Everything here runs offline. For anything I don’t want leaving the machine, that’s the entire point.
Step 2 — Voice that never leaves the machine
Voice that never leaves the machine.
A local Whisper speech-to-text container (OpenAI-compatible API) became a push-to-talk dictation tool: hold a key, talk, release, and the text drops into whatever app is focused. I verified the pipeline without even touching the mic — Windows text-to-speech generated a clip, the local Whisper transcribed it, and it round-tripped clean:
Spoken: “Testing one two three. This is the private local transcription engine.” Whisper heard: “Testing 1-2-3. This is the private local transcription engine.”
Windows has built-in dictation (Win+H) and Copilot voice too — but those ship your audio to the cloud. The local version does the same job, and your voice never leaves the laptop.
Step 3 — Turn your phone into the control panel
Using Tailscale (a private mesh network), every service on the laptop is reachable from my phone — without exposing anything to the public internet. I added a tiny web page (one small nginx container) as a mobile operator console: one tap to the local AI, automations, status, and finance dashboards. Pin it to the home screen and the laptop is in your pocket.
The honest scoreboard vs. a Copilot+ PC
Capability
Copilot+ PC ($1,000+)
This $400 laptop
Private AI running on the device
Limited (small NPU models)
✅ Full Ollama stack, 7 models
An AI that operates the machine
❌
✅ Runs commands, edits files, fixes things
Private, offline voice dictation
❌ (cloud)
✅ Local Whisper
Phone control panel
❌
✅ Tailscale operator console
Recall / Click to Do / Cocreator
✅ (needs the NPU)
❌
Screenshots everything you do
⚠️ Recall does, by design
✅ No — nothing is recorded
I’m being fair: the NPU-only features are genuinely off the table on cheap hardware. But for operating your computer — and for privacy — the architecture beats the chip.
Why this matters more than it looks
The quiet headline isn’t “I saved money.” It’s where the data lives. Microsoft’s flagship AI-PC feature, Recall, works by screenshotting everything you do. This build does the opposite: the sensitive payload stays on your machine, and the cloud is used only for the heavy thinking that doesn’t need your private files.
That’s not just a hobbyist’s preference. It’s the exact requirement for anyone in a regulated field — healthcare, legal, finance — who can’t send client data to a third party but still wants real AI leverage. The cheap laptop isn’t the story. The architecture is.
Frequently asked questions
Do I need a Copilot+ PC or an NPU to run local AI?
No. Any laptop with around 16 GB of RAM and a modern CPU can run small local models. An NPU accelerates certain Windows features but is not required for Ollama or local chat.
Is local AI actually private?
Yes. With Ollama, the model runs on your own machine and works with no internet connection — nothing is sent to a cloud service.
What is the difference between Ollama and open-webui?
Ollama is the engine that runs the models. open-webui is the friendly chat interface that sits in front of it.
How fast is a local model on a budget laptop?
On a CPU-only AMD Ryzen 5 with 16 GB of RAM, a 3-billion-parameter model answered at roughly 12 tokens per second — fine for quick questions, drafting, and summaries. Larger models run slower.
Can I use it from my phone?
Yes. Over a private Tailscale network you can reach your laptop’s AI and tools from your phone without exposing anything to the public internet.
Is this better than a Copilot+ PC?
For operating your machine and for privacy, this setup does more. For NPU-specific Windows features like Recall and Click to Do, a Copilot+ PC is required.
Want this on your machine?
Tygart Media builds privacy-first, local-AI operator setups — especially for teams in regulated industries that need real AI leverage without sending data to the cloud. Reach out and we’ll scope it to your hardware.
What Claude in Chrome can and can’t do on LinkedIn
What Claude in Chrome can and can’t do on LinkedIn.
Task
Verdict
Notes
Summarize a profile
✅ Safe and useful
Read-only, no automation signal
Draft a personalized DM
✅ Safe and useful
You review and send manually
Research a company page
✅ Safe and useful
Read-only extraction
Summarize a post or thread
✅ Safe and useful
Read-only, no interaction
Auto-post to your feed
❌ High risk
Violates ToS, triggers automation detection
Auto-connect with multiple people
❌ High risk
Account restriction risk
Bulk message sending
❌ High risk
Spam detection, potential ban
The Claude for Chrome extension lets Claude see and act inside your browser. The obvious temptation is to point it at LinkedIn and have it post for you. Do not do that. Here is what the extension is genuinely useful for on a professional network – and the one job you should never hand it.
What to avoid: automated feed posting
What to avoid — automated feed posting.
Driving the browser to auto-post feed content is a high-risk move. Professional networks actively detect automation, it violates their terms of service, and it can get an account throttled or suspended. If you want scheduled feed posts, use a social scheduler’s official API – that is the supported, durable path, and the one that will not get your account flagged. The browser is an assistant, not a posting robot.
What it is actually good for
1. Paste-assist for long-form Articles
This is the real opportunity. Social schedulers – and every third-party tool – can only push short feed posts through the official API. Native long-form Articles and Newsletters have no public publishing endpoint, so they stay a manual copy-paste. That matters because AI engines cite long-form Articles far more often than short posts, by a wide margin. The most citation-valuable format is the one no tool can automate. That is exactly where an in-browser assistant earns its place: with you in the loop, it can help move a finished, formatted draft into the Article composer and tidy the formatting – turning a tedious manual paste into a guided one.
2. Multi-account navigation
If you operate a personal profile plus several company pages, the extension can help you move between already-authenticated sessions and keep track of which identity you are acting as – reducing the “posted from the wrong account” mistakes that come with juggling many pages by hand.
3. Research, review, and drafting
Reading a profile and summarizing it, scanning a feed for the day’s relevant threads, or drafting a thoughtful comment for your approval are all squarely in bounds. The assistant prepares; you decide and click.
How to do it safely
How to do it safely.
Keep a human in the loop on anything that publishes or sends – review before you submit.
Never bulk-send connection requests, messages, or comments. That is the behavior detectors look for.
Use the official scheduler API for anything recurring; reserve the browser for the manual, assistive steps.
Treat the extension as read-and-prepare by default, act-and-publish only with your explicit click.
Frequently asked questions
Can Claude auto-post to LinkedIn for me?
Not safely, and you should not try. Use a social scheduler’s API for feed posts. The browser extension is for assistive, human-in-the-loop work – especially the long-form Articles that no API can publish.
Why can’t scheduling tools publish Articles or Newsletters?
Because the platform exposes no public API for them. Feed posts have an endpoint; long-form does not. That limitation is shared by every tool, which is why the manual paste persists.
Is browser automation against the rules?
Automated posting and bulk outreach generally violate the terms and risk the account. Assistive, human-approved use – drafting, summarizing, helping you paste – is the safe lane. When in doubt, keep a person on the trigger.
For the bigger picture of how this fits a full content operation, see The AI Operator’s Stack.
Frequently Asked Questions
What is the Claude for Chrome extension?
Claude for Chrome (Claude in Chrome) is a browser extension that lets Claude see and interact with the page currently open in your browser. It can read page content, summarize what’s visible, draft responses based on what it sees, and in some configurations take actions like clicking or filling forms — depending on what permissions are active.
Can I use Claude to automate LinkedIn posts?
You should not. Professional networks like LinkedIn actively detect browser automation, and auto-posting violates their Terms of Service. Using Claude in Chrome to drive automated feed posting can result in account throttling or permanent suspension. Claude is useful for drafting post content, but you should always review and publish manually.
What is Claude in Chrome actually useful for on LinkedIn?
Legitimate high-value uses include: summarizing a prospect’s profile before a sales call, researching a company page, drafting a personalized connection request or DM based on what you read on a profile, and summarizing a post or comment thread. All of these are read-and-assist operations that don’t trigger automation signals.
Does using Claude in Chrome on LinkedIn violate their terms of service?
Read-only operations (summarizing, researching, drafting) generally do not violate LinkedIn’s terms. Automated actions (clicking, posting, connecting, messaging at scale) do. The key distinction is whether Claude is taking actions on LinkedIn’s platform autonomously versus helping you draft content that you then review and submit yourself.
How is Claude in Chrome different from a LinkedIn scraper?
Claude in Chrome reads what’s visible on the page you have open — it is not a bulk scraper that crawls hundreds of profiles automatically. It operates within your active browser session, one page at a time, and does not bypass LinkedIn’s normal page rendering. A scraper typically makes API calls or headless browser requests at volume; Claude in Chrome is a single-session reading assistant.
What Claude model powers Claude in Chrome?
Claude in Chrome uses Anthropic’s Claude models — currently Claude Sonnet 4.6 is the primary model for browser interactions, balancing capability and speed. Anthropic may update the underlying model over time. You can check your current model in the extension settings.
Most “AI stack” articles hand you a list of tools. This one is about the wiring between them, because that is where the leverage lives. After running a multi-brand content operation end to end – research, writing, publishing, and distribution to a couple dozen destinations – one lesson keeps repeating: the tools are commodities, and the connective tissue is the moat. Here is the whole machine, and how the pieces talk to each other.
One machine, four jobs
One machine, four jobs.
The stack has four jobs: capture an idea, produce the content, remember everything, and distribute it where both people and AI engines will find it. Miss any one and the system stalls.
1. Intelligence and intake
The front door is an “AI as PR team” intake: you drop a raw thought, a link, or a voice memo, and the model turns it into the right shapes – an outline, a short post, a full brief. A lightweight signal scraper watches a professional network for the language practitioners actually use and feeds those angles back as prompts, so the writing starts from how people really talk instead of a blank page.
2. Production
Claude is the reasoning engine. A content pipeline turns a brief into a structured article; an image model generates the visuals; and a set of “beat desks” – small scheduled agents, each owning one topic – research, draft, quality-gate, and self-publish to WordPress through its REST API. Every desk has a freshness gate: if there is nothing genuinely new and sourceable, it skips the run rather than manufacture filler. A clean skip is a successful run.
3. Record and state
Notion is the control plane – the registries, the per-desk specs, the run logs, the system of record. The governing principle is load-bearing: the model is not the runtime. Claude supplies judgment; durable execution lives on schedulers and cloud jobs; Notion holds the state. Separate those three and the machine keeps running whether or not anyone is watching it.
4. Distribution and grounding
This is the layer most stacks forget, and the one that compounds. Publishing to your own site is half the job; the other half is getting that content into the indexes search engines and AI assistants actually read. Two moves do the heavy lifting. First, IndexNow pings the Bing index the moment anything changes – that is how new and updated content gets grounded fast instead of waiting on a crawl. Second, a social scheduler fans a tailored post out to a professional network – a personal profile plus company pages – drafted first for human approval, never blasted.
Here is the part worth internalizing: that professional network matters far more than its follower count suggests, because it is one of the most-cited domains in AI answers. Since it flows into the same index that feeds AI grounding, every post is also a citation asset. You are not chasing likes – you are seeding the corpus that AI engines quote back to the next person who asks.
The loop that compounds
The loop that compounds.
The layers are not a straight line; they form a loop. A researched social post is a compressed seed. Crack it open into a full article cluster – a core piece, audience-specific variants, an FAQ, schema, internal links – publish those, then queue the new URLs back to the scheduler as future posts. Social feeds the site; the site feeds social; both feed the grounding layer. Content you already made becomes the raw material for what you make next.
Why every layer optimizes for citation
Why every layer optimizes for citation.
AI engines do not cite broad overviews. They cite operational specifics, head-to-head comparisons, and fresh, dated facts. So the whole stack is tuned for that: specific over general, “this versus that” where it genuinely helps a reader decide, and same-day freshness on anything that changes. The pages that earn the most citations are the least glamorous – the exact limits, the real configuration, the honest comparison – because those are the answers nobody else keeps current.
The honest edges
This is maintained, not magic. Long-form articles on a professional network have no public API, so that step is a manual paste – and it happens to be the most citation-valuable format, which means the highest-value action is also the least automatable one. Auth tokens expire and quietly break distribution until someone notices. Account IDs drift, so you verify live before any bulk action. The wiring is powerful precisely because keeping it wired is real work.
No, but you need to be comfortable wiring tools together – connecting an API, editing a config file, reading a log. The reasoning model closes much of that gap, but the operator still has to understand how the pieces connect.
Why optimize for Bing and not just Google?
Because the AI assistants people increasingly ask their questions to are grounded substantially on the Bing index. Winning that index is how you get cited in AI answers – a different and faster game than ranking on a traditional results page.
Is the social distribution automated?
The drafting is. Publishing is draft-first: the system stages every post for a human to approve before it goes live. Automation writes; a person decides.
The short version: In Claude Code, the prompt that asks whether to “Always Allow” or “Allow Once” isn’t really about security. It’s a question about your own systems. If you keep choosing Always Allow, the work is recurring — go build the automaton. If it’s honestly Allow Once, it’s a one-off — let it go instead of trying to remember it.
I spend most of my day inside Claude Code, and a tiny piece of the interface has been living rent-free in my head. Every time the agent wants to run a command, edit a file, or hit an API, it stops and asks: Always Allow, or Allow Once?
On the surface that’s a permission prompt. Click the box, move on. But after the hundredth time, I started to notice the choice was telling me something about how I actually work — and where I was leaving time on the table.
“Always Allow” means: go build the automaton
Always Allow means: go build the automaton.
Always Allow vs Allow Once: quick reference
Always Allow vs Allow Once — quick reference.
Signal
Always Allow
Allow Once
Task type
Recurring, repeating work
One-off, situational
Right response
Build an automation
Let it go — don’t memorize it
Security posture
Persistent permission for that tool+action
Single-use, no persistent grant
What it reveals
A system worth building
An edge case not worth systemizing
Risk if overused
Broad standing permissions accumulate
Missed automation opportunity
Here’s the pattern. If I find myself reaching for Always Allow, it’s because I’ve seen this exact action before. I’ll see it again. I trust it enough to stop being asked.
That’s not a permission decision. That’s a build order.
If an action is safe, repeatable, and I do it constantly, the right move isn’t to keep approving it forever — it’s to take it out of the prompt entirely. Turn it into a tool. Wrap it in a script. Register it as a skill. Put it on a cron so it runs whether I’m at the desk or not. The “Always Allow” click is the moment the work earns its own piece of infrastructure.
Most people stop at the click. They grant the permission and feel productive because the friction went away. But friction that shows up every single day isn’t friction you should approve — it’s friction you should engineer out. Every “Always Allow” is a quiet little flag waving at you: this deserves to be an automaton.
“Allow Once” means: let it go on purpose
The other side is just as useful, and it’s the part people get wrong.
When the honest answer is Allow Once — this is a weird one-off, I’m not going to do it again — the temptation is to write it down. Save the command. Add it to a doc. File it away just in case it ever comes back.
Resist that. A one-off doesn’t deserve a permanent home in your memory or your system. The cost of storing it isn’t the disk space — it’s the upkeep. Every note you keep is something you now have to organize, search past, keep current, and trip over later. Knowledge you save but rarely touch quietly rots, and stale knowledge is worse than none.
The way I think about it: it’s more fit to sift through the dirt than to re-sift the knowledge. If a one-off ever does come back, re-deriving it from scratch is cheap — you dig through the dirt once and you’re done. But re-sifting a giant pile of “just in case” notes, over and over, every time you go looking for the thing you actually need? That’s the expensive part. Forgetting a one-off on purpose is a feature, not a failure.
Why re-deriving usually beats remembering
This is really a question of economics, and it’s the same math whether you’re managing an AI agent or your own head.
Storing knowledge has two costs people forget about: the cost to keep it accurate, and the cost to find the signal inside it later. A one-off has a low chance of ever being needed again, so the expected payoff of saving it is tiny — while the drag it adds to everything else you’ve stored is real and permanent. Recurring work is the opposite: high chance of reuse, so it’s worth paying once to encode it well and never think about it again.
So the rule of thumb falls out on its own:
Recurring → encode it. Build the tool, the skill, the cron. Pay once, reuse forever.
One-off → forget it on purpose. Do the thing, then let it go. If it ever comes back, dig it up fresh — it’ll be faster than you think.
The mistake is doing it backwards: hand-running the recurring stuff every day because you never built the automaton, while hoarding a graveyard of one-off notes you’ll never open again. That’s how you end up busy and buried at the same time.
How to act on the tell in Claude Code
How to act on the tell in Claude Code.
Next time that prompt pops up, treat it as a tiny decision point instead of a speed bump:
You reached for “Always Allow.” Stop for a second. Ask: what would it take to make this prompt never appear again? An orchestration step, a saved skill, a scheduled job, a hook? Put it on the list. The prompt just told you what to build next.
You reached for “Allow Once.” Do it, then genuinely drop it. Don’t screenshot it, don’t file it. Trust that if it matters, it’ll show up again — and the second sighting is your real signal to build.
You’re not sure. That’s fine — “Allow Once” is the safe default. Two or three “Allow Once” clicks for the same action is the universe telling you it was an “Always Allow” the whole time.
None of this is really about Claude Code. The tool just happens to put the decision right in front of you, every day, in a little box. Most systems make you guess where your time is leaking. This one points at it and asks you to choose. (It pairs well with knowing when to use Plan Mode and when to skip it — same instinct, a different prompt.)
Recurring work wants to become an automaton. One-off work wants to be forgotten. The prompt already knows which is which. The only question is whether you’re listening.
Frequently asked questions
What’s the difference between “Always Allow” and “Allow Once” in Claude Code?
“Allow Once” approves a single action one time; the next identical action prompts you again. “Always Allow” approves that action or pattern going forward, so Claude Code stops asking. Functionally, “Always Allow” is how you tell the tool an action is safe and routine.
Should I use “Always Allow” in Claude Code?
Use it when an action is safe, repeatable, and something you do often — but treat each “Always Allow” as a signal to eventually build that action into a tool, skill, hook, or scheduled job so it leaves the prompt entirely.
Is “Always Allow” a security risk?
It can be if you grant it to broad or destructive actions. Keep “Always Allow” for narrow, well-understood operations, and lean on “Allow Once” for anything unfamiliar, destructive, or outward-facing.
When should I turn a Claude Code action into an automation?
When you’ve granted — or wanted to grant — “Always Allow” for it. That’s the tell that the work is recurring, and recurring, trusted work is worth encoding once as a tool, skill, hook, or cron so you never approve it by hand again.
Why shouldn’t I save one-off commands?
Because storing knowledge has ongoing costs — keeping it accurate, and sifting past it to find what you actually need. A one-off has little chance of reuse, so it’s usually cheaper to re-derive it later than to maintain it forever.
What does “more fit to sift through the dirt than to re-sift the knowledge” mean?
It means re-deriving a rarely-needed answer from scratch — sifting the dirt once — is cheaper than maintaining and repeatedly searching a hoard of saved notes, which is re-sifting the knowledge every time. For one-offs, forgetting is the efficient choice.
Frequently Asked Questions
What does ‘Always Allow’ mean in Claude Code?
When Claude Code asks to run a tool or shell command, ‘Always Allow’ grants a persistent permission for that specific tool and action combination. Claude will not ask again for that combination in future sessions. ‘Allow Once’ grants permission only for the current request — Claude will ask again next time.
Is it safe to click Always Allow in Claude Code?
It depends on the action. Always Allow for read operations (reading files, querying a database) is generally low risk. Always Allow for write or execute operations (editing files, running shell commands) creates persistent permissions that compound over time. The best practice is to use Always Allow deliberately for actions you will genuinely repeat, and Allow Once for anything new or situational.
What is the deeper meaning of Always Allow vs Allow Once?
The choice is a signal about your own workflow. If you keep clicking Always Allow for the same action, that’s the system telling you the task is recurring and worth automating. If it’s genuinely Allow Once, the task is a one-off and you shouldn’t try to systemize it. The prompt is less about security and more about recognizing patterns in your own work.
How do I review or remove Always Allow permissions in Claude Code?
Run ‘claude permissions list’ to see what standing permissions you’ve granted. Use ‘claude permissions reset’ to clear them, or edit the .claude/settings.json file in your project directory to remove specific entries. Review these periodically — accumulated Always Allow grants are a common source of unexpected autonomous behavior.
Does Always Allow apply to a specific project or globally?
By default, permissions granted with Always Allow are scoped to the project where you granted them (stored in .claude/settings.json). If you use the –global flag, they apply across all projects. Be cautious with global Always Allow grants for write/execute operations — they persist across every codebase you open.
Google’s real superpower was never search or ads. It was the door home — and I learned that at 2 a.m., locked out of my own life.
I locked myself out of my own account a little after one in the morning. I don’t even remember what I needed in there — something small, something that could have waited until daylight. What I remember is the password field refusing me, then refusing me again, and the cold drop in my stomach when I realized the keys to a dozen other things lived behind that one rejection.
So I did what everyone does. I grabbed my phone. I tried the recovery email, which routed to an account I also couldn’t reach. I tried the text-message code. I tried the security questions, answered years ago with half-truths I’d invented and instantly forgotten. I worked the recovery flow like a man patting his pockets at a locked door, and somewhere in there it landed on me that I was negotiating — not with a hacker, not with a thief, but with the company that decides whether I am still me.
I got back in by morning. Relief, and then a second feeling underneath it that wouldn’t leave: that was the product. Not the search box. Not the ads. The way back in.
I build access layers for a living. Second brains. A life-ranking system I call the Compass. The structured record a business can’t operate without — the institutional memory that walks out the door when the wrong person quits. Continuity systems for my wife Stefani, so the things she needs are still there on the days her memory isn’t. I’d been filing all of it under content and tooling. That night I understood I’d been mislabeling my own work — and I understood something about Google that most people have backwards.
Two things, not one
Two things, not one — login vs search.
Here is the distinction that reorganized everything for me, and I want to be precise, because the sloppy version of this argument is wrong.
Search and ads are how Google makes money. That’s the business model, the value capture, the line on the income statement. Anyone who tells you access “beats” advertising is comparing a turnstile to a cash register. They don’t sit on the same axis.
But there are two things going on, and we only ever talk about one. Ads are how Google makes money. Access is why you can’t make Google stop. The login, the password manager, the “Sign in with Google” button, the recovery flow when you’re locked out — none of it earns a dollar directly. Google gives it all away. It exists to defend the surface where the money gets made.
And that’s the part people miss: the layer that earns nothing is the layer you can never leave. Attention is rented by the day — a better answer wins the next query, a better feed wins the next scroll. Access is owned by the year. So I won’t tell you access is more valuable than attention. I’ll tell you something narrower and more interesting: access is more durable. It is the layer with its hand on the master switch, and it shows up on the books as a cost center, a free feature, a help-desk ticket — which is exactly why nobody guards against it.
Why the door beats the window
Why the door beats the window.
The mechanics are almost embarrassingly simple once you see them.
You can change your default search engine in a single setting. One click, a coffee break, done. Now try changing the thing that holds the keys to everything else. Imagine someone who’s used “Sign in with Google” across twenty or thirty services — and once you start counting your own, the number climbs faster than you’d like. That account isn’t an account anymore. It’s the hinge the whole house swings on. Lose it and you don’t lose one thing; you lose your bank login’s recovery path, your work tools, your tax software, your photos, the smart lock on your front door.
That’s the asymmetry. Search is a window you can swap in an afternoon. Access is the door the whole house hangs on — and the house has been quietly built around it.
This is switching-cost economics, and it has a clean shape. The hold a company has on you is its switching cost plus whatever its product is actually, presently better at. Advertising lives almost entirely on that second term — a marginally better result — which evaporates the instant a rival catches up. Access lives on the first, and the first only grows. Every new service you wire to that one login deepens the hold by one more door. Adding a lock is a single pleasant click. Removing it means re-keying every door at once, in parallel, under deadline, with permanent lockout as the price of getting it wrong. The pain isn’t additive. It’s combinatorial. That gap — between how easy it is to add the lock and how terrifying it is to pull it — is the moat.
Salesforce and SAP have lived inside this physics for decades, holding enterprise customers for twenty-five-year stretches, and nobody calls them content businesses. Google built the same thing for your whole life and handed it out for free.
The institutions confirmed it by where they aimed. When the U.S. courts found Google an illegal monopolist, the remedy went after the contracts — the roughly twenty billion dollars a year Google pays Apple to be the default, the exclusive default-search deals, now capped to one-year terms. But the court declined to break off Chrome or Android. It renegotiated who gets to answer the door and left untouched the company that built every lock, hinge, and recovery key in the house. Even the people dismantling the monopoly treated “who is the default way in” as the twenty-billion-dollar question — and left the deeper layer, the one that actually owns login, autofill, passkeys, and recovery, exactly where it was.
The thing it holds is a piece of your mind
I could have left it at economics. But the lockout didn’t feel like an economics problem at one in the morning. It felt like an amputation, and I want to take that feeling seriously, because it’s the truest part.
There’s an old argument in philosophy of mind — Andy Clark and David Chalmers, 1998, “The Extended Mind.” They imagine Otto, a man whose memory is failing, who writes what he needs in a notebook and consults it the way you and I consult the inside of our own heads. Their claim isn’t that the notebook helps Otto’s mind. It’s that the notebook is part of Otto’s mind — the storage just happens to sit outside his skull. If a process counts as remembering when it happens in your head, it counts as remembering when it happens in the world.
I read that and thought about Stefani. “Remember for her when she can’t” is Otto’s notebook, almost word for word. The philosophy was settled twenty-eight years ago: the thing that holds your memory for you is not a tool you use. It is part of the mind doing the remembering.
Then the cognitive science caught up with the philosophy. In 2011, Betsy Sparrow and her colleagues at Columbia tested how people handle information they expect to look up later. We don’t retain the information, they found — we retain where to find it. The brain offloads the content and keeps the pointer. We are becoming, in their phrase, symbiotic with our tools. Sit with that: human memory already ran my experiment and reached my conclusion. It threw away the fact and kept the way back in. Access beating content isn’t a strategy I invented. It’s how your own head now works.
Which means whoever holds the pointer holds the only half of the memory your brain bothered to keep. You can swap a search engine in a second. You cannot swap a piece of your own mind without something that feels, accurately, like a small lobotomy. An ad interrupts you. A lockout unselfs you. And the entity that hands you back in isn’t selling you a service. It’s returning you to yourself.
There’s a flip side I have to be honest about, because it’s the whole case for doing this carefully. Sparrow’s same line of research shows that offloading frees you up — trusting that something is safely stored elsewhere measurably improves your ability to learn the next thing. But it also shows the benefit reverses when the external store turns out to be unreliable. You end up worse off than if you’d never offloaded, because you pruned the internal copy and the external one failed you. Reliability isn’t a feature of a continuity layer. It’s the entire product. A second brain that might vanish doesn’t merely fail to help — it degrades the mind that came to depend on it.
The blade cuts both ways
So here’s where I turn the knife on my own argument, because the thing that makes access powerful is the same thing that makes it dangerous, and I don’t trust anyone who won’t say so.
Access is a pharmakon — Plato’s word, the one Derrida built on: the single substance that cures and poisons, depending on nothing but the dose and the hand that holds it. The recovery flow that rescued me at 2 a.m. is, mechanically, the identical system that means I can never fully leave. Not two features in tension. One feature, seen from two sides.
Android makes it literal. Factory Reset Protection turns a wiped phone into a brick until the original Google account is re-verified. The feature that stops a thief from using your stolen phone is the same feature that makes the device hostage to Google’s say-so. Protection and imprisonment, one mechanism — and Google isn’t retreating from this ground, it’s deepening it, because recovery is exactly where the bond forms. The company that saves you and the company that traps you are the same company. You’re just meeting it at two different moments.
Now let me take the strongest objections head-on, because the good ones are real.
“Switching costs approach infinity.” No. I used to say it that way, and it was wrong. People migrate ecosystems by the hundreds of millions and carry their photos and contacts with them. Phone-number portability was mandated and it worked. Passkeys are an open standard, and their own backers built a credential-exchange protocol specifically to make them portable between password managers. Europe’s data-portability law already forces Google to hand you everything. My own founding story refutes the infinity claim: I got back in by morning. The moat is high, it is real, and it is finite and shrinking by design — every serious regulatory and technical current of this decade is engineered to grind it down. And that cuts in my favor. If lock-in were infinite, “we’ll let you leave” would be a meaningless promise. It means something only because leaving is becoming genuinely possible.
“Isn’t ‘access as care’ just what every captor says?” Yes. Company towns called themselves family. AOL called itself a community. Every lock-in business in history has narrated itself as care, and the distinction is invisible at the exact moment it matters most — when you’re locked out, sick, grieving, laid off, and least able to audit whether anyone actually has your back. This is the real soft spot, and I won’t paper over it. Care cannot be declared. It has to be engineered — and provable by someone who never read the terms. Words are free. I’ll come back to what isn’t.
“Gratitude isn’t a moat — the 2 a.m. plumber gets it too.” Correct. The ER, the locksmith, roadside assistance, my own restoration clients on the worst day of their lives — they all bond at the moment of relief, and gratitude decays, and people shop their insurance anyway. So gratitude isn’t the moat. It’s the on-ramp. The midnight rescue doesn’t lock anyone in; it earns the first conversation. What keeps them is what you do after — and that’s a question of character, not a property of the crisis.
Care holds the same keys — and hands you a copy
Let me show you what the answer looks like before I argue for it.
Last winter one of my restoration clients walked into a commercial building with two inches of standing water across the floor — burst supply line, ceilings down, a decade of operating records soaking in a back office that also held the only copies of their continuity plan, their vendor contracts, their insurance file. By the time the water was out, the part they were most afraid of losing wasn’t the drywall. It was the paper. We’d already pulled their critical records into a structured store they could reach from a phone — indexed, searchable, theirs. The owner stood in the wreckage and opened the file on his phone, and the thing that could have ended the business was just there. Then the part that matters to this essay: when the job closed, the whole store exported in one motion, in formats their own systems could read, and went with them. No call to me. No ransom for their own records. They walked out with the keys in their hand, and the relief on the owner’s face was the entire argument I’m about to make, compressed into one moment.
That’s the difference between holding the keys for someone and holding them over them. Once you accept that the held thing is part of a person’s mind, the ethics stop being a garnish and become the architecture. Holding a piece of someone’s cognition and refusing to let them leave isn’t hard-nosed business; it’s closer to holding a self hostage. Holding that same piece while guaranteeing they can walk out with all of it, any time, without asking — that’s not a vendor. That’s a trustee. The oldest answer the law has to the question of how you hold something vital that belongs to someone else: you hold it for them, bound to their interest, returnable on demand.
The whole thing collapses to one question. Not do you hold the keys — someone always holds the keys. The question is whether you hold them for her or over her. Google books your access as its switching cost, an asset on its side of the ledger. The humane version books it as your asset, merely held in trust. Same keys. Opposite politics.
Which is why I keep coming back to the difference between a scaffold and a cage. Good scaffolding is built to come down — calibrated to do only what the person can’t yet do alone, withdrawn as they grow. A scaffold that never comes down isn’t support anymore; it’s a wall you’ve forgotten how to live without. “Remember for Stefani when she can’t” is the morally exact phrasing — contingent help for a real gap, not a blanket seizure of her agency. Do everything for someone and you don’t make them safe. You teach them they can’t.
And I’ll admit the moat I’m choosing is the weaker one. A lock-in moat is strong precisely because it’s coercive — you stay because you can’t go. A trust moat is fragile; one breach and it’s gone overnight. I’m choosing the fragile one on purpose, and not only because it’s right. Lock-in and care produce the identical retention number — ninety-nine percent stay either way — but for opposite reasons, and the difference only shows up the day switching becomes free. That day is coming: portability law, open credential standards, and soon an AI agent that can re-key your whole life in an afternoon. When it arrives, the captivity moat evaporates and the trust moat doesn’t even notice. Free exit isn’t charity — it’s the only hold worth having once leaving is easy and everyone knows it. I’m not being generous. I’m being early.
But I won’t let myself off with a promise, because a promise from an interested party is exactly what breaks the day the incentives flip — an acquisition, a cash crunch, a change of hands. So the care has to be built into things that survive my intentions. Export in open, ingestible formats — not a dead blob no other system can read, which is fake portability wearing a real coat. A published exit that works without anyone calling me. A governance mechanism that binds the company after it’s sold. Don’t trust my intentions. Trust the mechanism that outlives them. That’s the only honest answer to “every captor says that.” The test was never the happy customer. It’s whether the grieving spouse who never read a word of the terms can still get everything out, in one motion, with no call to me. Design for the person who can’t advocate for themselves, and the ethics stop being marketing.
The door is moving — to the agent
The door is moving — to the agent.
This is also the shape of the next decade, and it’s why I work the way I work.
Google holds the keys to your accounts. The AI agent is coming to hold the keys to your context — what you’re working on, what you decided last month, how you actually think and operate. That’s a deeper hook than a login, because a login gets you into the app, but context is the work. Search was a query you typed and forgot. The agent is a relationship that accumulates.
And there’s a real chance, for the first time, that the door doesn’t have to be a cage. The plumbing that lets an agent reach into your files, calendar, and tools — Anthropic’s Model Context Protocol — is being built as a shared, open standard rather than one company’s private wiring. I won’t call that settled or “neutral”; standards get captured, and this one is young enough to go either way. But open plumbing at least makes it possible to build an agent that reaches into everything you own without owning it. Access without capture is finally buildable, not merely sayable.
The trap is moving too — and getting subtler. The new lock-in isn’t your data. It’s the agent’s learned understanding of you, accreted day after day. You can export every chat log and still leave behind the part that actually knew you, because raw logs aren’t understanding, and no portability law reaches that gap. Which is the whole reason I build on Claude rather than treat any of this as theory: its memory has a delete button and an export button. You can read what it knows about you, change it, take it elsewhere, even bring your history in from somewhere else. That’s not a feature. It’s a thesis with a receipt — own the payload, walk out anytime, shipped.
I have to name the obvious dark mirror, because it’s already shipping. Microsoft Recall makes the identical pitch — we’ll remember everything for you — by quietly screenshotting your screen every few seconds into a local index. Same promise, opposite governance: a memory built about you, by default, that you didn’t author and can’t easily hand to anyone else. The pointer to your own mind, held on someone else’s terms. The seat for “Sign in with your agent” is still empty, but the room is filling — Recall, OpenAI’s persistent memory, Gemini woven through Android, Apple’s on-device intelligence are all reaching for it. Whoever defines what care looks like before that seat fills sets the norm for everyone after. That’s not a forecast from the bleachers. It’s the work.
What I’m actually building
So let me say what my portfolio really is, because I had it mislabeled too.
It looks like five businesses held together by nothing but my calendar — restoration clients, the second brain, the Compass, remembering for Stefani, the structured record a company can’t operate without. It’s one product. Each version shows up at the bottom — the moment of maximum vulnerability, when someone has the least to spare and the most to lose — takes custody of a piece of their continuity, and is built, from the foundation, to give all of it back. Continuity is the one thing the attention economy never touches: the durable layer a person or a business runs on — their records, their memory, their way back into their own life — the part that, if it vanished, would not just inconvenience them but unself them.
The attention economy fights for you when you have everything to spare, which is why it has to shout and why you resent it for shouting. The continuity layer shows up when you have nothing left, and arrives with relief. Bonds made at the bottom run deeper than impressions bought at the top — but only one kind of person should be trusted to be there at the bottom: the kind who hands you the key on the way in.
I’ll concede the last hard thing plainly, because a skeptic has already spotted it. Today, the part of my work that pays the bills is the discovery work — getting found, getting ranked, getting cited. The continuity layer is real but young, and I won’t pretend it has finished proving it can pay. Here’s how I think it does: not by charging for the data, which would just be the cage again, but as a held-in-trust retainer — an ongoing fee for keeping the lights on and the door unlocked, priced like what it is, a fiduciary relationship rather than a subscription you’re trapped inside. You earn the right to charge it by first being useful enough to be found. Discovery isn’t a contradiction of the thesis; it’s the front door. Attention comes first. It always did. The mistake is thinking it’s the destination.
And here’s the part I can’t dodge, the one that keeps me honest. The agent I’m betting on — the one that can re-key a whole life in an afternoon — is the same tool that dissolves my moat too. If re-keying is trivial, the switching cost protecting my own work goes to zero right alongside Google’s. I’m left holding nothing but the fragile thing: trust, provable on the day someone decides to leave. That isn’t a bug in my bet. It’s the point of it. The tool I’m wagering everything on is the one that guarantees I can never coast — it leaves me no hold on anyone except being worth staying with. I’d rather build on that than on a lock.
Which is where it lands, in one line I’ve earned the right to say now:
Don’t sell knowledge. Don’t sell content. Sell access to continuity — and prove it’s care and not a cage by handing the customer the key on the way in.
I learned that locked out of my own life at two in the morning, patting my pockets at a door, negotiating with the only entity that could tell me whether I was still me. Google taught me how much that door is worth. It just never taught me to hand anyone a copy of the key. That part’s on us — and the copy is the whole job.
Most of what a working AI system does happens in silence. The operator sees the output. The operator does not see the labor. The labor — the prompts that ran, the data that was queried, the small decisions made hundreds of times across a session, the loops that were entered and exited — happens in a quiet room the operator usually does not enter.
There is a small but important practice in periodically going to the quiet room and watching the work happen.
Why most operators don’t do this
The quiet room is dull. The labor is repetitive. Watching the system work is much less satisfying than reviewing the system’s output. The dashboard is the highlight reel; the quiet room is the practice. Most operators, given the choice, watch the highlight reel.
This is reasonable in the short term. It is dangerous in the long term. The operator who only ever sees the output develops an intuition for the output and no intuition for the labor. When the output is wrong, the operator who has been watching the labor knows which step to look at. The operator who has been watching only the output is stuck.
What the quiet room teaches
It teaches the texture of the system’s reasoning. Where the system pauses. Where it overcommits. Which kinds of inputs produce which kinds of paths. What looks like efficiency is actually default behavior versus actual judgment.
It teaches what the system does badly. Every working system has a set of small recurring inefficiencies — wasted lookups, redundant verifications, paths that loop slightly more than necessary. Most of these are invisible from the output. They are visible from the labor. Watching them gives the operator a real sense of what to optimize and what to leave alone.
It teaches when to trust. The operator who has spent time in the quiet room has a calibrated sense of where the system is reliable and where it is reaching beyond its competence. That calibration is not in the output. It is only in watching the work.
The practice
The practice is small. Once a week, instead of reviewing only the output, spend twenty minutes in the labor. Read the trace of a session that produced something. Watch the prompts the system used, the tools it called, the decisions it made about which path to take. Note where the labor surprised you — positively or negatively. Update the working model.
This is unglamorous. It does not produce anything. It does not show up in the dashboard. It is a deposit in an account the operator will draw on six months from now when something does not look right and the operator has to decide whether to trust the system’s read.
The closing read
The output is the public face of the system. The quiet room is where the system is actually built. The operator who knows only the public face will, eventually, be surprised by the system. The operator who has been to the quiet room periodically — even briefly, even unsystematically — will not be. That is most of what calibration is. There is no shortcut for the labor of watching the labor.
The Architecture of Delegation: Moving Beyond the Chat Interface
The architecture of delegation — beyond the chat interface.
I spent today wiring Claude Code to boss around the Gemini CLI, clearing a 1,256-post WordPress tagging backlog without a single hallucinated tag. If you operate an agency or manage technical strategy at any reasonable scale, you already know the fundamental truth about current AI tools: the chat interface is a massive bottleneck. Copying, pasting, and waiting for a typing animation isn’t a workflow; it’s theater. Real, scalable throughput requires system-to-system communication and architectural delegation.
The goal for today wasn’t just to write a python script. The goal was to establish a functional hierarchy between two distinct AI systems operating locally on my machine. Claude Code, operating directly in my terminal, would act as the lead engineer and orchestrator. It would handle the logic, map out the API calls, write the Python bridges, and manage the error handling. Gemini, accessed via its official command-line interface, would act as the high-context, high-throughput worker.
The setup was brutally simple but effective. I installed the Gemini CLI using a standard node package manager command (npm install -g @google/gemini-cli) and authenticated it with a Google One AI Ultra account. This gave my local environment direct, command-line access to Google’s most capable models without needing to manage raw API keys or custom curl requests. From there, Claude Code was instructed to shell out via bash, calling the gemini command non-interactively to pass massive data payloads for processing, and then ingesting the structured output back into the orchestration pipeline.
It is an assembly line in the truest sense. Claude builds the machinery and defines the parameters; Gemini operates the heavy press, stamping out classifications at a volume that would break a standard chat context window.
Quantifying the Backlog and the Taxonomy Threat
Before you throw compute at a problem, you have to measure it accurately. I directed Claude to run a full audit of tygartmedia.com using the native WordPress REST API. The numbers came back clean, but the scale of the maintenance debt was daunting.
Total published posts: 2,529 individual pieces of content.
SEO infrastructure: RankMath confirmed healthy and active across the board.
Existing tag vocabulary: 931 distinct, strategically established tags.
The deficit: 1,256 posts sitting entirely untagged, orphaned from the site’s primary taxonomy.
In the past, solving this was a lose-lose proposition. It was either a job for a junior employee spending three agonizing weeks in the wp-admin panel, or it was a job for a messy automated script that inevitably hallucinates a thousand new, slightly misspelled tags. When you let an LLM tag 1,256 posts without strict, physical constraints, you don’t get an organized site. You get “Marketing”, “marketing”, “digital-marketing”, and “Digital Marketing Strategy” added as four completely separate taxonomy terms, permanently bloating your wp_terms table and diluting your internal link equity.
The constraint I set for this pipeline was absolute. The system had to read the 1,256 untagged posts, assign 5 to 8 highly relevant tags to each post, and only use tags from the exact 931-item vocabulary we already had. Zero deviation. Zero hallucination. If a perfect tag didn’t exist in the vocabulary, the system had to settle for the closest existing match rather than inventing a new one.
The Pilot Test and the Strict JSON Constraint
We started small to validate the pipeline. Claude pulled a pilot batch of 10 untagged posts from the WordPress API, along with the complete, raw list of 931 acceptable tags. It packaged this massive block of text into a single, dense prompt and fired it over to the Gemini CLI.
The instruction was clear and unforgiving: read the text of the posts, evaluate them against the vocabulary, and return ONLY a valid JSON object. I did not want markdown formatting. I did not want a polite introductory sentence. I needed a raw JSON string mapping each specific post_id to an array of its assigned tag IDs.
If you’ve spent any significant time wrestling with large language models, you know that asking for strict adherence to a vocabulary and strict, unformatted JSON output is exactly where things usually break down. Models inherently want to chat. They want to explain their reasoning. They want to invent a 932nd tag because it felt slightly more semantically accurate for a specific paragraph.
Gemini didn’t flinch. It processed the prompt and returned a raw, perfectly formatted JSON string directly to the standard output. Claude parsed it in memory, validated the suggested tags against the local vocabulary list, and found a 100% match rate. Every single tag suggested by Gemini was real. There was no conversational filler, no missing structural brackets, and no invented taxonomy. Claude immediately took that JSON, formatted the correct POST requests, and pushed the updates back to WordPress via the REST API.
Scaling Up: Hitting the Windows Bottlenecks
With the pilot completely successful, it was time to scale. Processing 1,256 posts one by one is inefficient, both in terms of time and system calls. We grouped the remaining posts into chunks of 25. This meant Claude would need to loop through roughly 50 distinct batches. For each batch, it would dynamically construct the prompt with the 931 tags and the 25 new post payloads, call Gemini, parse the resulting JSON, and patch the WordPress database.
That is where the friction started. Building a local orchestration pipeline means you are no longer just dealing with AI limitations; you are dealing with local OS limits. Windows had two specific, technical walls waiting for us.
Failure 1: WinError 2 (File Not Found)
The initial Python orchestration script used the standard subprocess.run(['gemini', '-p', prompt]) command to invoke the CLI. It failed almost immediately with a WinError 2. The issue? When npm installs global packages on a Windows machine, it doesn’t create a raw binary; it creates a .cmd wrapper. Python’s subprocess module doesn’t automatically resolve these wrappers unless you pass shell=True, which introduces a host of security and string parsing headaches. The clean, robust fix was forcing Claude to locate the executable and use the absolute, fully qualified path to gemini.cmd in the subprocess call. It’s a minor detail, but one that breaks entire automation pipelines if you don’t know what you’re looking at.
Failure 2: “The command line is too long”
Once the executable actually resolved, the script crashed again on the very first batch. Windows threw a fatal error: “The command line is too long.” Windows enforces a strict character limit on command-line arguments—roughly 8,191 characters depending on the exact environment. Our dynamically generated prompt, containing the full text of 25 blog posts and 931 taxonomy terms, hovered around 20KB. Trying to pass that payload via the standard -p argument flag was physically impossible for the operating system to handle.
The solution was architectural. Instead of trying to cram the prompt into an argument, Claude rewrote the Python script to pipe the prompt directly into Gemini’s standard input (stdin). By restructuring the workflow to write the 20KB payload to a temporary text file on disk, and then piping it via a standard input redirect (gemini < prompt.txt), we bypassed the OS argument limit entirely. The data flowed, and the pipeline spun back up to full speed.
The Verdict: The Orchestrator vs. The Worker
The orchestrator vs the worker.
Watching this script hum through 50 consecutive batches crystalized a specific, actionable opinion about the current state of local agentic workflows. You do not need one god-model to do everything; you need specialized roles operating within a hierarchy.
Claude Code is unmatched as an orchestrator. It understands the local filesystem, it navigates REST API documentation with ease, it writes robust, defensive Python, and it can dynamically debug Windows-specific OS errors on the fly. But using Claude for the repetitive, high-volume, token-heavy classification of thousands of posts is an expensive and slow use of a strategic brain. It is the equivalent of having your lead architect nailing drywall.
Gemini, operating locally via its CLI, proved to be the ultimate high-throughput worker. It absorbed the massive context window of 931 tags and 25 full articles simultaneously, over and over again, without degrading in quality. It maintained absolute discipline over the JSON output structure across 50 separate invocations. It didn’t need to understand how the WordPress API worked, and it didn’t need to know how to write Python. It only needed to process the classification task it was handed and get out of the way.
When Gemini acts as the worker and Claude acts as the boss, you get the absolute best of both architectures. You get the system-level problem-solving and environmental awareness of Claude, combined with the raw, reliable, high-context processing power of Gemini.
Tomorrow’s Takeaway
Tomorrow’s takeaway.
If you operate an agency and have a massive backlog of unstructured data—whether it is untagged content, uncategorized financial transactions, or messy CRM records—stop trying to fix it manually inside a browser window. The chat interface is dead for real, scalable work.
Tomorrow, install an agentic CLI like Claude Code. Give it access to a high-context execution model via a secondary CLI, like Gemini. Tell the orchestrator to write a local script that batches your data, hands the batches to the execution model, forces a strict, structured JSON return, and posts the results directly back to your database or CMS. Expect the script to break on local OS limits. Fix the pipes, use standard input instead of arguments for massive payloads, and let the machines clear the backlog while you focus on actual strategy.