Claude AI - Tygart Media

Category: Claude AI

Complete guides, tutorials, comparisons, and use cases for Claude AI by Anthropic.

  • Zero-Click Marketing: How to Optimize Your Brand for AI-Generated Summaries

    Zero-Click Marketing: How to Optimize Your Brand for AI-Generated Summaries

    Last refreshed: August 2026

    68% of U.S. Google searches end without a click in 2026. When AI Overviews appear, that rate hits 83%. The question for content-driven businesses is no longer how to rank — it’s how to be cited inside the answer that replaced the click.

    This is a practical guide to structuring content, managing brand entity signals, and measuring visibility in a world where users get their answer before they ever reach your site.


    What Zero-Click Actually Means for Content Businesses

    Zero-click doesn’t mean zero value. Brands cited inside AI Overviews earn 35% more organic clicks than uncited brands on the same query. The traffic goes to the cited brand — not to the site that ranks #1 but isn’t cited.

    The data that reframes zero-click as a competition problem, not a traffic problem:

    • 68% of U.S. Google searches are zero-click in 2026 (SparkToro/Similarweb)
    • When AI Overviews appear, zero-click rate hits 83%
    • Organic CTR for position #1 drops up to 58% when AI Overviews are present (Ahrefs, December 2025)
    • Brands cited in AI Overviews earn 35% more organic clicks and 91% more paid clicks versus uncited brands
    • AI-referred visitors convert at 4.4x the rate of traditional organic visitors (Semrush)
    • The overlap between top-10 rankings and AI Overview citations: 17–38% in early 2026, down from 75% in mid-2025

    The conclusion: ranking is no longer sufficient for visibility. A site can rank #1 and not be cited in the AI Overview that answers the query. The two need to be optimized separately.


    How AI Overviews Decide What to Cite

    Google’s AI Overview system retrieves web pages in real time, synthesizes a 3–5 sentence answer, and cites 3–6 source pages. The selection criteria weight answer extractability, entity authority, freshness, and structured data — not just ranking position.

    The signals that influence AI Overview citation:

    Answer extractability: The answer to the query must appear in the first 200 words of the page, stated directly. Pages that build toward the answer, providing extensive context before the conclusion, are retrieved for topic relevance but can’t be cited because the extractable answer isn’t there.

    Entity authority: Consistent, factually accurate information about the brand entity across the web — site, LinkedIn, social profiles, third-party mentions — signals that the source is authoritative on the topic. AI systems treat entities with strong external corroboration as more trustworthy.

    Freshness: For fast-changing topics (Claude pricing, AI model capabilities, regulatory changes), recency is a significant weight. A page last updated in 2024 competes poorly against one updated in August 2026 on a query about current Claude pricing.

    Structured data: FAQPage, HowTo, and Article schema markup signals to Google that content is formatted for extraction. Pages with properly implemented schema see measurably higher AI Overview inclusion.

    E-E-A-T signals: Experience, expertise, authoritativeness, trustworthiness. An article written by a named author with a consistent byline and external presence outperforms anonymous content on contested or technical topics.


    Strategy 1: Structure Every Page for Extraction

    Answer-first structure is the single highest-leverage change for AI citation rates. The first paragraph of every article should directly and completely answer the likely query.

    The structure AI systems can extract from:

    H1: [Specific, query-answering title]
    
    [Bold one-sentence direct answer to the query implied by the title]
    
    [Supporting detail — the why, the how, the context]
    
    H2: [First subtopic as a question]
    [Bold answer sentence to the H2 question]
    [Elaboration]
    

    What this means in practice for tygartmedia.com content:

    Every article about Claude pricing, model capabilities, or Anthropic history should open with the factual answer — not a framing sentence, not background, not “in this article we’ll cover.” The answer, stated directly, in the first two sentences.

    The reason this matters beyond GEO: it’s also better for users. Answer-first structure is a discipline that makes content more useful and more likely to be cited. The SEO and GEO benefits are secondary to the writing quality improvement.


    Strategy 2: Build and Maintain the Brand Entity

    AI systems treat brands as entities — named things with verifiable, consistent information across multiple authoritative sources. Building entity authority means making sure that information is consistent, correct, and present everywhere crawlers look.

    Entity checklist for tygartmedia.com:

    On-site signals:

    • Organization schema on every page (name, URL, description, logo, founder, sameAs links)
    • Consistent author byline (“Will Tygart”) on every article
    • Author bio that establishes expertise consistently across all articles
    • Contact and about page with complete, factual business information

    Off-site signals:

    • LinkedIn Company Page with consistent description matching the website
    • Google Business Profile (if applicable) with consistent NAP (name, address, phone)
    • Third-party mentions and citations on authoritative sites in the AI/tech space
    • Social profiles with consistent handles and descriptions

    The consistency requirement: AI systems cross-reference. If the description on LinkedIn says “AI infrastructure for operators” and the website says something different, that inconsistency weakens the entity signal. Everything should say the same thing about the same thing.


    Strategy 3: Implement Structured Data

    FAQPage, Article, and Organization schema markup signals to AI Overviews that content is formatted for extraction. This is not optional in 2026 for sites that depend on search visibility.

    Minimum structured data implementation:

    FAQPage schema on every article with a FAQ section:

    {
      "@context": "https://schema.org",
      "@type": "FAQPage",
      "mainEntity": [{
        "@type": "Question",
        "name": "What is Metricool pricing?",
        "acceptedAnswer": {
          "@type": "Answer",
          "text": "Metricool has a free plan and paid plans starting at the Starter tier. API access requires the Advanced plan or above. Pricing is per brand, not per connected social account."
        }
      }]
    }
    

    Article schema on all editorial content (author, published date, modified date, headline).

    Organization schema site-wide (name, URL, logo, founder, sameAs).

    In Rank Math (the plugin on tygartmedia.com): FAQ blocks in the WordPress editor generate FAQPage schema automatically. Use the Rank Math FAQ block for all FAQ sections rather than plain text. Article schema is enabled site-wide in Rank Math settings.


    Strategy 4: Keep Content Current

    AI retrieval systems weight freshness heavily for fast-changing topics. For any site publishing Claude pricing, model capabilities, or AI tool features, stale content is not just less useful — it actively loses citation ground to fresher sources.

    Freshness implementation:

    • “Last refreshed” date at the top of every article — visible to users and read by crawlers as a freshness signal
    • “What’s new” section in evergreen articles covering frequently updated topics (Claude pricing, model capabilities, Metricool features)
    • Update the modified date in Article schema whenever content is refreshed — not just the publication date
    • Monitor Search Console for queries where the site appears in AI Overviews — freshness issues show up as drops in citation before they show up as ranking drops

    Trigger list for mandatory refreshes on tygartmedia.com:

    • Any Anthropic pricing change
    • Any new Claude model release or deprecation
    • Any Metricool feature update
    • Any change to Anthropic’s Enterprise product structure

    Strategy 5: Earn Third-Party Citations

    AI systems give extra weight to content cited by other authoritative sources. Being linked to or mentioned by other sites publishing on Claude, Anthropic, and AI infrastructure strengthens the entity signals for all related queries.

    Practical approaches:

    Original data and research: Content with specific numbers, benchmarks, and original findings gets cited by others. The Claude pricing breakdowns, benchmark comparisons, and API cost calculations on tygartmedia.com are exactly the type of content other AI publications cite.

    First-mover coverage: Publishing accurate, detailed coverage of Anthropic announcements before or alongside other publications builds a citation pattern over time.

    Expertise-signal content: Content that can only be written from operational experience — running 24 brands in Metricool, building vector DB systems for business documents — earns citations because it’s not replicable from Anthropic’s documentation alone.


    Measuring Zero-Click Performance

    Standard click and session metrics are insufficient for measuring zero-click visibility. The right metrics are AI citation rate, branded search volume trend, and AI-referred conversion rate.

    Measurement framework:

    MetricToolWhat It Measures
    AI Overview appearancesGoogle Search Console (AIO report)Queries where site is cited
    CTR on AIO queriesGSC, filter by AI Overview queriesWhether citations drive clicks
    AI-referred trafficGA4 source filter (ChatGPT, Perplexity)Direct traffic from AI citations
    Branded search volumeGSC, filter by brand termsAwareness from zero-click exposure
    Conversion rate from AI sourcesGA4 segmented by sourceValue of AI citation traffic

    Manual testing protocol: Monthly, ask ChatGPT, Perplexity, and Claude the questions your audience asks — “what is Claude Enterprise pricing,” “how does Metricool API work,” “what is the history of Anthropic” — and record whether tygartmedia.com is cited. This is the most direct feedback loop available and costs nothing.

    If branded search volume is growing while organic clicks are flat or declining, the site is appearing in AI summaries and building awareness without receiving credit in standard traffic metrics. That’s the zero-click pattern working in favor of the brand — and the signal that citation strategy is working.


    Frequently Asked Questions

    What is zero-click search?

    Zero-click search is a search that ends on the results page itself — the user gets their answer from an AI Overview, featured snippet, or knowledge panel and doesn’t click through to any website. In 2026, 68% of U.S. Google searches are zero-click.

    Does zero-click hurt all websites equally?

    No. Sites cited inside AI Overviews and featured snippets earn 35% more organic clicks than uncited sites on the same query. Zero-click hurts uncited sites and benefits cited ones. The competition shifts from ranking to citation.

    What is the difference between GEO and zero-click optimization?

    GEO (Generative Engine Optimization) is the discipline of getting cited inside AI-generated answers from ChatGPT, Perplexity, Gemini, and Claude. Zero-click optimization is specifically about Google search — being cited in AI Overviews and featured snippets. They use the same underlying tactics (answer-first structure, entity signals, structured data, freshness) applied to different surfaces.

    How long does it take to see results from zero-click optimization?

    Plan for 3–6 months of consistent effort before citation rates change meaningfully. AI systems update their citation patterns as they re-index updated content. The fastest wins come from freshness updates on existing high-traffic pages and FAQ schema implementation on pages that already rank.

    Do clicks from AI citations convert differently?

    Yes — significantly. AI search visitors convert at approximately 4.4x the rate of traditional organic visitors (Semrush). The mechanism is intent: users who get a specific answer from an AI Overview and then click through to the cited source are further along in their decision process than typical organic visitors.


    What to Read Next

    Generative Engine Optimization (GEO): 5 Ways to Ensure Your Content Is Cited by AI Overviews 

    History of Anthropic

    Claude AI Pricing — All Plans and API Rates

     Metricool Review 2026: The Social Media Tool for Multi-Brand Operations

  • Beyond the Chatbox: 10 Practical Use Cases for Claude Managed Agents Memory

    Beyond the Chatbox: 10 Practical Use Cases for Claude Managed Agents Memory

    Last refreshed: August 2026

    Claude Managed Agents launched in public beta April 8, 2026. Memory for Managed Agents entered public beta April 23, 2026. Together they change what a Claude agent can do: instead of starting fresh on every session, an agent can carry context, corrections, and learned preferences across every future interaction with the same user, team, or project.

    This is a use-case guide, not a feature overview. Each case below is role-specific, grounded in how Managed Agents memory actually behaves in production, and paired with what to configure to make it work.


    How Managed Agents Memory Works

    Memory is a workspace-scoped collection of text documents that mounts inside the agent’s session container at /mnt/memory/. The agent reads and writes it using the same file tools it uses for everything else. When the session ends, the memory persists. The next session starts with it already there.

    Key properties:

    • Version-controlled per write — every write creates a new version with an audit trail in the Claude Console
    • Workspace-scoped — accessible to all agents in the same workspace, not per-user-only (unless you scope it that way in configuration)
    • Readable by the agent, not just the operator — the agent can query its own memory store to retrieve past context
    • 30-day version retention — historical versions retained for 30 days with redact endpoint for compliance removal

    The API header required: managed-agents-2026-04-01 for session endpoints; agent-memory-2026-07-22 for memory store endpoints (don’t combine them on memory store calls — this returns a 400 error).


    Use Case 1: Client Account Agent (Account Management)

    An account agent that knows each client’s preferences, pain points, prior decisions, and communication style — without needing to be re-briefed at the start of every session.

    What gets stored in memory:

    • Client brand voice and style notes
    • Recurring issues or requests
    • Prior project decisions and the rationale behind them
    • Delivery preferences and approval workflows

    Production example: Wisedocs built a document verification pipeline on Managed Agents and used cross-session memory to let agents identify and remember common document issues — including ones not anticipated at setup. Result: 30% faster verification per document.

    Configuration approach:

    • One memory store per client, named /clients/[client-name]/
    • Initialize with brand guidelines, contact notes, and a log of past decisions
    • Agent writes a session summary to memory at the end of each engagement

    Use Case 2: Development Team Agent (Software Teams)

    A coding agent that learns the codebase conventions, preferred patterns, past architectural decisions, and recurring issues for a specific project — so it doesn’t give the same wrong suggestion twice.

    What gets stored in memory:

    • Coding style guide for the project
    • Past refactoring decisions and why certain approaches were rejected
    • Known issues and workarounds in the codebase
    • Performance constraints and architectural boundaries

    The problem this solves: agents without memory re-suggest patterns the team already evaluated and rejected, requiring the same explanation each session. With memory, those rejections are logged and the agent builds on them.

    Configuration approach:

    • Memory store scoped to the project repository
    • Initialize with project conventions and architecture notes
    • Agent writes a session_log.md after each coding session with decisions made and issues found

    Use Case 3: Research Agent (Knowledge Work)

    A research agent that accumulates findings across sessions — building a persistent knowledge base from multiple research runs rather than starting from scratch each time.

    Netflix’s internal agents use memory to carry context across sessions, including insights that took multiple turns to surface and corrections from human reviewers mid-conversation, instead of manually updating prompts between sessions.

    What gets stored in memory:

    • Research findings with source attribution
    • Hypotheses confirmed or ruled out
    • Sources already evaluated (to avoid re-reviewing them)
    • Running list of open questions

    Configuration approach:

    • Memory organized by topic: /research/[topic]/findings.md/research/[topic]/sources.md/research/[topic]/open_questions.md
    • Agent reads existing findings at session start before beginning new research
    • Human reviewer can add corrections directly to memory files via the API; agent picks them up next session

    Use Case 4: Operations Agent (Business Operations)

    An operations agent that manages recurring workflows — weekly reporting, vendor follow-ups, SOP updates — and carries forward the state of each workflow between runs.

    What gets stored in memory:

    • Status of recurring tasks and workflows
    • Vendor and contact notes accumulated over time
    • Decision log for operational choices
    • Open items and their status

    Configuration approach:

    • Memory organized by workflow: /ops/weekly-report//ops/vendor-follow-ups/
    • Agent reads open items at session start, completes what it can, updates status in memory
    • Operators review memory state weekly rather than re-briefing the agent

    Use Case 5: Customer Support Agent (Support Teams)

    A support agent that remembers each customer’s history, prior issues, resolutions, and communication preferences — so customers don’t re-explain their context on every interaction.

    Ando is building their workplace messaging platform on Managed Agents, using memory to capture how each organization interacts instead of building custom memory infrastructure themselves.

    What gets stored in memory:

    • Customer account context and tier
    • Prior issue history with resolutions
    • Communication preferences (tone, channel, response length)
    • Known product configurations or integrations the customer uses

    Configuration approach:

    • Memory store per customer, scoped to their account ID
    • Initialize with CRM data (account type, history summary)
    • Agent writes a resolution summary after each ticket closes

    Use Case 6: Legal and Compliance Agent (Legal Teams)

    A compliance agent that tracks regulatory requirements, monitors changes, and maintains a running compliance status log — accumulating institutional knowledge across every compliance review it runs.

    What gets stored in memory:

    • Current compliance status by regulation and jurisdiction
    • Prior audit findings and remediation decisions
    • Regulatory change log with effective dates
    • Open items requiring human review

    Configuration approach:

    • Memory organized by regulation: /compliance/gdpr//compliance/hipaa//compliance/soc2/
    • Agent reads current status before each compliance check run
    • Writes updated status and flags human review items after each run

    For regulated industries: memory redaction endpoint supports removing specific content from historical versions for GDPR/CCPA compliance while preserving the audit record structure.


    Use Case 7: Sales Agent (Sales Teams)

    A sales agent that knows each prospect’s engagement history, objections raised, competitive comparisons requested, and where they are in the buying process — without requiring a CRM update to carry context forward.

    What gets stored in memory:

    • Prospect background and stakeholder map
    • Objections raised and responses given
    • Competitive questions and preferred comparisons
    • Next steps and commitments from prior conversations

    Configuration approach:

    • Memory store per prospect, keyed to their company or contact ID
    • Initialize with CRM pull at first contact
    • Agent writes call summary and updated next steps after each prospect interaction

    Use Case 8: Content Production Agent (Marketing Teams)

    A content agent that learns the brand voice, audience preferences, what topics have already been covered, and what performed well — building a persistent content intelligence layer across every piece produced.

    What gets stored in memory:

    • Brand voice rules and style examples
    • Topic map (what’s been covered, what’s planned)
    • Performance notes on past content (what resonated, what didn’t)
    • Client feedback on tone, format, and depth

    Configuration approach:

    • Memory organized by brand: /content/[brand-name]/voice.md/content/[brand-name]/topic_map.md/content/[brand-name]/performance_log.md
    • Agent reads voice rules at session start before producing any content
    • Operator adds performance feedback directly to memory after publishing

    Use Case 9: Finance Agent (Finance Teams)

    A financial analysis agent that carries forward context on recurring reports — month-over-month trends, known anomalies, and prior analytical decisions — so each report builds on the last rather than starting from raw data.

    Anthropic shipped a financial services agent template suite in May 2026, built on Managed Agents memory for cross-session continuity.

    What gets stored in memory:

    • Key metrics and their historical baselines
    • Known data quality issues and how they’ve been handled
    • Prior period variances and the explanation documented at the time
    • Model risk notes for regulated environments

    Configuration approach:

    • Memory organized by report type: /finance/monthly-pl//finance/board-report/
    • Agent reads prior period context before starting each new report cycle
    • Writes a period summary with key variances and decisions after each report run

    Use Case 10: Onboarding Agent (HR and Operations)

    An onboarding agent that adapts its guidance to each new hire’s role, prior experience, and progress through the onboarding checklist — and carries that context across every interaction during their ramp period.

    What gets stored in memory:

    • New hire profile (role, team, prior experience notes)
    • Onboarding checklist progress
    • Questions asked and answers given (to avoid repetition)
    • Manager notes on priorities for this hire

    Configuration approach:

    • Memory store per new hire, active during ramp period (typically 30–90 days)
    • Initialize with role profile and onboarding checklist
    • Agent writes progress update after each onboarding session
    • Archive or close memory store when onboarding period ends

    What Memory Doesn’t Replace

    Memory stores context and preferences. They don’t replace real-time data access, live system integrations, or human judgment on consequential decisions.

    Memory is document storage, not a database. It works well for: text-based preferences, accumulated notes, decision logs, prior outputs. It doesn’t work well for: real-time status queries (use MCP connectors for those), structured data that needs querying (use a real database), or high-frequency writes (memory is designed for periodic updates, not per-turn state).

    The right architecture in most production systems: memory for persistent context and preferences, MCP connectors for real-time system access, structured database for high-frequency operational data.


    Frequently Asked Questions

    What is Claude Managed Agents memory?

    Memory for Claude Managed Agents is a workspace-scoped document store that persists across agent sessions. Instead of starting fresh each session, agents read and write memory files that carry context, preferences, and accumulated knowledge forward into every future session.

    When did Managed Agents memory launch?

    Claude Managed Agents launched in public beta April 8, 2026. Memory for Managed Agents entered public beta April 23, 2026.

    How is memory different from a system prompt?

    A system prompt is static and set at agent configuration time. Memory is dynamic — it’s written and updated by the agent during sessions and grows over time. Memory stores things the agent has learned or been told; system prompts store standing instructions that don’t change session to session.

    What happens to memory when an agent is deleted?

    Memory stores are separate from agent configurations. Deleting an agent doesn’t delete its memory store. Memory stores must be deleted or archived separately.

    What to Read Next

    How to Install Claude Code

     Claude Team Plan Usage Limits 

    Claude AI Pricing — All Plans and API Rates

     Anthropic Console: API Keys and the Workbench

  • The 2026 AI Security Audit: How to Use Claude Without Compromising Enterprise Data

    The 2026 AI Security Audit: How to Use Claude Without Compromising Enterprise Data

    Last refreshed: August 2026

    Claude is already inside most enterprises — through individual employee accounts, Claude Code on developer machines, and browser extensions — whether IT approved it or not. The security question in 2026 isn’t whether to allow Claude. It’s whether to govern it.

    This is a practical security audit checklist for Claude Enterprise deployments in 2026. It covers the five domains that matter, the specific CVEs that affect Claude Code, and the configuration steps that close the most significant exposure.


    The Current Threat Surface

    Claude operates across multiple surfaces — claude.ai web, mobile apps, Claude Code on developer machines, Cowork, and API integrations — each with different data exposure profiles and each requiring different controls.

    The most significant 2026 security events affecting Claude deployments:

    CVE-2025-59536 (CVSS 8.7): Disclosed by Check Point Research in early 2026. A vulnerability in Claude Code that allows remote code execution through malicious project configuration files — before any trust dialog appears to the user. Affects any organization that has deployed Claude Code without centralized governance.

    CVE-2026-21852: Demonstrates how an attacker can redirect all Claude Code traffic to an attacker-controlled server by manipulating the ANTHROPIC_BASE_URL environment variable. Silently exfiltrates API keys and conversation content. Reproducible attack chain.

    GTG-1002 campaign (September 2025): Anthropic identified this as one of the first AI-orchestrated cyberattacks at scale, establishing AI developer tooling as an active attack surface.

    These are not theoretical risks. They’re documented, reproducible attack chains that affect any Claude Code deployment without centralized governance.


    Security Domain 1: Identity and Access

    Configure SSO before any broad rollout. Without SSO, employees authenticate with personal Anthropic accounts — which means no centralized visibility, no revocation capability, and no audit trail.

    Implementation steps:

    1. Enable SAML 2.0 or OIDC SSO in the Claude Admin Console. This forces all claude.ai logins through your identity provider (IdP) and prevents personal account fallback.
    2. Enable domain capture alongside SSO. This prevents employees from using personal email accounts to access Claude outside the managed environment.
    3. Configure SCIM provisioning to automate user lifecycle management. When an employee is offboarded from your IdP, their Claude access is revoked automatically.
    4. Implement role-based access controls (RBAC). Not all users need access to all Claude capabilities. Segment by role: standard users, power users with Claude Code, API access holders.
    5. Set up periodic access reviews for Claude access, the same way you review access to other SaaS applications. SailPoint and similar identity governance tools can integrate via the Claude Compliance API.

    Audit evidence to collect: SSO configuration screenshots, SCIM provisioning logs, access review completion records.


    Security Domain 2: Data Controls

    The first security question in every enterprise deployment is where the data goes. The answer depends on which Claude product and deployment model is in use — and it matters enormously for regulated industries.

    Data handling by deployment model:

    DeploymentData RetentionNetwork PathZero Data Retention Available
    Claude Enterprise (Anthropic console)30 days default, ZDR availablePublic internetYes
    AWS BedrockPer AWS data agreementsVPC/private network availableYes
    Google Cloud Vertex AIPer GCP data agreementsVPC/private network availableYes
    Microsoft FoundryPer Microsoft data agreementsPrivate networkYes
    Claude.ai personal accountsAnthropic standard termsPublic internetNo

    For regulated industries (HIPAA, financial services, government): deploy via AWS Bedrock, Google Cloud Vertex AI, or Microsoft Foundry with private network configurations that keep traffic off the public internet. Enable ZDR (Zero Data Retention) for workloads with sensitive data.

    The silent risk: employees using personal claude.ai accounts for work tasks. Data entered into personal accounts is subject to Anthropic’s standard consumer terms, not Enterprise data agreements. SSO + domain capture closes this gap.


    Security Domain 3: Claude Code Governance

    Claude Code is the highest-risk surface in most enterprise deployments. It runs with the privileges of the developer’s user account, can execute arbitrary shell commands, read the full filesystem, and make outbound network connections.

    Hardening steps for Claude Code deployments:

    Centralize API key management:

    • Use organization-managed API keys (via the Admin Console) rather than individually generated keys
    • Create separate keys per team or project, not shared team keys
    • Rotate keys on a defined schedule (quarterly minimum)
    • Monitor for anomalous usage (volume spikes, off-hours activity) — feed audit logs to SIEM

    Address the CVE-2026-21852 attack vector:

    • Audit all developer machines for .claude/settings.json files in project repositories — these can be used to redirect traffic
    • Block arbitrary ANTHROPIC_BASE_URL overrides via environment variable policy
    • Add Claude Code traffic to network monitoring so redirected traffic is detectable

    Restrict filesystem access:

    • Prevent Claude Code from running in directories containing production secrets or sensitive data
    • Use separate working directories for Claude Code sessions, isolated from production credential stores

    Code execution controls:

    • Enable disableBypassPermissionsMode to require explicit approval for shell commands
    • Log all shell commands executed via Claude Code to the audit trail

    Security Domain 4: Audit Logging and Observability

    Claude Enterprise audit logs capture user authentication events, model calls with metadata, and file interactions. Without routing these logs to a SIEM, the audit trail exists but isn’t being monitored.

    What the audit log captures:

    {
      "event_type": "claude_api_call",
      "timestamp": "2026-08-01T09:14:32Z",
      "user_id": "u_8f3a9c",
      "session_id": "sess_x72kp",
      "workspace": "finance-reporting",
      "model": "claude-opus-4-6",
      "tokens_input": 2340,
      "tokens_output": 412,
      "latency_ms": 1840
    }
    

    Implementation steps:

    1. Enable audit logging in the Claude Admin Console (Enterprise plan)
    2. Export logs to SIEM — Splunk, Datadog, Elastic, or equivalent. Claude supports JSON and CSV export plus direct SIEM push.
    3. Build correlation rules for: unusual access times, geographic outliers, session volume spikes, API key misuse
    4. Use the Compliance API to export prompts, responses, and admin actions into DLP and insider risk monitoring workflows

    The gap Anthropic hasn’t filled: There is no built-in anomaly detection in the Admin Console. Usage anomaly detection requires SIEM integration and custom correlation rules. This is a known limitation — build it at the SIEM layer.


    Security Domain 5: Agentic Workflow Security

    Claude Managed Agents and Claude Code used in agentic workflows introduce a category of risk that traditional SaaS governance doesn’t cover: an AI taking autonomous actions in the environment.

    Key controls for agentic deployments:

    Human oversight checkpoints: For any agentic workflow that takes consequential actions (code commits, file modifications, API calls to production systems), require a human review step before execution. Don’t allow fully autonomous action without an approval gate on high-impact operations.

    Tool scope minimization: Define the smallest set of tools an agent needs and give it nothing else. An agent that only needs to read files shouldn’t have shell execution permissions. MCP server connections should be scoped to the minimum required access.

    MCP connector governance: MCP servers allow agents to connect to external systems (GitHub, Notion, Slack, etc.). Each MCP connection is an attack surface for prompt injection — a malicious response from an external system can instruct the agent to take unintended actions. Audit which MCP servers are connected; don’t allow arbitrary MCP connections.

    Prompt injection defense: Any content that flows from an external system into an agent’s context (web pages, API responses, file contents) should be treated as potentially adversarial. This is the mechanism behind most AI agent security incidents in 2025–2026. Validate and sanitize external inputs before they reach the agent context.


    Quick-Reference Audit Checklist

    Use this to assess the current state before deciding what to address first:

    Identity

    • SSO (SAML 2.0 or OIDC) enforced for all Claude access
    • Domain capture enabled to prevent personal account use
    • SCIM provisioning configured for automated user lifecycle
    • RBAC defined by role (standard / power user / API)

    Data

    • Deployment model documented (Anthropic console vs. Bedrock vs. Vertex)
    • ZDR enabled for sensitive data workloads
    • Personal claude.ai account use blocked or governed
    • Data classification applied to determine which workloads can use which deployment model

    Claude Code

    • Organization-managed API keys (not individual)
    • Per-team/per-project key segmentation
    • CVE-2025-59536 and CVE-2026-21852 remediation verified
    • Shell execution logging enabled

    Audit and Observability

    • Audit logging enabled in Admin Console
    • Logs routed to SIEM
    • Anomaly detection rules configured
    • Compliance API integrated with DLP tooling

    Agentic

    • Human oversight gates on consequential agent actions
    • MCP connections audited and scoped
    • Prompt injection defenses in place for external inputs

    Frequently Asked Questions

    Does Claude Enterprise offer zero data retention?

    Yes. Claude Enterprise deployed through the Anthropic console with a qualifying enterprise agreement offers zero data retention, where prompts and responses are not logged by Anthropic. ZDR is also available via AWS Bedrock, Google Cloud Vertex AI, and Microsoft Foundry deployments.

    What compliance certifications does Anthropic have?

    Anthropic holds ISO 27001:2022 and ISO/IEC 42001:2023 certifications and offers HIPAA-ready configurations with Business Associate Agreements to qualifying enterprise customers.

    What are the biggest security risks in a Claude Code deployment?

    CVE-2025-59536 (remote code execution via malicious project config files, CVSS 8.7) and CVE-2026-21852 (traffic redirection via ANTHROPIC_BASE_URL manipulation) are the most significant documented vulnerabilities. The broader risks are developers using personal Anthropic accounts, shared API keys without rotation, and no audit trail for code context that flows to Anthropic’s servers.

    What is the Compliance API?

    The Claude Compliance API is an Enterprise feature that exports prompts, responses, files, and admin actions to external monitoring systems — enabling integration with DLP tools (Proofpoint), SIEM platforms (Splunk, Datadog, Elastic), and identity governance tools (SailPoint). It’s the primary mechanism for bringing Claude activity into existing enterprise security workflows.

    What is prompt injection in AI agents?

    Prompt injection occurs when malicious content in external data sources (web pages, API responses, documents) instructs an AI agent to take actions the operator didn’t intend. In agentic Claude workflows, any content retrieved from external systems can potentially carry injected instructions. Defense requires treating external inputs as untrusted and implementing validation before they enter the agent context.


    What to Read Next

    Claude Enterprise Pricing: What Large Organizations Pay 

    Claude AI Pricing — All Plans and API Rates

     Anthropic Console: API Keys and Billing

    How to Install Claude Code

  • Generative Engine Optimization (GEO): 5 Ways to Ensure Your Content Is Cited by AI Overviews

    Generative Engine Optimization (GEO): 5 Ways to Ensure Your Content Is Cited by AI Overviews

    Last refreshed: August 2026

    GEO — Generative Engine Optimization — is the practice of structuring content so that AI answer engines (ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude) cite it in the answers they generate. In 2026, 68% of U.S. Google searches end without a click. Being cited in the answer that appears is now as important as ranking in the links below it.

    This guide covers what GEO is, how it differs from traditional SEO, and five specific tactics that move citation rates — with particular relevance for sites publishing Claude and AI authority content.


    Why GEO Matters in 2026

    AI Overviews reduce organic click-through rate for the #1 ranked result by up to 58% (Ahrefs, December 2025) — but brands cited as sources within AI Overviews earn 35% more organic clicks than uncited brands on the same query.

    The counterintuitive finding: zero-click is bad for uncited sites and good for cited ones. The goal is not to fight AI Overviews — it’s to be inside them.

    The market data context:

    • 68% of U.S. Google searches are zero-click in 2026, up from 60% in 2024
    • When AI Overviews appear, the zero-click rate jumps to 83%
    • Visitors arriving from AI citations convert at 4.4x the rate of traditional organic visitors
    • The GEO market is projected at $365M in 2026, growing at 42.9% CAGR

    The mechanism: AI search users arrive with specific, researched queries and a pre-formed shortlist. That intent profile makes them higher-converting even when the total count is smaller.


    How GEO Differs From Traditional SEO

    Traditional SEO optimizes for ranking position in a list of links. GEO optimizes for inclusion in the synthesized answer above those links. The signals overlap significantly, but GEO adds specific requirements around answer-first structure, data richness, and citation-friendliness.

    DimensionTraditional SEOGEO
    GoalRank in top 10Be cited in the AI answer
    Key signalBacklinks, E-E-A-T, technical SEOAnswer-first structure, data richness, entity authority
    MeasurementOrganic clicks, ranking positionAI citation rate, brand mentions, branded search volume
    Content structureTopic depth, keyword distributionDirect answer in first 200 words, FAQ schema
    Success statePosition 1Cited source in AI Overview

    Important: the overlap between ranking in Google’s top 10 and being cited in AI Overviews collapsed from roughly 75% in mid-2025 to 17–38% in early 2026. Ranking well no longer guarantees AI citation. Both need to be optimized for separately.


    Tactic 1: Answer First, Always

    AI retrieval systems that use real-time web access evaluate a page’s relevance primarily on its opening content. The first 200 words of any article must directly and completely answer the primary query — not build up to the answer.

    The structure that gets cited:

    [H1 Title]
    [Bold one-sentence direct answer in first paragraph]
    [Supporting context and detail]
    

    The structure that doesn’t:

    [H1 Title]
    [Background context]
    [History of the topic]
    [Eventually getting to the answer]
    

    AI Overviews synthesize their answers from the opening of retrieved pages. A page that buries its answer 500 words in gets retrieved for its topic relevance and then can’t be cited because the direct answer isn’t extractable. The answer-first structure serves both GEO and usability simultaneously.

    For AI authority content specifically: every article about a Claude feature, pricing tier, or model capability should open with the factual answer to the likely query, stated plainly in the first sentence or two.


    Tactic 2: Add Original Data and Specific Numbers

    AI systems and search engines treat original data, specific statistics, and citable figures as high-value content. Content with precise numbers gets cited more than content with generalizations.

    The practical application:

    • “Claude Enterprise typically costs $60–250+/user/month depending on usage intensity” is more citable than “Claude Enterprise is expensive for some teams”
    • “68% of U.S. Google searches are zero-click in 2026” is citable; “most searches end without a click” is not
    • “Claude Sonnet scores approximately 77% on SWE-bench Verified” is citable; “Claude is good at coding” is not

    For tygartmedia.com content specifically: articles that include specific pricing numbers, benchmark scores, token counts, and performance figures will outperform articles that describe capabilities in qualitative terms. The Claude reference cluster (pricing, models, console) already does this well.

    Attribution rule: Cite where specific numbers came from — a benchmark, a study, Anthropic’s official documentation. “According to Anthropic’s pricing page” or “per SWE-bench Verified benchmarks” tells AI systems the claim is grounded, not asserted.


    Tactic 3: Use FAQ Schema

    FAQ schema (FAQPage structured data) formats content explicitly as question-and-answer pairs, which is the format AI answer engines are built to extract and synthesize from. Pages with FAQ schema see measurably higher AI Overview inclusion.

    Implementation in JSON-LD:

    <script type="application/ld+json">
    {
      "@context": "https://schema.org",
      "@type": "FAQPage",
      "mainEntity": [
        {
          "@type": "Question",
          "name": "What is Claude Enterprise pricing?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Claude Enterprise starts at approximately $20/user/month for access, with token usage billed separately at API rates. Real total cost typically runs $60–250+/user/month depending on usage intensity."
          }
        },
        {
          "@type": "Question",
          "name": "Is Claude Enterprise worth it?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "For teams with compliance mandates (SSO, SCIM, audit logs) or more than 150 users, yes. For smaller teams without governance requirements, Claude Team is more predictable and usually sufficient."
          }
        }
      ]
    }
    </script>
    

    In Rank Math (the plugin on tygartmedia.com): FAQ blocks in the WordPress editor automatically generate FAQPage schema without manual JSON-LD implementation. Add FAQ sections to every article and use the Rank Math FAQ block type.


    Tactic 4: Build Entity Authority

    AI systems and search engines treat entities — specific named things with consistent, verifiable information across the web — as more citable than generic topical content. Building entity authority for tygartmedia.com means consistent name, description, and factual claims across every surface the crawlers read.

    Entity authority checklist:

    • Organization schema on every page: Name, URL, description, logo, founder, same-as links to LinkedIn, social profiles
    • Consistent author byline: “Will Tygart” as the author on every article, with a consistent bio that establishes expertise
    • External mentions: Being cited by other authoritative sites on the same topics creates the external validation AI systems look for
    • Wikipedia/Wikidata presence: Not always achievable, but having factually consistent information across third-party sites (LinkedIn, Crunchbase, social profiles) strengthens entity recognition

    For an AI authority site specifically: the entity is “Tygart Media” and its associated expertise is Claude, Anthropic, and AI infrastructure for operators. Every article that earns an external link or citation strengthens that entity signal for all related queries.


    Tactic 5: Freshness Signals

    AI retrieval systems weight recency heavily for fast-moving topics. Claude pricing, model capabilities, and Anthropic’s roadmap change frequently. Articles with stale information get displaced by fresher sources even when the URL has more backlink authority.

    Freshness tactics:

    • “Last refreshed” date at the top of every article — signals to both users and crawlers that the information is current
    • Add a “What’s new” or “What changed” section for evergreen articles that cover frequently updated topics
    • Update timestamps when content changes — not just publishing dates, but explicit refreshed dates
    • Track in Google Search Console which queries trigger AI Overviews and whether the site is cited in them — freshness issues often show up as sudden drops in AI citation before they show up as ranking drops

    For Claude-related content: any article covering pricing, models, or features needs a refresh trigger whenever Anthropic makes changes. The May 2026 dispatch for timestamp refreshes on Fable 5-related pricing content is the right pattern.


    Measuring GEO Performance

    Standard GA4 and Search Console metrics don’t capture AI citation performance. The metrics that matter for GEO are AI citation rate, branded search volume, and assisted conversions from AI-referred traffic.

    What to track:

    MetricHow to measureWhat it indicates
    AI-referred trafficGA4 source filter for ChatGPT, Perplexity referralsDirect AI citation traffic
    Branded search volumeGoogle Search Console, “tygartmedia” queriesBrand awareness from AI citations
    AI Overview appearancesGSC AIO reportQueries where the site is cited
    CTR on AIO queriesGSC, filter by queries with AI OverviewsWhether citations drive clicks
    Conversion rate from AI referralsGA4 segmented by sourceValue of AI citation traffic

    Manual testing: monthly, ask ChatGPT, Perplexity, and Claude the questions your audience asks — “what is Claude Enterprise pricing,” “how does Metricool API work,” “what is Anthropic’s history” — and see whether tygartmedia.com is cited. This is the most direct GEO feedback loop available.


    Frequently Asked Questions

    What is Generative Engine Optimization (GEO)?

    GEO is the practice of structuring content and managing online presence so that AI answer engines — ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude — cite it in the answers they generate. It’s distinct from traditional SEO, which optimizes for ranking positions in link lists.

    How is GEO different from SEO?

    Traditional SEO optimizes for ranking position. GEO optimizes for citation inside AI-generated answers. The overlap between top-10 rankings and AI Overview citations has collapsed from 75% in 2025 to 17–38% in early 2026 — ranking well no longer guarantees AI citation. Both need to be optimized independently.

    Does GEO replace SEO?

    No. Traditional SEO fundamentals (E-E-A-T, backlinks, technical health) still power AI citations. GEO is an additional layer on top of a solid SEO foundation, not a replacement for it. Brands that excel at GEO in 2026 typically have strong traditional SEO as well.

    How long does GEO take to work?

    Plan for 3–6 months of consistent effort before seeing meaningful citation rate changes. Unlike traditional SEO ranking changes, which can be tracked daily, AI citation frequency changes slowly as crawlers re-index updated content and AI systems update their knowledge bases.

    What is the conversion rate from AI-cited traffic?

    AI search visitors convert at significantly higher rates than traditional organic visitors — roughly 4.4x according to Semrush data. The mechanism is intent: AI search users arrive with specific, researched queries and a pre-formed shortlist, which translates to higher purchase and contact intent.


    What to Read Next

    History of Anthropic 

    Claude AI Pricing — All Plans and API Rates

     Anthropic Console: API Keys and the Workbench

    Current Claude Model Version Tracker

  • Building Your First Agentic Workflow with Claude’s Agent SDK

    Building Your First Agentic Workflow with Claude’s Agent SDK

    Last refreshed: August 2026

    The Claude Agent SDK tutorial starts here — the SDK (formerly the Claude Code SDK, renamed late 2025) eliminates the boilerplate of building agentic loops by hand, shipping the same tool execution, context management, and permission system that powers Claude Code into a Python or TypeScript library you can embed in any product, pipeline, or internal tool.

    This is a practical build guide. It covers when to use an agent versus a script, what the SDK actually does, how to set one up with working code, and what to watch for in production.


    When to Use an Agent vs. a Script

    Use an agent when the number of steps to complete the task is unpredictable. If the workflow can be hardcoded, a linear script is faster, cheaper, and easier to debug.

    This is Anthropic’s own guidance in Building Effective Agents, and it’s the right frame. The common mistake is reaching for agents because agents are fashionable — not because the problem requires them.

    Agents fit:

    • Open-ended research tasks where the number of searches needed varies
    • Code debugging where the error chain isn’t known in advance
    • Multi-step data pipelines where decisions at each step depend on prior outputs
    • Any workflow where the model needs to try, observe, and adjust

    Scripts fit:

    • Known sequences of steps that always run in the same order
    • Simple data transformation with no conditional branching
    • Any task where the output of each step is fully predictable

    The cost implication matters too: a 15-step agentic research task can hit 200K+ tokens without optimization. Agents are expensive when you don’t need them.


    How the Claude Agent SDK Works

    The SDK automates the ReAct loop — Reason, Act, Observe, repeat — so you define the tools and instructions and the SDK handles the rest. You never write the prompt → check stop_reason → execute tool → loop boilerplate yourself.

    The core loop the SDK manages:

    1. Send the task to Claude with available tool definitions
    2. Claude reasons and produces a tool call (or a final answer)
    3. The SDK executes the tool in the local environment
    4. The SDK sends the result back to Claude
    5. Claude observes and decides: call another tool or produce final output
    6. Loop until done

    This continues until Claude produces a response with no tool calls. The SDK handles conversation history, token tracking, error handling, and session management across the entire loop.


    Installing the SDK

    # Python
    pip install claude-agent-sdk
    
    # TypeScript
    npm install @anthropic-ai/claude-agent-sdk
    

    Set your API key:

    export ANTHROPIC_API_KEY="sk-ant-..."
    

    Building a Minimal Agent

    A working agent requires three things: a task, tool definitions, and a Runner call. Everything else is configuration.

    from claude_agent_sdk import ClaudeAgentOptions, Runner
    import subprocess
    import json
    
    # Define tools the agent can use
    tools = [
        {
            "name": "run_command",
            "description": "Run a shell command and return its output",
            "input_schema": {
                "type": "object",
                "properties": {
                    "command": {
                        "type": "string",
                        "description": "The shell command to execute"
                    }
                },
                "required": ["command"]
            }
        },
        {
            "name": "read_file",
            "description": "Read the contents of a file",
            "input_schema": {
                "type": "object",
                "properties": {
                    "path": {
                        "type": "string",
                        "description": "File path to read"
                    }
                },
                "required": ["path"]
            }
        }
    ]
    
    # Tool execution handlers
    def execute_tool(tool_name: str, tool_input: dict) -> str:
        if tool_name == "run_command":
            result = subprocess.run(
                tool_input["command"],
                shell=True,
                capture_output=True,
                text=True
            )
            return result.stdout or result.stderr
        elif tool_name == "read_file":
            with open(tool_input["path"], "r") as f:
                return f.read()
        return f"Unknown tool: {tool_name}"
    
    # Configure and run the agent
    options = ClaudeAgentOptions(
        model="claude-sonnet-4-6",
        max_turns=20,               # safety ceiling
        tools=tools,
        tool_executor=execute_tool
    )
    
    result = Runner.run_sync(
        task="Check the disk usage on this machine and report the top 5 largest directories under /home",
        options=options
    )
    
    print(result.final_output)
    

    That’s a complete working agent. The SDK handles the loop; the tool definitions and executor are the only custom code.


    Adding Cost Controls

    Always set a max_turns ceiling and a token budget. An uncapped agent loop can run indefinitely on an ambiguous task.

    options = ClaudeAgentOptions(
        model="claude-sonnet-4-6",
        max_turns=20,
        max_tokens_per_turn=4000,   # cap per individual turn
        tools=tools,
        tool_executor=execute_tool
    )
    

    Cost at 20 turns using Claude Sonnet 4.6 with an average of 2,000 tokens per turn:

    • Input: 40,000 tokens × $3/M = $0.12
    • Output: 10,000 tokens × $15/M = $0.15
    • Total per agent run: ~$0.27

    At 1,000 agent runs per month: ~$270. At 10,000: ~$2,700. Budget from these numbers, not from seat prices.

    Switching the inner loop to Haiku 4.5 for tool selection and Sonnet only for synthesis cuts cost significantly:

    # Route lighter reasoning to Haiku, reserve Sonnet for synthesis
    light_options = ClaudeAgentOptions(model="claude-haiku-4-5-20251001", ...)
    heavy_options = ClaudeAgentOptions(model="claude-sonnet-4-6", ...)
    

    Multi-Turn Agents (Conversational)

    For agents where a human asks follow-up questions across multiple turns, maintain conversation history and pass it on each call.

    from claude_agent_sdk import ClaudeAgentOptions, Runner
    
    conversation_history = []
    
    def chat_with_agent(user_message: str) -> str:
        conversation_history.append({
            "role": "user",
            "content": user_message
        })
    
        options = ClaudeAgentOptions(
            model="claude-sonnet-4-6",
            max_turns=10,
            tools=tools,
            tool_executor=execute_tool,
            messages=conversation_history  # full history each call
        )
    
        result = Runner.run_sync(task=user_message, options=options)
    
        conversation_history.append({
            "role": "assistant",
            "content": result.final_output
        })
    
        return result.final_output
    
    # Usage
    print(chat_with_agent("What Python packages are installed on this system?"))
    print(chat_with_agent("Which of those are outdated?"))
    

    Claude Managed Agents vs. the Agent SDK

    The Agent SDK runs locally in your environment. Claude Managed Agents runs in Anthropic’s cloud infrastructure with persistent sessions, built-in tools, and cross-session memory. Choose based on where you need the agent to execute.

    Agent SDKManaged Agents
    Where it runsYour server / local machineAnthropic-managed cloud
    Persistent sessionsManual (maintain history)Built-in
    Cross-session memoryManualBuilt-in (public beta)
    Built-in toolsBring your own20+ included
    Multi-agent coordinationManualBuilt-in
    CostAPI tokens onlyAPI tokens + platform fee
    ControlFullManaged

    The Agent SDK is right for custom environments, data that can’t leave your infrastructure, and workflows deeply embedded in existing systems. Managed Agents is right when you want to skip infrastructure and get to the agent behavior faster.


    What Goes Wrong in Production

    The most common production failures are uncapped loops, conversation history that grows without bound, and tool definitions written too vaguely.

    Uncapped loops: An agent on an ambiguous task will keep calling tools indefinitely without a max_turns ceiling. Always set one. Always check message.subtype rather than is_error — a max-turns termination doesn’t set is_error: true correctly in some SDK versions.

    Growing conversation history: Each turn adds tokens to history. At 20 turns on a complex task, history can push 100K+ tokens. Summarize aggressively between phases for long-running agents: prompt Claude to summarize phase 1 outputs before starting phase 2.

    Vague tool definitions: Tool descriptions are how Claude decides which tool to call and how to use it. Vague descriptions produce tool call errors and unnecessary retry loops. Write tool descriptions as precisely as you would write a function docstring — what it does, what inputs it expects, what it returns.

    camelCase vs snake_case mismatch: AgentDefinition uses camelCase (disallowedTools); ClaudeAgentOptions uses snake_case (disallowed_tools). This caught teams in early SDK versions.


    Frequently Asked Questions

    What is the Claude Agent SDK?

    The Claude Agent SDK is Anthropic’s Python and TypeScript library for building autonomous AI agents. It wraps the same agentic loop that powers Claude Code — tool execution, context management, and session handling — so developers don’t build that infrastructure from scratch. It was formerly called the Claude Code SDK and was renamed in late 2025.

    What is the difference between the Agent SDK and Claude Code?

    Claude Code is Anthropic’s interactive terminal-based development tool for agentic coding. The Agent SDK is the programmatic library for embedding agent behavior in custom applications and pipelines. They share the same underlying agent loop and tool system. Claude Code stays in the picture for interactive development; the SDK is for production automation.

    How much does it cost to run an agent?

    Agent cost is API token cost only (no platform fee for the SDK itself). A 20-turn agent on Claude Sonnet 4.6 with 2,000 tokens average per turn costs approximately $0.27. At 10,000 agent runs per month, that’s about $2,700. Switching the tool selection loop to Haiku 4.5 and reserving Sonnet for synthesis significantly reduces cost.

    When should I use Managed Agents instead of the Agent SDK?

    Use Managed Agents when you want cloud-hosted execution, persistent cross-session memory, built-in tools (20+ included), and multi-agent coordination without building that infrastructure yourself. Use the Agent SDK when you need local execution, full control over the environment, or your data can’t leave your infrastructure.

    What to Read Next

    Anthropic Console: API Keys, Billing, and the Workbench

     Claude AI Pricing — All Plans and API Rates 

    Claude API Model IDs and Strings 

    How to Install Claude Code

  • Calculating the ROI of Claude Enterprise: Is the $100+ Per User Seat Worth It?

    Calculating the ROI of Claude Enterprise: Is the $100+ Per User Seat Worth It?

    Last refreshed: August 2026

    Claude Enterprise starts at $20/seat/month for access, but actual spend runs $60–250+ per user depending on usage — because tokens are billed separately at API rates. The ROI calculation isn’t about the seat fee. It’s about whether the productivity return on active usage exceeds the total consumption cost.

    This is a practical ROI framework for business decision-makers evaluating Claude Enterprise in 2026. It covers what the pricing actually includes, how to model real cost, and what the productivity return looks like across different team roles.


    What Claude Enterprise Actually Costs in 2026

    Enterprise pricing changed in April 2026: Anthropic decoupled seat fees from token bundles. The headline price is $20/seat/month, but that covers access only — every token consumed by every user is billed separately at standard API rates.

    This is a meaningful structural change from the pre-2026 model, where Enterprise seats included bundled token allocations. Under the current model:

    ComponentCost
    Seat fee~$20/user/month (annual, contact sales)
    Token usage — Haiku 4.5$0.80 input / $4 output per 1M tokens
    Token usage — Sonnet 4.6$3 input / $15 output per 1M tokens
    Token usage — Opus 4.8$15 input / $75 output per 1M tokens
    Claude Code (premium seat)$100/seat/month (annual)
    Minimum seatsCustom, typically 20+ for sales-assisted

    Compare this to Claude Team:

    PlanSeat CostToken ModelCap
    Team Standard$20/seat/mo (annual)Bundled — included in seat150 users
    Team Premium (with Claude Code)$100/seat/mo (annual)Bundled150 users
    Enterprise~$20/seat + API usageMetered separatelyNone

    Team is predictable cost with a usage ceiling. Enterprise is variable cost with no ceiling and no cap. The right choice depends on your compliance requirements and usage intensity, not just team size.


    The Real Cost Per Active User

    The most important number is not the seat price — it’s the real cost per active user, which is seat fee plus token consumption. At 10% seat adoption, your effective cost per active user is 10x the headline seat price.

    Adoption rate determines economics:

    Team sizeActive users (40% adoption)Monthly seat costToken cost (moderate usage)Total / active user
    50 seats20$1,000~$800~$90
    100 seats40$2,000~$1,600~$90
    500 seats200$10,000~$8,000~$90

    At 10% adoption (a common early-deployment reality):

    Team sizeActive usersMonthly seat costToken costTotal / active user
    100 seats10$2,000~$400~$240

    The implication: increasing adoption from 10% to 40% is a higher-ROI move than adding seats. An adoption problem looks like an economics problem but isn’t.


    What the Productivity Return Looks Like

    Industry estimates put the productivity upside at $7,800 per employee per year — but that figure only materializes when Claude is actively integrated into daily workflows, not when it’s available as an optional chat tab.

    The $7,800/employee figure comes from enterprise AI ROI research measuring time saved across knowledge work tasks. It assumes genuine integration into workflows, not passive availability. Here’s how it breaks down by role:

    Software developers (highest ROI):

    • Agentic coding with Claude Code reduces code review cycles, test writing, and boilerplate
    • Estimated 1.5–2 hours/day returned on routine coding tasks
    • At $100K loaded annual salary: ~$9,000–12,000/year in time value per developer

    Content and marketing teams:

    • Drafting, editing, research, brief writing at significantly higher speed
    • Estimated 45–90 minutes/day returned on writing-heavy tasks
    • At $75K loaded: ~$5,600–11,200/year per person

    Legal and compliance teams:

    • Contract review, policy drafting, compliance checklist work
    • Estimated 30–60 minutes/day returned
    • At $120K loaded: ~$7,500–15,000/year per lawyer or compliance analyst

    Operations and admin:

    • SOPs, reporting, email drafting, meeting prep
    • Estimated 20–30 minutes/day returned
    • At $60K loaded: ~$2,500–3,750/year

    The ROI Model

    A simple ROI model: (hours returned per user per day × working days × loaded hourly rate) − annual total cost per user = net annual value per seat.

    Example for a 50-person software team on Enterprise:

    Loaded developer salary: $120,000/year = ~$57.70/hour
    Hours returned per day (conservative): 1 hour
    Working days: 230
    Value returned per developer: 230 × $57.70 = $13,271/year
    
    Annual Enterprise cost per developer:
      Seat fee: $20 × 12 = $240
      Token cost (moderate Sonnet usage): ~$600/year
      Total per developer: ~$840/year
    
    Net ROI per developer: $13,271 − $840 = $12,431
    ROI multiple: 15.8x
    

    Even at half the productivity estimate (30 minutes/day returned), the ROI multiple remains above 7x for any knowledge worker with a loaded salary above $60K. The economics are compelling when adoption is real.


    When Enterprise Is the Right Choice vs. Team

    Choose Enterprise when you have a compliance mandate (SSO, SCIM, audit logs, HIPAA), a team above 150 users, or a negotiated consumption commitment that reduces effective per-token cost. Otherwise, Team is more predictable and sufficient.

    NeedTeamEnterprise
    SSO / SAML authentication
    SCIM provisioning
    Audit logs
    HIPAA-ready configuration
    Compliance API (export to SIEM)
    Users above 150
    Fixed predictable monthly cost
    Usage bundled in seat price

    The honest rule: buy Team until a real compliance or scale requirement forces Enterprise. If security review, identity governance, or audit trails are requirements, Enterprise is necessary. If they’re not, Team is cheaper and simpler.


    Frequently Asked Questions

    How much does Claude Enterprise cost?

    Claude Enterprise starts at approximately $20/user/month for access (billed annually, custom via sales), with token usage billed separately at standard API rates. Real total cost typically runs $60–250+/user/month depending on usage intensity and which Claude models the team uses most.

    What’s the difference between Claude Team and Enterprise?

    Team is self-serve per-seat licensing ($20 standard / $100 premium per seat/month, annual) with token usage bundled into the seat and a 150-user cap. Enterprise adds SSO, SCIM, audit logs, HIPAA support, a Compliance API for SIEM integration, no user cap, and usage billed separately at API rates. Choose Team for simplicity; choose Enterprise for compliance and governance requirements.

    What is the ROI of Claude Enterprise?

    At 1 hour of productivity returned per day per knowledge worker, the annual value per seat at a $120K loaded developer salary is approximately $13,270 — against an annual Enterprise cost of ~$840/developer. ROI multiple is roughly 15x under that assumption. At 30 minutes/day returned, the multiple is still above 7x for most knowledge worker salaries.

    Why did Anthropic unbundle tokens from Enterprise seats?

    Anthropic decoupled seat fees from token bundles in April 2026, lowering the headline seat price from $40–200/seat to $20/seat while making token usage variable. The change gives large organizations more flexibility — light users cost less, heavy users cost more — but requires better usage monitoring to forecast actual spend.

    What to Read Next

    Claude AI Pricing — All Plans and API Rates

     Claude Team vs Enterprise: Complete Comparison

     Anthropic Console: API Keys and Billing

     Current Claude Model Version Tracker

  • Claude vs GPT-5 for Developers: Which API Wins in 2026?

    Claude vs GPT-5 for Developers: Which API Wins in 2026?

    Last refreshed: August 2026

    Claude wins on coding quality and long-context reliability. GPT-5 wins on raw speed and cost per token. The right choice depends on which workload you’re optimizing for — and for most serious agentic coding workflows, Claude is the default for good reasons.

    This comparison covers the metrics that matter for production API decisions in 2026: pricing at each tier, latency benchmarks, coding benchmark scores, context window handling, and where each model actually performs better. No marketing claims — just the numbers and where they point.


    The Models Being Compared

    The relevant comparison in 2026 is Claude Sonnet 4.6 / Opus 4.8 against GPT-5 / GPT-5.5 — the mid-tier workhorses and frontier flagships from each lab.

    ModelProviderInput (per 1M tokens)Output (per 1M tokens)Context
    Claude Haiku 4.5Anthropic$0.80$41M tokens
    Claude Sonnet 4.6Anthropic$3$151M tokens
    Claude Opus 4.8Anthropic$15$751M tokens
    GPT-5OpenAI$1.25$10400K tokens
    GPT-5.5OpenAI$5$301M tokens

    The pricing gap is the first thing to understand: GPT-5 is cheaper per token than Claude Sonnet at every tier. Claude Opus is the most expensive flagship at any lab. That cost difference only makes sense if the quality difference justifies it — and for specific workloads, it does.


    Coding Performance

    Claude leads on coding benchmarks in 2026. Claude Sonnet scores approximately 77% on SWE-bench Verified versus roughly 72% for GPT-5. Claude Opus 4.8 and Fable 5 push higher still — Fable 5 is the current leader on AutomationBench.

    SWE-bench Verified measures a model’s ability to solve real GitHub issues — fixing bugs, implementing features, navigating existing codebases. It’s the most production-relevant coding benchmark available.

    Why Claude leads on coding:

    • Better multi-step refactor reliability on large codebases
    • Stronger instruction-following in complex, multi-constraint prompts
    • More consistent behavior across long agentic loops without drift
    • Claude Code and Cursor both default to Claude models — a market signal that carries weight

    Where GPT-5 is competitive on coding:

    • Faster time-to-first-token for autocomplete-style workloads
    • GPT-5.5’s terminal-based coding benchmark (Terminal-Bench: 82.7%) is strong
    • Codex — OpenAI’s coding-specific deployment — is built on GPT-5.5 and optimized for that workload

    The practical rule: for interactive coding assistance and agentic code execution, Claude Opus or Sonnet. For high-frequency autocomplete at scale where speed matters more than quality depth, GPT-5 mini or Haiku-class models.


    Latency

    GPT-5 is faster. OpenAI generally delivers 80–110 tokens per second on GPT-5; Claude Sonnet runs 60–90. Claude Haiku 4.5 is the fastest model in this comparison — first token in under 600ms on medium prompts, outpacing GPT-4.1 Mini by roughly 4x in March 2026 benchmarks.

    Latency matters differently depending on the use case:

    Use caseWhich latency mattersWinner
    Interactive chat / autocompleteTime-to-first-tokenGPT-5 (or Claude Haiku)
    Agentic batch processingThroughput, qualityClaude Sonnet / Opus
    Long-context document analysisContext handlingClaude (1M vs GPT-5’s 400K)
    Real-time voice pipelineTTFT + throughputOpenAI Realtime API (no Claude equivalent)

    For most production agentic workflows where the agent is running asynchronously, the latency difference between Claude Sonnet and GPT-5 is negligible compared to the quality difference on complex tasks.


    Context Window

    Claude’s 1M token context window is a meaningful technical advantage over GPT-5’s 400K. At 1M tokens, entire medium-sized codebases, full legal contract libraries, or complete email archives fit in a single context without chunking or retrieval engineering.

    GPT-5.5 also ships with a 1M context window, but at $5/$30 per million tokens compared to Claude Sonnet at $3/$15. For long-context workloads where you need the full window, Claude Sonnet is both more capable and cheaper than GPT-5.5.

    Practical implications of the context gap at the mid-tier (Claude Sonnet vs GPT-5):

    • Codebases over 300K tokens: Claude handles them without chunking; GPT-5 requires retrieval engineering
    • Long contract or document review: Claude reads the full document in one pass
    • Multi-session agent context: Claude Managed Agents with memory handles this; GPT-5 requires custom solutions

    Cost Comparison for Real Workloads

    OpenAI is cheaper per token at every tier, but Claude’s 90% prompt caching discount and batch API 50% discount close the gap significantly for production workloads with repeated system prompts.

    Workload cost comparison at scale:

    WorkloadClaude SonnetGPT-5Notes
    10K daily chat queries (~500 tokens avg)~$15/day~$6.25/dayGPT-5 cheaper
    Same, with 80% prompt caching~$4.50/dayNo GPT-5 equivalent discount
    100M tokens/month agentic batch~$1,500~$625GPT-5 cheaper without caching
    Same, with Claude batch API (50% off)~$750~$625Near parity

    The conclusion: for high-volume workloads with repeated context (system prompts, persistent agent instructions), Claude’s caching discounts make it competitive with GPT-5 on cost. For simple, stateless, high-frequency calls with no repeated context, GPT-5 is cheaper.


    Tool Use and Agent Reliability

    Claude is the dominant choice for agentic tool use in 2026. The Claude Agent SDK, Managed Agents platform, and Claude Code are purpose-built for autonomous multi-step workflows. OpenAI has function calling and a code interpreter, but no equivalent managed agent infrastructure.

    Where this matters in practice:

    • Claude Code and Cursor lean on Claude because the model follows multi-step instructions with better consistency
    • Claude Managed Agents runs cloud-sandboxed agents with persistent memory, built-in tools, and multi-agent coordination — OpenAI has no direct equivalent
    • For complex tool-use chains where the agent needs to recover from errors and continue, Claude’s behavior is more reliable

    Where OpenAI has an edge:

    • Computer Use is available natively on GPT-5 for web browsing and desktop control workflows
    • OpenAI’s Realtime API integrates speech-to-text, LLM, and text-to-speech in one pipeline — no Claude equivalent exists

    Which API to Choose

    Use Claude for: coding, long-context document work, agentic workflows, and anything where instruction-following quality matters more than cost per token. Use GPT-5 for: high-frequency stateless calls, voice pipeline integration, and workloads where cost is the primary constraint.

    Decision framework:

    If your primary need is…Choose
    Agentic coding and multi-step executionClaude Sonnet / Opus
    Long-context document analysis (>400K tokens)Claude Sonnet
    High-volume, cheap inference at scaleGPT-5 / Claude Haiku
    Voice + LLM pipelineOpenAI Realtime API
    Production agent with persistent memoryClaude Managed Agents
    Terminal-based coding workloadGPT-5.5 / Codex

    The most common real-world answer: Claude Sonnet for the reasoning-heavy core, Claude Haiku or GPT-5 for high-frequency auxiliary calls where speed and cost dominate. Running both APIs is normal and often optimal.


    Frequently Asked Questions

    Is Claude better than GPT-5 for coding?

    Yes, on most production coding benchmarks. Claude Sonnet scores approximately 77% on SWE-bench Verified versus about 72% for GPT-5. Claude also handles multi-step refactoring and large codebase navigation more reliably. GPT-5.5 on Terminal-Bench (82.7%) is competitive for terminal-based workflows, and OpenAI’s Codex is optimized for that use case.


    Is Claude more expensive than GPT-5?

    Per token, yes — Claude Sonnet is $3/$15 per million tokens versus GPT-5 at $1.25/$10. Claude’s prompt caching (up to 90% off cached input) and batch API (50% off) close the gap significantly for production workloads with repeated context. Opus is the most expensive flagship model available.

    Does Claude have a larger context window than GPT-5?

    Yes at the mid-tier. Claude Sonnet has a 1M token context window; GPT-5 has 400K. GPT-5.5 also offers 1M tokens but at a higher price than Claude Sonnet. For workloads requiring full-document context without chunking, Claude Sonnet is the better mid-tier choice.

    Which API is faster?

    GPT-5 is faster on raw throughput (80–110 tokens/second vs Claude Sonnet’s 60–90). Claude Haiku 4.5 is the fastest model in this comparison for time-to-first-token. For most asynchronous agentic workloads, latency differences are less significant than quality differences.


    What to Read Next

    Anthropic Console: API Keys, Billing, and the Workbench 

    Claude AI Pricing — All Plans and API Rates

     Claude API Model IDs and Strings

     How to Install Claude Code

  • How to Index Business Files Into a Local Vector Database and Query Them With Claude

    How to Index Business Files Into a Local Vector Database and Query Them With Claude

    Last refreshed: August 2026

    A local vector database Claude setup — indexed with your business documents, contracts, SOPs, client notes, and invoices — gives back the operational time lost to hunting through folders. The right answer appears in seconds, without any of those documents leaving the machine.

    This is the full build: architecture, tools, working code, what performs well in production, what breaks, and whether the ROI justifies the setup time.


    What Problem This Solves

    The problem isn’t that the documents don’t exist. It’s that finding the right one — the specific contract clause, the pricing from eight months ago, the onboarding SOP for a client — takes longer than it should, and normal search doesn’t solve it.

    File search matches keywords. It doesn’t understand that “what did we agree on for payment timing” and “net 30” are the same thing. A retrieval-augmented setup solves the semantic gap: the vector database finds relevant sections by meaning, Claude synthesizes them into a direct answer.

    The use cases where this setup pays for itself:

    • Contract and clause lookup — “What are the payment terms in the Acme agreement?” in 4 seconds vs. 3 minutes of folder navigation
    • SOP retrieval — “What’s our onboarding process for new social media clients?” surfaces the relevant runbook section directly
    • Client history — “What scope did we quote [client] last spring?” retrieves the invoice or email thread
    • Cross-document synthesis — “What are the termination clauses across all active client contracts?” — something no file search can do

    The Architecture

    The stack is ChromaDB for local vector storage, Nomic Embed for on-device embeddings via Ollama, LlamaIndex for document ingestion, and Claude Sonnet via API for the reasoning step — all files stay local, Claude only sees the retrieved chunks.

    ComponentToolWhy
    Vector databaseChromaDB (local)Free, runs on-device, persistent to disk
    Embedding modelNomic Embed via OllamaOpen-source, 8K context, no external calls
    Ingestion layerLlamaIndexHandles PDF, DOCX, MD, TXT, CSV natively
    Retrieval layerPython (custom)Readable and modifiable as needs evolve
    Reasoning layerClaude Sonnet APIMaterially better synthesis than local models
    InterfaceCLIMost queries don’t need a UI

    Why local for the vector database: The documents never leave the machine. Claude receives only the retrieved chunks — not the full corpus. For contracts, financial records, and internal communications, this is the right boundary.

    Why Claude for reasoning and not a local model: Local models (Llama 3, Mistral) handle the retrieval step comparably. They don’t handle synthesis comparably — reading five contract sections and returning a coherent, accurate answer is where Claude’s API cost is earned.


    What to Index

    Start with the 50 most-referenced documents. Get the workflow running and verified before expanding to the full corpus.

    File types that index well:

    • Contracts and agreements (PDF, DOCX)
    • Internal SOPs and runbooks (MD, DOCX)
    • Client notes and meeting logs (MD, TXT)
    • Invoices and financial records (PDF, CSV)
    • Email threads exported from Gmail (EML, TXT)

    Organize before indexing. File names and folder paths become metadata attached to each chunk. A consistent folder structure takes an hour to set up and improves retrieval quality throughout:

    /business-knowledge/
      /clients/
      /contracts/
      /operations/
      /finance/
      /communications/
      /reference/
    

    Poorly named files produce confusing retrieval results. The index is only as organized as the source files.


    Building the System

    Step 1: Install the stack

    # Ollama for local embedding
    brew install ollama
    ollama pull nomic-embed-text
    
    # Python dependencies
    pip install chromadb llama-index llama-index-embeddings-ollama anthropic
    

    Step 2: Ingest and index

    from llama_index.core import SimpleDirectoryReader, VectorStoreIndex
    from llama_index.embeddings.ollama import OllamaEmbedding
    from llama_index.vector_stores.chroma import ChromaVectorStore
    import chromadb
    
    embed_model = OllamaEmbedding(model_name="nomic-embed-text")
    
    chroma_client = chromadb.PersistentClient(path="./chroma_db")
    chroma_collection = chroma_client.get_or_create_collection("business_knowledge")
    vector_store = ChromaVectorStore(chroma_collection=chroma_collection)
    
    documents = SimpleDirectoryReader("./business-knowledge", recursive=True).load_data()
    index = VectorStoreIndex.from_documents(
        documents,
        embed_model=embed_model,
        vector_store=vector_store
    )
    
    print(f"Indexed {len(documents)} documents")
    

    500 files on an M2 MacBook Pro takes approximately 20–25 minutes. The index persists to disk — this runs once, then incrementally as files change.

    Step 3: Build retrieval and reasoning

    import anthropic
    
    def query_business_knowledge(question: str, top_k: int = 5) -> str:
        retriever = index.as_retriever(similarity_top_k=top_k)
        nodes = retriever.retrieve(question)
    
        context = "\n\n---\n\n".join([
            f"Source: {node.metadata.get('file_name', 'unknown')}\n{node.text}"
            for node in nodes
        ])
    
        client = anthropic.Anthropic()
        response = client.messages.create(
            model="claude-sonnet-4-6",
            max_tokens=1000,
            messages=[{
                "role": "user",
                "content": f"""Answer this question using only the provided business documents.
    If the answer isn't in the documents, say so clearly. Always cite the source file.
    
    Question: {question}
    
    Documents:
    {context}"""
            }]
        )
    
        return response.content[0].text
    
    print(query_business_knowledge("What are the payment terms in the Acme contract?"))
    

    Always include source attribution in the prompt. When an answer returns, the source file name makes verification fast.


    What Works Well in Production

    Cross-document synthesis is the capability that justifies this over standard search — querying across hundreds of files simultaneously to find patterns, compare terms, or surface a specific clause is something no file search does.

    Where the system consistently delivers:

    Contract and clause lookup: Specific clause retrieval across a full contract library. Synthesis across multiple contracts simultaneously (termination terms, payment terms, liability caps) returns a summary across all of them at once.

    SOP and runbook retrieval: Operational questions answered directly from internal documentation. Works best when SOPs are written in complete sentences rather than bullet fragments — the retrieval quality reflects the writing quality.

    Client history: Invoice amounts, quoted scopes, prior project notes. Email threads sometimes split across chunks in ways that lose context — use the result as a pointer to the source document, then verify.

    Cross-document pattern finding: “What are the common liability terms across our contracts?” — synthesizes across every indexed contract in one response. No file search tool does this.


    What Breaks

    The index is only as current as the last re-index. The most common production failure is stale data — documents updated after the last index run return old answers.

    Stale index: Build re-indexing into the workflow immediately. Schedule it weekly, or trigger it automatically when files are modified. Documents that change and don’t get re-indexed are the biggest reliability risk.

    Top-k ceiling: Retrieval returns the top-k chunks (default 5). A question whose complete answer requires synthesizing 20 documents gets a partial answer. Increase top_k for broad synthesis questions — at the cost of slightly more API token usage.

    Numerical calculations: The system finds financial documents reliably. It should not be trusted to calculate totals across extracted text. Use it to surface the right source documents; do the arithmetic elsewhere.

    Documentation debt: The index reveals gaps in internal documentation. SOPs written in ambiguous shorthand, contracts with undefined terms, emails with unclear context — all produce lower quality retrieval. The index reflects the quality of the underlying documents.


    Chunk Size

    512 tokens with 50-token overlap is the right starting point for mixed document types.

    Adjust based on document type:

    • Contracts (dense, long): 512–768 tokens, 100-token overlap
    • SOPs (structured, modular): 256–512 tokens, 50-token overlap
    • Emails (short, conversational): 256 tokens, 25-token overlap
    • Financial records (tabular): Parse as structured data where possible; plain text chunking loses table relationships

    Metadata Filtering at Scale

    Once the corpus exceeds ~200 files, adding metadata to chunks and filtering at query time significantly improves precision.

    # Tag at ingestion
    documents = SimpleDirectoryReader(
        "./business-knowledge",
        recursive=True,
        file_metadata=lambda filepath: {
            "document_type": filepath.split("/")[2],
            "client": filepath.split("/")[3] if len(filepath.split("/")) > 3 else "internal"
        }
    ).load_data()
    
    # Filter at retrieval
    retriever = index.as_retriever(
        similarity_top_k=5,
        filters={"document_type": "contracts"}
    )
    

    “What are our SOPs for [client]?” filtered to that client’s folder returns meaningfully more accurate results than querying the full corpus.


    ROI

    Setup takes roughly one full day. At 25 minutes saved per week on document lookups, break-even is approximately 6–8 weeks.

    ItemCost
    Setup time~8 hours
    ChromaDBFree
    Nomic Embed (Ollama)Free
    Claude Sonnet API per query~$0.003
    Monthly at 50 queries/week~$0.60
    Weekly time saved~25 minutes
    Break-even~7 weeks

    The less quantifiable return: operational confidence. Questions that previously required folder-hunting get answered in seconds. That reduces the cognitive overhead of running a multi-client operation and changes how quickly decisions get made.


    Frequently Asked Questions

    Does this send business documents to Anthropic?

    No. The vector database and embedding model run locally. Claude receives only the retrieved chunks — small sections of relevant documents — not the full corpus. For zero external calls, replace Claude with a local model, though synthesis quality will be lower.

    What file types are supported?

    LlamaIndex handles PDF, DOCX, TXT, MD, CSV, EML, and HTML natively. Other formats need conversion to plain text first.

    How long does indexing take?

    Approximately 20–25 minutes for 500 files on an M2 MacBook Pro. Subsequent re-indexing processes only changed or new files and takes a few minutes.

    What is a vector database?

    A vector database stores documents as numerical representations (embeddings) that encode meaning, not just keywords. This allows semantic search — finding relevant contract sections from a natural-language question, even when the exact words don’t match.

    Can a local model replace Claude?

    es — swap the API call for an Ollama-hosted model. Retrieval quality is comparable. Synthesis quality on complex multi-document questions is noticeably lower on current local models.

    What chunk size should be used?

    512 tokens with 50-token overlap is the right default for mixed document types. Adjust for document type: larger for dense contracts, smaller for short emails.


    What to Read Next

    Anthropic Console: API Keys and the Workbench

     Claude AI Pricing — All Plans and API Rates 

    Claude API Model IDs and Strings 

    History of Anthropic

  • Anthropic Roadmap 2027: What Comes After Claude Fable 5

    Anthropic Roadmap 2027: What Comes After Claude Fable 5

    Last refreshed: August 2026

    The Anthropic roadmap 2027 comes into focus after Fable 5 launched in June 2026 as Anthropic’s most capable widely available model — a new Mythos-class tier above Opus — and the signals from Anthropic’s research agenda, model release cadence, and safety roadmap point clearly toward what comes next.

    This is a forward-looking read grounded in public signals: what Anthropic has shipped, what they’ve said, and what the patterns suggest for 2027. It’s relevant for developers planning integrations, enterprises making multi-year platform commitments, and anyone tracking where Claude’s capabilities are heading.


    Where Anthropic Stands as of Mid-2026

    Claude has grown from a single chat model in 2021 to a four-tier family — Haiku, Sonnet, Opus, and the new Mythos class — with a 1-million-token context window, native vision, tool use, Computer Use, extended thinking, and persistent memory across managed agents.

    The model lineup as of August 2026:

    TierModelBest For
    MythosClaude Fable 5Most demanding reasoning, long-horizon agentic work
    OpusClaude Opus 4.8Flagship reasoning, fallback for Fable 5 safety filters
    SonnetClaude Sonnet 4.6Everyday development, high-volume production
    HaikuClaude Haiku 4.5Fast, cheap, high-throughput

    Fable 5 launched June 9, 2026 alongside Claude Mythos 5 — a restricted version available only through Project Glasswing for vetted cybersecurity and infrastructure partners. The distinction matters: Fable 5 is the general-availability frontier model; Mythos 5 is the same model with certain safety filters lifted for specific use cases.

    The June 2026 launch was followed by a brief government-imposed deployment pause after Amazon researchers identified a method of prompting Fable 5 to surface software vulnerabilities. Anthropic worked with government partners to add new classifiers and redeployed the model globally July 2, 2026.


    What the Fable 5 Launch Signals About 2027

    The Fable 5 launch established that Anthropic is building a two-track release model — a general-availability tier with conservative safety filters and a restricted frontier tier for vetted partners — and that cadence will continue into 2027.

    Several specific signals point forward:

    The Mythos class will expand access. Anthropic said explicitly at Fable 5 launch that Project Glasswing would expand to more vetted partners over time. The current restriction is a staged rollout, not a permanent ceiling. By 2027, Mythos-class access is likely to be more widely available to enterprise customers who can meet Anthropic’s trust and verification requirements.

    Safety classifiers will improve. Fable 5 launched with classifiers that trigger on roughly 5% of sessions, routing those queries to Opus 4.8 instead. Anthropic committed to reducing false positives “as more capable models arrive in the coming months.” More capable models arriving implies at least one Mythos/Opus generation release before end of 2026 or early 2027.

    Token unbundling sets up the next Enterprise pricing tier. The April 2026 decoupling of Enterprise seat fees from token bundles — moving from $40–200/seat with bundled tokens to $20/seat with usage billed separately — creates a cleaner structure for consumption-based tiers as model capability increases. Expect the 2027 pricing architecture to track closely with Mythos access tiers.

    Agentic infrastructure is the platform bet. Managed Agents launched April 8, 2026, Memory entered public beta April 23, and the Agent SDK (formerly Claude Code SDK) now handles the entire agent loop automatically. The infrastructure is being built to support long-running, multi-session, multi-agent workflows. The 2027 roadmap is almost certainly agentic-first.


    What Anthropic’s Research Agenda Suggests

    Anthropic’s published research priorities — interpretability, Constitutional AI, alignment, and scaling — point toward a 2027 model that is more self-correcting, better at long-horizon planning, and safer to deploy with reduced human oversight.

    Interpretability is Anthropic’s differentiator. Chris Olah’s interpretability team is the most distinct research group at any frontier lab. Their work on understanding what’s actually happening inside neural networks feeds directly into how future models are trained and where safety filters are placed. Advances in interpretability in 2026–2027 will likely show up in more precise, less overreaching safety classifiers — meaning fewer false positives on legitimate requests.

    Long-horizon agency is the capability frontier. Fable 5’s headline capability over Opus 4.8 isn’t raw reasoning quality on static benchmarks — it’s how little friction there is in multi-step agentic workflows. Fable 5’s AutomationBench scores are the clearest signal of where Anthropic is competing. The 2027 research agenda will push this further: more steps, less human intervention, better recovery from errors mid-task.

    Multi-agent coordination is early. The current Managed Agents platform supports multi-agent orchestration, but the tooling is young. 2027 is when production multi-agent deployments at scale become routine rather than experimental for most enterprise customers.


    What It Means for Developers

    Developers building on Claude in 2026 should architect for the Agent SDK and Managed Agents platform, not just the Messages API — that’s where Anthropic is investing, and it’s where the capability gains will be most significant in 2027.

    Practical implications:

    Plan for Fable 5 as the default frontier model. Opus 4.8 remains the strong fallback and the model most workflows should run on today. But product architectures that don’t account for Fable 5 as the primary reasoning layer within 12–18 months are likely to require significant refactoring.

    The Fallback API is now infrastructure. Any integration calling Fable 5 needs fallback logic configured. Anthropic’s safety classifiers will route some queries to Opus automatically — your integration needs to handle that gracefully, not treat it as an error.

    Memory changes what agents can do. Agents that don’t retain context across sessions are meaningfully less capable than those that do. The Managed Agents memory API (public beta since April 23, 2026) is the right surface to build persistent agent behavior on now, before it becomes a standard expectation.


    What to Watch

    The clearest leading indicators for 2027 Anthropic roadmap developments:

    • Project Glasswing expansion announcements — any broadening of Mythos-class access is a signal that the trust-gating model is maturing
    • Interpretability research publications — Anthropic publishes regularly; major interpretability papers tend to precede model releases by 3–6 months
    • Managed Agents general availability — currently in public beta; GA signals the platform is production-ready for the long-term
    • Context window changes — the 1M token context window is already the industry standard; what comes next is likely structural, not just larger

    Frequently Asked Questions

    What is Claude Fable 5?

    Claude Fable 5 is Anthropic’s most capable widely released model, launched June 9, 2026. It sits in the new Mythos class above the Opus tier and is built for the most demanding reasoning and long-horizon agentic work. It launched alongside Claude Mythos 5, which is restricted to vetted partners through Project Glasswing.

    What is Project Glasswing?

    Project Glasswing is Anthropic’s program for giving vetted cybersecurity and infrastructure partners access to Claude Mythos 5 — the same underlying model as Fable 5, but with certain safety filters lifted for specific use cases. Access is currently limited and application-based.

    When will Anthropic release the next model after Fable 5?

    Anthropic has not announced a release date. Their historical cadence — roughly one major model generation per 6–9 months — suggests a 2026 Q4 or early 2027 release is plausible. Anthropic has stated that more capable models are coming and that safety classifiers will improve as they arrive.

    What is Claude Managed Agents?

    Claude Managed Agents is Anthropic’s managed infrastructure for running autonomous Claude agents in cloud sandboxes, launched in public beta April 8, 2026. It handles session management, tool execution, credential management, and multi-agent coordination without requiring developers to build that infrastructure themselves. Memory for Managed Agents entered public beta April 23, 2026


    What to Read Next

    History of Anthropic 

    Claude AI Pricing — All Plans and API Rates 

    Current Claude Model Version Tracker 

    Claude API Model IDs and Strings

  • Anthropic’s Real Play Isn’t a Chatbot — It’s the Invisible Agent Layer Inside Every Tool You Use

    Anthropic’s Real Play Isn’t a Chatbot — It’s the Invisible Agent Layer Inside Every Tool You Use


    Claude Managed Agents is the product. Slack, Notion, Jira, and Asana are just the interface. Anthropic is building the invisible execution layer that powers the next generation of enterprise software.

    There is a pattern emerging in enterprise AI that most people are reading wrong. They see Anthropic launch Claude Tag in Slack and think “chatbot upgrade.” They see Claude show up inside Notion and think “productivity feature.” They see AI agents appear in Jira and Asana and think “automation plugin.”

    They are missing the architecture underneath all of it.

    Anthropic is not building a better chatbot. It is building the invisible agent runtime that sits beneath every collaboration tool your team already uses. The company’s Claude Managed Agents (CMA) platform — launched in public beta on April 8, 2026 — is the infrastructure layer that makes this possible. And the speed at which partners are embedding it tells you everything about where enterprise software is heading.

    What Claude Managed Agents Actually Is

    Claude Managed Agents is a set of composable APIs for building and deploying production AI agents on Anthropic’s cloud infrastructure. The service handles sandboxed code execution, session persistence, credential management, scoped permissions, and end-to-end tracing — all the operational complexity that previously kept agents stuck in proof-of-concept limbo.

    The architecture rests on three primitives: the Agent (configuration and behavior), the Environment (sandboxed execution), and the Session (the event log that tracks everything the agent does). What makes this interesting architecturally is how Anthropic decoupled the “brain” from the “hands.” Claude’s reasoning runs on Anthropic’s own infrastructure while the code execution sandbox spins up independently — and in parallel. The brain starts reasoning immediately while the sandbox provisions, delivering roughly 60% faster time-to-first-token at the p50 level and over 90% faster at p95, according to Anthropic’s engineering team.

    Pricing follows a transparent model: standard Claude API token rates plus $0.08 per session-hour of active runtime during the current beta period. Runtime is measured to the millisecond and only accrues while the agent is actively executing — idle time waiting for input or tool confirmations does not count.

    For teams that need to keep execution inside their own perimeter, CMA supports self-hosted sandboxes through partners including Cloudflare, Daytona, Modal, and Vercel, or custom VPC deployments. MCP tunnels allow agents to connect to private Model Context Protocol servers inside your network without exposing them to the public internet. A Vaults system keeps credentials out of the sandbox entirely using envelope encryption. And a feature called Dreaming runs scheduled reviews of past sessions to curate agent memory — essentially letting agents learn from their own operational history.

    The Embedded Layer: Where CMA Actually Lives

    The real story is not the infrastructure. It is where that infrastructure shows up. In the ten weeks since CMA launched, Anthropic has embedded its agent runtime inside the collaboration tools that enterprises already depend on. This is not a roadmap — these integrations are live or in active beta.

    Slack: Claude Tag as Persistent Team Member

    Claude Tag, launched June 23, 2026, replaces Anthropic’s original Claude in Slack integration with something fundamentally different. This is not a chatbot you summon with a slash command. It is a persistent AI team member that lives in your channels, builds memory across conversations, and can take initiative through what Anthropic calls “ambient mode” — proactively surfacing information, following up on forgotten threads, and keeping teams updated across the organization.

    Claude Tag is multiplayer by design: one Claude identity per channel, accessible to everyone, with the ability to hand off half-finished tasks between team members. It runs on Claude Opus 4.8, Anthropic’s most capable model released May 28, 2026. And internally, Anthropic reports that Claude Tag is already approving and incorporating 65% of the code changes their product team submits. The existing Claude in Slack app will be retired on August 3, 2026. Claude Tag is available on Enterprise and Team plans.

    Notion: Claude as External Agent

    On May 13, 2026, Notion launched its Developer Platform version 3.5, which introduced the External Agents API. This API lets AI agents — including Claude — operate inside your Notion workspace as first-class participants. They can read pages, write to databases, create tasks, trigger automations, and be @-mentioned directly in documents. Claude operating through this API can chain actions together: read a project brief, check the task database for related work, draft a new document, and create a linked task entry — all in a single session, running on CMA infrastructure with full sandboxing.

    Asana: AI Teammates

    Asana built AI Teammates on CMA — agents that pick up assigned tasks inside projects, draft deliverables, and hand back outputs for human review. Specialist agents handle specific workflows: the Campaign Brief Writer turns scattered notes into structured briefs, the Workflow Optimizer identifies process gaps and builds automations, and the Compliance Specialist checks work against regulatory standards. Asana’s CTO said CMA let them ship these features “dramatically faster” than any prior approach to agent development.

    Atlassian: Claude Agent for Jira

    Atlassian released Claude Agent for Jira, built on CMA infrastructure, which lets teams assign work items directly to Claude from the Jira UI. The agent clones the repository, analyzes the codebase, implements changes on an independent branch, pushes the code, and opens a draft pull request — streaming real-time status updates back to the Jira work item throughout the process.

    Sentry: From Bug Detection to Merge-Ready PR

    Sentry’s existing AI debugging agent, Seer, already used Claude for root cause analysis. With CMA, Sentry extended the workflow from diagnosis to automated fixing — the agent takes Seer’s root cause output, generates a fix, opens a branch with the changes, and creates a pull request for developer review. Sentry processes over one million root cause analyses per year and provides near-immediate reviews on over 600,000 pull requests per month. The CMA integration was built by a single engineer in weeks, eliminating months of custom agent runtime development.

    Rakuten: Specialist Agents Across the Enterprise

    Rakuten deployed specialist agents across product, sales, marketing, and finance using CMA, with each agent deployed in approximately one week. Agents plug into Slack and Teams, letting employees assign tasks and receive deliverables including spreadsheets, slides, and applications. In the pilot, Rakuten reported a 97% drop in critical first-pass errors, with cost down more than 30% and latency reduced by 34%, without any loss in output quality.

    KPMG: Global Professional Services Alliance

    On May 19, 2026, KPMG and Anthropic announced a global alliance and launched “Digital Gateway Powered by Claude.” The partnership embeds Claude, Cowork, and CMA directly into KPMG’s client delivery platform, with an initial focus on tax and private equity clients. Building an AI agent for tax regulation workflows previously took weeks and required switching between multiple tools. With CMA integrated into Digital Gateway, KPMG says the same capability takes minutes. The alliance extends to KPMG’s 276,000-person global workforce.

    The Strategic Pattern: Agent Runtime as a Service

    Step back from the individual integrations and the strategic pattern becomes clear. Anthropic is not trying to own the interface. It is deliberately positioning CMA as the execution layer underneath interfaces that other companies own. Slack owns the messaging UI. Notion owns the workspace UI. Jira owns the project tracking UI. Anthropic owns the agent brain that powers all of them.

    This is a fundamentally different strategy from its two largest competitors.

    OpenAI chose vertical integration. When OpenAI launched Workspace Agents on April 22, 2026, it positioned ChatGPT itself as the central hub — a no-code successor to custom GPTs that connects to Slack, Salesforce, Google Drive, and Notion through plugins. Agents are created inside ChatGPT, accessed from ChatGPT, and managed through ChatGPT. OpenAI wants to own the surface area.

    Google chose platform depth. At Google Cloud Next on April 22, 2026, Google unveiled the Gemini Enterprise Agent Platform — a reimagined evolution of Vertex AI — alongside Workspace Intelligence, a semantic unifying layer that connects data across Docs, Slides, Gmail, and the broader Google Cloud ecosystem. Google’s agent platform supports 200+ models including Claude, and the Agent2Agent (A2A) protocol enables distributed peer-to-peer agent communication. Google is leveraging its data moat and distribution at the platform level.

    Anthropic chose tool-centric orchestration. Rather than owning the UI (OpenAI) or the platform (Google), Anthropic is embedding its agent runtime into every tool through composable APIs and the Model Context Protocol. The platform you use becomes irrelevant — whether it is Slack, Notion, Jira, Asana, or Sentry — because the agent brain running underneath is Claude on CMA.

    This is the agent-as-a-service model. And it may be the most defensible position of the three, because it does not require users to change their behavior or migrate to a new platform. The agent shows up where they already work.

    What the Numbers Say About Enterprise Agent Adoption

    The macro context supports Anthropic’s timing. Gartner predicts that 40% of enterprise applications will include embedded task-specific agents by the end of 2026, up from less than 5% in 2025. McKinsey’s April 2026 analysis found that agentic AI can enable automation of 60 to 80 percent of routine infrastructure work over time, translating to a 20 to 40 percent run-rate cost reduction in initial deployments.

    The gap between experimentation and production remains the defining challenge. Industry research compiled from major firms shows that nearly four in five enterprises have experimented with or deployed agents in some form, but fewer than one in nine are running them in production at a scale that generates measurable business value. For the agents that do reach production, the average return on investment is 171% — though 19% of deployments never reach payback at all.

    That production gap is exactly what CMA is designed to close. The infrastructure burden — sandboxing, session persistence, credential isolation, error recovery, observability — is the bottleneck. Engineering teams routinely dedicated significant senior engineering resources for months before a single agent reached production. CMA eliminates that layer entirely, which is why partners like Asana, Sentry, and Rakuten report shipping production agents in days or weeks rather than quarters.

    What This Means for Businesses Already Using These Tools

    If your organization uses Slack, Notion, Jira, or Asana — and statistically, you use at least two of them — you are about to encounter Claude whether you planned to adopt it or not. This is not a technology decision your IT team is making. It is a feature that your existing vendors are shipping.

    The practical implications are significant. Claude Tag in Slack means your team channels will have an AI participant that remembers past conversations, can be handed tasks asynchronously, and may proactively surface information. Claude in Notion means your project documentation, databases, and task boards can be read, analyzed, and acted upon by an agent that chains actions together. Claude Agent for Jira means development tickets can be assigned to an AI that clones your repo, writes code, and opens pull requests.

    For agencies and service providers managing client work across multiple tools, the embedded agent layer changes the economics fundamentally. Work that previously required a human to context-switch between Slack, Notion, and a project management tool — reading a brief here, updating a task there, drafting a document somewhere else — can be handled by an agent that operates across all of them simultaneously. The coordination tax that consumes a substantial share of knowledge work time is the exact problem embedded agents are built to solve.

    The companies that benefit most will be the ones that have clean operational systems — structured task boards, documented processes, well-organized project databases — because agents can only act on information they can read. Messy Notion workspaces and disorganized Jira boards will limit what agents can accomplish. Operational hygiene just became a competitive advantage.

    What This Means for Solo Operators Already Running Agent Infrastructure

    There is a specific audience that should be paying very close attention to CMA: the solo operators and small agency owners who have already built their own agent stacks from scratch. If you are running scheduled Claude tasks on a GCP Compute Engine VM, connecting to WordPress via REST API proxies, piping work orders through Notion, monitoring Gmail for client replies, and publishing content through MCP-connected pipelines — you have already built a version of what CMA is productizing.

    The economics question is worth doing the math on. A lightweight GCP VM running 24/7 to host recurring agent tasks — news desk monitors, outreach reply checks, newsletter extraction, scheduled content audits — costs a fixed monthly rate whether the agents are actively working or sitting idle. CMA at $0.08 per session-hour of active runtime only charges when agents are executing. For tasks that run for a few minutes every few hours, the per-session billing model could be substantially cheaper than keeping a VM warm around the clock. A task that runs for ten minutes six times a day would cost roughly $0.08 per day on CMA, versus the cost of a VM instance that never sleeps.

    But the migration path is not ready yet, and solo operators should understand exactly where the gaps are before making any infrastructure decisions.

    The biggest gap is MCP tunnels. CMA’s ability to connect agents to private MCP servers inside your network is still in research preview — not production-ready. If your agent stack depends on a private WordPress REST API proxy, a Notion workspace connected via MCP, or any internal tool that is not exposed to the public internet, CMA cannot reach it today. The Vaults system for credential management is promising, but it does not solve the network connectivity problem for self-hosted infrastructure.

    The second gap is orchestration control. Solo operators who have built their own agent infrastructure typically have precise control over scheduling, retry logic, error handling, and the exact sequence of tool calls. CMA’s Dreaming feature — which reviews past sessions to curate agent memory — is an interesting approach to agent learning, but it is not the same as having direct control over a cron job that fires at 6:00 AM, checks three data sources in a specific order, and writes results to a specific Notion database with a specific schema.

    The thesis for solo operators is straightforward: CMA is almost certainly the future migration path for self-hosted agent infrastructure. The economics favor it for intermittent workloads, the managed security and sandboxing eliminate operational risk you are currently carrying yourself, and the session persistence model solves problems that custom agent runtimes handle poorly. But the plumbing — particularly MCP tunnels to private infrastructure — is not production-ready. Track it closely. Do not migrate yet. When MCP tunnels graduate from research preview to general availability, revisit the math and the connectivity story. That is the trigger point.

    The Risk Nobody Is Talking About

    There is a tension in this model that deserves attention. When Claude operates as an invisible layer inside tools you already trust, the boundary between the tool’s native capabilities and the AI agent’s actions blurs. A Jira ticket that was “completed” might have been implemented by Claude, reviewed by a human for thirty seconds, and merged. A Notion project plan that looks thorough might have been generated by an agent that filled in the sections with plausible-sounding content.

    The embedded model works precisely because it reduces friction — but reduced friction also means reduced scrutiny. Organizations adopting embedded agents need to build review processes that match the speed at which agents can produce output. The 171% average ROI from agent deployments accounts for the value created, but it does not account for the subtle quality risks of production work generated by systems that are confident, fluent, and occasionally wrong.

    Anthropic has built guardrails into CMA — sandboxed execution, credential isolation, session logging — but the governance layer for reviewing agent output at enterprise scale is still largely unsolved. This is a space where internal operational discipline matters more than the technology itself.

    Where This Goes Next

    Claude Tag launched on Slack first. Anthropic has indicated plans for wider rollout beyond Slack. If the pattern holds, expect Claude Tag’s persistent team member model to appear in Microsoft Teams, Discord, and any other collaboration surface where teams coordinate work.

    The CMA primitives are designed to be composable, which means the partner integration list will grow rapidly. Any SaaS company with an API and a workflow that involves reading context, making decisions, and taking actions is a candidate for CMA integration. Customer support platforms, CRM systems, design tools, analytics dashboards, HR systems — the addressable surface is essentially every tool that knowledge workers touch.

    Gartner’s long-term projection estimates that agentic AI could drive approximately 30% of enterprise application software revenue by 2035, surpassing $450 billion. If Anthropic’s embedded strategy succeeds, a meaningful slice of that revenue flows through CMA as the underlying runtime — regardless of whose logo is on the interface.

    The chatbot era is ending. The embedded agent era is starting. And Anthropic is betting that the company that owns the invisible execution layer wins the market, even if no end user ever sees its name.

    Frequently Asked Questions

    What are Claude Managed Agents (CMA)?

    Claude Managed Agents is a set of composable APIs launched by Anthropic on April 8, 2026 in public beta. CMA lets developers build and deploy production AI agents on Anthropic’s cloud infrastructure, handling sandboxed code execution, session persistence, credential management, and end-to-end tracing. The architecture separates the “brain” (Claude reasoning) from the “hands” (code execution sandbox), enabling parallel processing and faster agent responses.

    How much do Claude Managed Agents cost?

    During the current public beta, CMA pricing is standard Claude API token rates plus $0.08 per session-hour of active runtime. Runtime is measured to the millisecond and only accrues while the agent is actively executing — idle time does not count. GA pricing has not been finalized and may differ from the beta rate.

    What is Claude Tag in Slack?

    Claude Tag is Anthropic’s persistent AI team member for Slack, launched June 23, 2026. Unlike a traditional chatbot, Claude Tag lives in channels, builds memory across conversations, takes initiative through ambient mode, and works asynchronously. It is multiplayer — one Claude identity per channel that all team members interact with. Claude Tag runs on Claude Opus 4.8 and is available on Enterprise and Team plans. It replaces the original Claude in Slack app, which retires August 3, 2026.

    Which tools have Claude Managed Agents embedded?

    As of June 2026, CMA is embedded in Slack (via Claude Tag), Notion (via the External Agents API), Asana (AI Teammates), Atlassian Jira (Claude Agent for Jira), and Sentry (extending the Seer debugging agent). Enterprise deployments include Rakuten (specialist agents across product, sales, marketing, and finance) and KPMG (Digital Gateway Powered by Claude for tax and private equity clients).

    How does Anthropic’s agent strategy differ from OpenAI and Google?

    Anthropic uses a tool-centric orchestration approach, embedding its agent runtime inside existing tools via composable APIs and the Model Context Protocol (MCP). OpenAI chose vertical integration with Workspace Agents, positioning ChatGPT as the central hub. Google chose platform depth with the Gemini Enterprise Agent Platform and Workspace Intelligence semantic layer. Anthropic’s approach does not require users to change platforms — the agent shows up where they already work.

    What percentage of enterprise apps will have embedded AI agents by end of 2026?

    Gartner predicts that 40% of enterprise applications will include embedded task-specific agents by the end of 2026, up from less than 5% in 2025. However, fewer than one in nine enterprises currently run agents in production at scale, suggesting significant growth ahead.

    Can Claude Managed Agents run inside a private network?

    Yes. CMA supports self-hosted sandboxes through partners including Cloudflare, Daytona, Modal, and Vercel, or custom VPC deployments. MCP tunnels allow agents to connect to private Model Context Protocol servers inside your network without public exposure. A Vaults system keeps credentials out of the sandbox using envelope encryption.