Anchor fact: Workers for Agents is in developer preview as of April 2026, accessible via the Notion API but not exposed through any consumer-facing UI yet. Workers run server-side JavaScript and TypeScript, sandboxed via Vercel Sandbox, with a 30-second execution timeout, 128MB memory limit, no persistent state, and outbound HTTP restricted to approved domains.
What is Notion Workers for Agents?
Workers for Agents is Notion’s code execution environment for AI agents, in developer preview as of April 2026. Workers run server-side JavaScript and TypeScript functions that an agent calls when it needs to compute, query a database, transform data, or call an approved external API. Workers are sandboxed (30-second timeout, 128MB memory, no persistent state) and run on Vercel Sandbox infrastructure.
The 60-second version
Workers turn Notion AI from a text layer into a compute layer. Before Workers, Notion AI could read pages and write text. It couldn’t run code, couldn’t transform data, couldn’t reliably call external APIs. With Workers, an agent can offload computational tasks to a sandboxed JavaScript or TypeScript function — running for up to 30 seconds in 128MB of memory, with outbound HTTP restricted to approved domains. It’s the upgrade that makes Notion agents capable of real workflow automation, not just document assistance.
Why Workers matter
Why Workers matter.
Three things change when agents can call code:
1. Real database queries. Before Workers, an agent could read pages but couldn’t reliably do “give me all rows where date is in the next 7 days and owner is unassigned.” With Workers, that’s a one-line query that returns structured data the agent uses in its response.
2. Approved external API calls. An agent can fetch live exchange rates, look up shipping status, query an internal CRM, or pull from any service exposed through an approved domain. The agent doesn’t make the call directly — it delegates to a Worker that does the call and returns the result.
3. Multi-step transformation chains. Read CSV → transform → enrich → write back to a database. Each step is a Worker. The agent orchestrates the chain. This is the pattern that lets agents handle real ops workflows that previously required Zapier, n8n, or custom code.
The technical constraints worth knowing
Workers are not Lambda. They have intentional limits:
30-second execution timeout. Anything longer needs to be split into smaller Workers or moved off-platform. No long-running batch jobs.
128MB memory limit. Streams and chunked processing only for large data. No loading 500MB CSVs into memory.
No persistent state between calls. Each Worker invocation is fresh. State lives in Notion databases or external services, not in the Worker.
Outbound HTTP restricted to approved domains. You declare which domains a Worker can reach. This is a security feature, not a limitation to fight.
Sandboxed via Vercel Sandbox. Workers run on Vercel’s untrusted-code infrastructure. Performance is solid; cold starts exist.
What you need to use Workers
This is not a point-and-click feature. Requirements:
A Notion developer account
A Notion integration set up
Familiarity with the agent configuration format
API access — Workers are API-only as of April 2026
If you’ve never built on the Notion API, Workers aren’t your starting point. Standard agents and skills are. Workers are the next step once those don’t go far enough.
Three Worker patterns to start with
Three Worker patterns to start with.
1. The data-fetch Worker. Agent says “I need the current value of X.” Worker calls an approved external API, parses the response, returns a structured value. Common pattern: looking up live data the agent doesn’t have access to natively.
2. The transform-and-write Worker. Agent passes structured input to a Worker. Worker reshapes the data — formatting dates, normalizing strings, computing derived fields — and writes the result to a Notion database row. Common pattern: cleaning incoming form submissions before they land in the CRM.
3. The chain-orchestration Worker. A Worker that calls other Workers in sequence, collecting results and returning a synthesized output. Common pattern: a multi-step intake process where each step needs different logic.
Why this is the more interesting story than May 3
The May 3 credit cliff is the news story. Workers are the strategic story. Workers are why credits exist — Notion can’t ship “an agent that calls any code you want and any API you want” on a flat fee. Credits make Workers viable as a product. The pricing news is the boring infrastructure that supports the interesting capability.
If you’re a developer or an agency building on Notion, Workers reshape what’s possible. A custom Notion deployment for a client used to mean “we set up databases and trained the team.” Now it can mean “we set up databases, trained the team, and built five Workers that handle their specific workflows.”
What’s still missing
What’s still missing.
Three gaps in the current developer preview worth tracking:
No consumer UI. Workers are API-only. End users can’t build them in the Notion app. This will change.
Limited debugging. Errors in Workers surface as agent errors. Better tooling for inspecting Worker execution is on the roadmap.
Sandbox boundaries are evolving. Approved domain lists, memory limits, and timeout limits are likely to relax over time. Build with current limits; don’t bet on them staying fixed.
Workers turn Notion AI from a text layer into a compute layer.
Sources
Notion 3.4 part 2 release notes (April 14, 2026)
Vercel blog — How Notion Workers run untrusted code at scale with Vercel Sandbox
Notion API documentation — Workers for Agents (developer preview)
Continue the journey
This article is part of the May 3 Cliff Decision journey-pack on Tygart Media. Here’s where to go next:
Anchor fact: Custom Agents are powerful but inappropriate for tasks involving novel judgment, regulated content, sensitive personnel matters, or work where the cost of being wrong exceeds the cost of doing it manually.
When should you not use a Notion AI agent?
Don’t use Notion agents for tasks requiring novel judgment about people, compliance-sensitive output (legal, medical, financial guidance), one-off work that won’t repeat, or any decision where the cost of being wrong is higher than the cost of doing the work manually.
The 60-second version
Notion agents are a hammer. Not everything is a nail. The honest list of tasks that should stay manual is longer than most operators want to admit. Performance reviews. Hiring decisions. Compliance-sensitive drafting. Anything that gets sent to a regulator or a lawyer. One-off work. Anything where the value of doing it yourself is the thinking, not the output. The discipline of saying “not this one” is what separates operators who use AI from operators who use AI badly.
Five categories that stay manual
Five categories that stay manual.
1. Decisions about specific humans. Performance reviews, hiring choices, conflict mediation, layoff decisions. The agent can summarize and surface evidence; it shouldn’t draft the decision. The risk isn’t that the output is wrong — it’s that the decision-maker outsources the moral weight of the call. Don’t.
2. Regulated or compliance-sensitive output. Legal language, medical guidance, financial advice, anything that gets reviewed by a regulator. Use AI to draft inputs to a human reviewer. Never ship the AI output as final.
3. Novel work without precedent. “Plan our entry into a new market.” “Write our crisis response if X happens.” Agents synthesize from existing patterns. They struggle when the situation has no analog in your workspace.
4. One-off tasks. Building a Custom Agent for a task you’ll do once is more work than just doing the task. The investment in setup (prompt, scope, rubric, review) only pays back across many repetitions.
5. Work where doing it is the point. Strategic thinking. Writing meant to clarify your own ideas. Reflection journals. The output isn’t the value; the doing is. AI shortcuts the doing, which destroys the value.
The dangerous middle category
The dangerous middle category.
Worse than tasks that obviously shouldn’t be agent work are tasks that look like agent work but aren’t. Examples:
“Draft client emails” — sounds like a clear agent task, but the relationship cost of off-tone email outweighs the time saved
“Summarize our team’s wins for the board” — looks easy, but framing matters and an agent’s framing is generic
“Write our company values” — agents can produce values; only humans can mean them
The test: if the value of the output depends on being recognizably yours, agent involvement should be limited to research and drafting, not production.
How to decide
How to decide.
Three questions before launching a new Custom Agent:
Will I do this task at least 20 times in the next year? (No → don’t build an agent.)
Is the cost of a wrong output bounded? (No → don’t automate it.)
Is the value in the output, not the doing? (No → don’t outsource the doing.)
If any answer is no, the task stays manual. That’s not a failure of AI. That’s discipline.
AI shortcuts the doing, which destroys the value.
Sources
Tygart Media editorial line
Operator practice notes
Continue the journey
This article is part of the May 3 Cliff Decision journey-pack on Tygart Media. Here’s where to go next:
Anchor fact: Notion Custom Agents cost $10 per 1,000 credits starting May 4, 2026. Credits reset monthly with no rollover. Simple agent runs use a handful of credits; complex multi-step runs can use dozens to hundreds.
How do you calculate ROI on a Notion Custom Agent?
Multiply the human-equivalent time saved per agent run by the dollar value of that time, subtract the credit cost per run (at $10/1000 credits starting May 4, 2026), then multiply by run frequency. An agent that saves 30 minutes of work per run at $50/hour, costs 5 credits ($0.05) per run, and runs daily produces ~$700/month in net value.
The 60-second version
Most operators don’t do the math because the math feels small. It isn’t. A Custom Agent that runs daily and saves 30 minutes of $50-an-hour work produces about $750/month in time savings and costs maybe $1.50 in credits. The ratio is so favorable for the right agents that the real ROI question isn’t whether agents pay back — it’s which agents to retire because the math doesn’t clear. After May 4, the bottom of the agent fleet stops being free. That’s good. That’s how you stop running agents that weren’t earning their keep.
The simple formula
The simple formula.
For any Custom Agent:
Time saved per run (minutes) × frequency (runs per month) × hourly value ($/hour ÷ 60) = monthly value
Credits per run × frequency × $0.01 (since $10/1000 = $0.01/credit) = monthly cost
Monthly value − monthly cost = net ROI
Three worked examples:
Example 1 — The weekly digest agent. Saves 45 minutes/run, runs 4×/month, your hourly value is $75. Monthly value: 45 × 4 × ($75/60) = $225. Credits: ~20/run × 4 × $0.01 = $0.80. Net: $224.20/month. Keep it.
Example 2 — The lead enrichment agent. Saves 5 minutes/run, runs 200×/month (every new lead), hourly value $50. Monthly value: 5 × 200 × ($50/60) = $833. Credits: ~3/run × 200 × $0.01 = $6. Net: $827/month. Keep it.
Example 3 — The exploratory analysis agent. Saves 15 minutes/run, runs 2×/month, complex multi-step (~80 credits). Monthly value: 15 × 2 × ($50/60) = $25. Credits: 80 × 2 × $0.01 = $1.60. Net: $23.40/month. Keep it, but barely. If credit cost rises or run complexity grows, retire it.
Where the math turns negative
Where the math turns negative.
Three patterns where the ROI math fails:
The fancy agent that runs occasionally. Complex agents cost dozens to hundreds of credits per run. Low frequency means the per-month cost is small but so is the value. Net is small. Better as a manual prompt.
The agent that needs human review on every output. If you review 100% of the output anyway, the time saved is partial. Reduce the apparent monthly value by 40-60%. Many agents stop clearing the bar with that haircut.
The agent that runs but the output isn’t used. This is the silent killer. Credits consumed, no value extracted. The fix is monthly observation: which agent outputs do you actually open?
The portfolio approach
Treat your Custom Agents as a portfolio. Three categories:
Anchor fact: Custom Agents are available on Business and Enterprise plans only. They run autonomously on triggers or schedules, can work for up to 20 minutes per task across hundreds of pages, and starting May 4, 2026, consume Notion Credits at $10 per 1,000.
Do you need Notion Custom Agents or is basic Notion AI enough?
Basic Notion AI handles inline drafting, summaries, and reactive prompts within a page. Custom Agents add proactive execution — running on schedules or triggers, working autonomously for up to 20 minutes, and using skills and Workers. Choose Custom Agents only if you have recurring autonomous workflows that justify Business-plan pricing and Notion Credit consumption.
The 60-second version
Most operators don’t need Custom Agents. They think they do because the marketing makes Custom Agents sound essential, but the honest answer is that basic Notion AI plus standard agent prompts cover most knowledge-work needs. Custom Agents earn their cost only when you have specific, repeating, autonomous work — things that run on a schedule or trigger without you starting them. If you don’t have that pattern in your workflow, you’re paying for capability you won’t use.
The honest comparison
The honest comparison.
Basic Notion AI (included on Plus, Business, Enterprise plans):
Custom Agents (Business and Enterprise plans only):
Everything above, plus:
Runs on schedules or triggers without prompting
Can work autonomously for up to 20 minutes per task
Spans hundreds of pages in a single run
Skills can be attached for repeatable workflows
Workers integration (developer preview) for code execution
Can integrate with Calendar, Mail, Slack at agent level
After May 4, 2026: consumes Notion Credits at $10/1000
When Custom Agents are worth it
When Custom Agents are worth it.
Five workflow patterns where Custom Agents pay off:
1. Recurring deliverables. Weekly status reports, monthly board prep, daily standups. If you produce the same shape of document on a schedule, an agent that runs Friday at 4 PM and drops the draft in your inbox is worth real money in time saved.
2. Continuous database enrichment. A CRM that needs new leads scored, categorized, and routed within minutes of arrival. A content database that needs incoming articles tagged and summarized. An ops database that needs items checked for SLA breaches.
3. Cross-source synthesis on demand. “Pull everything from the last two weeks across Slack, Calendar, and our project pages and tell me what’s at risk.” This is a 20-minute autonomous task that would take a human two hours.
4. Multi-step workflows with handoffs. Triage incoming → route to owner → draft response → flag exceptions. The chain is what makes it agent work, not assistant work.
5. Off-hours and overnight work. If you’d benefit from work happening while you sleep, agents are the only Notion layer that can do it. Reactive AI sits idle until you arrive.
When basic Notion AI is enough
When basic Notion AI is enough.
Most knowledge workers fit here:
Solo writers and researchers who need help drafting and summarizing
Teams of fewer than 10 where work is mostly real-time collaborative
Workflows where the AI is occasional, not scheduled
Anyone on Plus plan (Custom Agents aren’t available anyway)
Anyone whose AI usage is “I ask, it answers” — that’s reactive, not agentic
If you’re in this group, upgrading to Business for Custom Agents is paying for capacity you won’t use. Stay with basic AI and revisit when the workflow pattern changes.
The cost calculus after May 4
Before May 4, 2026, Custom Agents are free to try on Business and Enterprise. After, every run consumes credits at $10 per 1,000. Real numbers:
A simple agent run (single-page summary): typically a handful of credits — pennies
A complex multi-step run (synthesis across many pages, multiple skills chained): can run into the dozens or hundreds of credits — measurable dollars
A daily scheduled agent that runs 30 days/month at moderate complexity: budget low tens of dollars per agent per month
Math gets serious when you have many agents running daily. A workspace with 10 active Custom Agents can easily consume hundreds of dollars per month in credits on top of Business-plan seat fees. That’s the ROI conversation that turns “I’m experimenting with agents” into “I run a small fleet on a budget.”
The decision framework
Walk yourself through these four questions:
Do you have recurring work on a schedule? No → basic AI is fine.
Are you on Business or Enterprise? No → Custom Agents aren’t available. Upgrade or stay with basic.
Does the time saved per agent run, multiplied by frequency, exceed the credit cost? No → basic AI plus manual prompts is cheaper.
Are you willing to manage the credit pool monthly? No → don’t take on the operational overhead.
If all four are yes, Custom Agents earn their place. If any is no, basic Notion AI is the right call.
Anchor fact: Custom Agents are free to try through May 3, 2026. Starting May 4, they require Notion Credits at $10 per 1,000 credits, and access stays gated to Business and Enterprise plans.
What changes for Notion Custom Agents on May 3, 2026?
Custom Agents are free to try through May 3, 2026 on Business and Enterprise plans. Starting May 4, agents require Notion Credits at $10 per 1,000 credits. Credits are workspace-shared, reset monthly, and don’t roll over. If credits hit zero, every Custom Agent in the workspace pauses until an admin tops up.
The 60-second version
If you’re running Notion Custom Agents on a free trial right now, you have until May 3, 2026 before the meter starts. On May 4, agents stop running unless your workspace admin has bought Notion Credits at $10 per 1,000 credits. Credits reset monthly. They don’t roll over. Custom Agents stay locked to Business and Enterprise plans only — Free and Plus plans don’t get them at all.
The decision in front of you isn’t “should I keep using Custom Agents.” It’s three smaller decisions stacked: whether to be on the right plan, whether to budget credits, and whether the agents you’ve already built earn their keep at the new price.
This article walks through each one in operator terms.
What actually changes on May 4
What actually changes on May 4.
Before May 3:
Custom Agents run for free on Business and Enterprise plans (including Business trials)
No credit accounting
You can build, test, and run as much as your plan allows
On and after May 4:
Custom Agents consume Notion Credits per task
Credits cost $10 per 1,000, billed as a workspace-level add-on
Credits are shared across the workspace, not per-seat
Credits reset every month with no rollover
If the credit pool empties, every Custom Agent in the workspace pauses until an admin tops up
Agents stay on Business and Enterprise plans only — no migration path to Free or Plus
The mechanic worth pausing on: shared, non-rolling, hard-pause-on-zero. That’s not a soft throttle. If your workspace runs out mid-month, the agent that drafts your weekly board update doesn’t degrade gracefully. It stops. An admin has to log in and add credits before anything resumes.
Why this matters more than it sounds
Most of the coverage of this transition reads it as a pricing announcement. It’s actually a posture announcement. Notion is saying: agents are real infrastructure, real infrastructure has metering, and metering changes how teams use it.
Three knock-on effects worth thinking about:
1. The “leave it running and forget about it” pattern dies. Free trial behavior — point an agent at a database, walk away, come back a week later, see what it did — becomes expensive behavior. Every autonomous run consumes credits. If you’ve built agents that run on schedules or triggers, that scheduled work is now a line item.
2. Agent ROI becomes a real conversation. Up to now, the question was “does this agent save me time?” Starting May 4, the question is “does this agent save me time at a credit cost lower than what my time is worth?” That’s a much sharper test, and a fair number of trial-era agents won’t survive it.
3. The build-vs-prompt decision shifts. A one-off prompt to Notion AI inside a doc still runs on plan-included AI. A Custom Agent — even doing similar work — runs on credits. For repetitive work that’s worth automating, the agent still wins. For occasional work, you may quietly retreat to manual prompts.
What you should do this week
What you should do this week.
This is the operator’s checklist, in priority order.
1. Audit every Custom Agent you’ve built
Open your workspace’s Custom Agents list. For each one, write down four things:
What does it do?
How often does it run?
Roughly how complex is each run (one step, multi-step, multi-page)?
What’s the human equivalent — how long would the task take a person?
Anything you can’t answer is a candidate to retire on May 3.
2. Identify your top 3 keepers
Sort the list by “human equivalent time saved per month.” The top three are your ROI anchors. Those are the agents you’ll actively budget credits for. Everything below the line is provisional — keep them running only if credit headroom allows.
3. Get on the right plan if you aren’t already
Custom Agents stay on Business and Enterprise. If your workspace is on Free or Plus and you’ve been using Custom Agents on a Business trial, the trial expiry is the cutoff. After that, agents disappear entirely unless you upgrade. Business is $20 per user per month billed annually, $24 monthly. Enterprise is custom-priced.
4. Have an admin set up the credit dashboard before May 4
The credit dashboard is where admins buy and track credits. The smart move is to provision a starter pack — somewhere in the hundreds-to-low-thousands range of credits — before the cutover, so your top-three agents don’t pause on the first morning of the new pricing era. You can scale credit purchases up or down monthly based on what actually gets consumed.
5. Set up usage observation
Once credits are running, treat the first 30 days as data collection. Watch which agents burn credits fastest. Watch which agents you actually open the output of. The gap between “credits consumed” and “output used” is where the next round of agent retirement happens.
The trap to avoid
The trap to avoid.
The natural temptation between now and May 3 is to build more agents while it’s still free. Don’t. The agents you build in a free-trial mindset are precisely the ones you’ll regret budgeting credits for in May.
A better use of the remaining trial window: harden the agents you already have. Tighten their scopes. Reduce the number of pages they touch. Cut the multi-step chains that don’t need to be multi-step. Every operation you can shave off a workflow today is a credit you don’t spend tomorrow.
This is the gates-before-volume principle applied to agents. You don’t scale by adding more agents. You scale by making each agent leaner before the meter starts.
What this signals about Notion’s roadmap
Reading the tea leaves: credit-based pricing for agents is the foundation for Workers for Agents (currently in developer preview as of April 2026). Workers let agents call code and external APIs. That’s the kind of capability that needs metering — you can’t ship “an agent that calls any API you want” on a flat fee. Credits make Workers possible at scale.
If you’re a developer or an agency, this is the more interesting story. The May 3 cliff is the boring part. The Workers preview is the part to watch, and credits are the pricing rail that makes Workers viable as a product.
The operator’s bottom line
May 3 is not a problem to solve. It’s a forcing function that turns “I’m experimenting with agents” into “I run a small fleet of agents on a budget.”
That’s a healthier place to be. Free trials produce sprawl. Metered usage produces discipline.
Decide your top three. Get on the right plan. Have an admin top up credits before May 4. Spend the next week tightening, not building. That’s the entire move.
Sources
Notion Help Center — Buy & track Notion credits for Custom Agents
Notion 3.3 release notes (February 24, 2026)
Notion Pricing page (April 2026 snapshot)
Continue the journey
This article is part of the May 3 Cliff Decision journey-pack on Tygart Media. Here’s where to go next:
Restoration company marketing in 2026 is multi-channel by default. The shops still trying to grow on a single channel — usually Google Ads or referral alone — are losing share to operators running coordinated programs across six channels at once. This is the working playbook.
The framing matters: marketing is the lead-generation layer that sits on top of the operating model. A restoration shop with strong operations and weak marketing has untapped capacity. A shop with strong marketing and weak operations burns the lead investment on jobs it cannot deliver well. The playbook below assumes the operating model is in place.
The Six Channels That Actually Move Restoration Lead Flow
The six channels that actually move restoration lead flow.
Restoration marketing in 2026 is built on six channels. Most shops operate two or three reasonably well and ignore the rest. Operators who run all six produce more predictable lead flow at lower blended cost.
Search engine optimization. The compounding channel. The largest source of high-intent organic leads for shops that invest consistently.
Paid search and local services ads. The fastest channel to turn on. The most price-sensitive in 2026 as competition has intensified.
Referral systems and partner networks. The highest-converting channel. Plumbers, insurance agents, property managers, real estate agents.
Content and AI-search visibility. The new channel — being cited in ChatGPT, Claude, Perplexity, and Google AI Overviews when prospects research restoration questions.
TPA and carrier program enrollment. The volume channel. Lower margin, predictable flow.
Direct outreach for commercial accounts. The relationship channel. Long cycle, high lifetime value.
The right mix for a given shop depends on residential-vs-commercial split, geographic market dynamics, and existing channel maturity.
Channel 1: SEO
SEO for restoration companies in 2026 has bifurcated. Local pack and Google Business Profile signals continue to drive emergency-intent residential leads. Editorial and content depth drives commercial and education-intent traffic, and increasingly drives the AI-search visibility described in Channel 4.
The high-leverage SEO investments for a restoration company in 2026:
Google Business Profile completeness — services, hours, service area, photos, posts, review velocity.
Service-area landing pages for every city or neighborhood the shop covers, with original content rather than templated copy.
Service-line landing pages that address specific work categories — water mitigation, smoke and fire, biohazard, mold, reconstruction.
Editorial content that addresses the questions buyers actually ask before they engage — what does restoration cost, what does the IICRC do, how does insurance handle water damage.
Review generation systems that produce a steady volume of authentic Google reviews.
Channel 2: Paid Search and Local Services Ads
Paid search produces the fastest lead flow but at the highest unit cost. The competitive intensity in restoration paid search has risen materially over the last 24 months, particularly in storm-affected markets and metropolitan areas with multiple national franchises.
Working principles for paid search in 2026:
Local Services Ads where available — the verified-vendor placement above traditional ads tends to produce higher-converting leads at competitive cost.
Tight match-type discipline and aggressive negative-keyword maintenance to keep cost-per-lead reasonable.
Landing pages built for the ad — not the home page. Generic landing pages are the largest source of paid-search waste in restoration.
Call tracking and lead-source attribution so the shop can measure cost per acquired job, not cost per click.
Channel 3: Referral Systems and Partner Networks
Referrals are the highest-converting source of restoration leads — and they are not free. They require a deliberate system. The partner categories that produce restoration referrals in 2026:
Insurance agents and brokers. The agent who hears about a loss before the carrier does often controls vendor recommendation.
Plumbers and HVAC contractors. The trades that arrive at water and smoke losses before restoration.
Property managers. Repeat referral source for water and reconstruction work.
Real estate agents. Pre-listing remediation work, mold and air-quality services.
Other restoration shops. Capacity-overflow referrals in busy seasons.
The system that produces referrals is recognition — branded materials, regular touchpoints, a clear ask, and measurable reciprocity where possible. Referral programs without a system tend to produce sporadic results.
Channel 4: AI Search Visibility
Channel 4 — AI search visibility.
The newest restoration marketing channel is appearance in AI-generated answers — ChatGPT, Claude, Perplexity, Google AI Overviews. Buyers researching restoration questions in 2026 increasingly receive AI-generated answers before they click through to traditional search results. Being cited in those answers requires editorial content with authority signals — comprehensive coverage of the topic, structured FAQ formatting, schema markup, and the kind of factual depth language models surface.
This channel does not replace traditional SEO. It rewards the same content investments and amplifies them. Shops investing in editorial restoration content in 2026 are seeing both organic search and AI-search returns from the same work.
Channel 5: TPA and Carrier Programs
TPA program enrollment is the most predictable lead flow available to a restoration shop, with the trade-off of compressed margin and dependency risk. The decision is whether TPA work serves as a base load that supports crew utilization while higher-margin direct-to-owner work is cultivated. For most shops, the answer is yes — but not as the entire pipeline.
Channel 6: Direct Outreach for Commercial
Channel 6 — direct outreach for commercial.
The commercial sales motion is its own channel — outbound, named-account, multi-persona, long-cycle. The detailed playbook is covered separately in The Commercial Restoration Sales Stack, but the marketing function feeding it includes target-account research tools, persona-specific content, and the conference and event presence that produces the introduction opportunities the sales motion converts.
Budget Framework
A working budget framework for restoration company marketing in 2026:
Total marketing investment: 4% to 8% of revenue, depending on growth ambition and competitive intensity.
Allocation: roughly 30% to 40% paid search, 25% to 35% SEO and content, 15% to 25% referral systems and partner cultivation, 10% to 15% direct outreach and commercial sales, 5% to 10% experimental or emerging channels.
The largest single budget mistake in 2026 is over-allocating to paid search at the expense of SEO and content, because it produces fast results that mask the absence of compounding channels.
Measurement
Each channel needs its own measurement, and the shop needs a blended view that ties marketing investment to acquired jobs. The metrics that matter:
Cost per acquired job by channel — not cost per lead, which obscures conversion quality.
Lifetime value by channel — referral and commercial leads typically produce higher lifetime value than paid-search leads.
Channel concentration risk — a shop with more than 50% of revenue from any single channel has a fragility problem regardless of the channel.
The Single Largest Marketing Mistake
The most common marketing mistake in the restoration industry in 2026 is treating channels as substitutes rather than complements. Paid search and SEO are not alternatives. Referral and direct outreach are not alternatives. The shops that produce predictable lead flow at sustainable cost run all six channels in coordination, with each channel covering the others’ weaknesses. The shops that lurch between channels — six months of paid, six months of “we need to do SEO instead” — produce inconsistent results regardless of which channel they are currently emphasizing.
What is the best marketing channel for restoration companies in 2026?
There is no single best channel. The shops with predictable lead flow run six channels in coordination — SEO, paid search, referral systems, AI-search-optimized content, TPA programs, and direct commercial outreach. Single-channel programs no longer produce reliable results.
How much should a restoration company spend on marketing?
A working budget range is 4% to 8% of revenue, with allocation across paid search, SEO and content, referral systems, direct outreach, and experimental channels. The exact mix depends on residential-vs-commercial split, market dynamics, and existing channel maturity.
Is paid search still worth it for restoration companies?
Yes, but with discipline. Competitive intensity has raised cost-per-click materially in 2026. Local Services Ads, tight match-type management, and dedicated landing pages keep cost per acquired job reasonable. Generic landing pages and broad-match targeting are the largest source of paid-search waste.
What is AI-search optimization for restoration companies?
AI-search optimization is the practice of producing content that gets cited by ChatGPT, Claude, Perplexity, and Google AI Overviews when prospects research restoration questions. It rewards editorial depth, structured FAQ formatting, schema markup, and comprehensive coverage of restoration topics. It complements rather than replaces traditional SEO.
How important are Google reviews for restoration companies?
Critical. Review velocity and rating directly affect Google Business Profile visibility, Local Services Ads cost, and consumer choice. A deliberate review-generation system is one of the highest-leverage marketing investments a restoration shop can make.
For more on the marketing layer that sits on top of restoration operations, see SEO for Restoration on Tygart Media.
“How do I increase restoration sales?” is usually answered with a list of marketing tactics. The honest answer is structural: three levers move restoration company revenue, and most growth that lasts comes from operating those three deliberately rather than chasing more leads.
The three levers are pricing discipline, mix shift toward higher-margin work, and capacity utilization. They compound. A restoration company that improves any one of them by 10% sees a meaningful revenue and margin lift. A company that improves all three simultaneously transforms its business in 18 months.
Lever 1: Pricing Discipline
Lever 1 — pricing discipline.
Pricing discipline is the most undervalued growth lever in the restoration industry. The reason is structural — most restoration revenue is priced by Xactimate or Symbility line items, which creates the illusion that pricing is fixed by the carrier. It is not.
The pricing levers that operators actually control:
Scope discipline. The most consequential pricing decision in any restoration job is whether the documented scope reflects the work performed. Under-scoping is the largest source of margin erosion in the industry.
Time and material work selection. Some categories of work — biohazard, contents, specialty services — can be billed on a time-and-material basis at materially higher margin than carrier-line-item rates. The mix question is whether your shop pursues this work or defaults to insurance-priced jobs.
Self-pay and direct-bill work. Cash work outside the insurance channel can be priced to market rather than to carrier line items. The discipline of building a direct-pay funnel produces a higher-margin revenue stream that compounds.
Estimating consistency. Two estimators on the same shop floor will produce different scopes for the same loss. The variance is pure margin leakage. Standardized estimating practice — checklist-driven, peer-reviewed — closes the variance.
Pricing discipline produces revenue without producing more jobs. It is the highest-margin growth lever a restoration shop has access to, and it is rarely the first one operators reach for.
Lever 2: Mix Shift
Mix shift is the deliberate movement of revenue from lower-margin work types to higher-margin work types. Not every job in a restoration shop produces the same gross margin. The honest accounting:
Carrier-driven residential water mitigation: stable volume, compressed margin, high competitive intensity.
TPA program work: predictable, lower margin, vendor-relationship dependent.
Direct-to-owner commercial work: longer cycle, higher margin, less price-sensitive.
Reconstruction: high revenue per job, complex margin dynamics, capacity-intensive.
The mix-shift question is which categories of work the shop is deliberately growing. Most restoration companies inherit their mix passively — they take what comes through the door. Companies that grow revenue without growing headcount tend to be operating mix shift deliberately, often by adding a single specialty service category that pulls margin upward.
The structural insight is that adding a higher-margin work category typically requires the same overhead as adding more of the existing mix, which means the incremental gross margin drops disproportionately to the bottom line.
Lever 3: Capacity Utilization
Lever 3 — capacity utilization.
Capacity utilization is the lever that determines whether existing assets produce more revenue. A restoration shop with 12 technicians, 6 trucks, and a fixed overhead is producing a specific level of revenue. The question is whether that level is constrained by lack of demand, lack of operational efficiency, or both.
The capacity levers that move revenue:
Dispatch efficiency. The minutes between FNOL and on-site arrival, and the routing efficiency across multiple jobs in a day, compound into measurable capacity gains.
Technician productivity. Documentation discipline, equipment readiness, and clean handoffs between production and reconstruction directly affect billable hours per technician per day.
Equipment turn rate. Restoration equipment that sits in the warehouse is not producing revenue. Equipment tracking and dispatch discipline produces meaningful utilization gains.
After-hours and weekend response. A 24/7 restoration operation that under-utilizes evening and weekend capacity is leaving the highest-urgency, lowest-competition work on the table.
Capacity utilization compounds with the other two levers. A shop with disciplined pricing and a deliberate mix shift, but poor capacity utilization, leaves substantial revenue uncaptured. A shop with strong utilization but weak pricing discipline is running hard for compressed margin.
The Multiplier Effect
The three levers multiply rather than add. A 10% improvement in pricing discipline, a 10% mix shift toward higher-margin work, and a 10% improvement in capacity utilization does not produce 30% revenue growth. It produces meaningfully more — typically in the range of 35% to 45% — because the higher-margin work earns higher prices on more efficient operations.
This is why operators who run all three levers deliberately can grow revenue and margin without growing the lead pipeline. The restoration industry’s default operating mode — chase more leads, take whatever comes through the door — leaves all three levers passive.
What to Measure
Each lever has a measurement that translates the abstract concept into operating discipline:
Pricing discipline: gross margin trend by job category, scope variance between estimators, percentage of revenue from time-and-material and direct-pay work.
Mix shift: revenue distribution across work categories, gross margin by category, year-over-year shift toward target categories.
Capacity utilization: billable hours per technician per day, equipment turn rate, percentage of jobs with arrival time within service-level commitment.
An operator who reviews these numbers monthly and can describe what is moving and why has a lever-driven business. An operator who reviews only top-line revenue is running on autopilot.
The Marketing Lever Is the Fourth, Not the First
The marketing lever is the fourth, not the first.
Marketing — SEO, paid advertising, referral systems, content — is a real lever, but it is the fourth one, not the first. A restoration company with disciplined pricing, deliberate mix shift, and strong capacity utilization will absorb marketing-driven leads at high efficiency. A company without those three will absorb marketing-driven leads at the same low efficiency they absorb existing leads, and the marketing investment will produce disappointing returns.
This is the structural reason that restoration owners who jump straight to “we need more leads” rarely produce sustained revenue growth. The leads land on a leaky operating model.
What is the highest-leverage way to increase restoration company revenue?
Pricing discipline — specifically scope discipline, deliberate inclusion of time-and-material and direct-pay work, and standardized estimating practice — is the highest-margin growth lever a restoration shop has. It produces revenue without producing more jobs.
How do I improve gross margin in a restoration business?
The three structural levers are pricing discipline, mix shift toward higher-margin work categories like biohazard or commercial direct-to-owner, and capacity utilization. Operating all three deliberately produces measurable margin lift in 12 to 18 months.
Should I add specialty services to my restoration business?
Specialty services — biohazard, trauma cleanup, contents, large-loss commercial — typically produce higher gross margin than carrier-driven residential water mitigation, and they pull mix toward the high-margin end. The decision depends on whether your shop has the operational capacity and certifications to deliver them well.
How do I know if my restoration company has a capacity utilization problem?
The diagnostic measures are billable hours per technician per day, equipment turn rate, and percentage of jobs with arrival time inside service-level commitment. A shop where these numbers are not measured monthly almost certainly has untapped capacity.
Is more marketing the answer to slow restoration sales?
Not by itself. Marketing-driven leads land on whatever operating model exists. A restoration company with weak pricing discipline, passive mix, and poor capacity utilization will absorb marketing leads at low efficiency and produce disappointing returns on marketing spend. Operating discipline first, marketing second.
For operator-focused playbooks on running and scaling a restoration company, see the Restoration Operator’s Playbook archive.
The honest answer to “where do restoration sales reps learn to sell?” is: from a patchwork of technical training, industry conferences, and outside sales programs that were not built for the restoration industry. There is no single program that produces a fully trained commercial restoration sales rep, and operators who pretend otherwise end up with reps who can talk about IICRC certifications but cannot run a buying-committee conversation.
This is a working map of the restoration sales training landscape as it exists in 2026, what each option teaches well, and where the gaps are. It is written for restoration owners and sales managers deciding where to spend training dollars.
Three Categories of Restoration Sales Training
Three categories of restoration sales training.
The training landscape splits into three categories that solve different problems:
IICRC and industry technical courses. Strong on the science, the standards, and the technical credibility that lets a sales rep hold a conversation with a facilities engineer or a risk manager.
Restoration industry conferences and sales tracks. Strong on community, peer learning, and tactical playbooks. Variable in depth.
Outside sales programs and sales coaching. Strong on the sales discipline itself — qualification, account management, negotiation, close mechanics — but generally not restoration-specific.
The reps who actually carry commercial restoration pipeline have typically drawn from all three. The reps who hold only one category tend to be one-dimensional in the field.
IICRC and Industry Technical Courses
IICRC courses — WRT, ASD, AMRT, FSRT, and the more advanced certifications — are the technical baseline. They are not sales courses, but they produce the technical fluency that lets a sales rep be taken seriously by buyers who care about standards. A rep who cannot speak to S500 category and class definitions, or who struggles to explain what an ASD-certified technician actually does on a job site, has a credibility ceiling in commercial restoration sales.
What technical courses do not teach: how to qualify a buying committee, how to map an account, how to run a quarterly cultivation cadence, or how to close a preferred-vendor agreement. The gap is structural — they were never intended as sales courses.
Industry Conferences and Sales Tracks
Restoration industry conferences — Experience Conference & Exchange, Restoration Industry Association events, and the various carrier and TPA-adjacent gatherings — are where tactical playbooks circulate. Sales tracks at these events typically run breakouts on commercial selling, marketing strategy, and account development.
The strength of conference-based learning is the peer-to-peer transfer. A sales rep who hears how a comparable operator runs their named-account program in a different market will absorb more in 45 minutes than from any structured curriculum. The weakness is depth — a 45-minute breakout cannot replace the cumulative skill of running a real commercial sales cycle.
Outside Sales Programs
Outside sales training programs — Sandler, Challenger, MEDDIC, and the various enterprise B2B sales methodologies — were not built for restoration but apply directly to the commercial restoration sales motion. Restoration-specific sales coaches and programs have emerged in the last five years that translate these methodologies into restoration language.
The strongest case for outside sales investment is for shops that have made the deliberate decision to pursue commercial accounts at scale. The structured discipline of a methodology like MEDDIC — identifying metrics, economic buyer, decision criteria, decision process, identify pain, and champion — maps cleanly onto the five-persona buying committee that controls commercial restoration vendor selection.
The risk is treating outside sales training as a silver bullet. A rep trained in MEDDIC who lacks the technical fluency to discuss S500 category determinations will lose credibility with the same buying committee the methodology is supposed to help them navigate.
The Internal Training That Actually Moves the Needle
The internal training that actually moves the needle.
The most undervalued sales training in the restoration industry is the internal kind — ride-alongs with the owner or senior sales leader, formal account reviews with critique, and structured debriefs after both wins and losses. Most restoration shops do not run this discipline because it requires senior time that is hard to carve out.
Operators who do run internal training cite a consistent pattern: a new sales rep who shadows the owner on twelve commercial cultivation meetings in the first 90 days will out-perform a rep who takes a six-week external program with no internal coaching. The mechanism is straightforward — the owner’s market-specific knowledge, account history, and judgment do not transfer through a course.
What to Look For in a Restoration Sales Training Investment
What to look for in a restoration sales training investment.
If you are an owner or sales manager evaluating where to spend training dollars in 2026, the framework that holds up:
Verify technical baseline through IICRC certifications appropriate to the work the rep will sell.
Build a structured methodology — Sandler, Challenger, or MEDDIC — into the rep’s first 90 days, with a clear application to commercial restoration buying committees.
Schedule conference attendance with deliberate breakout selection, not as a perk.
Run formal weekly sales reviews internally — pipeline, named-account progress, win/loss analysis — with the owner or sales leader present.
Treat the first six commercial cultivation meetings as paired ride-alongs, not solo selling attempts.
The total investment is meaningful but not extreme. The alternative — a rep who learns commercial restoration sales by burning through a year of pipeline — is far more expensive.
The Marketing Class Question
Restoration sales reps frequently search for “restoration sales marketing class” as if there is a single course that solves the gap. There is not. The functional substitute is the combination above, paired with a marketing program at the company level — content marketing, paid advertising, referral systems — that produces the qualified prospects the trained rep then converts. Sales training without a parallel marketing investment produces well-trained reps with empty pipelines.
Is there a single best restoration sales training program?
No. The reps who carry serious commercial restoration pipeline have typically combined IICRC technical courses, an outside sales methodology like Sandler or MEDDIC, structured internal coaching, and selective conference attendance. There is no single program that replaces this combination.
Do IICRC certifications teach sales skills?
IICRC certifications teach the technical and standards baseline that lets a sales rep be taken seriously by commercial buying committees. They do not teach sales skills — qualification, account mapping, cultivation cadence, or close mechanics — and were never intended to.
Should restoration sales reps take outside sales courses?
Yes, particularly for shops pursuing commercial accounts at scale. Methodologies like Challenger, Sandler, and MEDDIC translate directly to the multi-persona buying committee that controls commercial restoration vendor selection. The investment pays back in shorter cultivation cycles and higher win rates.
How long does it take to train a commercial restoration sales rep?
Most operators report that a new commercial sales rep needs nine to fifteen months to fully ramp — the time to complete one full cultivation cycle from cold prospect to first signed account. Compressing the ramp timeline below nine months is rarely realistic.
What is the highest-leverage internal sales training?
Paired ride-alongs with the owner or sales leader on the first six to twelve commercial cultivation meetings, paired with structured weekly pipeline reviews. This transfers market-specific knowledge and judgment that no external course can deliver.
For more on building the operational and sales infrastructure of a restoration company, see the Restoration Operator’s Playbook.
Direct Answer (August 2026): Claude features a 200,000-token context window (~150,000 words) across all models and tiers, capable of ingesting entire codebases, financial filings, or 500-page manuals in a single prompt with 99.5%+ needle-in-a-haystack recall.
Looking for quick answers? The FAQ version covers every common question directly.
Claude’s context window is one of those specs that sounds simple until you actually need to use it. “1 million tokens” means almost nothing without a frame of reference. This is the guide we wish existed when we started building on Claude — written from our own experience running it in production, with numbers pulled directly from Anthropic’s official documentation.
Quick Definition
The context window is Claude’s working memory for a conversation. It holds everything Claude can see and reason about at once: your messages, Claude’s responses, any documents you’ve shared, and system prompts. When the window fills up, earlier content drops out.
Current Context Window Sizes by Model (June 2026)
Current context window sizes by model.
These numbers come directly from Anthropic’s official models page, fetched May 9, 2026. Model strings are exact API identifiers:
Model
API String
Context Window
Max Output
Claude Fable 5
claude-fable-5
1,000,000 tokens
128,000 tokens
Claude Opus 4.8
claude-opus-4-8
1,000,000 tokens
128,000 tokens
Claude Sonnet 5
claude-sonnet-4-6
1,000,000 tokens
64,000 tokens
Claude Haiku 4.5
claude-haiku-4-5-20251001
200,000 tokens
64,000 tokens
Fable 5, Opus 4.8, and Sonnet 5 all have the full 1M token context window. Haiku 4.5 is 200K. The key difference between Opus 4.8 and Sonnet 5 in this table is the max output — Opus 4.8 can write up to 128K tokens in a single response, Sonnet 5 caps at 64K.
What Does 1 Million Tokens Actually Hold?
Token counts are an abstraction. Here’s what 1 million tokens translates to in practical terms:
About 750,000 words of English text — roughly 10 full-length novels, or 1,500 average blog posts
A full mid-size codebase — a 50,000-line Python project with comments fits comfortably
Hours of meeting transcripts — a full workday of recorded calls, transcribed, fits in one context window
Multiple large documents simultaneously — 10 research PDFs at 30 pages each, all in the same conversation
Long conversation histories — hundreds of back-and-forth exchanges before anything starts dropping off
We’ve loaded entire Notion exports, full project histories, and multi-document research packs into a single Claude session. At 1M tokens, you’re unlikely to hit the ceiling in a normal working session. You hit it when you’re doing things like: loading your entire codebase plus documentation plus conversation history and then asking Claude to do a full architectural review.
Context Window vs. Memory: What’s the Difference?
Context window vs memory — what’s the difference?
This is where a lot of people get confused. The context window and memory are not the same thing:
Context window: What Claude can see right now, in this session. Once a session ends, it’s gone.
Memory (in claude.ai): A separate system that extracts and stores key information from past sessions. It surfaces relevant facts into future conversations as a snippet in the context.
Managed Agents memory stores: A developer-layer construct where agents maintain and update knowledge bases across sessions — distinct from both the context window and the consumer memory feature.
The 1M token context window is your working memory for one session. It doesn’t persist. Memory systems are what carry information across sessions — but they work by injecting a summary into the context window of the new session, not by giving Claude access to the full history.
Does a Bigger Context Window Mean Better Performance?
Mostly yes, with one important nuance. More context means Claude has more information to reason about, which generally produces better outputs for tasks that benefit from full context — code reviews, document synthesis, long-form writing, multi-document comparison.
The nuance: performance can degrade on tasks involving specific information buried deep in a very long context. This is sometimes called the “lost in the middle” problem — models tend to pay more attention to the beginning and end of a long context than the middle. Anthropic has worked on this with Claude’s architecture, and it performs well on long-context tasks, but it’s worth structuring important information at natural reference points rather than burying it in the middle of a 500-page document.
How We Actually Use the 1M Token Window
How we actually use the 1M token window.
We run Claude in production for content operations, site management, and agentic coding workflows. Here’s where the 1M context window makes a concrete difference in our work:
Full site audits: Loading every post from a WordPress site (200+ posts worth of content) into one session for comprehensive SEO analysis — without having to chunk and re-prompt
Cross-session context: Pasting in long Notion briefings, prior session transcripts, and the current task in one go. The window is large enough that we don’t have to decide what to leave out.
Codebase-wide reasoning: In Claude Code, having the full project context means Claude can make changes that account for how files interact rather than reasoning only about the current file
Multi-document synthesis: Research projects where we load 10-15 source documents and ask Claude to synthesize across them — something that was impossible at 100K context windows
The practical shift from 200K to 1M tokens wasn’t just “more room.” It changed what we could ask Claude to do in a single session.
Context Window on the API: Batch Output Extension
For API users: on the Message Batches API, Fable 5, Opus 4.8, and Sonnet 5 support up to 300K output tokens using the output-300k-2026-03-24 beta header. This is relevant for batch generation tasks where you need very long outputs — documentation generation, large codebases, book-length content.
Frequently Asked Questions
What is Claude’s context window in 2026?
Claude Fable 5, Claude Opus 4.8, and Claude Sonnet 5 all have 1,000,000 token (1M token) context windows as of June 2026. Claude Haiku 4.5 has a 200,000 token context window. These are the current generally available models.
How many pages can Claude read at once?
At 1M tokens, Claude can hold roughly 750,000 words of English text — equivalent to approximately 3,000 average pages. In practice, a typical 20-page PDF is roughly 10,000-15,000 tokens, so you could load 60-100 such documents in a single session before approaching the limit.
Does the context window reset between messages?
No — the context window accumulates across an entire conversation session. Every message you send and every response Claude gives adds to the total. The window doesn’t reset between individual messages; it resets when you start a new conversation.
What happens when Claude hits the context window limit?
When a conversation reaches the context window limit, earlier messages begin to drop out of the active context. Claude can no longer reference information from those earlier messages — it effectively forgets that part of the conversation. In the claude.ai interface, you’ll see a notification when you’re approaching the limit.
Is the 1M context window available on the free plan?
The model available to free plan users has access to the 1M context window. However, free plan usage limits mean long-context sessions hit rate limits faster than paid plans. The window is technically available, but sustained heavy use of it is more practical on paid tiers.
What’s the difference between Claude Opus 4.8 and Sonnet 5 context windows?
Both have the same 1M token input context window. The difference is max output: Opus 4.8 can generate up to 128,000 tokens in a single response; Sonnet 5 caps at 64,000 tokens. For most tasks this distinction doesn’t matter, but for very long document generation or large code outputs, Opus 4.8 has the higher output ceiling.
💼 Deploying Claude or AI Infrastructure in Your Business?
Direct Answer (August 2026): Choose Sonnet 4.6 for everyday programming, writing, and operational agents (5x cheaper than Opus with near-parity reasoning). Choose Opus 4.8 for mission-critical architectural design, complex debugging, and legal synthesis. Choose Haiku 4.5 for high-volume classification, real-time chatbots, and automated extraction.
Claude Fable 5 launched June 9, 2026 as a new tier above Opus 4.8 — priced at $10/$50/MTok (2× Opus). This guide now covers all four models. Full Fable 5 breakdown →
Anthropic’s Claude model lineup in 2026 now spans four tiers: Fable 5 at the top for maximum capability ($10/$50/MTok), Opus 4.8 for serious production work ($5/$25), Sonnet 4.6 for the best balance of performance and cost ($3/$15), and Haiku 4.5 for speed and high-volume work ($1/$5). Picking the wrong model costs money or performance — sometimes both. This guide covers every meaningful difference so you can make the right call.
Quick answer: Sonnet 4.6 handles 80–90% of tasks at a fraction of the cost of higher tiers. Use Fable 5 for the hardest engineering and long-horizon agentic work ($10/$50/MTok). Use Opus 4.8 for serious production work with zero data retention requirements ($5/$25). Use Sonnet 4.6 as your daily driver ($3/$15). Use Haiku 4.5 when speed and cost dominate ($1/$5).
The Current Claude Model Lineup (June 2026)
Pick by the job shape. Version numbers move; the seats stay.
Claude Fable 5 vs Opus 4.8 vs Sonnet 4.6 vs Haiku 4.5: side-by-side
Feature
Claude Fable 5 🆕
Claude Opus 4.8
Claude Sonnet 4.6
Claude Haiku 4.5
Best for
Hardest engineering, long-horizon autonomy
Production work, zero-data-retention
Best speed/intelligence balance
Fastest responses, high-volume tasks
Input price
$10 / MTok
$5 / MTok
$3 / MTok
$1 / MTok
Output price
$50 / MTok
$25 / MTok
$15 / MTok
$5 / MTok
Context window
1M tokens
1M tokens
1M tokens
200k tokens
Max output
128k tokens
128k tokens
64k tokens
64k tokens
Extended thinking
No (adaptive always on)
No
Yes
Yes
Adaptive thinking
Always on
Yes
Yes
No
Zero data retention
No (30-day mandatory)
Yes
Yes
Yes
Latency
Slow–Moderate
Moderate
Fast
Fastest
API ID
claude-fable-5
claude-opus-4-8
claude-sonnet-4-6
claude-haiku-4-5
As of June 2026, Anthropic’s four current models are Claude Fable 5, Claude Opus 4.8, Claude Sonnet 4.6, and Claude Haiku 4.5. All four support text and image input, multilingual output, and vision processing. They differ significantly in pricing, context window, output limits, and capability.
Feature
Fable 5 🆕
Opus 4.8
Sonnet 4.6
Haiku 4.5
Input price
$10 / MTok
$5 / MTok
$3 / MTok
$1 / MTok
Output price
$50 / MTok
$25 / MTok
$15 / MTok
$5 / MTok
Context window
1M tokens
1M tokens
1M tokens
200K tokens
Max output
128K tokens
128K tokens
64K tokens
64K tokens
Extended thinking
No (adaptive always on)
No
Yes
Yes
Adaptive thinking
Always on
Yes
Yes
No
Latency
Slow–Moderate
Moderate
Fast
Fastest
Reliable knowledge cutoff
2026
Jan 2026
Aug 2025 (reliable)
Feb 2025 (reliable)
Pricing is per million tokens (MTok) via the Claude API. Source: Anthropic Models Overview, June 2026.
Claude Fable 5: The New Top Tier (June 9, 2026)
Fable 5 is Anthropic’s first Mythos-class model released for general availability. It landed June 9, 2026 and sits above Opus 4.8 in capability — scoring 95.0% on SWE-bench Verified (vs 88.6% for Opus 4.8) and 80.0% on SWE-bench Pro (vs 69.2%). On the Senior Engineer benchmark, Fable 5 scores 91/100 vs approximately 63/100 for Opus 4.8.
Key differentiators for Fable 5:
Adaptive thinking always on — Fable 5 doesn’t have an extended thinking toggle. It always reasons adaptively, scaling depth to task complexity.
128K max output — same as Opus 4.8, twice Sonnet’s 64K cap.
1M token context window — same as Opus 4.8 and Sonnet 4.6.
Two constraints that matter:
Mandatory 30-day data retention. Fable 5 is not available under zero data retention. If your use case requires ZDR (healthcare, legal, finance with strict data handling), use Opus 4.8.
Safety classifier routing. Prompts touching cybersecurity, biology, chemistry, and distillation route to an Opus 4.8 fallback — at Fable 5 pricing. If your workload is in these domains, the upgrade is less impactful.
Use Fable 5 for: large migrations or refactors, multi-agent orchestration at frontier quality, long-horizon agentic work, complex scientific analysis, and any task where quality on hard problems justifies 2x cost over Opus.
Skip Fable 5 for: well-scoped routine work, high-volume pipelines (2x cost compounds), ZDR-required use cases, or domains where the safety classifier fallback applies.
Claude Opus 4.8: The Production Standard
Opus 4.8 is Anthropic’s most capable model supporting zero data retention (ZDR) — the right default for most production API work. Fable 5 has since surpassed it in raw capability, but Opus 4.8 remains the better choice for ZDR workloads, cost-sensitive pipelines, and domains where Fable 5’s safety classifier routing applies. Anthropic describes it as a step-change improvement in agentic coding over Opus 4.8, with a new tokenizer that contributes to improved performance on a range of tasks. Note that this new tokenizer may use up to 35% more tokens for the same text compared to previous models — a cost consideration worth factoring in for high-volume workflows.
Key differentiators for Opus 4.8 over the other two models:
128K max output tokens — double Sonnet and Haiku’s 64K cap. This matters for generating long-form code, detailed reports, or complete document drafts in a single call.
1M token context window — same as Sonnet 4.6, meaning Opus can process entire codebases or book-length documents in a single session.
Adaptive thinking — Opus 4.8 and Sonnet 4.6 both support adaptive thinking, which lets the model adjust reasoning depth based on task complexity.
Most recent knowledge cutoff — January 2026, versus August 2025 (reliable) for Sonnet and February 2025 (reliable) for Haiku.
Opus does not support extended thinking — that capability lives on Sonnet 4.6 and Haiku 4.5 Extended thinking lets the model reason step-by-step before generating output, which is particularly useful for complex math, science, and multi-step logic problems.
Use Opus 4.8 for: complex architecture decisions, large codebase analysis, multi-agent orchestration tasks, outputs that require more than 64K tokens, tasks demanding the latest possible knowledge, and any work where you need Opus-tier reasoning with zero data retention (Fable 5 is the absolute frontier, but does not support ZDR).
Skip Opus 4.8 for: routine content generation, customer support pipelines, high-volume classification or extraction, real-time applications requiring low latency, or any task where Sonnet scores within your acceptable quality threshold.
Claude Sonnet 4.6: The Workhorse
Sonnet 4.6 is the model Anthropic recommends as the best combination of speed and intelligence. Released in February 2026, it delivers a 1M token context window at $3 input / $15 output per million tokens — the same context window as Opus at 40% lower cost.
Sonnet 4.6 also uniquely offers extended thinking, which Opus 4.8 does not. When extended thinking is enabled, Sonnet can perform additional internal reasoning before generating its response — useful for reasoning-heavy tasks like complex debugging, multi-step research, and technical problem-solving where chain-of-thought depth matters.
For developers and teams using Claude Code, Sonnet 4.6 is the standard daily driver. It handles tool calling, agentic workflows, and multi-file code reasoning reliably, at a price point that makes heavy daily use economically viable.
Use Sonnet 4.6 for: most production workloads, Claude Code sessions, long-document analysis, content generation, coding tasks, research synthesis, customer-facing applications, and any workflow requiring the 1M context window where Opus’s premium isn’t justified.
Skip Sonnet 4.6 for: high-volume pipelines where Haiku’s lower cost is acceptable, simple classification or extraction tasks, or real-time applications where Haiku’s faster latency is required.
Claude Haiku 4.5: Speed and Volume
Haiku 4.5 is the fastest model in the Claude family and the most cost-efficient at $1 input / $5 output per million tokens. It has a 200K token context window — smaller than Opus and Sonnet’s 1M, but still substantial for most single-task work. It supports extended thinking but not adaptive thinking.
The 200K context limit is the most important practical constraint. Most single-document, single-task workflows fit within 200K. Multi-file codebases, long books, or extended conversation histories that push past that threshold need Sonnet or Opus.
Haiku 4.5 has the oldest knowledge cutoff of the three: February 2025. For tasks requiring awareness of events or developments from mid-2025 onward, Haiku won’t have that context baked in.
Use Haiku 4.5 for: content moderation, classification pipelines, entity extraction, customer support triage, real-time chat interfaces, simple Q&A, high-volume API workflows where cost and speed dominate, and any task where quality requirements are modest.
Skip Haiku 4.5 for: complex reasoning, large codebase analysis, tasks requiring recent knowledge (post-February 2025), multi-step agent workflows, or any output requiring more than 200K tokens of input context.
Pricing: What the Numbers Actually Mean in Practice
Price is a consequence of the seat. Start from the task.
All three models price output tokens at 5x the input rate — a ratio that holds across the entire Claude lineup. This means verbose, long-form outputs cost significantly more than short, targeted responses. Minimizing generated output length is the highest-leverage cost optimization available before you touch model routing or caching.
To put the pricing in concrete terms: generating one million output tokens (roughly 750,000 words of generated text) costs $25 on Opus, $15 on Sonnet, and $5 on Haiku. For input-heavy workloads like document analysis where you’re feeding in large amounts of text but getting shorter responses, the cost gap narrows.
Three additional pricing levers apply across all models:
Prompt caching: Cuts cache-read input costs by up to 90% for repeated system prompts or documents. If your application reuses a large system prompt across many requests, caching is the single highest-impact cost reduction available.
Batch API: Provides a 50% discount for non-time-sensitive workloads processed asynchronously. Combine with prompt caching for up to 95% savings on qualifying workflows.
Model routing: Running a mix of Haiku for simple tasks, Sonnet for production workloads, and Opus for complex reasoning — rather than using one model for everything — can reduce total API costs by 60–70% without meaningful quality loss on the tasks that don’t require a flagship model.
Context Windows: 1M Tokens vs. 200K
Context is how much you can hold. Output is how much you can say back.
Opus 4.8 and Sonnet 4.6 both offer a 1M token context window at standard pricing — no premium surcharge for extended context. For reference, 1 million tokens is roughly 750,000 words, enough to hold a large codebase, a full academic textbook, or months of business communications in a single conversation.
Haiku 4.5 has a 200K token context window. That’s still roughly 150,000 words — sufficient for most single-document tasks, but it creates a hard ceiling for anything requiring multi-file code review, book-length document analysis, or lengthy conversation histories.
If your workflow consistently requires more than 200K tokens of input, Sonnet 4.6 is the cost-efficient choice. Opus 4.8 is the right call only when the input load requires the additional reasoning capability Opus provides, not just the context window size — because Sonnet gets you the same 1M window at 40% lower cost.
Extended Thinking vs. Adaptive Thinking
These are two distinct features that appear together in the comparison table but serve different purposes.
Extended thinking (available on Sonnet 4.6 and Haiku 4.5, not Opus 4.8) lets Claude perform additional internal reasoning before generating its response. When enabled, the model produces a “thinking” content block that exposes its reasoning process — step-by-step problem decomposition before the final answer. Extended thinking tokens are billed as standard output tokens at the model’s output rate. A minimum thinking budget of 1,024 tokens is required when enabling this feature.
Adaptive thinking (available on Opus 4.8 and Sonnet 4.6, not Haiku 4.5) adjusts reasoning depth dynamically based on task complexity — the model allocates more reasoning for harder problems and less for simpler ones, without requiring explicit configuration.
The practical implication: if you need transparent, controllable step-by-step reasoning that you can inspect and use in your application, Sonnet 4.6’s extended thinking is often the right tool — and at lower cost than Opus.
Which Claude Model Should You Choose?
The right framework for model selection in mid-2026 is a four-tier stack: Fable 5 for the hardest problems, Opus 4.8 as the production standard, Sonnet 4.6 as the daily driver, Haiku 4.5 for volume. Start with Sonnet 4.6 and escalate selectively. Most production workloads — coding, writing, analysis, customer-facing applications — are well-served by Sonnet. Opus 4.8 earns its premium when you need ZDR, outputs over 64K tokens, or the January 2026 knowledge cutoff. Fable 5 earns its 2x premium when the task is genuinely hard enough that 10+ percentage points on SWE-bench matters for your outcome.
Haiku 4.5 belongs in any pipeline where you’ve identified tasks that don’t require Sonnet’s capability. High-volume routing, triage, classification, and real-time response scenarios are Haiku’s natural territory. The optimal production routing split is roughly 70% Haiku 4.5, 20% Sonnet 4.6, 8% Opus 4.8, 2% Fable 5 — rather than using a single model for everything. That ratio cuts costs by 60–70% without meaningful quality loss on the tasks that don’t need a flagship model.
You picked your model tier. Now get the pre-built setup.
Claude Seed Kits are pre-configured skill files with 20 tested prompts and a setup guide for your specific use case. Pick the kit that matches how you work — $47 each.
What is the difference between Claude Opus 4.8, Sonnet, and Haiku?
Opus is Anthropic’s most capable model, optimized for complex reasoning, large outputs, and agentic tasks. Sonnet offers a balance of capability and cost, handling most production workloads at lower price. Haiku is the fastest and cheapest option, suited for high-volume, lower-complexity tasks. All three share the same core Claude architecture and safety training.
Is Claude Opus 4.8 worth the extra cost over Sonnet?
For most tasks, no. Sonnet 4.6 handles the majority of coding, writing, and analysis work at 40% lower cost. Opus 4.8 is worth the premium when you need outputs longer than 64K tokens, maximum agentic coding capability, or the most recent knowledge cutoff (January 2026 vs. Sonnet’s August 2025).
Which Claude model is best for coding?
Sonnet 4.6 is the standard recommendation for most coding work, including Claude Code sessions. Opus 4.8 is preferred for large codebase analysis, complex architecture decisions, or multi-agent coding workflows where maximum reasoning depth is required. Haiku 4.5 can handle simple code edits and explanations at much lower cost.
What is the Claude context window?
Claude Opus 4.8 and Sonnet 4.6 both have a 1 million token context window — roughly 750,000 words of combined input and conversation history. Claude Haiku 4.5 has a 200,000 token context window. Context window size determines how much information Claude can hold and reference in a single conversation.
Does Claude Opus 4.8 support extended thinking?
No. Extended thinking is available on Claude Sonnet 4.6 and Claude Haiku 4.5, but not on Claude Opus 4.8 Opus 4.8 supports adaptive thinking instead, which dynamically adjusts reasoning depth based on task complexity.
What is the cheapest Claude model?
Claude Haiku 4.5 is the least expensive model at $1 per million input tokens and $5 per million output tokens. It is also the fastest Claude model, making it well-suited for high-volume, latency-sensitive applications.
Can I use Claude through Amazon Bedrock or Google Vertex AI?
Yes. All three current Claude models — Opus 4.8, Sonnet 4.6, and Haiku 4.5 — are available through Amazon Bedrock and Google Vertex AI in addition to the direct Anthropic API. Bedrock and Vertex AI offer regional and global endpoint options. Pricing on third-party platforms may vary from direct Anthropic API rates.
Claude vs GPT-4o: Which Model Wins for Everyday Work?
Claude Sonnet 4.6 and GPT-4o are the primary head-to-head competitors in 2026 for professional daily use. They price similarly ($3 vs $3.00 per MTok input) but perform differently depending on task type.
Task Type
Claude Sonnet 4.6
GPT-4o
Long-document analysis (200K+ tokens)
✓ 1M context window
128K limit
Multi-step reasoning
Extended thinking available
o1 series for reasoning
Code generation
Strong; Claude Code natively
Strong; GitHub Copilot integration
Instruction following
Very consistent
Consistent
API cost (output)
$15/MTok
$10/MTok
Context window
1M tokens
128K tokens
The clearest differentiator is context window size. If your workflow involves analyzing full codebases, long contracts, or book-length documents in a single call, Claude Sonnet 4.6’s 1M token window eliminates chunking overhead that GPT-4o requires at 128K. For shorter tasks, either model performs comparably.
Claude vs Gemini 2.5 Pro: How Do They Compare?
Google’s Gemini 2.5 Pro competes directly with Claude Sonnet 4.6 on price and capability. Key differences:
Feature
Claude Sonnet 4.6
Gemini 2.5 Pro
Input price
$3.00/MTok
$3.00/MTok (under 200K tokens)
Output price
$15.00/MTok
$10.00/MTok
Context window
1M tokens
1M tokens
Extended thinking
Yes
Yes (2.5 Pro)
Agentic coding
Claude Code native
Via Gemini API / IDX
Gemini 2.5 Pro is cheaper on paper, especially for prompts under 200K tokens. Claude Sonnet 4.6’s advantage is instruction-following consistency on complex multi-step tasks and the Claude Code ecosystem for engineering teams already in the Anthropic stack.
Which Claude Model Should You Use in Claude Code?
Claude Code supports all four models. The recommended routing for most teams:
Fable 5 — Use for the hardest agentic tasks: large migrations, complex multi-file refactors, long-horizon autonomous workflows. Enable with claude --model claude-fable-5.
Opus 4.8 — Default for serious work: multi-agent orchestration, large codebase analysis, outputs over 64K tokens.
Sonnet 4.6 — Daily driver. Best cost-to-performance ratio for most coding tasks. Extended thinking handles complex architecture decisions.
Haiku 4.5 — High-frequency, low-complexity tasks: formatting, renaming, boilerplate, pipeline steps where speed matters more than depth.
The Max plan (available on claude.ai) unlocks 1M token context in Claude Code at no additional charge, which is the practical differentiator for large codebase work.
Frequently Asked Questions: Claude Model Comparison
What is the best Claude model in 2026?
Claude Sonnet 4.6 is the recommended default for most tasks — it delivers 80-90% of Opus 4.8’s capability at 40% lower cost. Use Opus 4.8 when you need maximum reasoning depth, outputs longer than 64K tokens, or the most recent knowledge cutoff (January 2026). Use Haiku 4.5 for high-volume, speed-sensitive work.
Is Claude Opus 4.8 better than Sonnet?
Claude Opus 4.8 has a higher capability ceiling than Sonnet 4.6: larger output window (128K vs 64K tokens), the most recent knowledge cutoff, and stronger performance on complex agentic coding tasks. However, Sonnet 4.6 uniquely offers extended thinking which Opus does not support, and it costs 40% less. For most users, Sonnet 4.6 is the better practical choice.
What is Claude Haiku 4.5 used for?
Claude Haiku 4.5 is optimized for speed and cost efficiency at $1 input / $5 output per million tokens. It is best suited for high-volume pipelines, classification, metadata generation, social media content, and any task where fast response time matters more than maximum reasoning depth. It has a 200K token context window.
Which Claude model supports extended thinking?
Claude Sonnet 4.6 and Claude Haiku 4.5 both support extended thinking. Claude Opus 4.8 does not. Extended thinking allows the model to reason step-by-step internally before generating output, which improves performance on complex math, science, and multi-step logic problems.
Frequently Asked Questions
What is the difference between Claude Opus, Sonnet, and Haiku?
Claude Opus 4.8 is the most capable model in the standard tier — best for complex reasoning, long-horizon agentic coding, and tasks requiring high autonomy. Claude Sonnet 4.6 balances intelligence and speed for production workloads — it supports extended thinking and adaptive thinking while costing less than Opus. Claude Haiku 4.5 is the fastest and cheapest option, suited for high-volume tasks where speed and cost matter more than maximum capability.
Which Claude model should I use in 2026?
Start with Claude Sonnet 4.6 for most production applications — it offers near-Opus intelligence at $3/$15 per million tokens and supports extended thinking. Use Claude Opus 4.8 for complex multi-step reasoning, long-horizon agentic work, or tasks where quality is worth the higher cost ($5/$25 per MTok). Use Claude Haiku 4.5 for high-volume, latency-sensitive tasks where cost is the primary concern. For maximum capability above Opus 4.8, Claude Fable 5 launched June 9, 2026.
How much does Claude Opus 4.8 cost?
Claude Opus 4.8 is priced at $5 per million input tokens and $25 per million output tokens on the Claude API (per platform.claude.com as of June 2026). Batch API offers 50% discounts. For comparison: Claude Sonnet 4.6 is $3/$15 per MTok and Claude Haiku 4.5 is $1/$5 per MTok.
Does Claude Sonnet support extended thinking?
Yes. Claude Sonnet 4.6 supports both extended thinking and adaptive thinking (per platform.claude.com/docs/en/about-claude/models/overview). Extended thinking lets the model reason through complex problems before answering. Claude Haiku 4.5 also supports extended thinking. Claude Opus 4.8 does not use extended thinking but does support adaptive thinking.
What is Claude Fable 5 and how does it compare to Opus?
Claude Fable 5 (API ID: claude-fable-5) is Anthropic’s most capable widely-released model as of June 9, 2026. It uses adaptive thinking (always on), has a 1M token context window, 128k max output, and is priced at $10 input / $50 output per million tokens. Fable 5 is positioned above Opus 4.8 in the model lineup for the most demanding reasoning and long-horizon agentic work.
What is the context window for each Claude model?
Claude Opus 4.8 and Claude Sonnet 4.6 both support 1 million token context windows. Claude Haiku 4.5 supports 200,000 tokens. All three are dramatically larger than the 200k context window that was standard in previous generations. The 1M context window allows Opus and Sonnet to process entire codebases, long research documents, or extended conversations without truncation.
We track Anthropic’s models, pricing, and limits daily and send a short note when something changes that affects what you pay or build. Occasional, no spam.
💼 Deploying Claude or AI Infrastructure in Your Business?