Tag: benchmarking

  • We Ran 23 AI Models on the Same Article. The Prompt Was Never the Point.

    We Ran 23 AI Models on the Same Article. The Prompt Was Never the Point.

    From the archive

    July 4, 2025 — I posted this on X:

    We tested 23 AI models on the same structured article scope — not just a prompt, but a data-backed framework with embeds, search mapping, and internal RAG.

    Each model got the same Google Doc. Same structure. Same instructions.

    July 5, 2025 — the follow-up, with the writeup:

    We gave the exact same prompt to 23 different AI models. Same input. 23 different outputs. 23 different personalities.

    The takeaway? It’s not just how you prompt — it’s who you’re talking to.

    Fourteen months later, the model names in the results table are history. The finding got more true, not less. Here’s the whole thing, with room to breathe.


    What happened since

    When I ran this test, Claude Sonnet 4 was the new hotness and Grok 3 was still in beta. The industry has turned over since then — the names below belong to mid-2025. Every specific ranking is a fossil, dated July 2025, preserved as-is.

    But the doctrine this experiment produced is now the operating system for everything I build:

    • Model selection beats prompt engineering. The gap between the best and worst model on the same scope was bigger than any prompt trick I knew.
    • Models have personalities. Not metaphorically — operationally. The same way you’d hand a delicate contents job to a different tech than a Category 3 demo.
    • Match the model to the job. Emails, SOPs, longform, code — different work, different worker.

    If you’ve read anything I’ve written about AI costs in 2026, you’ve seen this doctrine wearing different clothes. Model choice is cost choice. The most expensive mistake in AI isn’t a bad prompt — it’s the wrong model doing the wrong job at the wrong price.


    The experiment

    The setup mattered more than the scores. This wasn’t “ask 23 chatbots a question and see who sounds smartest.” Every model got:

    1. The same Google Doc — a structured article scope for “Kitchen Fire Cleanup: Expert Restoration and Prevention Tips,” a real topic from my industry, not a toy prompt.
    2. The same structure and instructions — no per-model tuning, no optimizing for quirks. Deliberately unfair in the fairest possible way.
    3. A data-backed framework — embeds, search mapping, and internal RAG behind the scope, so the test measured how models work with structure, not how they freestyle.

    Why a restoration topic? Because generic benchmarks test generic thinking. I wanted to know which models could handle domain work — regulated language, technical accuracy, a reader who’s trusting you with their home. That’s a harder test than poetry.

    Custom agents ran the harness. The models just had to do the job.


    The results (July 2025 — preserved)

    Same input, 23 different answers — and size wasn’t the differentiator.

    The 5 that stood out

    ModelScoreWhy it won
    Claude Sonnet 4 (Extended)5.0Calm, professional, field-ready tone
    GPT-4.15.0Sharp structure, publishable polish
    Claude 4 Opus (regular)5.0Great flow and clarity
    Claude Sonnet 45.0Excellent default performance
    o4 Mini High Effort5.0Lightweight but surprisingly strong

    These were the ones I’d have trusted, back in July 2025, with full-length articles, client emails, and education pieces. Note the last one: a mini model scored a perfect 5. Size wasn’t the differentiator. Fit was.

    Most improved, second round

    ModelBefore → After
    Claude 3.54.2 → 4.75
    o3 Mini High Effort4.5 → 4.8
    LLaMA 4 Maverick3.5 → 4.6

    Second-round testing with better scope alignment lifted every one of them. The lesson I keep coming back to: even machines do better when they’re understood. The fix usually isn’t a better model — it’s a better briefing.

    The full field

    ModelFinal score
    Claude 3.7 Extended4.9
    Gemini 2.5 Flash4.9
    Gemini 2.5 Pro4.9
    Grok 3 Mini High Effort4.9
    GPT-4o4.85
    GPT o34.85
    Qwen 2.5 72B4.85
    DeepSeek V34.85
    Qwen3 235B4.85
    Grok 3 Beta4.8
    Auto (router)4.8
    GPT-4.1 Mini4.6
    4o Mini4.6
    Mistral Large 24.5
    LLaMA 4 Scout4.1

    Two things worth noticing in the middle of the pack. First, the open models (DeepSeek V3, Qwen) hung with the frontier labs on structured domain work — the gap was narrower than the marketing suggested. Second, the auto-router scored 4.8, within spitting distance of the best hand-picked models. The machines were already learning to choose among themselves.


    The takeaway, then and now

    Then (July 2025): “It’s not just how you prompt — it’s who you’re talking to.” Learn the personalities the way you learn your field crew. Some need bullet points. Some need a whiteboard. Some just need to be trusted to go build.

    Now (September 2026): That sentence became a cost doctrine. Every model has a price per token and a personality per task, and the expensive failure mode is mismatch — a frontier model writing a two-line email, a mini model drafting your scope of work. The experiment’s real output wasn’t a ranking. It was a routing table.

    Here’s the version I’d hand a contractor today:

    1. Run your own version of this test. Not 23 models — three. Your best guess, the cheap one, and the weird one. Same job, same brief. You’ll learn more in an afternoon than in a month of prompt tweaking.
    2. Write down the personalities. Which model do you trust with numbers? With tone? With structure? That’s your routing table. Tape it to the wall.
    3. Re-run it when the generations turn. The names change; the method doesn’t. This is maintenance, not a one-time project.

    And if you ever wonder whether AI is broken or you’re just doing it wrong — remember the test. Same input, 23 different answers.

    It’s not about perfection. It’s about alignment.


    Series notes

    Episode 1 of The Working Years — ideas pulled from the 2022–2025 X archive (@willtygart, account since deleted), given room to breathe. The original posts are quoted verbatim from the archive export. The July 2025 beehiiv writeup “We Ran 23 AI Models on the Same Article Scope” carries the full original tables.

    Next in the series: The Lantern Principle — “the new gold isn’t answering the questions people ask.”