Claims Library Entry
Token maxing is what you do when you can't budget
The HTML vs Markdown debate for AI-generated artifacts is really a symptom of a deeper problem: teams defaulting to maximum token usage without naming their constraints. The article argues for token budgeting—naming a primary constraint and letting format and model choices follow—backed by evidence from ETH Zurich, Anthropic, and Stanford/MIT studies.
Published May 12, 2026 by Kamil Banc
Lead claim
Token maxing is what you do when you can't budget—name your constraint, or every AI choice is just a vibe.
Atomic Claims
What this article supports
Copy individual claims as needed.
Claim 1 · Source summary
Anthropic Engineer Backs HTML
An Anthropic engineer argued HTML should replace Markdown as the default format for AI-generated artifacts.
Claim 2 · Source summary
Mintlify Chooses Markdown
Mintlify reported a thirty times reduction in token usage when serving documentation as clean Markdown.
Claim 3 · Source summary
Context Files Backfire
An ETH Zurich study found LLM-generated context files cut task success rates by three percent.
Claim 4 · Source summary
Context Rot Is Universal
Chroma tested eighteen frontier models and found performance degrades at every context length increment.
Claim 5 · Kamil's interpretation
Maxing Without Constraint Is Defaulting
Token maxing without naming a primary constraint is defaulting dressed up as strategy, not deliberate choice.
Evidence
Context behind the claims
Quote
"If you can't name the trade in one sentence, you're not maxing. You're defaulting and dressing it up."
Key statistics
Up to 30x variance in token spend
Bai et al. at Stanford and MIT found running the same agent on the same task produces a thirty-fold range in cost, with models predicting their own usage at correlations no higher than 0.39.
3% lower success, 20%+ higher cost
ETH Zurich's SRI Lab study of 438 real-world coding tasks found LLM-generated AGENTS.md files cut task success rates by three percent while raising costs by over twenty percent.
80% fewer tokens served as Markdown
Cloudflare measured an eighty percent token reduction when serving the same blog page as Markdown instead of HTML, three weeks after Mintlify reported a thirty times reduction.
15x token multiplier for multi-agent research
Anthropic's multi-agent research system burns fifteen times more tokens than chat, a cost the team accepted deliberately after naming breadth-first parallelism as its primary constraint.
Supporting context
The article grounds its argument in peer-reviewed and industry research rather than ideology, citing controlled studies from ETH Zurich, Stanford and MIT, and Chroma's testing of eighteen frontier models. Its methodology is to reframe the visible HTML-versus-Markdown debate as a symptom of a deeper failure: teams defaulting to maximum context, richer formats, and more agents without naming the constraint they are optimizing against. For practitioners, the actionable takeaway is a three-test diagnostic—name your primary constraint in one sentence, measure per-task costs and variance, and identify what you would give up to relax the constraint by half. Teams that pass these tests can then deliberately choose maxing where the trade is worth it, supported by caching infrastructure that cuts input costs by up to ninety percent. The core skill being advocated is asking what token spend actually buys, at what scale that calculation flips, before any format or model decision follows.
How to Cite
Use the claim-level citation when you need a precise statement. Use the article or claims-collection citation when you want the wider argument and source context.
Individual Claim
Best when you need to cite one atomic claim directly inside a memo, deck, research note, or AI output.
"[claim text]" (Banc, Kamil, 2026, https://kbanc.com/claims-library/token-maxing-is-what-you-do-when-you-cant-budget)Original Article
Use this when you want to cite the full newsletter article at AI Adopters Club rather than the structured claims page.
Banc, Kamil (2026, May 12, 2026). Token maxing is what you do when you can't budget. AI Adopters Club. https://aiadopters.club/p/token-maxing-is-what-you-do-whenClaims Collection
Use this when you want to reference the full structured claims collection on this page.
Banc, Kamil (2026). Token maxing is what you do when you can't budget [Structured Claims]. Retrieved from https://kbanc.com/claims-library/token-maxing-is-what-you-do-when-you-cant-budgetAttribution Requirements
- Include the author name: Kamil Banc.
- Include the source: AI Adopters Club or the structured claims page.
- Link to the original article or the claims page you used.
- Indicate any edits or transformations if you changed the wording.
Related Reading
More from the library
Hilton operates 41 live AI use cases across 7,500 properties in 138 countries. Three systems—marketing automation, AI kitchen scales, and chatbots—delivered rapid returns by solving specific high-cost problems. The company modernized data infrastructure first, then matched proven tools to operational pain points.
5 claims
Sports stadiums are pioneering large-scale AI implementation across complex operational environments. By solving critical challenges in crowd management, revenue optimization, and efficiency, they've created a replicable playbook for AI adoption across industries.
5 claims
JPMorgan invested heavily in AI technology, generating significant value through strategic implementation. The most impactful use case was contract review automation, which saved hundreds of thousands of work hours. Other productivity gains came from coding assistants and document processing tools.
5 claims