Hi! A very interesting piece on reducing tokens. Tokenmaxxing has led to ROI taking center stage, with companies focusing heavily on model routing. Furthermore, companies are adopting alternatives, such as fine-tuning open-source models like Meta's Llama. AI usage per employee is also being limited. I think the tokenmaxxing becomes particularly important in document-intensive work like legal, where lawyers have to review multiple contracts, due diligence reports, draft terms sheets, and fundraising documents.
I write a blog titled "The LegalTech Thesis" wherein I track startups, trends, and opportunities in the space. Would love to know your thoughts!
The biggest savings often don't come from using a cheaper model. They come from making sure the model doesn't solve the same problem twice.
If you're repeatedly sending the same knowledge base, style guide, or documents, you're paying to rediscover context you already have. Structure the workflow so AI extracts the insight once, stores it, and reuses it.
Model routing matters. But reusable memory and better architecture can reduce costs while improving consistency at the same time.
The strategic read here is bigger than the cost savings. When CFOs start routing and caching their AI calls, the story stops being about one company's bill. It's the whole market repricing AI from a land grab into a utility.
Every one of these levers does the same thing in aggregate: it cuts tokens consumed per unit of work. Adoption keeps climbing, but tokens-per-task falls. That matters because the entire AI infrastructure build, the data centres and the listings lining up this year, is underwritten by an assumption that enterprise token demand grows more or less in step with adoption. The minute "tokenmaxxing is dead" becomes the default CFO posture, that demand curve gets a lot flatter than the capex is priced for.
The phase change is the real story here. I'd put maybe 50% that enterprise efficiency gains are large enough over the next two years to visibly soften frontier-model demand growth, even as headline adoption keeps rising. The land-grab pricing and the utility pricing can't both be right for long.
This is really helpful! Thank you
Let’s slow down ChatGPT and Claude growth rates by doing these things!
There is a tool that addresses the subject of the post. I've tried, and it is powerful: https://prismon.ai/
Hi! A very interesting piece on reducing tokens. Tokenmaxxing has led to ROI taking center stage, with companies focusing heavily on model routing. Furthermore, companies are adopting alternatives, such as fine-tuning open-source models like Meta's Llama. AI usage per employee is also being limited. I think the tokenmaxxing becomes particularly important in document-intensive work like legal, where lawyers have to review multiple contracts, due diligence reports, draft terms sheets, and fundraising documents.
I write a blog titled "The LegalTech Thesis" wherein I track startups, trends, and opportunities in the space. Would love to know your thoughts!
https://harshithviswanath.substack.com/
The biggest savings often don't come from using a cheaper model. They come from making sure the model doesn't solve the same problem twice.
If you're repeatedly sending the same knowledge base, style guide, or documents, you're paying to rediscover context you already have. Structure the workflow so AI extracts the insight once, stores it, and reuses it.
Model routing matters. But reusable memory and better architecture can reduce costs while improving consistency at the same time.
The strategic read here is bigger than the cost savings. When CFOs start routing and caching their AI calls, the story stops being about one company's bill. It's the whole market repricing AI from a land grab into a utility.
Every one of these levers does the same thing in aggregate: it cuts tokens consumed per unit of work. Adoption keeps climbing, but tokens-per-task falls. That matters because the entire AI infrastructure build, the data centres and the listings lining up this year, is underwritten by an assumption that enterprise token demand grows more or less in step with adoption. The minute "tokenmaxxing is dead" becomes the default CFO posture, that demand curve gets a lot flatter than the capex is priced for.
The phase change is the real story here. I'd put maybe 50% that enterprise efficiency gains are large enough over the next two years to visibly soften frontier-model demand growth, even as headline adoption keeps rising. The land-grab pricing and the utility pricing can't both be right for long.