Token Pricing Economics

Discussion of whether current API prices are subsidized, the gap between enterprise and consumer pricing, competition driving prices down, and whether AI companies can sustain current pricing models while recovering training costs

← Back to Uber's $1,500/month AI limit is a useful signal for AI tool pricing

The AI landscape is currently defined by a volatile "tokenomics" struggle, where heavily subsidized consumer subscriptions often mask the astronomical real-world costs of enterprise API usage, forcing major companies to impose strict per-developer token budgets. While aggressive competition from high-quality, low-cost Chinese models suggests a "race to the bottom" for inference prices, some observers warn that massive infrastructure debts and rising energy costs will eventually necessitate significant price hikes. To navigate this uncertainty, users are increasingly pivoting toward model routing, local hosting, and even adopting more concise programming languages to minimize their token footprints. This creates a precarious economic dance between the rapid commoditization of machine intelligence and the desperate need for frontier labs to recover trillions in research and development expenses before their investment capital runs dry.

124 comments tagged with this topic

View on HN · Topics
> I noted that my own token usage comes to about $1,000/month against each of Anthropic and OpenAI - which currently costs me just $100 per provider thanks to their generous subsidized plans for individual subscribers. Do we know that AI providers are going to keep these per-token prices, or eventually lower them because of competition from China? Many lower-budget individuals are now moving to China open weight models like DeepSeek. I wonder if China's really subsidising the providers, or if inferencing costs are actually much lower, and Anthropic/OpenAI are just making sure no money's left on the table for their eventual IPOs.
View on HN · Topics
We can tell that the inferencing costs for many of these models are low enough that these models are being sold close to real costs on the basis that many of them are open weight and available from third party providers who have no incentive to subsidize them. I think the frontier labs will need to drop their high per-token prices at least for their low and mid-level models for the reason that several Chinese models (at least Qwen, DeepSeek, Kimi and GLM) are "close enough" that with the right harness they are cost effective alternatives. They won't necessarily need to close the gap - at least not yet -, because these models won't necessarily compete at the same token counts . E.g. at least some of them need to do far more work to solve the same problems. But, yeah, the prices will come down one way or the other. At the same time, even the subscriptions for the cheap Chinese models are probably subsidised, and those subscriptions are likely to get less generous over time.
View on HN · Topics
One aspect Paul Kedrosky mentioned recently is the concept of „duration mismatch“. The price per token goes down over time (either because the AI vendor reduces due to competition pressure, or because customers are now incentivized to use older cheaper models). But datacenters are financed through debt, with the assumption their revenue increases over time. Quoting him: „[AI vendors are] paying for a fixed cost with a depreciating commodity“[0]. So you have on one end the token revenue trending down, on the other end the training cost going up for the next frontier models, and you need to pay back your 10y debt. 0: https://youtu.be/wGZboZcSGDY?is=64GuKyqBh_4aSjTE
View on HN · Topics
A few things, I think you’re missing the point here - most tasks do not require the latest frontier models, even if they are a magnitude more intelligent (we don’t actually know if that will be the case). Current Gemini flash is cheap, fast, and pretty capable with good guidance for most tasks - now that companies pay API costs instead of a subscription they will be setting restrictions on token use to not have their budget explode (like Uber in this submission), that’s a strong incentive to NOT use expensive models, and limit their thinking budget - there is competitive pressure from China and others who can offer very decent performances at a fraction of the token price - the price of tokens for the frontier models is likely to go up, but the price to access older models is what depreciates! The overall price per token is going down now that we are in a new world where companies understand that token maxing is one of the stupidest concept ever created by humankind.
View on HN · Topics
If you have a good model router, you can route to older, cheaper models that run on older hardware, for simpler tasks. That helps labs extend the economic life of their hardware investments. They will likely fight it at first though as they see it as reducing ASP. This is why I'm building role-model, a routing protocol and a router runtime: https://role-model.dev/
View on HN · Topics
The other part of that is that while price per token may be going down, tokens per task is going up
View on HN · Topics
For ~equivalent tasks/results, or because we’re expecting more or better from tasks? The real measure should be cost per ~equivalent task result, not cost per token nor tokens per task.
View on HN · Topics
For better performance of ~equivalent tasks. That's what all the harness tooling people are using does: (often) increasing output quality by significantly increasing token counts.
View on HN · Topics
There are data centers that use and rent out 10 year old server GPUs. They can't run larger modern models. They can't run smaller models as fast as newer servers. So their remaining market is applications where customers are okay with older, smaller models and slower performance. They have to price the service lower than competitors due to the lower performance. The older GPUs are less efficient so it costs them more to keep them running. They're paid off, but they're taking up valuable power, space, and cooling in a data center. Eventually there is a tipping point where it's better to replace that space and power budget with something new that has more demand. The parts are sold off on the open market. There's an equilibrium demand for the parts from other data centers keeping older servers running and from hobby people who are okay with a jet engine sounding toaster of a GPU running in their home.
View on HN · Topics
I would presume the reason they are overclocked is because they are trying to make up for the shortage. In time, the shortage of computing components will be remedied, and tokens produced at lower power pulls will be cheaper.
View on HN · Topics
GPU do depreciate indeed, but here the depreciating commodity is the token, not the hardware. You sell cheaper token with the same hardware
View on HN · Topics
When everything is said and done it'll be datacenters in American competing with ones in China that have several times lower electricity prices. Token prices will drop to a level that will be unprofitable for American data centers and they will need to close. Thats the main issue here.
View on HN · Topics
I don't think they'll offer open models for long. Since they've actually invested in power, cheap chips, cheap memory and can subsidize tokens - they'll keep undercutting big models to capture data forever. Bonus if they remove ridiculous safeguards and China will be unstoppable.
View on HN · Topics
Pretty sure they'll offer them at least so long as it takes to bring OpenAI and Anthropic into insolvency. Why wouldn't they? The Chinese models are way more nimble to train and run, bring in a ton of goodwill globally, and put immense pressure on the VC furnace that is the US AI sector. And apparently OpenAI and Anthropic think so, too - why else would they try so hard to ban them instead of outcompeting them?
View on HN · Topics
Raise them, more likely. NVidia says that GPU hardware prices won't decrease until at least 2030. The world is out of fab capacity.
View on HN · Topics
Seriously, they’re trying to justify trillion+ IPO’s while setting piles of money on fire, prices aren’t going DOWN.
View on HN · Topics
Today's frontier models will be tomorrows low-end option. I think whatever model you are using today will be less expensive to use a year or two from now.
View on HN · Topics
Last year's o3 was more expensive than 5.5 is. Whatever model we are using now is probably be more expensive than next year's leading models will be.
View on HN · Topics
Price per M/tokens is also a fuzzy metric when newer models reason longer, and then burn more tokens while doing so.
View on HN · Topics
Isn't 5.5 a router, though? As in, some prompts get automatically sent to a cheaper model?
View on HN · Topics
Why would I even pay for deepseek? I get deepseek v4 flash for free with opencode. If I somehow run out of tokens for the day, I can just then on my vpn
View on HN · Topics
I wonder if I could start a US-based company with good data regulation and just serve open-weight models at a competitive price. I feel like the real barrier is just that most companies willing to adopt AI usage enough to make it worth it at this point don't want to be using inferior models.
View on HN · Topics
Here's a free startup idea: operate an open-weight model service, and offer "Verified AI Integrity," which signs the input tokens, the seed for the randomness in selecting outputs, and the model ID, proving that the result of the call to AI was completely "organic" and was not interfered with. Your main audience would be snake oil salesmen trying to prove their AI products are unbiased and not under the thumb of any outside influence. This doesn't address the biases of the model itself, but that's not your business. Your business is selling tokens and security certificates. If you can get the right angel investor, you could maybe have your new standard required for some government applications.
View on HN · Topics
Yes, you can. There are multiple inference providers out there. The problem is, it’s hard to beat the Chinese providers in cost. And you also have to compete with frontier model providers’ subsidized offerings.
View on HN · Topics
They charge the exact same prices. So many people in these comments have no idea what they're talking about. Even if they did charge less, nobody is going to deal with the latency of sending requests to China. edit: Actually American inference providers are cheaper for Chinese models. There's way more competition here because the Chinese aren't idiots and investing every last dollar they have into data centers for llms that don't make money..
View on HN · Topics
Can you please link me DeepSeekV4 provider that's cheaper than their official offering? And not all tasks require low latency. Also, there are a lot of competition in China. Like a lot. You might know better than me as well, but although the biggest AI-labs are based in USA, the adoption is weirdly global. Like as a general sense of what's going on - you can see AI-related ads literally everywhere in Tokyo, almost all the time, in every single screen in public.
View on HN · Topics
Cro.ai seems to be: https://crof.ai/ Of course though they are not necessarily a viable solution for companies with security requirements etc. given it is just a single person project, but they still serve as a proof it can be done.
View on HN · Topics
This costs more.
View on HN · Topics
Not as far as I can tell. Are we seeing different things? For deepseek-v4-pro: - $0.350 in, $0.003000 cache, $0.80 out https://crof.ai/pricing - $0.435 in, $0.003625 cache, $0.87 out https://api-docs.deepseek.com/quick_start/pricing
View on HN · Topics
Deepseek's api platform for V4 Pro is the only example of this, and Deepseek V4 Flash is cheaper (usually) than from Deepseek itself on openrouter via DeepInfra. Deepseek shot themselves in the foot because they never intended to serve V4 Pro for .80c mm ouput, that was a promotional price that was meant to expire (and still might). They intended for v4 to cost $4.00 per million but Western inference providers drove down the price because they can operate at negative margins to try and push competition out. I can assure you they are losing a ton of money @ ~80cents. My point is, its Western inference providers that are establishing the floor price of inference. They are willing to operate at a loss in order to put their competition out of business. Chinese providers are typically at or above the prices set by American/western providers if you go looking on the Chinese internet. You aren't going to get deals from China for inference except through this one instance with Deepseek v4 Pro which wasn't even supposed to be permanent pricing.
View on HN · Topics
By "cost" I think the parent means the provider's own costs, not the cost of inference to the customer. The cost of land, labor, and electricity are significantly lower in China than in the US.
View on HN · Topics
There are plenty of US-based inference providers available, including AWS, that serve Chinese models at competitive prices (vs frontier US models). They also have lots of usage. Not necessarily for coding, but for other enterprise tasks.
View on HN · Topics
Have you heard of openrouter? There's 1000 of these companies already. Do something else.
View on HN · Topics
> Do we know that AI providers are going to keep these per-token prices, or eventually lower them because of competition from China? Raise, they are going to raise the prices. We will spend more on AI infrastructure in 2026 and 2027 than the gross sales of the entire global software and services sector. Current pricing is at a major loss for current providers.
View on HN · Topics
Per token costs will fall, but the harnesses will get more token hungry. Instead of just centering the div it’ll spin up a battery of agents to architect, critique, advise, code, review, refactor and so on.
View on HN · Topics
They're going to need to bring in a few trillion dollars fast to meet wall street expectations. Expect prices to rise.
View on HN · Topics
> Do we know that AI providers are going to keep these per-token prices, or eventually lower them because of competition from China? Are they even making money off them now ?
View on HN · Topics
> Do we know that AI providers are going to keep these per-token prices, or eventually lower them because of competition from China? I genuinely do not know how prices can get lower from the current major providers in NA without the whole market collapsing. Everyone is spending copious amounts of money to presumably make more money back.
View on HN · Topics
An inference only platform selling good open weight model inference without the research overhead could capture a-lot of market for lower size model uses (haiky, gemeni flash). Diffusion-transformers and clever cashing can drop inference even lower, which is improving at a high rate. The biggest reason large models are un-attainable for local applications is the lack hardware with large amount of unified/graphics memory (and the cost of the platforms that do). Once the memory slog goes back to normal and hardware manufacturers adapt to demand, we may see consumer hardware with large memory capacity effectively opening the door for slow but usable frontier model inference (assuming improvements in model efficiency and compute capacity) At that point, inference becomes a race to the bottom. The large labs hope they can attain a leap in capability (which is increasingly looking bleak, with a average catch-up of just a few months) or market dominance through integration (integration in platforms and OS, exclusive deals with companies or governments). For coding agents, i suspect no player will manage lock in enough market to enforce pricing much higher than the true inference cost, and catering to programmers becomes an unsustainable proposition. We will instead be further hit with a lot of AI integrated into our other tooling costs, such as GitHub, Microsoft suite, G-suite, forcing in AI functions as a value-ad into the total cost without giving the option to exclude them. (using their market position)
View on HN · Topics
AI may get so commoditized for certain use cases that you will not even be able sell inference at a profit. AI might be bundled in with other services, just like cursor bundles in their own AI model for auto complete with their editor. I.e. cameras might have AI for image recognition bundled in etc.
View on HN · Topics
Agreed, this is where google is really, really set up to win the market. They can combine gemini subscription with a moderately more expensive google workspace and steal MSFTs entire $50 billion enterprise productivity software market. MSFT is quickly trying to get copilot in a good enough state but without TPUs I think itll be tough for them to serve a good enough model at a price people will accept.
View on HN · Topics
I agree with all of this. So my question remains the same: How are the players investing 100s of billions in buildout going to hope to make this back? Market capture looks bleak, inference looks like a race to the bottom. End users look like they could be beneficiaries. Where do the big boys go?
View on HN · Topics
Prices can go down while tokens sold increases so that profit increases. The labs number one goal right now is moving past software engineers so that every white collar worker in the country finds ai assistants indispensable. Speculation here but I think openAI/antrhopic api inference is insanely profitable, it just needs more volume to amortize the training costs.
View on HN · Topics
API prices of Anthropic, OpenAI, and Google are massively inflated. https://martinalderson.com/posts/no-it-doesnt-cost-anthropic... There's no way that all AI inference providers are colluding and/or all running at a massive loss, meaning the cheap Chinese model prices must be the real cost it takes to run frontier-class models PLUS their margin. Look at Deepseek 4 Pro. https://openrouter.ai/deepseek/deepseek-v4-pro/providers Deepseek and Baidu are subsidising prices but they probably train on inputs. I have no model training and ZDR in OpenRouter enabled, and the first provider that shows up there is Deepinfra, significantly more expensive than Deepseek. BUT much cheaper than Sonnet 4.6 and ChatGPT GPT-5.4.
View on HN · Topics
It's pretty simple; organizations are willing to tolerate paying $1500/month/engineer, which seems to be roughly inline with "normal" consumption for most full-time engineers. If that number grows significantly, then I bet companies will start exploring flash models more, as you propose.
View on HN · Topics
> organizations are willing to tolerate paying $1500/month/engineer One organization, that is a software company > which seems to be roughly inline with "normal" consumption for most full-time engineers My peers are using $20/mo plans, only a handful are using more than $100/mo in tokens. We haven’t had any limits imposed yet.
View on HN · Topics
It’s also worth noting that’s the peak benefit. Expect most engineers to not hit those limits on the regular (if at all, since limiting this puts skills in focus again), and that limit to come down over time as the easy processes are automated and humans are re-tasked with harder problems relative to their TC. This is not a good bellwether for the AI industry, including its adherents. Their growth assumed a level of indispensability that’s not being reflected in hard numbers and real costs, which lends credence to the notion that these IPOs being fast-tracked are meant to try and cash out before the bubble really pops in earnest. There’s no way consuming enterprises are going to pay such insane costs for such minimal uplift in the long run, and the AI companies can’t keep offering subsidized tokens via subscription plans at their current pricing.
View on HN · Topics
I can see a corporate future where tokens are haggled over in department budgets just like any other line item. Some projects will get more of them, other projects will get less of them. "Use AI for everything" will become "use AI economically and build things that outlast our budget for it."
View on HN · Topics
> "unlimited tokens for every employees we don't even care if it actually ends up being a net positive financially" That was clearly a short-term trend that would obviously get fixed. Doesn't say much about AI coding as a business model.
View on HN · Topics
No disagreement on computing 2.0, but companies spending 3-5k per employee for hardware isn't generally a monthly cost. It's a at the time of hire, and then once every 3 to 5 years after that, for a monthly amortized cost of about $50/employee. I have my concerns with current inference pricing in that there's a non-zero possibility for a rug pull in the future for the subscription plans for organizations and individuals that can still use them. For now, its only companies larger than ~150 users that need to pay per token, but what if that wasn't the case? Not every company can afford over $1k/month/employee to give them access to AI tooling, further making it harder to compete against the behemoths. If we get to a point where an individual can no longer pay $100/month for nearly unlimited usage and instead must pay per token, that's going to be a problem. Personal computing eventually became an equalizer (until we started centralizing on mainframes again, aka the cloud) because it got cheap. My hope is that inference also gets just as, if not cheaper. I have high hopes for local AI and open weight models and we will continue the ethos of local, personal computing and not needing to offload everything to OpenAI/Anthropic/Google, etc. to get work done once the hardware and hardware availability catch up.
View on HN · Topics
I wonder if you will see app makers begin to open APIs (MCPs) up in ways that replace computer use. Computer use via human interfaces is pretty hacky IME, and if you can use an app that exposes spreadsheets in a way that reduces token costs by 90%. I'm optimistic that the demand for AI accessibility will drive programmatic interfaces in places where companies were previously reluctant to.
View on HN · Topics
It's cope. People desperately want to believe that AI coding is going away so that they can go back to partying like it's 2020. So there's a huge number of HN posters claiming that the price of tokens will go UP over time rather than down (that's how Moore's Law works, right???) or that code bases that AI contributes to will spontaneously combust, or something.
View on HN · Topics
> So there's a huge number of HN posters claiming that the price of tokens will go UP over time rather than down (that's how Moore's Law works, right???) I mean, Github Copilot's pricing just went up considerably, so I guess they were right?
View on HN · Topics
I don't think it is unreasonable to say both will happen, is it? In the long term, tokens will fall in price. Obviously. (If "tokens" continues to be the unit) In the short to medium term, for the IPOs to succeed, people have to start actually paying for what they are using, so the price will go up, and is going up, quite a lot. Once their value is set they will slowly fall from that point (or some point maybe halfway, depending on how much the market is willing to continue to subsidise). I am an AI cynic, but I am now an informed cynic; I am learning agentic tools so I know where they are useful and I know my enemy. I think the "fad" here is cloud-based, metered AI being a dominant work mode. Nothing, so far, has suggested to me that any other outcome is likely than edge- to local-scale, on-device, on-laptop, on-prem models getting good enough to the point where people use them by default and use the cloud models only when they need the extra oomph. I cannot believe that there is anything other than an enormous incentive for companies like Uber to find local, small model and on-premises solutions to their problems, not least while pricing is so changeable and people are getting nasty surprises. Betting on OpenAI and Anthropic being around over the long term in the form that they are now, that feels like valley hopium. Utility monopolies essentially always derive from physical/geograpical limitations, don't they?
View on HN · Topics
Token costs do go down over time for sure due to software optimizations (i.e. better attention kernals) but acting like hardware INFLATION isn't happening for at least a few more years is just nonsense. Objectively an A100 is more expensive to rent today than it was in 2024 (a 7 year old GPU - Big short guy is a turbo idiot) and rising. As such, over short time horizons, it's possible to see limited amounts of "price per token goes up" for the same model.
View on HN · Topics
I'd think for most companies the pace of change is too high at the moment. Give it a few years, a bit of a plateau in the improvements in frontier models and I can't see how many of these companies don't implode under the weight of competition on inference prices.
View on HN · Topics
Plenty of comparisons here between salaries and token costs. All fair but very much assumes that salaries are rational. Why do we pay some engineers 10x as much for the same role just because they are in a different location? The WFH discussion surfaced some of that. If money is cheap, all sorts of funny things are happening. Is it worth to spend 1500 USD on AI? I don’t know. Is it worth paying engineers 300k USD instead of 30k? Honestly, I don’t know
View on HN · Topics
$1500/mo is $18,000/seat/annum. Maybe Microsoft and Nvidia are on to something. 128 GB machines that can run local LLMs are a bargain even if priced $5-8k. Yes, tok/s is not quite there, but that's probably OK since the bottleneck really isn't the code; it's WTF did Uber build with all of that spend? How did it meaningfully impact their revenue in a positive direction?
View on HN · Topics
How is tok/s not a bottleneck I? I assume most people still use ai agents interactively rather than leaving them to do their own thing during the night. I find anything below 50 tps or so entirely unusable... Regardless its Apples to oranges anyway, inference is quite cheap for open weight models its just that Claude and OpenAI can charge very high margins compared to e.g. DeepSeek or various provider on OpenRouter since open models are a commodity.
View on HN · Topics
Sure, but has their rate of value added increased as a result? It's a good question to ask. They added value before LLM coding, and now are more expensive than before thanks to token costs.
View on HN · Topics
> How did it meaningfully impact their revenue in a positive direction? It probably allowed them to avoid hiring as many people to build a certain amount of software. Even if it didn't increase revenue, it could have lowered human labor costs. > 128 GB machines that can run local LLMs are a bargain even if priced $5-8k. Don't forget the energy costs. Searching around, advanced models use an average of 25 Wh/1000Tok. $1500/month gets you about 150M tokens. At the aforementioned energy/token, that's 3750kWh. What are your local office electricity rates/tariffs? (Hint: they are going up because of AI data centers). Even if my price and energy assumptions are wrong above, you probably aren't going to get the rates that the hyperscalers do. Even at cheap (i.e Texas) retail electricity rates, that many tokens will probably cost you hundreds per month. In most other electricity markets, probably far more.
View on HN · Topics
That's going to stop eventually, and I think at that point we're going to see business models more like the major CAD providers.
View on HN · Topics
I don't think they'll have a choice, open weights models are not far behind. At some point it's essentially a commodity game
View on HN · Topics
they also already do this… Anthropic and OpenAI license to the public clouds. Google reportedly licenses to Apple. licensing to Fortune 100 companies running on their own infra is an obvious next step it is a race to the bottom and I’m not sure the labs win that race. we’ll see!
View on HN · Topics
The problem isn't really Uber, Microsoft or Nvidia, it's all the smaller none IT companies that also have developers on staff. They are screwed. $1500 per seat per month is just way to expensive, but they also can't afford to build and maintain their own on-premise solution. If Microsoft can't afford to run CoPilot for their own developer, what chance does any of their customers stand? If the large, well founded IT companies in the world believes the current AI cost is to high, then Anthropic, OpenAI and CoPilot have no actual customer base. AI is then relegated to very profitable niche business, but that can't fund the R&D for the models.
View on HN · Topics
It's an extra 18k a year for developer tools when they're paying how much a year per developer? Having software developers at all isn't cheap. Also, I don't believe you need to spend $1500 a month on a coding agent if you optimize usage at all.
View on HN · Topics
In Latvia, the net salary for a Java dev is around 1729 - 4314 EUR, based on https://www.algas.lv/algu-informacija/informacijas-tehnologi... (crowd sourced data) For the employer those employees cost between 2945 - 7736 EUR per month based on https://kalkulatori.lv/lv/algas-kalkulators (income and social taxes). So on the lower end that's (1500 USD ~ 1300 EUR) close to half the total expenses of such a developer, on the high end here around 15-20%. That's quite significant, depends on whether their productivity also improves (if that's what the orgs care about). And we’re not even the country with the worst pay out there, but pay the same for tokens, cause regional pricing isn’t a thing!
View on HN · Topics
$18k a year is a non starter in most companies. Ive seen companies balk at Intellij.
View on HN · Topics
That depends on where you are. $18K is the equivalent of paying around 15% more for your developer.
View on HN · Topics
There's models for every price point. What was SOTA and stupid expensive to run a year ago is a cheap flash model today.
View on HN · Topics
That was badly worded on my part, my intend was to indicate that there was no way they can or will pay $1500 per month per seat.
View on HN · Topics
I don't see it. Leasing equipment and paying per seat license fees makes a lot of accounting and cash flow sense. Maybe when it gets to the point where you can run SOTA LLMs on consumer hardware. But that seems a solid decade and probably much more away. Even then it makes more sense to rent the bigger GPU and get your answer faster.
View on HN · Topics
18k/yr? None of the LLMs generate anything like that in value!
View on HN · Topics
Can you share some examples that you would say justify that price? Not a gotcha, I’m genuinely curious where you’re seeing a return at that level.
View on HN · Topics
I use the $100/mo sub but my 30 day API cost is about $1700/mo. It really depends how you use it, if you're using prompts to generate detailed designs, breaking those into lists of tasks, and then feeding those to multiple agents - it's really easy to burn through many thousands. If you're being more deliberate and using a few agents at a time interactively, having it review PRs/resolve issues, automated clean-ups and performance optimization, etc it could be more like $1500. If you're just throwing it one-off questions like a better stack-overflow that is well under a $100. I've really gotten into /goal, if you can find something verifiable and leave it overnight - it's kinda like christmas morning to see where it landed.
View on HN · Topics
Just to put this in context. If every company did this, all over the world, with that same limit, we are talking about something around $45B monthly in revenue for all AI companies to share.
View on HN · Topics
There are a lot of places in Europe where 1.5k$ is more than 50% of the total cost of an employee. And the obvious question: what it's the cost of that revenue? Because it looks huge but ...
View on HN · Topics
World bank says there are 3.7B employed humans. Putting the total addressable market at around 67T if all of us spend USD 1.5k on tokens every month. This lines up well with current forecasts from the major AI labs
View on HN · Topics
I think the main thing companies should try to understand is avoiding the use of 'claude -p'. I definitely have written a goal file, and then just ran claude in a loop over the goal in order to 'token max'... why not? I'm doing research and have some clear KPIs where research into all kinds of techniques / tuning can improve the results. I can spend my budget on a "experiment with blah blah blah to improve blah blah" or give it a list of things to try that I know will take awhile. Its no problem hitting hundreds of $ of API spend while sitting at a computer with 3 monitors have 6 windows of useful claude code interactive sessions, while working on 2 or 3 projects and using worktrees, and it's a little weird when you hit your limit by 2 o'clock and have to wait for token budgets to reset; god forbid, I manually edit code... which I did do for the first time in months. You can also start to generate a lot of token spend if you do something like "hey make me a stylized slide deck using internal skill / agent XYZ based on commits A through C", which as an engineer, makes presentations building much less painful . This uber limit is not high compared to the big SV companies.
View on HN · Topics
All new/renewing enterprise contracts with Claude Enterprise and ChatGPT Enterprise no longer offer usage-based subscriptions, but instead will charge API pricing for all tokens consumed, and as you've said, the subs are better deals than raw API pricing.
View on HN · Topics
I use Claude every day. Often for multiple hours a day. Basically doing my job not worrying how many tokens I spend (as in too many or too few). This is a pretty complex code base (database optimizer and related). Just looked at spent for the past 30 day, didn't even come to $600. 95% of my tokens are from cache. If I were to reach even $1500 I have to let claude run unsupervised over night (and with the amount of mistakes it still makes and guidance it needs, I do not believe we are there yet.)
View on HN · Topics
Why isn't self hosting (even just renting a GPU server, not necessarily on premise) at large companies or hosting via something like together AI to run the open weight models not more common? I've tried the open weight models and the premium models like Opus and Gemini Pro, and I find that the latter are a little better, but not nearly to the degree to justify the extreme price difference, since the differences largely don't matter for what I've tried them for, and I expect that many other users likely have similar use cases.
View on HN · Topics
If the premium models are just about 10% better - that could justify the price vs. self hosting a ~0.5-1T open weights model. Remember that utilization of these huge racks will not be 24h/7, and these are usually not GPU intensive shops that would train models on the spare compute. With prices of 100-200k USD and north with ~2 years lifetime, that would be hard to justify financially. Self hosting could easily amount to ~1000 USD a month amortized across many developers. In rush hours - there will be hard rate limits. Would that 1500-1000=500$ monthly USD justify the 10% decrease in "AI Productivity" ? I guess not. In most cases. For everyone that asks me around, I'd say that in short term, unless there's a really good reason to self host these coding assistant models, then the big 2/3 coding assistants providers are the better choice. No one got fired from licensing claude code.
View on HN · Topics
I would assume that at least one of two things are true: 1. They're costs are so so out of control that they need to impose a blanket cap immediately. Figuring out an allocation mechanism that can be deployed company wide is time consuming and they need to staunch the bleeding immediately, despite it being obviously suboptimal. 2. The few people who should have unlimited tokens were given exactly that. No reason to introduce such nuance to a public PR move. The hard-cap limit is a great negotiating posture with token providers.
View on HN · Topics
$1500/mth is token pricing. Your other plans are fixed price with rate limits where you get more tokens than the dollar equivalent you pay monthly. These plans are economical only if majority of users spend less tokens in $ than the plan's costs. This subsidizes the gap vs. power users who spend multiple k$ monthly in API tokens.
View on HN · Topics
> Your other plans are fixed price with rate limits where you get more tokens than the dollar equivalent you pay monthly. Or the fixed cost plans reflect the real cost and the people paying API prices give them the profit. Anyway, none of my customers will let me bill them $1500 more (about $75 per day) because I'm using AI. And what for? I'm not working to move money from the pockets of my customers to the pockets of AI companies.
View on HN · Topics
No, we know from the financials of these companies that API prices are close to being at cost and the individual developer plans are heavily subsidized (because they are roughly 10% of API cost per token[1]). If plans were at cost and API pricing was marked up that would mean there’s a 90%+ profit margin on tokens and instead of raising money and talking about revenue, Anthropic and OpenAI would be talking about their obscene profits. [1] the caveat is that the average plan user probably doesn’t use all of their quota, I guess maybe 30% is the average across all users.
View on HN · Topics
This completely ignores all the other huge costs the AI labs are paying in data center builds, researcher salaries, experiments, and training models. The fact that Anthropic is rumoured to have a profitable quarter indicates that their margins on API priced inference are very strong.
View on HN · Topics
Next to no one would be using less than the subscription price given how expensive Opus API is.
View on HN · Topics
Yea, I’m sure the personal plans are subsidized. I have $200 Claude Max at home and straight API pricing at work and equivalent work would easily cost me 5x if not more on the API.
View on HN · Topics
Uber is likely on an enterprise plan - these charge tokens at API cost, which can be much more expensive than the $20 flat rate.
View on HN · Topics
These are still at currently subsidized prices. We'll see if they think they're getting $1500/month of value when that buys significantly fewer tokens.
View on HN · Topics
There is no evidence that per-token inference prices (which is what Uber is setting a cap on) is subsidized.
View on HN · Topics
The evidence that per-token inference _is_ subsidized is (a) competition is a bloodbath (b) these companies are raising more money than any company has raised ever (c) a maybe-profitable quarter is maybe-coming for Anthropic after maybe-signing a compute deal with SpaceX that legitimizes both companies. The evidence that per-token inference _is not_ subsidized is... a quote or two from Dario and Sam Altman
View on HN · Topics
AI companies have more expenses than inference.
View on HN · Topics
yes, and theres no evidence that they arent (or can't) use profitable inference to subsidise those other expenses. Some companies will keep spending massively to train better models, and some other companies will not, and offer good api prices. Which will end up being used? That depends on whether the spending turns into better value models
View on HN · Topics
> theres no evidence that they arent (or can't) use profitable inference to subsidise those other expenses as far as we know there's no evidence that they can produce any profits at all
View on HN · Topics
Is there any evidence that it's not?
View on HN · Topics
The fact that Anthropic models are offered at the same API pricing by not just themselves but AWS, Azure and Vertex despite Anthropic taking a major slice on licensing along with the cost an open weight 1T parameter model like K2.6 costs to run on any third-party provider, make it unlikely that API inference cost are subsidized by the labs.
View on HN · Topics
Openrouter? i.e. Even excluding Deep Seek inference for very large open models is way cheaper. Maybe these providers are not very profitable but its highly unlikely that they are losing $4 for every $1 they make since selling inference is their only product...
View on HN · Topics
That's just market segmentation and them trying to maximize revenue it doesen't really say anything about their costs.
View on HN · Topics
That's not evidence. Very likely though, but the only evidence we get one way or another is when they IPO.
View on HN · Topics
This story isn't about those subscriptions - enterprise customers like Uber are paying the full API prices.
View on HN · Topics
afaik, enterprise plans are not subsidized. its 20$/seat+api pricing. Unless you are saying api pricing itself is subsidized.
View on HN · Topics
This is market introductory pricing that hasn't factored in cost recovery. Most of it has been run on early investment with the assumption they will recover costs in the long run. The prices are subsidized across the board and they will need to go up signficantly to recover them.
View on HN · Topics
Assuming this were accurate, then presumably the AI companies would be betting that inference costs come down before the bill is due - I don't see enterprises being willing to absorb another ~10x price increase for tokens (as they've just done going from subscription prices to per-token pricing)
View on HN · Topics
For claude shops this was a huge hit. But lets back this up. There are some companies that haven't even built a break-even model at this price because they are funded by investment. As soon as those investors lose patience the first dominos will fall. For those who have somewhat of a business model, will it survive a price increase? The bigger question is do the base model providers have enough runway and have a way to keep going as they need to recover costs.
View on HN · Topics
It's mostly R&D though, not inference. If LLM's effectively become a commodity then they are screwed anyway.
View on HN · Topics
True but they will raise prices slowly so people will optimize their workflow so they aren't just throwing as much inference as fast as possible like the current state. Right now you should do everything you wanted to try out because it is cheap (as long as you don't become dependent ... the risk).
View on HN · Topics
The inference prices for very large open models would indicate that Antrophic's and OpenAI's margins are quite large.
View on HN · Topics
It's not. They recently forced enterprise customers onto API billing instead of the cheap consumer pricing. Now the pricing is brutal.
View on HN · Topics
And $1500 a month is on the very high end of where most companies will land. When you run the numbers there isn’t a realistic path that connects the dots between likely market size and the claimed valuation of the AI companies. The math simply does not add up.
View on HN · Topics
ccusage for codex tells me the medium feature I prompted in codex, with a $200 subscription, running for 72 hours and still not delivering full result would have cost ~ $2200 at API rates. I also misconfigured something in my agent's configuration and a simple web tool request (maybe 4 turns) through OR went to GPT-5.5 accidentally and that cost me ~$0.4. I have no idea how any business can afford API rates without having a mindset of casually setting money on fire.
View on HN · Topics
Seems odd limit, especially since it highly dependant on Token provider used, with Opus this is not much and could easily be burnt in a week or less, but with something like deepseek the 1500 can literarily be an annual budget. That being said, I do have to wonder why someone as bug as say Uber, simply not rollout OSS model in the cloud for their team, I'd imagine that would be cheapest & most flexible option, while also keeping all the data shared with LLM private.
View on HN · Topics
eventually tokens will cost price of energy. and china is miles ahead. china will be major token exporter soon. mark my words.
View on HN · Topics
If I were paying API rates this year, I would have already burned through $20k in tokens. Looking forward to the costs of this level of capability coming down.
View on HN · Topics
I think a lot of people are missing that this is $1500 _per tool_ which is still rather a lot of money.
View on HN · Topics
Reading the headline Oh that's actually really economical! I wonder if they're doing a lot on locally running models or managing a shared context or knowledge-base in some clever way, maybe just encouraging employees to be efficient and mindful. ... > each employee ... > per AI coding tool ... > I noted that my own token usage comes to about $1,000/month against each of Anthropic and OpenAI What on this godforsaken earth are all you rich idiots doing???
View on HN · Topics
Token costs rising because data center build costs must be paid down.. is not the whole picture. It is actually possible for token costs to fall despite the spending frenzy. Naively you’d expect to always keep paying more - but growth in token usage is what changes the equation. Amortizing debt over an exponentially growing amount of spend across a growing customer base (not per customer) lets the debt be paid off & costs covered even as each individual’s spend stays steady or even goes down - but it only works if there’s growth beyond some threshold that makes the whole thing hang together. No one on the outside knows how much growth that is, and everyone chases maximum growth. Jevons Paradox ends up being your friend as well as the friend of the inference providers as well as the friend of the inference financiers. If it’s a strong enough effect, it has potential to cancel out all the circular financing too, and let everyone ride out the bursting of the bubble.
View on HN · Topics
They are also beholden to enterprise pricing and can't use the subsidized consumer max plans.
View on HN · Topics
The subscriptions are not available to enterprise users. Enterprise users must pay per-token. A $200 subscription gives you roughly the equivalent of $1500 in per-token billing.
View on HN · Topics
You are paying account pricing. Uber is paying API pricing. You're $100/m plan is likely equivalent to thousands of dollars of API pricing. You are being subsidized by the companies using AI.
View on HN · Topics
I have strong conviction that companies will now choose tech stack/programming languages based on 'tokenomics'. I am vibe coding using Clojure, a language I can read but cannot write and I never hit the usage limits even when using the latest model on Claude. I have similar experience with F#, which is a bit more verbose than clojure but absolutely beats every OOP language, Python, Typescript etc. The reason, I use F# & Clojure is they hit JVM and CLR, two popular enterprise stacks. In my not so humble opinion Lisp(Clojure) still remains the language of AI.
View on HN · Topics
if you have more than x seats, you have to use Enterprise pricing as far as I know which is pay as you go with a pool.