I love LLMs too, but I am concerned about their cost. They are all still very subsidised. Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices?
I think the opposite: I think the frontier labs have good margins on their inference unit costs.
We can already see what it costs to run near frontier-size models. There are independent business pivoting to serving these models at reasonable prices and they're competing on OpenRouter for costs much lower than frontier labs.
> Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices?
I would bet good money on prices going down significantly, not up.
If we get to the point where you can run an Opus 4.8 model on your local computer, it's going to be even cheaper for a datacenter to serve it on their hardware. That means prices crash, not that they're going to rise.
> 2. Unit costs are irrelevant when the labs don't price per unit, and instead charge, for instance, $200 / month for $10k worth of tokens.
Cost to generate all of the tokens divided by revenue generated by selling those tokens is what matters.
The subscription plans confuse a lot of people because that's what they see. They're not seeing the gigantic API bills from all of the tokens going into enterprise use cases.
The subscription plans are a small part of their income. Most users aren't maxing out 100% of their plan usage every week. I wouldn't be surprised if their average plan user was using less than 50% of their monthly quota each month.
Plans like that can produce a net increase in profit if they get consumers interested in the brand and pitching it at work. Giving them some extra token headroom on their $20/month or $100/month home plan is money well spent if it gets all of a company's developers advocating for enterprise plans with budgets exceeding $1000 per person.
enterprises are not dumb, they look at the cost of their ai investment and reevaluate it every quarter. Uber recently capped their AI spending per employee, and then there's this article a couple of days ago: https://finance.yahoo.com/technology/ai/articles/ceos-being-...
> Using a full Claude Max 20x plan to 100% of weekly usage
I doubt many of their customers are on the 20X plan. Of those, I doubt many of them are using 100% of their weekly usage regularly.
Comparing the 100% maximum usage scenario of their most discounted plan against the API cost has been a trap in this conversation since it came out. I bet if we saw their financials it would be a tiny sliver in a pie chart somewhere.
True, it'd be a whole other situation if the tokens limits were cumulative. I guess it would all come down to whether their Claude Code subscription plans are turning in a profit or not.
At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.
Yes, it is a 10x markup on the API prices. Depending on whether you factor in cooling costs, data center staff, etc. Or GPU costs and the electricity the GPUs are using only.
Either way, inference is very much where the money is made, training is where the money is lost.
Token prices are going down. Competition is global. A company could choose to keep their API prices high, but if another company comes in at 1/10th the price for 95% of the performance then they won't have many customers.
You can maybe run a local Sonnet-4.5-ish-level model (sort of) for less than the price of a new car, even at current massively inflated prices for fast RAM. This is probably not what you were looking for. But it's there. You could share one server between multiple developers. Maybe make a little AI co-op or something, with a pair of RTX Pro 6000 cards?
Also, DeepSeek V4 Pro is cheap via any commodity API, and DeepSeek V4 Flash is essentially free at API prices like $0.09/M, $0.18/M out. This is generally not subsidized.
For a more practical local setup, Qwen3.6 27B on a used Nvidia 3090 (US$1300) or two is surprisingly nice. It needs clear instructions and you can't use it for hands-off vibecoding, but it's actually quite reasonable for hands-on programmers.
I’ve got a pair of those cards and DS V4F is incredibly good. I’m happy I did what I did because I like this stuff but if you just want stuff then you are absolutely better off not spending $20k on two of these cards and using the API. This guy is absolutely correct.
"Of course, it will also probably cost somewhere around $50k..."
Whats to stop people remotely accesssing this? People already do this when working remotely in finance - they connect to a virtual environment that does their work in spreadsheets lmao. nobody cares about the lag, managers certainly dont care about sub-ordinates complaining about it - the same way nobody will care about a slight loss of quality if the economics make sense. frontier labs are screwed really.
Currently, because of the subsidies from the frontier models, demand is mostly for higher intelligence.
If subsidies do end, demand for price efficiency per unit of intelligence will go way up.And because there's many players in the market, this demand should be met by at least some of them.
GLM-5.2 is runnable and downloadable today on a MacBook studio that costs a stupid amount of money. No one can take that away from you except by force though, if you want to set it up today.