If you take away the cynical edge, though, if this actually happens we're talking about an age of mom-and-pop businesses in industries that have not traditionally had them. It can't/won't be all software, but mom and pop stores were only about serving their neighborhood to begin with.
When we've used software to disrupt other industries to deliver more with fewer people we called that a good thing. And it is, because consumers tend to win. Games are a hit-driven business, so they're going to feel the squeeze faster than, say, SaaS, but it's coming for all of us. Software is now being disrupted. Reap the whirlwind.
I worked at a (very) well known media company in Hollywood during the Netflix disruption. This really doesn't sound all that different from a Hollywood exec mocking internet media in the early 2000s. People are trying to understand it and leverage it. Quality is bad. Then one day you wake up and it's taken over, because it kept improving. Setting aside the debate on AI use in the creative process, even if models never get better there are years of optimizations to be done on the tools/business side.
As a gamer, I'm actually excited for it. I want more diversity in games beyond what the current indie scene provides. If we can give smaller teams the horsepower to make what was considered AAA 5-10 years ago I will be quite happy as a consumer.
Netflix disruption might be an apt metaphor, since it did severely damage the industry and reduce the quality of its output. Those Hollywood execs were right, even if they were ultimately doomed. It's another step in the Sheinification of everything. If it happens to software too, it won't be a victory for anyone except the owners of software companies.
I think people who long for pre-Netflix era are looking at the past with rose colored glasses. Distribution was terrible compared to now without a lot of choice. Anyone who thinks that time was better and the quality was higher don't actually remember it. They just remember the best of the era and TV shows they liked when they were younger.
Distribution got better for a time and then became the same old bullshit. When Netflix was the only dominant streaming service it was great. Now ever other person has their own custom streaming service and content is distributed across them. Back to the high seas for me. I don't mind paying. I mind the inconvenience. There was a brief window when Netflix was improving TV quality. But now it's almost all non-native US made originals that are cheap as fuck. They aren't all bad. My wife has a weakness for Spanish and Korean dramas for some reason. But it's all noise to me.
I honestly couldn't disagree more. With the advent of online services, the quality of both tv series and movies have massively increased. Budgets are larger the appetite is larger and grass roots non-Hollywood creators have a serious shot at disrupting the market.
I am personally no longer interested in the same shitty nostalgia rehash of previously garbage marvelslop and the same 15-20 actors on my screen and that has started to change because I have way more choice.
Sure. Who doesn't love 10 episode seasons every three years that are regularly and unceremoniously cancelled whenever the algorithm says so. I know folks in the industry absolutely love the total lack of job security...
I think that was absolutely true for a time. But now outside of Apple TV who is doing high budget quality TV? Certainly not Netflix. Best they can offer is a Korean or Spanish soap opera.
It's a huge mistake to think that indie teams would make better games if only they could balloon their projects to AAA-size.
Some of the most meaningful and engrossing gaming experiences of my life have come from 1-3 person teams, especially over the past few years. Meanwhile, the vast majority of AAA titles just feel like time-filling entertainment chow to me: fun-shaped but ultimately pointless. I would strongly argue that artistry and auterism is nearly impossible to preserve beyond a certain game size; the experience can't be focused by definition.
In any case, many of the most talented indie game devs are rejecting this technology outright.
Making things previously impossible possible? That's the business I joined.
Destroying people's livelihoods without care for the consequences while pursuing monopolies, capturing the value as rent seekers, and ultimately enshittifying entire industries?
Yeah, no, I know I didn't sign up for that, thanks.
And if you're thinking "wait that's not AI". Yes. That's precisely what these companies are trying to do.
That's honorable, but software has been used in no small way to disrupt traditional business. Uber. Doordash. Netflix. Tesla. Airbnb. Zillow. And yes, AI is coming for another round.
It's a part of the culture. You can say that isn't what you signed up for, but I find it hard to believe that you don't know it's part of the culture you're participating in, which is the "we" here.
Maybe first define "participating" and then perhaps I'll respond.
Because if this a case of "well you use a computer so you're part of the system, man" then I probably won't waste my time. Cynical fatalism usually involves moving goalposts IME.
I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets.
I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited.
I urge Anthropic to get better at this aspect of their business so I can come back to it.
IMO, if you depend on AI to make a living, I'd invest in hardware for local inference, and learn on how to effectively make a living using AI inference you control, on hardware you control. Sure, economically speaking it's way cheaper to use one of these heavily subsidised services (for now), and their models are faster and more capable, but if your livelihood depends on AI inference, and you are renting AI inference, you are a being a serf of the tokenlord. And your livelihood depends on the whims of the tokenlord. They can increase rent prices, they can decide you can no longer do whatever you are doing, and you have no recourse, because you are dependant on them to make a living.
There are a lot of things in my toolchain pre-AI that I did not own and relied on to make a living. Mobile developers are in even worse shape, and iOS developers doubly so. The idea we were somehow less beholden before AI, I think, is silly.
None of us can wholly do our trades without support. Local inference is a fun idea, but you'll be out-competed by the serfs, as you call them.
You listing things that make developers dependent doesn't mean they weren't less dependent before. Now they have all those things, and more.
It is debatable whether local inference is a competitive disadvantage. One key advantage is consistent performance. No unexpected model downgrades or yanks, no silly safeguards imposed, and no quotas is a lot of advantage.
And by the way, it is not an either-or decision. You can use local inference as your daily driver while still leaning on frontier models when you get stuck.
It's okay to be a prepper, but you don't have to be. Assuming access to internet, water, electricity and increasingly AI is a entirely fine way to live. It might not be anthropic which you will want to use but current capability models will be abundant and access readily available.
I don't read that recommendation as prepping. If I were to start an earthworks business I likely would rent heavy equipment at first, but once I get the business going I'd likely begin to buy my own machines.
That's not prepping in the sense of having a bunker full of canned beans. Its taking control of a key piece of my business, and likely saving money in the long run.
I don't think that analogy holds up. Owning heavy equipment requires lots of capital, and only makes sense if you keep utilization high. Even large companies will rent or lease equipment if it's something they use infrequently.
A large company might own their equipment, but an individual operator probably won't. So it might make sense for some large software companies to own their LLM hardware, but it probably won't make economic sense for individuals.
Of course the economics are different in different industries. Trucking owner operators account for ~15% of truckers, but buying a rig is six figures against 5-6 figure income. Buying a mac mini is 4 figures against a 6 figure income, so maybe lots of people will do it even if it's not economically optimal.
This may be a difference in location or urban vs rural? I live in a more rural area and many people I've hired over the years own their equipment (well, if having a loan on it counts). That goes for tractors obviously, but similarly for wheel loaders, excavators, etc that they use for hired work.
> once I get the business going I'd likely begin to buy my own machines
only when it's cheaper than to continue renting it in your calculations
same goes for renting vs buying a house, or anything.
the breakeven point here would come when they stop subsidizing the subscriptions, or when local ai gets to run on basically everything, making it ~free (sans electricity)
Except there's a huge gulf of self-hosting and using API hosts - no way you can reach the economics of a shared host. Privacy is a problem but you can chose who you host with and where it's hosted (which jurisdiction).
When privacy/compliance really starts to matter it's up to the client/business to provide you with tooling - you're not running that on your own hardware anyway.
So the local AI for individuals is just a hobby/gimmick at this point not a rational decision. Self-hosting for business is a different story.
I'm not sure. The problem with the cloud llm's is they are complete black boxes that change frequently and randomly day by day.
If you run Qwen 3.8 on your own hardware, every single day, it's the exact same model running in the exact same way.
Yes, it's nowhere near as "smart" as the cloud based models. But it's consistent.
So the workflows and "ways of working" you create will work mostly similar day to day.
With Claude/OpenAI you frequently find days where the models are useless, and days when they are out of this world.
So I guess the choice comes down to:
1. Randomly the smartest thing on the planet with unpredictable rate limits that is mostly amazing, but frequently messes with your workflows
2. A really good local coding model that is consistent every day with no rate limits
I'm not sure. My gut feeling is maybe the right answer is a mix of both.
Gambling on the biggest models, hoping they are working smart that day, when planning or doing very complex work. Then doing most of the tasks/daily work using local models??
You can run any open model on a shared API host via OpenRouter and pin to which host you want to go for the quant/privacy/etc. mix you care about. You can pay them directly if you don't want the OpenRouter overhead - but the convenience of switching, having one invoice, etc. is worth it IMO
It's not closed hosted models vs open local models, it's hosted open models vs local open models where the math doesn't work for local LLMs.
The only local inference use-case I can think of is porn generation (because most providers don't want to deal with it) and illegal shit like hacking to minimize the tracing.
And if you're super paranoid - but honestly giving sensitive info to LLMs in any scenario is a gamble.
If you game and can use your GPU I guess then it works as well but models that fit into a gaming GPU suck too much to bother IMO.
This is pretty terrible advice when there are dozens of AI inference providers out there serving great models with significantly more cost effectiveness than you'd get from buying your own hardware.
I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.
If you just say that you run out of tokens, it does not mean anything about the token quotas themselves being reasonable or not. That depends on how much you use it.
For instance, if you had 10s of agents running all the time, it is not that unexpected that you run out of tokens quickly.
I'm curious what your methodology is that results in that? Are you running multiple teams of agents all adversarially reviewing each others code? Lots of different projects in parallel?
I've only rarely maxed things out and then it's t through doing extreme things.
Agreed. The area I think will become more prevalent in the future for organizations are cost per intelligence -- effectively efficiency. An unoptimized model that costs 90x more than another that is only 10-15% less intelligent is something I would say is not a good deal.
Invest a bit of your time into optimising usage cost. Anthropic has first class docs, actually read it or ask llm to read them all for you and summarise most important points / ask to to reflect it on your .md files. Maybe silly thing like dropping your default thinking effort by one level or adding (sub)agent pinned to other model is all there it to completely fix it or maybe you have instructions that encourage big dumps in CLAUDE.md/AGENTS.md that needs splitting so progressive disclosure works correctly? Naively sending everything to the most expensive model on high thinking effort is anti pattern and will drain quota quickly.
My personal guess is that it's one of those. With effective context engineering it's hard to use all 20x quota, the limit becomes your own attention and time really.
You may argue that you're doing multiple, parallel extreme effort tasks – which may be true but then again, there will be results to actually look at sooner or later and that takes time.
I am on the Claude Max 20x plan, and this still happens when using Fable 5/Opus 5. I would run out of weekly quota in 2 days, whereas Opus 4.8 would last the entire week, and sit at about 80-90% at the end.
Same here. Over the last two weeks I switched back to Opus 4.8 and turns out that still seems to last the week like it used to. These new models must be eating tokens.
I am a huge enthusiast of running local models, but when multiple quality USA vendors provide models like GLM 5.3-flash, I run locally just for the fun of it.
For the purposes of comparing to Fable 5.1, I would mention GLM 5.3 that is about 1/12 the cost.
I'm with you, for what I usually do most models are already more than enough.
What I'm really keen on is better auto-reasoning so I don't have to constantly have the constant inner debate on which reasoning effort to pick for each task.
I seriously hate the none-low-medium-high-xhigh-max-ultra etc that we have now, with companies frequently recommending different ones on each new model release, etc.
It's apparently called Adaptive Test-Time Compute or Dynamic Test-Time Compute and companies are apparently working on it (according to some LLM :shrug:)
Adaptive reasoning is known to be an extremely hard problem to solve, though. It requires you to predict whether a certain LLM, with a certain effort level, with a certain prompt, will give you the right answer.
All of the subscription AI platforms are trimming down quotas across the board to push users into higher tiers. Whatever they can do. Local inference needs to meet pricing sooner
overthinks, been slow lately through the official api (slower than glm 5.3 somehow), and tries to run every conceivable e2e test once it does literally anything.
like yesterday it ran for like an hour to build a fairly basic frontend...
I have noticed fads in rhetoric as I've gotten older and lived through many of these moments. The rhetorical fads are mostly about structuring an argument where to disagree makes you appear beyond the pale.
What fascinating about these is that they eventually turn on themselves, because rhetoric has no logical foundation and can be twisted to mean anything.
There have been so, so many attempts to decentralize in my lifetime. I used to be a huge believer and worked on these problems, but have to admit I'm somewhat exhausted. The technology is not the problem. The problem is always quality, effort, and cost.
If decentralization is ever going to win, it needs to be turnkey, explained without showing a network diagram, be basically free to run on your laptop, and actually have the content people want to see and not just be a bunch of people who speak lojban. [1]
While I'm not sure how much of a legal leg X has to stand on, I understand why they'd rather Nitter not exist. Twitter tried (valiantly, imo) to stay open. What ultimately began the API lockdown was the need to stop bleeding financially. Ads were inevitable, and there being no reliable way to do that via API access.
Twitter solved those three problems, and no one spends nearly as much time talking or caring about the technology used to do it than people who talk about decentralization. Myself included in my younger days. Now they're just defending their moat.
Mechanical Turk had a good run, but not surprised it's shutting down. I'm sure the platform was getting flooded with people doing task arbitrage and using lots of AI anyway.
I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.
Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.
Giving out control of industrial machinery that interacts in the human environment without the physical interlocks (i.e. humanoid robots in a house) to random internet people seems like a problem.
I can imagine a carefully orchestrated plot to assassinate someone by having an embedded agent in the task delegation pool command the laundry bot to punch the target's head off their shoulders.
Giving out control of such things to LLMs is already complete madness, so once the first pleasure bot powered by Grok has dismembered a few thousand users, they'll get sophisticated safety mechanisms.
Though like as not you're still going to be right, after all, Stuxnet happened.
Why does a laundry folding bot have to have the physical strength to be able to kill somebody?
The principle of least privilege has been a hard learned lesson in cybersecurity. Why do we again need to first go through disasters to re-learn it in the physical world?
If a general-purpose humanoid robot has the strength to perform human actions at human speed, it will have that much torque in the motors. If it didn't, it can only move very slowly and/or jerkily. You need fast, high-torque feedback to stabilise the physical control system, especially if the robot is supposed to be able to lift and manipulate objects.
The idea here is Optimus-type robots that can do everything like a human and therefore don't need special infrastructure, as opposed to dedicated, immobile low-torque, low-velocity robots specifically for, say, laundry.
You can layer a safety system on top of the raw motor control but it's very complicated to define what is and isn't safe, especially when you consider what the robot is holding (and what it THINKS it is holding), where the robot is (or thinks it is) at the time, what is around it (what it thinks is about it) and so on.
Which isn't a robot-specific problem. Humans spend years learning how to subconsciously safely interact with their environment and even then fuck it up sometimes. What is robot-specific is being made of metal and having the "safe hand velocity within 10cm of a human head" function bypass be one bug or OTA update away.
giving out unrestricted control that is.
in the example of Waymo, a human can control some aspects of the car manually, but it can never override low level obstacle detection or say open the trunk/door when the car is moving.
This is what I meant by full stack AI companies. I don't think you could get humans into the loop fast enough if they didn't have some idea of the type of task involved. I don't want people to be asked to fold a tshirt one moment and do a difficult traffic merge the next.
There is training systems and validation of skills in mturk iirc: for tshirt folding, you'd be given fake setups to be able to get used to controlling the robot, if you can't do it, you won't ever get assignments to do it. For traffic overrides, you'd be tested on having correct knowledge, and once again given supervised tasks to show you can actually be trusted (and there would be safety systems, elevating tasks that can't be performed at your level to people who can, etc)
The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else
It doesn't really matter what you want though, only what CEOs want and that's low costs. I can see a combined shirt folding/traffic merging platform taking off.
but call center are still mostly single client even though it would be cheaper for any worker to be able to answer to any call. So clearly the expertise and context trade off is too big to be worthwhile
That's a remarkable idea. It could be heavily gamified, it could train models, and it might actually be mentally stimulating since you'd be facing different situations all the time.
Except, I'm a grown adult and I can't fold a t-shirt properly
There are already robotics models that can fold shirts and similar just fine. Progress in VLA models is good, I don't think this would form the foundation of a business. Humanoid robots are going to be another ChatGPT, it's going to seem to happen almost overnight because people aren't paying attention to the underlying research papers.
Most progress in data-driven robotics nowadays are done either in unicorn startups or corporate research labs - so you should follow the industry more than academia. The path to good robot performance isn't really in the models themselves - it's highly dependent on how much you can gather high-quality real-life data.
If their hero image video and the side by side at (5x) with a human at quote "1x" are "solved" I'm not impressed. I'm faster and more accurate and I am the worst folder in my house (kids included). The human looks like they are doing it slow mo to show children how to.
There are already a number of companies providing RLHF and SFT services for AI providers that does a lot of validation/prequalification of people that'd be well placed to take on tasks like that, but the big problem to solve would be latency if you don't have people contracted to carry out a task right now.
I used it before AI coding and it was getting rough. Lots of US interviews with proxy Chinese or Pakistani workers. Lots of bullshit “I’ve done that; I can do this” and instantly apparent that this was untrue. Just the outright lies… whew.
I haven’t touched it since AI coding.
It’s a bad contractor market now. IDK what I would do if I needed a contractor.
It is weird because the last time I've heard about MTurk was about developing countries being rather reliant on it for doing AI grunt work. If am not totally wrong this must mean that the data work has moved to other services.
I think the bigger issue is that a lot of the demand for labelling training data is now in highly-specialised fields (i.e. things like medical imaging), and mechanical turk's focus was on the generalist problems
To the best of my knowledge, that isn't true anymore, and nowadays you'd only get hired to do RLHF if you have particular skills beyond what can be achieved by just running other models against it.
I'm unsure. If it does, it will be work that is too expensive or inaccurate or regulatory for current AI methods. For example, you want a doctor to sign off on some AI output on a diagnosis.
However, I'm not sure a single platform will be how it emerges
A lot of what's been discussed in this thread is what we're tackling at Humwork (YC P26).
We're an MCP/API that connects AI agents to verified domain experts in real time (30s–3 min). Experts are vetted upfront by an AI interviewer that assesses and grades them, then get a mobile notification when a task matches their expertise and chat with the agent directly.
Soon we'll be verifying credentials for doctors, lawyers, CPAs etc on our platform for tasks where people are seeking credentialed experts to sign off and verify ai output.
We do our best to identify AI use and ban those experts - its not perfect just yet. Long term we are thinking of moving towards proctoring experts using their camera and screen capture. Hard to think of another reliable way.
You get matched with an expert in 30s to 3min, which then starts a back and forth with the AI agent and the expert which goes on for 10 to 60mins till the AI is happy.
This works well for many tasks, but for others a more async mechanism where the expert doesn't feel rushed might be better.
I have a lot of sympathy for the people working at GitHub trying to keep it going, but I do not have sympathy for GitHub. That's what they're paid for. I have things to get done, and it has no sympathy for me.
We came up with the idea of having a place where nerds could come together to hang out and learn from each other and we decided to call the concept "Hacker Dojo". I had a sign laser engraved in 2002 at a state fair to memorialize the idea but we didn't actually get to open a Hacker Dojo until 2009.
There's a conference room in it called Kaminsky.
Dan was not a very good roommate but man, his brain worked in weird and wonderful ways, and it was inspiring how he could just rabbit-hole on stuff other people didn't find interesting until he MADE it interesting and found things nobody else had.
I wasn't sure how specific to be in my comment, but I'm glad the people who knew him know who it is. Speaks a lot of the man. I'm also 100% ready to believe he was not a good roommate by the condition of his car ;)
Sorry about your friend. This kind of news hits very differently when you know someone who might have benefited from it. Hopefully continuous ketone monitoring becomes as ordinary and accessible as CGMs are becoming.
While this 1/2 typology is incomplete, diabetic ketoacidosis is a phenomenon generally relevant to type 1 family of conditions.
Normally (even with pure type 2!) glucose-ketone levels move in opposition: to put it simply ketone production is activated by low insulin, which is activated by low glucose. Ketone bodies are then consumed instead of glucose or excreted and generally homeostasis ensures that ketone levels do not approach anywhere close to ketoacidosis levels.
In cases where insulin production is broken (Type I diabetes, monogenic diabetes, mody) despite exogenous carbs raising blood glucose levels ketone body production is still active, however in most cases glucose is the preferred energy source, which leads to broken homeostasis of ketones, ketones accumulating and resulting diabetic ketoacidosis.
Type II family is characterized by high insulin. Some call it insufficient insulin, some call it insulin resistance, but the result is that insulin remains high and ketone production low. Regardless if the person is nominally healthy or has insulin resistance, the only way to trigger ketone production is to get insulin low for extended period, which is caused by lack of exogenous glucose. So either ketogenic diet or full-blown fasting.
Advanced and prolonged Type II diabetes (even if treated with insulin) can cause insulin production to drop along with insulin resistance, so the patient could develop Type I-like symptoms and risks, including DKA.
My alma mater, Cal Poly SLO, I think, already does a lot of this, and are working on getting better. They have the CIE, something I wish I had back in the day: https://cie.calpoly.edu/
I do wonder if entreprenuership would be higher if we gave students more gap years, because I find it rare that someone who has not been in the world knows they want to start their own business. Tech is an exception because it's more of a lifestyle and culture than other STEM-based business.
Similar in smell to political science degree --> career politician, I think pg nails it:
> The hard part of startups is product: knowing what to build, and being able to build it. And that kind of knowledge comes from studying computer science or mechanical engineering or molecular biology, not management or finance.
> Another thing that will tend to draw universities away from the optimal path is business schools, if they have them. Business schools were not designed to train founders. They were designed to train the managerial class of the large industrial companies that arose in the early 20th century; they're the West Points of industrial capitalism. That's why their official name is usually the School of Management. But while the skills they teach might be useful in running companies beyond a certain size, they're not the critical ingredient in founding them. And the skills that are are already taught by other departments. So to the extent business schools affect their parent university's strategy for preparing founders, it can only be by adding error.
> And that kind of knowledge comes from studying computer science or mechanical engineering or molecular biology, not management or finance.
But then he also states:
> Indeed, preparing students to start startups is closer to the ideal of liberal education than preparing them for almost any other kind of career.
This is pretty myopic and is obviously biased from someone embedded in the startup space, perhaps not surprising given how little value most startups provide. Not everything in life is a whiz-bang, fail fast, build-some-new-product-from-scratch startup. In fact, most things are not.
Pay attention to actual liberal arts universities. Their goal is on the long-term and not to satisfy some near-term need of startup founders making the next killer app or product. It can be argued that most startups build towards generating value only in the sense of monetary value, in that most startups' goal are purely about the exit. It really has little to do with building products. Most venture capitalists got their wealth from the dotcom bubble. Almost none of them know how to build actual products.