Hacker Newsnew | past | comments | ask | show | jobs | submit | bitexploder's commentslogin

Simpler view for me: this is one of the most capital intense technologies to exist. Europe does not have enough capital to compete. China is building its own chips. It’s own everything. Silicon up. How do you compete with that. Only Google and maybe Aamazon is doing it domestically.

No, it is actually very good. Qwen Flash 3.8 Next is fine. But you need ~128GB of RAM to get it going and not a lot of people have that or can serve it very quickly. I have been running it on an old gaming system around 25 t/s to do overnight work and it is very strong, even at 3 bit quant.

Qwen Flash Next 3.8 … even at 3 bit quant it is very solid.

They are just pushing for favorable legal environment before the Anthropic IPO.

This is an entirely pointless exercise without transparency into how these "unreleased" models are trained, what their RL goals and biases are and related RL data, what their system prompts are, what their environments are and its restrictions, etc. What good is it for the industry to say:

"Our unreleased model attempted to create a bioweapon", but "trust me bro, we didn't tell it to do that. We didn't train the model on a dataset that specializes in creating and glorifying bioweapons. We'd never stand to gain from misleading people about model capabilities in any way shape or form." - Anthropic are renowned for doing exactly this, for starters.

So this ends up resulting in more safety theater. You can't have anything fruitful come of this without transparency. Stop trying to protect your moat if you truly care about safety and actionable outcomes, and provide real transparency, otherwise this is as good as saying nothing at all.

I'm not even saying they're intentionally trying to do this by the way, but this is not sufficient if the goal is balanced incentives and accountability.


It's not only that but another aim is for them to ban the competition, US and domestic, including open source models.

Basically, the billionaire elites and their employees are panicking because they fear their funds are at risk when the bubble bursts.


I'll believe it when they confirm a date.

https://en.wikipedia.org/wiki/Chinese_room I think about this once in a while. At some point if it does the thing almost perfectly is it still not doing the thing?

I suppose if you all you need is a good enough opponent for the average person out there, sure this is good enough.

I was more talking in reference to why the LLMs in the above linked paper were producing so many illegal moves, and it is because they are not hard constrained by the rules of the game. Of course, a loop can prompt until a valid move is produced and then rendered on a screen. But why do this? I suppose, who am I to say what should be done or not, but a specialized tool being better than a general one at its specific job isn't particularly surprising.


MLX?

Then use oMLX or MTPLX or Rapid MLX.

20 years later… I don’t get joy from hacking things. But the 20 year wisdom is a lot of the time it doesn’t matter if it is safe. Just know when it does matter and worry about that :)

Things are different now on a Mac. Many small improvements make Ollama genuinely decent for many models now.

The fact that those issues existed for so long while the entire time where not issues with other inference runtimes is more the point. And I think are indicative of future incompetence.

Less. Probably 2-3K if you build right. Qwen 3.8 27B on constrained tasks is Opus 4.6-ish to me, it just doesn’t know enough, but when task is laid out just gets it done.

Comes down to how much of the ambiguity we expect out of the model.


What would you build? Flash Next on 128GB works ok, 3.8 27B might technicially work on less but it's just too slow at least on current APU style chips. I'm curious about the non-NVIDIA 32GB discrete GPU options but I haven't tried yet.

Flash next only needs 64GB for core model inference. If you really wanted it.

Need 64GB vram, 900+ GB/s speed, and a lot of system ram (128GB). Seems feasible. Hmm.

Maybe older GPUs work.


128GB on a Ryzen 395 is enough but not by a lot. Around 90GB for weights plus 10GB to 20GB for KV cache and checkpoints. I don't know about 64GB of VRAM for less than around $2500 by itself.

Tesla V100? Would have to be a 3-ish bit quant in 64GB ram

OMP. Opinionated but completely configurable. Probably the beat to have a lot of batteries and let you uninstall what you don’t want. Sadly Anthropic forbids its use on their subscriptions.

Thanks. I'll kick the tires on OMP a little harder...

Isn't OMP sort of Claude in Pi's clothing? I tried it and it seemed like I was using Claude. But if tweaking is needed there then why not stick to Pi and add/strip as needed?

> Anthropic forbids its use on their subscriptions

This! How are these companies even allowed to do this while they anyway charge for either API access or limit usage in the generic plans. It's blatantly just "I don't want you to spend less per generic task!".


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: