Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you.

What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.



Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing.

When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left

And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome


Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?


>So why would I want to switch to even worse model?

There would be no reason to if you are in the privileged position where cost isn't an issue.

For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.


Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.


This reply must have cost dozens of dollars.


This must be the new linked in strat. What I learned about using the best model after talking to a homeless person.


Imagine how much it must have cost to have him read your comment! I hope he sends you an invoice.


That’s the thing, the difference is so small that you won’t be wasting an hour per day with a model that’s 95% as good. In fact, you’d notice zero difference most of the days and when you do, maybe it’s an extra 30 minutes.

And the price difference is far greater than $2k/month once the API cost is no longer subsidized.

Would your employer be paying an extra $20k/month to Anthropic if it can save you 2 hours a month?


Imagine if we had this attitude for electricity. No point building a grid, just give it to a few factories that need all day lighting.


Pareto optimal dominant vs a human for the same task, not an unreasonable framing but that assumes that it can actually do the task, which the op was arguing it couldn’t at all. Which, I suppose you could model as the utility of task completion % as being non linear. I have heard many people argue that the nature of work is messy and complicated and many things they do could not easily be emulated or automated. I do wonder how many of those activities are actually something that are connected to a companies ability to generate revenue or are just the messy interactions between people.


That's why every benchmark should show the Pareto frontier against cost and latency.


> So why would I want to switch to even worse model? Even if it's just slightly worse?

Self-hosting is the biggest reason.


> They require constant supervision

I think that may be part of it.

LLMs can be autonomous to an extent. All of them need steering - which is why I feel they are more a superpower the more I am an expert on the subject matter.

The more you want it to be autonomous, than yeah, you may benefit from using the very best the industry has to offer, however slightly better it is.

But if you are always in the loop anyway, you may want to try DeepSeek. You will get similar results for a fraction of the price.


you seem to be mistaken, Opus and Sol are the worse models.

Stop trying to treat these things like a replacement for yourself and instead approach them as a limitless number of offshore developers. If you are willing to be endearingly literal they will make you happy


If they already require your constant supervision the reason is money.


I think americans assume when they see a chinese or asian person working at an american business that they "escaped" china as opposed to just being rich enough to go to school abroad. and has little to no bearing on the amount of intelligent going around.


If Americans see an Asian person working at a US business, they will assume that person is an American. They may even ask what state you are from. It's honestly one of the nice things about the place, you can belong even if you are not from there.


I will just assume they are American, unless they are an uber driver. Some cities in China very much are losers of the new tech plan. Otherwise why would they fly to mexico and cross the border


They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.


I don't think it's all Americans however there is a portion of them who are not able to grasp the world outside of their borders.

I think it's mainly due to poor education many receive and a very controlled media that suppresses information.

It's shocking considering how much money they spend on education compared to other nations.


> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.

Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians.

We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war.

Or we could talk about the number of nice palestinians killed since the beginning of the war in Gaza. Or we could go a bit further and talk about the joy and celebration in Gaza after their heroes brought back 200 hostages after having slaughtered 1200 civilians.

You may be living in a place that you think shields you from those but I know the ideologies behind these acts.

The fallacy of gray is just that: it's not true that there's always a nice middle ground and that there's no evil ideology out there.

Something something about the price of liberty being eternal vigilance. For there are people abusing your blind trust.


QED


+1 most people are too afraid to try something new. They've been roughly on par with their frontier offerings for 6 months if not more.


And I had the opposite experience. It's a really interesting phenomenon that I can't really explain. My co-founder swears by Deepseek and yet just the other day we were conversing and he was telling me about some of the issues with the way the AI was behaving and trying to show off the cool workarounds he came up with to limit it. I was like, "Interesting, yeah, I've literally never had that problem."

I suspect that the models are genuinely close and that certain experiences get felt across providers but are inconsistent enough to convince people one is superior to the other. I for one have tried Deepseek on and off since my co-founder is fond of it and I've stopped trying now because I never have a good experience.


{Black box} + {sunk financial/time cost} + {ambiguous rankings} + {marketing} = {irrational tribal loyalty}

On reddit et al., people talk about LLM brands like their sports teams.

I think the first-party ecosystem moats they're all trying to build are exacerbating this tendency, as now people have a lot of learning time sunk in a company-specific option.


That’s not surprising. I have been using Deepseek and it consistently produces excellent output given the right context howevwr. It depends on the task as it does have blindspots.


That may be true.

I switched to DeepSeek entirely once I decided to put 10 bucks on it and I realized that it could do whatever I was throwing at Claude or ChatGPT prior to that.

I recommended it to one of my friends, and he was surprised DeepSeek could solve task that Claude got stuck at. I was surprised at it too.

I know others that tried and were less impressed too.


Yeah DeepSeek has always been terrible until the new release version of 4.


Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.


They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.


I think it's more likely to be the effect of synchronization of launches, and the fact that models that do not challenge SOTA in some way do not get launched (think Gemini Pro delays), launched quietly or do not get any attention.


I gathered that the most recent advances haven’t been in capabilities of the model but more the way that it’s able to be employed (most recently agents).


If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.


- load bearing -


seams, boundaries, envelopes, etc


> LLM providers can distill all of human output into their models for 'free'

Not sure what part of being charged guilty and paying a fine you see as "free".


That’s definitely my impression of deepseek 0731 after a fair bit of use via ds4, it sounds like Claude.


What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.


I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.


Why isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?


Because that doesn't satisfy the narrative of American exceptionalism. It's easier to point at something and say it was stolen or copied than it is to compete, especially with the political climate in the US.

This isn't an anti-American sentiment. It is an anti-corporate/regulatory capture/embrace and extinguish sentiment (which probably reads the same to many people these days).


Because the companies did it underhandedly without prior consent? Imagine you walked into a hardware store to grab some lumber, didn't pay for it, and the store had to call the cops to swing around your house? That's hardly the typical shopping experience, now is it?


I think that’s my point: none of those things happened, so on what basis can we claim that it’s like they did. Who has to agree with what you’re doing in order for it not to be considered underhanded?


If the legal system declares the first thief’s theft not theft then all bets are off.


> If the legal system declares the first thief’s theft not theft

But they didn't find it. The Big LLM provider accepted guilt and paid a fine.

You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.


As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.


Copyright law only considers illegal ownership of a work, so the crime - or tort - was making/acquiring copies without permission or payment.

Training from copies has been ruled fair use because it's "transformative" and not simply "derivative."

This is obviously debatable, but that's where the debate is at the moment.


So basically because they just browsed and used the information that was mostly public on the internet and they didnt copy it, they just learned from it and thats fine. Which makes sense. None of the llms let u copy someones work exactly anyways... makes total sense honestly. So in this case what happens to distilling? Is that also learning or ur trying to get to their actual weights by kind of reverse engineering it? Where would the argument fall there?


> but that's where the debate is at the moment.

Because of the rulings of a couple of judges. Is that actually what the majority of people think?

> Copyright law only considers illegal ownership of a work

That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing.

Similarly, copyright has something to say if I read a legal copy of harry potter and then create a new work in that world.


> Because of the rulings of a couple of judges. Is that actually what the majority of people think?

There's a good reason for the law not to be based on what the majority thinks.


If you’re gonna equate copyright law being tied to a majority sense of morality, with some kind of mob rule, ya lost me.

Sure i don’t think the majority get to dictate things like who has rights or who the law applies to. That doesn’t apply here tho.


> They didn’t accept guilt for incorporating all of human output into their model without consent.

Because that use case is actually permitted by law.


I mean... that's one interpretation of the law, sure.

The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage.

So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.


> So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.

It's not an either-or.


Kinda is.


> But they didn't find it. The Big LLM provider accepted guilt and paid a fine.

That's not how it works. You have to give it back.

Otherwise, the distiller can just pay a fine (no larger than the original did) and be okay then, right ?


True. It's the courts that failed humanity. Or perhaps the shits that invented copyright to start


Is it theft if another thief steal's the first thief's theft?


question's phrasing made your judgement obvious


I seriously wasnt judging i maybe shouldve asked llm to frame it better cause i knew it might sound that way, thats why i added: im not judging, merely asking...


I don't think there's a ceiling to LLM capability. I do think that many software dev tasks are just far below that ceiling, and the gains to most dev work won't be that large from now on.

Where the new generation of LLMs (Fable, Sol) shines is tasks that are much harder than typical soft eng, yet that still have a verifiable answer, think mathematical proofs or exploits. I think there's still a good amount of low-hanging fruit in those (and similar) areas.

The next frontier after that is tasks that don't have automatically-verifiable answers, and may not even have correct and incorrect ones in the strictest sense of the word.

Reasonable lawyers might disagree on the question of "which trial strategy do I use given the following set of facts." There are answers that are clearly wrong, but being able to choose between many plausibly-correct ones requires many years of lawyering and seeing many trials play out. I do suspect that most lawyers are far below the ceiling that a hypothetical immortal lawyer that has practiced for an infinite amount of time would have achieved.


Or are humans more of a bottleneck than before, because to improve on the most complex problems that demonstrate intelligence you need some way to verify that they are correct. If it's hard for humans to even know if something is correct, wouldn't that slow everything down and simply put limits on the scaling speed of models based on human verification?

So instead of relying heavily on human bottlenecks, you focus on agentic task verification since that's the low hanging fruit and verifiable at scale?


Very interesting point you make! Before LLMs I had a theory that we cannot make something more intelligent/complex than us.

LLMs are certainly more knowledgable, but maybe not more intelligent, arguably. It's possible we're approacing a ceiling indeed.

Model capability might be on an asymptote appraching but never quite reaching parity with human intelligence.


There are many things that can have some decent level of automatic verification and those were some of the first for LLMs to excel at, like math and programming. Now, physics involves math, but verification would still often require some form of measurement to make sure that the math relates to the real world meaningfully.

Many extended kinds of verification can be done by LLMs, but they need to be able to follow instructions reliably and agentic task orchestration may be critical to that verification process.

There is no doubt they will surpass us as there is a lot of easy to reason about information that they can verify as incrementally proven by other knowledge. The trick is knowing what can be proven with existing knowledge and what needs human evaluation.


Don't say that too loud, you may burst the bubble prematurely.


the ceiling is to eval's quality


Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.


I find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me.

One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.


How do you do the review cycles? Is this some automated feature of your harness? Do you have a generic prompt for this or ask 27B yourself?


Usually something like dsv4 running them. Depends in the horizon. It’s kind of an overnight thing for the project. Use matt pocock skills or similar. Build good spec. Iterate on it with a big model or your brain. Break it down into pieces and features like you would for a human. Tell them to work on tickets. Have other agent review. Repeat. Stuff gets built. You can keep a smarter agent in a loop (think bash) to review and have a simple decision tree: implement, review, mark done, pick next ticket. When no tickets stop. If error or pathology detected, touch a stop file. That is what I run overnight. A3B likes a simple harness as does 27B, mostly unmodified Pi with a proxy to fix model bugs. It takes some investment. It isn’t batteries included. These small models are not very smart. They need a pretty narrow task domain.


Same! The only reason I'm not using it more is because it's summertime. I'm not in any hurry.

Setting the memory to "fast timings" is good for 8-12% more tokens/second if you haven't tried yet. I miss the slightly older days of AMD when powerplay tables were unlocked and we could configure the timings and voltages manually, there's another 30% being left on the table ez


What do you mean by 'Setting the memory to "fast timings"'? The only runtime I can get working for my GPUs is llama.cpp, which I haven't seen anything like that in its argument set. My perusal of the options for vllm and sglang didn't suggest anything similar either before failing miserably.


I think they are referring to the AMD drivers on windows, under the overclocking section you can enable fast memory timings. Not sure if this sort of thing is exposed on linux.


You can edit sys files or use amdmemorytweak on linux.


Works great with room to spare on my lenovo pgx too


I'm still skeptical of the smaller models after the talent exodus a few months ago.


At this point I feel like the only factor differentiating SOTA models now is who they’re propagandizing you on behalf of (not considering agentic tooling/state management, etc).


Who is going to break it to the Americans that China is more than a slight favorite to win an existential battle over which country is better at math?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: