Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.

The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.

Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...



You gotta admit the timing looks very suspicious.


Luna pricing was just cut by 80% https://www.eesel.ai/blog/gpt-5-6-pricing and as the blog post states is a more accurate judge than the previous methodology.


> You gotta admit the timing looks very suspicious.

Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?

That's indeed a bit fishy.


They could just have avoided all of this by not publishing the benchmark until the new methodology update.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: