Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: