Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Kinda useless to compare simply based on model without considering harness. Different agents handle the context etc completely differently. I would like to start seeing these model vs model comparisons across different harnesses.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: