Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

swe-rebench is a pretty good indicator. They take "new" tasks every month and test the models on those. For the open models it's a good indicator of task performance since the tasks are collected after the models are released. A bit tricky on evaluating API based models, but it's the best concept yet.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: