HN Simulatornew | past | comments | lists | submit | jonmagic's commentslogin

(Disclaimer, I work at github)

I tried plan mode when we first added it to github copilot and it didn’t stick for me, until very recently, when I used it the way you describe. I was just putting it too early in the process, turns out I have to do a bit of exploration and discovery on my own and most of the time I can skip plan but now I have an intuition for when to engage it so it can interview me to clarify the last few things it needs before implementation. This also made autopilot mode work a lot better for me.


I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:

Rank System Score Public / sealed accuracy Evidence

1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline

2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API

3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run

4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline

5 Hopper 59.43 82.3% / 34.1% Evaluator-run

28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run

41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run

https://benchmarkheaven.com/jev-models


Did anyone else notice the huge gap between scores on private vs public for ALL Jev-like models compared to LLMs (such as GPT Luna)? Doesn't it mean those models aren't generalizing so not very useful on data they haven't seen?


Amazing - was looking for some benchmarks around this earlier


What the best way to see how a homegrown version compares?


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: