Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
nylonstrung
10 months ago
|
parent
|
context
|
favorite
| on:
Kimi K2 Thinking, a SOTA open-source trillion-para...
Subjectively I find Kimi is far "smarter" than the benchmarks imply, maybe because they game then less than US labs
vessenes
10 months ago
|
next
[–]
I like Kimi too, but they definitely have some benchmark contamination: the blog post shows a substantial comparative drop in swebench verified vs open tests. I throw no shade - releasing these open weights is a service to humanity; really amazing.
rubymamis
10 months ago
|
prev
[–]
My impression as well!
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: