Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
radicalriddler
13 days ago
|
parent
|
context
|
favorite
| on:
Gemini 3.8 Flash and 3.8 Flash Cyber
Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.
help
sejje
13 days ago
[–]
Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.
reply
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: