Original caption
Gemini 3.8 Flash just edged ahead of Astra on DeepSWE. The comparison shared by Theo puts Flash at 73.8% against Astra’s 73.3%, but the win came with around 143,000 output tokens per task at high reasoning effort. That’s a lot of tokens for half a percentage point. One close benchmark doesn’t make Flash better at everything, but Google’s Flash model finishing above Astra is still a result worth to talk about #Gemini #Google #Astra #DeepSWE #AI