(无标题)
This is the most scaling-pilled project I've ever been part of, and the team really cooked.
TL;DR: With RL and inference scaling, Gemini perfectly solved 5 out of 6 problems, reaching a gold medal in IMO '25, all within the time constraints of 4.5hr.