Show HN

Gemini models getting better

by @favurdev

Gemini models are getting better – we can prove it We ran seven Google models through the same task, in the same multi-agent harness. Improvements to reasoning and tool calls are the main reasons for the improvements. We saw less and less reasoning and some of the lowest invalid tool calls across the entire evaluation suite. Github link for the run output and analysis. More at evals.favur.dev

Discover more builders

Builderlust is an endless, joyful scroll of real projects people are shipping right now. Get the app to keep finding your next spark of inspiration.

📱 Coming soon to iOS & AndroidOpen in the app