Google just put the latest AI models through a brutal coding test — here's how they did
Google's Android Bench 2.0 tests how well leading AI models handle complex Android development tasks that can take human engineers days.
What you need to know
- Google has launched Android Bench 2.0 to test AI models on complex Android development tasks that can take days.
- The new benchmark includes tasks like upgrading dependencies, adding major features, and building Android apps from scratch.
- GPT-6 Astra currently leads Google's new benchmark with a 28% pass rate, while Gemini 3.8 Flash scored just 8%.
Google has announced Android Bench 2.0, an updated version of its benchmark for evaluating how well large language models (LLMs) and AI agents handle complex Android development tasks.
Earlier this year, Google introduced the first version of Android Bench to measure how AI models perform on real-world Android development work. The company has now updated the benchmark with Android Bench 2.0, which is designed to evaluate models and agents against more complex tasks that better reflect actual software development.
One of the biggest additions is what Google calls long-horizon tasks (LHTs). These are significantly more complex development jobs that could take a human engineer several days or even a week to complete.
Google says the first version of Android Bench, along with many other early AI coding benchmarks, focused primarily on smaller, incremental changes. Android Bench 2.0 is designed to raise that bar with tasks such as upgrading dependencies, adding major new features, and even building Android apps from scratch.
With Android Bench 2.0, Google has changed how models are graded. Rather than relying entirely on a binary pass-or-fail system, Android Bench 2.0 uses "continuous scoring." The company says this provides a more "meaningful indication" of how well a model performed, even when it wasn't able to fully complete a task.
Google has already tested several of the latest AI models using the new benchmark, including Gemini 3.8 Flash, GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Claude Opus 5, among others. According to the results, GPT-6 Astra currently sits at the top of the benchmark with a 28% pass rate. Gemini 3.8 Flash, meanwhile, scored just 8%.
Google says testing models against the LHT dataset should give it a better understanding of their strengths and weaknesses, while also providing developers with more practical guidance about which models are better suited for different Android development tasks.
Get the latest news from Android Central, your trusted companion in the world of Android
The updated Android Bench 2.0 leaderboard is available now, and Google says it plans to continue expanding it with more models and results over time.
Sanuj is a tech writer who loves exploring smartphones, tablets, and wearables. He began his journey with a Nokia Lumia and later dived deep into Android and iPhone. He's been writing about tech since 2018, with bylines at Pocketnow, Android Police, Pocket-Lint, and MakeUseOf. When he's not testing gadgets, he's either sipping chai, watching football, or playing cricket.
You must confirm your public display name before commenting
Please logout and then login again, you will then be prompted to enter your display name.