Again in March, we launched Android Bench—our LLM leaderboard for real-world Android improvement duties. Our objective was to supply transparency round mannequin capabilities in Android improvement and to encourage mannequin enhancements, to offer you extra useful AI choices to your on a regular basis workflow. Since then, we’ve enhanced the benchmark based mostly in your suggestions, together with evaluating open-weight fashions and including value and effectivity dimensions to the leaderboard.
However AI capabilities are ever-evolving, and measurement must comply with go well with. As a part of our July launch, we’ve adopted the Harbor framework, which incorporates an up to date model of the benchmarking agent used to guage fashions.
Together with this alteration to our analysis, on this July launch we’re including 8 new fashions (Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max) to the leaderboard. We’re additionally sharing alternatives for you, the Android developer group, to contribute to the benchmark.
Upgrading our methodology with the Harbor framework
After we designed Android Bench, we anchored our methodology on main business requirements obtainable on the time. We used mini-swe-agent v1, a general-purpose benchmarking agent, and tailored it to the nuances of Android improvement to supply a baseline measurement for the capabilities of fashions for widespread Android improvement duties.
To proceed offering you with state-of-the-art evaluations that precisely measure the newest mannequin capabilities on Android improvement, we’re standardizing our benchmark to the Harbor framework. Harbor defines requirements and integrations that make it straightforward for anybody to run the benchmark, consider their most well-liked set-up, or share outcomes – offering you with extra transparency and visibility.
This improve permits us to extra rigorously consider fashions and their capabilities, and we re-ran the benchmark on all fashions to determine an up to date baseline. This implies there’s a minor shift in scoring, however you’ll nonetheless be capable to view historic scores inside the archive on our web site.
We wish to guarantee Android Bench is useful for you, so we’ll constantly replace it as our evaluations and the business mature.
Increasing the leaderboard with 8 new fashions
As a part of our dedication to protecting the leaderboard contemporary, we’ve added Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max to the Android Bench leaderboard.
You will note that Claude Fable 5 is on the prime of the leaderboard with a rating of 84.5, adopted by GPT 5.5 with 80.2, with Claude Sonnet 5 in third with a rating of 76.2.
When simply evaluating Open-weight fashions, GLM 5.2 is on the prime with 72.2, adopted by Kimi K2.7 Code with a rating of 70.4.
You possibly can try mannequin efficiency and effectivity metrics on the up to date leaderboard to see how these new and former fashions navigate Android-specific challenges like Jetpack Compose migrations, wearable networking, and platform API updates.
Opening Android Bench to group contributions
From the start, we’ve valued an open and clear method, which is why we made our authentic methodology and take a look at harness publicly obtainable on GitHub. You’ve requested for a approach to supply suggestions on our dataset, so now we’re taking collaboration a step additional by providing you with, the Android developer group, an opportunity to form Android Bench.
Beginning right this moment, you’ll be able to contribute to Android Bench in two methods:
We will likely be reviewing the submitted duties and will likely be assessing in the event that they get added to the benchmark. We hope to construct a benchmark that actually displays the varied, day-to-day realities of the worldwide Android developer group.
Trying forward
With an increasing number of choices for agentic improvement, sustaining a cutting-edge benchmark ensures that the AI help you depend on retains getting smarter, extra useful, and more practical. Head over to our GitHub repository to take a look at the duties. We invite you to submit a process to our crew for overview, and you may try Harbor Hub to discover the dataset or submit evaluations.
As all the time, yow will discover the up to date leaderboard, or learn the methodology on our web site.


