{"id":30091,"date":"2026-07-13T09:16:06","date_gmt":"2026-07-13T00:16:06","guid":{"rendered":"https:\/\/aireviewirush.com\/?p=30091"},"modified":"2026-07-13T09:16:06","modified_gmt":"2026-07-13T00:16:06","slug":"the-following-period-of-android-bench","status":"publish","type":"post","link":"https:\/\/aireviewirush.com\/?p=30091","title":{"rendered":"the following period of Android Bench"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<div class=\"separator\" style=\"clear: both; text-align: center;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEi49z_u9zPMjp-zyQ1yIpzLgDumtzUwZoprtIgPXv_kpF05e87KklDEguaKSJVhvV8dZJ7aVr98p-MG3FR4Sk37rcYTS91J3ADUQot-c-xnOuyIZ411VO4Hp43Yp7V_TwF6zO6RmAJpw51ZHPGbHfOwZxWgQ62SQeXblULcSc0RjMcZbLHGUZGgHzU6pEo\/s8583\/Bench%20July%20releas%20V01_Blog.png\" style=\"clear: left; float: left; margin-bottom: 1em; margin-right: 1em;\" target=\"_blank\" rel=\"noopener\"><img decoding=\"async\" border=\"0\" data-original-height=\"2601\" data-original-width=\"8583\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEi49z_u9zPMjp-zyQ1yIpzLgDumtzUwZoprtIgPXv_kpF05e87KklDEguaKSJVhvV8dZJ7aVr98p-MG3FR4Sk37rcYTS91J3ADUQot-c-xnOuyIZ411VO4Hp43Yp7V_TwF6zO6RmAJpw51ZHPGbHfOwZxWgQ62SQeXblULcSc0RjMcZbLHGUZGgHzU6pEo\/s1600\/Bench%20July%20releas%20V01_Blog.png\" alt=\"\"><\/a><\/div>\n<p><i><br \/><\/i><\/p>\n<p>Again in March, we launched <a href=\"http:\/\/d.android.com\/bench\" target=\"_blank\" rel=\"noopener\">Android Bench<\/a>\u2014our LLM leaderboard for real-world Android improvement duties. Our objective was to supply transparency round mannequin capabilities in Android improvement and to encourage mannequin enhancements, to offer you extra useful AI choices to your on a regular basis workflow. Since then, we&#8217;ve enhanced the benchmark based mostly in your suggestions, together with evaluating <a href=\"https:\/\/x.com\/AndroidDev\/status\/2064482677500080549\">open-weight fashions<\/a> and including value and effectivity dimensions to the leaderboard.<\/p>\n<p>However AI capabilities are ever-evolving, and measurement must comply with go well with. As a part of our July launch, we&#8217;ve adopted the <a href=\"https:\/\/www.harborframework.com\/\" target=\"_blank\" rel=\"noopener\">Harbor framework<\/a>, which incorporates an up to date model of the benchmarking agent used to guage fashions.<\/p>\n<p>Together with this alteration to our analysis, on this July launch we\u2019re including 8 new fashions (<b>Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max<\/b>) to the leaderboard. We\u2019re additionally sharing alternatives for you, the Android developer group, to contribute to the benchmark. <\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_53 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title \" >Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\" role=\"button\"><label for=\"item-6a6597b62a47f\" ><span class=\"\"><span style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input aria-label=\"Toggle\" aria-label=\"item-6a6597b62a47f\"  type=\"checkbox\" id=\"item-6a6597b62a47f\"><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/aireviewirush.com\/?p=30091\/#Upgrading_our_methodology_with_the_Harbor_framework\" title=\"Upgrading our methodology with the Harbor framework\">Upgrading our methodology with the Harbor framework<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/aireviewirush.com\/?p=30091\/#Increasing_the_leaderboard_with_8_new_fashions\" title=\"Increasing the leaderboard with 8 new fashions\">Increasing the leaderboard with 8 new fashions<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/aireviewirush.com\/?p=30091\/#Opening_Android_Bench_to_group_contributions\" title=\"Opening Android Bench to group contributions\">Opening Android Bench to group contributions<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/aireviewirush.com\/?p=30091\/#Trying_forward\" title=\"Trying forward\">Trying forward<\/a><\/li><\/ul><\/nav><\/div>\n<h2 style=\"margin-top: 10px;\"><span class=\"ez-toc-section\" id=\"Upgrading_our_methodology_with_the_Harbor_framework\"><\/span>Upgrading our methodology with the Harbor framework<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>After we designed Android Bench, we anchored our methodology on main business requirements obtainable on the time. We used mini-swe-agent v1, a general-purpose benchmarking agent, and tailored it to the nuances of Android improvement to supply a baseline measurement for the capabilities of fashions for widespread Android improvement duties.<\/p>\n<p>To proceed offering you with state-of-the-art evaluations that precisely measure the newest mannequin capabilities on Android improvement, we&#8217;re standardizing our benchmark to the <a href=\"https:\/\/www.harborframework.com\/\" target=\"_blank\" rel=\"noopener\">Harbor framework<\/a>. Harbor defines requirements and integrations that make it straightforward for anybody to run the benchmark, consider their most well-liked set-up, or share outcomes \u2013 offering you with extra transparency and visibility.<\/p>\n<p>This improve permits us to extra rigorously consider fashions and their capabilities, and we re-ran the benchmark on all fashions to determine an up to date baseline. This implies there&#8217;s a minor shift in scoring, however you&#8217;ll nonetheless be capable to view historic scores inside <a href=\"http:\/\/d.android.com\/bench\/archive\" target=\"_blank\" rel=\"noopener\">the archive<\/a> on our web site.<\/p>\n<p>We wish to guarantee Android Bench is useful for you, so we&#8217;ll constantly replace it as our evaluations and the business mature.<\/p>\n<h2 style=\"margin-top: 10px;\"><span class=\"ez-toc-section\" id=\"Increasing_the_leaderboard_with_8_new_fashions\"><\/span>Increasing the leaderboard with 8 new fashions<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>As a part of our dedication to protecting the leaderboard contemporary, we&#8217;ve added Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus and Qwen 3.7 Max to the Android Bench leaderboard.<\/p>\n<p>You will note that <b>Claude Fable 5<\/b> is on the prime of the leaderboard with a rating of 84.5, adopted by <b>GPT 5.5<\/b> with 80.2, with <b>Claude Sonnet 5<\/b> in third with a rating of 76.2.<\/p>\n<p>When simply evaluating Open-weight fashions,<b> GLM 5.2<\/b> is on the prime with 72.2, adopted by <b>Kimi K2.7 Code<\/b> with a rating of 70.4.<\/p>\n<p>You possibly can try mannequin efficiency and effectivity metrics on the up to date leaderboard to see how these new and former fashions navigate Android-specific challenges like Jetpack Compose migrations, wearable networking, and platform API updates.<\/p>\n<div class=\"separator\" style=\"clear: both; text-align: center;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEhQCbY3Td_I5gR8bC4uFSBTe4Sl-XuArNNdFU-27JP6-kwHycXt9AMpWfkLqjUIK37Zw18Tel6a7yOS9x0L_NabxBgYd9KIJKZ6dTLl6VxxJI4M7Zstqj12wvOFtF8LjnYrCIWnhCDdeGsgpQvFpFX8VOoSO0dFJcOW_gRc6eX7mXDq80sOwQAlQWNhlQg\/s1999\/image1.png\" style=\"clear: left; float: left; margin-bottom: 1em; margin-right: 1em;\" target=\"_blank\" rel=\"noopener\"><img decoding=\"async\" border=\"0\" data-original-height=\"890\" data-original-width=\"1999\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEhQCbY3Td_I5gR8bC4uFSBTe4Sl-XuArNNdFU-27JP6-kwHycXt9AMpWfkLqjUIK37Zw18Tel6a7yOS9x0L_NabxBgYd9KIJKZ6dTLl6VxxJI4M7Zstqj12wvOFtF8LjnYrCIWnhCDdeGsgpQvFpFX8VOoSO0dFJcOW_gRc6eX7mXDq80sOwQAlQWNhlQg\/s1600\/image1.png\" alt=\"\"><\/a><\/div>\n<h2><span class=\"ez-toc-section\" id=\"Opening_Android_Bench_to_group_contributions\"><\/span>Opening Android Bench to group contributions<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>From the start, we\u2019ve valued an open and clear method, which is why we made our authentic methodology and take a look at harness publicly obtainable on GitHub. You\u2019ve requested for a approach to supply suggestions on our dataset, so now we\u2019re taking collaboration a step additional by providing you with, the Android developer group, an opportunity to form Android Bench.<\/p>\n<p>Beginning right this moment, you&#8217;ll be able to contribute to Android Bench in two methods:<\/p>\n<p>We will likely be reviewing the submitted duties and will likely be assessing in the event that they get added to the benchmark. We hope to construct a benchmark that actually displays the varied, day-to-day realities of the worldwide Android developer group.<\/p>\n<h2 style=\"margin-top: 10px;\"><span class=\"ez-toc-section\" id=\"Trying_forward\"><\/span>Trying forward<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>With an increasing number of choices for agentic improvement, sustaining a cutting-edge benchmark ensures that the AI help you depend on retains getting smarter, extra useful, and more practical. Head over to our <a href=\"https:\/\/github.com\/android-bench\/android-bench\" target=\"_blank\" rel=\"noopener\">GitHub repository<\/a> to take a look at the duties. We invite you to submit a process to our crew for overview, and you may try <a href=\"https:\/\/hub.harborframework.com\/datasets\/android-bench\/android-bench\/latest\" target=\"_blank\" rel=\"noopener\">Harbor Hub<\/a> to discover the dataset or submit evaluations.<\/p>\n<p>As all the time, yow will discover the <a href=\"http:\/\/d.android.com\/bench\" target=\"_blank\" rel=\"noopener\">up to date leaderboard<\/a>, or learn the <a href=\"http:\/\/d.android.com\/bench\/methodology\" target=\"_blank\" rel=\"noopener\">methodology<\/a> on our web site.<\/p>\n<p>  <span style=\"display: none !important; visibility: hidden;\"><br \/>\n    Android Bench, LLM leaderboard, Harbor framework, Android improvement, Claude Fable 5, GPT 5.5, Claude Sonnet 5, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, Qwen 3.7 Max, AI benchmarking, Jetpack Compose migration, wearable networking, cell AI agent, Zoe Lopez-Latorre, mannequin analysis, open-weight fashions, developer group contributions.<br \/>\n<\/span>\n  <\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Again in March, we launched Android Bench\u2014our LLM leaderboard for real-world Android improvement duties. Our objective was to supply transparency round mannequin capabilities in Android improvement and to encourage mannequin enhancements, to offer you extra useful AI choices to your on a regular basis workflow. Since then, we&#8217;ve enhanced the benchmark based mostly in your [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":30093,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[23],"tags":[],"class_list":["post-30091","post","type-post","status-publish","format-standard","has-post-thumbnail","category-mobile"],"_links":{"self":[{"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/posts\/30091","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=30091"}],"version-history":[{"count":1,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/posts\/30091\/revisions"}],"predecessor-version":[{"id":30092,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/posts\/30091\/revisions\/30092"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=\/wp\/v2\/media\/30093"}],"wp:attachment":[{"href":"https:\/\/aireviewirush.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=30091"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=30091"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/aireviewirush.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=30091"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}