Live access is not confirmed for this model. No current purchase price is advertised.
Related provider subscriptions
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
17.2
Quality Score
---
Arena ELO
Undisclosed
Parameters
---
Context
Evidence profile
How complete is this record?
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
LaunchesGoogle1w ago
Our @GoogleResearch Connectomics team, in collaboration with @HHMIJanelia, has released the complete wiring diagram of a male fruit fly’s brain and central nervous system — the largest brain map by nu
Our @GoogleResearch Connectomics team, in collaboration with @HHMIJanelia, has released the complete wiring diagram of a male fruit fly’s brain and central nervous system — the largest brain map by number of proofread neurons to date. So, why do we care so much about a tiny
We’re launching AlphaGenome Atlas: an AI-powered searchable database mapping the predicted impact of all 9 billion possible single-letter DNA changes. Here’s how it could help researchers better under
We’re launching AlphaGenome Atlas: an AI-powered searchable database mapping the predicted impact of all 9 billion possible single-letter DNA changes. Here’s how it could help researchers better understand our biology 🧵 https://t.co/SABFZW5SiR
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model generates a forecast with high spatial resolution in order to catch https://t.co/XNDjd7eoHw
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.
How can we reconstruct a memory that was never filmed? Our team paired restored archival photos with pose control models to capture the mannerisms and micro-expressions of Burt and Ethelle. This helpe
How can we reconstruct a memory that was never filmed? Our team paired restored archival photos with pose control models to capture the mannerisms and micro-expressions of Burt and Ethelle. This helped bring the day they first met to life for Love, Rendered, a new documentary https://t.co/2mMxcmRyva
X/Twitter@GoogleDeepMindGoogleannouncementgeneral1w ago
From tracking hurricanes to optimizing renewable energy grids, @PeterWBattaglia and @fryrsquared explore how WeatherNext 3 is shaping how we model forecasts and prepare in a fast-changing climate. Pod
From tracking hurricanes to optimizing renewable energy grids, @PeterWBattaglia and @fryrsquared explore how WeatherNext 3 is shaping how we model forecasts and prepare in a fast-changing climate. Podcast timecodes: 00:00 Introduction 00:38 Hurricane Melissa 11:50 Why weather https://t.co/JmuAm6NdtE
Our @GoogleResearch Connectomics team, in collaboration with @HHMIJanelia, has released the complete wiring diagram of a male fruit fly’s brain and central nervous system — the largest brain map by nu
Our @GoogleResearch Connectomics team, in collaboration with @HHMIJanelia, has released the complete wiring diagram of a male fruit fly’s brain and central nervous system — the largest brain map by number of proofread neurons to date. So, why do we care so much about a tiny
We’re launching AlphaGenome Atlas: an AI-powered searchable database mapping the predicted impact of all 9 billion possible single-letter DNA changes. Here’s how it could help researchers better under
We’re launching AlphaGenome Atlas: an AI-powered searchable database mapping the predicted impact of all 9 billion possible single-letter DNA changes. Here’s how it could help researchers better understand our biology 🧵 https://t.co/SABFZW5SiR
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model generates a forecast with high spatial resolution in order to catch https://t.co/XNDjd7eoHw
X/Twitter@GoogleDeepMindGoogleannouncementgeneral2w ago
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅ Developed with @GoogleResearch, the model learns directly from real-world, real-time observations to give more localized high
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅ Developed with @GoogleResearch, the model learns directly from real-world, real-time observations to give more localized highly accurate predictions faster. 🧵 https://t.co/iI8c6uEN4n
X/Twitter@GoogleDeepMindGoogleannouncementgeneral2w ago
Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tas
Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable https://t.co/EEJDIMhRwp
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.