This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
LaunchesGoogleYesterday
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model generates a forecast with high spatial resolution in order to catch https://t.co/XNDjd7eoHw
We’re rolling out Gemini Omni 1.1 Flash to make generative video highly controllable, faster to iterate on, and more polished for production-grade use. Here’s how you can try it in @FlowbyGoogle and m
We’re rolling out Gemini Omni 1.1 Flash to make generative video highly controllable, faster to iterate on, and more polished for production-grade use. Here’s how you can try it in @FlowbyGoogle and more → https://t.co/BFHMropoby https://t.co/uCB46FHOs3
Last month, we released Lyria 3, enabling you to create tracks with lyrics from text, image, or video prompts. Now, we’re introducing Lyria 3 Pro, which expands upon our music generation model to offe
Last month, we released Lyria 3, enabling you to create tracks with lyrics from text, image, or video prompts. Now, we’re introducing Lyria 3 Pro, which expands upon our music generation model to offer additional advanced capabilities. What’s really special about this upgrade https://t.co/XeYjvVmUJC
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and https://t.co/puvIVxDjq7
Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model
Introducing WeatherNext 3️⃣— our most advanced global weather AI model yet from @GoogleDeepmind and @GoogleResearch With prediction capabilities that are up to 5x sharper than WeatherNext 2, the model generates a forecast with high spatial resolution in order to catch https://t.co/XNDjd7eoHw
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅ Developed with @GoogleResearch, the model learns directly from real-world, real-time observations to give more localized high
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅ Developed with @GoogleResearch, the model learns directly from real-world, real-time observations to give more localized highly accurate predictions faster. 🧵 https://t.co/iI8c6uEN4n
X/Twitter@GoogleDeepMindGoogleannouncementgeneral2d ago
Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tas
Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable https://t.co/EEJDIMhRwp
X/Twitter@GoogleDeepMindGoogleannouncementgeneral3d ago
We’re bringing agentic video understanding to our latest Gemini models. They can now analyze videos with better accuracy while using up to 88% fewer tokens. 🧵 https://t.co/ZTjlaLk7tX
We’re rolling out Gemini Omni 1.1 Flash to make generative video highly controllable, faster to iterate on, and more polished for production-grade use. Here’s how you can try it in @FlowbyGoogle and m
We’re rolling out Gemini Omni 1.1 Flash to make generative video highly controllable, faster to iterate on, and more polished for production-grade use. Here’s how you can try it in @FlowbyGoogle and more → https://t.co/BFHMropoby https://t.co/uCB46FHOs3
Meet Gemini Omni 1.1 Flash ⚡️ Our newest multimodal model for video generation and editing. It now features your favorite creative controls from Veo, plus brand new capabilities. Enjoy features like 4
Meet Gemini Omni 1.1 Flash ⚡️ Our newest multimodal model for video generation and editing. It now features your favorite creative controls from Veo, plus brand new capabilities. Enjoy features like 4K upscaling, first / last frame control, and fast 360p drafting. But, the https://t.co/eJiWsqFwxt
X/Twitter@GoogleDeepMindGoogleopen_sourceopen source1w ago
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and https://t.co/puvIVxDjq7
Last month, we released Lyria 3, enabling you to create tracks with lyrics from text, image, or video prompts. Now, we’re introducing Lyria 3 Pro, which expands upon our music generation model to offe
Last month, we released Lyria 3, enabling you to create tracks with lyrics from text, image, or video prompts. Now, we’re introducing Lyria 3 Pro, which expands upon our music generation model to offer additional advanced capabilities. What’s really special about this upgrade https://t.co/XeYjvVmUJC
Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.
Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption is restrictive for inference-only accelerators, frozen or third-party models, and memory-constrained deployments, and standard BatchNorm-based TTA configurations may also become inactive on architectures without BatchNorm. We study adaptation when the learned model must remain frozen. We introduce CASTER, a gradient-free method that stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification. CASTER requires no backward pass, optimizer state, or stored source feature bank. Across four backbones and seven datasets, it outperforms k-NN on identical frozen features in 27 of 28 backbone-dataset settings while retaining a median of 18x less state. Affine transport is not always reliable. On ImageNet-C, where batches contain only 64 samples for 1000 classes, unconditional transport loses 21.2 top-1 points. We therefore introduce an empirical residual-to-margin transportability certificate. Across 307 evaluation cells, every transport losing more than 10 points has certificate value above 3.9, although benign and destructive regimes are not perfectly separated. Gating converts an average -3.35-point effect of unconditional transport into a +1.69-point gain, and performance remains within 0.3 points of the best threshold over a broad threshold range. Finally, we show that this certificate is mechanism-specific: when applied to Tent, it accepts only 4.3% of updates and preserves 0.6% of Tent's available gain. These results position CASTER as a lightweight adaptation mechanism for frozen-model deployment, together with an explicit account of when its safety signal is informative and when it is not.
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions. Across families, construction challenges reflect not only modality but also how tightly decisions are coupled and whether failures become visible while they remain repairable. Decomposition can reduce local complexity while increasing coordination and reassembly costs. Learned judges may add little independent evidence when they share the generator's preferences or blind spots. We formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change. We also identify opportunities for sustaining coherent, accountable control as artifacts, creator intent, and construction systems evolve. A curated paper list is available at https://github.com/GeminiLight/awesome-agentic-artifact-creation.