A late-night release from Google, and a small model has just knocked the company's own flagship image-generation product off its perch.
The new release — call it Nano Banana 2.1 — is positioned as a Flash-tier model, the lighter and cheaper sibling in the Gemini 3 family. The headline numbers, though, are not particularly lightweight. The model supports masked local editing where only the selected region changes. It holds subject identity across multiple iterations, with up to fourteen reference images in a single prompt. Output reaches 4K. Per-image cost is half that of the previous generation. And on a popular public image-editing leaderboard, the model lands at number four globally — with only the three flagship models from the leading competitor ahead of it.
Availability began rolling out immediately across the Gemini app, the search engine's AI mode, the developer studio, and the company's video tool, with more surfaces following through the week.
How a Flash Model Beat the Flagship
The product belongs to the Gemini 3 family. Under the conventional naming scheme it would also be called Gemini 3.6 Flash Image. Its knowledge cutoff sits at March 2026.
On the official benchmark suites covering visual design, image editing, and subject consistency, the internal numbers are blunt. The new model leads comfortably. There is still room to improve in narrow areas — small-text rendering, three-dimensional spatial reasoning, and factuality — but for the workflows that depend on producing and iterating imagery, the upgrade is significant.
Use the model and the practical capabilities come into focus. It creates and edits images with professional-grade control, including rapid iteration across many rounds. It renders clear, sharp text suitable for posters and complex infographics. It carries long-context real-world knowledge into image composition. It localises text rendering across multiple languages.
Independent user testing confirms the picture. On the most-watched public leaderboard for image generation and editing, the new model places fourth on multi-image editing, fifth on text-to-image, and sixth on single-image editing. The previous generation was seventh, eleventh, and fourteenth respectively on those same boards. Only three models from the leading competitor sit ahead on the multi-image editing board — the most relevant surface for the kind of work most creators actually do.
Combined with the pricing change, the strategic signal is hard to miss. Google is opening a price war on the image-generation leaderboard.
Edit One Region, Keep the Rest Intact
Anyone who has used a mainstream image editor knows the pattern. Lasso the area you want to change. The rest of the image stays put.
Doing the same thing with a generative model has historically been unreliable. Tell the model to change the colour of the cup in the lower-left corner and it has to first guess which cup you mean, then re-generate the whole image and hope the rest of the scene survives intact.
The new release turns that into a single-step operation. Mark a region, describe the change, and only the marked region moves. Swap a shirt, replace a sign, airbrush a passerby out of the background. Other things stay where they were.
One widely-shared demonstration runs through three rounds. First, generate a banknote featuring a pensive philosopher whose hair dissolves into a labyrinth. Second, on the same image, edit the serial number, give the philosopher a banana for a telephone, and add a banner. Third, restyle the whole thing as an oil painting. Across all three, the labyrinth, the robes, and the beard remain untouched.
The model-card score for masked editing is the single largest jump in the release: a 1049 score, beating the previous generation by 84 points and the company's flagship model by 122. Even without the chain-of-thought mode enabled, the new model still beats the flagship on every dimension.
Subject Consistency: Stop the Face Drift
The second upgrade is the one most e-commerce studios and brand designers have been waiting for. Subject consistency is the polite term for: please do not let the model forget what the person looks like after round three.
The previous generation was passable at single-subject consistency and struggled with multi-subject scenes. The new model ships with developer-facing multi-reference slots. Up to ten object-reference images and four human-reference images can be fed into a single generation, and the model treats them as anchors.
A live demonstration from a working creator illustrates the impact. Feed in a single model portrait, then five separate product shots, and ask for six different scene compositions. The result is six images where the model, every garment, and the accessories hold their identity — good enough to drop straight into a product page.
The benchmark number for multi-human consistency has jumped by 128 points compared with the previous generation. For anyone whose business is producing variations of the same character, that single number is the whole story.
Cheaper to Draw, Pricier to Think
Per-image output has dropped from $0.067 to $0.034 for a 1K image, and from $0.151 to $0.076 for a 4K image. Other line items have moved the other way. Input-token costs are now three times higher than before. Reasoning and text output costs have roughly doubled and a half. Read together, the move is a deliberate reallocation of cost: less money for drawing, more money for thinking.
For developers, the new model is already selectable in the AI Studio under the identifier gemini-nano-banana-2.1. The previous-generation API will be retired by the end of the month. Consumer access is also live across the Gemini app, the AI mode inside search, the video tool, the design tool, the ads platform, and the enterprise Gemini tier.
The Image-Editing Arms Race
The model landed less than an hour before independent testers began putting it through torture tests. One prompt required a single image containing a crowd, loose objects, legible text, and the reflections of a wet street. Another asked for a tea-house menu board where the same tea item appeared in four languages in one image. A third targeted the known failure modes of image models: layered jewellery, tinted sunglasses, fine-line tattoos. Both the new Google model and the latest flagship from the leading competitor were given the same prompts.
Across the head-to-head, the new model placed detail where requested, rendered layered jewellery correctly, and missed the tint on the sunglasses. The competitor's output looked more cinematic but changed the subject's complexion across a quarter of the face. Speed and cost tell their own story: the new Google model produced a single image in about eleven seconds at $0.044, while the leading competitor's equivalent needed fifty seconds and cost $0.069.
Both labs are now optimising for the same target — editable, controllable, identity-stable image generation — because that is what real workloads demand. A model that survives twenty edits without breaking matters more, in practice, than a model that produces one breathtaking image and then collapses. That is why the new release is being pushed first into the company's advertising platform and enterprise tier.
The unspoken question at the bottom of all this is what happens next at the top of the line. If a Flash-tier model can already outperform the flagship on editing, the flagship's replacement is going to be a serious event when it arrives.