From characters to infographics, consistency takes priority in Nano Banana 2.1

Image generation and editing become more consistent with Nano Banana 2.1, according to Google’s evaluations. The update improves character preservation and infographic composition, while text rendering and instruction following still have documented limitations.

Keeping several characters recognizable after an edit is one of Nano Banana 2.1’s clearest improvements in Google’s tests. Built on Gemini 3.6 Flash, the image generation and editing model outperforms its predecessors across the criteria published in its model card. In Thinking mode, it achieves an Elo score of 1,106 for consistency across multiple characters, compared with 978 for Nano Banana 2 and 1,011 for Nano Banana Pro. These results come from comparisons run by Google, with human evaluators assessing the images. They do not constitute an independent ranking.

Working from reference images is a substantial part of this version. Up to 14 images can be combined, with output available at 1K, 2K, and 4K resolutions. The API documentation also describes support for grounding generation in web and image searches. Measured improvements cover product preservation, style changes, and edits guided by a mask or sketch.

Infographics receive higher scores for composition and factual accuracy in the company’s evaluation protocol. Their factual content is assessed automatically, while visual comparisons rely on human judgments. The model card still acknowledges difficulties with small lettering, long paragraphs, and full pages of text. Character preservation remains imperfect, some editing instructions are only partially followed, and traces of a guiding sketch can remain in the result. Confusion between left and right is also documented.

Access is listed through the Gemini app, Google AI Studio, and the Gemini API, alongside Search AI Mode, Ads, Flow, and Stitch in the model card’s distribution channels. For integrations, the documentation lists `gemini-nano-banana-2.1` as a stable version and specifies three Thinking levels: minimal, medium by default, and high.