For over a year, the term natively multimodal has resonated in the world of artificial intelligence, but few have managed to fully leverage these capabilities. Now, Google has made its move with the launch of its latest model, Gemini 2.0 Flash Experimental, which allows not only generating images but also editing them natively. In other words, it has given Photoshop a pinch 🤏🏼…
Why is image generation so important? Although AI image generation has been available through chatbots like ChatGPT, these often rely on specialized models like Dall-E 3 or Imagen 3, which are extensions of the main model and not an integral part of it. In contrast, models like Gemini are natively multimodal, meaning they can understand and create both text and images intrinsically.
Natively Image Generation with Gemini 2.0 Flash Experimental
Currently, this native image generation feature is not available to all users. The Gemini 2.0 Flash Experimental model can be tested for free at the Google AI Studio and will soon be available to a wider audience. After experimenting with this model, I can say that the experience was truly remarkable.
I started by asking Gemini to create a visual guide on how to make Bolognese macaroni. The results were astonishing, showing a remarkable consistency among the generated images, from the pan to the ingredients. Each image maintains the same resolution of 1024 x 680, making it easy to create visual guides on any topic.
Then, I asked Gemini to generate an empty room, and I kept asking for modifications on the decoration and utility of the room. The continuity it maintained was astonishing.
Natively Image Editing with Gemini 2.0 Flash Experimental
To demonstrate the image editing feature, I uploaded a photo of my garage and asked it to change my car to a white Tesla, and the result was impressive. Finally, I asked it to add some tables with computers, and this showed me the potential of image editing thanks to Gemini’s native multimodal capability. They weren’t perfect, but they were very good. Additionally, I asked Gemini to colorize an old black and white photo, and the result exceeded my expectations, with optimal visual quality and no visible errors.
The possibilities with Gemini are vast and exciting. Google has done an admirable job of integrating image generation and editing natively. With the recent launch of Veo 2 for video generation and Imagen 3 for specialized image generation, it seems that Google has surpassed OpenAI in several aspects, not just in text generation. It will be interesting to see how OpenAI responds to this advancement with its ChatGPT.
In IT since 2002: systems, engineering, web development and SEO. I have been doing business with AI since February 2023, when ChatGPT could first be bought in Spain.
Related guides
-
Business applications
Discover artificial intelligence agents and their functioning in the new digital era
Artificial intelligence is changing rapidly, and it is no longer just about chatbots that answer questions. Since the arrival of ChatGPT at the end of…
-
ChatGPT -
Gemini
-
-
Essentials
Discover how to create your own custom AI applications for specific tasks
Tired of writing bibles in ChatGPT just to get it to understand what you want? Well, there’s good news: creating your own AI app is easier than you…
-
ChatGPT -
Gemini
-
-
Business applications
Discover the 10 examples of AI agents that will transform your life in 2025
Artificial intelligence is advancing by leaps and bounds, leaving behind old chatbots and giving way to action-oriented AI agents. These not only…
-
Gemini
-
Anything left unclear? Ask me
I, Miguel Ángel, answer right here. And if something is out of date, tell me and I’ll fix it.