image model · by Google DeepMind
Nano Banana 2 Lite AI Image Generator
Generate and edit images fast with Nano Banana 2 Lite, Google DeepMind's Gemini 3.1 Flash-Lite Image model released on June 30, 2026. Google calls it the fastest, most cost-efficient Gemini image model, with text-to-image in about 4 seconds while keeping the Nano Banana family's character consistency, precise editing, real-world knowledge and legible in-image text. Available on Fliki Standard and Premium plans.
Generated with Nano Banana 2 Lite
A handful of Nano Banana 2 Lite samples generated inside Fliki. No edits, no post.

Prompt
A vertical TikTok hook frame inside a busy classic Chicago pizzeria on a Friday night. A 24 year old Puerto Rican woman with curly dark hair in a high puff, a vintage Chicago Cubs t shirt and silver hoop earrings leans toward the camera over a red and white checkered tablecloth. In front of her sits a whole Chicago style deep dish pizza in a dark seasoned steel pan, one slice lifted out with a pie server so the tall buttery crust edge, the thick layer of chunky tomato sauce on top and the long stretch of melted mozzarella underneath are clearly visible. She holds the slice up with both hands on the server, the cheese pulling in a long string back to the pan, and gives the lens a wide eyed, daring grin as if challenging the viewer. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. Around her, the restaurant has dark wood booths, framed old black and white photos of the city on the walls, a neon beer sign glowing red and blue in the background, and other diners of different ages and backgrounds laughing out of focus. A pitcher of iced water with beaded condensation and two red plastic cups sit near the pan. At the top of the frame, a bold white caption with a black outline in a thick sans serif reads "CAN I FINISH THIS ALONE?" on two lines, and a small yellow tag near the pan reads "3 LBS". Those are the only words. Warm overhead pendant lighting makes the cheese glisten and the crust look golden, with a soft cool fill from the neon sign on the side of her face. Shot like a phone at arm length on a 26mm lens, slightly from above, with the pizza and her face both sharp and the room softly blurred. Vivid, fun, appetizing creator footage. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A calm, faceless quote frame for an Instagram Reel, photographed from the passenger seat of a car parked on a pullout along California State Route 1 in Big Sur, just after sunrise. Through the open passenger window and windshield, the rugged coastline curves away: green and golden hills dropping into a misty blue Pacific, white surf lines at the base of the cliffs, and the recognizable concrete arch of Bixby Creek Bridge in the middle distance catching the first warm light. Thin morning fog lies low over the water. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. In the lower part of the frame, the dashboard edge is visible in soft focus with a folded paper road map of the California coast and a travel mug of coffee resting on it. No people, no hands. In the upper middle, over the soft sky, clean graphic type in a light, elegant serif in warm white, centered on three lines: "The best views" / "come after" / "the longest drives." Beneath it, a thin short line and then small tracked uppercase sans serif text: "ROAD NOTES". Nothing else is written. Gentle, low golden side light from the east, soft haze, pastel sky shifting from peach near the horizon to pale blue above. Shot on a 35mm lens at f/4, focus on the bridge and coastline with the dashboard slightly soft. Quiet, nostalgic, film inspired color with soft grain. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A vertical product teaser photograph for a new trail running shoe, shot outdoors on a red sandstone ledge in Sedona, Arizona, at golden hour. A single right shoe stands on the rock in a three quarter side view, angled slightly toward the camera. The shoe has a breathable engineered mesh upper in a bright tangerine orange that fades into sand beige at the heel, a chunky charcoal midsole with visible foam texture, and an aggressive black rubber outsole with deep lugs, a few grains of red dust caught in them. Flat reflective laces in black. On the side of the midsole, in small raised white letters, the model name "RIDGELINE 3". FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. Around the shoe, a few loose pebbles and a small dry juniper twig on the rock. Behind it, out of focus, the famous layered red rock buttes of Sedona glow orange and crimson, with a clear deep blue sky and a thin wisp of cloud. At the top of the frame, bold condensed sans serif text in white with a subtle shadow reads "BUILT FOR THE CLIMB", and near the bottom, above the safe area, a small white pill tag with black text reads "OCT 1". Those plus the midsole name are the only words. Strong low sun from behind camera left rakes across the mesh, revealing every knit texture and throwing a long warm shadow across the rock, with a crisp orange rim on the heel. Shot low to the ground on an 85mm lens at f/4 so the shoe is sharp and the buttes blur into shapes. Rich, warm, energetic outdoor product photography. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A vertical educational hook frame for a fun facts Short, filmed on the Marin Headlands side of San Francisco looking back at the Golden Gate Bridge on a clear, windy morning. In the foreground on the left third, a 45 year old white man with a salt and pepper beard, a navy windbreaker and a canvas bucket hat stands at the overlook railing, turned toward the camera and pointing back over his shoulder at the bridge with a curious, teasing expression, as if he is about to share a surprising answer. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. The bridge fills the background: the two tall Art Deco towers in their distinctive International Orange, the sweeping main cables and vertical suspenders, the deck crossing to the city, the skyline and the bay with a few sailboats beyond. Low fog rolls in under the deck near the city side, and the hills in front are covered in dry golden grass and green coastal scrub. At the top, bold white sans serif with a black outline, two lines: "WHY IS IT ORANGE?" and below it, in smaller yellow type, "It was never supposed to be." No other text. Bright morning sun from the right lights his face and the bridge towers, with crisp shadows and a deep blue sky. Shot on a 35mm lens at f/8 so both the presenter and the bridge are sharp. Clean, vivid, trustworthy explainer look. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A UGC style ad still for a set of glass meal prep containers, shot as a real customer video frame on a Sunday afternoon in a small apartment kitchen in Houston. A 37 year old Vietnamese American man with short black hair, rectangular black glasses and a heather grey hoodie stands at the counter, smiling proudly at the phone camera propped against the backsplash. In front of him, five clear rectangular glass containers with snap lock bright teal lids sit in a neat row, each filled with the same meal: jasmine rice, sliced lemongrass chicken, steamed broccoli and a lime wedge. An open cardboard shipping box with crinkled kraft paper fill sits to one side, one teal lid still inside. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. He holds one container up at chest height with both hands, tilting it slightly toward the lens so the food and the lid are clearly visible. On the lid, a small embossed word reads "PREPLOCK". A stainless steel rice cooker, a bamboo cutting board and a bottle of sriracha are on the counter behind him, and a window over the sink lets in soft daylight. Near the top of the frame, a social caption on a white rounded rectangle in black sans serif reads "5 lunches. 40 minutes." and below it, small white text with a soft shadow reads "my coworkers keep asking". Those plus the lid word are the only text. Soft natural window light from the left, a warm household ceiling light as fill, phone camera look on a 26mm lens, slight wide angle, true colors and realistic food texture. Authentic, cheerful and relatable. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A widescreen YouTube thumbnail for a video ranking the best bagels in New York City. On the right third, a 28 year old Black woman with a short tapered natural cut, gold hoops and a cozy burgundy crewneck sweatshirt holds a huge everything bagel sandwich close to the camera: a toasted everything bagel packed with thick scallion cream cheese, silky lox, capers, red onion and tomato, with seeds scattered on the paper wrap. She gives an exaggerated delighted face, eyes closed, mouth open mid bite reaction. FRAME: Horizontal 16:9 widescreen frame, readable at thumbnail size on a phone. Use the left and right thirds deliberately, one side for the subject and one side for the headline, with strong separation between them. On the left and center, an iconic Manhattan street corner deli storefront, blurred just enough: a green awning, a window full of bagels on wooden trays, steam on the glass, a yellow taxi passing, brownstone steps on the side street. Early morning autumn light with a few fallen leaves on the sidewalk. Across the upper left, huge bold white condensed sans serif text with a thick black outline reads "I ATE 20 BAGELS", and underneath, in a smaller red rectangle with white text, "ONE WINNER". A hand drawn style yellow arrow points from the text to the sandwich. No other words. Bright crisp key light on her face from the left, a warm rim from behind, saturated but natural colors, strong contrast for readability at small sizes. Shot on a 35mm lens at f/2.8, the sandwich and her face sharp and the storefront soft. Clean, punchy thumbnail composition with nothing important in the bottom right corner. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A square Instagram feed post for a family farm in Vermont, featuring its recurring mascot character so the same character can be reused across a whole fall campaign. The mascot is a friendly scarecrow named Hollis: a stitched burlap face with a warm stitched smile, round black button eyes, rosy painted cheeks, a floppy straw hat with a red plaid band, a patched blue denim overall over a mustard flannel shirt, and straw poking out of his sleeves and collar. He is a soft, handmade looking figure with a slightly lopsided hat, photographed as a real crafted character, not a cartoon. FRAME: Square 1:1 feed frame. Center of gravity sits in the middle, with balanced breathing room on all four sides and nothing important in the corners, which get cropped in grid previews. Hollis sits on a hay bale in the center of a pumpkin patch, holding a small round orange pumpkin in both hands on his lap. Around him, dozens of pumpkins in orange, white and pale green, a wooden wheelbarrow, and a hand painted wooden sign leaning against the bale. Behind him, a classic red Vermont barn with white trim and rolling hills of maple trees in bright red, orange and yellow peak foliage. The wooden sign reads, in hand painted white letters on weathered red boards, "PICK YOUR OWN" on the top line and "OPEN DAILY" on the second line. That is the only text in the image. Late afternoon golden sun from the left, long soft shadows across the pumpkins, a little haze in the air, warm autumn color that is rich but natural. Shot on a 50mm lens at f/4, Hollis centered and sharp, the barn softly out of focus. Cozy, wholesome, seasonal. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A vertical TikTok hook frame inside a busy classic Chicago pizzeria on a Friday night. A 24 year old Puerto Rican woman with curly dark hair in a high puff, a vintage Chicago Cubs t shirt and silver hoop earrings leans toward the camera over a red and white checkered tablecloth. In front of her sits a whole Chicago style deep dish pizza in a dark seasoned steel pan, one slice lifted out with a pie server so the tall buttery crust edge, the thick layer of chunky tomato sauce on top and the long stretch of melted mozzarella underneath are clearly visible. She holds the slice up with both hands on the server, the cheese pulling in a long string back to the pan, and gives the lens a wide eyed, daring grin as if challenging the viewer. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. Around her, the restaurant has dark wood booths, framed old black and white photos of the city on the walls, a neon beer sign glowing red and blue in the background, and other diners of different ages and backgrounds laughing out of focus. A pitcher of iced water with beaded condensation and two red plastic cups sit near the pan. At the top of the frame, a bold white caption with a black outline in a thick sans serif reads "CAN I FINISH THIS ALONE?" on two lines, and a small yellow tag near the pan reads "3 LBS". Those are the only words. Warm overhead pendant lighting makes the cheese glisten and the crust look golden, with a soft cool fill from the neon sign on the side of her face. Shot like a phone at arm length on a 26mm lens, slightly from above, with the pizza and her face both sharp and the room softly blurred. Vivid, fun, appetizing creator footage. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A calm, faceless quote frame for an Instagram Reel, photographed from the passenger seat of a car parked on a pullout along California State Route 1 in Big Sur, just after sunrise. Through the open passenger window and windshield, the rugged coastline curves away: green and golden hills dropping into a misty blue Pacific, white surf lines at the base of the cliffs, and the recognizable concrete arch of Bixby Creek Bridge in the middle distance catching the first warm light. Thin morning fog lies low over the water. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. In the lower part of the frame, the dashboard edge is visible in soft focus with a folded paper road map of the California coast and a travel mug of coffee resting on it. No people, no hands. In the upper middle, over the soft sky, clean graphic type in a light, elegant serif in warm white, centered on three lines: "The best views" / "come after" / "the longest drives." Beneath it, a thin short line and then small tracked uppercase sans serif text: "ROAD NOTES". Nothing else is written. Gentle, low golden side light from the east, soft haze, pastel sky shifting from peach near the horizon to pale blue above. Shot on a 35mm lens at f/4, focus on the bridge and coastline with the dashboard slightly soft. Quiet, nostalgic, film inspired color with soft grain. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A widescreen YouTube thumbnail for a video ranking the best bagels in New York City. On the right third, a 28 year old Black woman with a short tapered natural cut, gold hoops and a cozy burgundy crewneck sweatshirt holds a huge everything bagel sandwich close to the camera: a toasted everything bagel packed with thick scallion cream cheese, silky lox, capers, red onion and tomato, with seeds scattered on the paper wrap. She gives an exaggerated delighted face, eyes closed, mouth open mid bite reaction. FRAME: Horizontal 16:9 widescreen frame, readable at thumbnail size on a phone. Use the left and right thirds deliberately, one side for the subject and one side for the headline, with strong separation between them. On the left and center, an iconic Manhattan street corner deli storefront, blurred just enough: a green awning, a window full of bagels on wooden trays, steam on the glass, a yellow taxi passing, brownstone steps on the side street. Early morning autumn light with a few fallen leaves on the sidewalk. Across the upper left, huge bold white condensed sans serif text with a thick black outline reads "I ATE 20 BAGELS", and underneath, in a smaller red rectangle with white text, "ONE WINNER". A hand drawn style yellow arrow points from the text to the sandwich. No other words. Bright crisp key light on her face from the left, a warm rim from behind, saturated but natural colors, strong contrast for readability at small sizes. Shot on a 35mm lens at f/2.8, the sandwich and her face sharp and the storefront soft. Clean, punchy thumbnail composition with nothing important in the bottom right corner. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A vertical product teaser photograph for a new trail running shoe, shot outdoors on a red sandstone ledge in Sedona, Arizona, at golden hour. A single right shoe stands on the rock in a three quarter side view, angled slightly toward the camera. The shoe has a breathable engineered mesh upper in a bright tangerine orange that fades into sand beige at the heel, a chunky charcoal midsole with visible foam texture, and an aggressive black rubber outsole with deep lugs, a few grains of red dust caught in them. Flat reflective laces in black. On the side of the midsole, in small raised white letters, the model name "RIDGELINE 3". FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. Around the shoe, a few loose pebbles and a small dry juniper twig on the rock. Behind it, out of focus, the famous layered red rock buttes of Sedona glow orange and crimson, with a clear deep blue sky and a thin wisp of cloud. At the top of the frame, bold condensed sans serif text in white with a subtle shadow reads "BUILT FOR THE CLIMB", and near the bottom, above the safe area, a small white pill tag with black text reads "OCT 1". Those plus the midsole name are the only words. Strong low sun from behind camera left rakes across the mesh, revealing every knit texture and throwing a long warm shadow across the rock, with a crisp orange rim on the heel. Shot low to the ground on an 85mm lens at f/4 so the shoe is sharp and the buttes blur into shapes. Rich, warm, energetic outdoor product photography. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A square Instagram feed post for a family farm in Vermont, featuring its recurring mascot character so the same character can be reused across a whole fall campaign. The mascot is a friendly scarecrow named Hollis: a stitched burlap face with a warm stitched smile, round black button eyes, rosy painted cheeks, a floppy straw hat with a red plaid band, a patched blue denim overall over a mustard flannel shirt, and straw poking out of his sleeves and collar. He is a soft, handmade looking figure with a slightly lopsided hat, photographed as a real crafted character, not a cartoon. FRAME: Square 1:1 feed frame. Center of gravity sits in the middle, with balanced breathing room on all four sides and nothing important in the corners, which get cropped in grid previews. Hollis sits on a hay bale in the center of a pumpkin patch, holding a small round orange pumpkin in both hands on his lap. Around him, dozens of pumpkins in orange, white and pale green, a wooden wheelbarrow, and a hand painted wooden sign leaning against the bale. Behind him, a classic red Vermont barn with white trim and rolling hills of maple trees in bright red, orange and yellow peak foliage. The wooden sign reads, in hand painted white letters on weathered red boards, "PICK YOUR OWN" on the top line and "OPEN DAILY" on the second line. That is the only text in the image. Late afternoon golden sun from the left, long soft shadows across the pumpkins, a little haze in the air, warm autumn color that is rich but natural. Shot on a 50mm lens at f/4, Hollis centered and sharp, the barn softly out of focus. Cozy, wholesome, seasonal. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A vertical educational hook frame for a fun facts Short, filmed on the Marin Headlands side of San Francisco looking back at the Golden Gate Bridge on a clear, windy morning. In the foreground on the left third, a 45 year old white man with a salt and pepper beard, a navy windbreaker and a canvas bucket hat stands at the overlook railing, turned toward the camera and pointing back over his shoulder at the bridge with a curious, teasing expression, as if he is about to share a surprising answer. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. The bridge fills the background: the two tall Art Deco towers in their distinctive International Orange, the sweeping main cables and vertical suspenders, the deck crossing to the city, the skyline and the bay with a few sailboats beyond. Low fog rolls in under the deck near the city side, and the hills in front are covered in dry golden grass and green coastal scrub. At the top, bold white sans serif with a black outline, two lines: "WHY IS IT ORANGE?" and below it, in smaller yellow type, "It was never supposed to be." No other text. Bright morning sun from the right lights his face and the bridge towers, with crisp shadows and a deep blue sky. Shot on a 35mm lens at f/8 so both the presenter and the bridge are sharp. Clean, vivid, trustworthy explainer look. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.

Prompt
A UGC style ad still for a set of glass meal prep containers, shot as a real customer video frame on a Sunday afternoon in a small apartment kitchen in Houston. A 37 year old Vietnamese American man with short black hair, rectangular black glasses and a heather grey hoodie stands at the counter, smiling proudly at the phone camera propped against the backsplash. In front of him, five clear rectangular glass containers with snap lock bright teal lids sit in a neat row, each filled with the same meal: jasmine rice, sliced lemongrass chicken, steamed broccoli and a lime wedge. An open cardboard shipping box with crinkled kraft paper fill sits to one side, one teal lid still inside. FRAME: Vertical 9:16 phone frame. Keep the face, product and headline inside the middle two thirds. Leave the top eighth and the bottom fifth free of critical detail, because app interface and captions sit there. Favor vertical depth: foreground, subject, background stacked top to bottom. He holds one container up at chest height with both hands, tilting it slightly toward the lens so the food and the lid are clearly visible. On the lid, a small embossed word reads "PREPLOCK". A stainless steel rice cooker, a bamboo cutting board and a bottle of sriracha are on the counter behind him, and a window over the sink lets in soft daylight. Near the top of the frame, a social caption on a white rounded rectangle in black sans serif reads "5 lunches. 40 minutes." and below it, small white text with a soft shadow reads "my coworkers keep asking". Those plus the lid word are the only text. Soft natural window light from the left, a warm household ceiling light as fill, phone camera look on a 26mm lens, slight wide angle, true colors and realistic food texture. Authentic, cheerful and relatable. FINISH: Describe it as one coherent, believable photograph. Keep real-world objects, places and foods accurate to how they actually look. Faces are natural and specific, with real skin texture and relaxed expressions; hands have five fingers and hold objects convincingly. Every quoted string appears exactly once, spelled exactly as written, in clean legible lettering with even spacing. The image contains no other words, no logos and no watermark marks, and the look stays crisp at 1K with clean edges and true color. Build depth in three layers: a clearly readable foreground detail, a sharp main subject, and a softer background that still reads as the real place. Keep the palette controlled, with one dominant color family, one accent color that draws the eye to the subject, and neutral tones everywhere else. Let the light tell the time of day and the season without any extra props, and make sure small details such as seams, crumbs, dust, fabric weave, water droplets and paint texture stay physically believable when viewed full screen on a phone.
What Nano Banana 2 Lite delivers
The fastest Gemini image model
Google reports text-to-image in about 4 seconds, which it puts at roughly 2.7 times faster than Nano Banana 2. Good for rapid iteration and batches of drafts.
Legible in-image text
Google says the Lite model keeps legible in-image text rendering despite the speed focus. Quote your headline and keep it short for the cleanest lettering.
Native 1K frames
Outputs are generated at 1K. Fliki renders true 9:16 at 768 x 1376, 16:9 at 1376 x 768 and 1:1 at 1024 x 1024, the same frame sizes as Nano Banana 2.
Character consistency
Nano Banana 2 Lite keeps the family's character consistency, so the same person or mascot can appear across a series of frames. Fliki accepts up to 14 references.
Edit, generate and compose in one model
A single model handles text-to-image, image editing and multi-image composition. Swap a background, change a product color or merge references without switching tools.
Real-world knowledge
Google says the Lite model keeps the real-world knowledge of the Nano Banana family, so landmarks, foods, objects and everyday scenes come out recognizable.
Half the credits of Nano Banana 2
On Fliki, Nano Banana 2 Lite costs half the credits of Nano Banana 2 per image, which makes it the everyday choice for drafts and quick edits.
SynthID watermarking
Every output carries Google's invisible SynthID watermark, so images can be identified as AI-generated.
How it works
How to generate an image with Nano Banana 2 Lite
Nano Banana 2 Lite runs inside Fliki's AI image generator. Here's the six-step flow.

Write your prompt
Describe the scene like a brief: subject, setting, lighting, camera and any words that should appear on the image in quotes. Nano Banana 2 Lite follows detailed prompts, so specifics pay off.

Select Nano Banana 2 Lite as your model
Open Fliki's model selector and choose Nano Banana 2 Lite. Your prompt routes to Google DeepMind's Gemini 3.1 Flash-Lite Image model.

Pick your aspect ratio
Choose 9:16 for vertical video, 16:9 for YouTube and slides, or 1:1 for feed posts.

Add reference images (optional)
Upload up to 14 references to edit a photo, keep a character consistent or combine several images into one composition.

Review the output size
Nano Banana 2 Lite outputs at 1K. Fliki renders 1024 x 1024 for square, 1376 x 768 for landscape and 768 x 1376 for vertical.

Generate
Hit Generate. Outputs carry Google's invisible SynthID watermark. Nano Banana 2 Lite is available on Fliki Standard and Premium plans.
AI MODEL GALLERY
Built on the best AI models - ready inside Fliki
Every leading video, voice, and image model - integrated, unified, and tuned for creators. Generate with the latest AI video, AI voice, and AI image models from OpenAI, Google, Kling, Bytedance, ElevenLabs, and more - all from one place.
















Nano Banana 2 Lite FAQ
Frequently asked questions
Everything you need to know about generating images with Nano Banana 2 Lite inside Fliki.
Nano Banana 2 Lite is Google DeepMind's Gemini 3.1 Flash-Lite Image model, released on June 30, 2026. Google describes it as its fastest, most cost-efficient Gemini image model.
Nano Banana 2 Lite is available on Fliki Standard and Premium plans. It is not included on the Free or Basic plans or on lifetime deal plans.
Nano Banana 2 (Gemini 3.1 Flash Image) is the higher-fidelity model with up to 4K output from Google. Nano Banana 2 Lite trades some fidelity for speed: about 4 seconds per image at 1K. On Fliki, Lite costs half the credits of Nano Banana 2 and both accept up to 14 references.
1K. Fliki renders 1024 x 1024 for 1:1, 1376 x 768 for 16:9 and 768 x 1376 for 9:16.
Yes. The same model handles text-to-image, editing and multi-image composition. Upload a photo as a reference and describe the change.
Yes. Google says it keeps legible in-image text rendering. Short, quoted headlines work best. For dense, text-heavy layouts, Nano Banana 2 or Meta Muse are stronger picks.
Yes. Outputs carry Google's invisible SynthID watermark so they can be identified as AI-generated.
Pick Nano Banana 2 Lite for fast drafts, quick edits and large batches at a lower credit cost. Pick GPT Image 2.5 when you need maximum photoreal detail and very long prompts.
Still curious?
Try Fliki free in your browser, no credit card required.
Start free→AI image models
Discover more
Tools
Discover features
Generate your next image with Nano Banana 2 Lite.
Upgrade to a Standard or Premium Fliki plan to unlock Nano Banana 2 Lite.
Upgrade to generateFree forever plan · No credit card required · Cancel anytime