Pollo AI Turns Creative Ideas into High-Converting Multimodal Campaigns with OpenAI
Pollo AI, an AI video and image platform with more than 26 million users, is the subject of an OpenAI customer story published on October 8, 2026. Its newest product, Pollo Agent, uses GPT-5.6 for routing, GPT-6 Astra for harder work such as narrative development and scene planning, and GPT-Image-2.5 for image generation. It turns rough ideas into posters, storyboards and video ads of about 15 seconds. The company says users spend more than 50% less time choosing or switching models, and that 30% of creator sessions begin from templates. All figures are self-reported.
Pollo AI, a platform that helps creators and marketers produce AI-generated video and images, now serves more than 26 million users. On October 8, 2026, OpenAI published a customer story describing how the startup uses GPT-5.6, GPT-6 Astra and GPT-Image-2.5 to turn rough creative ideas into detailed images and cinematic video ads. The center of the story is Pollo Agent, the company's newest product. It compresses the chain of choosing a model, writing a script, planning scenes, generating images and assembling a video into one guided workflow. 1. Positioning: "What Canva did for design" Bill Zhu, founder and CEO of Pollo AI, states the goal plainly: "We want to do for video what Canva did for design. We want to make it easy enough that more people can actually use it, not just the ones who already know how." The company started from a simple observation. Even with generative AI, creators and marketers still have to move between different tools and pick between models to get the result they want. Some abandon their projects in frustration. Pollo AI aims to remove that friction. The first breakthrough was video templates. According to the company, 30% of creator sessions now start from templates or trending-format workflows. Pollo Agent builds on that. Users can begin with simpler inputs and let the system handle the rest of the process. Zhu says that as OpenAI's models have improved, more of the workflows Pollo wanted to support have become practical inside the product. 2. How it works: a model mix matched to task difficulty Pollo Agent first interprets the user. It reads rough ideas and visual references, then develops storylines, scene flows and scripts. CTO Peter Zhou says OpenAI models stood out for their ability to understand ambiguous creative inputs and for their multimodal capability: "GPT-5.6 performed especially well there." The team also values the stability and release cadence of the OpenAI API.
To balance reasoning depth, speed and cost, Pollo AI splits the work across three roles: 1. GPT-5.6 handles routing. It decides what kind of task a request is and where to send it.
2. GPT-6 Astra handles the harder work, such as developing a narrative or planning scene revisions.
3. GPT-Image-2.5 handles image generation, including posters, storyboards and campaign visuals. In Zhou's words: "If we receive any task from a user to generate an image, we always use GPT-Image-2.5. It's especially good at working with different types of characters and letters."
This lets the system match the level of reasoning to the user's prompt, so simple requests do not pay for heavy models. Pollo AI reports that users spend more than 50% less time choosing or switching models, which leaves more time for creative work. For model selection, the team tests generative models against real user prompts and workflows. It evaluates output quality, consistency, stylistic range, speed and cost. One caution applies here: these figures are self-reported by the company. The story gives no test method and no third-party verification. 3. Three use cases from the story Photo to video ads. The Photo to Video Ads tool in Pollo's Marketing Studio suite uses OpenAI models to turn a still image and a prompt into a short video. It is designed to help small businesses and direct-to-consumer (DTC) brands generate social media product ads in minutes rather than days. In the example, a user wanted a luxurious fragrance ad from a product photo. The static photo could not convey the mood, and a professional shoot was beyond the budget. The tool produced a roughly 15-second video with velvet textures, warm spotlights and close-ups of the bottle, ending with the fragrance framed against deep red fabric.
A luxury campaign visual. A jewelry brand wanted a bold editorial campaign image, but its budget could not stretch to a high-concept studio shoot. Using Pollo Agent and GPT-Image-2.5, the brand realized its idea of statement jewelry on a model seated beside a cheetah. The tool framed the scene with dark drapery and warm lighting and kept the jewelry clearly visible. According to the story, the customer got a finished luxury visual with lower production effort, time and cost than a live shoot. Product shots turned into magical worlds. A beverage brand wanted its peach soda to feel playful and imaginative, which a standard studio shot could not deliver. With Pollo Agent and GPT-Image-2.5, which the story credits with sharper details and faster generation, the team turned a simple bottle image into a snowy miniature landscape with a train, lizards and sparkling droplets. 4. What this means for developers, enterprises and the ecosystem For developers, the story shows an architecture pattern that is easy to copy: a lighter model for routing, a stronger model for hard reasoning, and a dedicated model for images. Layering like this controls cost and latency at the same time. For marketing teams and small businesses, the value is a lower barrier to production. A usable social ad no longer needs a crew. For platform builders, the pairing of templates and an agent suggests that repeatable formats and less trial and error may matter more than adding more models.
5. Limits and outlook Readers should stay careful. OpenAI published this as a customer story, and every number and outcome comes from the company's own account. The article gives no conversion rates, no per-asset cost figures and no quality evaluations. The headline claim of "high-converting" campaigns therefore cannot be checked independently. Industry-level questions about brand consistency, asset rights and disclosure of AI-generated ads also go unaddressed. Looking ahead, Pollo wants AI video to become part of everyday content creation over the next 12 months. Zhu says that depends on making it even easier to turn an idea into an asset ready to post, test or use in a campaign. His view: "We think the AI shift in content creation will favor products that make creation more repeatable from the start, with stronger formats, less trial and error, and a simpler workflow around the models themselves." In other words, the next stage of competition may be less about whose single model is strongest, and more about who can orchestrate several models into a steady, low-friction production line.