Skip to content

Create multimodal applications with OpenAI models in Microsoft Foundry

Imagine you are building a home design application. A customer says, “Show me a modern kitchen with dark cabinets and a large center island.” Images appear in seconds. Before the application finishes describing the design, the customer interrupts: “Make it brighter and add natural wood accents.” The application listens, adapts, updates the visuals, and continues the conversation without losing context. OpenAI’s GPT-Live-1 and GPT-image-2.5 models make this type of conversational and creative experience possible, and they are now available in Microsoft Foundry.


Many applications today operate through a sequence of prompts and responses, one after the other, no overlap. The user asks a question, the AI answers, and the interaction starts again. Human collaboration does not work that way. People interrupt, clarify what they mean, react to new ideas, change direction, and build on concepts together in real time.


GPT-Live-1 brings that natural rhythm to AI applications through real-time, full-duplex conversation. The model can listen and speak simultaneously, helping developers create experiences that respond to interruptions, pauses, overlap, and other conversational signals. GPT-image-2.5 extends that interaction into visual creation, allowing applications to generate and edit images as ideas emerge.


From AI responses to AI collaboration


Until now, developers often had to assemble multimodal experiences from separate systems. A voice interface captured a request, another system processed it, an image model generated the output, and the user moved between different screens or workflows to refine the result.


Voice experiences were also largely built around a familiar turn-based pattern: the user spoke, the system waited for silence, processed the request, and then delivered a response. Even when the answer was useful, the interaction could still feel rigid.


GPT-Live-1 changes that dynamic. Instead of treating conversation as a series of isolated requests, the model can continuously evaluate whether to listen, speak, pause, acknowledge the user, respond to an interruption, or call a tool. This makes conversation more than an input method. It becomes the interface through which the user explores an idea and directs an application.


When that real-time conversation is combined with GPT-image-2.5, the application can turn ideas into visual output as the discussion unfolds. The user does not need to translate every thought into a carefully structured prompt or restart the workflow whenever a requirement changes. They can speak naturally, see the result, react, and continue creating.


What is GPT-Live-1?


Voice is becoming an important interface for customer engagement, employee productivity, accessibility, education, and hands-free work. However, many voice applications still depend on a rigid exchange in which one person speaks and the system waits for silence before responding.


GPT-Live-1 is designed for continuous, real-time interaction. Its bidirectional, full-duplex design allows it to process incoming audio while generating speech, creating experiences that can listen and speak at the same time. The model can respond to interruptions, adapt its pacing, recognize pauses, and use conversational timing to determine when to listen or respond.


The result is not simply faster voice AI. It is a different interaction model. Instead of requiring users to adapt their behavior to the application, developers can build applications that adapt to the natural flow of a conversation.


Key capabilities



  • Simultaneous audio input and output: Support continuous conversation without requiring strict turn-taking.

  • Natural interruption handling: Allow users to clarify, redirect, or add information while the application is speaking.

  • Conversational timing: Respond to pauses, pacing, and other signals that help conversations feel less mechanical.

  • Backchannel responses: Signal that the application is listening without unnecessarily taking over the conversation.

  • Contextual memory: Maintain continuity as users revisit ideas and refine what they need.

  • Multimodal inputs: Work with voice, text, and images to support richer application experiences.


How developers are using GPT-live-1


GPT-Live can support voice-first experiences wherever natural timing, continuous listening, and fast responses matter. Developers can combine its real-time speech capabilities with tools and application context to create experiences that move beyond rigid turn-by-turn interactions.



  • Customer service and contact centers: Build voice agents that answer questions, guide customers through tasks, respond to interruptions, and call tools to retrieve account or order information.

  • Voice-enabled productivity assistants: Let employees capture notes, find information, update records, or complete routine workflows through hands-free conversation.

  • Conversational commerce and discovery: Help customers refine product, travel, or service searches by describing preferences naturally and adjusting them throughout the conversation.

  • Education and coaching: Create interactive tutors, language-practice partners, and training simulations that adapt explanations, pacing, and follow-up questions in real time.

  • Accessibility experiences: Add responsive voice interaction to applications for people who prefer or rely on speech as an alternative to typing and screen-based navigation.

  • Field and frontline support: Deliver hands-free guidance for technicians, warehouse teams, and other mobile workers who need immediate answers while completing physical tasks.

  • Multilingual experiences: Support conversational applications that serve users across languages and use cases such as travel assistance, public services, and global customer engagement.


Audio models in action


CoStar Group is a leader in demonstrating how conversational audio can reshape a familiar digital experience. On Homes.com and Apartments.com, the company has developed an AI-powered housing-search experience that enables buyers and renters to describe what they want conversationally rather than relying only on fixed filters.



“We saw real-time conversational AI as the future of how people search for homes, and we wanted to lead that shift rather than follow it. Early adoption of Azure OpenAI’s Realtime models through Microsoft Foundry gave CoStar Group a foundation to bring that experience to millions of homes.com users.” – Andy Ventura, Vice President Applied AI and Enterprise Architecture



Using Azure Open AI models in Microsoft Foundry, CoStar designed the solution for responsive audio interaction, strong security, and the ability to scale across a marketplace used by more than 100 million visitors each month. A buyer can progressively refine a search through dialogue and explore details that can be difficult to express through traditional search fields, from neighborhood preferences to property and Matterport-powered visual insights.


Learn more in the CoStar Group customer story


What are GPT‑Image‑2.5 Flare and GPT‑Image‑2.5 Sunburst?


The GPT‑Image‑2.5 family gives developers two options for building image generation and editing into user-facing applications. GPT‑Image‑2.5 Flare is a fast, high-throughput model designed for most production workloads. As the smaller model, Flare delivers improved image quality, editing accuracy, and responsiveness for content creation, marketing assets, product experiences, visual search, and image generation at scale. GPT‑Image‑2.5 Sunburst is a premium model optimized for creative workflows that require greater precision and control. Sunburst is designed for high-fidelity campaign assets, product imagery, and other polished visual experiences where quality takes priority over generation speed. Together they bring:



  • Significantly faster generation and editing. GPT‑Image‑2.5 Flare delivers higher-quality images than GPT‑Image‑2 at 50% lower latency, helping keep users in a live, conversational iteration loop.

  • More accurate editing. Update targeted elements while maintaining the rest of the image.

  • Stronger multi-turn editing. Refine images through successive instructions without losing consistency or image quality.

  • Better instruction following. Interpret complex visual instructions, layouts, and stylistic directions more accurately.

  • Two models for different workflows. Use Sunburst for premium creative and editing workflows and Flare, the smaller model, for faster iteration.


See it in action:


Below is a series of images that demonstrates the improvements in multi-turn editing with GPT-images 2.5 Sunburst. The sequence shows the model responding to a series of editing instructions over multiple turns. As changes accumulate, the model maintains visual consistency, preserves details from earlier edits, and incorporates new requests without degrading image quality. These examples highlight how GPT-images 2.5 Sunburst enables more reliable iterative editing workflows, making it easier to refine assets through a sequence of targeted updates.



Prompt: A pair of running shoes placed on a forest trail at sunrise, sitting on packed dirt scattered with fallen leaves and pine needles. Soft golden sunlight filters through tall trees in the background, casting long shadows across the path. Morning mist hovers low between the trunks. Shallow depth of field, the shoes in sharp focus, photorealistic, natural color palette, cinematic lighting.



Base image

Edit 1 – lighting

Edit 2 – color

Edit 3 – composition


 


For developers and creators, this means less time recreating assets and more time refining them, with the confidence that each edit will build on the last.


How developers can use GPT‑Image‑2.5 models


Retail and media companies can use GPT‑Image‑2.5 Flare to accelerate content creation and production across high-volume digital channels, while GPT‑Image‑2.5 Sunburst supports premium creative and editing workflows that require tighter control across successive edits. Commercial teams can move from an initial concept to channel-ready variations in an interactive workflow without treating every adaptation as a separate production cycle.



  • Scale retail campaign creative. A retailer can create and refine seasonal product imagery for websites, retail media, email, social, and stores, testing settings and promotions while keeping the product and campaign direction consistent.

  • Personalize streaming artwork by viewer segment. A streaming service can tailor thumbnails for the same show to viewers’ preferred genres, then test which versions drive more visits and plays without separate photo shoots.

  • Create personalized virtual try-on experiences. An e-commerce retailer can let shoppers upload a photo or select a model to preview clothing and shoes, then adjust color, style, or setting to compare options before buying.


Bring conversation and creation together


Real-time conversation and image generation enable developers to build applications where users can describe ideas, refine them through dialogue, and see results as the conversation evolves. A traveler might explore destinations by describing preferences and reacting to generated imagery. A homeowner could talk through renovation ideas while new designs are generated in real time. A shopper might explain a preferred style and receive visual recommendations that adapt as requirements become more specific. In these experiences, conversation becomes the interface. GPT-Live-1 handles the interaction while GPT-image-2.5 generates and refines visual content, allowing applications to listen, respond, and create within a single workflow.



  • Travel planning: Users can explore locations, accommodations, activities, or itineraries through an ongoing dialogue while simultaneously viewing personalized imagery.

  • Education and learning: Students can discuss ideas with an AI tutor while receiving generated diagrams, illustrations, visual explanations, and supporting content that evolves throughout the lesson.

  • Customer support: Users can verbally describe an issue, provide photos, receive generated visual guidance, and continue asking questions within a single multimodal interaction.


Pricing


Pricing details will be available on the Microsoft Foundry pricing page within the coming weeks.


Model

Modality

Input

Cached input

Output

gpt-image-2.5-sunburst

Image

$8.00

$2.00

$30.00

Text

$5.00

$1.25

gpt-image-2.5-flare

Image

$8.00

$2.00

$30.00

Text

$5.00

$1.25

GPT-live-1

Audio

$3.00 per hour


Pricing is per 1 million tokens unless otherwise specified.


Get Started


GPT-Live-1 and GPT-image-2.5 are available today in Microsoft Foundry, giving developers everything they need to build applications that can listen, speak, generate, and edit content within a single multimodal experience.


To stay current on new model releases and discover implementation examples across model families, visit the Microsoft Foundry Model Releases repository , where you’ll find release resources, samples, and technical content accompanying new model announcements. Developers looking to deepen their understanding of models and AI application design can also explore the Model Mastery repository, which provides hands-on workshops and learning paths covering partner models, model routing, optimization, and Microsoft Foundry development experiences. These resources are a great way to continue exploring new capabilities across the rapidly expanding Microsoft Foundry ecosystem.


You can also join the community through Model Mondays and the Foundry Discord to learn from product teams, watch technical deep dives, and connect with other developers building on Microsoft Foundry.

Microsoft Tech Community originally posted this article on 10 September 2026 at 9:14 PM.

Leave a Reply