Loading…
Loading…
Combine vision and language AI. Build applications that understand and generate content across multiple data types.
Created by Alex Rivera · Staff Engineer, LLM Systems
30-day money-back guarantee
This course includes
The next wave of AI isn't limited to text. Multimodal AI bridges the gap between different data formats like images, audio, and text, enabling richer, more intuitive interactions. This course provides a comprehensive introduction to the concepts, architectures, and practical applications of multimodal AI. You'll explore state-of-the-art models and techniques that allow AI to 'see' and 'understand' visual information, then reason about it using natural language. We'll cover key areas such as image captioning, visual question answering (VQA), and text-to-image generation, demonstrating how these capabilities can be integrated into real-world products. Get hands-on experience with leading libraries and APIs. This intermediate-level course is designed for developers and AI enthusiasts eager to build the next generation of intelligent applications that interact seamlessly with the world around us.
5 modules · 15 lessons · 4h 17m
Alex Rivera
Staff Engineer, LLM Systems
Alex builds and operates retrieval and agent systems in production. He writes about evaluation, latency and the unglamorous parts of shipping LLM applications.
30-day money-back guarantee
This course includes