Get in Touch

Course Outline

Introduction to Multimodal LLMs in Vertex AI

  • Explore the multimodal capabilities available in Vertex AI
  • Review Gemini models and their supported modalities
  • Examine enterprise and research use cases

Setting Up the Development Environment

  • Configure Vertex AI specifically for multimodal workflows
  • Manage datasets across different modalities
  • Hands-on lab: Setting up the environment and preparing datasets

Long Context Windows and Advanced Reasoning

  • Gain insight into long-context workflow mechanics
  • Apply concepts to planning and decision-making scenarios
  • Hands-on lab: Implementing long-context analysis

Cross-Modal Workflow Design

  • Synthesize text, audio, and image analysis techniques
  • Chain multimodal steps effectively within pipelines
  • Hands-on lab: Designing a comprehensive multimodal pipeline

Working with Gemini API Parameters

  • Configure multimodal inputs and outputs
  • Optimize inference speed and operational efficiency
  • Hands-on lab: Tuning Gemini API parameters

Advanced Applications and Integrations

  • Create interactive multimodal agents and assistants
  • Integrate external APIs and tools into the workflow
  • Hands-on lab: Building a functional multimodal application

Evaluation and Iteration

  • Test the performance of multimodal models
  • Apply metrics for accuracy, alignment, and drift detection
  • Hands-on lab: Evaluating multimodal workflows

Summary and Next Steps

Requirements

  • Solid proficiency in Python programming
  • Background in machine learning model development
  • Understanding of multimodal data types (text, audio, image)

Target Audience

  • AI researchers
  • Advanced developers
  • ML scientists
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories