GCPSPARKSAIINFRA

AI Infrastructure Essentials Training

This course provides a foundational overview of the hardware, software, and networking components required to develop and manage AI models at scale. It explores Google Cloud's AI Hypercomputer architecture, compares compute accelerators like GPUs and TPUs, and examines the critical data pipelines and storage solutions necessary to maximize training performance.

Google Cloud
✓ Official training Google CloudLevel Intermediate⏱️ 0.5 day (3h)

What you will learn

  • Differentiate between the layers of the AI Hypercomputer.
  • Select appropriate accelerators for the most cost-effective AI workloads.
  • Evaluate storage and networking solutions to maximize training goodput.
  • Compare various deployment and consumption models for resource optimization.

Prerequisites

  • Familiarity with cloud computing concepts and general data center infrastructure.

Target audience

  • IT decision-makers and infrastructure architects looking to understand the technical requirements and the AI Hypercomputer's offerings for enterprise-grade AI deployment.

Training Program

6 modules to master the fundamentals

Topics covered
  • →Definition of AI infrastructure
  • →The evolution of computing demands
  • →The need for new computing power
Objectives
  • Differentiate between the layers of the AI Hypercomputer.
Topics covered
  • →The AI Hypercomputer
  • →The 3 layers of the AI Hypercomputer: Overview
Topics covered
  • →Graphics Processing Units (GPU architecture, Google Cloud GPU family, Selecting GPUs)
  • →Tensor Processing Units (TPU architecture, Google Cloud TPU family, Best practices and considerations)
Activities

1x exercise/discussion

Objectives
  • Evaluate storage and networking solutions to maximize training goodput.
Topics covered
  • →Maximizing goodput
  • →Networking for data ingestion and training
  • →Storage for data preparation and training
  • →Architecture for inference
Activities

1x discussion

Objectives
  • Compare various deployment and consumption models for resource optimization.
Topics covered
  • →Deployment options
  • →Flexible consumption
Objectives
  • Differentiate between the layers of the AI Hypercomputer.
  • Select appropriate accelerators for the most cost-effective AI workloads.
  • Evaluate storage and networking solutions to maximize training goodput.
  • Compare various deployment and consumption models for resource optimization.
Topics covered
  • →Course summary
  • →Q&A
  • →Quiz
Activities

1x quiz with 4 MCQs

Related Trainings

Upcoming sessions

No date suits you?

We regularly organize new sessions. Contact us to find out about upcoming dates or to schedule a session at a date of your choice.

Register for a custom date

Quality Process

SFEIR Institute's commitment: an excellence approach to ensure the quality and success of all our training programs. Learn more about our quality approach

Teaching Methods Used
  • Lectures / Theoretical Slides — Presentation of concepts using visual aids (PowerPoint, PDF).
  • Technical Demonstration (Demos) — The instructor performs a task or procedure while students observe.
  • Quiz / MCQ — Quick knowledge check (paper-based or digital via tools like Kahoot/Klaxoon).
Evaluation and Monitoring System

The achievement of training objectives is evaluated at multiple levels to ensure quality:

  • Continuous Knowledge Assessment : Verification of knowledge throughout the training via participatory methods (quizzes, practical exercises, case studies) under instructor supervision.
  • Progress Measurement : Comparative self-assessment system including an initial diagnostic to determine the starting level, followed by a final evaluation to validate skills development.
  • Quality Evaluation : End-of-session satisfaction questionnaire to measure the relevance and effectiveness of the training as perceived by participants.

Frequently Asked Questions

It is designed for cloud architects, ML engineers and technical practitioners who design or manage the infrastructure needed to develop and run AI models at scale.
Basic familiarity with cloud computing is helpful. This is a level 200 course, so some technical background makes the content easier to follow.
You will learn to differentiate the layers of the AI Hypercomputer, select accelerators cost-effectively, evaluate storage and networking to maximize training goodput, and compare deployment and consumption models.
Yes. A dedicated module on compute accelerators compares GPUs and TPUs so you can choose the most cost-effective option for a given AI workload.
The course lasts 3 hours and is delivered in an instructor-led format, alternating theory and demonstrations.
Yes. This is an official Google Cloud Sparks course delivered by SFEIR Institute, a certified Google Cloud Training Partner.

395€ excl. VAT

per learner