All interviews
Google logo

Google

Mid

Optimize Google's GPU software stack from ML compiler cost models down to kernel-level performance

Join Google's Core ML organization to build optimizations for the latest generation of GPUs powering Google's ML infrastructure at massive scale. This role spans the full GPU software stack — from compiler cost model design to high-performance kernel tuning to cross-node model serving configuration.

Practice this interview

Free · a live voice mock calibrated to this exact role

Start the mock interview

What this interview tests

  • Low-level GPU programming (CUDA/Triton/CUTLASS)
  • GPU performance engineering and profiling
  • GPU memory hierarchy and architecture
  • ML compiler optimization (OpenXLA/MLIR)
  • LLM deployment on accelerators

Common question themes

Walk through profiling and fixing a GPU kernel bottleneck

Memory-bound vs. compute-bound kernel diagnosis

How a compiler lowers/schedules ops for a GPU target

Tradeoffs in kernel tuning (occupancy, register pressure, tiling)

Cross-node serving considerations for large models

How candidates describe it

Real Software Engineer interview stories — retold from candidates' public write-ups, with sources.

View the original posting

All Google Software Engineer interviews

All Google interviews

Related interviews