GPU Programming in Swift Using Metal 4 Get Book
★ ★ ★ ★ ★ | Be the first to review (0 reviews) | Metal 4 & Swift

GPU Programming in Swift Using Metal 4

An in-depth, hands-on guide to mastering GPU computing on Apple Silicon.

100% handwritten. No LLM slop or vibe coded garbage included.

ParallelKernel.metal
Metal 4 / MSL 3.1
#include <metal_stdlib>
using namespace metal;

kernel void parallelComputeKernel(
    device const float* inData  [[buffer(0)]],
    device float* outData [[buffer(1)]],
    uint threadPositionInGrid [[thread_position_in_grid]]) {
    outData[tid] = inData[tid] * 2.0f;
}
Book Overview

About the Book

GPU Programming in Swift Using Metal 4 provides an introduction to GPU programming using Apple's Metal framework. It begins with an overview of the relevant parts of Metal 4 that are used for the parallel computations in the later chapters. It then moves on to a series of examples that demonstrate parallelizing computations on the GPU using Metal 4.
Detailed Syllabus

Table of Contents

Metal 4

  • Kernels
    • Compilation
    • Metal IR
  • Thread Organization
    • Thread Grid
    • Threadgroup
    • SIMD-Groups
  • Thread Dispatching
    • Grids (1D, 2D, 3D)
    • Threadgroups (1D, 2D, 3D, Dispatching Threads vs Threadgroups)
    • Grid and Threadgroup Sizing
    • Nonuniform Threadgroups
    • SIMD-Groups
  • Memory
    • Unified Memory
    • Address Spaces (Device, Thread, Threadgroup, Constant)
    • Storage Modes (Shared, Private)
    • Heaps
    • Memory Bandwidth & Limits
    • Random Device Memory Access
  • Thread Coordination
    • Threadgroup Synchronization
    • Synchronization Between Multiple Compute Passes
    • Fences, Barriers, and Events
  • Indirect Command Buffers

Mandelbrot Set

  • Visualizing the Mandelbrot Set
  • Sequential Mandelbrot Set
  • Parallel Mandelbrot Set

K-Means Clustering

  • Sequential K-Means
  • Parallel K-Means

Conway’s Game of Life

  • Tile-Based Sparse Simulation
  • Tile-Based Sparse Simulation Output
  • Tile-Based Sparse Simulation With an Indirect Command Buffer
  • Tile-Based Sparse Simulation With an Indirect Command Buffer Output

Gray-Scott Reaction Diffusion Model

  • Gray-Scott Simulation
  • Gray-Scott Simulation Output
  • Gray-Scott Multi-Environment Simulation
  • Gray-Scott Multi-Environment Simulation Output
  • When To Use Function Pointers
  • SIMD Divergence

Array Max

  • Array Max Threadgroup Memory
  • Array Max SIMD-Groups

Summations

  • Generating Data on the GPU
  • Single Threadgroup Kernels
    • Storing Partial Sums in Device Memory
    • Storing Partial Sums in Threadgroup Memory
    • Threadgroup-Stride
    • Atomic Float Output Buffer
  • Multi-Threadgroup Kernels
    • Storing Partial Sums in Device Memory
    • Atomic Float Output Buffer
    • SIMD Reductions

Sorting

  • Odd-Even Mergesort
    • Odd-Even Mergesort Device Memory
    • Odd-Even Mergesort Threadgroup Memory
  • Bitonic Mergesort
    • Bitonic Mergesort Device Memory
    • Bitonic Mergesort Threadgroup Memory

Matrix Transpose

  • Matrix Transpose Tiled Threadgroup Memory

Matrix Multiplication

  • Sequential Matrix Multiplication
  • Parallel Matrix Multiplication
    • Naive Matrix Multiplication & Performance
    • Tiled Threadgroup Memory Matrix Multiplication
    • Tiled SIMD Operations Matrix Multiplication

Systolic Arrays

  • Threadgroup Memory GEMM
  • SIMD-Group GEMM

Fast Fourier Transform

  • Image Filtering Pipeline
  • Discrete Fourier Transform
  • Cooley-Tukey Algorithm
    • Cooley-Tukey Threadgroup Memory
    • Cooley-Tukey SIMD
  • Stockham Algorithm
    • Radix-2 Stockham
    • Radix-4 Stockham

Attention

  • Single-head Attention
  • Multi-head Attention
  • Multi-head Online Softmax

Profiling

  • Profiling Basics
  • Capturing GPU Information
  • Calculating TFLOPS

Get Instant Access Today

Includes the DRM-free 244-page ebook, with links to the complete Metal 4 sample code.

Buy Now — $34.99