GPU Programming in Swift Using Metal 4
An in-depth, hands-on guide to mastering GPU computing on Apple Silicon.
100% handwritten. No LLM slop or vibe coded garbage included.
#include <metal_stdlib>
using namespace metal;
kernel void parallelComputeKernel(
device const float* inData [[buffer(0)]],
device float* outData [[buffer(1)]],
uint threadPositionInGrid [[thread_position_in_grid]]) {
outData[tid] = inData[tid] * 2.0f;
}
Book Overview
About the Book
GPU Programming in Swift Using Metal 4 provides an introduction to GPU programming using Apple's Metal framework. It begins with an overview of the relevant parts of Metal 4 that are used for the parallel computations in the later chapters. It then moves on to a series of examples that demonstrate parallelizing
computations on the GPU using Metal 4.
Detailed Syllabus
Table of Contents
Metal 4
- Kernels
- Compilation
- Metal IR
- Thread Organization
- Thread Grid
- Threadgroup
- SIMD-Groups
- Thread Dispatching
- Grids (1D, 2D, 3D)
- Threadgroups (1D, 2D, 3D, Dispatching Threads vs Threadgroups)
- Grid and Threadgroup Sizing
- Nonuniform Threadgroups
- SIMD-Groups
- Memory
- Unified Memory
- Address Spaces (Device, Thread, Threadgroup, Constant)
- Storage Modes (Shared, Private)
- Heaps
- Memory Bandwidth & Limits
- Random Device Memory Access
- Thread Coordination
- Threadgroup Synchronization
- Synchronization Between Multiple Compute Passes
- Fences, Barriers, and Events
- Indirect Command Buffers
Mandelbrot Set
- Visualizing the Mandelbrot Set
- Sequential Mandelbrot Set
- Parallel Mandelbrot Set
K-Means Clustering
- Sequential K-Means
- Parallel K-Means
Conway’s Game of Life
- Tile-Based Sparse Simulation
- Tile-Based Sparse Simulation Output
- Tile-Based Sparse Simulation With an Indirect Command Buffer
- Tile-Based Sparse Simulation With an Indirect Command Buffer Output
Gray-Scott Reaction Diffusion Model
- Gray-Scott Simulation
- Gray-Scott Simulation Output
- Gray-Scott Multi-Environment Simulation
- Gray-Scott Multi-Environment Simulation Output
- When To Use Function Pointers
- SIMD Divergence
Array Max
- Array Max Threadgroup Memory
- Array Max SIMD-Groups
Summations
- Generating Data on the GPU
- Single Threadgroup Kernels
- Storing Partial Sums in Device Memory
- Storing Partial Sums in Threadgroup Memory
- Threadgroup-Stride
- Atomic Float Output Buffer
- Multi-Threadgroup Kernels
- Storing Partial Sums in Device Memory
- Atomic Float Output Buffer
- SIMD Reductions
Sorting
- Odd-Even Mergesort
- Odd-Even Mergesort Device Memory
- Odd-Even Mergesort Threadgroup Memory
- Bitonic Mergesort
- Bitonic Mergesort Device Memory
- Bitonic Mergesort Threadgroup Memory
Matrix Transpose
- Matrix Transpose Tiled Threadgroup Memory
Matrix Multiplication
- Sequential Matrix Multiplication
- Parallel Matrix Multiplication
- Naive Matrix Multiplication & Performance
- Tiled Threadgroup Memory Matrix Multiplication
- Tiled SIMD Operations Matrix Multiplication
Systolic Arrays
- Threadgroup Memory GEMM
- SIMD-Group GEMM
Fast Fourier Transform
- Image Filtering Pipeline
- Discrete Fourier Transform
- Cooley-Tukey Algorithm
- Cooley-Tukey Threadgroup Memory
- Cooley-Tukey SIMD
- Stockham Algorithm
- Radix-2 Stockham
- Radix-4 Stockham
Attention
- Single-head Attention
- Multi-head Attention
- Multi-head Online Softmax
Profiling
- Profiling Basics
- Capturing GPU Information
- Calculating TFLOPS
Get Instant Access Today
Includes the DRM-free 244-page ebook, with links to the complete Metal 4 sample code.
Buy Now — $34.99