This advanced course develops the skills to design, optimize, and integrate custom AI inference accelerator kernels using Altera HLS Compiler Pro, the Altera FPGA SDK for OpenCL, and RTL/SystemVerilog on Altera FPGAs. Participants work at the lowest level of the acceleration stack: implementing Conv2D, GEMM, depth-wise separable convolution, multi-head attention, and activation function kernels from scratch; applying advanced HLS optimization techniques (pipelining, unrolling, dataflow, memory partitioning); designing systolic array architectures; implementing Transformer and attention mechanisms in hardware; and integrating all components into complete Platform Designer systems through AXI4 and AXI4-Stream interfaces.