The course assignment was to implement a 128-point 1D FFT on the Basys3 FPGA board. One constraint shaped everything else: every multiplication had to be performed using Xilinx DSP48 IP blocks — not LUT-based multipliers. This is actually the correct way to do it. DSP48 slices are faster, more power-efficient, and have better timing characteristics than equivalent LUT multiplications. The constraint taught the right habit.
The problem: Basys3 carries an Artix-7 xc7a35t with exactly 90 DSP48 slices. A 128-point Radix-2 FFT requires N/2 × log₂(N) = 64 × 7 = 448 complex multiplications per transform. With the DSP48 constraint and a naive fully-parallel implementation, that means 448 DSP instantiations. The board has 90. The obvious implementation doesn't fit.