Programming With CUDA Print

  • specialisedtechnology, specialised, performance, guide, howto, solution, zillionkinghost, hosting
  • 0

Vendor-specific parallel computing.

WHAT IT IS

A platform for general computation on one vendor's graphics processors.

WHY IT DOMINATES

Mature tooling, extensive libraries, and the ecosystem built around it.

WHAT THE MODEL IS

Kernels executed by many threads, organised into blocks and grids.

WHAT A BLOCK PROVIDES

Threads that can share fast memory and synchronise with each other.

WHAT THREADS IN DIFFERENT BLOCKS CANNOT DO

Synchronise, within a kernel.

WHAT TO CHOOSE CAREFULLY

Block size, which affects how well the hardware is occupied.

WHAT THE LIBRARIES PROVIDE

Linear algebra Fourier transforms Deep learning primitives Sorting and scanning

WHAT TO USE FIRST

Those libraries.

WHY

They are heavily optimised, and hand-written kernels rarely match them.

WHAT TO PROFILE WITH

The vendor's profiling tools, reporting occupancy, memory throughput and divergence.

WHAT COMMONLY LIMITS PERFORMANCE

Memory access patterns Low occupancy Transfers dominating

WHAT TO MEASURE

End-to-end time, including transfers.

WHAT THE LOCK-IN IS

Code written for it runs only on that vendor's hardware.


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot