Key Takeaway

The first thing a DMA interview question checks is the concept itself. DMA lets a DMA controller handle the transfer between a peripheral and memory, so the CPU only has to handle setup and completion. In an interview, the key is to clearly explain the difference from CPU copy and to also cover cache coherency.

This post organizes the core concepts and answer points for DMA from the perspective of preparing for embedded/firmware job interviews.

What is DMA, and how does it differ from the CPU copying data directly?

One-line answer: A DMA controller performs the transfer between a peripheral and memory instead of the CPU, and the CPU only handles setup and completion.

  1. DMA definition: a DMA controller performs the transfer between a peripheral and memory (or memory and memory) instead of the CPU, and the CPU only handles setup and completion.
  2. CPU copy (PIO/memcpy): the CPU reads and writes one word at a time directly, so the CPU stays occupied for the entire transfer.
  3. DMA flow: once the source, destination, and length (transfer size) are configured, the DMA controller performs the transfer and notifies the CPU with a completion interrupt. For interrupt handling itself, see Interrupt vs Polling.

Data path of CPU copy vs DMA: CPU copy goes device → CPU → RAM and keeps the CPU busy; DMA goes device → DMA controller → RAM while the CPU does other work and receives a completion IRQ

  1. Interview opening-line template: “Instead of the CPU, a DMA controller moves data between a peripheral and memory, and the CPU only handles setup and the completion interrupt, freeing the CPU up for other work.”

When is DMA the right fit for large, continuous transfers, and when is CPU copy better for short transfers?

One-line answer: DMA is better suited for large, continuous transfers, while CPU copy can be better for short transfers where setup/completion overhead is smaller.

  1. DMA is favorable for large, continuous block transfers, audio/ADC streams, and storage/network buffer transfers.
  2. CPU copy is favorable for short transfers of just a few bytes, where the DMA setup and completion overhead can outweigh the transfer itself.
  3. Terminology: scatter-gather links several scattered buffers together into a single DMA operation, and bus arbitration refers to arbitrating access order when DMA and the CPU share the same bus/memory — during which the CPU may also have to wait briefly.
  4. The table below compares CPU copy and DMA.
ItemCPU CopyDMA
CPU occupancyEntire transferSetup/completion only
Setup costAlmost noneChannel/descriptor setup required
Suitable sizeShort transfersLarge, continuous transfers
ComplexitySimpleMust consider cache, alignment, completion handling

How do you address cache coherency and buffer alignment in an interview?

One-line answer: Clean (flush) the cache before sending to a device, and invalidate it after receiving from a device, so the CPU and memory stay in sync.

  1. The problem: the contents of the CPU cache and memory can diverge, which can cause DMA to read stale data from memory or the CPU to see stale data in its cache. This is a general principle, and SoCs with hardware cache coherency support may behave differently.
  2. Direction-specific principle: when sending to a device (TX, memory → device), clean (flush) the cache before starting DMA so memory reflects the latest data; when receiving from a device (RX, device → memory), invalidate the relevant cache region after DMA completes and before the CPU reads it.

DMA and cache: for TX, clean (flush) the cache then start DMA; for RX, invalidate the cache after DMA completes before the CPU reads; align the DMA buffer to cache-line boundaries

  1. Buffer alignment: DMA buffers should be aligned to, and sized to, cache-line boundaries. Otherwise they can share a cache line with neighboring data, and CPU writes and DMA writes can overwrite each other and corrupt the data.
  2. On Linux, the kernel DMA API (dma_map_single / dma_unmap_single, dma_sync_single_for_cpu / dma_sync_single_for_device, direction flags DMA_TO_DEVICE / DMA_FROM_DEVICE, and dma_alloc_coherent for coherent memory) handles this synchronization. See the kernel.org Dynamic DMA mapping Guide for details.
  3. Interview closing-answer framework: “Start with the definition and the difference in CPU occupancy, then cover which is appropriate based on transfer size, and finish with cache coherency (TX clean, RX invalidate) and buffer alignment as a single flow.”

Interview one-liner: A DMA controller performs transfers on the CPU’s behalf to reduce CPU occupancy, but implementations must also account for cache coherency synchronization and buffer alignment.

What are the frequently confused questions (FAQ)?

Q. If you use DMA, does the CPU rest completely?

No. Channel setup before the transfer and completion-interrupt handling after it are still required, and bus arbitration can make the CPU wait briefly as well.

Q. If there is cache-coherency hardware, is flush/invalidate unnecessary?

It depends on the SoC. Some hardware does support cache coherency, but on Linux the kernel’s DMA API usually handles this synchronization appropriately for the environment.