Key Takeaway
The first thing a DMA interview question checks is the concept itself. DMA lets a DMA controller handle the transfer between a peripheral and memory, so the CPU only has to handle setup and completion. In an interview, the key is to clearly explain the difference from CPU copy and to also cover cache coherency.
This post organizes the core concepts and answer points for DMA from the perspective of preparing for embedded/firmware job interviews.
What is DMA, and how does it differ from the CPU copying data directly?
One-line answer: A DMA controller performs the transfer between a peripheral and memory instead of the CPU, and the CPU only handles setup and completion.
- DMA definition: a DMA controller performs the transfer between a peripheral and memory (or memory and memory) instead of the CPU, and the CPU only handles setup and completion.
- CPU copy (PIO/memcpy): the CPU reads and writes one word at a time directly, so the CPU stays occupied for the entire transfer.
- DMA flow: once the source, destination, and length (transfer size) are configured, the DMA controller performs the transfer and notifies the CPU with a completion interrupt. For interrupt handling itself, see Interrupt vs Polling.

- Interview opening-line template: “Instead of the CPU, a DMA controller moves data between a peripheral and memory, and the CPU only handles setup and the completion interrupt, freeing the CPU up for other work.”
When is DMA the right fit for large, continuous transfers, and when is CPU copy better for short transfers?
One-line answer: DMA is better suited for large, continuous transfers, while CPU copy can be better for short transfers where setup/completion overhead is smaller.
- DMA is favorable for large, continuous block transfers, audio/ADC streams, and storage/network buffer transfers.
- CPU copy is favorable for short transfers of just a few bytes, where the DMA setup and completion overhead can outweigh the transfer itself.
- Terminology: scatter-gather links several scattered buffers together into a single DMA operation, and bus arbitration refers to arbitrating access order when DMA and the CPU share the same bus/memory — during which the CPU may also have to wait briefly.
- The table below compares CPU copy and DMA.
| Item | CPU Copy | DMA |
|---|---|---|
| CPU occupancy | Entire transfer | Setup/completion only |
| Setup cost | Almost none | Channel/descriptor setup required |
| Suitable size | Short transfers | Large, continuous transfers |
| Complexity | Simple | Must consider cache, alignment, completion handling |
How do you address cache coherency and buffer alignment in an interview?
One-line answer: Clean (flush) the cache before sending to a device, and invalidate it after receiving from a device, so the CPU and memory stay in sync.
- The problem: the contents of the CPU cache and memory can diverge, which can cause DMA to read stale data from memory or the CPU to see stale data in its cache. This is a general principle, and SoCs with hardware cache coherency support may behave differently.
- Direction-specific principle: when sending to a device (TX, memory → device), clean (flush) the cache before starting DMA so memory reflects the latest data; when receiving from a device (RX, device → memory), invalidate the relevant cache region after DMA completes and before the CPU reads it.

- Buffer alignment: DMA buffers should be aligned to, and sized to, cache-line boundaries. Otherwise they can share a cache line with neighboring data, and CPU writes and DMA writes can overwrite each other and corrupt the data.
- On Linux, the kernel DMA API (
dma_map_single/dma_unmap_single,dma_sync_single_for_cpu/dma_sync_single_for_device, direction flagsDMA_TO_DEVICE/DMA_FROM_DEVICE, anddma_alloc_coherentfor coherent memory) handles this synchronization. See the kernel.org Dynamic DMA mapping Guide for details. - Interview closing-answer framework: “Start with the definition and the difference in CPU occupancy, then cover which is appropriate based on transfer size, and finish with cache coherency (TX clean, RX invalidate) and buffer alignment as a single flow.”
Interview one-liner: A DMA controller performs transfers on the CPU’s behalf to reduce CPU occupancy, but implementations must also account for cache coherency synchronization and buffer alignment.
What are the frequently confused questions (FAQ)?
Q. If you use DMA, does the CPU rest completely?
No. Channel setup before the transfer and completion-interrupt handling after it are still required, and bus arbitration can make the CPU wait briefly as well.
Q. If there is cache-coherency hardware, is flush/invalidate unnecessary?
It depends on the SoC. Some hardware does support cache coherency, but on Linux the kernel’s DMA API usually handles this synchronization appropriately for the environment.