← Documents Documentation/core-api/dma-api-howto.rst GitHub 원문 ↗

Linux 6.18.37 · Core API

Dynamic DMA mapping Guide

Device driver에서 dynamic DMA mapping을 올바르게 사용하는 데 필요한 주소 공간, mask, coherent/streaming mapping, 동기화, 오류 처리 규칙을 설명합니다.

Source pathDocumentation/core-api/dma-api-howto.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

dma-api-howto.rst:1-935

DMA driver는 CPU virtual address, CPU physical address, device bus address를 구분해야 합니다. `dma_map_*()` API는 필요한 IOMMU mapping을 만들고 device가 실제로 사용할 `dma_addr_t`를 반환하므로 bus 전용 API 대신 공통 DMA API를 사용해야 합니다.

Coherent mapping은 CPU와 device가 서로의 갱신을 즉시 볼 수 있는 장기 공유 영역에, streaming mapping은 개별 transfer에 적합합니다. Coherent memory도 CPU store ordering을 위한 memory barrier가 필요하며 streaming buffer를 transfer 사이에 CPU가 만지면 `dma_sync_*_for_cpu()`와 `dma_sync_*_for_device()`로 ownership을 전환해야 합니다.

모든 mapping 결과는 오류를 검사하고 반드시 대응하는 unmap을 수행해야 합니다. DMA mask는 device 능력에 맞게 probe하고, scatterlist의 `nents`에는 mapping 입력 개수와 반환된 segment 개수의 차이를 엄격히 지켜야 합니다. 이 규칙을 어기면 address space 고갈, panic, 조용한 data corruption으로 이어질 수 있습니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 =========================
2 Dynamic DMA mapping Guide
3 =========================
4
5 :Author: David S. Miller <[email protected]>
6 :Author: Richard Henderson <[email protected]>
7 :Author: Jakub Jelinek <[email protected]>
8
9 This is a guide to device driver writers on how to use the DMA API
10 with example pseudo-code. For a concise description of the API, see
11 Documentation/core-api/dma-api.rst.
12
13 CPU and DMA addresses
14 =====================
15
16 There are several kinds of addresses involved in the DMA API, and it's
17 important to understand the differences.
18
19 The kernel normally uses virtual addresses. Any address returned by
20 kmalloc(), vmalloc(), and similar interfaces is a virtual address and can
21 be stored in a ``void *``.
22
23 The virtual memory system (TLB, page tables, etc.) translates virtual
24 addresses to CPU physical addresses, which are stored as "phys_addr_t" or
25 "resource_size_t". The kernel manages device resources like registers as
26 physical addresses. These are the addresses in /proc/iomem. The physical
27 address is not directly useful to a driver; it must use ioremap() to map
28 the space and produce a virtual address.
29
30 I/O devices use a third kind of address: a "bus address". If a device has
31 registers at an MMIO address, or if it performs DMA to read or write system
32 memory, the addresses used by the device are bus addresses. In some
33 systems, bus addresses are identical to CPU physical addresses, but in
34 general they are not. IOMMUs and host bridges can produce arbitrary
35 mappings between physical and bus addresses.
36
37 From a device's point of view, DMA uses the bus address space, but it may
38 be restricted to a subset of that space. For example, even if a system
39 supports 64-bit addresses for main memory and PCI BARs, it may use an IOMMU
40 so devices only need to use 32-bit DMA addresses.
41
42 Here's a picture and some examples::
43
44 CPU CPU Bus
45 Virtual Physical Address
46 Address Address Space
47 Space Space
48
49 +-------+ +------+ +------+
50 | | |MMIO | Offset | |
51 | | Virtual |Space | applied | |
52 C +-------+ --------> B +------+ ----------> +------+ A
53 | | mapping | | by host | |
54 +-----+ | | | | bridge | | +--------+
55 | | | | +------+ | | | |
56 | CPU | | | | RAM | | | | Device |
57 | | | | | | | | | |
58 +-----+ +-------+ +------+ +------+ +--------+
59 | | Virtual |Buffer| Mapping | |
60 X +-------+ --------> Y +------+ <---------- +------+ Z
61 | | mapping | RAM | by IOMMU
62 | | | |
63 | | | |
64 +-------+ +------+
65
66 During the enumeration process, the kernel learns about I/O devices and
67 their MMIO space and the host bridges that connect them to the system. For
68 example, if a PCI device has a BAR, the kernel reads the bus address (A)
69 from the BAR and converts it to a CPU physical address (B). The address B
70 is stored in a struct resource and usually exposed via /proc/iomem. When a
71 driver claims a device, it typically uses ioremap() to map physical address
72 B at a virtual address (C). It can then use, e.g., ioread32(C), to access
73 the device registers at bus address A.
74
75 If the device supports DMA, the driver sets up a buffer using kmalloc() or
76 a similar interface, which returns a virtual address (X). The virtual
77 memory system maps X to a physical address (Y) in system RAM. The driver
78 can use virtual address X to access the buffer, but the device itself
79 cannot because DMA doesn't go through the CPU virtual memory system.
80
81 In some simple systems, the device can do DMA directly to physical address
82 Y. But in many others, there is IOMMU hardware that translates DMA
83 addresses to physical addresses, e.g., it translates Z to Y. This is part
84 of the reason for the DMA API: the driver can give a virtual address X to
85 an interface like dma_map_single(), which sets up any required IOMMU
86 mapping and returns the DMA address Z. The driver then tells the device to
87 do DMA to Z, and the IOMMU maps it to the buffer at address Y in system
88 RAM.
89
90 So that Linux can use the dynamic DMA mapping, it needs some help from the
91 drivers, namely it has to take into account that DMA addresses should be
92 mapped only for the time they are actually used and unmapped after the DMA
93 transfer.
94
95 The following API will work of course even on platforms where no such
96 hardware exists.
97
98 Note that the DMA API works with any bus independent of the underlying
99 microprocessor architecture. You should use the DMA API rather than the
100 bus-specific DMA API, i.e., use the dma_map_*() interfaces rather than the
101 pci_map_*() interfaces.
102
103 First of all, you should make sure::
104
105 #include <linux/dma-mapping.h>
106
107 is in your driver, which provides the definition of dma_addr_t. This type
108 can hold any valid DMA address for the platform and should be used
109 everywhere you hold a DMA address returned from the DMA mapping functions.
110
111 What memory is DMA'able?
112 ========================
113
114 The first piece of information you must know is what kernel memory can
115 be used with the DMA mapping facilities. There has been an unwritten
116 set of rules regarding this, and this text is an attempt to finally
117 write them down.
118
119 If you acquired your memory via the page allocator
120 (i.e. __get_free_page*()) or the generic memory allocators
121 (i.e. kmalloc() or kmem_cache_alloc()) then you may DMA to/from
122 that memory using the addresses returned from those routines.
123
124 This means specifically that you may _not_ use the memory/addresses
125 returned from vmalloc() for DMA. It is possible to DMA to the
126 _underlying_ memory mapped into a vmalloc() area, but this requires
127 walking page tables to get the physical addresses, and then
128 translating each of those pages back to a kernel address using
129 something like __va(). [ EDIT: Update this when we integrate
130 Gerd Knorr's generic code which does this. ]
131
132 This rule also means that you may use neither kernel image addresses
133 (items in data/text/bss segments), nor module image addresses, nor
134 stack addresses for DMA. These could all be mapped somewhere entirely
135 different than the rest of physical memory. Even if those classes of
136 memory could physically work with DMA, you'd need to ensure the I/O
137 buffers were cacheline-aligned. Without that, you'd see cacheline
138 sharing problems (data corruption) on CPUs with DMA-incoherent caches.
139 (The CPU could write to one word, DMA would write to a different one
140 in the same cache line, and one of them could be overwritten.)
141
142 Also, this means that you cannot take the return of a kmap()
143 call and DMA to/from that. This is similar to vmalloc().
144
145 What about block I/O and networking buffers? The block I/O and
146 networking subsystems make sure that the buffers they use are valid
147 for you to DMA from/to.
148
149 DMA addressing capabilities
150 ===========================
151
152 By default, the kernel assumes that your device can address 32-bits of DMA
153 addressing. For a 64-bit capable device, this needs to be increased, and for
154 a device with limitations, it needs to be decreased.
155
156 Special note about PCI: PCI-X specification requires PCI-X devices to support
157 64-bit addressing (DAC) for all transactions. And at least one platform (SGI
158 SN2) requires 64-bit coherent allocations to operate correctly when the IO
159 bus is in PCI-X mode.
160
161 For correct operation, you must set the DMA mask to inform the kernel about
162 your devices DMA addressing capabilities.
163
164 This is performed via a call to dma_set_mask_and_coherent()::
165
166 int dma_set_mask_and_coherent(struct device *dev, u64 mask);
167
168 which will set the mask for both streaming and coherent APIs together. If you
169 have some special requirements, then the following two separate calls can be
170 used instead:
171
172 The setup for streaming mappings is performed via a call to
173 dma_set_mask()::
174
175 int dma_set_mask(struct device *dev, u64 mask);
176
177 The setup for coherent allocations is performed via a call
178 to dma_set_coherent_mask()::
179
180 int dma_set_coherent_mask(struct device *dev, u64 mask);
181
182 Here, dev is a pointer to the device struct of your device, and mask is a bit
183 mask describing which bits of an address your device supports. Often the
184 device struct of your device is embedded in the bus-specific device struct of
185 your device. For example, &pdev->dev is a pointer to the device struct of a
186 PCI device (pdev is a pointer to the PCI device struct of your device).
187
188 These calls usually return zero to indicate your device can perform DMA
189 properly on the machine given the address mask you provided, but they might
190 return an error if the mask is too small to be supportable on the given
191 system. If it returns non-zero, your device cannot perform DMA properly on
192 this platform, and attempting to do so will result in undefined behavior.
193 You must not use DMA on this device unless the dma_set_mask family of
194 functions has returned success.
195
196 This means that in the failure case, you have two options:
197
198 1) Use some non-DMA mode for data transfer, if possible.
199 2) Ignore this device and do not initialize it.
200
201 It is recommended that your driver print a kernel KERN_WARNING message when
202 setting the DMA mask fails. In this manner, if a user of your driver reports
203 that performance is bad or that the device is not even detected, you can ask
204 them for the kernel messages to find out exactly why.
205
206 The 24-bit addressing device would do something like this::
207
208 if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(24))) {
209 dev_warn(dev, "mydev: No suitable DMA available\n");
210 goto ignore_this_device;
211 }
212
213 The standard 64-bit addressing device would do something like this::
214
215 dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64))
216
217 dma_set_mask_and_coherent() never return fail when DMA_BIT_MASK(64). Typical
218 error code like::
219
220 /* Wrong code */
221 if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64)))
222 dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32))
223
224 dma_set_mask_and_coherent() will never return failure when bigger than 32.
225 So typical code like::
226
227 /* Recommended code */
228 if (support_64bit)
229 dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
230 else
231 dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32));
232
233 If the device only supports 32-bit addressing for descriptors in the
234 coherent allocations, but supports full 64-bits for streaming mappings
235 it would look like this::
236
237 if (dma_set_mask(dev, DMA_BIT_MASK(64))) {
238 dev_warn(dev, "mydev: No suitable DMA available\n");
239 goto ignore_this_device;
240 }
241
242 The coherent mask will always be able to set the same or a smaller mask as
243 the streaming mask. However for the rare case that a device driver only
244 uses coherent allocations, one would have to check the return value from
245 dma_set_coherent_mask().
246
247 Finally, if your device can only drive the low 24-bits of
248 address you might do something like::
249
250 if (dma_set_mask(dev, DMA_BIT_MASK(24))) {
251 dev_warn(dev, "mydev: 24-bit DMA addressing not available\n");
252 goto ignore_this_device;
253 }
254
255 When dma_set_mask() or dma_set_mask_and_coherent() is successful, and
256 returns zero, the kernel saves away this mask you have provided. The
257 kernel will use this information later when you make DMA mappings.
258
259 There is a case which we are aware of at this time, which is worth
260 mentioning in this documentation. If your device supports multiple
261 functions (for example a sound card provides playback and record
262 functions) and the various different functions have _different_
263 DMA addressing limitations, you may wish to probe each mask and
264 only provide the functionality which the machine can handle. It
265 is important that the last call to dma_set_mask() be for the
266 most specific mask.
267
268 Here is pseudo-code showing how this might be done::
269
270 #define PLAYBACK_ADDRESS_BITS DMA_BIT_MASK(32)
271 #define RECORD_ADDRESS_BITS DMA_BIT_MASK(24)
272
273 struct my_sound_card *card;
274 struct device *dev;
275
276 ...
277 if (!dma_set_mask(dev, PLAYBACK_ADDRESS_BITS)) {
278 card->playback_enabled = 1;
279 } else {
280 card->playback_enabled = 0;
281 dev_warn(dev, "%s: Playback disabled due to DMA limitations\n",
282 card->name);
283 }
284 if (!dma_set_mask(dev, RECORD_ADDRESS_BITS)) {
285 card->record_enabled = 1;
286 } else {
287 card->record_enabled = 0;
288 dev_warn(dev, "%s: Record disabled due to DMA limitations\n",
289 card->name);
290 }
291
292 A sound card was used as an example here because this genre of PCI
293 devices seems to be littered with ISA chips given a PCI front end,
294 and thus retaining the 16MB DMA addressing limitations of ISA.
295
296 Types of DMA mappings
297 =====================
298
299 There are two types of DMA mappings:
300
301 - Coherent DMA mappings which are usually mapped at driver
302 initialization, unmapped at the end and for which the hardware should
303 guarantee that the device and the CPU can access the data
304 in parallel and will see updates made by each other without any
305 explicit software flushing.
306
307 Think of "coherent" as "synchronous".
308
309 The current default is to return coherent memory in the low 32
310 bits of the DMA space. However, for future compatibility you should
311 set the coherent mask even if this default is fine for your
312 driver.
313
314 Good examples of what to use coherent mappings for are:
315
316 - Network card DMA ring descriptors.
317 - SCSI adapter mailbox command data structures.
318 - Device firmware microcode executed out of
319 main memory.
320
321 The invariant these examples all require is that any CPU store
322 to memory is immediately visible to the device, and vice
323 versa. Coherent mappings guarantee this.
324
325 .. important::
326
327 Coherent DMA memory does not preclude the usage of
328 proper memory barriers. The CPU may reorder stores to
329 coherent memory just as it may normal memory. Example:
330 if it is important for the device to see the first word
331 of a descriptor updated before the second, you must do
332 something like::
333
334 desc->word0 = address;
335 wmb();
336 desc->word1 = DESC_VALID;
337
338 in order to get correct behavior on all platforms.
339
340 Also, on some platforms your driver may need to flush CPU write
341 buffers in much the same way as it needs to flush write buffers
342 found in PCI bridges (such as by reading a register's value
343 after writing it).
344
345 - Streaming DMA mappings which are usually mapped for one DMA
346 transfer, unmapped right after it (unless you use dma_sync_* below)
347 and for which hardware can optimize for sequential accesses.
348
349 Think of "streaming" as "asynchronous" or "outside the coherency
350 domain".
351
352 Good examples of what to use streaming mappings for are:
353
354 - Networking buffers transmitted/received by a device.
355 - Filesystem buffers written/read by a SCSI device.
356
357 The interfaces for using this type of mapping were designed in
358 such a way that an implementation can make whatever performance
359 optimizations the hardware allows. To this end, when using
360 such mappings you must be explicit about what you want to happen.
361
362 Neither type of DMA mapping has alignment restrictions that come from
363 the underlying bus, although some devices may have such restrictions.
364 Also, systems with caches that aren't DMA-coherent will work better
365 when the underlying buffers don't share cache lines with other data.
366
367
368 Using Coherent DMA mappings
369 ===========================
370
371 To allocate and map large (PAGE_SIZE or so) coherent DMA regions,
372 you should do::
373
374 dma_addr_t dma_handle;
375
376 cpu_addr = dma_alloc_coherent(dev, size, &dma_handle, gfp);
377
378 where device is a ``struct device *``. This may be called in interrupt
379 context with the GFP_ATOMIC flag.
380
381 Size is the length of the region you want to allocate, in bytes.
382
383 This routine will allocate RAM for that region, so it acts similarly to
384 __get_free_pages() (but takes size instead of a page order). If your
385 driver needs regions sized smaller than a page, you may prefer using
386 the dma_pool interface, described below.
387
388 The coherent DMA mapping interfaces, will by default return a DMA address
389 which is 32-bit addressable. Even if the device indicates (via the DMA mask)
390 that it may address the upper 32-bits, coherent allocation will only
391 return > 32-bit addresses for DMA if the coherent DMA mask has been
392 explicitly changed via dma_set_coherent_mask(). This is true of the
393 dma_pool interface as well.
394
395 dma_alloc_coherent() returns two values: the virtual address which you
396 can use to access it from the CPU and dma_handle which you pass to the
397 card.
398
399 The CPU virtual address and the DMA address are both
400 guaranteed to be aligned to the smallest PAGE_SIZE order which
401 is greater than or equal to the requested size. This invariant
402 exists (for example) to guarantee that if you allocate a chunk
403 which is smaller than or equal to 64 kilobytes, the extent of the
404 buffer you receive will not cross a 64K boundary.
405
406 To unmap and free such a DMA region, you call::
407
408 dma_free_coherent(dev, size, cpu_addr, dma_handle);
409
410 where dev, size are the same as in the above call and cpu_addr and
411 dma_handle are the values dma_alloc_coherent() returned to you.
412 This function may not be called in interrupt context.
413
414 If your driver needs lots of smaller memory regions, you can write
415 custom code to subdivide pages returned by dma_alloc_coherent(),
416 or you can use the dma_pool API to do that. A dma_pool is like
417 a kmem_cache, but it uses dma_alloc_coherent(), not __get_free_pages().
418 Also, it understands common hardware constraints for alignment,
419 like queue heads needing to be aligned on N byte boundaries.
420
421 Create a dma_pool like this::
422
423 struct dma_pool *pool;
424
425 pool = dma_pool_create(name, dev, size, align, boundary);
426
427 The "name" is for diagnostics (like a kmem_cache name); dev and size
428 are as above. The device's hardware alignment requirement for this
429 type of data is "align" (which is expressed in bytes, and must be a
430 power of two). If your device has no boundary crossing restrictions,
431 pass 0 for boundary; passing 4096 says memory allocated from this pool
432 must not cross 4KByte boundaries (but at that time it may be better to
433 use dma_alloc_coherent() directly instead).
434
435 Allocate memory from a DMA pool like this::
436
437 cpu_addr = dma_pool_alloc(pool, flags, &dma_handle);
438
439 flags are GFP_KERNEL if blocking is permitted (not in_interrupt nor
440 holding SMP locks), GFP_ATOMIC otherwise. Like dma_alloc_coherent(),
441 this returns two values, cpu_addr and dma_handle.
442
443 Free memory that was allocated from a dma_pool like this::
444
445 dma_pool_free(pool, cpu_addr, dma_handle);
446
447 where pool is what you passed to dma_pool_alloc(), and cpu_addr and
448 dma_handle are the values dma_pool_alloc() returned. This function
449 may be called in interrupt context.
450
451 Destroy a dma_pool by calling::
452
453 dma_pool_destroy(pool);
454
455 Make sure you've called dma_pool_free() for all memory allocated
456 from a pool before you destroy the pool. This function may not
457 be called in interrupt context.
458
459 DMA Direction
460 =============
461
462 The interfaces described in subsequent portions of this document
463 take a DMA direction argument, which is an integer and takes on
464 one of the following values::
465
466 DMA_BIDIRECTIONAL
467 DMA_TO_DEVICE
468 DMA_FROM_DEVICE
469 DMA_NONE
470
471 You should provide the exact DMA direction if you know it.
472
473 DMA_TO_DEVICE means "from main memory to the device"
474 DMA_FROM_DEVICE means "from the device to main memory"
475 It is the direction in which the data moves during the DMA
476 transfer.
477
478 You are _strongly_ encouraged to specify this as precisely
479 as you possibly can.
480
481 If you absolutely cannot know the direction of the DMA transfer,
482 specify DMA_BIDIRECTIONAL. It means that the DMA can go in
483 either direction. The platform guarantees that you may legally
484 specify this, and that it will work, but this may be at the
485 cost of performance for example.
486
487 The value DMA_NONE is to be used for debugging. One can
488 hold this in a data structure before you come to know the
489 precise direction, and this will help catch cases where your
490 direction tracking logic has failed to set things up properly.
491
492 Another advantage of specifying this value precisely (outside of
493 potential platform-specific optimizations of such) is for debugging.
494 Some platforms actually have a write permission boolean which DMA
495 mappings can be marked with, much like page protections in the user
496 program address space. Such platforms can and do report errors in the
497 kernel logs when the DMA controller hardware detects violation of the
498 permission setting.
499
500 Only streaming mappings specify a direction, coherent mappings
501 implicitly have a direction attribute setting of
502 DMA_BIDIRECTIONAL.
503
504 The SCSI subsystem tells you the direction to use in the
505 'sc_data_direction' member of the SCSI command your driver is
506 working on.
507
508 For Networking drivers, it's a rather simple affair. For transmit
509 packets, map/unmap them with the DMA_TO_DEVICE direction
510 specifier. For receive packets, just the opposite, map/unmap them
511 with the DMA_FROM_DEVICE direction specifier.
512
513 Using Streaming DMA mappings
514 ============================
515
516 The streaming DMA mapping routines can be called from interrupt
517 context. There are two versions of each map/unmap, one which will
518 map/unmap a single memory region, and one which will map/unmap a
519 scatterlist.
520
521 To map a single region, you do::
522
523 struct device *dev = &my_dev->dev;
524 dma_addr_t dma_handle;
525 void *addr = buffer->ptr;
526 size_t size = buffer->len;
527
528 dma_handle = dma_map_single(dev, addr, size, direction);
529 if (dma_mapping_error(dev, dma_handle)) {
530 /*
531 * reduce current DMA mapping usage,
532 * delay and try again later or
533 * reset driver.
534 */
535 goto map_error_handling;
536 }
537
538 and to unmap it::
539
540 dma_unmap_single(dev, dma_handle, size, direction);
541
542 You should call dma_mapping_error() as dma_map_single() could fail and return
543 error. Doing so will ensure that the mapping code will work correctly on all
544 DMA implementations without any dependency on the specifics of the underlying
545 implementation. Using the returned address without checking for errors could
546 result in failures ranging from panics to silent data corruption. The same
547 applies to dma_map_page() as well.
548
549 You should call dma_unmap_single() when the DMA activity is finished, e.g.,
550 from the interrupt which told you that the DMA transfer is done.
551
552 Using CPU pointers like this for single mappings has a disadvantage:
553 you cannot reference HIGHMEM memory in this way. Thus, there is a
554 map/unmap interface pair akin to dma_{map,unmap}_single(). These
555 interfaces deal with page/offset pairs instead of CPU pointers.
556 Specifically::
557
558 struct device *dev = &my_dev->dev;
559 dma_addr_t dma_handle;
560 struct page *page = buffer->page;
561 unsigned long offset = buffer->offset;
562 size_t size = buffer->len;
563
564 dma_handle = dma_map_page(dev, page, offset, size, direction);
565 if (dma_mapping_error(dev, dma_handle)) {
566 /*
567 * reduce current DMA mapping usage,
568 * delay and try again later or
569 * reset driver.
570 */
571 goto map_error_handling;
572 }
573
574 ...
575
576 dma_unmap_page(dev, dma_handle, size, direction);
577
578 Here, "offset" means byte offset within the given page.
579
580 You should call dma_mapping_error() as dma_map_page() could fail and return
581 error as outlined under the dma_map_single() discussion.
582
583 You should call dma_unmap_page() when the DMA activity is finished, e.g.,
584 from the interrupt which told you that the DMA transfer is done.
585
586 With scatterlists, you map a region gathered from several regions by::
587
588 int i, count = dma_map_sg(dev, sglist, nents, direction);
589 struct scatterlist *sg;
590
591 for_each_sg(sglist, sg, count, i) {
592 hw_address[i] = sg_dma_address(sg);
593 hw_len[i] = sg_dma_len(sg);
594 }
595
596 where nents is the number of entries in the sglist.
597
598 The implementation is free to merge several consecutive sglist entries
599 into one (e.g. if DMA mapping is done with PAGE_SIZE granularity, any
600 consecutive sglist entries can be merged into one provided the first one
601 ends and the second one starts on a page boundary - in fact this is a huge
602 advantage for cards which either cannot do scatter-gather or have very
603 limited number of scatter-gather entries) and returns the actual number
604 of sg entries it mapped them to. On failure 0 is returned.
605
606 Then you should loop count times (note: this can be less than nents times)
607 and use sg_dma_address() and sg_dma_len() macros where you previously
608 accessed sg->address and sg->length as shown above.
609
610 To unmap a scatterlist, just call::
611
612 dma_unmap_sg(dev, sglist, nents, direction);
613
614 Again, make sure DMA activity has already finished.
615
616 .. note::
617
618 The 'nents' argument to the dma_unmap_sg call must be
619 the _same_ one you passed into the dma_map_sg call,
620 it should _NOT_ be the 'count' value _returned_ from the
621 dma_map_sg call.
622
623 Every dma_map_{single,sg}() call should have its dma_unmap_{single,sg}()
624 counterpart, because the DMA address space is a shared resource and
625 you could render the machine unusable by consuming all DMA addresses.
626
627 If you need to use the same streaming DMA region multiple times and touch
628 the data in between the DMA transfers, the buffer needs to be synced
629 properly in order for the CPU and device to see the most up-to-date and
630 correct copy of the DMA buffer.
631
632 So, firstly, just map it with dma_map_{single,sg}(), and after each DMA
633 transfer call either::
634
635 dma_sync_single_for_cpu(dev, dma_handle, size, direction);
636
637 or::
638
639 dma_sync_sg_for_cpu(dev, sglist, nents, direction);
640
641 as appropriate.
642
643 Then, if you wish to let the device get at the DMA area again,
644 finish accessing the data with the CPU, and then before actually
645 giving the buffer to the hardware call either::
646
647 dma_sync_single_for_device(dev, dma_handle, size, direction);
648
649 or::
650
651 dma_sync_sg_for_device(dev, sglist, nents, direction);
652
653 as appropriate.
654
655 .. note::
656
657 The 'nents' argument to dma_sync_sg_for_cpu() and
658 dma_sync_sg_for_device() must be the same passed to
659 dma_map_sg(). It is _NOT_ the count returned by
660 dma_map_sg().
661
662 After the last DMA transfer call one of the DMA unmap routines
663 dma_unmap_{single,sg}(). If you don't touch the data from the first
664 dma_map_*() call till dma_unmap_*(), then you don't have to call the
665 dma_sync_*() routines at all.
666
667 Here is pseudo code which shows a situation in which you would need
668 to use the dma_sync_*() interfaces::
669
670 my_card_setup_receive_buffer(struct my_card *cp, char *buffer, int len)
671 {
672 dma_addr_t mapping;
673
674 mapping = dma_map_single(cp->dev, buffer, len, DMA_FROM_DEVICE);
675 if (dma_mapping_error(cp->dev, mapping)) {
676 /*
677 * reduce current DMA mapping usage,
678 * delay and try again later or
679 * reset driver.
680 */
681 goto map_error_handling;
682 }
683
684 cp->rx_buf = buffer;
685 cp->rx_len = len;
686 cp->rx_dma = mapping;
687
688 give_rx_buf_to_card(cp);
689 }
690
691 ...
692
693 my_card_interrupt_handler(int irq, void *devid, struct pt_regs *regs)
694 {
695 struct my_card *cp = devid;
696
697 ...
698 if (read_card_status(cp) == RX_BUF_TRANSFERRED) {
699 struct my_card_header *hp;
700
701 /* Examine the header to see if we wish
702 * to accept the data. But synchronize
703 * the DMA transfer with the CPU first
704 * so that we see updated contents.
705 */
706 dma_sync_single_for_cpu(&cp->dev, cp->rx_dma,
707 cp->rx_len,
708 DMA_FROM_DEVICE);
709
710 /* Now it is safe to examine the buffer. */
711 hp = (struct my_card_header *) cp->rx_buf;
712 if (header_is_ok(hp)) {
713 dma_unmap_single(&cp->dev, cp->rx_dma, cp->rx_len,
714 DMA_FROM_DEVICE);
715 pass_to_upper_layers(cp->rx_buf);
716 make_and_setup_new_rx_buf(cp);
717 } else {
718 /* CPU should not write to
719 * DMA_FROM_DEVICE-mapped area,
720 * so dma_sync_single_for_device() is
721 * not needed here. It would be required
722 * for DMA_BIDIRECTIONAL mapping if
723 * the memory was modified.
724 */
725 give_rx_buf_to_card(cp);
726 }
727 }
728 }
729
730 Handling Errors
731 ===============
732
733 DMA address space is limited on some architectures and an allocation
734 failure can be determined by:
735
736 - checking if dma_alloc_coherent() returns NULL or dma_map_sg returns 0
737
738 - checking the dma_addr_t returned from dma_map_single() and dma_map_page()
739 by using dma_mapping_error()::
740
741 dma_addr_t dma_handle;
742
743 dma_handle = dma_map_single(dev, addr, size, direction);
744 if (dma_mapping_error(dev, dma_handle)) {
745 /*
746 * reduce current DMA mapping usage,
747 * delay and try again later or
748 * reset driver.
749 */
750 goto map_error_handling;
751 }
752
753 - unmap pages that are already mapped, when mapping error occurs in the middle
754 of a multiple page mapping attempt. These example are applicable to
755 dma_map_page() as well.
756
757 Example 1::
758
759 dma_addr_t dma_handle1;
760 dma_addr_t dma_handle2;
761
762 dma_handle1 = dma_map_single(dev, addr, size, direction);
763 if (dma_mapping_error(dev, dma_handle1)) {
764 /*
765 * reduce current DMA mapping usage,
766 * delay and try again later or
767 * reset driver.
768 */
769 goto map_error_handling1;
770 }
771 dma_handle2 = dma_map_single(dev, addr, size, direction);
772 if (dma_mapping_error(dev, dma_handle2)) {
773 /*
774 * reduce current DMA mapping usage,
775 * delay and try again later or
776 * reset driver.
777 */
778 goto map_error_handling2;
779 }
780
781 ...
782
783 map_error_handling2:
784 dma_unmap_single(dma_handle1);
785 map_error_handling1:
786
787 Example 2::
788
789 /*
790 * if buffers are allocated in a loop, unmap all mapped buffers when
791 * mapping error is detected in the middle
792 */
793
794 dma_addr_t dma_addr;
795 dma_addr_t array[DMA_BUFFERS];
796 int save_index = 0;
797
798 for (i = 0; i < DMA_BUFFERS; i++) {
799
800 ...
801
802 dma_addr = dma_map_single(dev, addr, size, direction);
803 if (dma_mapping_error(dev, dma_addr)) {
804 /*
805 * reduce current DMA mapping usage,
806 * delay and try again later or
807 * reset driver.
808 */
809 goto map_error_handling;
810 }
811 array[i].dma_addr = dma_addr;
812 save_index++;
813 }
814
815 ...
816
817 map_error_handling:
818
819 for (i = 0; i < save_index; i++) {
820
821 ...
822
823 dma_unmap_single(array[i].dma_addr);
824 }
825
826 Networking drivers must call dev_kfree_skb() to free the socket buffer
827 and return NETDEV_TX_OK if the DMA mapping fails on the transmit hook
828 (ndo_start_xmit). This means that the socket buffer is just dropped in
829 the failure case.
830
831 SCSI drivers must return SCSI_MLQUEUE_HOST_BUSY if the DMA mapping
832 fails in the queuecommand hook. This means that the SCSI subsystem
833 passes the command to the driver again later.
834
835 Optimizing Unmap State Space Consumption
836 ========================================
837
838 On many platforms, dma_unmap_{single,page}() is simply a nop.
839 Therefore, keeping track of the mapping address and length is a waste
840 of space. Instead of filling your drivers up with ifdefs and the like
841 to "work around" this (which would defeat the whole purpose of a
842 portable API) the following facilities are provided.
843
844 Actually, instead of describing the macros one by one, we'll
845 transform some example code.
846
847 1) Use DEFINE_DMA_UNMAP_{ADDR,LEN} in state saving structures.
848 Example, before::
849
850 struct ring_state {
851 struct sk_buff *skb;
852 dma_addr_t mapping;
853 __u32 len;
854 };
855
856 after::
857
858 struct ring_state {
859 struct sk_buff *skb;
860 DEFINE_DMA_UNMAP_ADDR(mapping);
861 DEFINE_DMA_UNMAP_LEN(len);
862 };
863
864 2) Use dma_unmap_{addr,len}_set() to set these values.
865 Example, before::
866
867 ringp->mapping = FOO;
868 ringp->len = BAR;
869
870 after::
871
872 dma_unmap_addr_set(ringp, mapping, FOO);
873 dma_unmap_len_set(ringp, len, BAR);
874
875 3) Use dma_unmap_{addr,len}() to access these values.
876 Example, before::
877
878 dma_unmap_single(dev, ringp->mapping, ringp->len,
879 DMA_FROM_DEVICE);
880
881 after::
882
883 dma_unmap_single(dev,
884 dma_unmap_addr(ringp, mapping),
885 dma_unmap_len(ringp, len),
886 DMA_FROM_DEVICE);
887
888 It really should be self-explanatory. We treat the ADDR and LEN
889 separately, because it is possible for an implementation to only
890 need the address in order to perform the unmap operation.
891
892 Platform Issues
893 ===============
894
895 If you are just writing drivers for Linux and do not maintain
896 an architecture port for the kernel, you can safely skip down
897 to "Closing".
898
899 1) Struct scatterlist requirements.
900
901 You need to enable CONFIG_NEED_SG_DMA_LENGTH if the architecture
902 supports IOMMUs (including software IOMMU).
903
904 2) ARCH_DMA_MINALIGN
905
906 Architectures must ensure that kmalloc'ed buffer is
907 DMA-safe. Drivers and subsystems depend on it. If an architecture
908 isn't fully DMA-coherent (i.e. hardware doesn't ensure that data in
909 the CPU cache is identical to data in main memory),
910 ARCH_DMA_MINALIGN must be set so that the memory allocator
911 makes sure that kmalloc'ed buffer doesn't share a cache line with
912 the others. See arch/arm/include/asm/cache.h as an example.
913
914 Note that ARCH_DMA_MINALIGN is about DMA memory alignment
915 constraints. You don't need to worry about the architecture data
916 alignment constraints (e.g. the alignment constraints about 64-bit
917 objects).
918
919 Closing
920 =======
921
922 This document, and the API itself, would not be in its current
923 form without the feedback and suggestions from numerous individuals.
924 We would like to specifically mention, in no particular order, the
925 following people::
926
927 Russell King <[email protected]>
928 Leo Dagum <[email protected]>
929 Ralf Baechle <[email protected]>
930 Grant Grundler <[email protected]>
931 Jay Estabrook <[email protected]>
932 Thomas Sailer <[email protected]>
933 Andrea Arcangeli <[email protected]>
934 Jens Axboe <[email protected]>
935 David Mosberger-Tang <[email protected]>
936

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Dynamic DMA mapping 안내서

1-12

Dynamic DMA mapping Guide (동적 DMA mapping 안내서)

저자는 David S. Miller <[email protected]>, Richard Henderson <[email protected]>, Jakub Jelinek <[email protected]>입니다.

이 문서는 device driver 작성자가 DMA API를 사용하는 방법을 예제 pseudo-code와 함께 설명합니다. API의 간결한 설명은 `Documentation/core-api/dma-api.rst`를 참고하십시오.

CPU 주소와 DMA 주소

13-110

CPU 주소와 DMA 주소

DMA API에는 여러 종류의 주소가 관여하므로 그 차이를 이해하는 것이 중요합니다.

Kernel은 보통 virtual address를 사용합니다. `kmalloc()`, `vmalloc()` 및 비슷한 interface가 반환한 주소는 virtual address이며 `void *`에 저장할 수 있습니다.

Virtual memory system인 TLB와 page table 등은 virtual address를 CPU physical address로 변환합니다. CPU physical address는 `phys_addr_t` 또는 `resource_size_t`에 저장됩니다. Kernel은 register 같은 device resource를 physical address로 관리하고 `/proc/iomem`에 노출합니다. Physical address는 driver가 직접 사용할 수 없으므로 `ioremap()`으로 공간을 mapping해 virtual address를 만들어야 합니다.

I/O device는 세 번째 종류인 bus address를 사용합니다. Device의 register가 MMIO address에 있거나 device가 DMA로 system memory를 읽고 쓸 때 사용하는 주소가 bus address입니다. 일부 시스템에서는 bus address와 CPU physical address가 같지만 일반적으로는 다르며, IOMMU와 host bridge가 둘 사이에 임의의 mapping을 만들 수 있습니다.

Device 관점에서 DMA는 bus address space를 사용하지만 그 일부로 제한될 수 있습니다. 예를 들어 main memory와 PCI BAR가 64-bit address를 지원해도 IOMMU를 사용하면 device는 32-bit DMA address만 다루면 됩니다.

CPU와 device 사이의 주소 변환
Device MMIO bus address AHost bridge offsetCPU physical MMIO address Bioremap virtual mappingCPU virtual address C
Device DMA address ZIOMMU mappingSystem RAM buffer YVirtual memory mappingCPU virtual buffer X

MMIO 접근은 bus address A에서 host bridge offset을 거쳐 CPU physical address B로, 다시 virtual mapping을 거쳐 CPU virtual address C로 이어집니다. DMA buffer는 CPU virtual address X가 physical RAM address Y로 mapping되고, IOMMU가 device의 DMA address Z를 Y로 변환합니다.

Enumeration 과정에서 kernel은 I/O device, MMIO space, 시스템에 연결하는 host bridge를 파악합니다. PCI device에 BAR가 있으면 kernel은 BAR의 bus address A를 읽어 CPU physical address B로 변환합니다. B는 `struct resource`에 저장되고 보통 `/proc/iomem`에 표시됩니다. Driver가 device를 점유하면 `ioremap()`으로 B를 virtual address C에 mapping하고, 예를 들어 `ioread32(C)`로 bus address A의 register에 접근합니다.

Device가 DMA를 지원하면 driver는 `kmalloc()` 같은 interface로 virtual address X를 반환받아 buffer를 준비합니다. Virtual memory system은 X를 system RAM의 physical address Y로 mapping합니다. Driver는 X로 buffer에 접근하지만 DMA는 CPU virtual memory system을 통과하지 않으므로 device는 X를 사용할 수 없습니다.

단순한 시스템에서는 device가 Y에 직접 DMA할 수 있습니다. 많은 시스템에서는 IOMMU가 DMA address를 physical address로 변환하며, 예를 들어 Z를 Y로 mapping합니다. 이것이 DMA API가 필요한 이유입니다. Driver가 virtual address X를 `dma_map_single()` 같은 interface에 주면 필요한 IOMMU mapping을 만들고 DMA address Z를 반환합니다. Driver는 device에 Z로 DMA하도록 지시하고 IOMMU는 이를 system RAM의 Y buffer로 연결합니다.

Linux의 dynamic DMA mapping을 사용하려면 driver가 DMA address를 실제 사용하는 동안에만 mapping하고 DMA transfer 뒤에는 unmap해야 합니다. 아래 API는 이런 hardware가 없는 platform에서도 동작합니다.

DMA API는 microprocessor architecture와 무관하게 모든 bus에서 동작합니다. Bus 전용 DMA API인 `pci_map_*()` 대신 `dma_map_*()` interface를 사용해야 합니다.

먼저 driver에 다음 header가 포함되어 있는지 확인합니다.

#include <linux/dma-mapping.h>

이 header는 `dma_addr_t`를 정의합니다. 이 type은 platform에서 유효한 모든 DMA address를 담을 수 있으므로 DMA mapping function이 반환한 DMA address를 보관하는 모든 곳에서 사용해야 합니다.

DMA에 사용할 수 있는 메모리

111-148

어떤 메모리를 DMA에 사용할 수 있는가?

먼저 DMA mapping 기능에 사용할 수 있는 kernel memory를 알아야 합니다. 과거에는 기록되지 않은 규칙이었으며 이 절에서 이를 명문화합니다.

Page allocator인 `__get_free_page*()` 또는 generic memory allocator인 `kmalloc()`, `kmem_cache_alloc()`으로 얻은 메모리는 해당 routine이 반환한 주소로 DMA 송수신에 사용할 수 있습니다.

`vmalloc()`이 반환한 메모리나 주소는 DMA에 사용할 수 없습니다. `vmalloc()` 영역에 mapping된 실제 하부 메모리에 DMA하는 것은 가능하지만, page table을 순회해 physical address를 얻고 각 page를 `__va()` 같은 방법으로 다시 kernel address로 변환해야 합니다. 원문의 편집 메모리는 Gerd Knorr의 generic code가 통합되면 이 설명을 갱신하라고 적고 있습니다.

Kernel image의 data/text/bss segment 주소, module image 주소, stack 주소 역시 DMA에 사용할 수 없습니다. 이들은 나머지 physical memory와 완전히 다른 곳에 mapping될 수 있습니다. 물리적으로 DMA가 가능하더라도 I/O buffer를 cacheline에 맞춰 정렬해야 합니다. 그렇지 않으면 DMA-incoherent cache를 쓰는 CPU에서 cacheline sharing으로 data corruption이 생길 수 있습니다. CPU와 DMA가 같은 cache line의 서로 다른 word를 쓰면 둘 중 하나가 덮어써질 수 있습니다.

`kmap()` 반환값도 DMA 송수신에 사용할 수 없습니다. 이는 `vmalloc()`과 같은 종류의 제약입니다.

Block I/O와 networking subsystem은 자신들이 제공하는 buffer가 DMA 송수신에 유효하도록 보장합니다.

DMA 주소 지정 능력과 mask 설정

149-205

DMA 주소 지정 능력

기본적으로 kernel은 device가 32-bit DMA address를 지정할 수 있다고 가정합니다. 64-bit device는 범위를 늘려야 하고 제약이 있는 device는 줄여야 합니다.

PCI에 관한 특별한 주의 사항으로, PCI-X 명세는 PCI-X device가 모든 transaction에서 64-bit addressing인 DAC를 지원하도록 요구합니다. 적어도 SGI SN2 platform은 I/O bus가 PCI-X mode일 때 올바르게 동작하려면 64-bit coherent allocation이 필요합니다.

올바른 동작을 위해 DMA mask를 설정하여 device의 DMA addressing capability를 kernel에 알려야 합니다. Streaming과 coherent API의 mask를 함께 설정하려면 다음을 호출합니다.

int dma_set_mask_and_coherent(struct device *dev, u64 mask);

특별한 요구가 있으면 두 API를 나누어 설정할 수 있습니다. Streaming mapping에는 다음을 사용합니다.

int dma_set_mask(struct device *dev, u64 mask);

Coherent allocation에는 다음을 사용합니다.

int dma_set_coherent_mask(struct device *dev, u64 mask);

`dev`는 device의 `struct device` pointer이고 `mask`는 device가 지원하는 address bit를 나타내는 bit mask입니다. `struct device`는 보통 bus 전용 device structure 안에 들어 있습니다. 예를 들어 PCI device structure pointer가 `pdev`라면 `&pdev->dev`가 해당 device structure의 pointer입니다.

이 함수들은 주어진 address mask로 해당 시스템에서 DMA를 올바르게 수행할 수 있으면 보통 0을 반환합니다. Mask가 시스템에서 지원하기에 너무 작으면 오류를 반환할 수 있습니다. 0이 아니면 이 platform에서 DMA를 수행할 수 없고 시도 결과는 undefined behavior입니다. `dma_set_mask` 계열이 성공하기 전에는 해당 device에서 DMA를 사용하면 안 됩니다.

실패하면 가능한 선택은 두 가지입니다.

  • 가능하다면 data transfer에 DMA가 아닌 mode를 사용합니다.
  • 해당 device를 무시하고 초기화하지 않습니다.

DMA mask 설정 실패 시 driver가 kernel `KERN_WARNING` message를 출력하는 것이 좋습니다. 사용자가 성능 저하나 device 미검출을 보고할 때 kernel message를 통해 정확한 원인을 확인할 수 있습니다.

DMA mask 설정 예제와 다기능 device

206-295

24-bit addressing device는 다음처럼 설정합니다.

if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(24))) {
        dev_warn(dev, "mydev: No suitable DMA available\n");
        goto ignore_this_device;
}

표준 64-bit addressing device는 다음과 같이 설정합니다.

dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64))

`dma_set_mask_and_coherent()`는 `DMA_BIT_MASK(64)`에서 실패를 반환하지 않습니다. 따라서 다음과 같은 fallback code는 잘못되었습니다.

/* Wrong code */
if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64)))
        dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32))

`dma_set_mask_and_coherent()`는 32-bit보다 큰 mask에서 실패하지 않으므로 device의 capability를 알고 다음과 같이 선택하는 것이 권장됩니다.

/* Recommended code */
if (support_64bit)
        dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
else
        dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32));

Device가 coherent allocation의 descriptor에는 32-bit addressing만 지원하지만 streaming mapping에는 전체 64-bit를 지원한다면 streaming mask를 다음처럼 설정합니다.

if (dma_set_mask(dev, DMA_BIT_MASK(64))) {
        dev_warn(dev, "mydev: No suitable DMA available\n");
        goto ignore_this_device;
}

Coherent mask는 항상 streaming mask와 같거나 더 작은 mask로 설정할 수 있습니다. 드물게 coherent allocation만 사용하는 driver라면 `dma_set_coherent_mask()`의 반환값을 검사해야 합니다.

Device가 주소의 낮은 24-bit만 구동한다면 다음처럼 설정합니다.

if (dma_set_mask(dev, DMA_BIT_MASK(24))) {
        dev_warn(dev, "mydev: 24-bit DMA addressing not available\n");
        goto ignore_this_device;
}

`dma_set_mask()` 또는 `dma_set_mask_and_coherent()`가 성공해 0을 반환하면 kernel은 제공된 mask를 저장하고 이후 DMA mapping을 만들 때 사용합니다.

Sound card의 playback과 record처럼 여러 기능의 DMA addressing limit가 서로 다르면 각 mask를 probe하고 시스템이 처리할 수 있는 기능만 제공할 수 있습니다. 마지막 `dma_set_mask()` 호출에는 가장 구체적인 mask를 사용해야 합니다.

다음 pseudo-code가 그 방법을 보여 줍니다.

#define PLAYBACK_ADDRESS_BITS        DMA_BIT_MASK(32)
#define RECORD_ADDRESS_BITS        DMA_BIT_MASK(24)

struct my_sound_card *card;
struct device *dev;

...
if (!dma_set_mask(dev, PLAYBACK_ADDRESS_BITS)) {
        card->playback_enabled = 1;
} else {
        card->playback_enabled = 0;
        dev_warn(dev, "%s: Playback disabled due to DMA limitations\n",
               card->name);
}
if (!dma_set_mask(dev, RECORD_ADDRESS_BITS)) {
        card->record_enabled = 1;
} else {
        card->record_enabled = 0;
        dev_warn(dev, "%s: Record disabled due to DMA limitations\n",
               card->name);
}

Sound card를 예로 든 이유는 이런 PCI device에 PCI front end를 붙인 ISA chip이 흔하여 ISA의 16MB DMA addressing limit를 그대로 유지하는 경우가 많기 때문입니다.

Coherent와 streaming DMA mapping

296-367

DMA mapping의 종류

DMA mapping은 `Coherent DMA mappings`와 `Streaming DMA mappings` 두 종류입니다.

Coherent DMA mapping은 보통 driver initialization 때 mapping하고 종료 때 unmap합니다. Hardware는 device와 CPU가 data에 병렬로 접근하면서 명시적인 software flush 없이도 서로의 update를 볼 수 있도록 보장해야 합니다. `coherent`를 `synchronous`로 이해하면 됩니다.

현재 기본값은 DMA space의 낮은 32-bit에서 coherent memory를 반환하는 것입니다. Driver가 이 기본값으로 충분하더라도 향후 호환성을 위해 coherent mask를 설정해야 합니다.

Coherent mapping의 좋은 사용 예는 다음과 같습니다.

  • Network card DMA ring descriptor
  • SCSI adapter mailbox command data structure
  • Main memory에서 실행되는 device firmware microcode

이 예들의 공통 불변 조건은 CPU의 memory store가 device에 즉시 보이고 그 반대도 성립해야 한다는 것입니다. Coherent mapping이 이를 보장합니다.

Coherent DMA memory에서도 올바른 memory barrier가 필요합니다. CPU는 일반 memory와 마찬가지로 coherent memory에 대한 store 순서를 바꿀 수 있습니다. Device가 descriptor의 첫 word 갱신을 두 번째 word보다 먼저 봐야 한다면 다음 순서를 사용해야 합니다.

desc->word0 = address;
wmb();
desc->word1 = DESC_VALID;

이렇게 해야 모든 platform에서 올바르게 동작합니다. 일부 platform에서는 PCI bridge의 write buffer를 flush하듯 register를 쓴 뒤 값을 읽는 방식 등으로 CPU write buffer도 flush해야 할 수 있습니다.

Streaming DMA mapping은 보통 한 번의 DMA transfer를 위해 mapping하고 직후 unmap합니다. 아래의 `dma_sync_*`를 쓰는 경우는 예외입니다. Hardware는 sequential access에 맞게 최적화할 수 있습니다. `streaming`을 `asynchronous` 또는 `coherency domain 밖`으로 이해하면 됩니다.

Streaming mapping의 좋은 사용 예는 다음과 같습니다.

  • Device가 송수신하는 networking buffer
  • SCSI device가 읽고 쓰는 filesystem buffer

Streaming interface는 구현이 hardware가 허용하는 모든 성능 최적화를 할 수 있도록 설계되었습니다. 따라서 사용자는 원하는 동작을 명시해야 합니다.

두 종류 모두 하부 bus에서 비롯된 alignment 제한은 없지만 device 자체의 제한은 있을 수 있습니다. DMA-coherent가 아닌 cache를 쓰는 시스템에서는 buffer가 다른 data와 cache line을 공유하지 않을 때 더 잘 동작합니다.

Coherent DMA mapping 사용법

368-458

Coherent DMA mapping 사용

`PAGE_SIZE` 정도의 큰 coherent DMA region을 allocate하고 mapping하려면 다음과 같이 합니다.

dma_addr_t dma_handle;

cpu_addr = dma_alloc_coherent(dev, size, &dma_handle, gfp);

`dev`는 `struct device *`입니다. `GFP_ATOMIC` flag를 사용하면 interrupt context에서 호출할 수 있습니다. `size`는 allocate할 region의 byte 길이입니다.

이 routine은 해당 region의 RAM을 allocate하므로 page order 대신 size를 받는 `__get_free_pages()`와 비슷합니다. 한 page보다 작은 region이 필요하면 아래의 `dma_pool` interface가 더 적합할 수 있습니다.

Coherent DMA mapping interface는 기본적으로 32-bit로 address 가능한 DMA address를 반환합니다. Device가 DMA mask로 상위 32-bit 접근을 표시해도 `dma_set_coherent_mask()`로 coherent DMA mask를 명시적으로 바꾼 경우에만 32-bit보다 큰 DMA address를 반환합니다. `dma_pool`에도 같은 규칙이 적용됩니다.

`dma_alloc_coherent()`는 CPU에서 접근할 virtual address와 card에 전달할 `dma_handle` 두 값을 반환합니다.

CPU virtual address와 DMA address는 모두 요청 size 이상인 가장 작은 `PAGE_SIZE` order에 맞춰 정렬됩니다. 예를 들어 64KB 이하 chunk를 allocate하면 반환된 buffer 범위가 64K boundary를 넘지 않습니다.

이 DMA region을 unmap하고 free하려면 `dma_free_coherent()`를 호출합니다.

dma_free_coherent(dev, size, cpu_addr, dma_handle);

`dev`와 `size`는 allocate 때와 같고 `cpu_addr`와 `dma_handle`은 `dma_alloc_coherent()`의 반환값입니다. 이 함수는 interrupt context에서 호출할 수 없습니다.

작은 memory region이 많이 필요하면 `dma_alloc_coherent()`가 반환한 page를 직접 나누거나 `dma_pool` API를 사용할 수 있습니다. `dma_pool`은 `kmem_cache`와 비슷하지만 `__get_free_pages()` 대신 `dma_alloc_coherent()`를 사용하며, queue head가 N-byte boundary에 정렬되어야 하는 것 같은 hardware alignment constraint도 이해합니다.

다음과 같이 `dma_pool`을 만듭니다.

struct dma_pool *pool;

pool = dma_pool_create(name, dev, size, align, boundary);

`name`은 `kmem_cache` 이름처럼 진단에 쓰입니다. `dev`와 `size`는 위와 같습니다. `align`은 이 data type에 대한 device hardware의 byte 단위 alignment requirement이며 2의 거듭제곱이어야 합니다. Boundary crossing 제한이 없으면 `boundary`에 0을 전달합니다. 4096은 pool allocation이 4KByte boundary를 넘지 못한다는 뜻이며, 이 경우에는 `dma_alloc_coherent()`를 직접 쓰는 편이 나을 수 있습니다.

DMA pool에서 memory를 allocate하려면 다음을 사용합니다.

cpu_addr = dma_pool_alloc(pool, flags, &dma_handle);

Blocking이 허용되면, 즉 interrupt 안이나 SMP lock 보유 중이 아니라면 `flags`는 `GFP_KERNEL`이고 그 밖에는 `GFP_ATOMIC`입니다. `dma_alloc_coherent()`처럼 `cpu_addr`와 `dma_handle`을 반환합니다.

Pool에서 allocate한 memory는 다음과 같이 free합니다.

dma_pool_free(pool, cpu_addr, dma_handle);

`pool`은 `dma_pool_alloc()`에 전달한 값이고 나머지는 그 함수가 반환한 값입니다. 이 함수는 interrupt context에서 호출할 수 있습니다.

DMA pool은 다음과 같이 destroy합니다.

dma_pool_destroy(pool);

Pool을 destroy하기 전에 그 pool에서 allocate한 모든 memory에 `dma_pool_free()`를 호출했는지 확인해야 합니다. `dma_pool_destroy()`는 interrupt context에서 호출할 수 없습니다.

DMA 전송 방향

459-512

DMA 방향

이 문서 뒤에서 설명하는 interface는 정수형 DMA direction argument로 다음 값 중 하나를 받습니다.

DMA_BIDIRECTIONAL
DMA_TO_DEVICE
DMA_FROM_DEVICE
DMA_NONE

방향을 안다면 정확히 지정해야 합니다. `DMA_TO_DEVICE`는 main memory에서 device로, `DMA_FROM_DEVICE`는 device에서 main memory로 data가 이동한다는 뜻이며 DMA transfer의 실제 이동 방향을 가리킵니다.

가능한 한 정확한 방향을 지정할 것을 강하게 권장합니다.

DMA transfer 방향을 정말 알 수 없다면 양방향을 뜻하는 `DMA_BIDIRECTIONAL`을 지정합니다. Platform은 이 값이 합법적이며 동작함을 보장하지만 성능 비용이 들 수 있습니다.

`DMA_NONE`은 debugging에 사용합니다. 정확한 방향을 알기 전 data structure에 넣어 두면 direction tracking logic이 올바르게 설정하지 못한 경우를 잡는 데 도움이 됩니다.

정확한 방향은 platform 전용 최적화뿐 아니라 debugging에도 유리합니다. 일부 platform은 user address space의 page protection처럼 DMA mapping에 write permission boolean을 표시할 수 있고, DMA controller가 permission 위반을 감지하면 kernel log에 오류를 보고합니다.

Direction을 지정하는 것은 streaming mapping뿐입니다. Coherent mapping은 암묵적으로 `DMA_BIDIRECTIONAL` direction attribute를 가집니다.

SCSI subsystem은 driver가 처리 중인 SCSI command의 `sc_data_direction` member로 사용할 방향을 알려 줍니다.

Networking driver에서는 transmit packet을 `DMA_TO_DEVICE`로 map/unmap하고 receive packet은 반대로 `DMA_FROM_DEVICE`로 map/unmap합니다.

Streaming single 및 page mapping

513-585

Streaming DMA mapping 사용

Streaming DMA mapping routine은 interrupt context에서 호출할 수 있습니다. 각 map/unmap에는 단일 memory region용과 scatterlist용 두 형태가 있습니다.

단일 region은 다음과 같이 mapping합니다.

struct device *dev = &my_dev->dev;
dma_addr_t dma_handle;
void *addr = buffer->ptr;
size_t size = buffer->len;

dma_handle = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
        /*
         * reduce current DMA mapping usage,
         * delay and try again later or
         * reset driver.
         */
        goto map_error_handling;
}

다음과 같이 unmap합니다.

dma_unmap_single(dev, dma_handle, size, direction);

`dma_map_single()`은 실패해 오류를 반환할 수 있으므로 `dma_mapping_error()`를 호출해야 합니다. 그래야 하부 구현 세부 사항에 의존하지 않고 모든 DMA 구현에서 올바르게 동작합니다. 오류 검사 없이 반환 주소를 사용하면 panic부터 조용한 data corruption까지 발생할 수 있습니다. `dma_map_page()`에도 같은 규칙이 적용됩니다.

DMA activity가 끝나면, 예를 들어 DMA transfer 완료를 알린 interrupt에서 `dma_unmap_single()`을 호출해야 합니다.

CPU pointer를 쓰는 single mapping은 HIGHMEM memory를 참조할 수 없다는 단점이 있습니다. 이를 위해 `dma_{map,unmap}_single()`과 비슷하지만 CPU pointer 대신 page/offset pair를 다루는 interface가 있습니다.

struct device *dev = &my_dev->dev;
dma_addr_t dma_handle;
struct page *page = buffer->page;
unsigned long offset = buffer->offset;
size_t size = buffer->len;

dma_handle = dma_map_page(dev, page, offset, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
        /*
         * reduce current DMA mapping usage,
         * delay and try again later or
         * reset driver.
         */
        goto map_error_handling;
}

...

dma_unmap_page(dev, dma_handle, size, direction);

`offset`은 주어진 page 안의 byte offset입니다. `dma_map_page()`도 실패할 수 있으므로 `dma_mapping_error()`를 호출해야 하며, DMA activity가 끝나면 `dma_unmap_page()`를 호출합니다.

Streaming scatterlist mapping

586-626

여러 region에서 모은 영역은 scatterlist로 다음처럼 mapping합니다.

int i, count = dma_map_sg(dev, sglist, nents, direction);
struct scatterlist *sg;

for_each_sg(sglist, sg, count, i) {
        hw_address[i] = sg_dma_address(sg);
        hw_len[i] = sg_dma_len(sg);
}

`nents`는 `sglist` entry 수입니다. 구현은 연속된 여러 entry를 하나로 합칠 수 있습니다. 예를 들어 `PAGE_SIZE` 단위로 mapping할 때 앞 entry가 page boundary에서 끝나고 다음 entry가 그 boundary에서 시작하면 합칠 수 있습니다. Scatter-gather를 지원하지 않거나 entry 수가 매우 제한된 card에 큰 장점입니다. 반환값은 실제 mapping된 sg entry 수이며 실패하면 0입니다.

그 뒤에는 `nents`가 아니라 반환된 `count`만큼 순회하며, 기존의 `sg->address`, `sg->length` 대신 `sg_dma_address()`와 `sg_dma_len()` macro를 사용합니다.

Scatterlist는 다음과 같이 unmap합니다.

dma_unmap_sg(dev, sglist, nents, direction);

이때도 DMA activity가 이미 끝났는지 확인해야 합니다.

`dma_unmap_sg()`의 `nents` argument는 `dma_map_sg()`에 전달했던 값과 같아야 합니다. `dma_map_sg()`가 반환한 `count`를 사용하면 안 됩니다.

DMA address space는 공유 자원이므로 모든 `dma_map_{single,sg}()` 호출에는 대응하는 `dma_unmap_{single,sg}()`가 있어야 합니다. DMA address를 모두 소비하면 시스템을 사용할 수 없게 만들 수 있습니다.

반복 사용하는 streaming buffer 동기화

627-729

같은 streaming DMA region을 여러 번 사용하고 transfer 사이에 data를 만진다면 CPU와 device가 최신의 올바른 DMA buffer 사본을 보도록 buffer를 정확히 동기화해야 합니다.

먼저 `dma_map_{single,sg}()`로 mapping하고 각 DMA transfer 뒤 CPU가 접근하기 전에 상황에 맞는 다음 함수 중 하나를 호출합니다.

dma_sync_single_for_cpu(dev, dma_handle, size, direction);
dma_sync_sg_for_cpu(dev, sglist, nents, direction);

Device가 DMA area에 다시 접근하게 하려면 CPU 접근을 마친 뒤 buffer를 hardware에 넘기기 전에 다음 중 맞는 함수를 호출합니다.

dma_sync_single_for_device(dev, dma_handle, size, direction);
dma_sync_sg_for_device(dev, sglist, nents, direction);

`dma_sync_sg_for_cpu()`와 `dma_sync_sg_for_device()`의 `nents`는 `dma_map_sg()`에 전달한 값과 같아야 하며, `dma_map_sg()`가 반환한 `count`가 아닙니다.

마지막 DMA transfer 뒤에는 `dma_unmap_{single,sg}()` 중 하나를 호출합니다. 첫 `dma_map_*()`부터 `dma_unmap_*()`까지 CPU가 data를 건드리지 않는다면 `dma_sync_*()`를 호출할 필요가 없습니다.

다음 pseudo-code는 `dma_sync_*()` interface가 필요한 상황을 보여 줍니다.

my_card_setup_receive_buffer(struct my_card *cp, char *buffer, int len)
{
        dma_addr_t mapping;

        mapping = dma_map_single(cp->dev, buffer, len, DMA_FROM_DEVICE);
        if (dma_mapping_error(cp->dev, mapping)) {
                /*
                 * reduce current DMA mapping usage,
                 * delay and try again later or
                 * reset driver.
                 */
                goto map_error_handling;
        }

        cp->rx_buf = buffer;
        cp->rx_len = len;
        cp->rx_dma = mapping;

        give_rx_buf_to_card(cp);
}

...

my_card_interrupt_handler(int irq, void *devid, struct pt_regs *regs)
{
        struct my_card *cp = devid;

        ...
        if (read_card_status(cp) == RX_BUF_TRANSFERRED) {
                struct my_card_header *hp;

                /* Examine the header to see if we wish
                 * to accept the data.  But synchronize
                 * the DMA transfer with the CPU first
                 * so that we see updated contents.
                 */
                dma_sync_single_for_cpu(&cp->dev, cp->rx_dma,
                                        cp->rx_len,
                                        DMA_FROM_DEVICE);

                /* Now it is safe to examine the buffer. */
                hp = (struct my_card_header *) cp->rx_buf;
                if (header_is_ok(hp)) {
                        dma_unmap_single(&cp->dev, cp->rx_dma, cp->rx_len,
                                         DMA_FROM_DEVICE);
                        pass_to_upper_layers(cp->rx_buf);
                        make_and_setup_new_rx_buf(cp);
                } else {
                        /* CPU should not write to
                         * DMA_FROM_DEVICE-mapped area,
                         * so dma_sync_single_for_device() is
                         * not needed here. It would be required
                         * for DMA_BIDIRECTIONAL mapping if
                         * the memory was modified.
                         */
                        give_rx_buf_to_card(cp);
                }
        }
}

예제는 receive buffer를 `DMA_FROM_DEVICE`로 mapping하고 card에 넘깁니다. Interrupt handler는 header를 검사하기 전에 `dma_sync_single_for_cpu()`로 CPU 관찰 상태를 갱신합니다. Header가 유효하면 unmap하고 상위 layer로 넘깁니다. 유효하지 않으면 buffer를 card에 다시 주며, CPU가 `DMA_FROM_DEVICE` 영역을 쓰지 않았으므로 `dma_sync_single_for_device()`는 필요 없습니다. `DMA_BIDIRECTIONAL` mapping에서 memory를 수정했다면 필요합니다.

DMA mapping 오류 처리

730-834

오류 처리

일부 architecture에서는 DMA address space가 제한되어 있습니다. `dma_alloc_coherent()`가 `NULL`을 반환하거나 `dma_map_sg()`가 0을 반환하는지 검사하여 allocation failure를 확인합니다.

`dma_map_single()`과 `dma_map_page()`가 반환한 `dma_addr_t`는 `dma_mapping_error()`로 검사합니다.

dma_addr_t dma_handle;

dma_handle = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
        /*
         * reduce current DMA mapping usage,
         * delay and try again later or
         * reset driver.
         */
        goto map_error_handling;
}

여러 page를 mapping하는 도중 오류가 나면 이미 mapping한 page를 unmap해야 합니다. 다음 예들은 `dma_map_page()`에도 적용됩니다.

예제 1은 두 번째 mapping 실패 시 첫 번째 mapping을 해제합니다.

dma_addr_t dma_handle1;
dma_addr_t dma_handle2;

dma_handle1 = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle1)) {
        /*
         * reduce current DMA mapping usage,
         * delay and try again later or
         * reset driver.
         */
        goto map_error_handling1;
}
dma_handle2 = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle2)) {
        /*
         * reduce current DMA mapping usage,
         * delay and try again later or
         * reset driver.
         */
        goto map_error_handling2;
}

...

map_error_handling2:
        dma_unmap_single(dma_handle1);
map_error_handling1:

예제 2는 loop 도중 오류가 나면 성공한 index 수를 이용해 앞서 mapping한 모든 buffer를 해제합니다.

/*
 * if buffers are allocated in a loop, unmap all mapped buffers when
 * mapping error is detected in the middle
 */

dma_addr_t dma_addr;
dma_addr_t array[DMA_BUFFERS];
int save_index = 0;

for (i = 0; i < DMA_BUFFERS; i++) {

        ...

        dma_addr = dma_map_single(dev, addr, size, direction);
        if (dma_mapping_error(dev, dma_addr)) {
                /*
                 * reduce current DMA mapping usage,
                 * delay and try again later or
                 * reset driver.
                 */
                goto map_error_handling;
        }
        array[i].dma_addr = dma_addr;
        save_index++;
}

...

map_error_handling:

for (i = 0; i < save_index; i++) {

        ...

        dma_unmap_single(array[i].dma_addr);
}

Networking driver는 transmit hook인 `ndo_start_xmit`에서 DMA mapping이 실패하면 `dev_kfree_skb()`로 socket buffer를 free하고 `NETDEV_TX_OK`를 반환해야 합니다. 실패한 socket buffer는 drop됩니다.

SCSI driver는 `queuecommand` hook에서 DMA mapping이 실패하면 `SCSI_MLQUEUE_HOST_BUSY`를 반환해야 합니다. 그러면 SCSI subsystem이 나중에 command를 다시 driver에 전달합니다.

Unmap 상태 공간 사용량 최적화

835-891

Unmap 상태 공간 사용량 최적화

많은 platform에서 `dma_unmap_{single,page}()`는 단순한 nop입니다. 이 경우 mapping address와 length를 계속 저장하는 것은 공간 낭비입니다. Portable API의 목적을 해치는 `#ifdef` 우회 대신 다음 기능을 사용합니다.

Macro를 하나씩 설명하는 대신 예제 code를 변환해 보겠습니다. 첫째, state 저장 structure에서 `DEFINE_DMA_UNMAP_{ADDR,LEN}`을 사용합니다. 변경 전:

struct ring_state {
        struct sk_buff *skb;
        dma_addr_t mapping;
        __u32 len;
};

변경 후:

struct ring_state {
        struct sk_buff *skb;
        DEFINE_DMA_UNMAP_ADDR(mapping);
        DEFINE_DMA_UNMAP_LEN(len);
};

둘째, 값을 설정할 때 `dma_unmap_{addr,len}_set()`을 사용합니다. 변경 전:

ringp->mapping = FOO;
ringp->len = BAR;

변경 후:

dma_unmap_addr_set(ringp, mapping, FOO);
dma_unmap_len_set(ringp, len, BAR);

셋째, 값에 접근할 때 `dma_unmap_{addr,len}()`을 사용합니다. 변경 전:

dma_unmap_single(dev, ringp->mapping, ringp->len,
                 DMA_FROM_DEVICE);

변경 후:

dma_unmap_single(dev,
                 dma_unmap_addr(ringp, mapping),
                 dma_unmap_len(ringp, len),
                 DMA_FROM_DEVICE);

동작은 이름 그대로입니다. 구현에 따라 unmap operation에 address만 필요할 수 있으므로 `ADDR`과 `LEN`은 별도로 취급합니다.

Architecture port의 platform 고려 사항

892-918

Platform 문제

Linux driver만 작성하고 kernel architecture port를 유지하지 않는다면 이 절을 건너뛰고 `Closing`으로 가도 됩니다.

1) `struct scatterlist` 요구 사항: Architecture가 software IOMMU를 포함한 IOMMU를 지원한다면 `CONFIG_NEED_SG_DMA_LENGTH`를 활성화해야 합니다.

2) `ARCH_DMA_MINALIGN`: Architecture는 `kmalloc()` buffer가 DMA-safe하도록 보장해야 하며 driver와 subsystem은 이를 전제로 합니다. Hardware가 CPU cache의 data와 main memory의 data가 같음을 보장하지 않는 등 완전히 DMA-coherent하지 않다면 memory allocator가 `kmalloc()` buffer를 다른 data와 같은 cache line에 두지 않도록 `ARCH_DMA_MINALIGN`을 설정해야 합니다. 예제는 `arch/arm/include/asm/cache.h`를 참고하십시오.

`ARCH_DMA_MINALIGN`은 DMA memory alignment constraint를 위한 것입니다. 64-bit object 정렬 같은 architecture data alignment constraint는 여기서 고려할 필요가 없습니다.

맺음말과 기여자

919-935

맺음말

수많은 사람의 feedback과 제안이 없었다면 이 문서와 API는 현재 형태가 되지 못했을 것입니다. 특별히 다음 기여자들을 순서 없이 언급합니다.

Russell King <[email protected]>
Leo Dagum <[email protected]>
Ralf Baechle <[email protected]>
Grant Grundler <[email protected]>
Jay Estabrook <[email protected]>
Thomas Sailer <[email protected]>
Andrea Arcangeli <[email protected]>
Jens Axboe <[email protected]>
David Mosberger-Tang <[email protected]>