← Documents Documentation/crypto/async-tx-api.rst GitHub 원문 ↗

Linux 6.18.37 · Crypto

Asynchronous Transfers/Transforms API

dmaengine 기반 async_tx dependency chain, 지원 memory·RAID 연산, descriptor 수명, 실행·완료 시점, driver 준수 사항과 독점 channel 할당을 설명합니다.

Source pathDocumentation/crypto/async-tx-api.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

async-tx-api.rst:1-270

`async_tx`는 `xor->copy->xor`처럼 서로 의존하는 bulk memory 연산을 hardware offload와 software fallback 사이에서 같은 API로 실행합니다. Descriptor acknowledge와 dependency 연결이 재활용 시점을 제어합니다.

Driver는 tasklet 완료 callback, IRQ context 제약, dependency 정리 계약을 지켜야 합니다. Device-to-memory처럼 공유할 수 없는 channel은 `dma_request_channel()`과 `filter_fn`, `DMA_PRIVATE`로 독점 할당합니다. 원문의 연산 표와 source path 표를 구조화된 표로 보존했습니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 =====================================
4 Asynchronous Transfers/Transforms API
5 =====================================
6
7 .. Contents
8
9 1. INTRODUCTION
10
11 2 GENEALOGY
12
13 3 USAGE
14 3.1 General format of the API
15 3.2 Supported operations
16 3.3 Descriptor management
17 3.4 When does the operation execute?
18 3.5 When does the operation complete?
19 3.6 Constraints
20 3.7 Example
21
22 4 DMAENGINE DRIVER DEVELOPER NOTES
23 4.1 Conformance points
24 4.2 "My application needs exclusive control of hardware channels"
25
26 5 SOURCE
27
28 1. Introduction
29 ===============
30
31 The async_tx API provides methods for describing a chain of asynchronous
32 bulk memory transfers/transforms with support for inter-transactional
33 dependencies. It is implemented as a dmaengine client that smooths over
34 the details of different hardware offload engine implementations. Code
35 that is written to the API can optimize for asynchronous operation and
36 the API will fit the chain of operations to the available offload
37 resources.
38
39 2.Genealogy
40 ===========
41
42 The API was initially designed to offload the memory copy and
43 xor-parity-calculations of the md-raid5 driver using the offload engines
44 present in the Intel(R) Xscale series of I/O processors. It also built
45 on the 'dmaengine' layer developed for offloading memory copies in the
46 network stack using Intel(R) I/OAT engines. The following design
47 features surfaced as a result:
48
49 1. implicit synchronous path: users of the API do not need to know if
50 the platform they are running on has offload capabilities. The
51 operation will be offloaded when an engine is available and carried out
52 in software otherwise.
53 2. cross channel dependency chains: the API allows a chain of dependent
54 operations to be submitted, like xor->copy->xor in the raid5 case. The
55 API automatically handles cases where the transition from one operation
56 to another implies a hardware channel switch.
57 3. dmaengine extensions to support multiple clients and operation types
58 beyond 'memcpy'
59
60 3. Usage
61 ========
62
63 3.1 General format of the API
64 -----------------------------
65
66 ::
67
68 struct dma_async_tx_descriptor *
69 async_<operation>(<op specific parameters>, struct async_submit_ctl *submit)
70
71 3.2 Supported operations
72 ------------------------
73
74 ======== ====================================================================
75 memcpy memory copy between a source and a destination buffer
76 memset fill a destination buffer with a byte value
77 xor xor a series of source buffers and write the result to a
78 destination buffer
79 xor_val xor a series of source buffers and set a flag if the
80 result is zero. The implementation attempts to prevent
81 writes to memory
82 pq generate the p+q (raid6 syndrome) from a series of source buffers
83 pq_val validate that a p and or q buffer are in sync with a given series of
84 sources
85 datap (raid6_datap_recov) recover a raid6 data block and the p block
86 from the given sources
87 2data (raid6_2data_recov) recover 2 raid6 data blocks from the given
88 sources
89 ======== ====================================================================
90
91 3.3 Descriptor management
92 -------------------------
93
94 The return value is non-NULL and points to a 'descriptor' when the operation
95 has been queued to execute asynchronously. Descriptors are recycled
96 resources, under control of the offload engine driver, to be reused as
97 operations complete. When an application needs to submit a chain of
98 operations it must guarantee that the descriptor is not automatically recycled
99 before the dependency is submitted. This requires that all descriptors be
100 acknowledged by the application before the offload engine driver is allowed to
101 recycle (or free) the descriptor. A descriptor can be acked by one of the
102 following methods:
103
104 1. setting the ASYNC_TX_ACK flag if no child operations are to be submitted
105 2. submitting an unacknowledged descriptor as a dependency to another
106 async_tx call will implicitly set the acknowledged state.
107 3. calling async_tx_ack() on the descriptor.
108
109 3.4 When does the operation execute?
110 ------------------------------------
111
112 Operations do not immediately issue after return from the
113 async_<operation> call. Offload engine drivers batch operations to
114 improve performance by reducing the number of mmio cycles needed to
115 manage the channel. Once a driver-specific threshold is met the driver
116 automatically issues pending operations. An application can force this
117 event by calling async_tx_issue_pending_all(). This operates on all
118 channels since the application has no knowledge of channel to operation
119 mapping.
120
121 3.5 When does the operation complete?
122 -------------------------------------
123
124 There are two methods for an application to learn about the completion
125 of an operation.
126
127 1. Call dma_wait_for_async_tx(). This call causes the CPU to spin while
128 it polls for the completion of the operation. It handles dependency
129 chains and issuing pending operations.
130 2. Specify a completion callback. The callback routine runs in tasklet
131 context if the offload engine driver supports interrupts, or it is
132 called in application context if the operation is carried out
133 synchronously in software. The callback can be set in the call to
134 async_<operation>, or when the application needs to submit a chain of
135 unknown length it can use the async_trigger_callback() routine to set a
136 completion interrupt/callback at the end of the chain.
137
138 3.6 Constraints
139 ---------------
140
141 1. Calls to async_<operation> are not permitted in IRQ context. Other
142 contexts are permitted provided constraint #2 is not violated.
143 2. Completion callback routines cannot submit new operations. This
144 results in recursion in the synchronous case and spin_locks being
145 acquired twice in the asynchronous case.
146
147 3.7 Example
148 -----------
149
150 Perform a xor->copy->xor operation where each operation depends on the
151 result from the previous operation::
152
153 #include <linux/async_tx.h>
154
155 static void callback(void *param)
156 {
157 complete(param);
158 }
159
160 #define NDISKS 2
161
162 static void run_xor_copy_xor(struct page **xor_srcs,
163 struct page *xor_dest,
164 size_t xor_len,
165 struct page *copy_src,
166 struct page *copy_dest,
167 size_t copy_len)
168 {
169 struct dma_async_tx_descriptor *tx;
170 struct async_submit_ctl submit;
171 addr_conv_t addr_conv[NDISKS];
172 struct completion cmp;
173
174 init_async_submit(&submit, ASYNC_TX_XOR_DROP_DST, NULL, NULL, NULL,
175 addr_conv);
176 tx = async_xor(xor_dest, xor_srcs, 0, NDISKS, xor_len, &submit);
177
178 submit.depend_tx = tx;
179 tx = async_memcpy(copy_dest, copy_src, 0, 0, copy_len, &submit);
180
181 init_completion(&cmp);
182 init_async_submit(&submit, ASYNC_TX_XOR_DROP_DST | ASYNC_TX_ACK, tx,
183 callback, &cmp, addr_conv);
184 tx = async_xor(xor_dest, xor_srcs, 0, NDISKS, xor_len, &submit);
185
186 async_tx_issue_pending_all();
187
188 wait_for_completion(&cmp);
189 }
190
191 See include/linux/async_tx.h for more information on the flags. See the
192 ops_run_* and ops_complete_* routines in drivers/md/raid5.c for more
193 implementation examples.
194
195 4. Driver Development Notes
196 ===========================
197
198 4.1 Conformance points
199 ----------------------
200
201 There are a few conformance points required in dmaengine drivers to
202 accommodate assumptions made by applications using the async_tx API:
203
204 1. Completion callbacks are expected to happen in tasklet context
205 2. dma_async_tx_descriptor fields are never manipulated in IRQ context
206 3. Use async_tx_run_dependencies() in the descriptor clean up path to
207 handle submission of dependent operations
208
209 4.2 "My application needs exclusive control of hardware channels"
210 -----------------------------------------------------------------
211
212 Primarily this requirement arises from cases where a DMA engine driver
213 is being used to support device-to-memory operations. A channel that is
214 performing these operations cannot, for many platform specific reasons,
215 be shared. For these cases the dma_request_channel() interface is
216 provided.
217
218 The interface is::
219
220 struct dma_chan *dma_request_channel(dma_cap_mask_t mask,
221 dma_filter_fn filter_fn,
222 void *filter_param);
223
224 Where dma_filter_fn is defined as::
225
226 typedef bool (*dma_filter_fn)(struct dma_chan *chan, void *filter_param);
227
228 When the optional 'filter_fn' parameter is set to NULL
229 dma_request_channel simply returns the first channel that satisfies the
230 capability mask. Otherwise, when the mask parameter is insufficient for
231 specifying the necessary channel, the filter_fn routine can be used to
232 disposition the available channels in the system. The filter_fn routine
233 is called once for each free channel in the system. Upon seeing a
234 suitable channel filter_fn returns DMA_ACK which flags that channel to
235 be the return value from dma_request_channel. A channel allocated via
236 this interface is exclusive to the caller, until dma_release_channel()
237 is called.
238
239 The DMA_PRIVATE capability flag is used to tag dma devices that should
240 not be used by the general-purpose allocator. It can be set at
241 initialization time if it is known that a channel will always be
242 private. Alternatively, it is set when dma_request_channel() finds an
243 unused "public" channel.
244
245 A couple caveats to note when implementing a driver and consumer:
246
247 1. Once a channel has been privately allocated it will no longer be
248 considered by the general-purpose allocator even after a call to
249 dma_release_channel().
250 2. Since capabilities are specified at the device level a dma_device
251 with multiple channels will either have all channels public, or all
252 channels private.
253
254 5. Source
255 ---------
256
257 include/linux/dmaengine.h:
258 core header file for DMA drivers and api users
259 drivers/dma/dmaengine.c:
260 offload engine channel management routines
261 drivers/dma/:
262 location for offload engine drivers
263 include/linux/async_tx.h:
264 core header file for the async_tx api
265 crypto/async_tx/async_tx.c:
266 async_tx interface to dmaengine and common code
267 crypto/async_tx/async_memcpy.c:
268 copy offload
269 crypto/async_tx/async_xor.c:
270 xor and xor zero sum offload
271

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

비동기 transfer·transform API

1-27

SPDX 라이선스 식별자: `GPL-2.0`

비동기 transfer·transform API

문서 목차는 다음과 같습니다.

  • 1. 소개
  • 2. 계보
  • 3. 사용법: API 일반 형식, 지원 연산, descriptor 관리, 실행·완료 시점, 제약, 예제
  • 4. DMAENGINE driver 개발자 참고 사항: 준수 항목과 hardware channel 독점 제어
  • 5. Source

1. 소개

28-38

1. 소개

`async_tx` API는 transaction 사이의 dependency를 지원하면서 비동기 bulk memory transfer·transform chain을 기술하는 방법을 제공합니다. 서로 다른 hardware offload engine 구현의 세부 차이를 감추는 dmaengine client로 구현됩니다. 이 API를 사용하는 코드는 비동기 연산에 맞게 최적화할 수 있고, API가 연산 chain을 사용 가능한 offload resource에 맞춥니다.

2. 계보

39-59

2. 계보

이 API는 처음에 Intel(R) Xscale I/O processor의 offload engine을 사용하여 `md-raid5` driver의 memory copy와 XOR parity 계산을 offload하기 위해 설계되었습니다. 또한 Intel(R) I/OAT engine으로 network stack의 memory copy를 offload하기 위해 개발된 `dmaengine` 계층을 기반으로 했습니다.

그 결과 다음 설계 특성이 생겼습니다.

  • 1. 암시적 동기 경로: 사용자는 실행 중인 platform의 offload capability를 알 필요가 없습니다. Engine이 있으면 연산을 offload하고, 없으면 software로 수행합니다.
  • 2. Channel 간 dependency chain: RAID5의 `xor->copy->xor`처럼 서로 의존하는 연산 chain을 제출할 수 있습니다. 연산 전환에 hardware channel 변경이 필요하면 API가 자동으로 처리합니다.
  • 3. 여러 client와 `memcpy` 이외의 연산 유형을 지원하도록 dmaengine을 확장했습니다.

3. 사용법과 API 일반 형식

60-70

3. 사용법

3.1 API의 일반 형식

::

  struct dma_async_tx_descriptor *
  async_<operation>(<op specific parameters>, struct async_submit_ctl *submit)

각 `async_<operation>()` 함수는 연산별 매개변수와 submission 제어 정보인 `struct async_submit_ctl *submit`을 받고 `struct dma_async_tx_descriptor *`를 반환합니다.

3.2 지원 연산

71-90

3.2 지원 연산

연산설명
`memcpy`source buffer와 destination buffer 사이의 memory copy
`memset`destination buffer를 하나의 byte 값으로 채움
`xor`여러 source buffer를 XOR하고 결과를 destination buffer에 기록
`xor_val`여러 source buffer를 XOR하여 결과가 0이면 flag를 설정하며, 구현은 memory write를 피하려고 시도함
`pq`여러 source buffer에서 p+q, 즉 RAID6 syndrome을 생성
`pq_val`p 또는 q buffer가 주어진 source 집합과 동기화되어 있는지 검증
`datap``raid6_datap_recov`: 주어진 source에서 RAID6 data block 하나와 p block을 복구
`2data``raid6_2data_recov`: 주어진 source에서 RAID6 data block 두 개를 복구

3.3 Descriptor 관리

91-108

3.3 Descriptor 관리

연산이 비동기 실행 queue에 들어가면 반환값은 NULL이 아니며 `descriptor`를 가리킵니다. Descriptor는 offload engine driver가 관리하는 재활용 resource로, 연산이 완료되면 다시 사용됩니다.

Application이 연산 chain을 제출하려면 dependency를 제출하기 전에 descriptor가 자동 재활용되지 않음을 보장해야 합니다. Offload engine driver가 descriptor를 재활용하거나 해제할 수 있기 전에 application이 모든 descriptor를 acknowledge해야 합니다.

Descriptor를 acknowledge하는 방법은 다음과 같습니다.

  • 1. Child 연산을 제출하지 않을 경우 `ASYNC_TX_ACK` flag를 설정합니다.
  • 2. Acknowledge되지 않은 descriptor를 다른 `async_tx` 호출의 dependency로 제출하면 acknowledge 상태가 암시적으로 설정됩니다.
  • 3. Descriptor에 `async_tx_ack()`를 호출합니다.

3.4 연산 실행 시점

109-120

3.4 연산은 언제 실행되는가?

`async_<operation>` 호출이 반환된 직후 연산이 발행되는 것은 아닙니다. Offload engine driver는 channel 관리에 필요한 MMIO cycle 수를 줄여 성능을 높이기 위해 연산을 batch 처리합니다. Driver별 threshold에 도달하면 pending 연산을 자동 발행합니다.

Application은 `async_tx_issue_pending_all()`을 호출하여 이 event를 강제할 수 있습니다. Application은 channel과 연산의 mapping을 모르므로 이 함수는 모든 channel에 작동합니다.

3.5 연산 완료 시점

121-137

3.5 연산은 언제 완료되는가?

Application이 연산 완료를 알아내는 방법은 두 가지입니다.

  • 1. `dma_wait_for_async_tx()`를 호출합니다. CPU가 연산 완료를 polling하며 spin하도록 하고 dependency chain과 pending 연산 발행을 처리합니다.
  • 2. 완료 callback을 지정합니다. Offload engine driver가 interrupt를 지원하면 callback은 tasklet context에서 실행되고, 연산이 software에서 동기적으로 수행되면 application context에서 호출됩니다. `async_<operation>` 호출에서 callback을 설정할 수 있습니다. 길이를 알 수 없는 chain이라면 `async_trigger_callback()`으로 chain 끝에 완료 interrupt·callback을 설정할 수 있습니다.

3.6 제약

138-146

3.6 제약

  • 1. IRQ context에서는 `async_<operation>`을 호출할 수 없습니다. 2번 제약을 위반하지 않는 다른 context에서는 호출할 수 있습니다.
  • 2. 완료 callback은 새 연산을 제출할 수 없습니다. 동기 경로에서는 recursion이 발생하고 비동기 경로에서는 `spin_lock`을 두 번 획득하게 됩니다.

3.7 예제

147-194

3.7 예제

각 연산이 이전 연산의 결과에 의존하는 `xor->copy->xor` 연산을 수행합니다.

#include <linux/async_tx.h>

static void callback(void *param)
{
        complete(param);
}

#define NDISKS  2

static void run_xor_copy_xor(struct page **xor_srcs,
                             struct page *xor_dest,
                             size_t xor_len,
                             struct page *copy_src,
                             struct page *copy_dest,
                             size_t copy_len)
{
        struct dma_async_tx_descriptor *tx;
        struct async_submit_ctl submit;
        addr_conv_t addr_conv[NDISKS];
        struct completion cmp;

        init_async_submit(&submit, ASYNC_TX_XOR_DROP_DST, NULL, NULL, NULL,
                        addr_conv);
        tx = async_xor(xor_dest, xor_srcs, 0, NDISKS, xor_len, &submit);

        submit.depend_tx = tx;
        tx = async_memcpy(copy_dest, copy_src, 0, 0, copy_len, &submit);

        init_completion(&cmp);
        init_async_submit(&submit, ASYNC_TX_XOR_DROP_DST | ASYNC_TX_ACK, tx,
                        callback, &cmp, addr_conv);
        tx = async_xor(xor_dest, xor_srcs, 0, NDISKS, xor_len, &submit);

        async_tx_issue_pending_all();

        wait_for_completion(&cmp);
}

Flag에 관한 자세한 내용은 `include/linux/async_tx.h`를 참조하십시오. 추가 구현 예제는 `drivers/md/raid5.c`의 `ops_run_*` 및 `ops_complete_*` routine을 참조하십시오.

4. Driver 개발 참고 사항과 준수 항목

195-208

4. Driver 개발 참고 사항

4.1 준수 항목

`async_tx` API application의 가정을 수용하기 위해 dmaengine driver가 지켜야 하는 항목은 다음과 같습니다.

  • 1. 완료 callback은 tasklet context에서 발생해야 합니다.
  • 2. `dma_async_tx_descriptor` field는 IRQ context에서 조작하지 않습니다.
  • 3. Descriptor 정리 경로에서 `async_tx_run_dependencies()`를 사용해 dependent 연산 제출을 처리합니다.

4.2 Hardware channel 독점 제어

209-253

4.2 "Application이 hardware channel을 독점 제어해야 하는 경우"

이 요구는 주로 DMA engine driver가 device-to-memory 연산을 지원하는 경우에 생깁니다. 이러한 연산을 수행하는 channel은 platform별 여러 이유로 공유할 수 없습니다. 이를 위해 `dma_request_channel()` 인터페이스가 제공됩니다.

인터페이스는 다음과 같습니다.

struct dma_chan *dma_request_channel(dma_cap_mask_t mask,
                                     dma_filter_fn filter_fn,
                                     void *filter_param);

`dma_filter_fn`은 다음과 같이 정의됩니다.

typedef bool (*dma_filter_fn)(struct dma_chan *chan, void *filter_param);

선택적 `filter_fn` 매개변수가 NULL이면 `dma_request_channel()`은 capability mask를 만족하는 첫 channel을 반환합니다. Mask만으로 필요한 channel을 지정하기 부족하면 `filter_fn`으로 시스템의 사용 가능한 channel을 선별할 수 있습니다.

`filter_fn`은 시스템의 free channel마다 한 번 호출됩니다. 적합한 channel을 찾으면 `DMA_ACK`를 반환하여 그 channel을 `dma_request_channel()`의 반환값으로 표시합니다. 이 인터페이스로 할당한 channel은 `dma_release_channel()`을 호출할 때까지 호출자가 독점합니다.

`DMA_PRIVATE` capability flag는 범용 allocator가 사용해서는 안 되는 DMA device를 표시합니다. Channel이 항상 private일 것을 안다면 초기화 시 설정할 수 있고, `dma_request_channel()`이 사용하지 않는 public channel을 찾을 때 설정되기도 합니다.

Driver와 consumer 구현 시 다음 주의 사항이 있습니다.

  • 1. Channel이 한 번 private으로 할당되면 `dma_release_channel()` 이후에도 범용 allocator의 고려 대상이 되지 않습니다.
  • 2. Capability는 device 수준에서 지정되므로 여러 channel이 있는 `dma_device`는 모든 channel이 public이거나 모두 private입니다.

5. Source

254-270

5. Source

Source path역할
`include/linux/dmaengine.h`DMA driver와 API 사용자를 위한 core header
`drivers/dma/dmaengine.c`Offload engine channel 관리 routine
`drivers/dma/`Offload engine driver 위치
`include/linux/async_tx.h``async_tx` API core header
`crypto/async_tx/async_tx.c`dmaengine에 대한 `async_tx` interface와 공통 code
`crypto/async_tx/async_memcpy.c`Copy offload
`crypto/async_tx/async_xor.c`XOR 및 XOR zero-sum offload