← Documents Documentation/mm/hmm.rst GitHub 원문 ↗

Linux 6.18.37 · Memory management

Heterogeneous Memory Management (HMM)

CPU와 device의 shared virtual address space, page-table mirroring, ZONE_DEVICE 표현, DMA migration과 exclusive access를 설명합니다.

Source pathDocumentation/mm/hmm.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

hmm.rst:1-441

HMM은 CPU에서 유효한 pointer를 GPU·DSP·FPGA 같은 device에서도 같은 주소로 사용하게 하고, active data를 device memory로 옮기면서 일반 MM 경로와 통합합니다. Page-table mirroring은 주소 의미를 공유하고, 특수 `ZONE_DEVICE struct page`는 CPU가 직접 mapping할 수 없는 memory도 migration과 회계 체계 안에 넣습니다.

Split와 shared address space
모델CPU와 device의 주소결과
Split서로 다른 allocation·pointer복사와 pointer remap 필요
Shared같은 virtual address가 같은 memory 지시Library·compiler가 투명하게 device 활용

전용 allocator 복제 문제를 shared pointer 모델로 바꾸는 HMM의 목표입니다.

Page-table mirroring
CPU virtual addressCPU page tablemmu_interval_notifierDriver update lockDevice page table같은 physical memory

CPU page-table 변경을 notifier로 추적해 device page table과 동기화합니다.

HMM migration 경로
mmap lockmigrate_vma_setup()dst page 할당·DMA copymigrate_vma_pages()device MMU 갱신migrate_vma_finalize()unlock

Device DMA가 copy를 수행하고 migrate_vma helper가 공통 MM 상태 전환을 담당합니다.

Migration 선택과 상태 bit
Symbol역할
`MIGRATE_VMA_SELECT_SYSTEM`System-memory source만 선택
`MIGRATE_VMA_SELECT_DEVICE_PRIVATE`Device-private source만 선택
`MIGRATE_PFN_MIGRATE`해당 entry가 migration 중임
`MIGRATE_PFN_WRITE`Destination mapping에 write 허용
`MMU_NOTIFY_MIGRATE`MMU notifier에 migration event 전달

Source 선택, 진행 여부, write 권한을 구분하는 주요 symbol입니다.

Exclusive device access 종료
make_device_exclusive()특수 swap entryDevice exclusive accessCPU faultMMU notifier원래 mapping 복원

사용자 mapping을 swap entry로 바꾼 뒤 CPU fault가 원래 mapping을 복원합니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 =====================================
2 Heterogeneous Memory Management (HMM)
3 =====================================
4
5 Provide infrastructure and helpers to integrate non-conventional memory (device
6 memory like GPU on board memory) into regular kernel path, with the cornerstone
7 of this being specialized struct page for such memory (see sections 5 to 7 of
8 this document).
9
10 HMM also provides optional helpers for SVM (Share Virtual Memory), i.e.,
11 allowing a device to transparently access program addresses coherently with
12 the CPU meaning that any valid pointer on the CPU is also a valid pointer
13 for the device. This is becoming mandatory to simplify the use of advanced
14 heterogeneous computing where GPU, DSP, or FPGA are used to perform various
15 computations on behalf of a process.
16
17 This document is divided as follows: in the first section I expose the problems
18 related to using device specific memory allocators. In the second section, I
19 expose the hardware limitations that are inherent to many platforms. The third
20 section gives an overview of the HMM design. The fourth section explains how
21 CPU page-table mirroring works and the purpose of HMM in this context. The
22 fifth section deals with how device memory is represented inside the kernel.
23 Finally, the last section presents a new migration helper that allows
24 leveraging the device DMA engine.
25
26 .. contents:: :local:
27
28 Problems of using a device specific memory allocator
29 ====================================================
30
31 Devices with a large amount of on board memory (several gigabytes) like GPUs
32 have historically managed their memory through dedicated driver specific APIs.
33 This creates a disconnect between memory allocated and managed by a device
34 driver and regular application memory (private anonymous, shared memory, or
35 regular file backed memory). From here on I will refer to this aspect as split
36 address space. I use shared address space to refer to the opposite situation:
37 i.e., one in which any application memory region can be used by a device
38 transparently.
39
40 Split address space happens because devices can only access memory allocated
41 through a device specific API. This implies that all memory objects in a program
42 are not equal from the device point of view which complicates large programs
43 that rely on a wide set of libraries.
44
45 Concretely, this means that code that wants to leverage devices like GPUs needs
46 to copy objects between generically allocated memory (malloc, mmap private, mmap
47 share) and memory allocated through the device driver API (this still ends up
48 with an mmap but of the device file).
49
50 For flat data sets (array, grid, image, ...) this isn't too hard to achieve but
51 for complex data sets (list, tree, ...) it's hard to get right. Duplicating a
52 complex data set needs to re-map all the pointer relations between each of its
53 elements. This is error prone and programs get harder to debug because of the
54 duplicate data set and addresses.
55
56 Split address space also means that libraries cannot transparently use data
57 they are getting from the core program or another library and thus each library
58 might have to duplicate its input data set using the device specific memory
59 allocator. Large projects suffer from this and waste resources because of the
60 various memory copies.
61
62 Duplicating each library API to accept as input or output memory allocated by
63 each device specific allocator is not a viable option. It would lead to a
64 combinatorial explosion in the library entry points.
65
66 Finally, with the advance of high level language constructs (in C++ but in
67 other languages too) it is now possible for the compiler to leverage GPUs and
68 other devices without programmer knowledge. Some compiler identified patterns
69 are only doable with a shared address space. It is also more reasonable to use
70 a shared address space for all other patterns.
71
72
73 I/O bus, device memory characteristics
74 ======================================
75
76 I/O buses cripple shared address spaces due to a few limitations. Most I/O
77 buses only allow basic memory access from device to main memory; even cache
78 coherency is often optional. Access to device memory from a CPU is even more
79 limited. More often than not, it is not cache coherent.
80
81 If we only consider the PCIE bus, then a device can access main memory (often
82 through an IOMMU) and be cache coherent with the CPUs. However, it only allows
83 a limited set of atomic operations from the device on main memory. This is worse
84 in the other direction: the CPU can only access a limited range of the device
85 memory and cannot perform atomic operations on it. Thus device memory cannot
86 be considered the same as regular memory from the kernel point of view.
87
88 Another crippling factor is the limited bandwidth (~32GBytes/s with PCIE 4.0
89 and 16 lanes). This is 33 times less than the fastest GPU memory (1 TBytes/s).
90 The final limitation is latency. Access to main memory from the device has an
91 order of magnitude higher latency than when the device accesses its own memory.
92
93 Some platforms are developing new I/O buses or additions/modifications to PCIE
94 to address some of these limitations (OpenCAPI, CCIX). They mainly allow
95 two-way cache coherency between CPU and device and allow all atomic operations the
96 architecture supports. Sadly, not all platforms are following this trend and
97 some major architectures are left without hardware solutions to these problems.
98
99 So for shared address space to make sense, not only must we allow devices to
100 access any memory but we must also permit any memory to be migrated to device
101 memory while the device is using it (blocking CPU access while it happens).
102
103
104 Shared address space and migration
105 ==================================
106
107 HMM intends to provide two main features. The first one is to share the address
108 space by duplicating the CPU page table in the device page table so the same
109 address points to the same physical memory for any valid main memory address in
110 the process address space.
111
112 To achieve this, HMM offers a set of helpers to populate the device page table
113 while keeping track of CPU page table updates. Device page table updates are
114 not as easy as CPU page table updates. To update the device page table, you must
115 allocate a buffer (or use a pool of pre-allocated buffers) and write GPU
116 specific commands in it to perform the update (unmap, cache invalidations, and
117 flush, ...). This cannot be done through common code for all devices. Hence
118 why HMM provides helpers to factor out everything that can be while leaving the
119 hardware specific details to the device driver.
120
121 The second mechanism HMM provides is a new kind of ZONE_DEVICE memory that
122 allows allocating a struct page for each page of device memory. Those pages
123 are special because the CPU cannot map them. However, they allow migrating
124 main memory to device memory using existing migration mechanisms and everything
125 looks like a page that is swapped out to disk from the CPU point of view. Using a
126 struct page gives the easiest and cleanest integration with existing mm
127 mechanisms. Here again, HMM only provides helpers, first to hotplug new ZONE_DEVICE
128 memory for the device memory and second to perform migration. Policy decisions
129 of what and when to migrate is left to the device driver.
130
131 Note that any CPU access to a device page triggers a page fault and a migration
132 back to main memory. For example, when a page backing a given CPU address A is
133 migrated from a main memory page to a device page, then any CPU access to
134 address A triggers a page fault and initiates a migration back to main memory.
135
136 With these two features, HMM not only allows a device to mirror process address
137 space and keeps both CPU and device page tables synchronized, but also
138 leverages device memory by migrating the part of the data set that is actively being
139 used by the device.
140
141
142 Address space mirroring implementation and API
143 ==============================================
144
145 Address space mirroring's main objective is to allow duplication of a range of
146 CPU page table into a device page table; HMM helps keep both synchronized. A
147 device driver that wants to mirror a process address space must start with the
148 registration of a mmu_interval_notifier::
149
150 int mmu_interval_notifier_insert(struct mmu_interval_notifier *interval_sub,
151 struct mm_struct *mm, unsigned long start,
152 unsigned long length,
153 const struct mmu_interval_notifier_ops *ops);
154
155 During the ops->invalidate() callback the device driver must perform the
156 update action to the range (mark range read only, or fully unmap, etc.). The
157 device must complete the update before the driver callback returns.
158
159 When the device driver wants to populate a range of virtual addresses, it can
160 use::
161
162 int hmm_range_fault(struct hmm_range *range);
163
164 It will trigger a page fault on missing or read-only entries if write access is
165 requested (see below). Page faults use the generic mm page fault code path just
166 like a CPU page fault. The usage pattern is::
167
168 int driver_populate_range(...)
169 {
170 struct hmm_range range;
171 ...
172
173 range.notifier = &interval_sub;
174 range.start = ...;
175 range.end = ...;
176 range.hmm_pfns = ...;
177
178 if (!mmget_not_zero(interval_sub->notifier.mm))
179 return -EFAULT;
180
181 again:
182 range.notifier_seq = mmu_interval_read_begin(&interval_sub);
183 mmap_read_lock(mm);
184 ret = hmm_range_fault(&range);
185 if (ret) {
186 mmap_read_unlock(mm);
187 if (ret == -EBUSY)
188 goto again;
189 return ret;
190 }
191 mmap_read_unlock(mm);
192
193 take_lock(driver->update);
194 if (mmu_interval_read_retry(&ni, range.notifier_seq) {
195 release_lock(driver->update);
196 goto again;
197 }
198
199 /* Use pfns array content to update device page table,
200 * under the update lock */
201
202 release_lock(driver->update);
203 return 0;
204 }
205
206 The driver->update lock is the same lock that the driver takes inside its
207 invalidate() callback. That lock must be held before calling
208 mmu_interval_read_retry() to avoid any race with a concurrent CPU page table
209 update.
210
211 Leverage default_flags and pfn_flags_mask
212 =========================================
213
214 The hmm_range struct has 2 fields, default_flags and pfn_flags_mask, that specify
215 fault or snapshot policy for the whole range instead of having to set them
216 for each entry in the pfns array.
217
218 For instance if the device driver wants pages for a range with at least read
219 permission, it sets::
220
221 range->default_flags = HMM_PFN_REQ_FAULT;
222 range->pfn_flags_mask = 0;
223
224 and calls hmm_range_fault() as described above. This will fill fault all pages
225 in the range with at least read permission.
226
227 Now let's say the driver wants to do the same except for one page in the range for
228 which it wants to have write permission. Now driver set::
229
230 range->default_flags = HMM_PFN_REQ_FAULT;
231 range->pfn_flags_mask = HMM_PFN_REQ_WRITE;
232 range->pfns[index_of_write] = HMM_PFN_REQ_WRITE;
233
234 With this, HMM will fault in all pages with at least read (i.e., valid) and for the
235 address == range->start + (index_of_write << PAGE_SHIFT) it will fault with
236 write permission i.e., if the CPU pte does not have write permission set then HMM
237 will call handle_mm_fault().
238
239 After hmm_range_fault completes the flag bits are set to the current state of
240 the page tables, ie HMM_PFN_VALID | HMM_PFN_WRITE will be set if the page is
241 writable.
242
243
244 Represent and manage device memory from core kernel point of view
245 =================================================================
246
247 Several different designs were tried to support device memory. The first one
248 used a device specific data structure to keep information about migrated memory
249 and HMM hooked itself in various places of mm code to handle any access to
250 addresses that were backed by device memory. It turns out that this ended up
251 replicating most of the fields of struct page and also needed many kernel code
252 paths to be updated to understand this new kind of memory.
253
254 Most kernel code paths never try to access the memory behind a page
255 but only care about struct page contents. Because of this, HMM switched to
256 directly using struct page for device memory which left most kernel code paths
257 unaware of the difference. We only need to make sure that no one ever tries to
258 map those pages from the CPU side.
259
260 Migration to and from device memory
261 ===================================
262
263 Because the CPU cannot access device memory directly, the device driver must
264 use hardware DMA or device specific load/store instructions to migrate data.
265 The migrate_vma_setup(), migrate_vma_pages(), and migrate_vma_finalize()
266 functions are designed to make drivers easier to write and to centralize common
267 code across drivers.
268
269 Before migrating pages to device private memory, special device private
270 ``struct page`` needs to be created. These will be used as special "swap"
271 page table entries so that a CPU process will fault if it tries to access
272 a page that has been migrated to device private memory.
273
274 These can be allocated and freed with::
275
276 struct resource *res;
277 struct dev_pagemap pagemap;
278
279 res = request_free_mem_region(&iomem_resource, /* number of bytes */,
280 "name of driver resource");
281 pagemap.type = MEMORY_DEVICE_PRIVATE;
282 pagemap.range.start = res->start;
283 pagemap.range.end = res->end;
284 pagemap.nr_range = 1;
285 pagemap.ops = &device_devmem_ops;
286 memremap_pages(&pagemap, numa_node_id());
287
288 memunmap_pages(&pagemap);
289 release_mem_region(pagemap.range.start, range_len(&pagemap.range));
290
291 There are also devm_request_free_mem_region(), devm_memremap_pages(),
292 devm_memunmap_pages(), and devm_release_mem_region() when the resources can
293 be tied to a ``struct device``.
294
295 The overall migration steps are similar to migrating NUMA pages within system
296 memory (see Documentation/mm/page_migration.rst) but the steps are split
297 between device driver specific code and shared common code:
298
299 1. ``mmap_read_lock()``
300
301 The device driver has to pass a ``struct vm_area_struct`` to
302 migrate_vma_setup() so the mmap_read_lock() or mmap_write_lock() needs to
303 be held for the duration of the migration.
304
305 2. ``migrate_vma_setup(struct migrate_vma *args)``
306
307 The device driver initializes the ``struct migrate_vma`` fields and passes
308 the pointer to migrate_vma_setup(). The ``args->flags`` field is used to
309 filter which source pages should be migrated. For example, setting
310 ``MIGRATE_VMA_SELECT_SYSTEM`` will only migrate system memory and
311 ``MIGRATE_VMA_SELECT_DEVICE_PRIVATE`` will only migrate pages residing in
312 device private memory. If the latter flag is set, the ``args->pgmap_owner``
313 field is used to identify device private pages owned by the driver. This
314 avoids trying to migrate device private pages residing in other devices.
315 Currently only anonymous private VMA ranges can be migrated to or from
316 system memory and device private memory.
317
318 One of the first steps migrate_vma_setup() does is to invalidate other
319 device's MMUs with the ``mmu_notifier_invalidate_range_start(()`` and
320 ``mmu_notifier_invalidate_range_end()`` calls around the page table
321 walks to fill in the ``args->src`` array with PFNs to be migrated.
322 The ``invalidate_range_start()`` callback is passed a
323 ``struct mmu_notifier_range`` with the ``event`` field set to
324 ``MMU_NOTIFY_MIGRATE`` and the ``owner`` field set to
325 the ``args->pgmap_owner`` field passed to migrate_vma_setup(). This
326 allows the device driver to skip the invalidation callback and only
327 invalidate device private MMU mappings that are actually migrating.
328 This is explained more in the next section.
329
330 While walking the page tables, a ``pte_none()`` or ``is_zero_pfn()``
331 entry results in a valid "zero" PFN stored in the ``args->src`` array.
332 This lets the driver allocate device private memory and clear it instead
333 of copying a page of zeros. Valid PTE entries to system memory or
334 device private struct pages will be locked with ``lock_page()``, isolated
335 from the LRU (if system memory since device private pages are not on
336 the LRU), unmapped from the process, and a special migration PTE is
337 inserted in place of the original PTE.
338 migrate_vma_setup() also clears the ``args->dst`` array.
339
340 3. The device driver allocates destination pages and copies source pages to
341 destination pages.
342
343 The driver checks each ``src`` entry to see if the ``MIGRATE_PFN_MIGRATE``
344 bit is set and skips entries that are not migrating. The device driver
345 can also choose to skip migrating a page by not filling in the ``dst``
346 array for that page.
347
348 The driver then allocates either a device private struct page or a
349 system memory page, locks the page with ``lock_page()``, and fills in the
350 ``dst`` array entry with::
351
352 dst[i] = migrate_pfn(page_to_pfn(dpage));
353
354 Now that the driver knows that this page is being migrated, it can
355 invalidate device private MMU mappings and copy device private memory
356 to system memory or another device private page. The core Linux kernel
357 handles CPU page table invalidations so the device driver only has to
358 invalidate its own MMU mappings.
359
360 The driver can use ``migrate_pfn_to_page(src[i])`` to get the
361 ``struct page`` of the source and either copy the source page to the
362 destination or clear the destination device private memory if the pointer
363 is ``NULL`` meaning the source page was not populated in system memory.
364
365 4. ``migrate_vma_pages()``
366
367 This step is where the migration is actually "committed".
368
369 If the source page was a ``pte_none()`` or ``is_zero_pfn()`` page, this
370 is where the newly allocated page is inserted into the CPU's page table.
371 This can fail if a CPU thread faults on the same page. However, the page
372 table is locked and only one of the new pages will be inserted.
373 The device driver will see that the ``MIGRATE_PFN_MIGRATE`` bit is cleared
374 if it loses the race.
375
376 If the source page was locked, isolated, etc. the source ``struct page``
377 information is now copied to destination ``struct page`` finalizing the
378 migration on the CPU side.
379
380 5. Device driver updates device MMU page tables for pages still migrating,
381 rolling back pages not migrating.
382
383 If the ``src`` entry still has ``MIGRATE_PFN_MIGRATE`` bit set, the device
384 driver can update the device MMU and set the write enable bit if the
385 ``MIGRATE_PFN_WRITE`` bit is set.
386
387 6. ``migrate_vma_finalize()``
388
389 This step replaces the special migration page table entry with the new
390 page's page table entry and releases the reference to the source and
391 destination ``struct page``.
392
393 7. ``mmap_read_unlock()``
394
395 The lock can now be released.
396
397 Exclusive access memory
398 =======================
399
400 Some devices have features such as atomic PTE bits that can be used to implement
401 atomic access to system memory. To support atomic operations to a shared virtual
402 memory page such a device needs access to that page which is exclusive of any
403 userspace access from the CPU. The ``make_device_exclusive()`` function
404 can be used to make a memory range inaccessible from userspace.
405
406 This replaces all mappings for pages in the given range with special swap
407 entries. Any attempt to access the swap entry results in a fault which is
408 resolved by replacing the entry with the original mapping. A driver gets
409 notified that the mapping has been changed by MMU notifiers, after which point
410 it will no longer have exclusive access to the page. Exclusive access is
411 guaranteed to last until the driver drops the page lock and page reference, at
412 which point any CPU faults on the page may proceed as described.
413
414 Memory cgroup (memcg) and rss accounting
415 ========================================
416
417 For now, device memory is accounted as any regular page in rss counters (either
418 anonymous if device page is used for anonymous, file if device page is used for
419 file backed page, or shmem if device page is used for shared memory). This is a
420 deliberate choice to keep existing applications, that might start using device
421 memory without knowing about it, running unimpacted.
422
423 A drawback is that the OOM killer might kill an application using a lot of
424 device memory and not a lot of regular system memory and thus not freeing much
425 system memory. We want to gather more real world experience on how applications
426 and system react under memory pressure in the presence of device memory before
427 deciding to account device memory differently.
428
429
430 Same decision was made for memory cgroup. Device memory pages are accounted
431 against same memory cgroup a regular page would be accounted to. This does
432 simplify migration to and from device memory. This also means that migration
433 back from device memory to regular memory cannot fail because it would
434 go above memory cgroup limit. We might revisit this choice later on once we
435 get more experience in how device memory is used and its impact on memory
436 resource control.
437
438
439 Note that device memory can never be pinned by a device driver nor through GUP
440 and thus such memory is always free upon process exit. Or when last reference
441 is dropped in case of shared memory or file backed memory.
442

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

HMM의 목적과 문서 구성

1-26

Heterogeneous Memory Management, 즉 HMM은 GPU board memory 같은 비전통적 device memory를 일반 커널 경로에 통합하는 infrastructure와 helper를 제공합니다. 핵심은 이런 메모리를 위한 특수 `struct page`이며 문서의 5~7절에서 다룹니다.

HMM은 선택적으로 SVM, 즉 Shared Virtual Memory helper도 제공합니다. Device가 CPU와 coherent하게 프로그램 주소에 투명하게 접근해 CPU에서 유효한 pointer가 device에서도 유효하도록 합니다. 프로세스를 대신해 GPU, DSP, FPGA가 여러 계산을 수행하는 고급 heterogeneous computing을 단순화하려면 이 기능이 점점 필수가 되고 있습니다.

문서는 device 전용 allocator 문제, 여러 platform의 hardware 한계, HMM 설계 개요, CPU page-table mirroring과 그 안에서 HMM의 역할, kernel의 device memory 표현, device DMA engine을 활용하는 migration helper 순서로 구성됩니다.

=====================================
Heterogeneous Memory Management (HMM)
=====================================

Provide infrastructure and helpers to integrate non-conventional memory (device
memory like GPU on board memory) into regular kernel path, with the cornerstone
of this being specialized struct page for such memory (see sections 5 to 7 of
this document).

HMM also provides optional helpers for SVM (Share Virtual Memory), i.e.,
allowing a device to transparently access program addresses coherently with
the CPU meaning that any valid pointer on the CPU is also a valid pointer
for the device. This is becoming mandatory to simplify the use of advanced
heterogeneous computing where GPU, DSP, or FPGA are used to perform various
computations on behalf of a process.

This document is divided as follows: in the first section I expose the problems
related to using device specific memory allocators. In the second section, I
expose the hardware limitations that are inherent to many platforms. The third
section gives an overview of the HMM design. The fourth section explains how
CPU page-table mirroring works and the purpose of HMM in this context. The
fifth section deals with how device memory is represented inside the kernel.
Finally, the last section presents a new migration helper that allows
leveraging the device DMA engine.

.. contents:: :local:

Device 전용 allocator와 split address space 문제

27-72

수 GiB의 on-board memory를 가진 GPU 같은 device는 역사적으로 driver 전용 API로 메모리를 관리했습니다. 이 방식은 device driver가 할당·관리하는 메모리와 private anonymous, shared memory, 일반 file-backed memory 같은 보통 application memory를 분리합니다. 문서는 이를 `split address space`라고 하고, 임의의 application memory region을 device가 투명하게 사용할 수 있는 반대 상황을 `shared address space`라고 부릅니다.

Split address space에서는 device가 전용 API로 할당한 메모리만 접근할 수 있어 프로그램의 모든 memory object가 device 관점에서 동등하지 않습니다. 많은 library에 의존하는 큰 프로그램은 이 때문에 복잡해집니다.

GPU를 활용하려는 코드는 `malloc`, private `mmap`, shared `mmap`으로 일반 할당한 메모리와 device driver API로 할당한 메모리 사이에서 object를 복사해야 합니다. 후자도 결국 device file에 대한 `mmap`으로 나타납니다.

Array, grid, image 같은 평평한 data set은 비교적 복사하기 쉽지만 list나 tree 같은 복잡한 data set은 어렵습니다. 복제본의 각 element 사이 pointer 관계를 모두 다시 mapping해야 하므로 오류가 생기기 쉽고, data set과 주소가 중복돼 debugging도 어려워집니다.

Library도 core program이나 다른 library에서 받은 data를 투명하게 사용할 수 없으므로 각 library가 device 전용 allocator로 입력 data set을 복제해야 할 수 있습니다. 큰 project는 여러 memory copy 때문에 자원을 낭비합니다.

각 device 전용 allocator가 만든 메모리를 입력 또는 출력으로 받도록 library API를 모두 복제하는 방법은 entry point가 조합적으로 폭증하므로 실용적이지 않습니다.

C++를 비롯한 고수준 언어 구성의 발전으로 compiler가 programmer의 개입 없이 GPU와 다른 device를 활용할 수 있게 됐습니다. Compiler가 찾은 일부 pattern은 shared address space에서만 구현할 수 있고, 나머지 pattern에도 shared address space가 더 합리적입니다.


Problems of using a device specific memory allocator
====================================================

Devices with a large amount of on board memory (several gigabytes) like GPUs
have historically managed their memory through dedicated driver specific APIs.
This creates a disconnect between memory allocated and managed by a device
driver and regular application memory (private anonymous, shared memory, or
regular file backed memory). From here on I will refer to this aspect as split
address space. I use shared address space to refer to the opposite situation:
i.e., one in which any application memory region can be used by a device
transparently.

Split address space happens because devices can only access memory allocated
through a device specific API. This implies that all memory objects in a program
are not equal from the device point of view which complicates large programs
that rely on a wide set of libraries.

Concretely, this means that code that wants to leverage devices like GPUs needs
to copy objects between generically allocated memory (malloc, mmap private, mmap
share) and memory allocated through the device driver API (this still ends up
with an mmap but of the device file).

For flat data sets (array, grid, image, ...) this isn't too hard to achieve but
for complex data sets (list, tree, ...) it's hard to get right. Duplicating a
complex data set needs to re-map all the pointer relations between each of its
elements. This is error prone and programs get harder to debug because of the
duplicate data set and addresses.

Split address space also means that libraries cannot transparently use data
they are getting from the core program or another library and thus each library
might have to duplicate its input data set using the device specific memory
allocator. Large projects suffer from this and waste resources because of the
various memory copies.

Duplicating each library API to accept as input or output memory allocated by
each device specific allocator is not a viable option. It would lead to a
combinatorial explosion in the library entry points.

Finally, with the advance of high level language constructs (in C++ but in
other languages too) it is now possible for the compiler to leverage GPUs and
other devices without programmer knowledge. Some compiler identified patterns
are only doable with a shared address space. It is also more reasonable to use
a shared address space for all other patterns.

I/O bus와 device memory의 제약

73-103

I/O bus의 제약은 shared address space를 어렵게 합니다. 대부분의 I/O bus는 device에서 main memory로 기본적인 memory access만 허용하고 cache coherency도 선택 사항인 경우가 많습니다. CPU에서 device memory로의 접근은 더 제한적이며 대개 cache coherent하지 않습니다.

PCIe만 보면 device는 흔히 IOMMU를 거쳐 main memory에 접근하고 CPU와 cache coherent할 수 있습니다. 그러나 device가 main memory에 수행할 수 있는 atomic operation 집합은 제한됩니다. 반대 방향은 더 나빠서 CPU는 device memory의 제한된 범위만 접근하고 atomic operation은 수행할 수 없습니다. 따라서 kernel 관점에서 device memory를 일반 memory와 같게 취급할 수 없습니다.

대역폭도 PCIe 4.0 16 lane에서 약 32 GBytes/s로 제한되며, 가장 빠른 GPU memory의 1 TBytes/s보다 33배 낮습니다. Device가 main memory에 접근할 때의 latency도 자체 memory 접근보다 한 자릿수 order만큼 큽니다.

OpenCAPI와 CCIX 같은 새 I/O bus 또는 PCIe 확장은 CPU와 device 사이 양방향 cache coherency와 아키텍처가 지원하는 모든 atomic operation을 제공해 일부 한계를 해결합니다. 하지만 모든 platform이 이 흐름을 따르지는 않아 주요 아키텍처 중에도 hardware 해법이 없는 경우가 있습니다.

Shared address space가 의미 있으려면 device가 임의의 memory에 접근하게 하는 것뿐 아니라, device가 memory를 사용하는 동안 CPU 접근을 막고 그 memory를 device memory로 migration할 수 있어야 합니다.

I/O bus, device memory characteristics
======================================

I/O buses cripple shared address spaces due to a few limitations. Most I/O
buses only allow basic memory access from device to main memory; even cache
coherency is often optional. Access to device memory from a CPU is even more
limited. More often than not, it is not cache coherent.

If we only consider the PCIE bus, then a device can access main memory (often
through an IOMMU) and be cache coherent with the CPUs. However, it only allows
a limited set of atomic operations from the device on main memory. This is worse
in the other direction: the CPU can only access a limited range of the device
memory and cannot perform atomic operations on it. Thus device memory cannot
be considered the same as regular memory from the kernel point of view.

Another crippling factor is the limited bandwidth (~32GBytes/s with PCIE 4.0
and 16 lanes). This is 33 times less than the fastest GPU memory (1 TBytes/s).
The final limitation is latency. Access to main memory from the device has an
order of magnitude higher latency than when the device accesses its own memory.

Some platforms are developing new I/O buses or additions/modifications to PCIE
to address some of these limitations (OpenCAPI, CCIX). They mainly allow
two-way cache coherency between CPU and device and allow all atomic operations the
architecture supports. Sadly, not all platforms are following this trend and
some major architectures are left without hardware solutions to these problems.

So for shared address space to make sense, not only must we allow devices to
access any memory but we must also permit any memory to be migrated to device
memory while the device is using it (blocking CPU access while it happens).

Shared address space와 migration 설계

104-141

HMM의 첫 번째 핵심 기능은 CPU page table을 device page table에 복제해 주소 공간을 공유하는 것입니다. 프로세스 주소 공간의 유효한 main-memory 주소에 대해 같은 주소가 CPU와 device에서 같은 물리 memory를 가리킵니다.

HMM은 CPU page table 갱신을 추적하면서 device page table을 채우는 helper를 제공합니다. Device page table 갱신은 buffer를 할당하거나 미리 할당한 pool을 사용하고, unmap·cache invalidation·flush 등을 수행하는 GPU 전용 command를 써야 하므로 CPU page table 갱신만큼 단순하지 않습니다. 모든 device에 공통 코드로 만들 수 없어 HMM은 공통 부분만 추출하고 hardware 세부 사항은 driver에 맡깁니다.

두 번째 기능은 device memory의 각 page마다 `struct page`를 할당하는 새로운 `ZONE_DEVICE` memory 유형입니다. CPU는 이 특수 page를 mapping할 수 없지만 기존 migration 메커니즘으로 main memory를 device memory로 옮길 수 있고, CPU 관점에서는 disk로 swap out된 page처럼 보입니다.

`struct page`를 사용하면 기존 MM 메커니즘과 가장 쉽고 깔끔하게 통합됩니다. HMM은 device memory를 새 `ZONE_DEVICE` memory로 hotplug하고 migration을 수행하는 helper만 제공하며, 무엇을 언제 migration할지에 대한 policy는 device driver가 결정합니다.

CPU가 device page에 접근하면 page fault가 발생해 main memory로 다시 migration됩니다. 예를 들어 CPU 주소 A를 뒷받침하던 page가 main memory에서 device page로 이동했다면 CPU의 A 접근이 fault를 일으키고 복귀 migration을 시작합니다.

이 두 기능으로 HMM은 device가 프로세스 주소 공간을 mirror하고 CPU·device page table을 동기화할 뿐 아니라, device가 활발히 사용하는 data set 일부를 device memory로 옮겨 활용하게 합니다.

Shared address space and migration
==================================

HMM intends to provide two main features. The first one is to share the address
space by duplicating the CPU page table in the device page table so the same
address points to the same physical memory for any valid main memory address in
the process address space.

To achieve this, HMM offers a set of helpers to populate the device page table
while keeping track of CPU page table updates. Device page table updates are
not as easy as CPU page table updates. To update the device page table, you must
allocate a buffer (or use a pool of pre-allocated buffers) and write GPU
specific commands in it to perform the update (unmap, cache invalidations, and
flush, ...). This cannot be done through common code for all devices. Hence
why HMM provides helpers to factor out everything that can be while leaving the
hardware specific details to the device driver.

The second mechanism HMM provides is a new kind of ZONE_DEVICE memory that
allows allocating a struct page for each page of device memory. Those pages
are special because the CPU cannot map them. However, they allow migrating
main memory to device memory using existing migration mechanisms and everything
looks like a page that is swapped out to disk from the CPU point of view. Using a
struct page gives the easiest and cleanest integration with existing mm
mechanisms. Here again, HMM only provides helpers, first to hotplug new ZONE_DEVICE
memory for the device memory and second to perform migration. Policy decisions
of what and when to migrate is left to the device driver.

Note that any CPU access to a device page triggers a page fault and a migration
back to main memory. For example, when a page backing a given CPU address A is
migrated from a main memory page to a device page, then any CPU access to
address A triggers a page fault and initiates a migration back to main memory.

With these two features, HMM not only allows a device to mirror process address
space and keeps both CPU and device page tables synchronized, but also
leverages device memory by migrating the part of the data set that is actively being
used by the device.

주소 공간 mirroring 구현과 API

142-210

Address-space mirroring의 주목적은 CPU page table 범위를 device page table에 복제하고 둘을 동기화하는 것입니다. 프로세스 주소 공간을 mirror하려는 driver는 `mmu_interval_notifier` 등록부터 시작합니다.

int mmu_interval_notifier_insert(struct mmu_interval_notifier *interval_sub,
				  struct mm_struct *mm, unsigned long start,
				  unsigned long length,
				  const struct mmu_interval_notifier_ops *ops);

`ops->invalidate()` callback 동안 driver는 해당 범위를 read-only로 표시하거나 완전히 unmap하는 등 갱신 동작을 수행해야 합니다. Callback이 반환되기 전에 device가 갱신을 완료해야 합니다.

Driver가 virtual address 범위를 채우려면 `hmm_range_fault()`를 사용합니다.

int hmm_range_fault(struct hmm_range *range);

Write access가 요청된 경우 missing 또는 read-only entry에 page fault를 발생시킵니다. Page fault는 CPU fault와 같은 generic MM page-fault 경로를 사용합니다. 일반적인 사용 pattern은 다음과 같습니다.

int driver_populate_range(...)
{
struct hmm_range range;
...

range.notifier = &interval_sub;
range.start = ...;
range.end = ...;
range.hmm_pfns = ...;

if (!mmget_not_zero(interval_sub->notifier.mm))
return -EFAULT;

again:
range.notifier_seq = mmu_interval_read_begin(&interval_sub);
mmap_read_lock(mm);
ret = hmm_range_fault(&range);
if (ret) {
mmap_read_unlock(mm);
if (ret == -EBUSY)
goto again;
return ret;
}
mmap_read_unlock(mm);

take_lock(driver->update);
if (mmu_interval_read_retry(&ni, range.notifier_seq) {
release_lock(driver->update);
goto again;
}

/* Use pfns array content to update device page table,
* under the update lock */

release_lock(driver->update);
return 0;
}

`driver->update` lock은 driver가 `invalidate()` callback 안에서 잡는 lock과 같습니다. 동시 CPU page-table 갱신과 race가 생기지 않도록 `mmu_interval_read_retry()`를 호출하기 전에 이 lock을 보유해야 합니다.

Address space mirroring implementation and API
==============================================

Address space mirroring's main objective is to allow duplication of a range of
CPU page table into a device page table; HMM helps keep both synchronized. A
device driver that wants to mirror a process address space must start with the
registration of a mmu_interval_notifier::

 int mmu_interval_notifier_insert(struct mmu_interval_notifier *interval_sub,
                                  struct mm_struct *mm, unsigned long start,
                                  unsigned long length,
                                  const struct mmu_interval_notifier_ops *ops);

During the ops->invalidate() callback the device driver must perform the
update action to the range (mark range read only, or fully unmap, etc.). The
device must complete the update before the driver callback returns.

When the device driver wants to populate a range of virtual addresses, it can
use::

  int hmm_range_fault(struct hmm_range *range);

It will trigger a page fault on missing or read-only entries if write access is
requested (see below). Page faults use the generic mm page fault code path just
like a CPU page fault. The usage pattern is::

 int driver_populate_range(...)
 {
      struct hmm_range range;
      ...

      range.notifier = &interval_sub;
      range.start = ...;
      range.end = ...;
      range.hmm_pfns = ...;

      if (!mmget_not_zero(interval_sub->notifier.mm))
          return -EFAULT;

 again:
      range.notifier_seq = mmu_interval_read_begin(&interval_sub);
      mmap_read_lock(mm);
      ret = hmm_range_fault(&range);
      if (ret) {
          mmap_read_unlock(mm);
          if (ret == -EBUSY)
                 goto again;
          return ret;
      }
      mmap_read_unlock(mm);

      take_lock(driver->update);
      if (mmu_interval_read_retry(&ni, range.notifier_seq) {
          release_lock(driver->update);
          goto again;
      }

      /* Use pfns array content to update device page table,
       * under the update lock */

      release_lock(driver->update);
      return 0;
 }

The driver->update lock is the same lock that the driver takes inside its
invalidate() callback. That lock must be held before calling
mmu_interval_read_retry() to avoid any race with a concurrent CPU page table
update.

default_flags와 pfn_flags_mask 활용

211-243

`struct hmm_range`의 `default_flags`와 `pfn_flags_mask`는 `pfns` array entry마다 설정하는 대신 전체 range의 fault 또는 snapshot policy를 지정합니다.

Driver가 range의 모든 page에 적어도 read permission을 원하면 다음처럼 설정하고 `hmm_range_fault()`를 호출합니다. 그러면 range의 모든 page가 최소 read permission으로 fault-in됩니다.

range->default_flags = HMM_PFN_REQ_FAULT;
range->pfn_flags_mask = 0;

같은 range에서 page 하나만 write permission이 필요하다면 mask와 해당 `pfns` entry를 다음처럼 설정합니다.

range->default_flags = HMM_PFN_REQ_FAULT;
range->pfn_flags_mask = HMM_PFN_REQ_WRITE;
range->pfns[index_of_write] = HMM_PFN_REQ_WRITE;

이 설정으로 모든 page는 최소 read, 즉 valid 상태로 fault-in되고 `range->start + (index_of_write << PAGE_SHIFT)` 주소는 write permission으로 fault-in됩니다. CPU PTE에 write permission이 없다면 HMM이 `handle_mm_fault()`를 호출합니다.

`hmm_range_fault()`가 끝나면 flag bit가 현재 page-table 상태로 설정됩니다. 예를 들어 page가 writable이면 `HMM_PFN_VALID | HMM_PFN_WRITE`가 설정됩니다.

Leverage default_flags and pfn_flags_mask
=========================================

The hmm_range struct has 2 fields, default_flags and pfn_flags_mask, that specify
fault or snapshot policy for the whole range instead of having to set them
for each entry in the pfns array.

For instance if the device driver wants pages for a range with at least read
permission, it sets::

    range->default_flags = HMM_PFN_REQ_FAULT;
    range->pfn_flags_mask = 0;

and calls hmm_range_fault() as described above. This will fill fault all pages
in the range with at least read permission.

Now let's say the driver wants to do the same except for one page in the range for
which it wants to have write permission. Now driver set::

    range->default_flags = HMM_PFN_REQ_FAULT;
    range->pfn_flags_mask = HMM_PFN_REQ_WRITE;
    range->pfns[index_of_write] = HMM_PFN_REQ_WRITE;

With this, HMM will fault in all pages with at least read (i.e., valid) and for the
address == range->start + (index_of_write << PAGE_SHIFT) it will fault with
write permission i.e., if the CPU pte does not have write permission set then HMM
will call handle_mm_fault().

After hmm_range_fault completes the flag bits are set to the current state of
the page tables, ie HMM_PFN_VALID | HMM_PFN_WRITE will be set if the page is
writable.

Core kernel의 device memory 표현

244-259

Device memory를 지원하기 위해 여러 설계가 시도됐습니다. 첫 설계는 migration된 memory 정보를 device 전용 data structure에 보관하고, device memory가 뒷받침하는 주소 접근을 처리하도록 HMM을 여러 MM code 위치에 연결했습니다.

하지만 이 방식은 `struct page` field 대부분을 복제했고, 많은 kernel code path가 새 memory 유형을 이해하도록 수정해야 했습니다.

대부분의 kernel code path는 page 뒤의 memory에 직접 접근하지 않고 `struct page` 내용만 봅니다. 그래서 HMM은 device memory에도 `struct page`를 직접 사용하도록 바꿨고, 대부분의 code path는 차이를 알 필요가 없어졌습니다. 단 CPU 쪽에서 이 page를 mapping하려는 시도가 없도록 보장해야 합니다.

Represent and manage device memory from core kernel point of view
=================================================================

Several different designs were tried to support device memory. The first one
used a device specific data structure to keep information about migrated memory
and HMM hooked itself in various places of mm code to handle any access to
addresses that were backed by device memory. It turns out that this ended up
replicating most of the fields of struct page and also needed many kernel code
paths to be updated to understand this new kind of memory.

Most kernel code paths never try to access the memory behind a page
but only care about struct page contents. Because of this, HMM switched to
directly using struct page for device memory which left most kernel code paths
unaware of the difference. We only need to make sure that no one ever tries to
map those pages from the CPU side.

Device-private memory 준비

260-294

CPU는 device memory에 직접 접근할 수 없으므로 driver는 hardware DMA 또는 device 전용 load/store instruction으로 data를 migration해야 합니다. `migrate_vma_setup()`, `migrate_vma_pages()`, `migrate_vma_finalize()`는 driver 작성을 쉽게 하고 공통 코드를 한곳에 모읍니다.

Page를 device-private memory로 옮기기 전에 특수 device-private `struct page`를 만들어야 합니다. 이 page는 특수 swap page-table entry로 사용되므로 CPU process가 device-private memory로 이동한 page에 접근하면 fault가 발생합니다.

할당과 해제 예는 다음과 같습니다.

struct resource *res;
struct dev_pagemap pagemap;

res = request_free_mem_region(&iomem_resource, /* number of bytes */,
"name of driver resource");
pagemap.type = MEMORY_DEVICE_PRIVATE;
pagemap.range.start = res->start;
pagemap.range.end = res->end;
pagemap.nr_range = 1;
pagemap.ops = &device_devmem_ops;
memremap_pages(&pagemap, numa_node_id());

memunmap_pages(&pagemap);
release_mem_region(pagemap.range.start, range_len(&pagemap.range));

Resource 수명을 `struct device`에 연결할 수 있다면 `devm_request_free_mem_region()`, `devm_memremap_pages()`, `devm_memunmap_pages()`, `devm_release_mem_region()`도 사용할 수 있습니다.

Migration to and from device memory
===================================

Because the CPU cannot access device memory directly, the device driver must
use hardware DMA or device specific load/store instructions to migrate data.
The migrate_vma_setup(), migrate_vma_pages(), and migrate_vma_finalize()
functions are designed to make drivers easier to write and to centralize common
code across drivers.

Before migrating pages to device private memory, special device private
``struct page`` needs to be created. These will be used as special "swap"
page table entries so that a CPU process will fault if it tries to access
a page that has been migrated to device private memory.

These can be allocated and freed with::

    struct resource *res;
    struct dev_pagemap pagemap;

    res = request_free_mem_region(&iomem_resource, /* number of bytes */,
                                  "name of driver resource");
    pagemap.type = MEMORY_DEVICE_PRIVATE;
    pagemap.range.start = res->start;
    pagemap.range.end = res->end;
    pagemap.nr_range = 1;
    pagemap.ops = &device_devmem_ops;
    memremap_pages(&pagemap, numa_node_id());

    memunmap_pages(&pagemap);
    release_mem_region(pagemap.range.start, range_len(&pagemap.range));

There are also devm_request_free_mem_region(), devm_memremap_pages(),
devm_memunmap_pages(), and devm_release_mem_region() when the resources can
be tied to a ``struct device``.

Device memory migration 7단계

295-396

전체 migration 절차는 system memory 안에서 NUMA page를 옮기는 과정과 비슷하지만 device driver 전용 코드와 공유 공통 코드로 나뉩니다.

1. `mmap_read_lock()`: Driver는 `migrate_vma_setup()`에 `struct vm_area_struct`를 전달해야 하므로 migration 전체 동안 `mmap_read_lock()` 또는 `mmap_write_lock()`을 유지합니다.

2. `migrate_vma_setup(struct migrate_vma *args)`: Driver가 `struct migrate_vma` field를 초기화해 전달합니다. `args->flags`는 migration할 source page를 거릅니다. `MIGRATE_VMA_SELECT_SYSTEM`은 system memory만, `MIGRATE_VMA_SELECT_DEVICE_PRIVATE`은 device-private page만 선택합니다. 후자를 쓰면 `args->pgmap_owner`로 driver가 소유한 page를 식별해 다른 device의 private page를 잘못 옮기지 않습니다. 현재 system memory와 device-private memory 사이에서는 anonymous private VMA range만 migration할 수 있습니다.

`migrate_vma_setup()`은 page table을 순회해 migration PFN으로 `args->src`를 채우는 동안 `mmu_notifier_invalidate_range_start()`와 `mmu_notifier_invalidate_range_end()`로 다른 device의 MMU를 invalidate합니다. `invalidate_range_start()`에는 `event=MMU_NOTIFY_MIGRATE`, `owner=args->pgmap_owner`인 `struct mmu_notifier_range`가 전달됩니다. Driver는 이를 이용해 자기 migration인 callback을 건너뛰고 실제 이동 중인 device-private MMU mapping만 invalidate할 수 있습니다.

Page-table 순회 중 `pte_none()` 또는 `is_zero_pfn()` entry는 유효한 zero PFN으로 `args->src`에 저장됩니다. Driver는 zero page를 복사하지 않고 device-private memory를 할당해 clear할 수 있습니다. System memory나 device-private `struct page`를 가리키는 유효 PTE는 `lock_page()`로 잠그고, system memory라면 LRU에서 격리하며, process에서 unmap한 뒤 원래 PTE 대신 특수 migration PTE를 넣습니다. Device-private page는 LRU에 없습니다. `migrate_vma_setup()`은 `args->dst`도 지웁니다.

3. Driver가 destination page를 할당하고 source를 복사합니다. 각 `src` entry의 `MIGRATE_PFN_MIGRATE` bit를 확인하고 이동하지 않는 entry는 건너뜁니다. 특정 page의 `dst`를 채우지 않아 선택적으로 migration을 취소할 수도 있습니다.

Driver는 device-private `struct page` 또는 system-memory page를 할당하고 `lock_page()`로 잠근 뒤 다음처럼 `dst` entry를 채웁니다.

dst[i] = migrate_pfn(page_to_pfn(dpage));

이제 driver는 이동 대상임을 알고 device-private MMU mapping을 invalidate한 뒤 device-private memory를 system memory 또는 다른 device-private page로 복사할 수 있습니다. Core kernel이 CPU page-table invalidation을 처리하므로 driver는 자신의 MMU mapping만 invalidate하면 됩니다.

`migrate_pfn_to_page(src[i])`로 source `struct page`를 얻어 destination으로 복사합니다. Pointer가 `NULL`이면 source page가 system memory에 populate되지 않은 것이므로 destination device-private memory를 clear합니다.

4. `migrate_vma_pages()`: 실제 migration을 commit합니다. Source가 `pte_none()` 또는 `is_zero_pfn()`이었다면 새 page를 CPU page table에 넣습니다. 같은 page에서 CPU thread가 fault를 일으키면 실패할 수 있지만 page table이 잠겨 있어 새 page 하나만 삽입됩니다. Driver가 race에서 지면 `MIGRATE_PFN_MIGRATE` bit가 지워진 것을 확인합니다. Source가 lock·isolate된 page였다면 source `struct page` 정보를 destination에 복사해 CPU 쪽 migration을 마무리합니다.

5. Driver는 아직 이동 중인 page에 대해 device MMU page table을 갱신하고 이동하지 않는 page는 rollback합니다. `src` entry에 `MIGRATE_PFN_MIGRATE`가 남아 있으면 MMU를 갱신하며, `MIGRATE_PFN_WRITE`가 설정된 경우 write-enable bit도 켤 수 있습니다.

6. `migrate_vma_finalize()`: 특수 migration PTE를 새 page의 PTE로 바꾸고 source와 destination `struct page` reference를 해제합니다.

7. `mmap_read_unlock()`: 이제 migration 동안 유지한 lock을 해제할 수 있습니다.

The overall migration steps are similar to migrating NUMA pages within system
memory (see Documentation/mm/page_migration.rst) but the steps are split
between device driver specific code and shared common code:

1. ``mmap_read_lock()``

   The device driver has to pass a ``struct vm_area_struct`` to
   migrate_vma_setup() so the mmap_read_lock() or mmap_write_lock() needs to
   be held for the duration of the migration.

2. ``migrate_vma_setup(struct migrate_vma *args)``

   The device driver initializes the ``struct migrate_vma`` fields and passes
   the pointer to migrate_vma_setup(). The ``args->flags`` field is used to
   filter which source pages should be migrated. For example, setting
   ``MIGRATE_VMA_SELECT_SYSTEM`` will only migrate system memory and
   ``MIGRATE_VMA_SELECT_DEVICE_PRIVATE`` will only migrate pages residing in
   device private memory. If the latter flag is set, the ``args->pgmap_owner``
   field is used to identify device private pages owned by the driver. This
   avoids trying to migrate device private pages residing in other devices.
   Currently only anonymous private VMA ranges can be migrated to or from
   system memory and device private memory.

   One of the first steps migrate_vma_setup() does is to invalidate other
   device's MMUs with the ``mmu_notifier_invalidate_range_start(()`` and
   ``mmu_notifier_invalidate_range_end()`` calls around the page table
   walks to fill in the ``args->src`` array with PFNs to be migrated.
   The ``invalidate_range_start()`` callback is passed a
   ``struct mmu_notifier_range`` with the ``event`` field set to
   ``MMU_NOTIFY_MIGRATE`` and the ``owner`` field set to
   the ``args->pgmap_owner`` field passed to migrate_vma_setup(). This
   allows the device driver to skip the invalidation callback and only
   invalidate device private MMU mappings that are actually migrating.
   This is explained more in the next section.

   While walking the page tables, a ``pte_none()`` or ``is_zero_pfn()``
   entry results in a valid "zero" PFN stored in the ``args->src`` array.
   This lets the driver allocate device private memory and clear it instead
   of copying a page of zeros. Valid PTE entries to system memory or
   device private struct pages will be locked with ``lock_page()``, isolated
   from the LRU (if system memory since device private pages are not on
   the LRU), unmapped from the process, and a special migration PTE is
   inserted in place of the original PTE.
   migrate_vma_setup() also clears the ``args->dst`` array.

3. The device driver allocates destination pages and copies source pages to
   destination pages.

   The driver checks each ``src`` entry to see if the ``MIGRATE_PFN_MIGRATE``
   bit is set and skips entries that are not migrating. The device driver
   can also choose to skip migrating a page by not filling in the ``dst``
   array for that page.

   The driver then allocates either a device private struct page or a
   system memory page, locks the page with ``lock_page()``, and fills in the
   ``dst`` array entry with::

     dst[i] = migrate_pfn(page_to_pfn(dpage));

   Now that the driver knows that this page is being migrated, it can
   invalidate device private MMU mappings and copy device private memory
   to system memory or another device private page. The core Linux kernel
   handles CPU page table invalidations so the device driver only has to
   invalidate its own MMU mappings.

   The driver can use ``migrate_pfn_to_page(src[i])`` to get the
   ``struct page`` of the source and either copy the source page to the
   destination or clear the destination device private memory if the pointer
   is ``NULL`` meaning the source page was not populated in system memory.

4. ``migrate_vma_pages()``

   This step is where the migration is actually "committed".

   If the source page was a ``pte_none()`` or ``is_zero_pfn()`` page, this
   is where the newly allocated page is inserted into the CPU's page table.
   This can fail if a CPU thread faults on the same page. However, the page
   table is locked and only one of the new pages will be inserted.
   The device driver will see that the ``MIGRATE_PFN_MIGRATE`` bit is cleared
   if it loses the race.

   If the source page was locked, isolated, etc. the source ``struct page``
   information is now copied to destination ``struct page`` finalizing the
   migration on the CPU side.

5. Device driver updates device MMU page tables for pages still migrating,
   rolling back pages not migrating.

   If the ``src`` entry still has ``MIGRATE_PFN_MIGRATE`` bit set, the device
   driver can update the device MMU and set the write enable bit if the
   ``MIGRATE_PFN_WRITE`` bit is set.

6. ``migrate_vma_finalize()``

   This step replaces the special migration page table entry with the new
   page's page table entry and releases the reference to the source and
   destination ``struct page``.

7. ``mmap_read_unlock()``

   The lock can now be released.

Exclusive access memory

397-413

일부 device는 system memory에 대한 atomic access를 구현할 수 있는 atomic PTE bit 같은 기능을 갖습니다. Shared virtual memory page에 atomic operation을 지원하려면 해당 device가 CPU 사용자 공간 접근과 배타적으로 page를 사용해야 합니다.

`make_device_exclusive()`는 memory range를 사용자 공간에서 접근할 수 없게 만듭니다. 주어진 range의 모든 page mapping을 특수 swap entry로 바꿉니다.

Swap entry에 접근하면 fault가 발생하고 원래 mapping으로 entry를 되돌려 해결합니다. Mapping 변경은 MMU notifier로 driver에 통지되며, 그 시점부터 driver는 page에 대한 exclusive access를 잃습니다. Driver가 page lock과 page reference를 놓을 때까지 exclusive access가 보장되고, 놓은 뒤에는 CPU page fault가 앞의 방식으로 진행할 수 있습니다.

Exclusive access memory
=======================

Some devices have features such as atomic PTE bits that can be used to implement
atomic access to system memory. To support atomic operations to a shared virtual
memory page such a device needs access to that page which is exclusive of any
userspace access from the CPU. The ``make_device_exclusive()`` function
can be used to make a memory range inaccessible from userspace.

This replaces all mappings for pages in the given range with special swap
entries. Any attempt to access the swap entry results in a fault which is
resolved by replacing the entry with the original mapping. A driver gets
notified that the mapping has been changed by MMU notifiers, after which point
it will no longer have exclusive access to the page. Exclusive access is
guaranteed to last until the driver drops the page lock and page reference, at
which point any CPU faults on the page may proceed as described.

Memcg와 RSS 회계

414-441

현재 device memory는 일반 page처럼 RSS counter에 집계됩니다. Anonymous 용도라면 anonymous, file-backed page 용도라면 file, shared memory 용도라면 shmem으로 계산합니다. Device memory 사용을 알지 못한 채 시작하는 기존 application에 영향을 주지 않기 위한 의도적인 선택입니다.

단점은 OOM killer가 system memory는 적게 쓰고 device memory를 많이 쓰는 application을 죽여 실제 system memory를 거의 확보하지 못할 수 있다는 점입니다. Device memory가 있는 상태에서 memory pressure에 application과 system이 어떻게 반응하는지 실제 경험을 더 모은 뒤 다른 회계 방식을 결정하려 합니다.

Memory cgroup에도 같은 결정을 적용합니다. Device-memory page는 일반 page라면 귀속됐을 동일한 memcg에 집계됩니다. 이 방식은 device memory로 오가는 migration을 단순화하고, device memory에서 일반 memory로 돌아오는 migration이 memcg limit 초과 때문에 실패하지 않게 합니다. Device memory 사용과 resource control 영향에 대한 경험이 쌓이면 재검토할 수 있습니다.

Device memory는 driver나 GUP로 pin할 수 없으므로 process가 종료되면 항상 해제됩니다. Shared memory나 file-backed memory라면 마지막 reference가 사라질 때 해제됩니다.

Memory cgroup (memcg) and rss accounting
========================================

For now, device memory is accounted as any regular page in rss counters (either
anonymous if device page is used for anonymous, file if device page is used for
file backed page, or shmem if device page is used for shared memory). This is a
deliberate choice to keep existing applications, that might start using device
memory without knowing about it, running unimpacted.

A drawback is that the OOM killer might kill an application using a lot of
device memory and not a lot of regular system memory and thus not freeing much
system memory. We want to gather more real world experience on how applications
and system react under memory pressure in the presence of device memory before
deciding to account device memory differently.


Same decision was made for memory cgroup. Device memory pages are accounted
against same memory cgroup a regular page would be accounted to. This does
simplify migration to and from device memory. This also means that migration
back from device memory to regular memory cannot fail because it would
go above memory cgroup limit. We might revisit this choice later on once we
get more experience in how device memory is used and its impact on memory
resource control.


Note that device memory can never be pinned by a device driver nor through GUP
and thus such memory is always free upon process exit. Or when last reference
is dropped in case of shared memory or file backed memory.