← Documents Documentation/admin-guide/mm/concepts.rst GitHub 원문 ↗

Linux 6.18.37 · Administration / Memory Management

Concepts overview

virtual memory부터 huge page, zone, NUMA node, reclaim, compaction, OOM killer까지 Linux memory management의 기본 개념을 설명합니다.

Source pathDocumentation/admin-guide/mm/concepts.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

Memory management 개념 지도

concepts.rst:1-220

주소 변환과 page 계층에서 시작해 memory 배치, file·anonymous data의 수명, pressure 대응과 최종 OOM 복구까지 이어지는 관계를 한 문서에서 설명합니다.

개념역할
Virtual memoryphysical address range를 추상화하고 demand paging, 보호, process 간 제어된 공유를 제공합니다.
Huge pages상위 page-table level에서 큰 page를 mapping해 TLB pressure와 miss를 줄입니다.
ZonesDMA 가능 여부와 kernel mapping 방식 같은 hardware 제약에 따라 page를 분류합니다.
NUMA nodesprocessor와의 거리별 memory bank마다 독립적인 zone, page list, 통계를 둡니다.
Page cachefile read/write 데이터를 RAM에 보관하고 dirty page를 backing storage와 동기화합니다.
Anonymous memoryfilesystem backing 없이 stack, heap, mmap(2)에 쓰이며 필요하면 swap됩니다.
Reclaim재사용 가능한 page를 free하거나 backing storage로 evict해 allocation 여유를 만듭니다.
Compaction사용 중인 page를 이동해 큰 physically contiguous free range를 만듭니다.
OOM killerreclaim으로도 진행할 수 없을 때 task 하나를 종료해 시스템 전체를 살립니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 =================
2 Concepts overview
3 =================
4
5 The memory management in Linux is a complex system that evolved over the
6 years and included more and more functionality to support a variety of
7 systems from MMU-less microcontrollers to supercomputers. The memory
8 management for systems without an MMU is called ``nommu`` and it
9 definitely deserves a dedicated document, which hopefully will be
10 eventually written. Yet, although some of the concepts are the same,
11 here we assume that an MMU is available and a CPU can translate a virtual
12 address to a physical address.
13
14 .. contents:: :local:
15
16 Virtual Memory Primer
17 =====================
18
19 The physical memory in a computer system is a limited resource and
20 even for systems that support memory hotplug there is a hard limit on
21 the amount of memory that can be installed. The physical memory is not
22 necessarily contiguous; it might be accessible as a set of distinct
23 address ranges. Besides, different CPU architectures, and even
24 different implementations of the same architecture have different views
25 of how these address ranges are defined.
26
27 All this makes dealing directly with physical memory quite complex and
28 to avoid this complexity a concept of virtual memory was developed.
29
30 The virtual memory abstracts the details of physical memory from the
31 application software, allows to keep only needed information in the
32 physical memory (demand paging) and provides a mechanism for the
33 protection and controlled sharing of data between processes.
34
35 With virtual memory, each and every memory access uses a virtual
36 address. When the CPU decodes an instruction that reads (or
37 writes) from (or to) the system memory, it translates the `virtual`
38 address encoded in that instruction to a `physical` address that the
39 memory controller can understand.
40
41 The physical system memory is divided into page frames, or pages. The
42 size of each page is architecture specific. Some architectures allow
43 selection of the page size from several supported values; this
44 selection is performed at the kernel build time by setting an
45 appropriate kernel configuration option.
46
47 Each physical memory page can be mapped as one or more virtual
48 pages. These mappings are described by page tables that allow
49 translation from a virtual address used by programs to the physical
50 memory address. The page tables are organized hierarchically.
51
52 The tables at the lowest level of the hierarchy contain physical
53 addresses of actual pages used by the software. The tables at higher
54 levels contain physical addresses of the pages belonging to the lower
55 levels. The pointer to the top level page table resides in a
56 register. When the CPU performs the address translation, it uses this
57 register to access the top level page table. The high bits of the
58 virtual address are used to index an entry in the top level page
59 table. That entry is then used to access the next level in the
60 hierarchy with the next bits of the virtual address as the index to
61 that level page table. The lowest bits in the virtual address define
62 the offset inside the actual page.
63
64 Huge Pages
65 ==========
66
67 The address translation requires several memory accesses and memory
68 accesses are slow relatively to CPU speed. To avoid spending precious
69 processor cycles on the address translation, CPUs maintain a cache of
70 such translations called Translation Lookaside Buffer (or
71 TLB). Usually TLB is pretty scarce resource and applications with
72 large memory working set will experience performance hit because of
73 TLB misses.
74
75 Many modern CPU architectures allow mapping of the memory pages
76 directly by the higher levels in the page table. For instance, on x86,
77 it is possible to map 2M and even 1G pages using entries in the second
78 and the third level page tables. In Linux such pages are called
79 `huge`. Usage of huge pages significantly reduces pressure on TLB,
80 improves TLB hit-rate and thus improves overall system performance.
81
82 There are two mechanisms in Linux that enable mapping of the physical
83 memory with the huge pages. The first one is `HugeTLB filesystem`, or
84 hugetlbfs. It is a pseudo filesystem that uses RAM as its backing
85 store. For the files created in this filesystem the data resides in
86 the memory and mapped using huge pages. The hugetlbfs is described at
87 Documentation/admin-guide/mm/hugetlbpage.rst.
88
89 Another, more recent, mechanism that enables use of the huge pages is
90 called `Transparent HugePages`, or THP. Unlike the hugetlbfs that
91 requires users and/or system administrators to configure what parts of
92 the system memory should and can be mapped by the huge pages, THP
93 manages such mappings transparently to the user and hence the
94 name. See Documentation/admin-guide/mm/transhuge.rst for more details
95 about THP.
96
97 Zones
98 =====
99
100 Often hardware poses restrictions on how different physical memory
101 ranges can be accessed. In some cases, devices cannot perform DMA to
102 all the addressable memory. In other cases, the size of the physical
103 memory exceeds the maximal addressable size of virtual memory and
104 special actions are required to access portions of the memory. Linux
105 groups memory pages into `zones` according to their possible
106 usage. For example, ZONE_DMA will contain memory that can be used by
107 devices for DMA, ZONE_HIGHMEM will contain memory that is not
108 permanently mapped into kernel's address space and ZONE_NORMAL will
109 contain normally addressed pages.
110
111 The actual layout of the memory zones is hardware dependent as not all
112 architectures define all zones, and requirements for DMA are different
113 for different platforms.
114
115 Nodes
116 =====
117
118 Many multi-processor machines are NUMA - Non-Uniform Memory Access -
119 systems. In such systems the memory is arranged into banks that have
120 different access latency depending on the "distance" from the
121 processor. Each bank is referred to as a `node` and for each node Linux
122 constructs an independent memory management subsystem. A node has its
123 own set of zones, lists of free and used pages and various statistics
124 counters. You can find more details about NUMA in
125 Documentation/mm/numa.rst` and in
126 Documentation/admin-guide/mm/numa_memory_policy.rst.
127
128 Page cache
129 ==========
130
131 The physical memory is volatile and the common case for getting data
132 into the memory is to read it from files. Whenever a file is read, the
133 data is put into the `page cache` to avoid expensive disk access on
134 the subsequent reads. Similarly, when one writes to a file, the data
135 is placed in the page cache and eventually gets into the backing
136 storage device. The written pages are marked as `dirty` and when Linux
137 decides to reuse them for other purposes, it makes sure to synchronize
138 the file contents on the device with the updated data.
139
140 Anonymous Memory
141 ================
142
143 The `anonymous memory` or `anonymous mappings` represent memory that
144 is not backed by a filesystem. Such mappings are implicitly created
145 for program's stack and heap or by explicit calls to mmap(2) system
146 call. Usually, the anonymous mappings only define virtual memory areas
147 that the program is allowed to access. The read accesses will result
148 in creation of a page table entry that references a special physical
149 page filled with zeroes. When the program performs a write, a regular
150 physical page will be allocated to hold the written data. The page
151 will be marked dirty and if the kernel decides to repurpose it,
152 the dirty page will be swapped out.
153
154 Reclaim
155 =======
156
157 Throughout the system lifetime, a physical page can be used for storing
158 different types of data. It can be kernel internal data structures,
159 DMA'able buffers for device drivers use, data read from a filesystem,
160 memory allocated by user space processes etc.
161
162 Depending on the page usage it is treated differently by the Linux
163 memory management. The pages that can be freed at any time, either
164 because they cache the data available elsewhere, for instance, on a
165 hard disk, or because they can be swapped out, again, to the hard
166 disk, are called `reclaimable`. The most notable categories of the
167 reclaimable pages are page cache and anonymous memory.
168
169 In most cases, the pages holding internal kernel data and used as DMA
170 buffers cannot be repurposed, and they remain pinned until freed by
171 their user. Such pages are called `unreclaimable`. However, in certain
172 circumstances, even pages occupied with kernel data structures can be
173 reclaimed. For instance, in-memory caches of filesystem metadata can
174 be re-read from the storage device and therefore it is possible to
175 discard them from the main memory when system is under memory
176 pressure.
177
178 The process of freeing the reclaimable physical memory pages and
179 repurposing them is called (surprise!) `reclaim`. Linux can reclaim
180 pages either asynchronously or synchronously, depending on the state
181 of the system. When the system is not loaded, most of the memory is free
182 and allocation requests will be satisfied immediately from the free
183 pages supply. As the load increases, the amount of the free pages goes
184 down and when it reaches a certain threshold (low watermark), an
185 allocation request will awaken the ``kswapd`` daemon. It will
186 asynchronously scan memory pages and either just free them if the data
187 they contain is available elsewhere, or evict to the backing storage
188 device (remember those dirty pages?). As memory usage increases even
189 more and reaches another threshold - min watermark - an allocation
190 will trigger `direct reclaim`. In this case allocation is stalled
191 until enough memory pages are reclaimed to satisfy the request.
192
193 Compaction
194 ==========
195
196 As the system runs, tasks allocate and free the memory and it becomes
197 fragmented. Although with virtual memory it is possible to present
198 scattered physical pages as virtually contiguous range, sometimes it is
199 necessary to allocate large physically contiguous memory areas. Such
200 need may arise, for instance, when a device driver requires a large
201 buffer for DMA, or when THP allocates a huge page. Memory `compaction`
202 addresses the fragmentation issue. This mechanism moves occupied pages
203 from the lower part of a memory zone to free pages in the upper part
204 of the zone. When a compaction scan is finished free pages are grouped
205 together at the beginning of the zone and allocations of large
206 physically contiguous areas become possible.
207
208 Like reclaim, the compaction may happen asynchronously in the ``kcompactd``
209 daemon or synchronously as a result of a memory allocation request.
210
211 OOM killer
212 ==========
213
214 It is possible that on a loaded machine memory will be exhausted and the
215 kernel will be unable to reclaim enough memory to continue to operate. In
216 order to save the rest of the system, it invokes the `OOM killer`.
217
218 The `OOM killer` selects a task to sacrifice for the sake of the overall
219 system health. The selected task is killed in a hope that after it exits
220 enough memory will be freed to continue normal operation.
221

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Linux memory management 개요

1-15

Linux memory management는 MMU가 없는 microcontroller부터 supercomputer까지 지원하도록 오랜 시간 확장된 복잡한 system입니다. MMU가 없는 시스템의 `nommu` memory management는 별도 문서가 필요하지만, 여기서는 MMU가 있고 CPU가 virtual address를 physical address로 변환할 수 있다고 가정합니다.

원문의 `contents:: :local:` 지시자는 이 문서 안의 절을 대상으로 로컬 목차를 만듭니다.

Virtual Memory Primer

16-63

physical memory는 한정된 자원이며 memory hotplug를 지원해도 설치 가능한 양에는 상한이 있습니다. 반드시 contiguous하지 않고 여러 address range로 나뉠 수 있으며, range 정의 방식도 CPU architecture와 구현마다 다릅니다.

virtual memory는 이런 physical memory 세부 사항을 application에서 숨깁니다. 필요한 정보만 physical memory에 두는 demand paging을 가능하게 하고 process 사이의 데이터 보호와 제어된 공유를 제공합니다.

CPU가 system memory를 읽거나 쓰는 instruction을 해석할 때 instruction의 `virtual` address를 memory controller가 이해하는 `physical` address로 변환합니다.

physical memory는 architecture별 크기의 page frame 또는 page로 나뉩니다. 여러 page size를 지원하는 architecture에서는 kernel build 때 configuration option으로 크기를 선택합니다.

physical page 하나는 하나 이상의 virtual page에 mapping될 수 있습니다. 계층형 page table이 program의 virtual address를 physical address로 변환하는 mapping을 설명합니다.

최하위 table entry는 실제 page의 physical address를 담고 상위 table은 하위 table page의 physical address를 담습니다. CPU는 register의 top-level table pointer에서 시작해 virtual address의 높은 bit부터 각 level index로 사용하고, 가장 낮은 bit를 실제 page 내부 offset으로 사용합니다.

Huge Pages

64-96

주소 변환은 여러 번의 느린 memory access가 필요하므로 CPU는 Translation Lookaside Buffer(TLB)에 변환 결과를 cache합니다. TLB는 제한적이어서 큰 working set을 가진 application은 TLB miss로 성능이 떨어질 수 있습니다.

현대 CPU는 상위 page-table level에서 memory page를 직접 mapping할 수 있습니다. x86의 second/third-level entry로 2M 또는 1G page를 mapping하는 것이 예이며 Linux에서는 이를 `huge` page라 부릅니다. huge page는 TLB pressure를 줄이고 hit rate와 전체 성능을 높입니다.

Linux의 첫 mechanism은 RAM을 backing store로 쓰는 pseudo filesystem `HugeTLB filesystem` 또는 hugetlbfs입니다. 여기에 만든 file 데이터는 memory에 있고 huge page로 mapping됩니다. 자세한 내용은 `Documentation/admin-guide/mm/hugetlbpage.rst`에 있습니다.

더 최근의 `Transparent HugePages`(THP)는 관리자가 huge-page 대상 memory를 미리 구성해야 하는 hugetlbfs와 달리 mapping을 투명하게 관리합니다. 자세한 내용은 `Documentation/admin-guide/mm/transhuge.rst`를 참고합니다.

Zones

97-114

hardware는 physical memory range 접근에 제약을 둘 수 있습니다. 어떤 device는 모든 addressable memory에 DMA를 할 수 없고, physical memory가 virtual address 가능 범위를 넘어 특별한 접근 절차가 필요하기도 합니다.

Linux는 가능한 용도에 따라 page를 `zone`으로 묶습니다. ZONE_DMA는 device DMA에 쓸 memory, ZONE_HIGHMEM은 kernel address space에 영구 mapping되지 않은 memory, ZONE_NORMAL은 일반적으로 address할 수 있는 page를 담습니다. 실제 zone layout은 architecture와 platform의 DMA 요구에 따라 달라집니다.

Nodes

115-127

많은 multiprocessor machine은 NUMA(Non-Uniform Memory Access) system입니다. memory bank마다 processor와의 `distance`에 따른 access latency가 다르며 각 bank를 `node`라고 합니다.

Linux는 node마다 독립적인 memory management subsystem을 만들고 각 node는 고유한 zone 집합, free/used page list, statistics counter를 가집니다. 자세한 내용은 `Documentation/mm/numa.rst`와 `Documentation/admin-guide/mm/numa_memory_policy.rst`에 있습니다.

Page cache

128-139

physical memory는 volatile이므로 보통 file에서 데이터를 읽어옵니다. file read 데이터는 다음 접근의 비싼 disk I/O를 피하려고 `page cache`에 둡니다. write 데이터도 먼저 page cache에 들어간 뒤 backing storage로 전달됩니다.

write된 page는 `dirty`로 표시됩니다. Linux가 이를 다른 용도로 재사용하려 할 때 device의 file 내용과 갱신된 데이터를 먼저 동기화합니다.

Anonymous Memory

140-153

`anonymous memory` 또는 `anonymous mapping`은 filesystem이 backing하지 않는 memory입니다. program stack과 heap에 암묵적으로 만들어지거나 mmap(2) system call로 명시적으로 생성됩니다.

처음에는 program이 접근할 수 있는 virtual memory area만 정의됩니다. read하면 zero로 채운 특수 physical page를 가리키는 page-table entry가 생기고 write하면 실제 데이터를 보관할 일반 physical page가 할당됩니다. 이 page는 dirty로 표시되며 kernel이 재사용하려 하면 swap out됩니다.

Reclaim

154-192

system 수명 동안 physical page는 kernel 내부 구조, device driver의 DMA buffer, filesystem data, userspace allocation 등 여러 용도로 바뀌어 사용됩니다.

다른 곳에서 다시 읽거나 disk로 swap out할 수 있어 언제든 free 가능한 page를 `reclaimable`이라 합니다. 대표적으로 page cache와 anonymous memory가 있습니다.

kernel 내부 데이터나 DMA buffer는 보통 사용자가 해제할 때까지 pinned되어 `unreclaimable`입니다. 다만 storage에서 다시 읽을 수 있는 filesystem metadata cache 같은 일부 kernel 구조는 memory pressure에서 버릴 수 있습니다.

reclaimable physical page를 free해 다른 용도로 바꾸는 과정을 `reclaim`이라 합니다. system이 한가할 때 allocation은 free page에서 즉시 충족됩니다. free page가 low watermark에 이르면 allocation이 `kswapd`를 깨우고, daemon은 비동기로 page를 scan해 즉시 free하거나 dirty 데이터를 backing storage로 evict합니다.

memory 사용량이 더 올라 min watermark에 이르면 allocation이 `direct reclaim`을 시작합니다. 이때 요청은 필요한 page를 충분히 회수할 때까지 멈춥니다.

Compaction

193-210

task가 memory를 할당하고 해제하면서 physical memory는 fragmentation됩니다. virtual memory는 흩어진 page를 virtually contiguous하게 보일 수 있지만 DMA의 큰 buffer나 THP huge page에는 큰 physically contiguous 영역이 필요합니다.

`compaction`은 memory zone 아래쪽의 사용 중 page를 위쪽 free page로 옮깁니다. scan이 끝나면 free page가 zone 시작 부분에 모여 큰 contiguous allocation이 가능해집니다.

reclaim과 마찬가지로 `kcompactd` daemon에서 비동기로 실행하거나 memory allocation 요청 때문에 동기로 실행할 수 있습니다.

OOM killer

211-220

부하가 큰 machine에서 memory가 고갈되고 kernel이 동작을 계속할 만큼 reclaim하지 못할 수 있습니다. 나머지 system을 살리기 위해 `OOM killer`를 호출합니다.

`OOM killer`는 전체 system health를 위해 희생할 task를 선택해 종료합니다. 그 task가 빠져나간 뒤 정상 동작을 계속할 만큼 memory가 해제되기를 기대합니다.