← Documents Documentation/power/freezing-of-tasks.rst GitHub 원문 ↗

Linux 6.18.37 · Power

Task freezing

Hibernation·system suspend에서 userspace와 선택된 kernel thread를 freeze하는 상태 전이, filesystem·memory·device 일관성의 이유, 교착과 firmware·locking 주의점을 설명합니다.

Source pathDocumentation/power/freezing-of-tasks.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

freezing-of-tasks.rst:1-256

Task freezer는 image와 disk·device 상태의 일관성을 지키기 위해 userspace를 먼저, 참여한 kernel thread를 다음으로 멈춥니다. Kernel thread의 opt-in 규칙과 dependency deadlock, firmware 선로딩, freezer-aware system sleep lock 사용이 핵심입니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 =================
2 Freezing of tasks
3 =================
4
5 (C) 2007 Rafael J. Wysocki <[email protected]>, GPL
6
7 I. What is the freezing of tasks?
8 =================================
9
10 The freezing of tasks is a mechanism by which user space processes and some
11 kernel threads are controlled during hibernation or system-wide suspend (on some
12 architectures).
13
14 II. How does it work?
15 =====================
16
17 There is one per-task flag (PF_NOFREEZE) and three per-task states
18 (TASK_FROZEN, TASK_FREEZABLE and __TASK_FREEZABLE_UNSAFE) used for that.
19 The tasks that have PF_NOFREEZE unset (all user space tasks and some kernel
20 threads) are regarded as 'freezable' and treated in a special way before the
21 system enters a sleep state as well as before a hibernation image is created
22 (hibernation is directly covered by what follows, but the description applies
23 to system-wide suspend too).
24
25 Namely, as the first step of the hibernation procedure the function
26 freeze_processes() (defined in kernel/power/process.c) is called. A system-wide
27 static key freezer_active (as opposed to a per-task flag or state) is used to
28 indicate whether the system is to undergo a freezing operation. And
29 freeze_processes() sets this static key. After this, it executes
30 try_to_freeze_tasks() that sends a fake signal to all user space processes, and
31 wakes up all the kernel threads. All freezable tasks must react to that by
32 calling try_to_freeze(), which results in a call to __refrigerator() (defined
33 in kernel/freezer.c), which changes the task's state to TASK_FROZEN, and makes
34 it loop until it is woken by an explicit TASK_FROZEN wakeup. Then, that task
35 is regarded as 'frozen' and so the set of functions handling this mechanism is
36 referred to as 'the freezer' (these functions are defined in
37 kernel/power/process.c, kernel/freezer.c & include/linux/freezer.h). User space
38 tasks are generally frozen before kernel threads.
39
40 __refrigerator() must not be called directly. Instead, use the
41 try_to_freeze() function (defined in include/linux/freezer.h), that checks
42 if the task is to be frozen and makes the task enter __refrigerator().
43
44 For user space processes try_to_freeze() is called automatically from the
45 signal-handling code, but the freezable kernel threads need to call it
46 explicitly in suitable places or use the wait_event_freezable() or
47 wait_event_freezable_timeout() macros (defined in include/linux/wait.h)
48 that put the task to sleep (TASK_INTERRUPTIBLE) or freeze it (TASK_FROZEN) if
49 freezer_active is set. The main loop of a freezable kernel thread may look
50 like the following one::
51
52 set_freezable();
53
54 while (true) {
55 struct task_struct *tsk = NULL;
56
57 wait_event_freezable(oom_reaper_wait, oom_reaper_list != NULL);
58 spin_lock_irq(&oom_reaper_lock);
59 if (oom_reaper_list != NULL) {
60 tsk = oom_reaper_list;
61 oom_reaper_list = tsk->oom_reaper_list;
62 }
63 spin_unlock_irq(&oom_reaper_lock);
64
65 if (tsk)
66 oom_reap_task(tsk);
67 }
68
69 (from mm/oom_kill.c::oom_reaper()).
70
71 If a freezable kernel thread is not put to the frozen state after the freezer
72 has initiated a freezing operation, the freezing of tasks will fail and the
73 entire system-wide transition will be cancelled. For this reason, freezable
74 kernel threads must call try_to_freeze() somewhere or use one of the
75 wait_event_freezable() and wait_event_freezable_timeout() macros.
76
77 After the system memory state has been restored from a hibernation image and
78 devices have been reinitialized, the function thaw_processes() is called in
79 order to wake up each frozen task. Then, the tasks that have been frozen leave
80 __refrigerator() and continue running.
81
82
83 Rationale behind the functions dealing with freezing and thawing of tasks
84 -------------------------------------------------------------------------
85
86 freeze_processes():
87 - freezes only userspace tasks
88
89 freeze_kernel_threads():
90 - freezes all tasks (including kernel threads) because we can't freeze
91 kernel threads without freezing userspace tasks
92
93 thaw_kernel_threads():
94 - thaws only kernel threads; this is particularly useful if we need to do
95 anything special in between thawing of kernel threads and thawing of
96 userspace tasks, or if we want to postpone the thawing of userspace tasks
97
98 thaw_processes():
99 - thaws all tasks (including kernel threads) because we can't thaw userspace
100 tasks without thawing kernel threads
101
102
103 III. Which kernel threads are freezable?
104 ========================================
105
106 Kernel threads are not freezable by default. However, a kernel thread may clear
107 PF_NOFREEZE for itself by calling set_freezable() (the resetting of PF_NOFREEZE
108 directly is not allowed). From this point it is regarded as freezable
109 and must call try_to_freeze() or variants of wait_event_freezable() in a
110 suitable place.
111
112 IV. Why do we do that?
113 ======================
114
115 Generally speaking, there is a couple of reasons to use the freezing of tasks:
116
117 1. The principal reason is to prevent filesystems from being damaged after
118 hibernation. At the moment we have no simple means of checkpointing
119 filesystems, so if there are any modifications made to filesystem data and/or
120 metadata on disks, we cannot bring them back to the state from before the
121 modifications. At the same time each hibernation image contains some
122 filesystem-related information that must be consistent with the state of the
123 on-disk data and metadata after the system memory state has been restored
124 from the image (otherwise the filesystems will be damaged in a nasty way,
125 usually making them almost impossible to repair). We therefore freeze
126 tasks that might cause the on-disk filesystems' data and metadata to be
127 modified after the hibernation image has been created and before the
128 system is finally powered off. The majority of these are user space
129 processes, but if any of the kernel threads may cause something like this
130 to happen, they have to be freezable.
131
132 2. Next, to create the hibernation image we need to free a sufficient amount of
133 memory (approximately 50% of available RAM) and we need to do that before
134 devices are deactivated, because we generally need them for swapping out.
135 Then, after the memory for the image has been freed, we don't want tasks
136 to allocate additional memory and we prevent them from doing that by
137 freezing them earlier. [Of course, this also means that device drivers
138 should not allocate substantial amounts of memory from their .suspend()
139 callbacks before hibernation, but this is a separate issue.]
140
141 3. The third reason is to prevent user space processes and some kernel threads
142 from interfering with the suspending and resuming of devices. A user space
143 process running on a second CPU while we are suspending devices may, for
144 example, be troublesome and without the freezing of tasks we would need some
145 safeguards against race conditions that might occur in such a case.
146
147 Although Linus Torvalds doesn't like the freezing of tasks, he said this in one
148 of the discussions on LKML (https://lore.kernel.org/r/[email protected]):
149
150 "RJW:> Why we freeze tasks at all or why we freeze kernel threads?
151
152 Linus: In many ways, 'at all'.
153
154 I **do** realize the IO request queue issues, and that we cannot actually do
155 s2ram with some devices in the middle of a DMA. So we want to be able to
156 avoid *that*, there's no question about that. And I suspect that stopping
157 user threads and then waiting for a sync is practically one of the easier
158 ways to do so.
159
160 So in practice, the 'at all' may become a 'why freeze kernel threads?' and
161 freezing user threads I don't find really objectionable."
162
163 Still, there are kernel threads that may want to be freezable. For example, if
164 a kernel thread that belongs to a device driver accesses the device directly, it
165 in principle needs to know when the device is suspended, so that it doesn't try
166 to access it at that time. However, if the kernel thread is freezable, it will
167 be frozen before the driver's .suspend() callback is executed and it will be
168 thawed after the driver's .resume() callback has run, so it won't be accessing
169 the device while it's suspended.
170
171 4. Another reason for freezing tasks is to prevent user space processes from
172 realizing that hibernation (or suspend) operation takes place. Ideally, user
173 space processes should not notice that such a system-wide operation has
174 occurred and should continue running without any problems after the restore
175 (or resume from suspend). Unfortunately, in the most general case this
176 is quite difficult to achieve without the freezing of tasks. Consider,
177 for example, a process that depends on all CPUs being online while it's
178 running. Since we need to disable nonboot CPUs during the hibernation,
179 if this process is not frozen, it may notice that the number of CPUs has
180 changed and may start to work incorrectly because of that.
181
182 V. Are there any problems related to the freezing of tasks?
183 ===========================================================
184
185 Yes, there are.
186
187 First of all, the freezing of kernel threads may be tricky if they depend one
188 on another. For example, if kernel thread A waits for a completion (in the
189 TASK_UNINTERRUPTIBLE state) that needs to be done by freezable kernel thread B
190 and B is frozen in the meantime, then A will be blocked until B is thawed, which
191 may be undesirable. That's why kernel threads are not freezable by default.
192
193 Second, there are the following two problems related to the freezing of user
194 space processes:
195
196 1. Putting processes into an uninterruptible sleep distorts the load average.
197 2. Now that we have FUSE, plus the framework for doing device drivers in
198 userspace, it gets even more complicated because some userspace processes are
199 now doing the sorts of things that kernel threads do
200 (https://lists.linux-foundation.org/pipermail/linux-pm/2007-May/012309.html).
201
202 The problem 1. seems to be fixable, although it hasn't been fixed so far. The
203 other one is more serious, but it seems that we can work around it by using
204 hibernation (and suspend) notifiers (in that case, though, we won't be able to
205 avoid the realization by the user space processes that the hibernation is taking
206 place).
207
208 There are also problems that the freezing of tasks tends to expose, although
209 they are not directly related to it. For example, if request_firmware() is
210 called from a device driver's .resume() routine, it will timeout and eventually
211 fail, because the user land process that should respond to the request is frozen
212 at this point. So, seemingly, the failure is due to the freezing of tasks.
213 Suppose, however, that the firmware file is located on a filesystem accessible
214 only through another device that hasn't been resumed yet. In that case,
215 request_firmware() will fail regardless of whether or not the freezing of tasks
216 is used. Consequently, the problem is not really related to the freezing of
217 tasks, since it generally exists anyway.
218
219 A driver must have all firmwares it may need in RAM before suspend() is called.
220 If keeping them is not practical, for example due to their size, they must be
221 requested early enough using the suspend notifier API described in
222 Documentation/driver-api/pm/notifiers.rst.
223
224 VI. Are there any precautions to be taken to prevent freezing failures?
225 =======================================================================
226
227 Yes, there are.
228
229 First of all, grabbing the 'system_transition_mutex' lock to mutually exclude a
230 piece of code from system-wide sleep such as suspend/hibernation is not
231 encouraged. If possible, that piece of code must instead hook onto the
232 suspend/hibernation notifiers to achieve mutual exclusion. Look at the
233 CPU-Hotplug code (kernel/cpu.c) for an example.
234
235 However, if that is not feasible, and grabbing 'system_transition_mutex' is
236 deemed necessary, it is strongly discouraged to directly call
237 mutex_[un]lock(&system_transition_mutex) since that could lead to freezing
238 failures, because if the suspend/hibernate code successfully acquired the
239 'system_transition_mutex' lock, and hence that other entity failed to acquire
240 the lock, then that task would get blocked in TASK_UNINTERRUPTIBLE state. As a
241 consequence, the freezer would not be able to freeze that task, leading to
242 freezing failure.
243
244 However, the [un]lock_system_sleep() APIs are safe to use in this scenario,
245 since they ask the freezer to skip freezing this task, since it is anyway
246 "frozen enough" as it is blocked on 'system_transition_mutex', which will be
247 released only after the entire suspend/hibernation sequence is complete. So, to
248 summarize, use [un]lock_system_sleep() instead of directly using
249 mutex_[un]lock(&system_transition_mutex). That would prevent freezing failures.
250
251 V. Miscellaneous
252 ================
253
254 /sys/power/pm_freeze_timeout controls how long it will cost at most to freeze
255 all user space processes or all freezable kernel threads, in unit of
256 millisecond. The default value is 20000, with range of unsigned integer.
257

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Task freezing이란

1-13

Task freezing은 hibernation 또는 일부 architecture의 system-wide suspend 동안 userspace process와 일부 kernel thread의 실행을 통제하는 mechanism입니다.

목적은 단순히 모든 task를 멈추는 것이 아니라, sleep transition이 의존하는 filesystem·memory·device 상태가 바뀌지 않도록 freezable task만 정해진 순서로 정지시키는 데 있습니다.

Freezer 적용 범위
hibernation / system suspendtask freezeruserspace processes
hibernation / system suspendtask freezerselected kernel threads

System-wide power transition 중 userspace 전체와 명시적으로 참여한 kernel thread를 멈춥니다.

=================
Freezing of tasks
=================

(C) 2007 Rafael J. Wysocki <[email protected]>, GPL

I. What is the freezing of tasks?
=================================

The freezing of tasks is a mechanism by which user space processes and some
kernel threads are controlled during hibernation or system-wide suspend (on some
architectures).

Flag·state와 freeze 진입

14-43

이 mechanism은 per-task flag `PF_NOFREEZE`와 세 per-task state `TASK_FROZEN`, `TASK_FREEZABLE`, `__TASK_FREEZABLE_UNSAFE`를 사용합니다. `PF_NOFREEZE`가 설정되지 않은 모든 userspace task와 일부 kernel thread가 freezable입니다.

이 task들은 system이 sleep state로 들어가기 전과 hibernation image를 만들기 전에 특별히 처리됩니다. 아래 설명은 hibernation을 직접 다루지만 system-wide suspend에도 적용됩니다.

Hibernation의 첫 단계에서 `kernel/power/process.c`의 `freeze_processes()`가 system-wide static key `freezer_active`를 설정합니다. 이는 per-task flag나 state가 아니라 system 전체가 freeze operation 중임을 나타냅니다.

이어서 `try_to_freeze_tasks()`가 모든 userspace process에 fake signal을 보내고 모든 kernel thread를 깨웁니다. Userspace의 signal 경로에서는 `TIF_SIGPENDING`으로 대표되는 pending signal 상태를 통해 freeze 검사가 이루어지며, freezable task는 `try_to_freeze()`를 호출해야 합니다.

`try_to_freeze()`는 `kernel/freezer.c`의 `__refrigerator()`로 들어가 task state를 `TASK_FROZEN`으로 바꾸고, 명시적인 `TASK_FROZEN` wakeup을 받을 때까지 loop합니다. 이 상태를 frozen이라고 하며 관련 함수 집합을 freezer라고 부릅니다. 구현은 `kernel/power/process.c`, `kernel/freezer.c`, `include/linux/freezer.h`에 있습니다. 일반적으로 userspace task를 kernel thread보다 먼저 freeze합니다.

`__refrigerator()`를 직접 호출하면 안 됩니다. 대신 task가 실제로 freeze 대상인지 검사한 뒤 진입시키는 `include/linux/freezer.h`의 `try_to_freeze()`를 사용합니다.

Freeze 진입 경로
freeze_processes()freezer_active = truetry_to_freeze_tasks()fake signal / wake kernel threadstry_to_freeze()__refrigerator()TASK_FROZEN

System-wide static key를 켠 뒤 각 freezable task가 refrigerator loop에 들어갑니다.

II. How does it work?
=====================

There is one per-task flag (PF_NOFREEZE) and three per-task states
(TASK_FROZEN, TASK_FREEZABLE and __TASK_FREEZABLE_UNSAFE) used for that.
The tasks that have PF_NOFREEZE unset (all user space tasks and some kernel
threads) are regarded as 'freezable' and treated in a special way before the
system enters a sleep state as well as before a hibernation image is created
(hibernation is directly covered by what follows, but the description applies
to system-wide suspend too).

Namely, as the first step of the hibernation procedure the function
freeze_processes() (defined in kernel/power/process.c) is called.  A system-wide
static key freezer_active (as opposed to a per-task flag or state) is used to
indicate whether the system is to undergo a freezing operation. And
freeze_processes() sets this static key.  After this, it executes
try_to_freeze_tasks() that sends a fake signal to all user space processes, and
wakes up all the kernel threads. All freezable tasks must react to that by
calling try_to_freeze(), which results in a call to __refrigerator() (defined
in kernel/freezer.c), which changes the task's state to TASK_FROZEN, and makes
it loop until it is woken by an explicit TASK_FROZEN wakeup. Then, that task
is regarded as 'frozen' and so the set of functions handling this mechanism is
referred to as 'the freezer' (these functions are defined in
kernel/power/process.c, kernel/freezer.c & include/linux/freezer.h). User space
tasks are generally frozen before kernel threads.

__refrigerator() must not be called directly.  Instead, use the
try_to_freeze() function (defined in include/linux/freezer.h), that checks
if the task is to be frozen and makes the task enter __refrigerator().

Freezable kernel thread와 thaw

44-82

Userspace process에서는 signal handling code가 `try_to_freeze()`를 자동 호출합니다. Freezable kernel thread는 적절한 위치에서 직접 호출하거나 `include/linux/wait.h`의 `wait_event_freezable()` 또는 `wait_event_freezable_timeout()`을 사용해야 합니다.

이 macro들은 평소 task를 `TASK_INTERRUPTIBLE`로 재우고, `freezer_active`가 설정되면 `TASK_FROZEN`으로 freeze합니다.

예제는 `mm/oom_kill.c::oom_reaper()`의 main loop입니다. `set_freezable()`로 참여를 선언하고 `wait_event_freezable(oom_reaper_wait, oom_reaper_list != NULL)`에서 work를 기다립니다. 깨어나면 lock 아래 list에서 task를 꺼내 `oom_reap_task(tsk)`를 실행합니다.

Freezer가 시작됐는데 freezable kernel thread가 frozen state에 들어가지 않으면 task freezing이 실패하고 system-wide transition 전체가 취소됩니다. 따라서 반드시 `try_to_freeze()` 또는 freezable wait macro 중 하나를 사용해야 합니다.

Hibernation image에서 system memory가 복원되고 device가 재초기화되면 `thaw_processes()`가 frozen task를 깨웁니다. Task는 `__refrigerator()`를 빠져나와 실행을 계속합니다.

Userspace와 kernel thread
Task 종류Freeze 확인대기 상태
Userspace processSignal handling code가 자동으로 try_to_freeze() 호출TASK_FROZEN
Freezable kernel thread명시적 try_to_freeze() 또는 wait_event_freezable*()TASK_INTERRUPTIBLE 또는 TASK_FROZEN

Freezer 요청을 관찰하는 경로가 다릅니다.

For user space processes try_to_freeze() is called automatically from the
signal-handling code, but the freezable kernel threads need to call it
explicitly in suitable places or use the wait_event_freezable() or
wait_event_freezable_timeout() macros (defined in include/linux/wait.h)
that put the task to sleep (TASK_INTERRUPTIBLE) or freeze it (TASK_FROZEN) if
freezer_active is set. The main loop of a freezable kernel thread may look
like the following one::

        set_freezable();

        while (true) {
                struct task_struct *tsk = NULL;

                wait_event_freezable(oom_reaper_wait, oom_reaper_list != NULL);
                spin_lock_irq(&oom_reaper_lock);
                if (oom_reaper_list != NULL) {
                        tsk = oom_reaper_list;
                        oom_reaper_list = tsk->oom_reaper_list;
                }
                spin_unlock_irq(&oom_reaper_lock);

                if (tsk)
                        oom_reap_task(tsk);
        }

(from mm/oom_kill.c::oom_reaper()).

If a freezable kernel thread is not put to the frozen state after the freezer
has initiated a freezing operation, the freezing of tasks will fail and the
entire system-wide transition will be cancelled.  For this reason, freezable
kernel threads must call try_to_freeze() somewhere or use one of the
wait_event_freezable() and wait_event_freezable_timeout() macros.

After the system memory state has been restored from a hibernation image and
devices have been reinitialized, the function thaw_processes() is called in
order to wake up each frozen task.  Then, the tasks that have been frozen leave
__refrigerator() and continue running.

Freeze·thaw 함수의 역할

83-102

Freeze와 thaw 함수를 나눈 이유는 userspace와 kernel thread의 순서를 제어하기 위해서입니다.

Freezer 함수별 범위
함수대상의도
freeze_processes()Userspace task만먼저 userspace 정지
freeze_kernel_threads()Kernel thread를 포함한 모든 taskUserspace가 frozen인 상태에서 kernel thread 정지
thaw_kernel_threads()Kernel thread만Userspace보다 먼저 kernel 작업 재개
thaw_processes()Kernel thread를 포함한 모든 task마지막으로 userspace까지 재개

Kernel thread는 userspace를 freeze한 뒤에만 freeze할 수 있고, userspace는 kernel thread를 thaw한 뒤에만 thaw할 수 있습니다.

`thaw_kernel_threads()`와 `thaw_processes()`를 분리하면 kernel thread를 깨운 뒤 userspace를 깨우기 전에 특별한 작업을 수행하거나 userspace thaw를 늦출 수 있습니다.

권장 순서
freeze_processes()freeze_kernel_threads()sleep / image workthaw_kernel_threads()special work if neededthaw_processes()

정지할 때는 userspace→kernel thread, 재개할 때는 kernel thread→userspace 순서입니다.

Rationale behind the functions dealing with freezing and thawing of tasks
-------------------------------------------------------------------------

freeze_processes():
  - freezes only userspace tasks

freeze_kernel_threads():
  - freezes all tasks (including kernel threads) because we can't freeze
    kernel threads without freezing userspace tasks

thaw_kernel_threads():
  - thaws only kernel threads; this is particularly useful if we need to do
    anything special in between thawing of kernel threads and thawing of
    userspace tasks, or if we want to postpone the thawing of userspace tasks

thaw_processes():
  - thaws all tasks (including kernel threads) because we can't thaw userspace
    tasks without thawing kernel threads

어떤 kernel thread가 freezable인가

103-111

Kernel thread는 기본적으로 freezable이 아닙니다. 참여하려는 thread는 `set_freezable()`을 호출해 자신의 `PF_NOFREEZE`를 clear할 수 있습니다. `PF_NOFREEZE`를 직접 reset하는 것은 허용되지 않습니다.

그 시점부터 thread는 freezable로 간주되며 적절한 위치에서 `try_to_freeze()` 또는 `wait_event_freezable()` 계열을 호출해야 합니다.

Kernel thread opt-in
kernel thread: PF_NOFREEZE setset_freezable()PF_NOFREEZE clearedtry_to_freeze() / wait_event_freezable*()

명시적으로 참여를 선언한 thread만 freezer protocol을 따라야 합니다.

III. Which kernel threads are freezable?
========================================

Kernel threads are not freezable by default.  However, a kernel thread may clear
PF_NOFREEZE for itself by calling set_freezable() (the resetting of PF_NOFREEZE
directly is not allowed).  From this point it is regarded as freezable
and must call try_to_freeze() or variants of wait_event_freezable() in a
suitable place.

Freezer가 필요한 핵심 이유

112-146

첫째이자 가장 중요한 이유는 hibernation 뒤 filesystem 손상을 막는 것입니다. 현재 filesystem을 간단히 checkpoint하는 방법이 없으므로 disk의 data나 metadata가 바뀌면 변경 전 상태로 되돌릴 수 없습니다.

Hibernation image에는 filesystem 관련 상태가 들어 있고, 복원 뒤 disk의 data·metadata와 일치해야 합니다. 그렇지 않으면 filesystem이 심하게 손상되어 복구가 거의 불가능할 수 있습니다. 따라서 image 생성 뒤 최종 power-off 전까지 disk filesystem을 바꿀 수 있는 task를 freeze합니다. 대부분 userspace process지만 그런 변경을 일으킬 수 있는 kernel thread도 freezable이어야 합니다.

둘째, hibernation image를 만들려면 available RAM의 약 50%를 비워야 합니다. Swap-out에 device가 필요하므로 device를 deactivate하기 전에 memory를 확보합니다. 확보한 뒤 task가 다시 memory를 할당하지 못하도록 먼저 freeze합니다. 별개로 driver의 `.suspend()` callback도 hibernation 전에 많은 memory를 할당하지 않아야 합니다.

셋째, userspace process와 일부 kernel thread가 device suspend·resume를 방해하지 못하게 합니다. Device를 suspend하는 동안 두 번째 CPU에서 process가 실행되면 race condition이 생길 수 있으므로 freezer가 없으면 별도의 safeguard가 필요합니다.

Freezer의 세 핵심 목적
목적방지하는 문제
Filesystem 일관성Image 생성 뒤 disk data·metadata 변경
Image memory 확보RAM 약 50% 확보 뒤 task의 재할당
Device transition 격리Suspend·resume 중 task 접근과 race

Hibernation image와 외부 상태 사이의 일관성을 지킵니다.

IV. Why do we do that?
======================

Generally speaking, there is a couple of reasons to use the freezing of tasks:

1. The principal reason is to prevent filesystems from being damaged after
   hibernation.  At the moment we have no simple means of checkpointing
   filesystems, so if there are any modifications made to filesystem data and/or
   metadata on disks, we cannot bring them back to the state from before the
   modifications.  At the same time each hibernation image contains some
   filesystem-related information that must be consistent with the state of the
   on-disk data and metadata after the system memory state has been restored
   from the image (otherwise the filesystems will be damaged in a nasty way,
   usually making them almost impossible to repair).  We therefore freeze
   tasks that might cause the on-disk filesystems' data and metadata to be
   modified after the hibernation image has been created and before the
   system is finally powered off. The majority of these are user space
   processes, but if any of the kernel threads may cause something like this
   to happen, they have to be freezable.

2. Next, to create the hibernation image we need to free a sufficient amount of
   memory (approximately 50% of available RAM) and we need to do that before
   devices are deactivated, because we generally need them for swapping out.
   Then, after the memory for the image has been freed, we don't want tasks
   to allocate additional memory and we prevent them from doing that by
   freezing them earlier. [Of course, this also means that device drivers
   should not allocate substantial amounts of memory from their .suspend()
   callbacks before hibernation, but this is a separate issue.]

3. The third reason is to prevent user space processes and some kernel threads
   from interfering with the suspending and resuming of devices.  A user space
   process running on a second CPU while we are suspending devices may, for
   example, be troublesome and without the freezing of tasks we would need some
   safeguards against race conditions that might occur in such a case.

Kernel thread와 userspace 투명성

147-181

문서는 Linus Torvalds의 LKML 논의를 인용합니다. 핵심은 DMA 중인 일부 device로는 s2ram을 진행할 수 없으므로 I/O request를 멈춰야 하고, userspace thread를 멈춘 뒤 sync를 기다리는 방식은 실용적이라는 것입니다. 반면 모든 kernel thread까지 freeze해야 하는지는 더 의문스럽다는 취지입니다.

그래도 freezable이어야 하는 kernel thread가 있습니다. Device driver 소속 thread가 hardware에 직접 접근한다면 device가 suspend된 동안 접근해서는 안 됩니다. Thread를 freezable로 만들면 driver의 `.suspend()` callback 전에 freeze되고 `.resume()` callback 뒤에 thaw되므로 정지된 device에 접근하지 않습니다.

넷째 이유는 userspace process가 hibernation 또는 suspend가 일어났다는 사실을 알아차리지 못하게 하는 것입니다. 이상적으로 process는 system-wide transition 뒤에도 문제없이 이어서 실행해야 합니다.

예를 들어 실행 중 모든 CPU가 online이라고 가정하는 process가 있습니다. Hibernation 중 nonboot CPU를 disable해야 하는데 process를 freeze하지 않으면 CPU 수 변경을 감지하고 잘못 동작할 수 있습니다.

Device thread 보호
kernel thread runningfreeze before driver .suspend()device suspended: no accessdriver .resume()thaw kernel thread

Driver callback을 경계로 kernel thread를 멈춰 suspended device 접근을 차단합니다.

Although Linus Torvalds doesn't like the freezing of tasks, he said this in one
of the discussions on LKML (https://lore.kernel.org/r/[email protected]):

"RJW:> Why we freeze tasks at all or why we freeze kernel threads?

Linus: In many ways, 'at all'.

I **do** realize the IO request queue issues, and that we cannot actually do
s2ram with some devices in the middle of a DMA.  So we want to be able to
avoid *that*, there's no question about that.  And I suspect that stopping
user threads and then waiting for a sync is practically one of the easier
ways to do so.

So in practice, the 'at all' may become a 'why freeze kernel threads?' and
freezing user threads I don't find really objectionable."

Still, there are kernel threads that may want to be freezable.  For example, if
a kernel thread that belongs to a device driver accesses the device directly, it
in principle needs to know when the device is suspended, so that it doesn't try
to access it at that time.  However, if the kernel thread is freezable, it will
be frozen before the driver's .suspend() callback is executed and it will be
thawed after the driver's .resume() callback has run, so it won't be accessing
the device while it's suspended.

4. Another reason for freezing tasks is to prevent user space processes from
   realizing that hibernation (or suspend) operation takes place.  Ideally, user
   space processes should not notice that such a system-wide operation has
   occurred and should continue running without any problems after the restore
   (or resume from suspend).  Unfortunately, in the most general case this
   is quite difficult to achieve without the freezing of tasks.  Consider,
   for example, a process that depends on all CPUs being online while it's
   running.  Since we need to disable nonboot CPUs during the hibernation,
   if this process is not frozen, it may notice that the number of CPUs has
   changed and may start to work incorrectly because of that.

Freezing이 드러내는 문제

182-223

Kernel thread끼리 의존하면 freezing이 까다롭습니다. Thread A가 `TASK_UNINTERRUPTIBLE` 상태로 completion을 기다리고 그 completion을 수행할 freezable thread B가 먼저 frozen되면, A는 B가 thaw될 때까지 막힙니다. 이것이 kernel thread가 기본적으로 freezable이 아닌 이유입니다.

Userspace freezing에는 두 문제가 있습니다. Process를 uninterruptible sleep에 넣으면 load average가 왜곡됩니다. 또 FUSE와 userspace device-driver framework 때문에 일부 userspace process가 kernel thread와 같은 역할을 수행하므로 의존 관계가 복잡해집니다.

Load average 문제는 고칠 수 있을 것으로 보이지만 아직 해결되지 않았습니다. 두 번째 문제는 더 심각하며 hibernation·suspend notifier로 우회할 수 있지만, 그러면 userspace가 transition을 알아차리지 못하게 한다는 목표는 포기해야 합니다.

Freezer가 원인처럼 보이지만 원래 존재하던 문제를 드러내기도 합니다. Driver `.resume()`에서 `request_firmware()`를 호출하면 응답할 userland process가 frozen이어서 timeout될 수 있습니다. 그러나 firmware file이 아직 resume되지 않은 다른 device를 통해서만 접근 가능한 filesystem에 있다면 freezer가 없어도 실패합니다.

Driver는 필요할 모든 firmware를 `suspend()` 호출 전에 RAM에 보유해야 합니다. 크기 때문에 계속 유지하기 어렵다면 `Documentation/driver-api/pm/notifiers.rst`의 suspend notifier API를 이용해 충분히 일찍 요청해야 합니다.

Freezer 관련 위험
상황결과대응
A가 frozen B의 completion 대기TASK_UNINTERRUPTIBLE 교착Kernel thread는 기본 non-freezable
FUSE/userspace driverUserspace freeze가 kernel service 중단Notifier 설계 검토
.resume()의 request_firmware()Userland 또는 storage 미복귀로 timeoutsuspend 전 RAM 확보 또는 early notifier

직접적인 freezer 교착과 원래 존재하던 resume 의존 문제를 구분해야 합니다.

V. Are there any problems related to the freezing of tasks?
===========================================================

Yes, there are.

First of all, the freezing of kernel threads may be tricky if they depend one
on another.  For example, if kernel thread A waits for a completion (in the
TASK_UNINTERRUPTIBLE state) that needs to be done by freezable kernel thread B
and B is frozen in the meantime, then A will be blocked until B is thawed, which
may be undesirable.  That's why kernel threads are not freezable by default.

Second, there are the following two problems related to the freezing of user
space processes:

1. Putting processes into an uninterruptible sleep distorts the load average.
2. Now that we have FUSE, plus the framework for doing device drivers in
   userspace, it gets even more complicated because some userspace processes are
   now doing the sorts of things that kernel threads do
   (https://lists.linux-foundation.org/pipermail/linux-pm/2007-May/012309.html).

The problem 1. seems to be fixable, although it hasn't been fixed so far.  The
other one is more serious, but it seems that we can work around it by using
hibernation (and suspend) notifiers (in that case, though, we won't be able to
avoid the realization by the user space processes that the hibernation is taking
place).

There are also problems that the freezing of tasks tends to expose, although
they are not directly related to it.  For example, if request_firmware() is
called from a device driver's .resume() routine, it will timeout and eventually
fail, because the user land process that should respond to the request is frozen
at this point.  So, seemingly, the failure is due to the freezing of tasks.
Suppose, however, that the firmware file is located on a filesystem accessible
only through another device that hasn't been resumed yet.  In that case,
request_firmware() will fail regardless of whether or not the freezing of tasks
is used.  Consequently, the problem is not really related to the freezing of
tasks, since it generally exists anyway.

A driver must have all firmwares it may need in RAM before suspend() is called.
If keeping them is not practical, for example due to their size, they must be
requested early enough using the suspend notifier API described in
Documentation/driver-api/pm/notifiers.rst.

Freezing 실패를 막는 locking

224-250

System-wide suspend·hibernation과 code 구간을 상호 배제하려고 `system_transition_mutex`를 직접 잡는 방식은 권장되지 않습니다. 가능하면 suspend/hibernation notifier를 연결해야 하며 `kernel/cpu.c`의 CPU-Hotplug code가 예입니다.

직접 `mutex_[un]lock(&system_transition_mutex)`을 호출하면 freezing 실패가 생길 수 있습니다. Suspend code가 lock을 먼저 얻은 상태에서 다른 task가 lock을 기다리면 그 task는 `TASK_UNINTERRUPTIBLE`로 막히고 freezer가 freeze하지 못합니다.

이 경우 `[un]lock_system_sleep()` API는 안전합니다. 이 API는 해당 task가 `system_transition_mutex`에 막혀 이미 충분히 frozen된 상태이므로 freezer가 freeze count에서 건너뛰도록 표시합니다. 구현상 이런 경로는 `PF_FREEZER_SKIP` 같은 freezer skip 상태를 사용하고, mutex는 suspend/hibernation sequence 전체가 끝난 뒤에만 풀립니다.

정리하면 `mutex_[un]lock(&system_transition_mutex)`을 직접 쓰지 말고 `[un]lock_system_sleep()`을 사용해야 freezing 실패를 막을 수 있습니다.

System sleep 상호 배제
need mutual exclusionprefer suspend/hibernation notifier
mutex unavoidable[un]lock_system_sleep()mark task freezer-skip / frozen enoughavoid freezing failure
do not usemutex_[un]lock(&system_transition_mutex)TASK_UNINTERRUPTIBLE waiterfreezer failure

Notifier가 우선이며 mutex가 꼭 필요하면 freezer-aware wrapper를 사용합니다.

VI. Are there any precautions to be taken to prevent freezing failures?
=======================================================================

Yes, there are.

First of all, grabbing the 'system_transition_mutex' lock to mutually exclude a
piece of code from system-wide sleep such as suspend/hibernation is not
encouraged.  If possible, that piece of code must instead hook onto the
suspend/hibernation notifiers to achieve mutual exclusion. Look at the
CPU-Hotplug code (kernel/cpu.c) for an example.

However, if that is not feasible, and grabbing 'system_transition_mutex' is
deemed necessary, it is strongly discouraged to directly call
mutex_[un]lock(&system_transition_mutex) since that could lead to freezing
failures, because if the suspend/hibernate code successfully acquired the
'system_transition_mutex' lock, and hence that other entity failed to acquire
the lock, then that task would get blocked in TASK_UNINTERRUPTIBLE state. As a
consequence, the freezer would not be able to freeze that task, leading to
freezing failure.

However, the [un]lock_system_sleep() APIs are safe to use in this scenario,
since they ask the freezer to skip freezing this task, since it is anyway
"frozen enough" as it is blocked on 'system_transition_mutex', which will be
released only after the entire suspend/hibernation sequence is complete.  So, to
summarize, use [un]lock_system_sleep() instead of directly using
mutex_[un]lock(&system_transition_mutex). That would prevent freezing failures.

pm_freeze_timeout

251-256

원문은 마지막 절 번호를 다시 V로 표기합니다. `/sys/power/pm_freeze_timeout`은 모든 userspace process 또는 모든 freezable kernel thread를 freeze하는 데 허용할 최대 시간을 밀리초 단위로 제어합니다.

기본값은 `20000`, 즉 20초이며 unsigned integer 범위를 사용합니다.

pm_freeze_timeout
경로단위기본값범위
/sys/power/pm_freeze_timeoutmillisecond20000unsigned integer

Userspace와 kernel-thread freeze 단계 각각의 최대 대기 시간을 제한합니다.

V. Miscellaneous
================

/sys/power/pm_freeze_timeout controls how long it will cost at most to freeze
all user space processes or all freezable kernel threads, in unit of
millisecond.  The default value is 20000, with range of unsigned integer.