요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
=================
Freezing of tasks
=================
(C) 2007 Rafael J. Wysocki <[email protected]>, GPL
I. What is the freezing of tasks?
=================================
The freezing of tasks is a mechanism by which user space processes and some
kernel threads are controlled during hibernation or system-wide suspend (on some
architectures).
II. How does it work?
=====================
There is one per-task flag (PF_NOFREEZE) and three per-task states
(TASK_FROZEN, TASK_FREEZABLE and __TASK_FREEZABLE_UNSAFE) used for that.
The tasks that have PF_NOFREEZE unset (all user space tasks and some kernel
threads) are regarded as 'freezable' and treated in a special way before the
system enters a sleep state as well as before a hibernation image is created
(hibernation is directly covered by what follows, but the description applies
to system-wide suspend too).
Namely, as the first step of the hibernation procedure the function
freeze_processes() (defined in kernel/power/process.c) is called. A system-wide
static key freezer_active (as opposed to a per-task flag or state) is used to
indicate whether the system is to undergo a freezing operation. And
freeze_processes() sets this static key. After this, it executes
try_to_freeze_tasks() that sends a fake signal to all user space processes, and
wakes up all the kernel threads. All freezable tasks must react to that by
calling try_to_freeze(), which results in a call to __refrigerator() (defined
in kernel/freezer.c), which changes the task's state to TASK_FROZEN, and makes
it loop until it is woken by an explicit TASK_FROZEN wakeup. Then, that task
is regarded as 'frozen' and so the set of functions handling this mechanism is
referred to as 'the freezer' (these functions are defined in
kernel/power/process.c, kernel/freezer.c & include/linux/freezer.h). User space
tasks are generally frozen before kernel threads.
__refrigerator() must not be called directly. Instead, use the
try_to_freeze() function (defined in include/linux/freezer.h), that checks
if the task is to be frozen and makes the task enter __refrigerator().
For user space processes try_to_freeze() is called automatically from the
signal-handling code, but the freezable kernel threads need to call it
explicitly in suitable places or use the wait_event_freezable() or
wait_event_freezable_timeout() macros (defined in include/linux/wait.h)
that put the task to sleep (TASK_INTERRUPTIBLE) or freeze it (TASK_FROZEN) if
freezer_active is set. The main loop of a freezable kernel thread may look
like the following one::
set_freezable();
while (true) {
struct task_struct *tsk = NULL;
wait_event_freezable(oom_reaper_wait, oom_reaper_list != NULL);
spin_lock_irq(&oom_reaper_lock);
if (oom_reaper_list != NULL) {
tsk = oom_reaper_list;
oom_reaper_list = tsk->oom_reaper_list;
}
spin_unlock_irq(&oom_reaper_lock);
if (tsk)
oom_reap_task(tsk);
}
(from mm/oom_kill.c::oom_reaper()).
If a freezable kernel thread is not put to the frozen state after the freezer
has initiated a freezing operation, the freezing of tasks will fail and the
entire system-wide transition will be cancelled. For this reason, freezable
kernel threads must call try_to_freeze() somewhere or use one of the
wait_event_freezable() and wait_event_freezable_timeout() macros.
After the system memory state has been restored from a hibernation image and
devices have been reinitialized, the function thaw_processes() is called in
order to wake up each frozen task. Then, the tasks that have been frozen leave
__refrigerator() and continue running.
Rationale behind the functions dealing with freezing and thawing of tasks
-------------------------------------------------------------------------
freeze_processes():
- freezes only userspace tasks
freeze_kernel_threads():
- freezes all tasks (including kernel threads) because we can't freeze
kernel threads without freezing userspace tasks
thaw_kernel_threads():
- thaws only kernel threads; this is particularly useful if we need to do
anything special in between thawing of kernel threads and thawing of
userspace tasks, or if we want to postpone the thawing of userspace tasks
thaw_processes():
- thaws all tasks (including kernel threads) because we can't thaw userspace
tasks without thawing kernel threads
III. Which kernel threads are freezable?
========================================
Kernel threads are not freezable by default. However, a kernel thread may clear
PF_NOFREEZE for itself by calling set_freezable() (the resetting of PF_NOFREEZE
directly is not allowed). From this point it is regarded as freezable
and must call try_to_freeze() or variants of wait_event_freezable() in a
suitable place.
IV. Why do we do that?
======================
Generally speaking, there is a couple of reasons to use the freezing of tasks:
1. The principal reason is to prevent filesystems from being damaged after
hibernation. At the moment we have no simple means of checkpointing
filesystems, so if there are any modifications made to filesystem data and/or
metadata on disks, we cannot bring them back to the state from before the
modifications. At the same time each hibernation image contains some
filesystem-related information that must be consistent with the state of the
on-disk data and metadata after the system memory state has been restored
from the image (otherwise the filesystems will be damaged in a nasty way,
usually making them almost impossible to repair). We therefore freeze
tasks that might cause the on-disk filesystems' data and metadata to be
modified after the hibernation image has been created and before the
system is finally powered off. The majority of these are user space
processes, but if any of the kernel threads may cause something like this
to happen, they have to be freezable.
2. Next, to create the hibernation image we need to free a sufficient amount of
memory (approximately 50% of available RAM) and we need to do that before
devices are deactivated, because we generally need them for swapping out.
Then, after the memory for the image has been freed, we don't want tasks
to allocate additional memory and we prevent them from doing that by
freezing them earlier. [Of course, this also means that device drivers
should not allocate substantial amounts of memory from their .suspend()
callbacks before hibernation, but this is a separate issue.]
3. The third reason is to prevent user space processes and some kernel threads
from interfering with the suspending and resuming of devices. A user space
process running on a second CPU while we are suspending devices may, for
example, be troublesome and without the freezing of tasks we would need some
safeguards against race conditions that might occur in such a case.
Although Linus Torvalds doesn't like the freezing of tasks, he said this in one
of the discussions on LKML (https://lore.kernel.org/r/[email protected]):
"RJW:> Why we freeze tasks at all or why we freeze kernel threads?
Linus: In many ways, 'at all'.
I **do** realize the IO request queue issues, and that we cannot actually do
s2ram with some devices in the middle of a DMA. So we want to be able to
avoid *that*, there's no question about that. And I suspect that stopping
user threads and then waiting for a sync is practically one of the easier
ways to do so.
So in practice, the 'at all' may become a 'why freeze kernel threads?' and
freezing user threads I don't find really objectionable."
Still, there are kernel threads that may want to be freezable. For example, if
a kernel thread that belongs to a device driver accesses the device directly, it
in principle needs to know when the device is suspended, so that it doesn't try
to access it at that time. However, if the kernel thread is freezable, it will
be frozen before the driver's .suspend() callback is executed and it will be
thawed after the driver's .resume() callback has run, so it won't be accessing
the device while it's suspended.
4. Another reason for freezing tasks is to prevent user space processes from
realizing that hibernation (or suspend) operation takes place. Ideally, user
space processes should not notice that such a system-wide operation has
occurred and should continue running without any problems after the restore
(or resume from suspend). Unfortunately, in the most general case this
is quite difficult to achieve without the freezing of tasks. Consider,
for example, a process that depends on all CPUs being online while it's
running. Since we need to disable nonboot CPUs during the hibernation,
if this process is not frozen, it may notice that the number of CPUs has
changed and may start to work incorrectly because of that.
V. Are there any problems related to the freezing of tasks?
===========================================================
Yes, there are.
First of all, the freezing of kernel threads may be tricky if they depend one
on another. For example, if kernel thread A waits for a completion (in the
TASK_UNINTERRUPTIBLE state) that needs to be done by freezable kernel thread B
and B is frozen in the meantime, then A will be blocked until B is thawed, which
may be undesirable. That's why kernel threads are not freezable by default.
Second, there are the following two problems related to the freezing of user
space processes:
1. Putting processes into an uninterruptible sleep distorts the load average.
2. Now that we have FUSE, plus the framework for doing device drivers in
userspace, it gets even more complicated because some userspace processes are
now doing the sorts of things that kernel threads do
(https://lists.linux-foundation.org/pipermail/linux-pm/2007-May/012309.html).
The problem 1. seems to be fixable, although it hasn't been fixed so far. The
other one is more serious, but it seems that we can work around it by using
hibernation (and suspend) notifiers (in that case, though, we won't be able to
avoid the realization by the user space processes that the hibernation is taking
place).
There are also problems that the freezing of tasks tends to expose, although
they are not directly related to it. For example, if request_firmware() is
called from a device driver's .resume() routine, it will timeout and eventually
fail, because the user land process that should respond to the request is frozen
at this point. So, seemingly, the failure is due to the freezing of tasks.
Suppose, however, that the firmware file is located on a filesystem accessible
only through another device that hasn't been resumed yet. In that case,
request_firmware() will fail regardless of whether or not the freezing of tasks
is used. Consequently, the problem is not really related to the freezing of
tasks, since it generally exists anyway.
A driver must have all firmwares it may need in RAM before suspend() is called.
If keeping them is not practical, for example due to their size, they must be
requested early enough using the suspend notifier API described in
Documentation/driver-api/pm/notifiers.rst.
VI. Are there any precautions to be taken to prevent freezing failures?
=======================================================================
Yes, there are.
First of all, grabbing the 'system_transition_mutex' lock to mutually exclude a
piece of code from system-wide sleep such as suspend/hibernation is not
encouraged. If possible, that piece of code must instead hook onto the
suspend/hibernation notifiers to achieve mutual exclusion. Look at the
CPU-Hotplug code (kernel/cpu.c) for an example.
However, if that is not feasible, and grabbing 'system_transition_mutex' is
deemed necessary, it is strongly discouraged to directly call
mutex_[un]lock(&system_transition_mutex) since that could lead to freezing
failures, because if the suspend/hibernate code successfully acquired the
'system_transition_mutex' lock, and hence that other entity failed to acquire
the lock, then that task would get blocked in TASK_UNINTERRUPTIBLE state. As a
consequence, the freezer would not be able to freeze that task, leading to
freezing failure.
However, the [un]lock_system_sleep() APIs are safe to use in this scenario,
since they ask the freezer to skip freezing this task, since it is anyway
"frozen enough" as it is blocked on 'system_transition_mutex', which will be
released only after the entire suspend/hibernation sequence is complete. So, to
summarize, use [un]lock_system_sleep() instead of directly using
mutex_[un]lock(&system_transition_mutex). That would prevent freezing failures.
V. Miscellaneous
================
/sys/power/pm_freeze_timeout controls how long it will cost at most to freeze
all user space processes or all freezable kernel threads, in unit of
millisecond. The default value is 20000, with range of unsigned integer.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
Task freezing이란
1-13Task freezing은 hibernation 또는 일부 architecture의 system-wide suspend 동안 userspace process와 일부 kernel thread의 실행을 통제하는 mechanism입니다.
목적은 단순히 모든 task를 멈추는 것이 아니라, sleep transition이 의존하는 filesystem·memory·device 상태가 바뀌지 않도록 freezable task만 정해진 순서로 정지시키는 데 있습니다.
System-wide power transition 중 userspace 전체와 명시적으로 참여한 kernel thread를 멈춥니다.
=================
Freezing of tasks
=================
(C) 2007 Rafael J. Wysocki <[email protected]>, GPL
I. What is the freezing of tasks?
=================================
The freezing of tasks is a mechanism by which user space processes and some
kernel threads are controlled during hibernation or system-wide suspend (on some
architectures).
Flag·state와 freeze 진입
14-43이 mechanism은 per-task flag `PF_NOFREEZE`와 세 per-task state `TASK_FROZEN`, `TASK_FREEZABLE`, `__TASK_FREEZABLE_UNSAFE`를 사용합니다. `PF_NOFREEZE`가 설정되지 않은 모든 userspace task와 일부 kernel thread가 freezable입니다.
이 task들은 system이 sleep state로 들어가기 전과 hibernation image를 만들기 전에 특별히 처리됩니다. 아래 설명은 hibernation을 직접 다루지만 system-wide suspend에도 적용됩니다.
Hibernation의 첫 단계에서 `kernel/power/process.c`의 `freeze_processes()`가 system-wide static key `freezer_active`를 설정합니다. 이는 per-task flag나 state가 아니라 system 전체가 freeze operation 중임을 나타냅니다.
이어서 `try_to_freeze_tasks()`가 모든 userspace process에 fake signal을 보내고 모든 kernel thread를 깨웁니다. Userspace의 signal 경로에서는 `TIF_SIGPENDING`으로 대표되는 pending signal 상태를 통해 freeze 검사가 이루어지며, freezable task는 `try_to_freeze()`를 호출해야 합니다.
`try_to_freeze()`는 `kernel/freezer.c`의 `__refrigerator()`로 들어가 task state를 `TASK_FROZEN`으로 바꾸고, 명시적인 `TASK_FROZEN` wakeup을 받을 때까지 loop합니다. 이 상태를 frozen이라고 하며 관련 함수 집합을 freezer라고 부릅니다. 구현은 `kernel/power/process.c`, `kernel/freezer.c`, `include/linux/freezer.h`에 있습니다. 일반적으로 userspace task를 kernel thread보다 먼저 freeze합니다.
`__refrigerator()`를 직접 호출하면 안 됩니다. 대신 task가 실제로 freeze 대상인지 검사한 뒤 진입시키는 `include/linux/freezer.h`의 `try_to_freeze()`를 사용합니다.
System-wide static key를 켠 뒤 각 freezable task가 refrigerator loop에 들어갑니다.
II. How does it work?
=====================
There is one per-task flag (PF_NOFREEZE) and three per-task states
(TASK_FROZEN, TASK_FREEZABLE and __TASK_FREEZABLE_UNSAFE) used for that.
The tasks that have PF_NOFREEZE unset (all user space tasks and some kernel
threads) are regarded as 'freezable' and treated in a special way before the
system enters a sleep state as well as before a hibernation image is created
(hibernation is directly covered by what follows, but the description applies
to system-wide suspend too).
Namely, as the first step of the hibernation procedure the function
freeze_processes() (defined in kernel/power/process.c) is called. A system-wide
static key freezer_active (as opposed to a per-task flag or state) is used to
indicate whether the system is to undergo a freezing operation. And
freeze_processes() sets this static key. After this, it executes
try_to_freeze_tasks() that sends a fake signal to all user space processes, and
wakes up all the kernel threads. All freezable tasks must react to that by
calling try_to_freeze(), which results in a call to __refrigerator() (defined
in kernel/freezer.c), which changes the task's state to TASK_FROZEN, and makes
it loop until it is woken by an explicit TASK_FROZEN wakeup. Then, that task
is regarded as 'frozen' and so the set of functions handling this mechanism is
referred to as 'the freezer' (these functions are defined in
kernel/power/process.c, kernel/freezer.c & include/linux/freezer.h). User space
tasks are generally frozen before kernel threads.
__refrigerator() must not be called directly. Instead, use the
try_to_freeze() function (defined in include/linux/freezer.h), that checks
if the task is to be frozen and makes the task enter __refrigerator().
Freezable kernel thread와 thaw
44-82Userspace process에서는 signal handling code가 `try_to_freeze()`를 자동 호출합니다. Freezable kernel thread는 적절한 위치에서 직접 호출하거나 `include/linux/wait.h`의 `wait_event_freezable()` 또는 `wait_event_freezable_timeout()`을 사용해야 합니다.
이 macro들은 평소 task를 `TASK_INTERRUPTIBLE`로 재우고, `freezer_active`가 설정되면 `TASK_FROZEN`으로 freeze합니다.
예제는 `mm/oom_kill.c::oom_reaper()`의 main loop입니다. `set_freezable()`로 참여를 선언하고 `wait_event_freezable(oom_reaper_wait, oom_reaper_list != NULL)`에서 work를 기다립니다. 깨어나면 lock 아래 list에서 task를 꺼내 `oom_reap_task(tsk)`를 실행합니다.
Freezer가 시작됐는데 freezable kernel thread가 frozen state에 들어가지 않으면 task freezing이 실패하고 system-wide transition 전체가 취소됩니다. 따라서 반드시 `try_to_freeze()` 또는 freezable wait macro 중 하나를 사용해야 합니다.
Hibernation image에서 system memory가 복원되고 device가 재초기화되면 `thaw_processes()`가 frozen task를 깨웁니다. Task는 `__refrigerator()`를 빠져나와 실행을 계속합니다.
Freezer 요청을 관찰하는 경로가 다릅니다.
For user space processes try_to_freeze() is called automatically from the
signal-handling code, but the freezable kernel threads need to call it
explicitly in suitable places or use the wait_event_freezable() or
wait_event_freezable_timeout() macros (defined in include/linux/wait.h)
that put the task to sleep (TASK_INTERRUPTIBLE) or freeze it (TASK_FROZEN) if
freezer_active is set. The main loop of a freezable kernel thread may look
like the following one::
set_freezable();
while (true) {
struct task_struct *tsk = NULL;
wait_event_freezable(oom_reaper_wait, oom_reaper_list != NULL);
spin_lock_irq(&oom_reaper_lock);
if (oom_reaper_list != NULL) {
tsk = oom_reaper_list;
oom_reaper_list = tsk->oom_reaper_list;
}
spin_unlock_irq(&oom_reaper_lock);
if (tsk)
oom_reap_task(tsk);
}
(from mm/oom_kill.c::oom_reaper()).
If a freezable kernel thread is not put to the frozen state after the freezer
has initiated a freezing operation, the freezing of tasks will fail and the
entire system-wide transition will be cancelled. For this reason, freezable
kernel threads must call try_to_freeze() somewhere or use one of the
wait_event_freezable() and wait_event_freezable_timeout() macros.
After the system memory state has been restored from a hibernation image and
devices have been reinitialized, the function thaw_processes() is called in
order to wake up each frozen task. Then, the tasks that have been frozen leave
__refrigerator() and continue running.
Freeze·thaw 함수의 역할
83-102Freeze와 thaw 함수를 나눈 이유는 userspace와 kernel thread의 순서를 제어하기 위해서입니다.
Kernel thread는 userspace를 freeze한 뒤에만 freeze할 수 있고, userspace는 kernel thread를 thaw한 뒤에만 thaw할 수 있습니다.
`thaw_kernel_threads()`와 `thaw_processes()`를 분리하면 kernel thread를 깨운 뒤 userspace를 깨우기 전에 특별한 작업을 수행하거나 userspace thaw를 늦출 수 있습니다.
정지할 때는 userspace→kernel thread, 재개할 때는 kernel thread→userspace 순서입니다.
Rationale behind the functions dealing with freezing and thawing of tasks
-------------------------------------------------------------------------
freeze_processes():
- freezes only userspace tasks
freeze_kernel_threads():
- freezes all tasks (including kernel threads) because we can't freeze
kernel threads without freezing userspace tasks
thaw_kernel_threads():
- thaws only kernel threads; this is particularly useful if we need to do
anything special in between thawing of kernel threads and thawing of
userspace tasks, or if we want to postpone the thawing of userspace tasks
thaw_processes():
- thaws all tasks (including kernel threads) because we can't thaw userspace
tasks without thawing kernel threads
어떤 kernel thread가 freezable인가
103-111Kernel thread는 기본적으로 freezable이 아닙니다. 참여하려는 thread는 `set_freezable()`을 호출해 자신의 `PF_NOFREEZE`를 clear할 수 있습니다. `PF_NOFREEZE`를 직접 reset하는 것은 허용되지 않습니다.
그 시점부터 thread는 freezable로 간주되며 적절한 위치에서 `try_to_freeze()` 또는 `wait_event_freezable()` 계열을 호출해야 합니다.
명시적으로 참여를 선언한 thread만 freezer protocol을 따라야 합니다.
III. Which kernel threads are freezable?
========================================
Kernel threads are not freezable by default. However, a kernel thread may clear
PF_NOFREEZE for itself by calling set_freezable() (the resetting of PF_NOFREEZE
directly is not allowed). From this point it is regarded as freezable
and must call try_to_freeze() or variants of wait_event_freezable() in a
suitable place.
Freezer가 필요한 핵심 이유
112-146첫째이자 가장 중요한 이유는 hibernation 뒤 filesystem 손상을 막는 것입니다. 현재 filesystem을 간단히 checkpoint하는 방법이 없으므로 disk의 data나 metadata가 바뀌면 변경 전 상태로 되돌릴 수 없습니다.
Hibernation image에는 filesystem 관련 상태가 들어 있고, 복원 뒤 disk의 data·metadata와 일치해야 합니다. 그렇지 않으면 filesystem이 심하게 손상되어 복구가 거의 불가능할 수 있습니다. 따라서 image 생성 뒤 최종 power-off 전까지 disk filesystem을 바꿀 수 있는 task를 freeze합니다. 대부분 userspace process지만 그런 변경을 일으킬 수 있는 kernel thread도 freezable이어야 합니다.
둘째, hibernation image를 만들려면 available RAM의 약 50%를 비워야 합니다. Swap-out에 device가 필요하므로 device를 deactivate하기 전에 memory를 확보합니다. 확보한 뒤 task가 다시 memory를 할당하지 못하도록 먼저 freeze합니다. 별개로 driver의 `.suspend()` callback도 hibernation 전에 많은 memory를 할당하지 않아야 합니다.
셋째, userspace process와 일부 kernel thread가 device suspend·resume를 방해하지 못하게 합니다. Device를 suspend하는 동안 두 번째 CPU에서 process가 실행되면 race condition이 생길 수 있으므로 freezer가 없으면 별도의 safeguard가 필요합니다.
Hibernation image와 외부 상태 사이의 일관성을 지킵니다.
IV. Why do we do that?
======================
Generally speaking, there is a couple of reasons to use the freezing of tasks:
1. The principal reason is to prevent filesystems from being damaged after
hibernation. At the moment we have no simple means of checkpointing
filesystems, so if there are any modifications made to filesystem data and/or
metadata on disks, we cannot bring them back to the state from before the
modifications. At the same time each hibernation image contains some
filesystem-related information that must be consistent with the state of the
on-disk data and metadata after the system memory state has been restored
from the image (otherwise the filesystems will be damaged in a nasty way,
usually making them almost impossible to repair). We therefore freeze
tasks that might cause the on-disk filesystems' data and metadata to be
modified after the hibernation image has been created and before the
system is finally powered off. The majority of these are user space
processes, but if any of the kernel threads may cause something like this
to happen, they have to be freezable.
2. Next, to create the hibernation image we need to free a sufficient amount of
memory (approximately 50% of available RAM) and we need to do that before
devices are deactivated, because we generally need them for swapping out.
Then, after the memory for the image has been freed, we don't want tasks
to allocate additional memory and we prevent them from doing that by
freezing them earlier. [Of course, this also means that device drivers
should not allocate substantial amounts of memory from their .suspend()
callbacks before hibernation, but this is a separate issue.]
3. The third reason is to prevent user space processes and some kernel threads
from interfering with the suspending and resuming of devices. A user space
process running on a second CPU while we are suspending devices may, for
example, be troublesome and without the freezing of tasks we would need some
safeguards against race conditions that might occur in such a case.
Kernel thread와 userspace 투명성
147-181문서는 Linus Torvalds의 LKML 논의를 인용합니다. 핵심은 DMA 중인 일부 device로는 s2ram을 진행할 수 없으므로 I/O request를 멈춰야 하고, userspace thread를 멈춘 뒤 sync를 기다리는 방식은 실용적이라는 것입니다. 반면 모든 kernel thread까지 freeze해야 하는지는 더 의문스럽다는 취지입니다.
그래도 freezable이어야 하는 kernel thread가 있습니다. Device driver 소속 thread가 hardware에 직접 접근한다면 device가 suspend된 동안 접근해서는 안 됩니다. Thread를 freezable로 만들면 driver의 `.suspend()` callback 전에 freeze되고 `.resume()` callback 뒤에 thaw되므로 정지된 device에 접근하지 않습니다.
넷째 이유는 userspace process가 hibernation 또는 suspend가 일어났다는 사실을 알아차리지 못하게 하는 것입니다. 이상적으로 process는 system-wide transition 뒤에도 문제없이 이어서 실행해야 합니다.
예를 들어 실행 중 모든 CPU가 online이라고 가정하는 process가 있습니다. Hibernation 중 nonboot CPU를 disable해야 하는데 process를 freeze하지 않으면 CPU 수 변경을 감지하고 잘못 동작할 수 있습니다.
Driver callback을 경계로 kernel thread를 멈춰 suspended device 접근을 차단합니다.
Although Linus Torvalds doesn't like the freezing of tasks, he said this in one
of the discussions on LKML (https://lore.kernel.org/r/[email protected]):
"RJW:> Why we freeze tasks at all or why we freeze kernel threads?
Linus: In many ways, 'at all'.
I **do** realize the IO request queue issues, and that we cannot actually do
s2ram with some devices in the middle of a DMA. So we want to be able to
avoid *that*, there's no question about that. And I suspect that stopping
user threads and then waiting for a sync is practically one of the easier
ways to do so.
So in practice, the 'at all' may become a 'why freeze kernel threads?' and
freezing user threads I don't find really objectionable."
Still, there are kernel threads that may want to be freezable. For example, if
a kernel thread that belongs to a device driver accesses the device directly, it
in principle needs to know when the device is suspended, so that it doesn't try
to access it at that time. However, if the kernel thread is freezable, it will
be frozen before the driver's .suspend() callback is executed and it will be
thawed after the driver's .resume() callback has run, so it won't be accessing
the device while it's suspended.
4. Another reason for freezing tasks is to prevent user space processes from
realizing that hibernation (or suspend) operation takes place. Ideally, user
space processes should not notice that such a system-wide operation has
occurred and should continue running without any problems after the restore
(or resume from suspend). Unfortunately, in the most general case this
is quite difficult to achieve without the freezing of tasks. Consider,
for example, a process that depends on all CPUs being online while it's
running. Since we need to disable nonboot CPUs during the hibernation,
if this process is not frozen, it may notice that the number of CPUs has
changed and may start to work incorrectly because of that.
Freezing이 드러내는 문제
182-223Kernel thread끼리 의존하면 freezing이 까다롭습니다. Thread A가 `TASK_UNINTERRUPTIBLE` 상태로 completion을 기다리고 그 completion을 수행할 freezable thread B가 먼저 frozen되면, A는 B가 thaw될 때까지 막힙니다. 이것이 kernel thread가 기본적으로 freezable이 아닌 이유입니다.
Userspace freezing에는 두 문제가 있습니다. Process를 uninterruptible sleep에 넣으면 load average가 왜곡됩니다. 또 FUSE와 userspace device-driver framework 때문에 일부 userspace process가 kernel thread와 같은 역할을 수행하므로 의존 관계가 복잡해집니다.
Load average 문제는 고칠 수 있을 것으로 보이지만 아직 해결되지 않았습니다. 두 번째 문제는 더 심각하며 hibernation·suspend notifier로 우회할 수 있지만, 그러면 userspace가 transition을 알아차리지 못하게 한다는 목표는 포기해야 합니다.
Freezer가 원인처럼 보이지만 원래 존재하던 문제를 드러내기도 합니다. Driver `.resume()`에서 `request_firmware()`를 호출하면 응답할 userland process가 frozen이어서 timeout될 수 있습니다. 그러나 firmware file이 아직 resume되지 않은 다른 device를 통해서만 접근 가능한 filesystem에 있다면 freezer가 없어도 실패합니다.
Driver는 필요할 모든 firmware를 `suspend()` 호출 전에 RAM에 보유해야 합니다. 크기 때문에 계속 유지하기 어렵다면 `Documentation/driver-api/pm/notifiers.rst`의 suspend notifier API를 이용해 충분히 일찍 요청해야 합니다.
직접적인 freezer 교착과 원래 존재하던 resume 의존 문제를 구분해야 합니다.
V. Are there any problems related to the freezing of tasks?
===========================================================
Yes, there are.
First of all, the freezing of kernel threads may be tricky if they depend one
on another. For example, if kernel thread A waits for a completion (in the
TASK_UNINTERRUPTIBLE state) that needs to be done by freezable kernel thread B
and B is frozen in the meantime, then A will be blocked until B is thawed, which
may be undesirable. That's why kernel threads are not freezable by default.
Second, there are the following two problems related to the freezing of user
space processes:
1. Putting processes into an uninterruptible sleep distorts the load average.
2. Now that we have FUSE, plus the framework for doing device drivers in
userspace, it gets even more complicated because some userspace processes are
now doing the sorts of things that kernel threads do
(https://lists.linux-foundation.org/pipermail/linux-pm/2007-May/012309.html).
The problem 1. seems to be fixable, although it hasn't been fixed so far. The
other one is more serious, but it seems that we can work around it by using
hibernation (and suspend) notifiers (in that case, though, we won't be able to
avoid the realization by the user space processes that the hibernation is taking
place).
There are also problems that the freezing of tasks tends to expose, although
they are not directly related to it. For example, if request_firmware() is
called from a device driver's .resume() routine, it will timeout and eventually
fail, because the user land process that should respond to the request is frozen
at this point. So, seemingly, the failure is due to the freezing of tasks.
Suppose, however, that the firmware file is located on a filesystem accessible
only through another device that hasn't been resumed yet. In that case,
request_firmware() will fail regardless of whether or not the freezing of tasks
is used. Consequently, the problem is not really related to the freezing of
tasks, since it generally exists anyway.
A driver must have all firmwares it may need in RAM before suspend() is called.
If keeping them is not practical, for example due to their size, they must be
requested early enough using the suspend notifier API described in
Documentation/driver-api/pm/notifiers.rst.
Freezing 실패를 막는 locking
224-250System-wide suspend·hibernation과 code 구간을 상호 배제하려고 `system_transition_mutex`를 직접 잡는 방식은 권장되지 않습니다. 가능하면 suspend/hibernation notifier를 연결해야 하며 `kernel/cpu.c`의 CPU-Hotplug code가 예입니다.
직접 `mutex_[un]lock(&system_transition_mutex)`을 호출하면 freezing 실패가 생길 수 있습니다. Suspend code가 lock을 먼저 얻은 상태에서 다른 task가 lock을 기다리면 그 task는 `TASK_UNINTERRUPTIBLE`로 막히고 freezer가 freeze하지 못합니다.
이 경우 `[un]lock_system_sleep()` API는 안전합니다. 이 API는 해당 task가 `system_transition_mutex`에 막혀 이미 충분히 frozen된 상태이므로 freezer가 freeze count에서 건너뛰도록 표시합니다. 구현상 이런 경로는 `PF_FREEZER_SKIP` 같은 freezer skip 상태를 사용하고, mutex는 suspend/hibernation sequence 전체가 끝난 뒤에만 풀립니다.
정리하면 `mutex_[un]lock(&system_transition_mutex)`을 직접 쓰지 말고 `[un]lock_system_sleep()`을 사용해야 freezing 실패를 막을 수 있습니다.
Notifier가 우선이며 mutex가 꼭 필요하면 freezer-aware wrapper를 사용합니다.
VI. Are there any precautions to be taken to prevent freezing failures?
=======================================================================
Yes, there are.
First of all, grabbing the 'system_transition_mutex' lock to mutually exclude a
piece of code from system-wide sleep such as suspend/hibernation is not
encouraged. If possible, that piece of code must instead hook onto the
suspend/hibernation notifiers to achieve mutual exclusion. Look at the
CPU-Hotplug code (kernel/cpu.c) for an example.
However, if that is not feasible, and grabbing 'system_transition_mutex' is
deemed necessary, it is strongly discouraged to directly call
mutex_[un]lock(&system_transition_mutex) since that could lead to freezing
failures, because if the suspend/hibernate code successfully acquired the
'system_transition_mutex' lock, and hence that other entity failed to acquire
the lock, then that task would get blocked in TASK_UNINTERRUPTIBLE state. As a
consequence, the freezer would not be able to freeze that task, leading to
freezing failure.
However, the [un]lock_system_sleep() APIs are safe to use in this scenario,
since they ask the freezer to skip freezing this task, since it is anyway
"frozen enough" as it is blocked on 'system_transition_mutex', which will be
released only after the entire suspend/hibernation sequence is complete. So, to
summarize, use [un]lock_system_sleep() instead of directly using
mutex_[un]lock(&system_transition_mutex). That would prevent freezing failures.
pm_freeze_timeout
251-256원문은 마지막 절 번호를 다시 V로 표기합니다. `/sys/power/pm_freeze_timeout`은 모든 userspace process 또는 모든 freezable kernel thread를 freeze하는 데 허용할 최대 시간을 밀리초 단위로 제어합니다.
기본값은 `20000`, 즉 20초이며 unsigned integer 범위를 사용합니다.
Userspace와 kernel-thread freeze 단계 각각의 최대 대기 시간을 제한합니다.
V. Miscellaneous
================
/sys/power/pm_freeze_timeout controls how long it will cost at most to freeze
all user space processes or all freezable kernel threads, in unit of
millisecond. The default value is 20000, with range of unsigned integer.
요약·해설
freezing-of-tasks.rst:1-256Task freezer는 image와 disk·device 상태의 일관성을 지키기 위해 userspace를 먼저, 참여한 kernel thread를 다음으로 멈춥니다. Kernel thread의 opt-in 규칙과 dependency deadlock, firmware 선로딩, freezer-aware system sleep lock 사용이 핵심입니다.