← Documents Documentation/power/pci.rst GitHub 원문 ↗

Linux 6.18.37 · Power

PCI Power Management

Native PCI PM과 ACPI firmware, PME·GPE wakeup, PCI subsystem의 system sleep·hibernation·runtime callback, PCI driver의 dev_pm_ops·flag·usage-counter 규칙을 설명합니다.

Source pathDocumentation/power/pci.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

pci.rst:1-1132

PCI PM은 native D-state와 platform firmware를 함께 사용하고, PCI subsystem이 register 저장·D0 복귀·wakeup 준비를 driver 대신 처리합니다. Driver는 device quiesce·기능 복원과 runtime usage counter 균형에 집중하는 것이 핵심입니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 ====================
2 PCI Power Management
3 ====================
4
5 Copyright (c) 2010 Rafael J. Wysocki <[email protected]>, Novell Inc.
6
7 An overview of concepts and the Linux kernel's interfaces related to PCI power
8 management. Based on previous work by Patrick Mochel <[email protected]>
9 (and others).
10
11 This document only covers the aspects of power management specific to PCI
12 devices. For general description of the kernel's interfaces related to device
13 power management refer to Documentation/driver-api/pm/devices.rst and
14 Documentation/power/runtime_pm.rst.
15
16 .. contents:
17
18 1. Hardware and Platform Support for PCI Power Management
19 2. PCI Subsystem and Device Power Management
20 3. PCI Device Drivers and Power Management
21 4. Resources
22
23
24 1. Hardware and Platform Support for PCI Power Management
25 =========================================================
26
27 1.1. Native and Platform-Based Power Management
28 -----------------------------------------------
29
30 In general, power management is a feature allowing one to save energy by putting
31 devices into states in which they draw less power (low-power states) at the
32 price of reduced functionality or performance.
33
34 Usually, a device is put into a low-power state when it is underutilized or
35 completely inactive. However, when it is necessary to use the device once
36 again, it has to be put back into the "fully functional" state (full-power
37 state). This may happen when there are some data for the device to handle or
38 as a result of an external event requiring the device to be active, which may
39 be signaled by the device itself.
40
41 PCI devices may be put into low-power states in two ways, by using the device
42 capabilities introduced by the PCI Bus Power Management Interface Specification,
43 or with the help of platform firmware, such as an ACPI BIOS. In the first
44 approach, that is referred to as the native PCI power management (native PCI PM)
45 in what follows, the device power state is changed as a result of writing a
46 specific value into one of its standard configuration registers. The second
47 approach requires the platform firmware to provide special methods that may be
48 used by the kernel to change the device's power state.
49
50 Devices supporting the native PCI PM usually can generate wakeup signals called
51 Power Management Events (PMEs) to let the kernel know about external events
52 requiring the device to be active. After receiving a PME the kernel is supposed
53 to put the device that sent it into the full-power state. However, the PCI Bus
54 Power Management Interface Specification doesn't define any standard method of
55 delivering the PME from the device to the CPU and the operating system kernel.
56 It is assumed that the platform firmware will perform this task and therefore,
57 even though a PCI device is set up to generate PMEs, it also may be necessary to
58 prepare the platform firmware for notifying the CPU of the PMEs coming from the
59 device (e.g. by generating interrupts).
60
61 In turn, if the methods provided by the platform firmware are used for changing
62 the power state of a device, usually the platform also provides a method for
63 preparing the device to generate wakeup signals. In that case, however, it
64 often also is necessary to prepare the device for generating PMEs using the
65 native PCI PM mechanism, because the method provided by the platform depends on
66 that.
67
68 Thus in many situations both the native and the platform-based power management
69 mechanisms have to be used simultaneously to obtain the desired result.
70
71 1.2. Native PCI Power Management
72 --------------------------------
73
74 The PCI Bus Power Management Interface Specification (PCI PM Spec) was
75 introduced between the PCI 2.1 and PCI 2.2 Specifications. It defined a
76 standard interface for performing various operations related to power
77 management.
78
79 The implementation of the PCI PM Spec is optional for conventional PCI devices,
80 but it is mandatory for PCI Express devices. If a device supports the PCI PM
81 Spec, it has an 8 byte power management capability field in its PCI
82 configuration space. This field is used to describe and control the standard
83 features related to the native PCI power management.
84
85 The PCI PM Spec defines 4 operating states for devices (D0-D3) and for buses
86 (B0-B3). The higher the number, the less power is drawn by the device or bus
87 in that state. However, the higher the number, the longer the latency for
88 the device or bus to return to the full-power state (D0 or B0, respectively).
89
90 There are two variants of the D3 state defined by the specification. The first
91 one is D3hot, referred to as the software accessible D3, because devices can be
92 programmed to go into it. The second one, D3cold, is the state that PCI devices
93 are in when the supply voltage (Vcc) is removed from them. It is not possible
94 to program a PCI device to go into D3cold, although there may be a programmable
95 interface for putting the bus the device is on into a state in which Vcc is
96 removed from all devices on the bus.
97
98 PCI bus power management, however, is not supported by the Linux kernel at the
99 time of this writing and therefore it is not covered by this document.
100
101 Note that every PCI device can be in the full-power state (D0) or in D3cold,
102 regardless of whether or not it implements the PCI PM Spec. In addition to
103 that, if the PCI PM Spec is implemented by the device, it must support D3hot
104 as well as D0. The support for the D1 and D2 power states is optional.
105
106 PCI devices supporting the PCI PM Spec can be programmed to go to any of the
107 supported low-power states (except for D3cold). While in D1-D3hot the
108 standard configuration registers of the device must be accessible to software
109 (i.e. the device is required to respond to PCI configuration accesses), although
110 its I/O and memory spaces are then disabled. This allows the device to be
111 programmatically put into D0. Thus the kernel can switch the device back and
112 forth between D0 and the supported low-power states (except for D3cold) and the
113 possible power state transitions the device can undergo are the following:
114
115 +----------------------------+
116 | Current State | New State |
117 +----------------------------+
118 | D0 | D1, D2, D3 |
119 +----------------------------+
120 | D1 | D2, D3 |
121 +----------------------------+
122 | D2 | D3 |
123 +----------------------------+
124 | D1, D2, D3 | D0 |
125 +----------------------------+
126
127 The transition from D3cold to D0 occurs when the supply voltage is provided to
128 the device (i.e. power is restored). In that case the device returns to D0 with
129 a full power-on reset sequence and the power-on defaults are restored to the
130 device by hardware just as at initial power up.
131
132 PCI devices supporting the PCI PM Spec can be programmed to generate PMEs
133 while in any power state (D0-D3), but they are not required to be capable
134 of generating PMEs from all supported power states. In particular, the
135 capability of generating PMEs from D3cold is optional and depends on the
136 presence of additional voltage (3.3Vaux) allowing the device to remain
137 sufficiently active to generate a wakeup signal.
138
139 1.3. ACPI Device Power Management
140 ---------------------------------
141
142 The platform firmware support for the power management of PCI devices is
143 system-specific. However, if the system in question is compliant with the
144 Advanced Configuration and Power Interface (ACPI) Specification, like the
145 majority of x86-based systems, it is supposed to implement device power
146 management interfaces defined by the ACPI standard.
147
148 For this purpose the ACPI BIOS provides special functions called "control
149 methods" that may be executed by the kernel to perform specific tasks, such as
150 putting a device into a low-power state. These control methods are encoded
151 using special byte-code language called the ACPI Machine Language (AML) and
152 stored in the machine's BIOS. The kernel loads them from the BIOS and executes
153 them as needed using an AML interpreter that translates the AML byte code into
154 computations and memory or I/O space accesses. This way, in theory, a BIOS
155 writer can provide the kernel with a means to perform actions depending
156 on the system design in a system-specific fashion.
157
158 ACPI control methods may be divided into global control methods, that are not
159 associated with any particular devices, and device control methods, that have
160 to be defined separately for each device supposed to be handled with the help of
161 the platform. This means, in particular, that ACPI device control methods can
162 only be used to handle devices that the BIOS writer knew about in advance. The
163 ACPI methods used for device power management fall into that category.
164
165 The ACPI specification assumes that devices can be in one of four power states
166 labeled as D0, D1, D2, and D3 that roughly correspond to the native PCI PM
167 D0-D3 states (although the difference between D3hot and D3cold is not taken
168 into account by ACPI). Moreover, for each power state of a device there is a
169 set of power resources that have to be enabled for the device to be put into
170 that state. These power resources are controlled (i.e. enabled or disabled)
171 with the help of their own control methods, _ON and _OFF, that have to be
172 defined individually for each of them.
173
174 To put a device into the ACPI power state Dx (where x is a number between 0 and
175 3 inclusive) the kernel is supposed to (1) enable the power resources required
176 by the device in this state using their _ON control methods and (2) execute the
177 _PSx control method defined for the device. In addition to that, if the device
178 is going to be put into a low-power state (D1-D3) and is supposed to generate
179 wakeup signals from that state, the _DSW (or _PSW, replaced with _DSW by ACPI
180 3.0) control method defined for it has to be executed before _PSx. Power
181 resources that are not required by the device in the target power state and are
182 not required any more by any other device should be disabled (by executing their
183 _OFF control methods). If the current power state of the device is D3, it can
184 only be put into D0 this way.
185
186 However, quite often the power states of devices are changed during a
187 system-wide transition into a sleep state or back into the working state. ACPI
188 defines four system sleep states, S1, S2, S3, and S4, and denotes the system
189 working state as S0. In general, the target system sleep (or working) state
190 determines the highest power (lowest number) state the device can be put
191 into and the kernel is supposed to obtain this information by executing the
192 device's _SxD control method (where x is a number between 0 and 4 inclusive).
193 If the device is required to wake up the system from the target sleep state, the
194 lowest power (highest number) state it can be put into is also determined by the
195 target state of the system. The kernel is then supposed to use the device's
196 _SxW control method to obtain the number of that state. It also is supposed to
197 use the device's _PRW control method to learn which power resources need to be
198 enabled for the device to be able to generate wakeup signals.
199
200 1.4. Wakeup Signaling
201 ---------------------
202
203 Wakeup signals generated by PCI devices, either as native PCI PMEs, or as
204 a result of the execution of the _DSW (or _PSW) ACPI control method before
205 putting the device into a low-power state, have to be caught and handled as
206 appropriate. If they are sent while the system is in the working state
207 (ACPI S0), they should be translated into interrupts so that the kernel can
208 put the devices generating them into the full-power state and take care of the
209 events that triggered them. In turn, if they are sent while the system is
210 sleeping, they should cause the system's core logic to trigger wakeup.
211
212 On ACPI-based systems wakeup signals sent by conventional PCI devices are
213 converted into ACPI General-Purpose Events (GPEs) which are hardware signals
214 from the system core logic generated in response to various events that need to
215 be acted upon. Every GPE is associated with one or more sources of potentially
216 interesting events. In particular, a GPE may be associated with a PCI device
217 capable of signaling wakeup. The information on the connections between GPEs
218 and event sources is recorded in the system's ACPI BIOS from where it can be
219 read by the kernel.
220
221 If a PCI device known to the system's ACPI BIOS signals wakeup, the GPE
222 associated with it (if there is one) is triggered. The GPEs associated with PCI
223 bridges may also be triggered in response to a wakeup signal from one of the
224 devices below the bridge (this also is the case for root bridges) and, for
225 example, native PCI PMEs from devices unknown to the system's ACPI BIOS may be
226 handled this way.
227
228 A GPE may be triggered when the system is sleeping (i.e. when it is in one of
229 the ACPI S1-S4 states), in which case system wakeup is started by its core logic
230 (the device that was the source of the signal causing the system wakeup to occur
231 may be identified later). The GPEs used in such situations are referred to as
232 wakeup GPEs.
233
234 Usually, however, GPEs are also triggered when the system is in the working
235 state (ACPI S0) and in that case the system's core logic generates a System
236 Control Interrupt (SCI) to notify the kernel of the event. Then, the SCI
237 handler identifies the GPE that caused the interrupt to be generated which,
238 in turn, allows the kernel to identify the source of the event (that may be
239 a PCI device signaling wakeup). The GPEs used for notifying the kernel of
240 events occurring while the system is in the working state are referred to as
241 runtime GPEs.
242
243 Unfortunately, there is no standard way of handling wakeup signals sent by
244 conventional PCI devices on systems that are not ACPI-based, but there is one
245 for PCI Express devices. Namely, the PCI Express Base Specification introduced
246 a native mechanism for converting native PCI PMEs into interrupts generated by
247 root ports. For conventional PCI devices native PMEs are out-of-band, so they
248 are routed separately and they need not pass through bridges (in principle they
249 may be routed directly to the system's core logic), but for PCI Express devices
250 they are in-band messages that have to pass through the PCI Express hierarchy,
251 including the root port on the path from the device to the Root Complex. Thus
252 it was possible to introduce a mechanism by which a root port generates an
253 interrupt whenever it receives a PME message from one of the devices below it.
254 The PCI Express Requester ID of the device that sent the PME message is then
255 recorded in one of the root port's configuration registers from where it may be
256 read by the interrupt handler allowing the device to be identified. [PME
257 messages sent by PCI Express endpoints integrated with the Root Complex don't
258 pass through root ports, but instead they cause a Root Complex Event Collector
259 (if there is one) to generate interrupts.]
260
261 In principle the native PCI Express PME signaling may also be used on ACPI-based
262 systems along with the GPEs, but to use it the kernel has to ask the system's
263 ACPI BIOS to release control of root port configuration registers. The ACPI
264 BIOS, however, is not required to allow the kernel to control these registers
265 and if it doesn't do that, the kernel must not modify their contents. Of course
266 the native PCI Express PME signaling cannot be used by the kernel in that case.
267
268
269 2. PCI Subsystem and Device Power Management
270 ============================================
271
272 2.1. Device Power Management Callbacks
273 --------------------------------------
274
275 The PCI Subsystem participates in the power management of PCI devices in a
276 number of ways. First of all, it provides an intermediate code layer between
277 the device power management core (PM core) and PCI device drivers.
278 Specifically, the pm field of the PCI subsystem's struct bus_type object,
279 pci_bus_type, points to a struct dev_pm_ops object, pci_dev_pm_ops, containing
280 pointers to several device power management callbacks::
281
282 const struct dev_pm_ops pci_dev_pm_ops = {
283 .prepare = pci_pm_prepare,
284 .complete = pci_pm_complete,
285 .suspend = pci_pm_suspend,
286 .resume = pci_pm_resume,
287 .freeze = pci_pm_freeze,
288 .thaw = pci_pm_thaw,
289 .poweroff = pci_pm_poweroff,
290 .restore = pci_pm_restore,
291 .suspend_noirq = pci_pm_suspend_noirq,
292 .resume_noirq = pci_pm_resume_noirq,
293 .freeze_noirq = pci_pm_freeze_noirq,
294 .thaw_noirq = pci_pm_thaw_noirq,
295 .poweroff_noirq = pci_pm_poweroff_noirq,
296 .restore_noirq = pci_pm_restore_noirq,
297 .runtime_suspend = pci_pm_runtime_suspend,
298 .runtime_resume = pci_pm_runtime_resume,
299 .runtime_idle = pci_pm_runtime_idle,
300 };
301
302 These callbacks are executed by the PM core in various situations related to
303 device power management and they, in turn, execute power management callbacks
304 provided by PCI device drivers. They also perform power management operations
305 involving some standard configuration registers of PCI devices that device
306 drivers need not know or care about.
307
308 The structure representing a PCI device, struct pci_dev, contains several fields
309 that these callbacks operate on::
310
311 struct pci_dev {
312 ...
313 pci_power_t current_state; /* Current operating state. */
314 int pm_cap; /* PM capability offset in the
315 configuration space */
316 unsigned int pme_support:5; /* Bitmask of states from which PME#
317 can be generated */
318 unsigned int pme_poll:1; /* Poll device's PME status bit */
319 unsigned int d1_support:1; /* Low power state D1 is supported */
320 unsigned int d2_support:1; /* Low power state D2 is supported */
321 unsigned int no_d1d2:1; /* D1 and D2 are forbidden */
322 unsigned int wakeup_prepared:1; /* Device prepared for wake up */
323 unsigned int d3hot_delay; /* D3hot->D0 transition time in ms */
324 ...
325 };
326
327 They also indirectly use some fields of the struct device that is embedded in
328 struct pci_dev.
329
330 2.2. Device Initialization
331 --------------------------
332
333 The PCI subsystem's first task related to device power management is to
334 prepare the device for power management and initialize the fields of struct
335 pci_dev used for this purpose. This happens in two functions defined in
336 drivers/pci/, pci_pm_init() and pci_acpi_setup().
337
338 The first of these functions checks if the device supports native PCI PM
339 and if that's the case the offset of its power management capability structure
340 in the configuration space is stored in the pm_cap field of the device's struct
341 pci_dev object. Next, the function checks which PCI low-power states are
342 supported by the device and from which low-power states the device can generate
343 native PCI PMEs. The power management fields of the device's struct pci_dev and
344 the struct device embedded in it are updated accordingly and the generation of
345 PMEs by the device is disabled.
346
347 The second function checks if the device can be prepared to signal wakeup with
348 the help of the platform firmware, such as the ACPI BIOS. If that is the case,
349 the function updates the wakeup fields in struct device embedded in the
350 device's struct pci_dev and uses the firmware-provided method to prevent the
351 device from signaling wakeup.
352
353 At this point the device is ready for power management. For driverless devices,
354 however, this functionality is limited to a few basic operations carried out
355 during system-wide transitions to a sleep state and back to the working state.
356
357 2.3. Runtime Device Power Management
358 ------------------------------------
359
360 The PCI subsystem plays a vital role in the runtime power management of PCI
361 devices. For this purpose it uses the general runtime power management
362 (runtime PM) framework described in Documentation/power/runtime_pm.rst.
363 Namely, it provides subsystem-level callbacks::
364
365 pci_pm_runtime_suspend()
366 pci_pm_runtime_resume()
367 pci_pm_runtime_idle()
368
369 that are executed by the core runtime PM routines. It also implements the
370 entire mechanics necessary for handling runtime wakeup signals from PCI devices
371 in low-power states, which at the time of this writing works for both the native
372 PCI Express PME signaling and the ACPI GPE-based wakeup signaling described in
373 Section 1.
374
375 First, a PCI device is put into a low-power state, or suspended, with the help
376 of pm_schedule_suspend() or pm_runtime_suspend() which for PCI devices call
377 pci_pm_runtime_suspend() to do the actual job. For this to work, the device's
378 driver has to provide a pm->runtime_suspend() callback (see below), which is
379 run by pci_pm_runtime_suspend() as the first action. If the driver's callback
380 returns successfully, the device's standard configuration registers are saved,
381 the device is prepared to generate wakeup signals and, finally, it is put into
382 the target low-power state.
383
384 The low-power state to put the device into is the lowest-power (highest number)
385 state from which it can signal wakeup. The exact method of signaling wakeup is
386 system-dependent and is determined by the PCI subsystem on the basis of the
387 reported capabilities of the device and the platform firmware. To prepare the
388 device for signaling wakeup and put it into the selected low-power state, the
389 PCI subsystem can use the platform firmware as well as the device's native PCI
390 PM capabilities, if supported.
391
392 It is expected that the device driver's pm->runtime_suspend() callback will
393 not attempt to prepare the device for signaling wakeup or to put it into a
394 low-power state. The driver ought to leave these tasks to the PCI subsystem
395 that has all of the information necessary to perform them.
396
397 A suspended device is brought back into the "active" state, or resumed,
398 with the help of pm_request_resume() or pm_runtime_resume() which both call
399 pci_pm_runtime_resume() for PCI devices. Again, this only works if the device's
400 driver provides a pm->runtime_resume() callback (see below). However, before
401 the driver's callback is executed, pci_pm_runtime_resume() brings the device
402 back into the full-power state, prevents it from signaling wakeup while in that
403 state and restores its standard configuration registers. Thus the driver's
404 callback need not worry about the PCI-specific aspects of the device resume.
405
406 Note that generally pci_pm_runtime_resume() may be called in two different
407 situations. First, it may be called at the request of the device's driver, for
408 example if there are some data for it to process. Second, it may be called
409 as a result of a wakeup signal from the device itself (this sometimes is
410 referred to as "remote wakeup"). Of course, for this purpose the wakeup signal
411 is handled in one of the ways described in Section 1 and finally converted into
412 a notification for the PCI subsystem after the source device has been
413 identified.
414
415 The pci_pm_runtime_idle() function, called for PCI devices by pm_runtime_idle()
416 and pm_request_idle(), executes the device driver's pm->runtime_idle()
417 callback, if defined, and if that callback doesn't return error code (or is not
418 present at all), suspends the device with the help of pm_runtime_suspend().
419 Sometimes pci_pm_runtime_idle() is called automatically by the PM core (for
420 example, it is called right after the device has just been resumed), in which
421 cases it is expected to suspend the device if that makes sense. Usually,
422 however, the PCI subsystem doesn't really know if the device really can be
423 suspended, so it lets the device's driver decide by running its
424 pm->runtime_idle() callback.
425
426 2.4. System-Wide Power Transitions
427 ----------------------------------
428 There are a few different types of system-wide power transitions, described in
429 Documentation/driver-api/pm/devices.rst. Each of them requires devices to be
430 handled in a specific way and the PM core executes subsystem-level power
431 management callbacks for this purpose. They are executed in phases such that
432 each phase involves executing the same subsystem-level callback for every device
433 belonging to the given subsystem before the next phase begins. These phases
434 always run after tasks have been frozen.
435
436 2.4.1. System Suspend
437 ^^^^^^^^^^^^^^^^^^^^^
438
439 When the system is going into a sleep state in which the contents of memory will
440 be preserved, such as one of the ACPI sleep states S1-S3, the phases are:
441
442 prepare, suspend, suspend_noirq.
443
444 The following PCI bus type's callbacks, respectively, are used in these phases::
445
446 pci_pm_prepare()
447 pci_pm_suspend()
448 pci_pm_suspend_noirq()
449
450 The pci_pm_prepare() routine first puts the device into the "fully functional"
451 state with the help of pm_runtime_resume(). Then, it executes the device
452 driver's pm->prepare() callback if defined (i.e. if the driver's struct
453 dev_pm_ops object is present and the prepare pointer in that object is valid).
454
455 The pci_pm_suspend() routine first checks if the device's driver implements
456 legacy PCI suspend routines (see Section 3), in which case the driver's legacy
457 suspend callback is executed, if present, and its result is returned. Next, if
458 the device's driver doesn't provide a struct dev_pm_ops object (containing
459 pointers to the driver's callbacks), pci_pm_default_suspend() is called, which
460 simply turns off the device's bus master capability and runs
461 pcibios_disable_device() to disable it, unless the device is a bridge (PCI
462 bridges are ignored by this routine). Next, the device driver's pm->suspend()
463 callback is executed, if defined, and its result is returned if it fails.
464 Finally, pci_fixup_device() is called to apply hardware suspend quirks related
465 to the device if necessary.
466
467 Note that the suspend phase is carried out asynchronously for PCI devices, so
468 the pci_pm_suspend() callback may be executed in parallel for any pair of PCI
469 devices that don't depend on each other in a known way (i.e. none of the paths
470 in the device tree from the root bridge to a leaf device contains both of them).
471
472 The pci_pm_suspend_noirq() routine is executed after suspend_device_irqs() has
473 been called, which means that the device driver's interrupt handler won't be
474 invoked while this routine is running. It first checks if the device's driver
475 implements legacy PCI suspend routines (Section 3), in which case the legacy
476 late suspend routine is called and its result is returned (the standard
477 configuration registers of the device are saved if the driver's callback hasn't
478 done that). Second, if the device driver's struct dev_pm_ops object is not
479 present, the device's standard configuration registers are saved and the routine
480 returns success. Otherwise the device driver's pm->suspend_noirq() callback is
481 executed, if present, and its result is returned if it fails. Next, if the
482 device's standard configuration registers haven't been saved yet (one of the
483 device driver's callbacks executed before might do that), pci_pm_suspend_noirq()
484 saves them, prepares the device to signal wakeup (if necessary) and puts it into
485 a low-power state.
486
487 The low-power state to put the device into is the lowest-power (highest number)
488 state from which it can signal wakeup while the system is in the target sleep
489 state. Just like in the runtime PM case described above, the mechanism of
490 signaling wakeup is system-dependent and determined by the PCI subsystem, which
491 is also responsible for preparing the device to signal wakeup from the system's
492 target sleep state as appropriate.
493
494 PCI device drivers (that don't implement legacy power management callbacks) are
495 generally not expected to prepare devices for signaling wakeup or to put them
496 into low-power states. However, if one of the driver's suspend callbacks
497 (pm->suspend() or pm->suspend_noirq()) saves the device's standard configuration
498 registers, pci_pm_suspend_noirq() will assume that the device has been prepared
499 to signal wakeup and put into a low-power state by the driver (the driver is
500 then assumed to have used the helper functions provided by the PCI subsystem for
501 this purpose). PCI device drivers are not encouraged to do that, but in some
502 rare cases doing that in the driver may be the optimum approach.
503
504 2.4.2. System Resume
505 ^^^^^^^^^^^^^^^^^^^^
506
507 When the system is undergoing a transition from a sleep state in which the
508 contents of memory have been preserved, such as one of the ACPI sleep states
509 S1-S3, into the working state (ACPI S0), the phases are:
510
511 resume_noirq, resume, complete.
512
513 The following PCI bus type's callbacks, respectively, are executed in these
514 phases::
515
516 pci_pm_resume_noirq()
517 pci_pm_resume()
518 pci_pm_complete()
519
520 The pci_pm_resume_noirq() routine first puts the device into the full-power
521 state, restores its standard configuration registers and applies early resume
522 hardware quirks related to the device, if necessary. This is done
523 unconditionally, regardless of whether or not the device's driver implements
524 legacy PCI power management callbacks (this way all PCI devices are in the
525 full-power state and their standard configuration registers have been restored
526 when their interrupt handlers are invoked for the first time during resume,
527 which allows the kernel to avoid problems with the handling of shared interrupts
528 by drivers whose devices are still suspended). If legacy PCI power management
529 callbacks (see Section 3) are implemented by the device's driver, the legacy
530 early resume callback is executed and its result is returned. Otherwise, the
531 device driver's pm->resume_noirq() callback is executed, if defined, and its
532 result is returned.
533
534 The pci_pm_resume() routine first checks if the device's standard configuration
535 registers have been restored and restores them if that's not the case (this
536 only is necessary in the error path during a failing suspend). Next, resume
537 hardware quirks related to the device are applied, if necessary, and if the
538 device's driver implements legacy PCI power management callbacks (see
539 Section 3), the driver's legacy resume callback is executed and its result is
540 returned. Otherwise, the device's wakeup signaling mechanisms are blocked and
541 its driver's pm->resume() callback is executed, if defined (the callback's
542 result is then returned).
543
544 The resume phase is carried out asynchronously for PCI devices, like the
545 suspend phase described above, which means that if two PCI devices don't depend
546 on each other in a known way, the pci_pm_resume() routine may be executed for
547 both of them in parallel.
548
549 The pci_pm_complete() routine only executes the device driver's pm->complete()
550 callback, if defined.
551
552 2.4.3. System Hibernation
553 ^^^^^^^^^^^^^^^^^^^^^^^^^
554
555 System hibernation is more complicated than system suspend, because it requires
556 a system image to be created and written into a persistent storage medium. The
557 image is created atomically and all devices are quiesced, or frozen, before that
558 happens.
559
560 The freezing of devices is carried out after enough memory has been freed (at
561 the time of this writing the image creation requires at least 50% of system RAM
562 to be free) in the following three phases:
563
564 prepare, freeze, freeze_noirq
565
566 that correspond to the PCI bus type's callbacks::
567
568 pci_pm_prepare()
569 pci_pm_freeze()
570 pci_pm_freeze_noirq()
571
572 This means that the prepare phase is exactly the same as for system suspend.
573 The other two phases, however, are different.
574
575 The pci_pm_freeze() routine is quite similar to pci_pm_suspend(), but it runs
576 the device driver's pm->freeze() callback, if defined, instead of pm->suspend(),
577 and it doesn't apply the suspend-related hardware quirks. It is executed
578 asynchronously for different PCI devices that don't depend on each other in a
579 known way.
580
581 The pci_pm_freeze_noirq() routine, in turn, is similar to
582 pci_pm_suspend_noirq(), but it calls the device driver's pm->freeze_noirq()
583 routine instead of pm->suspend_noirq(). It also doesn't attempt to prepare the
584 device for signaling wakeup and put it into a low-power state. Still, it saves
585 the device's standard configuration registers if they haven't been saved by one
586 of the driver's callbacks.
587
588 Once the image has been created, it has to be saved. However, at this point all
589 devices are frozen and they cannot handle I/O, while their ability to handle
590 I/O is obviously necessary for the image saving. Thus they have to be brought
591 back to the fully functional state and this is done in the following phases:
592
593 thaw_noirq, thaw, complete
594
595 using the following PCI bus type's callbacks::
596
597 pci_pm_thaw_noirq()
598 pci_pm_thaw()
599 pci_pm_complete()
600
601 respectively.
602
603 The first of them, pci_pm_thaw_noirq(), is analogous to pci_pm_resume_noirq().
604 It puts the device into the full power state and restores its standard
605 configuration registers. It also executes the device driver's pm->thaw_noirq()
606 callback, if defined, instead of pm->resume_noirq().
607
608 The pci_pm_thaw() routine is similar to pci_pm_resume(), but it runs the device
609 driver's pm->thaw() callback instead of pm->resume(). It is executed
610 asynchronously for different PCI devices that don't depend on each other in a
611 known way.
612
613 The complete phase is the same as for system resume.
614
615 After saving the image, devices need to be powered down before the system can
616 enter the target sleep state (ACPI S4 for ACPI-based systems). This is done in
617 three phases:
618
619 prepare, poweroff, poweroff_noirq
620
621 where the prepare phase is exactly the same as for system suspend. The other
622 two phases are analogous to the suspend and suspend_noirq phases, respectively.
623 The PCI subsystem-level callbacks they correspond to::
624
625 pci_pm_poweroff()
626 pci_pm_poweroff_noirq()
627
628 work in analogy with pci_pm_suspend() and pci_pm_suspend_noirq(), respectively,
629 although they don't attempt to save the device's standard configuration
630 registers.
631
632 2.4.4. System Restore
633 ^^^^^^^^^^^^^^^^^^^^^
634
635 System restore requires a hibernation image to be loaded into memory and the
636 pre-hibernation memory contents to be restored before the pre-hibernation system
637 activity can be resumed.
638
639 As described in Documentation/driver-api/pm/devices.rst, the hibernation image
640 is loaded into memory by a fresh instance of the kernel, called the boot kernel,
641 which in turn is loaded and run by a boot loader in the usual way. After the
642 boot kernel has loaded the image, it needs to replace its own code and data with
643 the code and data of the "hibernated" kernel stored within the image, called the
644 image kernel. For this purpose all devices are frozen just like before creating
645 the image during hibernation, in the
646
647 prepare, freeze, freeze_noirq
648
649 phases described above. However, the devices affected by these phases are only
650 those having drivers in the boot kernel; other devices will still be in whatever
651 state the boot loader left them.
652
653 Should the restoration of the pre-hibernation memory contents fail, the boot
654 kernel would go through the "thawing" procedure described above, using the
655 thaw_noirq, thaw, and complete phases (that will only affect the devices having
656 drivers in the boot kernel), and then continue running normally.
657
658 If the pre-hibernation memory contents are restored successfully, which is the
659 usual situation, control is passed to the image kernel, which then becomes
660 responsible for bringing the system back to the working state. To achieve this,
661 it must restore the devices' pre-hibernation functionality, which is done much
662 like waking up from the memory sleep state, although it involves different
663 phases:
664
665 restore_noirq, restore, complete
666
667 The first two of these are analogous to the resume_noirq and resume phases
668 described above, respectively, and correspond to the following PCI subsystem
669 callbacks::
670
671 pci_pm_restore_noirq()
672 pci_pm_restore()
673
674 These callbacks work in analogy with pci_pm_resume_noirq() and pci_pm_resume(),
675 respectively, but they execute the device driver's pm->restore_noirq() and
676 pm->restore() callbacks, if available.
677
678 The complete phase is carried out in exactly the same way as during system
679 resume.
680
681
682 3. PCI Device Drivers and Power Management
683 ==========================================
684
685 3.1. Power Management Callbacks
686 -------------------------------
687
688 PCI device drivers participate in power management by providing callbacks to be
689 executed by the PCI subsystem's power management routines described above and by
690 controlling the runtime power management of their devices.
691
692 At the time of this writing there are two ways to define power management
693 callbacks for a PCI device driver, the recommended one, based on using a
694 dev_pm_ops structure described in Documentation/driver-api/pm/devices.rst, and
695 the "legacy" one, in which the .suspend() and .resume() callbacks from struct
696 pci_driver are used. The legacy approach, however, doesn't allow one to define
697 runtime power management callbacks and is not really suitable for any new
698 drivers. Therefore it is not covered by this document (refer to the source code
699 to learn more about it).
700
701 It is recommended that all PCI device drivers define a struct dev_pm_ops object
702 containing pointers to power management (PM) callbacks that will be executed by
703 the PCI subsystem's PM routines in various circumstances. A pointer to the
704 driver's struct dev_pm_ops object has to be assigned to the driver.pm field in
705 its struct pci_driver object. Once that has happened, the "legacy" PM callbacks
706 in struct pci_driver are ignored (even if they are not NULL).
707
708 The PM callbacks in struct dev_pm_ops are not mandatory and if they are not
709 defined (i.e. the respective fields of struct dev_pm_ops are unset) the PCI
710 subsystem will handle the device in a simplified default manner. If they are
711 defined, though, they are expected to behave as described in the following
712 subsections.
713
714 3.1.1. prepare()
715 ^^^^^^^^^^^^^^^^
716
717 The prepare() callback is executed during system suspend, during hibernation
718 (when a hibernation image is about to be created), during power-off after
719 saving a hibernation image and during system restore, when a hibernation image
720 has just been loaded into memory.
721
722 This callback is only necessary if the driver's device has children that in
723 general may be registered at any time. In that case the role of the prepare()
724 callback is to prevent new children of the device from being registered until
725 one of the resume_noirq(), thaw_noirq(), or restore_noirq() callbacks is run.
726
727 In addition to that the prepare() callback may carry out some operations
728 preparing the device to be suspended, although it should not allocate memory
729 (if additional memory is required to suspend the device, it has to be
730 preallocated earlier, for example in a suspend/hibernate notifier as described
731 in Documentation/driver-api/pm/notifiers.rst).
732
733 3.1.2. suspend()
734 ^^^^^^^^^^^^^^^^
735
736 The suspend() callback is only executed during system suspend, after prepare()
737 callbacks have been executed for all devices in the system.
738
739 This callback is expected to quiesce the device and prepare it to be put into a
740 low-power state by the PCI subsystem. It is not required (in fact it even is
741 not recommended) that a PCI driver's suspend() callback save the standard
742 configuration registers of the device, prepare it for waking up the system, or
743 put it into a low-power state. All of these operations can very well be taken
744 care of by the PCI subsystem, without the driver's participation.
745
746 However, in some rare case it is convenient to carry out these operations in
747 a PCI driver. Then, pci_save_state(), pci_prepare_to_sleep(), and
748 pci_set_power_state() should be used to save the device's standard configuration
749 registers, to prepare it for system wakeup (if necessary), and to put it into a
750 low-power state, respectively. Moreover, if the driver calls pci_save_state(),
751 the PCI subsystem will not execute either pci_prepare_to_sleep(), or
752 pci_set_power_state() for its device, so the driver is then responsible for
753 handling the device as appropriate.
754
755 While the suspend() callback is being executed, the driver's interrupt handler
756 can be invoked to handle an interrupt from the device, so all suspend-related
757 operations relying on the driver's ability to handle interrupts should be
758 carried out in this callback.
759
760 3.1.3. suspend_noirq()
761 ^^^^^^^^^^^^^^^^^^^^^^
762
763 The suspend_noirq() callback is only executed during system suspend, after
764 suspend() callbacks have been executed for all devices in the system and
765 after device interrupts have been disabled by the PM core.
766
767 The difference between suspend_noirq() and suspend() is that the driver's
768 interrupt handler will not be invoked while suspend_noirq() is running. Thus
769 suspend_noirq() can carry out operations that would cause race conditions to
770 arise if they were performed in suspend().
771
772 3.1.4. freeze()
773 ^^^^^^^^^^^^^^^
774
775 The freeze() callback is hibernation-specific and is executed in two situations,
776 during hibernation, after prepare() callbacks have been executed for all devices
777 in preparation for the creation of a system image, and during restore,
778 after a system image has been loaded into memory from persistent storage and the
779 prepare() callbacks have been executed for all devices.
780
781 The role of this callback is analogous to the role of the suspend() callback
782 described above. In fact, they only need to be different in the rare cases when
783 the driver takes the responsibility for putting the device into a low-power
784 state.
785
786 In that cases the freeze() callback should not prepare the device system wakeup
787 or put it into a low-power state. Still, either it or freeze_noirq() should
788 save the device's standard configuration registers using pci_save_state().
789
790 3.1.5. freeze_noirq()
791 ^^^^^^^^^^^^^^^^^^^^^
792
793 The freeze_noirq() callback is hibernation-specific. It is executed during
794 hibernation, after prepare() and freeze() callbacks have been executed for all
795 devices in preparation for the creation of a system image, and during restore,
796 after a system image has been loaded into memory and after prepare() and
797 freeze() callbacks have been executed for all devices. It is always executed
798 after device interrupts have been disabled by the PM core.
799
800 The role of this callback is analogous to the role of the suspend_noirq()
801 callback described above and it very rarely is necessary to define
802 freeze_noirq().
803
804 The difference between freeze_noirq() and freeze() is analogous to the
805 difference between suspend_noirq() and suspend().
806
807 3.1.6. poweroff()
808 ^^^^^^^^^^^^^^^^^
809
810 The poweroff() callback is hibernation-specific. It is executed when the system
811 is about to be powered off after saving a hibernation image to a persistent
812 storage. prepare() callbacks are executed for all devices before poweroff() is
813 called.
814
815 The role of this callback is analogous to the role of the suspend() and freeze()
816 callbacks described above, although it does not need to save the contents of
817 the device's registers. In particular, if the driver wants to put the device
818 into a low-power state itself instead of allowing the PCI subsystem to do that,
819 the poweroff() callback should use pci_prepare_to_sleep() and
820 pci_set_power_state() to prepare the device for system wakeup and to put it
821 into a low-power state, respectively, but it need not save the device's standard
822 configuration registers.
823
824 3.1.7. poweroff_noirq()
825 ^^^^^^^^^^^^^^^^^^^^^^^
826
827 The poweroff_noirq() callback is hibernation-specific. It is executed after
828 poweroff() callbacks have been executed for all devices in the system.
829
830 The role of this callback is analogous to the role of the suspend_noirq() and
831 freeze_noirq() callbacks described above, but it does not need to save the
832 contents of the device's registers.
833
834 The difference between poweroff_noirq() and poweroff() is analogous to the
835 difference between suspend_noirq() and suspend().
836
837 3.1.8. resume_noirq()
838 ^^^^^^^^^^^^^^^^^^^^^
839
840 The resume_noirq() callback is only executed during system resume, after the
841 PM core has enabled the non-boot CPUs. The driver's interrupt handler will not
842 be invoked while resume_noirq() is running, so this callback can carry out
843 operations that might race with the interrupt handler.
844
845 Since the PCI subsystem unconditionally puts all devices into the full power
846 state in the resume_noirq phase of system resume and restores their standard
847 configuration registers, resume_noirq() is usually not necessary. In general
848 it should only be used for performing operations that would lead to race
849 conditions if carried out by resume().
850
851 3.1.9. resume()
852 ^^^^^^^^^^^^^^^
853
854 The resume() callback is only executed during system resume, after
855 resume_noirq() callbacks have been executed for all devices in the system and
856 device interrupts have been enabled by the PM core.
857
858 This callback is responsible for restoring the pre-suspend configuration of the
859 device and bringing it back to the fully functional state. The device should be
860 able to process I/O in a usual way after resume() has returned.
861
862 3.1.10. thaw_noirq()
863 ^^^^^^^^^^^^^^^^^^^^
864
865 The thaw_noirq() callback is hibernation-specific. It is executed after a
866 system image has been created and the non-boot CPUs have been enabled by the PM
867 core, in the thaw_noirq phase of hibernation. It also may be executed if the
868 loading of a hibernation image fails during system restore (it is then executed
869 after enabling the non-boot CPUs). The driver's interrupt handler will not be
870 invoked while thaw_noirq() is running.
871
872 The role of this callback is analogous to the role of resume_noirq(). The
873 difference between these two callbacks is that thaw_noirq() is executed after
874 freeze() and freeze_noirq(), so in general it does not need to modify the
875 contents of the device's registers.
876
877 3.1.11. thaw()
878 ^^^^^^^^^^^^^^
879
880 The thaw() callback is hibernation-specific. It is executed after thaw_noirq()
881 callbacks have been executed for all devices in the system and after device
882 interrupts have been enabled by the PM core.
883
884 This callback is responsible for restoring the pre-freeze configuration of
885 the device, so that it will work in a usual way after thaw() has returned.
886
887 3.1.12. restore_noirq()
888 ^^^^^^^^^^^^^^^^^^^^^^^
889
890 The restore_noirq() callback is hibernation-specific. It is executed in the
891 restore_noirq phase of hibernation, when the boot kernel has passed control to
892 the image kernel and the non-boot CPUs have been enabled by the image kernel's
893 PM core.
894
895 This callback is analogous to resume_noirq() with the exception that it cannot
896 make any assumption on the previous state of the device, even if the BIOS (or
897 generally the platform firmware) is known to preserve that state over a
898 suspend-resume cycle.
899
900 For the vast majority of PCI device drivers there is no difference between
901 resume_noirq() and restore_noirq().
902
903 3.1.13. restore()
904 ^^^^^^^^^^^^^^^^^
905
906 The restore() callback is hibernation-specific. It is executed after
907 restore_noirq() callbacks have been executed for all devices in the system and
908 after the PM core has enabled device drivers' interrupt handlers to be invoked.
909
910 This callback is analogous to resume(), just like restore_noirq() is analogous
911 to resume_noirq(). Consequently, the difference between restore_noirq() and
912 restore() is analogous to the difference between resume_noirq() and resume().
913
914 For the vast majority of PCI device drivers there is no difference between
915 resume() and restore().
916
917 3.1.14. complete()
918 ^^^^^^^^^^^^^^^^^^
919
920 The complete() callback is executed in the following situations:
921
922 - during system resume, after resume() callbacks have been executed for all
923 devices,
924 - during hibernation, before saving the system image, after thaw() callbacks
925 have been executed for all devices,
926 - during system restore, when the system is going back to its pre-hibernation
927 state, after restore() callbacks have been executed for all devices.
928
929 It also may be executed if the loading of a hibernation image into memory fails
930 (in that case it is run after thaw() callbacks have been executed for all
931 devices that have drivers in the boot kernel).
932
933 This callback is entirely optional, although it may be necessary if the
934 prepare() callback performs operations that need to be reversed.
935
936 3.1.15. runtime_suspend()
937 ^^^^^^^^^^^^^^^^^^^^^^^^^
938
939 The runtime_suspend() callback is specific to device runtime power management
940 (runtime PM). It is executed by the PM core's runtime PM framework when the
941 device is about to be suspended (i.e. quiesced and put into a low-power state)
942 at run time.
943
944 This callback is responsible for freezing the device and preparing it to be
945 put into a low-power state, but it must allow the PCI subsystem to perform all
946 of the PCI-specific actions necessary for suspending the device.
947
948 3.1.16. runtime_resume()
949 ^^^^^^^^^^^^^^^^^^^^^^^^
950
951 The runtime_resume() callback is specific to device runtime PM. It is executed
952 by the PM core's runtime PM framework when the device is about to be resumed
953 (i.e. put into the full-power state and programmed to process I/O normally) at
954 run time.
955
956 This callback is responsible for restoring the normal functionality of the
957 device after it has been put into the full-power state by the PCI subsystem.
958 The device is expected to be able to process I/O in the usual way after
959 runtime_resume() has returned.
960
961 3.1.17. runtime_idle()
962 ^^^^^^^^^^^^^^^^^^^^^^
963
964 The runtime_idle() callback is specific to device runtime PM. It is executed
965 by the PM core's runtime PM framework whenever it may be desirable to suspend
966 the device according to the PM core's information. In particular, it is
967 automatically executed right after runtime_resume() has returned in case the
968 resume of the device has happened as a result of a spurious event.
969
970 This callback is optional, but if it is not implemented or if it returns 0, the
971 PCI subsystem will call pm_runtime_suspend() for the device, which in turn will
972 cause the driver's runtime_suspend() callback to be executed.
973
974 3.1.18. Pointing Multiple Callback Pointers to One Routine
975 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
976
977 Although in principle each of the callbacks described in the previous
978 subsections can be defined as a separate function, it often is convenient to
979 point two or more members of struct dev_pm_ops to the same routine. There are
980 a few convenience macros that can be used for this purpose.
981
982 The DEFINE_SIMPLE_DEV_PM_OPS() declares a struct dev_pm_ops object with one
983 suspend routine pointed to by the .suspend(), .freeze(), and .poweroff()
984 members and one resume routine pointed to by the .resume(), .thaw(), and
985 .restore() members. The other function pointers in this struct dev_pm_ops are
986 unset.
987
988 The DEFINE_RUNTIME_DEV_PM_OPS() is similar to DEFINE_SIMPLE_DEV_PM_OPS(), but it
989 additionally sets the .runtime_resume() pointer to pm_runtime_force_resume()
990 and the .runtime_suspend() pointer to pm_runtime_force_suspend().
991
992 The SYSTEM_SLEEP_PM_OPS() can be used inside of a declaration of struct
993 dev_pm_ops to indicate that one suspend routine is to be pointed to by the
994 .suspend(), .freeze(), and .poweroff() members and one resume routine is to
995 be pointed to by the .resume(), .thaw(), and .restore() members.
996
997 3.1.19. Driver Flags for Power Management
998 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
999
1000 The PM core allows device drivers to set flags that influence the handling of
1001 power management for the devices by the core itself and by middle layer code
1002 including the PCI bus type. The flags should be set once at the driver probe
1003 time with the help of the dev_pm_set_driver_flags() function and they should not
1004 be updated directly afterwards.
1006 The DPM_FLAG_NO_DIRECT_COMPLETE flag prevents the PM core from using the
1007 direct-complete mechanism allowing device suspend/resume callbacks to be skipped
1008 if the device is in runtime suspend when the system suspend starts. That also
1009 affects all of the ancestors of the device, so this flag should only be used if
1010 absolutely necessary.
1012 The DPM_FLAG_SMART_PREPARE flag causes the PCI bus type to return a positive
1013 value from pci_pm_prepare() only if the ->prepare callback provided by the
1014 driver of the device returns a positive value. That allows the driver to opt
1015 out from using the direct-complete mechanism dynamically (whereas setting
1016 DPM_FLAG_NO_DIRECT_COMPLETE means permanent opt-out).
1018 The DPM_FLAG_SMART_SUSPEND flag tells the PCI bus type that from the driver's
1019 perspective the device can be safely left in runtime suspend during system
1020 suspend. That causes pci_pm_suspend(), pci_pm_freeze() and pci_pm_poweroff()
1021 to avoid resuming the device from runtime suspend unless there are PCI-specific
1022 reasons for doing that. Also, it causes pci_pm_suspend_late/noirq() and
1023 pci_pm_poweroff_late/noirq() to return early if the device remains in runtime
1024 suspend during the "late" phase of the system-wide transition under way.
1025 Moreover, if the device is in runtime suspend in pci_pm_resume_noirq() or
1026 pci_pm_restore_noirq(), its runtime PM status will be changed to "active" (as it
1027 is going to be put into D0 going forward).
1029 Setting the DPM_FLAG_MAY_SKIP_RESUME flag means that the driver allows its
1030 "noirq" and "early" resume callbacks to be skipped if the device can be left
1031 in suspend after a system-wide transition into the working state. This flag is
1032 taken into consideration by the PM core along with the power.may_skip_resume
1033 status bit of the device which is set by pci_pm_suspend_noirq() in certain
1034 situations. If the PM core determines that the driver's "noirq" and "early"
1035 resume callbacks should be skipped, the dev_pm_skip_resume() helper function
1036 will return "true" and that will cause pci_pm_resume_noirq() and
1037 pci_pm_resume_early() to return upfront without touching the device and
1038 executing the driver callbacks.
1040 3.2. Device Runtime Power Management
1041 ------------------------------------
1043 In addition to providing device power management callbacks PCI device drivers
1044 are responsible for controlling the runtime power management (runtime PM) of
1045 their devices.
1047 The PCI device runtime PM is optional, but it is recommended that PCI device
1048 drivers implement it at least in the cases where there is a reliable way of
1049 verifying that the device is not used (like when the network cable is detached
1050 from an Ethernet adapter or there are no devices attached to a USB controller).
1052 To support the PCI runtime PM the driver first needs to implement the
1053 runtime_suspend() and runtime_resume() callbacks. It also may need to implement
1054 the runtime_idle() callback to prevent the device from being suspended again
1055 every time right after the runtime_resume() callback has returned
1056 (alternatively, the runtime_suspend() callback will have to check if the
1057 device should really be suspended and return -EAGAIN if that is not the case).
1059 The runtime PM of PCI devices is enabled by default by the PCI core. PCI
1060 device drivers do not need to enable it and should not attempt to do so.
1061 However, it is blocked by pci_pm_init() that runs the pm_runtime_forbid()
1062 helper function. In addition to that, the runtime PM usage counter of
1063 each PCI device is incremented by local_pci_probe() before executing the
1064 probe callback provided by the device's driver.
1066 If a PCI driver implements the runtime PM callbacks and intends to use the
1067 runtime PM framework provided by the PM core and the PCI subsystem, it needs
1068 to decrement the device's runtime PM usage counter in its probe callback
1069 function. If it doesn't do that, the counter will always be different from
1070 zero for the device and it will never be runtime-suspended. The simplest
1071 way to do that is by calling pm_runtime_put_noidle(), but if the driver
1072 wants to schedule an autosuspend right away, for example, it may call
1073 pm_runtime_put_autosuspend() instead for this purpose. Generally, it
1074 just needs to call a function that decrements the devices usage counter
1075 from its probe routine to make runtime PM work for the device.
1077 It is important to remember that the driver's runtime_suspend() callback
1078 may be executed right after the usage counter has been decremented, because
1079 user space may already have caused the pm_runtime_allow() helper function
1080 unblocking the runtime PM of the device to run via sysfs, so the driver must
1081 be prepared to cope with that.
1083 The driver itself should not call pm_runtime_allow(), though. Instead, it
1084 should let user space or some platform-specific code do that (user space can
1085 do it via sysfs as stated above), but it must be prepared to handle the
1086 runtime PM of the device correctly as soon as pm_runtime_allow() is called
1087 (which may happen at any time, even before the driver is loaded).
1089 When the driver's remove callback runs, it has to balance the decrementation
1090 of the device's runtime PM usage counter at the probe time. For this reason,
1091 if it has decremented the counter in its probe callback, it must run
1092 pm_runtime_get_noresume() in its remove callback. [Since the core carries
1093 out a runtime resume of the device and bumps up the device's usage counter
1094 before running the driver's remove callback, the runtime PM of the device
1095 is effectively disabled for the duration of the remove execution and all
1096 runtime PM helper functions incrementing the device's usage counter are
1097 then effectively equivalent to pm_runtime_get_noresume().]
1099 The runtime PM framework works by processing requests to suspend or resume
1100 devices, or to check if they are idle (in which cases it is reasonable to
1101 subsequently request that they be suspended). These requests are represented
1102 by work items put into the power management workqueue, pm_wq. Although there
1103 are a few situations in which power management requests are automatically
1104 queued by the PM core (for example, after processing a request to resume a
1105 device the PM core automatically queues a request to check if the device is
1106 idle), device drivers are generally responsible for queuing power management
1107 requests for their devices. For this purpose they should use the runtime PM
1108 helper functions provided by the PM core, discussed in
1109 Documentation/power/runtime_pm.rst.
1111 Devices can also be suspended and resumed synchronously, without placing a
1112 request into pm_wq. In the majority of cases this also is done by their
1113 drivers that use helper functions provided by the PM core for this purpose.
1115 For more information on the runtime PM of devices refer to
1116 Documentation/power/runtime_pm.rst.
1119 4. Resources
1120 ============
1122 PCI Local Bus Specification, Rev. 3.0
1124 PCI Bus Power Management Interface Specification, Rev. 1.2
1126 Advanced Configuration and Power Interface (ACPI) Specification, Rev. 3.0b
1128 PCI Express Base Specification, Rev. 2.0
1130 Documentation/driver-api/pm/devices.rst
1132 Documentation/power/runtime_pm.rst

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Native PCI PM과 platform PM

1-70

이 문서는 PCI device에 특화된 power management 개념과 Linux kernel interface를 설명합니다. 일반 device PM은 `Documentation/driver-api/pm/devices.rst`, runtime PM은 `Documentation/power/runtime_pm.rst`를 참조합니다.

Power management는 사용량이 적거나 inactive인 device를 기능·성능을 줄인 low-power state에 넣어 energy를 절약하고, data 처리나 외부 event가 필요할 때 full-power state로 되돌리는 기능입니다.

PCI device의 low-power 전환은 PCI Bus Power Management Interface Specification의 표준 capability를 쓰는 native PCI PM과 ACPI BIOS 같은 platform firmware method를 쓰는 방식이 있습니다. Native 방식은 standard configuration register에 값을 쓰고, platform 방식은 firmware가 제공한 system-specific method를 kernel이 호출합니다.

Native PCI PM device는 외부 event를 알리는 Power Management Event(PME)를 생성할 수 있습니다. Kernel은 PME source를 full-power로 복귀시켜야 하지만 PCI PM specification은 PME를 CPU와 OS kernel에 전달하는 표준 방식을 정의하지 않습니다. 따라서 firmware가 interrupt 등의 전달 경로도 준비해야 할 수 있습니다.

반대로 firmware method로 device state를 바꾸는 경우에도 platform wakeup method가 native PME 설정에 의존할 수 있어 device 자체의 PCI PM mechanism을 함께 준비해야 합니다. 원하는 결과를 얻으려면 두 mechanism을 동시에 사용하는 경우가 많습니다.

PCI power control 방식
방식State 전환Wakeup 준비
Native PCI PMStandard configuration registerDevice PME + platform delivery 필요
Platform-based PMACPI BIOS 등 firmware methodPlatform method가 native PME에 의존 가능

State 전환 주체와 wake signal 준비가 서로 보완적입니다.

Low-power 왕복
device underusedlow-power statedata / external event / PMEfull-power state

Inactive일 때 낮추고 work 또는 wake event가 생기면 D0로 복귀합니다.

====================
PCI Power Management
====================

Copyright (c) 2010 Rafael J. Wysocki <[email protected]>, Novell Inc.

An overview of concepts and the Linux kernel's interfaces related to PCI power
management.  Based on previous work by Patrick Mochel <[email protected]>
(and others).

This document only covers the aspects of power management specific to PCI
devices.  For general description of the kernel's interfaces related to device
power management refer to Documentation/driver-api/pm/devices.rst and
Documentation/power/runtime_pm.rst.

.. contents:

   1. Hardware and Platform Support for PCI Power Management
   2. PCI Subsystem and Device Power Management
   3. PCI Device Drivers and Power Management
   4. Resources


1. Hardware and Platform Support for PCI Power Management
=========================================================

1.1. Native and Platform-Based Power Management
-----------------------------------------------

In general, power management is a feature allowing one to save energy by putting
devices into states in which they draw less power (low-power states) at the
price of reduced functionality or performance.

Usually, a device is put into a low-power state when it is underutilized or
completely inactive.  However, when it is necessary to use the device once
again, it has to be put back into the "fully functional" state (full-power
state).  This may happen when there are some data for the device to handle or
as a result of an external event requiring the device to be active, which may
be signaled by the device itself.

PCI devices may be put into low-power states in two ways, by using the device
capabilities introduced by the PCI Bus Power Management Interface Specification,
or with the help of platform firmware, such as an ACPI BIOS.  In the first
approach, that is referred to as the native PCI power management (native PCI PM)
in what follows, the device power state is changed as a result of writing a
specific value into one of its standard configuration registers.  The second
approach requires the platform firmware to provide special methods that may be
used by the kernel to change the device's power state.

Devices supporting the native PCI PM usually can generate wakeup signals called
Power Management Events (PMEs) to let the kernel know about external events
requiring the device to be active.  After receiving a PME the kernel is supposed
to put the device that sent it into the full-power state.  However, the PCI Bus
Power Management Interface Specification doesn't define any standard method of
delivering the PME from the device to the CPU and the operating system kernel.
It is assumed that the platform firmware will perform this task and therefore,
even though a PCI device is set up to generate PMEs, it also may be necessary to
prepare the platform firmware for notifying the CPU of the PMEs coming from the
device (e.g. by generating interrupts).

In turn, if the methods provided by the platform firmware are used for changing
the power state of a device, usually the platform also provides a method for
preparing the device to generate wakeup signals.  In that case, however, it
often also is necessary to prepare the device for generating PMEs using the
native PCI PM mechanism, because the method provided by the platform depends on
that.

Thus in many situations both the native and the platform-based power management
mechanisms have to be used simultaneously to obtain the desired result.

Native PCI power state와 transition

71-138

PCI PM Spec은 PCI 2.1과 2.2 사이에 도입된 표준 PM interface입니다. Conventional PCI에서는 구현이 선택 사항이지만 PCI Express에서는 필수입니다. 지원 device는 configuration space에 8-byte power management capability field를 가집니다.

Device state D0~D3와 bus state B0~B3에서 숫자가 클수록 소비 power는 낮고 full-power D0 또는 B0로 돌아오는 latency는 길어집니다. Linux kernel은 이 문서 작성 시점에 PCI bus power management를 지원하지 않으므로 device state만 다룹니다.

D3hot은 software가 register로 진입시킬 수 있는 D3이고, D3cold는 Vcc가 제거된 상태입니다. Device를 D3cold로 직접 programming할 수는 없지만 bus 전체의 Vcc를 제거하는 interface가 있을 수 있습니다.

모든 PCI device는 PCI PM Spec 지원 여부와 무관하게 D0와 D3cold에 있을 수 있습니다. Spec을 구현하면 D0와 D3hot은 필수이고 D1·D2는 선택 사항입니다.

D1~D3hot에서도 standard configuration register는 접근 가능해야 하지만 I/O와 memory space는 disable됩니다. 따라서 kernel이 programming으로 D0로 복귀시킬 수 있습니다.

허용되는 native transition
CurrentNew
D0D1, D2, D3
D1D2, D3
D2D3
D1, D2, D3D0

Low-power state 사이에는 더 낮은 power 방향으로만 이동하고, 복귀는 D0로 합니다.

D3cold→D0는 Vcc가 복원될 때 full power-on reset sequence로 일어나며 hardware가 초기 power-on default를 복원합니다.

PCI PM device는 D0~D3에서 PME를 생성하도록 programming할 수 있지만 모든 지원 state에서 PME 생성 능력이 필수인 것은 아닙니다. D3cold PME는 선택 사항이며 device가 신호를 만들 만큼 활성 상태를 유지하는 3.3Vaux가 필요합니다.

D-state 성격
State전원·접근복귀
D0Full power이미 active
D1/D2Optional low power, config access 가능Programming으로 D0
D3hotSoftware-accessible D3, config access 가능Programming으로 D0
D3coldVcc 제거, optional 3.3Vaux PMEVcc 복원 + full reset

D3hot과 D3cold의 software 접근성과 전원 조건이 다릅니다.

1.2. Native PCI Power Management
--------------------------------

The PCI Bus Power Management Interface Specification (PCI PM Spec) was
introduced between the PCI 2.1 and PCI 2.2 Specifications.  It defined a
standard interface for performing various operations related to power
management.

The implementation of the PCI PM Spec is optional for conventional PCI devices,
but it is mandatory for PCI Express devices.  If a device supports the PCI PM
Spec, it has an 8 byte power management capability field in its PCI
configuration space.  This field is used to describe and control the standard
features related to the native PCI power management.

The PCI PM Spec defines 4 operating states for devices (D0-D3) and for buses
(B0-B3).  The higher the number, the less power is drawn by the device or bus
in that state.  However, the higher the number, the longer the latency for
the device or bus to return to the full-power state (D0 or B0, respectively).

There are two variants of the D3 state defined by the specification.  The first
one is D3hot, referred to as the software accessible D3, because devices can be
programmed to go into it.  The second one, D3cold, is the state that PCI devices
are in when the supply voltage (Vcc) is removed from them.  It is not possible
to program a PCI device to go into D3cold, although there may be a programmable
interface for putting the bus the device is on into a state in which Vcc is
removed from all devices on the bus.

PCI bus power management, however, is not supported by the Linux kernel at the
time of this writing and therefore it is not covered by this document.

Note that every PCI device can be in the full-power state (D0) or in D3cold,
regardless of whether or not it implements the PCI PM Spec.  In addition to
that, if the PCI PM Spec is implemented by the device, it must support D3hot
as well as D0.  The support for the D1 and D2 power states is optional.

PCI devices supporting the PCI PM Spec can be programmed to go to any of the
supported low-power states (except for D3cold).  While in D1-D3hot the
standard configuration registers of the device must be accessible to software
(i.e. the device is required to respond to PCI configuration accesses), although
its I/O and memory spaces are then disabled.  This allows the device to be
programmatically put into D0.  Thus the kernel can switch the device back and
forth between D0 and the supported low-power states (except for D3cold) and the
possible power state transitions the device can undergo are the following:

+----------------------------+
| Current State | New State  |
+----------------------------+
| D0            | D1, D2, D3 |
+----------------------------+
| D1            | D2, D3     |
+----------------------------+
| D2            | D3         |
+----------------------------+
| D1, D2, D3    | D0         |
+----------------------------+

The transition from D3cold to D0 occurs when the supply voltage is provided to
the device (i.e. power is restored).  In that case the device returns to D0 with
a full power-on reset sequence and the power-on defaults are restored to the
device by hardware just as at initial power up.

PCI devices supporting the PCI PM Spec can be programmed to generate PMEs
while in any power state (D0-D3), but they are not required to be capable
of generating PMEs from all supported power states.  In particular, the
capability of generating PMEs from D3cold is optional and depends on the
presence of additional voltage (3.3Vaux) allowing the device to remain
sufficiently active to generate a wakeup signal.

ACPI device power management

139-199

PCI PM의 platform firmware 지원은 system-specific이지만 ACPI-compliant system은 ACPI 표준 device PM interface를 구현해야 합니다.

ACPI BIOS는 ACPI Machine Language(AML) bytecode로 된 control method를 저장합니다. Kernel은 BIOS에서 method를 load하고 AML interpreter로 계산과 memory·I/O 접근을 수행하여 system design에 맞는 작업을 실행합니다.

Control method는 특정 device에 속하지 않는 global method와 device마다 따로 정의하는 device method로 나뉩니다. Device method는 BIOS 작성자가 미리 알고 있던 device에만 사용할 수 있고 PCI power method도 이 범주입니다.

ACPI의 D0~D3는 native PCI state와 대략 대응하지만 D3hot·D3cold를 구분하지 않습니다. 각 state에는 필요한 power resource 집합이 있고 resource마다 `_ON`, `_OFF` method가 있습니다.

Device를 Dx로 전환할 때 kernel은 필요한 resource의 `_ON`을 실행하고 device의 `_PSx`를 호출합니다. D1~D3에서 wake signal을 만들어야 하면 `_PSx`보다 먼저 `_DSW` 또는 구형 `_PSW`를 실행합니다. Target에 필요 없고 다른 device도 쓰지 않는 resource는 `_OFF`로 끕니다. Current D3에서는 이 방식으로 D0에만 갈 수 있습니다.

System-wide sleep에서는 ACPI S0가 working state, S1~S4가 sleep state입니다. `_SxD`는 target system state에서 device가 가질 수 있는 highest-power state를, wakeup device의 `_SxW`는 허용되는 lowest-power state를 알려 줍니다. `_PRW`는 wake signal에 필요한 power resource를 알려 줍니다.

ACPI Dx 전환
target Dxenable required resources: _ONif wakeup: _DSW / _PSWdevice method: _PSxdisable unused resources: _OFF

Resource와 device method를 순서대로 실행하고 필요 없는 shared resource만 끕니다.

System state 관련 ACPI method
Method정보
_SxDTarget S-state에서 device의 highest-power D-state
_SxWWakeup 가능성을 유지하는 lowest-power D-state
_PRWWake signal에 필요한 power resources

Target S-state와 wakeup 요구가 device D-state 경계를 결정합니다.

1.3. ACPI Device Power Management
---------------------------------

The platform firmware support for the power management of PCI devices is
system-specific.  However, if the system in question is compliant with the
Advanced Configuration and Power Interface (ACPI) Specification, like the
majority of x86-based systems, it is supposed to implement device power
management interfaces defined by the ACPI standard.

For this purpose the ACPI BIOS provides special functions called "control
methods" that may be executed by the kernel to perform specific tasks, such as
putting a device into a low-power state.  These control methods are encoded
using special byte-code language called the ACPI Machine Language (AML) and
stored in the machine's BIOS.  The kernel loads them from the BIOS and executes
them as needed using an AML interpreter that translates the AML byte code into
computations and memory or I/O space accesses.  This way, in theory, a BIOS
writer can provide the kernel with a means to perform actions depending
on the system design in a system-specific fashion.

ACPI control methods may be divided into global control methods, that are not
associated with any particular devices, and device control methods, that have
to be defined separately for each device supposed to be handled with the help of
the platform.  This means, in particular, that ACPI device control methods can
only be used to handle devices that the BIOS writer knew about in advance.  The
ACPI methods used for device power management fall into that category.

The ACPI specification assumes that devices can be in one of four power states
labeled as D0, D1, D2, and D3 that roughly correspond to the native PCI PM
D0-D3 states (although the difference between D3hot and D3cold is not taken
into account by ACPI).  Moreover, for each power state of a device there is a
set of power resources that have to be enabled for the device to be put into
that state.  These power resources are controlled (i.e. enabled or disabled)
with the help of their own control methods, _ON and _OFF, that have to be
defined individually for each of them.

To put a device into the ACPI power state Dx (where x is a number between 0 and
3 inclusive) the kernel is supposed to (1) enable the power resources required
by the device in this state using their _ON control methods and (2) execute the
_PSx control method defined for the device.  In addition to that, if the device
is going to be put into a low-power state (D1-D3) and is supposed to generate
wakeup signals from that state, the _DSW (or _PSW, replaced with _DSW by ACPI
3.0) control method defined for it has to be executed before _PSx.  Power
resources that are not required by the device in the target power state and are
not required any more by any other device should be disabled (by executing their
_OFF control methods).  If the current power state of the device is D3, it can
only be put into D0 this way.

However, quite often the power states of devices are changed during a
system-wide transition into a sleep state or back into the working state.  ACPI
defines four system sleep states, S1, S2, S3, and S4, and denotes the system
working state as S0.  In general, the target system sleep (or working) state
determines the highest power (lowest number) state the device can be put
into and the kernel is supposed to obtain this information by executing the
device's _SxD control method (where x is a number between 0 and 4 inclusive).
If the device is required to wake up the system from the target sleep state, the
lowest power (highest number) state it can be put into is also determined by the
target state of the system.  The kernel is then supposed to use the device's
_SxW control method to obtain the number of that state.  It also is supposed to
use the device's _PRW control method to learn which power resources need to be
enabled for the device to be able to generate wakeup signals.

PME·GPE·PCIe wakeup signaling

200-268

Native PME 또는 `_DSW`·`_PSW`로 준비한 wake signal은 system이 S0이면 interrupt로 변환되어 source device를 full-power로 복귀시키고 event를 처리해야 합니다. System이 sleeping이면 core logic가 system wakeup을 시작해야 합니다.

ACPI system에서 conventional PCI wake signal은 General-Purpose Event(GPE)로 변환됩니다. BIOS에는 GPE와 event source의 연결 정보가 기록됩니다. PCI bridge 또는 root bridge의 GPE가 아래 device의 wake signal에 반응할 수 있어 BIOS가 모르는 endpoint의 native PME도 처리할 수 있습니다.

S1~S4에서 발생해 core logic의 wakeup을 시작하는 GPE를 wakeup GPE라고 합니다. S0에서 발생하면 core logic가 System Control Interrupt(SCI)를 만들고 handler가 GPE와 source를 식별하며, 이때의 GPE를 runtime GPE라고 합니다.

Non-ACPI conventional PCI에는 표준 전달 방식이 없지만 PCI Express는 root port가 in-band PME message를 interrupt로 바꾸는 native mechanism을 정의합니다. Root port register에 Requester ID가 기록되어 handler가 source를 식별합니다. Root Complex-integrated endpoint는 root port를 지나지 않고 Root Complex Event Collector가 interrupt를 생성할 수 있습니다.

ACPI system에서도 native PCIe PME와 GPE를 함께 쓸 수 있지만 kernel이 ACPI BIOS에 root-port configuration register control을 넘겨 달라고 요청해야 합니다. BIOS가 거부하면 kernel은 register를 수정할 수 없고 native PCIe PME signaling도 사용할 수 없습니다.

ACPI wake signal
PCI wake / bridge GPES1-S4wakeup GPEcore logic wakes system
PCI wake / bridge GPES0SCIGPE handleridentify source

System state에 따라 같은 GPE가 wakeup 또는 runtime interrupt 경로로 처리됩니다.

Native PCIe PME
PCIe endpoint PMEPCIe hierarchyroot portinterrupt + Requester IDkernel identifies device
Root Complex-integrated endpointRoot Complex Event Collectorinterrupt

In-band message가 hierarchy를 지나 root port interrupt와 Requester ID로 source를 드러냅니다.

1.4. Wakeup Signaling
---------------------

Wakeup signals generated by PCI devices, either as native PCI PMEs, or as
a result of the execution of the _DSW (or _PSW) ACPI control method before
putting the device into a low-power state, have to be caught and handled as
appropriate.  If they are sent while the system is in the working state
(ACPI S0), they should be translated into interrupts so that the kernel can
put the devices generating them into the full-power state and take care of the
events that triggered them.  In turn, if they are sent while the system is
sleeping, they should cause the system's core logic to trigger wakeup.

On ACPI-based systems wakeup signals sent by conventional PCI devices are
converted into ACPI General-Purpose Events (GPEs) which are hardware signals
from the system core logic generated in response to various events that need to
be acted upon.  Every GPE is associated with one or more sources of potentially
interesting events.  In particular, a GPE may be associated with a PCI device
capable of signaling wakeup.  The information on the connections between GPEs
and event sources is recorded in the system's ACPI BIOS from where it can be
read by the kernel.

If a PCI device known to the system's ACPI BIOS signals wakeup, the GPE
associated with it (if there is one) is triggered.  The GPEs associated with PCI
bridges may also be triggered in response to a wakeup signal from one of the
devices below the bridge (this also is the case for root bridges) and, for
example, native PCI PMEs from devices unknown to the system's ACPI BIOS may be
handled this way.

A GPE may be triggered when the system is sleeping (i.e. when it is in one of
the ACPI S1-S4 states), in which case system wakeup is started by its core logic
(the device that was the source of the signal causing the system wakeup to occur
may be identified later).  The GPEs used in such situations are referred to as
wakeup GPEs.

Usually, however, GPEs are also triggered when the system is in the working
state (ACPI S0) and in that case the system's core logic generates a System
Control Interrupt (SCI) to notify the kernel of the event.  Then, the SCI
handler identifies the GPE that caused the interrupt to be generated which,
in turn, allows the kernel to identify the source of the event (that may be
a PCI device signaling wakeup).  The GPEs used for notifying the kernel of
events occurring while the system is in the working state are referred to as
runtime GPEs.

Unfortunately, there is no standard way of handling wakeup signals sent by
conventional PCI devices on systems that are not ACPI-based, but there is one
for PCI Express devices.  Namely, the PCI Express Base Specification introduced
a native mechanism for converting native PCI PMEs into interrupts generated by
root ports.  For conventional PCI devices native PMEs are out-of-band, so they
are routed separately and they need not pass through bridges (in principle they
may be routed directly to the system's core logic), but for PCI Express devices
they are in-band messages that have to pass through the PCI Express hierarchy,
including the root port on the path from the device to the Root Complex.  Thus
it was possible to introduce a mechanism by which a root port generates an
interrupt whenever it receives a PME message from one of the devices below it.
The PCI Express Requester ID of the device that sent the PME message is then
recorded in one of the root port's configuration registers from where it may be
read by the interrupt handler allowing the device to be identified.  [PME
messages sent by PCI Express endpoints integrated with the Root Complex don't
pass through root ports, but instead they cause a Root Complex Event Collector
(if there is one) to generate interrupts.]

In principle the native PCI Express PME signaling may also be used on ACPI-based
systems along with the GPEs, but to use it the kernel has to ask the system's
ACPI BIOS to release control of root port configuration registers.  The ACPI
BIOS, however, is not required to allow the kernel to control these registers
and if it doesn't do that, the kernel must not modify their contents.  Of course
the native PCI Express PME signaling cannot be used by the kernel in that case.

PCI subsystem callback과 pci_dev PM field

269-329

PCI subsystem은 PM core와 PCI driver 사이의 middle layer입니다. `pci_bus_type.pm`은 `pci_dev_pm_ops`를 가리키며 system sleep, hibernation, noirq, runtime PM callback을 모두 제공합니다.

pci_dev_pm_ops callback
범주Callbacks
경계prepare, complete
System suspendsuspend, suspend_noirq, resume_noirq, resume
Hibernation imagefreeze, freeze_noirq, thaw_noirq, thaw
Hibernation powerpoweroff, poweroff_noirq, restore_noirq, restore
Runtime PMruntime_suspend, runtime_resume, runtime_idle

PM core phase를 PCI subsystem routine으로 연결합니다.

이 callback은 driver PM callback을 호출하면서 driver가 알 필요 없는 PCI standard configuration register 관련 작업도 수행합니다.

`struct pci_dev`의 PM field에는 `current_state`, configuration space의 PM capability offset `pm_cap`, PME 가능 state bitmask `pme_support`, PME status polling `pme_poll`, D1·D2 지원 bit, D1/D2 금지 `no_d1d2`, wakeup 준비 `wakeup_prepared`, D3hot→D0 delay `d3hot_delay`가 있습니다. Embedded `struct device`의 PM field도 간접 사용합니다.

pci_dev PM state
Field의미
current_state현재 PCI D-state
pm_capPM capability offset
pme_support / pme_pollPME 가능 state와 polling
d1_support / d2_support / no_d1d2저전력 state 지원·금지
wakeup_preparedWake signaling 준비 여부
d3hot_delayD3hot→D0 지연 ms

Capability, wakeup, transition latency를 subsystem이 추적합니다.

2. PCI Subsystem and Device Power Management
============================================

2.1. Device Power Management Callbacks
--------------------------------------

The PCI Subsystem participates in the power management of PCI devices in a
number of ways.  First of all, it provides an intermediate code layer between
the device power management core (PM core) and PCI device drivers.
Specifically, the pm field of the PCI subsystem's struct bus_type object,
pci_bus_type, points to a struct dev_pm_ops object, pci_dev_pm_ops, containing
pointers to several device power management callbacks::

  const struct dev_pm_ops pci_dev_pm_ops = {
        .prepare = pci_pm_prepare,
        .complete = pci_pm_complete,
        .suspend = pci_pm_suspend,
        .resume = pci_pm_resume,
        .freeze = pci_pm_freeze,
        .thaw = pci_pm_thaw,
        .poweroff = pci_pm_poweroff,
        .restore = pci_pm_restore,
        .suspend_noirq = pci_pm_suspend_noirq,
        .resume_noirq = pci_pm_resume_noirq,
        .freeze_noirq = pci_pm_freeze_noirq,
        .thaw_noirq = pci_pm_thaw_noirq,
        .poweroff_noirq = pci_pm_poweroff_noirq,
        .restore_noirq = pci_pm_restore_noirq,
        .runtime_suspend = pci_pm_runtime_suspend,
        .runtime_resume = pci_pm_runtime_resume,
        .runtime_idle = pci_pm_runtime_idle,
  };

These callbacks are executed by the PM core in various situations related to
device power management and they, in turn, execute power management callbacks
provided by PCI device drivers.  They also perform power management operations
involving some standard configuration registers of PCI devices that device
drivers need not know or care about.

The structure representing a PCI device, struct pci_dev, contains several fields
that these callbacks operate on::

  struct pci_dev {
        ...
        pci_power_t     current_state;  /* Current operating state. */
        int                pm_cap;                /* PM capability offset in the
                                           configuration space */
        unsigned int        pme_support:5;        /* Bitmask of states from which PME#
                                           can be generated */
        unsigned int        pme_poll:1;        /* Poll device's PME status bit */
        unsigned int        d1_support:1;        /* Low power state D1 is supported */
        unsigned int        d2_support:1;        /* Low power state D2 is supported */
        unsigned int        no_d1d2:1;        /* D1 and D2 are forbidden */
        unsigned int        wakeup_prepared:1;  /* Device prepared for wake up */
        unsigned int        d3hot_delay;        /* D3hot->D0 transition time in ms */
        ...
  };

They also indirectly use some fields of the struct device that is embedded in
struct pci_dev.

Device PM 초기화

330-356

PCI subsystem은 `drivers/pci/`의 `pci_pm_init()`과 `pci_acpi_setup()`으로 device를 PM에 준비하고 `struct pci_dev` field를 초기화합니다.

`pci_pm_init()`은 native PCI PM 지원 여부를 검사하고 capability offset을 `pm_cap`에 저장합니다. 지원 D-state와 각 low-power state의 native PME 가능 여부를 조사해 `pci_dev` 및 embedded `struct device` field를 갱신하고 PME 생성을 disable합니다.

`pci_acpi_setup()`은 ACPI BIOS 같은 firmware로 wakeup을 준비할 수 있는지 검사합니다. 가능하면 embedded `struct device`의 wakeup field를 갱신하고 firmware method로 wake signaling을 막은 초기 상태를 만듭니다.

이후 device는 PM 준비가 끝납니다. Driver가 없는 device는 system-wide sleep 왕복 중 PCI core가 수행하는 몇 가지 기본 operation만 받을 수 있습니다.

PCI PM 초기화
pci_pm_init()detect PM cap / D1 D2 / PME statesdisable PME
pci_acpi_setup()detect firmware wakeupupdate struct devicedisable platform wake signaling

Native capability와 platform wakeup을 모두 조사한 뒤 초기 signal을 차단합니다.

2.2. Device Initialization
--------------------------

The PCI subsystem's first task related to device power management is to
prepare the device for power management and initialize the fields of struct
pci_dev used for this purpose.  This happens in two functions defined in
drivers/pci/, pci_pm_init() and pci_acpi_setup().

The first of these functions checks if the device supports native PCI PM
and if that's the case the offset of its power management capability structure
in the configuration space is stored in the pm_cap field of the device's struct
pci_dev object.  Next, the function checks which PCI low-power states are
supported by the device and from which low-power states the device can generate
native PCI PMEs.  The power management fields of the device's struct pci_dev and
the struct device embedded in it are updated accordingly and the generation of
PMEs by the device is disabled.

The second function checks if the device can be prepared to signal wakeup with
the help of the platform firmware, such as the ACPI BIOS.  If that is the case,
the function updates the wakeup fields in struct device embedded in the
device's struct pci_dev and uses the firmware-provided method to prevent the
device from signaling wakeup.

At this point the device is ready for power management.  For driverless devices,
however, this functionality is limited to a few basic operations carried out
during system-wide transitions to a sleep state and back to the working state.

PCI subsystem의 runtime PM

357-425

PCI subsystem은 일반 runtime PM framework 위에 `pci_pm_runtime_suspend()`, `pci_pm_runtime_resume()`, `pci_pm_runtime_idle()`을 제공하고 native PCIe PME와 ACPI GPE runtime wakeup 처리를 구현합니다.

`pm_schedule_suspend()` 또는 `pm_runtime_suspend()`가 PCI suspend를 요청하면 먼저 driver의 `runtime_suspend()`를 호출합니다. 성공하면 standard configuration register를 저장하고 wakeup을 준비한 뒤 target low-power state로 전환합니다.

Target은 wake signal을 보낼 수 있는 state 가운데 lowest-power, 즉 번호가 가장 큰 state입니다. Device capability와 firmware를 바탕으로 PCI subsystem이 signaling mechanism을 선택하고 native·platform method를 함께 사용할 수 있습니다.

Driver의 `runtime_suspend()`는 wakeup 준비나 D-state 전환을 직접 시도하지 않고 필요한 정보를 모두 가진 PCI subsystem에 맡겨야 합니다.

`pm_request_resume()` 또는 `pm_runtime_resume()`는 `pci_pm_runtime_resume()`을 호출합니다. PCI core가 먼저 D0 복귀, full-power wake signaling 차단, configuration 복원을 수행한 뒤 driver의 `runtime_resume()`을 호출합니다. 요청은 driver work 때문이거나 device의 remote wakeup 때문일 수 있습니다.

`pci_pm_runtime_idle()`은 driver의 `runtime_idle()`을 호출하고 error가 없거나 callback이 없으면 `pm_runtime_suspend()`를 요청합니다. Resume 직후 자동 호출될 수도 있으며 실제 suspend 가능 여부는 driver가 판단합니다.

Runtime suspend
pm_runtime_suspend()driver runtime_suspend()save configprepare wakeupselect lowest wake-capable D-state

Driver는 quiesce만 하고 PCI-specific save·wakeup·D-state 작업은 subsystem이 수행합니다.

Runtime resume
driver request / remote wakeuppci_pm_runtime_resume()D0 + block wake + restore configdriver runtime_resume()

PCI core가 먼저 hardware를 접근 가능한 상태로 만든 뒤 driver 기능을 복원합니다.

2.3. Runtime Device Power Management
------------------------------------

The PCI subsystem plays a vital role in the runtime power management of PCI
devices.  For this purpose it uses the general runtime power management
(runtime PM) framework described in Documentation/power/runtime_pm.rst.
Namely, it provides subsystem-level callbacks::

        pci_pm_runtime_suspend()
        pci_pm_runtime_resume()
        pci_pm_runtime_idle()

that are executed by the core runtime PM routines.  It also implements the
entire mechanics necessary for handling runtime wakeup signals from PCI devices
in low-power states, which at the time of this writing works for both the native
PCI Express PME signaling and the ACPI GPE-based wakeup signaling described in
Section 1.

First, a PCI device is put into a low-power state, or suspended, with the help
of pm_schedule_suspend() or pm_runtime_suspend() which for PCI devices call
pci_pm_runtime_suspend() to do the actual job.  For this to work, the device's
driver has to provide a pm->runtime_suspend() callback (see below), which is
run by pci_pm_runtime_suspend() as the first action.  If the driver's callback
returns successfully, the device's standard configuration registers are saved,
the device is prepared to generate wakeup signals and, finally, it is put into
the target low-power state.

The low-power state to put the device into is the lowest-power (highest number)
state from which it can signal wakeup.  The exact method of signaling wakeup is
system-dependent and is determined by the PCI subsystem on the basis of the
reported capabilities of the device and the platform firmware.  To prepare the
device for signaling wakeup and put it into the selected low-power state, the
PCI subsystem can use the platform firmware as well as the device's native PCI
PM capabilities, if supported.

It is expected that the device driver's pm->runtime_suspend() callback will
not attempt to prepare the device for signaling wakeup or to put it into a
low-power state.  The driver ought to leave these tasks to the PCI subsystem
that has all of the information necessary to perform them.

A suspended device is brought back into the "active" state, or resumed,
with the help of pm_request_resume() or pm_runtime_resume() which both call
pci_pm_runtime_resume() for PCI devices.  Again, this only works if the device's
driver provides a pm->runtime_resume() callback (see below).  However, before
the driver's callback is executed, pci_pm_runtime_resume() brings the device
back into the full-power state, prevents it from signaling wakeup while in that
state and restores its standard configuration registers.  Thus the driver's
callback need not worry about the PCI-specific aspects of the device resume.

Note that generally pci_pm_runtime_resume() may be called in two different
situations.  First, it may be called at the request of the device's driver, for
example if there are some data for it to process.  Second, it may be called
as a result of a wakeup signal from the device itself (this sometimes is
referred to as "remote wakeup").  Of course, for this purpose the wakeup signal
is handled in one of the ways described in Section 1 and finally converted into
a notification for the PCI subsystem after the source device has been
identified.

The pci_pm_runtime_idle() function, called for PCI devices by pm_runtime_idle()
and pm_request_idle(), executes the device driver's pm->runtime_idle()
callback, if defined, and if that callback doesn't return error code (or is not
present at all), suspends the device with the help of pm_runtime_suspend().
Sometimes pci_pm_runtime_idle() is called automatically by the PM core (for
example, it is called right after the device has just been resumed), in which
cases it is expected to suspend the device if that makes sense.  Usually,
however, the PCI subsystem doesn't really know if the device really can be
suspended, so it lets the device's driver decide by running its
pm->runtime_idle() callback.

System suspend phase

426-503

System-wide transition은 task freezing 뒤 phase별로 같은 subsystem callback을 모든 device에 실행하고 다음 phase로 넘어갑니다. Memory를 보존하는 ACPI S1~S3 suspend 순서는 `prepare`, `suspend`, `suspend_noirq`이며 PCI callback은 각각 `pci_pm_prepare()`, `pci_pm_suspend()`, `pci_pm_suspend_noirq()`입니다.

`pci_pm_prepare()`는 `pm_runtime_resume()`로 device를 fully functional하게 만든 뒤 driver `prepare()`가 있으면 호출합니다.

`pci_pm_suspend()`는 legacy callback을 우선 처리합니다. `dev_pm_ops`가 없으면 bridge를 제외하고 bus mastering을 끄고 `pcibios_disable_device()`를 실행하는 default suspend를 사용합니다. 그렇지 않으면 driver `suspend()`를 호출하고 성공 뒤 hardware suspend quirk를 적용합니다.

Suspend phase는 dependency path를 공유하지 않는 PCI device끼리 asynchronous parallel 실행될 수 있습니다.

`pci_pm_suspend_noirq()`는 `suspend_device_irqs()` 뒤 실행되어 driver interrupt handler가 호출되지 않습니다. Legacy late callback 또는 driver `suspend_noirq()`를 호출하고, 아직 저장되지 않았다면 standard config를 저장한 뒤 wakeup을 준비하고 low-power state로 전환합니다.

Target은 system target sleep state에서 wake signal이 가능한 가장 낮은 power state입니다. Mechanism 선택과 wake 준비는 PCI subsystem 책임입니다.

일반 driver는 wakeup 준비·D-state 전환을 직접 하지 않습니다. 다만 driver callback이 `pci_save_state()`를 호출하면 subsystem은 driver가 `pci_prepare_to_sleep()`과 `pci_set_power_state()`까지 책임졌다고 간주합니다. 드물게 유용하지만 권장되지는 않습니다.

System suspend
PhasePCI callback핵심 작업
preparepci_pm_prepareRuntime resume + driver prepare
suspendpci_pm_suspendDriver quiesce, async 가능, quirks
suspend_noirqpci_pm_suspend_noirqNo interrupts, config save, wakeup, D-state

모든 device가 각 phase를 끝낸 뒤 다음 phase로 이동합니다.

2.4. System-Wide Power Transitions
----------------------------------
There are a few different types of system-wide power transitions, described in
Documentation/driver-api/pm/devices.rst.  Each of them requires devices to be
handled in a specific way and the PM core executes subsystem-level power
management callbacks for this purpose.  They are executed in phases such that
each phase involves executing the same subsystem-level callback for every device
belonging to the given subsystem before the next phase begins.  These phases
always run after tasks have been frozen.

2.4.1. System Suspend
^^^^^^^^^^^^^^^^^^^^^

When the system is going into a sleep state in which the contents of memory will
be preserved, such as one of the ACPI sleep states S1-S3, the phases are:

        prepare, suspend, suspend_noirq.

The following PCI bus type's callbacks, respectively, are used in these phases::

        pci_pm_prepare()
        pci_pm_suspend()
        pci_pm_suspend_noirq()

The pci_pm_prepare() routine first puts the device into the "fully functional"
state with the help of pm_runtime_resume().  Then, it executes the device
driver's pm->prepare() callback if defined (i.e. if the driver's struct
dev_pm_ops object is present and the prepare pointer in that object is valid).

The pci_pm_suspend() routine first checks if the device's driver implements
legacy PCI suspend routines (see Section 3), in which case the driver's legacy
suspend callback is executed, if present, and its result is returned.  Next, if
the device's driver doesn't provide a struct dev_pm_ops object (containing
pointers to the driver's callbacks), pci_pm_default_suspend() is called, which
simply turns off the device's bus master capability and runs
pcibios_disable_device() to disable it, unless the device is a bridge (PCI
bridges are ignored by this routine).  Next, the device driver's pm->suspend()
callback is executed, if defined, and its result is returned if it fails.
Finally, pci_fixup_device() is called to apply hardware suspend quirks related
to the device if necessary.

Note that the suspend phase is carried out asynchronously for PCI devices, so
the pci_pm_suspend() callback may be executed in parallel for any pair of PCI
devices that don't depend on each other in a known way (i.e. none of the paths
in the device tree from the root bridge to a leaf device contains both of them).

The pci_pm_suspend_noirq() routine is executed after suspend_device_irqs() has
been called, which means that the device driver's interrupt handler won't be
invoked while this routine is running.  It first checks if the device's driver
implements legacy PCI suspend routines (Section 3), in which case the legacy
late suspend routine is called and its result is returned (the standard
configuration registers of the device are saved if the driver's callback hasn't
done that).  Second, if the device driver's struct dev_pm_ops object is not
present, the device's standard configuration registers are saved and the routine
returns success.  Otherwise the device driver's pm->suspend_noirq() callback is
executed, if present, and its result is returned if it fails.  Next, if the
device's standard configuration registers haven't been saved yet (one of the
device driver's callbacks executed before might do that), pci_pm_suspend_noirq()
saves them, prepares the device to signal wakeup (if necessary) and puts it into
a low-power state.

The low-power state to put the device into is the lowest-power (highest number)
state from which it can signal wakeup while the system is in the target sleep
state.  Just like in the runtime PM case described above, the mechanism of
signaling wakeup is system-dependent and determined by the PCI subsystem, which
is also responsible for preparing the device to signal wakeup from the system's
target sleep state as appropriate.

PCI device drivers (that don't implement legacy power management callbacks) are
generally not expected to prepare devices for signaling wakeup or to put them
into low-power states.  However, if one of the driver's suspend callbacks
(pm->suspend() or pm->suspend_noirq()) saves the device's standard configuration
registers, pci_pm_suspend_noirq() will assume that the device has been prepared
to signal wakeup and put into a low-power state by the driver (the driver is
then assumed to have used the helper functions provided by the PCI subsystem for
this purpose).  PCI device drivers are not encouraged to do that, but in some
rare cases doing that in the driver may be the optimum approach.

System resume phase

504-551

S1~S3에서 S0로 복귀하는 순서는 `resume_noirq`, `resume`, `complete`이며 PCI callback은 `pci_pm_resume_noirq()`, `pci_pm_resume()`, `pci_pm_complete()`입니다.

`pci_pm_resume_noirq()`는 모든 device를 무조건 D0로 올리고 standard configuration register를 복원하며 early resume quirk를 적용합니다. Driver interrupt가 처음 호출될 때 모든 PCI device가 접근 가능해야 shared interrupt 문제를 피할 수 있기 때문입니다. 이후 legacy early resume 또는 driver `resume_noirq()`를 호출합니다.

`pci_pm_resume()`은 실패한 suspend error path에서 config가 복원되지 않았으면 복원하고 resume quirk를 적용합니다. Legacy resume가 아니면 wake signaling을 block하고 driver `resume()`을 호출합니다.

Resume phase도 dependency가 없는 device끼리 asynchronous parallel 실행될 수 있습니다. `pci_pm_complete()`는 driver `complete()`만 호출합니다.

System resume
PhasePCI callback핵심 작업
resume_noirqpci_pm_resume_noirqD0, config restore, early quirks
resumepci_pm_resumeBlock wake, driver restore, async 가능
completepci_pm_completeDriver completion cleanup

Noirq 단계에서 먼저 모든 PCI hardware를 D0·config 복원 상태로 만듭니다.

2.4.2. System Resume
^^^^^^^^^^^^^^^^^^^^

When the system is undergoing a transition from a sleep state in which the
contents of memory have been preserved, such as one of the ACPI sleep states
S1-S3, into the working state (ACPI S0), the phases are:

        resume_noirq, resume, complete.

The following PCI bus type's callbacks, respectively, are executed in these
phases::

        pci_pm_resume_noirq()
        pci_pm_resume()
        pci_pm_complete()

The pci_pm_resume_noirq() routine first puts the device into the full-power
state, restores its standard configuration registers and applies early resume
hardware quirks related to the device, if necessary.  This is done
unconditionally, regardless of whether or not the device's driver implements
legacy PCI power management callbacks (this way all PCI devices are in the
full-power state and their standard configuration registers have been restored
when their interrupt handlers are invoked for the first time during resume,
which allows the kernel to avoid problems with the handling of shared interrupts
by drivers whose devices are still suspended).  If legacy PCI power management
callbacks (see Section 3) are implemented by the device's driver, the legacy
early resume callback is executed and its result is returned.  Otherwise, the
device driver's pm->resume_noirq() callback is executed, if defined, and its
result is returned.

The pci_pm_resume() routine first checks if the device's standard configuration
registers have been restored and restores them if that's not the case (this
only is necessary in the error path during a failing suspend).  Next, resume
hardware quirks related to the device are applied, if necessary, and if the
device's driver implements legacy PCI power management callbacks (see
Section 3), the driver's legacy resume callback is executed and its result is
returned.  Otherwise, the device's wakeup signaling mechanisms are blocked and
its driver's pm->resume() callback is executed, if defined (the callback's
result is then returned).

The resume phase is carried out asynchronously for PCI devices, like the
suspend phase described above, which means that if two PCI devices don't depend
on each other in a known way, the pci_pm_resume() routine may be executed for
both of them in parallel.

The pci_pm_complete() routine only executes the device driver's pm->complete()
callback, if defined.

Hibernation image와 poweroff phase

552-631

Hibernation은 persistent storage에 system image를 원자적으로 만들고 저장해야 하므로 suspend보다 복잡합니다. Image 생성에 system RAM의 최소 약 50% free가 필요하며 device를 `prepare`, `freeze`, `freeze_noirq` phase로 quiesce합니다.

`pci_pm_freeze()`는 `pci_pm_suspend()`와 비슷하지만 driver `freeze()`를 호출하고 suspend hardware quirk를 적용하지 않습니다. Dependency 없는 device끼리 asynchronous 실행됩니다.

`pci_pm_freeze_noirq()`는 `freeze_noirq()`를 호출하고 config를 저장하지만 wakeup 준비나 low-power 전환은 하지 않습니다.

Image를 storage에 쓰려면 frozen device가 다시 I/O를 처리해야 하므로 `thaw_noirq`, `thaw`, `complete`로 복귀합니다. `pci_pm_thaw_noirq()`는 D0와 config 복원을 하고 driver `thaw_noirq()`를, `pci_pm_thaw()`는 driver `thaw()`를 호출하며 async 실행 가능합니다.

Image 저장 뒤 ACPI S4 같은 target sleep으로 가기 전에 `prepare`, `poweroff`, `poweroff_noirq`를 실행합니다. 뒤 두 callback은 suspend 계열과 비슷하지만 standard config를 저장하지 않습니다.

Hibernation phase 집합
목적PhasesPCI callbacks
Image 생성prepare → freeze → freeze_noirqpci_pm_prepare / freeze / freeze_noirq
Image 저장용 I/O 복귀thaw_noirq → thaw → completepci_pm_thaw_noirq / thaw / complete
최종 S4 poweroffprepare → poweroff → poweroff_noirqpci_pm_prepare / poweroff / poweroff_noirq

Image 생성·저장·최종 poweroff에 서로 다른 callback 집합을 사용합니다.

2.4.3. System Hibernation
^^^^^^^^^^^^^^^^^^^^^^^^^

System hibernation is more complicated than system suspend, because it requires
a system image to be created and written into a persistent storage medium.  The
image is created atomically and all devices are quiesced, or frozen, before that
happens.

The freezing of devices is carried out after enough memory has been freed (at
the time of this writing the image creation requires at least 50% of system RAM
to be free) in the following three phases:

        prepare, freeze, freeze_noirq

that correspond to the PCI bus type's callbacks::

        pci_pm_prepare()
        pci_pm_freeze()
        pci_pm_freeze_noirq()

This means that the prepare phase is exactly the same as for system suspend.
The other two phases, however, are different.

The pci_pm_freeze() routine is quite similar to pci_pm_suspend(), but it runs
the device driver's pm->freeze() callback, if defined, instead of pm->suspend(),
and it doesn't apply the suspend-related hardware quirks.  It is executed
asynchronously for different PCI devices that don't depend on each other in a
known way.

The pci_pm_freeze_noirq() routine, in turn, is similar to
pci_pm_suspend_noirq(), but it calls the device driver's pm->freeze_noirq()
routine instead of pm->suspend_noirq().  It also doesn't attempt to prepare the
device for signaling wakeup and put it into a low-power state.  Still, it saves
the device's standard configuration registers if they haven't been saved by one
of the driver's callbacks.

Once the image has been created, it has to be saved.  However, at this point all
devices are frozen and they cannot handle I/O, while their ability to handle
I/O is obviously necessary for the image saving.  Thus they have to be brought
back to the fully functional state and this is done in the following phases:

        thaw_noirq, thaw, complete

using the following PCI bus type's callbacks::

        pci_pm_thaw_noirq()
        pci_pm_thaw()
        pci_pm_complete()

respectively.

The first of them, pci_pm_thaw_noirq(), is analogous to pci_pm_resume_noirq().
It puts the device into the full power state and restores its standard
configuration registers.  It also executes the device driver's pm->thaw_noirq()
callback, if defined, instead of pm->resume_noirq().

The pci_pm_thaw() routine is similar to pci_pm_resume(), but it runs the device
driver's pm->thaw() callback instead of pm->resume().  It is executed
asynchronously for different PCI devices that don't depend on each other in a
known way.

The complete phase is the same as for system resume.

After saving the image, devices need to be powered down before the system can
enter the target sleep state (ACPI S4 for ACPI-based systems).  This is done in
three phases:

        prepare, poweroff, poweroff_noirq

where the prepare phase is exactly the same as for system suspend.  The other
two phases are analogous to the suspend and suspend_noirq phases, respectively.
The PCI subsystem-level callbacks they correspond to::

        pci_pm_poweroff()
        pci_pm_poweroff_noirq()

work in analogy with pci_pm_suspend() and pci_pm_suspend_noirq(), respectively,
although they don't attempt to save the device's standard configuration
registers.

Boot kernel과 image kernel restore

632-681

System restore에서는 boot loader가 새 boot kernel을 실행하고, 이 kernel이 hibernation image를 memory에 load한 뒤 자신의 code·data를 image 안의 hibernated image kernel 것으로 교체합니다.

교체 전 boot kernel은 `prepare`, `freeze`, `freeze_noirq`로 자신이 driver를 가진 device만 freeze합니다. 다른 device는 boot loader가 남긴 상태 그대로입니다.

Memory restore가 실패하면 boot kernel이 `thaw_noirq`, `thaw`, `complete`를 실행하고 정상 동작을 계속합니다.

성공하면 control이 image kernel로 넘어가고 `restore_noirq`, `restore`, `complete`로 pre-hibernation 기능을 복원합니다. 앞의 두 phase는 resume 계열과 비슷하지만 driver의 `restore_noirq()`와 `restore()`를 호출합니다.

System restore
boot loaderboot kernelload imagefreeze boot-kernel devicesrestore memoryimage kernelrestore_noirq → restore → complete
memory restore failurethaw_noirq → thaw → completeboot kernel continues

Boot kernel이 image를 load·freeze하고 성공 시 image kernel이 device를 복원합니다.

2.4.4. System Restore
^^^^^^^^^^^^^^^^^^^^^

System restore requires a hibernation image to be loaded into memory and the
pre-hibernation memory contents to be restored before the pre-hibernation system
activity can be resumed.

As described in Documentation/driver-api/pm/devices.rst, the hibernation image
is loaded into memory by a fresh instance of the kernel, called the boot kernel,
which in turn is loaded and run by a boot loader in the usual way.  After the
boot kernel has loaded the image, it needs to replace its own code and data with
the code and data of the "hibernated" kernel stored within the image, called the
image kernel.  For this purpose all devices are frozen just like before creating
the image during hibernation, in the

        prepare, freeze, freeze_noirq

phases described above.  However, the devices affected by these phases are only
those having drivers in the boot kernel; other devices will still be in whatever
state the boot loader left them.

Should the restoration of the pre-hibernation memory contents fail, the boot
kernel would go through the "thawing" procedure described above, using the
thaw_noirq, thaw, and complete phases (that will only affect the devices having
drivers in the boot kernel), and then continue running normally.

If the pre-hibernation memory contents are restored successfully, which is the
usual situation, control is passed to the image kernel, which then becomes
responsible for bringing the system back to the working state.  To achieve this,
it must restore the devices' pre-hibernation functionality, which is done much
like waking up from the memory sleep state, although it involves different
phases:

        restore_noirq, restore, complete

The first two of these are analogous to the resume_noirq and resume phases
described above, respectively, and correspond to the following PCI subsystem
callbacks::

        pci_pm_restore_noirq()
        pci_pm_restore()

These callbacks work in analogy with pci_pm_resume_noirq() and pci_pm_resume(),
respectively, but they execute the device driver's pm->restore_noirq() and
pm->restore() callbacks, if available.

The complete phase is carried out in exactly the same way as during system
resume.

Driver dev_pm_ops와 prepare()

682-732

PCI driver는 subsystem callback이 호출할 PM callback을 제공하고 device runtime PM을 제어합니다. 권장 방식은 `Documentation/driver-api/pm/devices.rst`의 `struct dev_pm_ops`이며 `struct pci_driver`의 legacy `suspend()`·`resume()`은 runtime PM을 지원하지 않아 새 driver에 부적합합니다.

Driver의 `dev_pm_ops` pointer를 `pci_driver.driver.pm`에 지정하면 legacy callback은 non-NULL이어도 무시됩니다. 개별 callback은 선택 사항이며 없으면 PCI subsystem이 단순 default 방식으로 처리합니다.

`prepare()`는 system suspend, image 생성 전 hibernation, image 저장 뒤 poweroff, image load 직후 restore에서 실행됩니다.

Runtime 중 언제든 child가 등록될 수 있는 device라면 `prepare()`가 새 child 등록을 막고 `resume_noirq()`, `thaw_noirq()`, `restore_noirq()` 중 하나가 실행될 때까지 유지해야 합니다.

Suspend 준비 작업도 할 수 있지만 memory를 할당하면 안 됩니다. 추가 memory가 필요하면 `Documentation/driver-api/pm/notifiers.rst`의 notifier 등에서 미리 할당해야 합니다.

Driver PM 선택
방식Runtime PM권장
struct dev_pm_ops via driver.pm지원권장
legacy pci_driver suspend/resume미지원새 driver에 부적합

새 PCI driver는 dev_pm_ops를 사용하고 필요한 callback만 구현합니다.

3. PCI Device Drivers and Power Management
==========================================

3.1. Power Management Callbacks
-------------------------------

PCI device drivers participate in power management by providing callbacks to be
executed by the PCI subsystem's power management routines described above and by
controlling the runtime power management of their devices.

At the time of this writing there are two ways to define power management
callbacks for a PCI device driver, the recommended one, based on using a
dev_pm_ops structure described in Documentation/driver-api/pm/devices.rst, and
the "legacy" one, in which the .suspend() and .resume() callbacks from struct
pci_driver are used.  The legacy approach, however, doesn't allow one to define
runtime power management callbacks and is not really suitable for any new
drivers.  Therefore it is not covered by this document (refer to the source code
to learn more about it).

It is recommended that all PCI device drivers define a struct dev_pm_ops object
containing pointers to power management (PM) callbacks that will be executed by
the PCI subsystem's PM routines in various circumstances.  A pointer to the
driver's struct dev_pm_ops object has to be assigned to the driver.pm field in
its struct pci_driver object.  Once that has happened, the "legacy" PM callbacks
in struct pci_driver are ignored (even if they are not NULL).

The PM callbacks in struct dev_pm_ops are not mandatory and if they are not
defined (i.e. the respective fields of struct dev_pm_ops are unset) the PCI
subsystem will handle the device in a simplified default manner.  If they are
defined, though, they are expected to behave as described in the following
subsections.

3.1.1. prepare()
^^^^^^^^^^^^^^^^

The prepare() callback is executed during system suspend, during hibernation
(when a hibernation image is about to be created), during power-off after
saving a hibernation image and during system restore, when a hibernation image
has just been loaded into memory.

This callback is only necessary if the driver's device has children that in
general may be registered at any time.  In that case the role of the prepare()
callback is to prevent new children of the device from being registered until
one of the resume_noirq(), thaw_noirq(), or restore_noirq() callbacks is run.

In addition to that the prepare() callback may carry out some operations
preparing the device to be suspended, although it should not allocate memory
(if additional memory is required to suspend the device, it has to be
preallocated earlier, for example in a suspend/hibernate notifier as described
in Documentation/driver-api/pm/notifiers.rst).

suspend()와 suspend_noirq()

733-771

`suspend()`는 system suspend에서 모든 device의 `prepare()` 뒤 실행됩니다. Device를 quiesce하고 PCI subsystem이 low-power state에 넣을 수 있도록 준비합니다.

Driver가 standard configuration register 저장, system wakeup 준비, D-state 전환을 할 필요가 없고 권장되지도 않습니다. PCI subsystem이 모두 처리할 수 있습니다.

드문 경우 직접 처리한다면 `pci_save_state()`, `pci_prepare_to_sleep()`, `pci_set_power_state()`를 순서에 맞게 사용합니다. Driver가 `pci_save_state()`를 호출하면 subsystem은 나머지 wakeup·D-state 작업도 driver 책임이라고 보고 수행하지 않습니다.

`suspend()` 중에는 interrupt handler가 실행될 수 있으므로 interrupt 처리 능력에 의존하는 suspend 작업은 여기서 수행합니다.

`suspend_noirq()`는 모든 device의 `suspend()` 뒤 PM core가 device interrupt를 disable한 다음 실행됩니다. Handler와 race가 날 작업을 interrupt가 없는 이 callback에서 안전하게 수행할 수 있습니다.

suspend vs suspend_noirq
CallbackInterrupt handler작업
suspend호출 가능Quiesce, interrupt 의존 작업
suspend_noirq호출 안 됨Handler와 race 가능한 마지막 작업

Interrupt handler 실행 가능 여부가 두 callback의 핵심 차이입니다.

3.1.2. suspend()
^^^^^^^^^^^^^^^^

The suspend() callback is only executed during system suspend, after prepare()
callbacks have been executed for all devices in the system.

This callback is expected to quiesce the device and prepare it to be put into a
low-power state by the PCI subsystem.  It is not required (in fact it even is
not recommended) that a PCI driver's suspend() callback save the standard
configuration registers of the device, prepare it for waking up the system, or
put it into a low-power state.  All of these operations can very well be taken
care of by the PCI subsystem, without the driver's participation.

However, in some rare case it is convenient to carry out these operations in
a PCI driver.  Then, pci_save_state(), pci_prepare_to_sleep(), and
pci_set_power_state() should be used to save the device's standard configuration
registers, to prepare it for system wakeup (if necessary), and to put it into a
low-power state, respectively.  Moreover, if the driver calls pci_save_state(),
the PCI subsystem will not execute either pci_prepare_to_sleep(), or
pci_set_power_state() for its device, so the driver is then responsible for
handling the device as appropriate.

While the suspend() callback is being executed, the driver's interrupt handler
can be invoked to handle an interrupt from the device, so all suspend-related
operations relying on the driver's ability to handle interrupts should be
carried out in this callback.

3.1.3. suspend_noirq()
^^^^^^^^^^^^^^^^^^^^^^

The suspend_noirq() callback is only executed during system suspend, after
suspend() callbacks have been executed for all devices in the system and
after device interrupts have been disabled by the PM core.

The difference between suspend_noirq() and suspend() is that the driver's
interrupt handler will not be invoked while suspend_noirq() is running.  Thus
suspend_noirq() can carry out operations that would cause race conditions to
arise if they were performed in suspend().

freeze()와 freeze_noirq()

772-806

`freeze()`는 image 생성 전 hibernation과 image load 뒤 restore에서 `prepare()` 다음 실행되는 hibernation 전용 callback입니다. 역할은 `suspend()`와 비슷합니다.

Driver가 직접 low-power 전환을 책임지는 드문 경우에만 `freeze()`와 `suspend()`의 동작이 달라집니다. `freeze()`는 system wakeup 준비나 low-power 전환을 하지 않아야 하지만, `freeze()` 또는 `freeze_noirq()` 중 하나가 `pci_save_state()`로 standard config를 저장해야 합니다.

`freeze_noirq()`는 모든 device의 prepare·freeze 뒤 interrupt disable 상태에서 실행됩니다. 역할과 두 callback의 차이는 suspend_noirq·suspend 관계와 같으며 별도 구현이 필요한 경우는 매우 드뭅니다.

Hibernation freeze callbacks
CallbackInterrupt필수·금지
freeze가능Quiesce; 직접 PM이면 suspend와 구분
freeze_noirq불가Race-free final work
둘 중 하나-pci_save_state() 필요
둘 모두-Wakeup 준비·low-power 전환 불필요

Image 생성에는 config snapshot만 필요하고 wakeup·D-state 전환은 하지 않습니다.

3.1.4. freeze()
^^^^^^^^^^^^^^^

The freeze() callback is hibernation-specific and is executed in two situations,
during hibernation, after prepare() callbacks have been executed for all devices
in preparation for the creation of a system image, and during restore,
after a system image has been loaded into memory from persistent storage and the
prepare() callbacks have been executed for all devices.

The role of this callback is analogous to the role of the suspend() callback
described above.  In fact, they only need to be different in the rare cases when
the driver takes the responsibility for putting the device into a low-power
state.

In that cases the freeze() callback should not prepare the device system wakeup
or put it into a low-power state.  Still, either it or freeze_noirq() should
save the device's standard configuration registers using pci_save_state().

3.1.5. freeze_noirq()
^^^^^^^^^^^^^^^^^^^^^

The freeze_noirq() callback is hibernation-specific.  It is executed during
hibernation, after prepare() and freeze() callbacks have been executed for all
devices in preparation for the creation of a system image, and during restore,
after a system image has been loaded into memory and after prepare() and
freeze() callbacks have been executed for all devices.  It is always executed
after device interrupts have been disabled by the PM core.

The role of this callback is analogous to the role of the suspend_noirq()
callback described above and it very rarely is necessary to define
freeze_noirq().

The difference between freeze_noirq() and freeze() is analogous to the
difference between suspend_noirq() and suspend().

poweroff 계열과 resume_noirq()

807-850

`poweroff()`는 hibernation image를 persistent storage에 저장한 뒤 system power-off 직전에 실행됩니다. 역할은 suspend·freeze와 비슷하지만 register contents를 저장할 필요가 없습니다.

Driver가 PCI subsystem 대신 low-power 전환을 맡는다면 `pci_prepare_to_sleep()`과 `pci_set_power_state()`로 wakeup과 D-state를 준비하되 standard config는 저장하지 않아도 됩니다.

`poweroff_noirq()`는 모든 poweroff 뒤 실행되며 suspend_noirq·freeze_noirq와 비슷하지만 register 저장이 필요 없습니다.

`resume_noirq()`는 system resume에서 nonboot CPU를 enable한 뒤, driver interrupt handler가 호출되지 않는 상태에서 실행됩니다. PCI subsystem이 모든 device를 D0로 올리고 config를 복원하므로 보통 필요하지 않으며, `resume()`에서 하면 interrupt handler와 race가 날 작업에만 사용합니다.

Poweroff와 early resume
Callback주요 규칙
poweroffWakeup·D-state는 필요 시 직접, config save 불필요
poweroff_noirqInterrupt 없는 마지막 poweroff, config save 불필요
resume_noirqD0·config restore 뒤 race-sensitive work만

Hibernation 종료에는 저장이 필요 없고 resume noirq에는 PCI core 복원이 선행됩니다.

3.1.6. poweroff()
^^^^^^^^^^^^^^^^^

The poweroff() callback is hibernation-specific.  It is executed when the system
is about to be powered off after saving a hibernation image to a persistent
storage.  prepare() callbacks are executed for all devices before poweroff() is
called.

The role of this callback is analogous to the role of the suspend() and freeze()
callbacks described above, although it does not need to save the contents of
the device's registers.  In particular, if the driver wants to put the device
into a low-power state itself instead of allowing the PCI subsystem to do that,
the poweroff() callback should use pci_prepare_to_sleep() and
pci_set_power_state() to prepare the device for system wakeup and to put it
into a low-power state, respectively, but it need not save the device's standard
configuration registers.

3.1.7. poweroff_noirq()
^^^^^^^^^^^^^^^^^^^^^^^

The poweroff_noirq() callback is hibernation-specific.  It is executed after
poweroff() callbacks have been executed for all devices in the system.

The role of this callback is analogous to the role of the suspend_noirq() and
freeze_noirq() callbacks described above, but it does not need to save the
contents of the device's registers.

The difference between poweroff_noirq() and poweroff() is analogous to the
difference between suspend_noirq() and suspend().

3.1.8. resume_noirq()
^^^^^^^^^^^^^^^^^^^^^

The resume_noirq() callback is only executed during system resume, after the
PM core has enabled the non-boot CPUs.  The driver's interrupt handler will not
be invoked while resume_noirq() is running, so this callback can carry out
operations that might race with the interrupt handler.

Since the PCI subsystem unconditionally puts all devices into the full power
state in the resume_noirq phase of system resume and restores their standard
configuration registers, resume_noirq() is usually not necessary.  In general
it should only be used for performing operations that would lead to race
conditions if carried out by resume().

resume·thaw·restore·complete

851-935

`resume()`은 모든 `resume_noirq()` 뒤 interrupt가 enable된 상태에서 pre-suspend configuration과 fully functional I/O를 복원합니다.

`thaw_noirq()`는 image 생성 뒤 또는 restore image load 실패 뒤 nonboot CPU가 enable된 hibernation thaw phase에서 interrupt 없이 실행됩니다. `resume_noirq()`와 비슷하지만 freeze 계열 뒤이므로 일반적으로 register를 수정할 필요가 없습니다.

`thaw()`는 모든 thaw_noirq 뒤 interrupt enable 상태에서 pre-freeze configuration을 복원해 정상 동작하게 합니다.

`restore_noirq()`는 boot kernel에서 image kernel로 control이 넘어간 뒤 image kernel PM core가 nonboot CPU를 enable한 restore phase에서 실행됩니다. 이전 device state를 전혀 가정할 수 없다는 점만 제외하면 resume_noirq와 같습니다.

`restore()`는 모든 restore_noirq 뒤 interrupt handler가 enable된 상태에서 실행되며 resume과 유사합니다. 대부분 PCI driver에서는 resume_noirq와 restore_noirq, resume과 restore 사이에 차이가 없습니다.

`complete()`는 system resume의 resume 뒤, hibernation image 생성 전 thaw 뒤, restore의 restore 뒤에 실행됩니다. Image load 실패 때도 boot kernel driver가 있는 device의 thaw 뒤 실행될 수 있습니다. 완전히 선택 사항이지만 prepare 작업을 되돌려야 하면 필요합니다.

복귀 callback
경로NoirqNormal마무리
System resumeresume_noirqresumecomplete
Image 생성 뒤thaw_noirqthawcomplete
Image kernel restorerestore_noirqrestorecomplete

경로별 이름은 다르지만 noirq→normal→complete의 역할이 대응합니다.

3.1.9. resume()
^^^^^^^^^^^^^^^

The resume() callback is only executed during system resume, after
resume_noirq() callbacks have been executed for all devices in the system and
device interrupts have been enabled by the PM core.

This callback is responsible for restoring the pre-suspend configuration of the
device and bringing it back to the fully functional state.  The device should be
able to process I/O in a usual way after resume() has returned.

3.1.10. thaw_noirq()
^^^^^^^^^^^^^^^^^^^^

The thaw_noirq() callback is hibernation-specific.  It is executed after a
system image has been created and the non-boot CPUs have been enabled by the PM
core, in the thaw_noirq phase of hibernation.  It also may be executed if the
loading of a hibernation image fails during system restore (it is then executed
after enabling the non-boot CPUs).  The driver's interrupt handler will not be
invoked while thaw_noirq() is running.

The role of this callback is analogous to the role of resume_noirq().  The
difference between these two callbacks is that thaw_noirq() is executed after
freeze() and freeze_noirq(), so in general it does not need to modify the
contents of the device's registers.

3.1.11. thaw()
^^^^^^^^^^^^^^

The thaw() callback is hibernation-specific.  It is executed after thaw_noirq()
callbacks have been executed for all devices in the system and after device
interrupts have been enabled by the PM core.

This callback is responsible for restoring the pre-freeze configuration of
the device, so that it will work in a usual way after thaw() has returned.

3.1.12. restore_noirq()
^^^^^^^^^^^^^^^^^^^^^^^

The restore_noirq() callback is hibernation-specific.  It is executed in the
restore_noirq phase of hibernation, when the boot kernel has passed control to
the image kernel and the non-boot CPUs have been enabled by the image kernel's
PM core.

This callback is analogous to resume_noirq() with the exception that it cannot
make any assumption on the previous state of the device, even if the BIOS (or
generally the platform firmware) is known to preserve that state over a
suspend-resume cycle.

For the vast majority of PCI device drivers there is no difference between
resume_noirq() and restore_noirq().

3.1.13. restore()
^^^^^^^^^^^^^^^^^

The restore() callback is hibernation-specific.  It is executed after
restore_noirq() callbacks have been executed for all devices in the system and
after the PM core has enabled device drivers' interrupt handlers to be invoked.

This callback is analogous to resume(), just like restore_noirq() is analogous
to resume_noirq().  Consequently, the difference between restore_noirq() and
restore() is analogous to the difference between resume_noirq() and resume().

For the vast majority of PCI device drivers there is no difference between
resume() and restore().

3.1.14. complete()
^^^^^^^^^^^^^^^^^^

The complete() callback is executed in the following situations:

  - during system resume, after resume() callbacks have been executed for all
    devices,
  - during hibernation, before saving the system image, after thaw() callbacks
    have been executed for all devices,
  - during system restore, when the system is going back to its pre-hibernation
    state, after restore() callbacks have been executed for all devices.

It also may be executed if the loading of a hibernation image into memory fails
(in that case it is run after thaw() callbacks have been executed for all
devices that have drivers in the boot kernel).

This callback is entirely optional, although it may be necessary if the
prepare() callback performs operations that need to be reversed.

Runtime callback과 PM macro

936-996

`runtime_suspend()`는 runtime에서 device를 quiesce하고 low-power 전환을 준비하지만 PCI-specific action은 subsystem이 수행하도록 허용해야 합니다.

`runtime_resume()`은 PCI subsystem이 D0로 올린 뒤 device의 정상 기능을 복원하며 반환 후 usual I/O가 가능해야 합니다.

`runtime_idle()`은 suspend가 바람직할 수 있을 때, 특히 spurious event로 resume된 직후 자동 호출될 수 있습니다. Callback이 없거나 0을 반환하면 PCI subsystem이 `pm_runtime_suspend()`를 호출해 runtime_suspend로 이어집니다.

여러 system callback을 같은 routine에 연결할 수 있습니다. `DEFINE_SIMPLE_DEV_PM_OPS()`는 suspend routine을 `.suspend/.freeze/.poweroff`에, resume routine을 `.resume/.thaw/.restore`에 연결합니다.

`DEFINE_RUNTIME_DEV_PM_OPS()`는 여기에 `pm_runtime_force_suspend()`와 `pm_runtime_force_resume()` runtime pointer도 설정합니다. `SYSTEM_SLEEP_PM_OPS()`는 `struct dev_pm_ops` 선언 안에서 같은 system-sleep mapping을 제공합니다.

Callback mapping macro
MacroSystem callbacksRuntime callbacks
DEFINE_SIMPLE_DEV_PM_OPSsuspend/freeze/poweroff + resume/thaw/restore미설정
DEFINE_RUNTIME_DEV_PM_OPS동일force_suspend / force_resume
SYSTEM_SLEEP_PM_OPSdev_pm_ops 내부 system mapping별도

반복되는 system-sleep callback mapping을 한 쌍의 routine으로 선언합니다.

3.1.15. runtime_suspend()
^^^^^^^^^^^^^^^^^^^^^^^^^

The runtime_suspend() callback is specific to device runtime power management
(runtime PM).  It is executed by the PM core's runtime PM framework when the
device is about to be suspended (i.e. quiesced and put into a low-power state)
at run time.

This callback is responsible for freezing the device and preparing it to be
put into a low-power state, but it must allow the PCI subsystem to perform all
of the PCI-specific actions necessary for suspending the device.

3.1.16. runtime_resume()
^^^^^^^^^^^^^^^^^^^^^^^^

The runtime_resume() callback is specific to device runtime PM.  It is executed
by the PM core's runtime PM framework when the device is about to be resumed
(i.e. put into the full-power state and programmed to process I/O normally) at
run time.

This callback is responsible for restoring the normal functionality of the
device after it has been put into the full-power state by the PCI subsystem.
The device is expected to be able to process I/O in the usual way after
runtime_resume() has returned.

3.1.17. runtime_idle()
^^^^^^^^^^^^^^^^^^^^^^

The runtime_idle() callback is specific to device runtime PM.  It is executed
by the PM core's runtime PM framework whenever it may be desirable to suspend
the device according to the PM core's information.  In particular, it is
automatically executed right after runtime_resume() has returned in case the
resume of the device has happened as a result of a spurious event.

This callback is optional, but if it is not implemented or if it returns 0, the
PCI subsystem will call pm_runtime_suspend() for the device, which in turn will
cause the driver's runtime_suspend() callback to be executed.

3.1.18. Pointing Multiple Callback Pointers to One Routine
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

Although in principle each of the callbacks described in the previous
subsections can be defined as a separate function, it often is convenient to
point two or more members of struct dev_pm_ops to the same routine.  There are
a few convenience macros that can be used for this purpose.

The DEFINE_SIMPLE_DEV_PM_OPS() declares a struct dev_pm_ops object with one
suspend routine pointed to by the .suspend(), .freeze(), and .poweroff()
members and one resume routine pointed to by the .resume(), .thaw(), and
.restore() members.  The other function pointers in this struct dev_pm_ops are
unset.

The DEFINE_RUNTIME_DEV_PM_OPS() is similar to DEFINE_SIMPLE_DEV_PM_OPS(), but it
additionally sets the .runtime_resume() pointer to pm_runtime_force_resume()
and the .runtime_suspend() pointer to pm_runtime_force_suspend().

The SYSTEM_SLEEP_PM_OPS() can be used inside of a declaration of struct
dev_pm_ops to indicate that one suspend routine is to be pointed to by the
.suspend(), .freeze(), and .poweroff() members and one resume routine is to
be pointed to by the .resume(), .thaw(), and .restore() members.

Power management driver flag

997-1039

Driver는 probe 때 `dev_pm_set_driver_flags()`로 PM core와 PCI middle layer 동작에 영향을 주는 flag를 한 번 설정해야 하며 이후 직접 갱신하면 안 됩니다.

`DPM_FLAG_NO_DIRECT_COMPLETE`는 system suspend 시작 때 device가 runtime-suspended이면 system callback을 건너뛰는 direct-complete를 영구 차단합니다. Ancestor에도 영향을 주므로 꼭 필요할 때만 사용합니다.

`DPM_FLAG_SMART_PREPARE`는 driver `prepare()`가 양수를 반환한 경우에만 `pci_pm_prepare()`도 양수를 반환하게 하여 runtime마다 direct-complete를 동적으로 거부할 수 있게 합니다.

`DPM_FLAG_SMART_SUSPEND`는 driver 관점에서 system suspend 동안 device를 runtime suspend 상태로 둬도 안전함을 알립니다. PCI callback은 특별한 이유가 없으면 resume하지 않고 late/noirq phase에서 그대로 runtime-suspended면 일찍 반환합니다. Resume_noirq·restore_noirq에서 runtime-suspended면 곧 D0로 갈 것이므로 runtime PM status를 active로 바꿉니다.

`DPM_FLAG_MAY_SKIP_RESUME`는 working state로 돌아온 뒤에도 suspend 상태를 유지할 수 있으면 noirq·early resume callback 생략을 허용합니다. PM core는 `power.may_skip_resume`과 함께 판단하며 `dev_pm_skip_resume()`이 true이면 PCI resume callback이 device를 건드리지 않고 반환합니다.

DPM driver flags
Flag효과
NO_DIRECT_COMPLETEDirect-complete 영구 금지, ancestor 영향
SMART_PREPAREprepare 결과로 동적 direct-complete opt-out
SMART_SUSPENDSystem suspend 중 runtime-suspended 유지 허용
MAY_SKIP_RESUME조건 충족 시 noirq·early resume 생략 허용

Direct-complete와 runtime-suspend 유지, resume 생략의 범위를 구분합니다.

3.1.19. Driver Flags for Power Management
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

The PM core allows device drivers to set flags that influence the handling of
power management for the devices by the core itself and by middle layer code
including the PCI bus type.  The flags should be set once at the driver probe
time with the help of the dev_pm_set_driver_flags() function and they should not
be updated directly afterwards.

The DPM_FLAG_NO_DIRECT_COMPLETE flag prevents the PM core from using the
direct-complete mechanism allowing device suspend/resume callbacks to be skipped
if the device is in runtime suspend when the system suspend starts.  That also
affects all of the ancestors of the device, so this flag should only be used if
absolutely necessary.

The DPM_FLAG_SMART_PREPARE flag causes the PCI bus type to return a positive
value from pci_pm_prepare() only if the ->prepare callback provided by the
driver of the device returns a positive value.  That allows the driver to opt
out from using the direct-complete mechanism dynamically (whereas setting
DPM_FLAG_NO_DIRECT_COMPLETE means permanent opt-out).

The DPM_FLAG_SMART_SUSPEND flag tells the PCI bus type that from the driver's
perspective the device can be safely left in runtime suspend during system
suspend.  That causes pci_pm_suspend(), pci_pm_freeze() and pci_pm_poweroff()
to avoid resuming the device from runtime suspend unless there are PCI-specific
reasons for doing that.  Also, it causes pci_pm_suspend_late/noirq() and
pci_pm_poweroff_late/noirq() to return early if the device remains in runtime
suspend during the "late" phase of the system-wide transition under way.
Moreover, if the device is in runtime suspend in pci_pm_resume_noirq() or
pci_pm_restore_noirq(), its runtime PM status will be changed to "active" (as it
is going to be put into D0 going forward).

Setting the DPM_FLAG_MAY_SKIP_RESUME flag means that the driver allows its
"noirq" and "early" resume callbacks to be skipped if the device can be left
in suspend after a system-wide transition into the working state.  This flag is
taken into consideration by the PM core along with the power.may_skip_resume
status bit of the device which is set by pci_pm_suspend_noirq() in certain
situations.  If the PM core determines that the driver's "noirq" and "early"
resume callbacks should be skipped, the dev_pm_skip_resume() helper function
will return "true" and that will cause pci_pm_resume_noirq() and
pci_pm_resume_early() to return upfront without touching the device and
executing the driver callbacks.

Driver의 runtime PM 책임

1040-1118

PCI runtime PM은 선택 사항이지만 Ethernet cable 분리나 USB child 부재처럼 device 미사용을 신뢰성 있게 확인할 수 있으면 구현을 권장합니다.

Driver는 `runtime_suspend()`와 `runtime_resume()`을 구현해야 합니다. Resume 직후 매번 다시 suspend되는 것을 막으려면 `runtime_idle()`도 구현하거나 `runtime_suspend()`가 실제 suspend 불가 시 `-EAGAIN`을 반환해야 합니다.

PCI core는 runtime PM을 기본 enable하지만 `pci_pm_init()`이 `pm_runtime_forbid()`로 block합니다. 또 `local_pci_probe()`가 driver probe 전에 usage counter를 증가시킵니다.

Runtime PM을 쓰려는 driver는 probe에서 usage counter를 감소시켜야 합니다. 가장 간단한 함수는 `pm_runtime_put_noidle()`이며 즉시 autosuspend를 예약하려면 `pm_runtime_put_autosuspend()`를 사용할 수 있습니다. 그렇지 않으면 counter가 0이 되지 않아 절대 runtime suspend되지 않습니다.

Counter를 내린 직후 userspace가 sysfs로 이미 `pm_runtime_allow()`를 호출했을 수 있어 `runtime_suspend()`가 즉시 실행될 수 있습니다. Driver는 직접 `pm_runtime_allow()`를 호출하지 말고 userspace나 platform code에 맡기되 언제 호출되어도 올바르게 처리해야 합니다.

Remove에서는 probe의 decrement를 맞추기 위해 `pm_runtime_get_noresume()`를 호출합니다. Core가 remove 전에 runtime resume와 counter 증가를 수행하므로 remove 동안 runtime PM은 사실상 disable되어 있고 counter 증가 helper는 같은 효과입니다.

Runtime PM request는 `pm_wq` workqueue의 work item으로 suspend, resume, idle check를 처리합니다. PM core가 일부 요청을 자동 queue하지만 일반적으로 driver가 `Documentation/power/runtime_pm.rst`의 helper로 요청합니다. 필요하면 `pm_wq` 없이 synchronous suspend·resume도 가능합니다.

PCI driver runtime PM 수명
local_pci_probe(): usage++driver probepm_runtime_put_noidle() / put_autosuspend()runtime PM availabledriver removepm_runtime_get_noresume()

Probe에서 core가 올린 counter를 내리고 remove에서 다시 맞춥니다.

Runtime PM 책임
주체책임
PCI coreRuntime PM enable, 초기 forbid, probe 전 usage++
Driver probeCallbacks 구현, usage--
Userspace/platformpm_runtime_allow policy
Driver runtimeHelper로 pm_wq request 또는 synchronous transition
Driver removeusage counter 균형 복원

Core enable 상태와 policy unblock, usage counter를 구분해야 합니다.

3.2. Device Runtime Power Management
------------------------------------

In addition to providing device power management callbacks PCI device drivers
are responsible for controlling the runtime power management (runtime PM) of
their devices.

The PCI device runtime PM is optional, but it is recommended that PCI device
drivers implement it at least in the cases where there is a reliable way of
verifying that the device is not used (like when the network cable is detached
from an Ethernet adapter or there are no devices attached to a USB controller).

To support the PCI runtime PM the driver first needs to implement the
runtime_suspend() and runtime_resume() callbacks.  It also may need to implement
the runtime_idle() callback to prevent the device from being suspended again
every time right after the runtime_resume() callback has returned
(alternatively, the runtime_suspend() callback will have to check if the
device should really be suspended and return -EAGAIN if that is not the case).

The runtime PM of PCI devices is enabled by default by the PCI core.  PCI
device drivers do not need to enable it and should not attempt to do so.
However, it is blocked by pci_pm_init() that runs the pm_runtime_forbid()
helper function.  In addition to that, the runtime PM usage counter of
each PCI device is incremented by local_pci_probe() before executing the
probe callback provided by the device's driver.

If a PCI driver implements the runtime PM callbacks and intends to use the
runtime PM framework provided by the PM core and the PCI subsystem, it needs
to decrement the device's runtime PM usage counter in its probe callback
function.  If it doesn't do that, the counter will always be different from
zero for the device and it will never be runtime-suspended.  The simplest
way to do that is by calling pm_runtime_put_noidle(), but if the driver
wants to schedule an autosuspend right away, for example, it may call
pm_runtime_put_autosuspend() instead for this purpose.  Generally, it
just needs to call a function that decrements the devices usage counter
from its probe routine to make runtime PM work for the device.

It is important to remember that the driver's runtime_suspend() callback
may be executed right after the usage counter has been decremented, because
user space may already have caused the pm_runtime_allow() helper function
unblocking the runtime PM of the device to run via sysfs, so the driver must
be prepared to cope with that.

The driver itself should not call pm_runtime_allow(), though.  Instead, it
should let user space or some platform-specific code do that (user space can
do it via sysfs as stated above), but it must be prepared to handle the
runtime PM of the device correctly as soon as pm_runtime_allow() is called
(which may happen at any time, even before the driver is loaded).

When the driver's remove callback runs, it has to balance the decrementation
of the device's runtime PM usage counter at the probe time.  For this reason,
if it has decremented the counter in its probe callback, it must run
pm_runtime_get_noresume() in its remove callback.  [Since the core carries
out a runtime resume of the device and bumps up the device's usage counter
before running the driver's remove callback, the runtime PM of the device
is effectively disabled for the duration of the remove execution and all
runtime PM helper functions incrementing the device's usage counter are
then effectively equivalent to pm_runtime_get_noresume().]

The runtime PM framework works by processing requests to suspend or resume
devices, or to check if they are idle (in which cases it is reasonable to
subsequently request that they be suspended).  These requests are represented
by work items put into the power management workqueue, pm_wq.  Although there
are a few situations in which power management requests are automatically
queued by the PM core (for example, after processing a request to resume a
device the PM core automatically queues a request to check if the device is
idle), device drivers are generally responsible for queuing power management
requests for their devices.  For this purpose they should use the runtime PM
helper functions provided by the PM core, discussed in
Documentation/power/runtime_pm.rst.

Devices can also be suspended and resumed synchronously, without placing a
request into pm_wq.  In the majority of cases this also is done by their
drivers that use helper functions provided by the PM core for this purpose.

For more information on the runtime PM of devices refer to
Documentation/power/runtime_pm.rst.

참고 자료

1119-1132

문서가 참조하는 규격은 PCI Local Bus Specification Rev. 3.0, PCI Bus Power Management Interface Specification Rev. 1.2, ACPI Specification Rev. 3.0b, PCI Express Base Specification Rev. 2.0입니다.

Kernel 내부 일반 PM 설명은 `Documentation/driver-api/pm/devices.rst`, runtime PM 상세는 `Documentation/power/runtime_pm.rst`를 참조합니다.

Resources
종류자료
PCIPCI Local Bus 3.0, PCI PM 1.2
PlatformACPI 3.0b
PCIePCI Express Base 2.0
Kerneldriver-api/pm/devices.rst, power/runtime_pm.rst

Hardware specification과 Linux PM 문서를 함께 확인합니다.

4. Resources
============

PCI Local Bus Specification, Rev. 3.0

PCI Bus Power Management Interface Specification, Rev. 1.2

Advanced Configuration and Power Interface (ACPI) Specification, Rev. 3.0b

PCI Express Base Specification, Rev. 2.0

Documentation/driver-api/pm/devices.rst

Documentation/power/runtime_pm.rst