← Documents Documentation/watchdog/watchdog-kernel-api.rst GitHub 원문 ↗

Linux 6.18.37 · Watchdog / Kernel API

The Linux WatchDog Timer Driver Core kernel API

watchdog_device, watchdog_ops, 상태 비트와 core helper의 계약을 설명합니다.

Source pathDocumentation/watchdog/watchdog-kernel-api.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

watchdog-kernel-api.rst:1-350

Watchdog Driver Core를 사용하는 커널 드라이버가 구현·설정해야 할 구조체, operation, 상태와 helper를 정리합니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 ===============================================
2 The Linux WatchDog Timer Driver Core kernel API
3 ===============================================
4
5 Last reviewed: 12-Feb-2013
6
7 Wim Van Sebroeck <[email protected]>
8
9 Introduction
10 ------------
11 This document does not describe what a WatchDog Timer (WDT) Driver or Device is.
12 It also does not describe the API which can be used by user space to communicate
13 with a WatchDog Timer. If you want to know this then please read the following
14 file: Documentation/watchdog/watchdog-api.rst .
15
16 So what does this document describe? It describes the API that can be used by
17 WatchDog Timer Drivers that want to use the WatchDog Timer Driver Core
18 Framework. This framework provides all interfacing towards user space so that
19 the same code does not have to be reproduced each time. This also means that
20 a watchdog timer driver then only needs to provide the different routines
21 (operations) that control the watchdog timer (WDT).
22
23 The API
24 -------
25 Each watchdog timer driver that wants to use the WatchDog Timer Driver Core
26 must #include <linux/watchdog.h> (you would have to do this anyway when
27 writing a watchdog device driver). This include file contains following
28 register/unregister routines::
29
30 extern int watchdog_register_device(struct watchdog_device *);
31 extern void watchdog_unregister_device(struct watchdog_device *);
32
33 The watchdog_register_device routine registers a watchdog timer device.
34 The parameter of this routine is a pointer to a watchdog_device structure.
35 This routine returns zero on success and a negative errno code for failure.
36
37 The watchdog_unregister_device routine deregisters a registered watchdog timer
38 device. The parameter of this routine is the pointer to the registered
39 watchdog_device structure.
40
41 The watchdog subsystem includes an registration deferral mechanism,
42 which allows you to register an watchdog as early as you wish during
43 the boot process.
44
45 The watchdog device structure looks like this::
46
47 struct watchdog_device {
48 int id;
49 struct device *parent;
50 const struct attribute_group **groups;
51 const struct watchdog_info *info;
52 const struct watchdog_ops *ops;
53 const struct watchdog_governor *gov;
54 unsigned int bootstatus;
55 unsigned int timeout;
56 unsigned int pretimeout;
57 unsigned int min_timeout;
58 unsigned int max_timeout;
59 unsigned int min_hw_heartbeat_ms;
60 unsigned int max_hw_heartbeat_ms;
61 struct notifier_block reboot_nb;
62 struct notifier_block restart_nb;
63 void *driver_data;
64 struct watchdog_core_data *wd_data;
65 unsigned long status;
66 struct list_head deferred;
67 };
68
69 It contains following fields:
70
71 * id: set by watchdog_register_device, id 0 is special. It has both a
72 /dev/watchdog0 cdev (dynamic major, minor 0) as well as the old
73 /dev/watchdog miscdev. The id is set automatically when calling
74 watchdog_register_device.
75 * parent: set this to the parent device (or NULL) before calling
76 watchdog_register_device.
77 * groups: List of sysfs attribute groups to create when creating the watchdog
78 device.
79 * info: a pointer to a watchdog_info structure. This structure gives some
80 additional information about the watchdog timer itself. (Like its unique name)
81 * ops: a pointer to the list of watchdog operations that the watchdog supports.
82 * gov: a pointer to the assigned watchdog device pretimeout governor or NULL.
83 * timeout: the watchdog timer's timeout value (in seconds).
84 This is the time after which the system will reboot if user space does
85 not send a heartbeat request if WDOG_ACTIVE is set.
86 * pretimeout: the watchdog timer's pretimeout value (in seconds).
87 * min_timeout: the watchdog timer's minimum timeout value (in seconds).
88 If set, the minimum configurable value for 'timeout'.
89 * max_timeout: the watchdog timer's maximum timeout value (in seconds),
90 as seen from userspace. If set, the maximum configurable value for
91 'timeout'. Not used if max_hw_heartbeat_ms is non-zero.
92 * min_hw_heartbeat_ms: Hardware limit for minimum time between heartbeats,
93 in milli-seconds. This value is normally 0; it should only be provided
94 if the hardware can not tolerate lower intervals between heartbeats.
95 * max_hw_heartbeat_ms: Maximum hardware heartbeat, in milli-seconds.
96 If set, the infrastructure will send heartbeats to the watchdog driver
97 if 'timeout' is larger than max_hw_heartbeat_ms, unless WDOG_ACTIVE
98 is set and userspace failed to send a heartbeat for at least 'timeout'
99 seconds. max_hw_heartbeat_ms must be set if a driver does not implement
100 the stop function.
101 * reboot_nb: notifier block that is registered for reboot notifications, for
102 internal use only. If the driver calls watchdog_stop_on_reboot, watchdog core
103 will stop the watchdog on such notifications.
104 * restart_nb: notifier block that is registered for machine restart, for
105 internal use only. If a watchdog is capable of restarting the machine, it
106 should define ops->restart. Priority can be changed through
107 watchdog_set_restart_priority.
108 * bootstatus: status of the device after booting (reported with watchdog
109 WDIOF_* status bits).
110 * driver_data: a pointer to the drivers private data of a watchdog device.
111 This data should only be accessed via the watchdog_set_drvdata and
112 watchdog_get_drvdata routines.
113 * wd_data: a pointer to watchdog core internal data.
114 * status: this field contains a number of status bits that give extra
115 information about the status of the device (Like: is the watchdog timer
116 running/active, or is the nowayout bit set).
117 * deferred: entry in wtd_deferred_reg_list which is used to
118 register early initialized watchdogs.
119
120 The list of watchdog operations is defined as::
121
122 struct watchdog_ops {
123 struct module *owner;
124 /* mandatory operations */
125 int (*start)(struct watchdog_device *);
126 /* optional operations */
127 int (*stop)(struct watchdog_device *);
128 int (*ping)(struct watchdog_device *);
129 unsigned int (*status)(struct watchdog_device *);
130 int (*set_timeout)(struct watchdog_device *, unsigned int);
131 int (*set_pretimeout)(struct watchdog_device *, unsigned int);
132 unsigned int (*get_timeleft)(struct watchdog_device *);
133 int (*restart)(struct watchdog_device *);
134 long (*ioctl)(struct watchdog_device *, unsigned int, unsigned long);
135 };
136
137 It is important that you first define the module owner of the watchdog timer
138 driver's operations. This module owner will be used to lock the module when
139 the watchdog is active. (This to avoid a system crash when you unload the
140 module and /dev/watchdog is still open).
141
142 Some operations are mandatory and some are optional. The mandatory operations
143 are:
144
145 * start: this is a pointer to the routine that starts the watchdog timer
146 device.
147 The routine needs a pointer to the watchdog timer device structure as a
148 parameter. It returns zero on success or a negative errno code for failure.
149
150 Not all watchdog timer hardware supports the same functionality. That's why
151 all other routines/operations are optional. They only need to be provided if
152 they are supported. These optional routines/operations are:
153
154 * stop: with this routine the watchdog timer device is being stopped.
155
156 The routine needs a pointer to the watchdog timer device structure as a
157 parameter. It returns zero on success or a negative errno code for failure.
158 Some watchdog timer hardware can only be started and not be stopped. A
159 driver supporting such hardware does not have to implement the stop routine.
160
161 If a driver has no stop function, the watchdog core will set WDOG_HW_RUNNING
162 and start calling the driver's keepalive pings function after the watchdog
163 device is closed.
164
165 If a watchdog driver does not implement the stop function, it must set
166 max_hw_heartbeat_ms.
167 * ping: this is the routine that sends a keepalive ping to the watchdog timer
168 hardware.
169
170 The routine needs a pointer to the watchdog timer device structure as a
171 parameter. It returns zero on success or a negative errno code for failure.
172
173 Most hardware that does not support this as a separate function uses the
174 start function to restart the watchdog timer hardware. And that's also what
175 the watchdog timer driver core does: to send a keepalive ping to the watchdog
176 timer hardware it will either use the ping operation (when available) or the
177 start operation (when the ping operation is not available).
178
179 (Note: the WDIOC_KEEPALIVE ioctl call will only be active when the
180 WDIOF_KEEPALIVEPING bit has been set in the option field on the watchdog's
181 info structure).
182 * status: this routine checks the status of the watchdog timer device. The
183 status of the device is reported with watchdog WDIOF_* status flags/bits.
184
185 WDIOF_MAGICCLOSE and WDIOF_KEEPALIVEPING are reported by the watchdog core;
186 it is not necessary to report those bits from the driver. Also, if no status
187 function is provided by the driver, the watchdog core reports the status bits
188 provided in the bootstatus variable of struct watchdog_device.
189
190 * set_timeout: this routine checks and changes the timeout of the watchdog
191 timer device. It returns 0 on success, -EINVAL for "parameter out of range"
192 and -EIO for "could not write value to the watchdog". On success this
193 routine should set the timeout value of the watchdog_device to the
194 achieved timeout value (which may be different from the requested one
195 because the watchdog does not necessarily have a 1 second resolution).
196
197 Drivers implementing max_hw_heartbeat_ms set the hardware watchdog heartbeat
198 to the minimum of timeout and max_hw_heartbeat_ms. Those drivers set the
199 timeout value of the watchdog_device either to the requested timeout value
200 (if it is larger than max_hw_heartbeat_ms), or to the achieved timeout value.
201 (Note: the WDIOF_SETTIMEOUT needs to be set in the options field of the
202 watchdog's info structure).
203
204 If the watchdog driver does not have to perform any action but setting the
205 watchdog_device.timeout, this callback can be omitted.
206
207 If set_timeout is not provided but, WDIOF_SETTIMEOUT is set, the watchdog
208 infrastructure updates the timeout value of the watchdog_device internally
209 to the requested value.
210
211 If the pretimeout feature is used (WDIOF_PRETIMEOUT), then set_timeout must
212 also take care of checking if pretimeout is still valid and set up the timer
213 accordingly. This can't be done in the core without races, so it is the
214 duty of the driver.
215 * set_pretimeout: this routine checks and changes the pretimeout value of
216 the watchdog. It is optional because not all watchdogs support pretimeout
217 notification. The timeout value is not an absolute time, but the number of
218 seconds before the actual timeout would happen. It returns 0 on success,
219 -EINVAL for "parameter out of range" and -EIO for "could not write value to
220 the watchdog". A value of 0 disables pretimeout notification.
221
222 (Note: the WDIOF_PRETIMEOUT needs to be set in the options field of the
223 watchdog's info structure).
224
225 If the watchdog driver does not have to perform any action but setting the
226 watchdog_device.pretimeout, this callback can be omitted. That means if
227 set_pretimeout is not provided but WDIOF_PRETIMEOUT is set, the watchdog
228 infrastructure updates the pretimeout value of the watchdog_device internally
229 to the requested value.
230
231 * get_timeleft: this routines returns the time that's left before a reset.
232 * restart: this routine restarts the machine. It returns 0 on success or a
233 negative errno code for failure.
234 * ioctl: if this routine is present then it will be called first before we do
235 our own internal ioctl call handling. This routine should return -ENOIOCTLCMD
236 if a command is not supported. The parameters that are passed to the ioctl
237 call are: watchdog_device, cmd and arg.
238
239 The status bits should (preferably) be set with the set_bit and clear_bit alike
240 bit-operations. The status bits that are defined are:
241
242 * WDOG_ACTIVE: this status bit indicates whether or not a watchdog timer device
243 is active or not from user perspective. User space is expected to send
244 heartbeat requests to the driver while this flag is set.
245 * WDOG_NO_WAY_OUT: this bit stores the nowayout setting for the watchdog.
246 If this bit is set then the watchdog timer will not be able to stop.
247 * WDOG_HW_RUNNING: Set by the watchdog driver if the hardware watchdog is
248 running. The bit must be set if the watchdog timer hardware can not be
249 stopped. The bit may also be set if the watchdog timer is running after
250 booting, before the watchdog device is opened. If set, the watchdog
251 infrastructure will send keepalives to the watchdog hardware while
252 WDOG_ACTIVE is not set.
253 Note: when you register the watchdog timer device with this bit set,
254 then opening /dev/watchdog will skip the start operation but send a keepalive
255 request instead.
256
257 To set the WDOG_NO_WAY_OUT status bit (before registering your watchdog
258 timer device) you can either:
259
260 * set it statically in your watchdog_device struct with
261
262 .status = WATCHDOG_NOWAYOUT_INIT_STATUS,
263
264 (this will set the value the same as CONFIG_WATCHDOG_NOWAYOUT) or
265 * use the following helper function::
266
267 static inline void watchdog_set_nowayout(struct watchdog_device *wdd,
268 int nowayout)
269
270 Note:
271 The WatchDog Timer Driver Core supports the magic close feature and
272 the nowayout feature. To use the magic close feature you must set the
273 WDIOF_MAGICCLOSE bit in the options field of the watchdog's info structure.
274
275 The nowayout feature will overrule the magic close feature.
276
277 To get or set driver specific data the following two helper functions should be
278 used::
279
280 static inline void watchdog_set_drvdata(struct watchdog_device *wdd,
281 void *data)
282 static inline void *watchdog_get_drvdata(struct watchdog_device *wdd)
283
284 The watchdog_set_drvdata function allows you to add driver specific data. The
285 arguments of this function are the watchdog device where you want to add the
286 driver specific data to and a pointer to the data itself.
287
288 The watchdog_get_drvdata function allows you to retrieve driver specific data.
289 The argument of this function is the watchdog device where you want to retrieve
290 data from. The function returns the pointer to the driver specific data.
291
292 To initialize the timeout field, the following function can be used::
293
294 extern int watchdog_init_timeout(struct watchdog_device *wdd,
295 unsigned int timeout_parm,
296 struct device *dev);
297
298 The watchdog_init_timeout function allows you to initialize the timeout field
299 using the module timeout parameter or by retrieving the timeout-sec property from
300 the device tree (if the module timeout parameter is invalid). Best practice is
301 to set the default timeout value as timeout value in the watchdog_device and
302 then use this function to set the user "preferred" timeout value.
303 This routine returns zero on success and a negative errno code for failure.
304
305 To disable the watchdog on reboot, the user must call the following helper::
306
307 static inline void watchdog_stop_on_reboot(struct watchdog_device *wdd);
308
309 To disable the watchdog when unregistering the watchdog, the user must call
310 the following helper. Note that this will only stop the watchdog if the
311 nowayout flag is not set.
312
313 ::
314
315 static inline void watchdog_stop_on_unregister(struct watchdog_device *wdd);
316
317 To change the priority of the restart handler the following helper should be
318 used::
319
320 void watchdog_set_restart_priority(struct watchdog_device *wdd, int priority);
321
322 User should follow the following guidelines for setting the priority:
323
324 * 0: should be called in last resort, has limited restart capabilities
325 * 128: default restart handler, use if no other handler is expected to be
326 available, and/or if restart is sufficient to restart the entire system
327 * 255: highest priority, will preempt all other restart handlers
328
329 To raise a pretimeout notification, the following function should be used::
330
331 void watchdog_notify_pretimeout(struct watchdog_device *wdd)
332
333 The function can be called in the interrupt context. If watchdog pretimeout
334 governor framework (kbuild CONFIG_WATCHDOG_PRETIMEOUT_GOV symbol) is enabled,
335 an action is taken by a preconfigured pretimeout governor preassigned to
336 the watchdog device. If watchdog pretimeout governor framework is not
337 enabled, watchdog_notify_pretimeout() prints a notification message to
338 the kernel log buffer.
339
340 To set the last known HW keepalive time for a watchdog, the following function
341 should be used::
342
343 int watchdog_set_last_hw_keepalive(struct watchdog_device *wdd,
344 unsigned int last_ping_ms)
345
346 This function must be called immediately after watchdog registration. It
347 sets the last known hardware heartbeat to have happened last_ping_ms before
348 current time. Calling this is only needed if the watchdog is already running
349 when probe is called, and the watchdog can only be pinged after the
350 min_hw_heartbeat_ms time has passed from the last ping.
351

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

문서 정보

1-8

이 문서는 Linux WatchDog Timer Driver Core 커널 API를 설명합니다. 마지막 검토일은 2013-02-12이며 작성자는 Wim Van Sebroeck입니다.

===============================================
The Linux WatchDog Timer Driver Core kernel API
===============================================

Last reviewed: 12-Feb-2013

Wim Van Sebroeck <[email protected]>

공통 사용자 공간 인터페이스

9-22

이 문서는 WDT 드라이버나 장치 자체, 사용자 공간 통신 API를 설명하지 않습니다. 사용자 API는 `Documentation/watchdog/watchdog-api.rst`를 참조합니다.

대신 WatchDog Timer Driver Core Framework를 사용하는 WDT 드라이버용 API를 설명합니다. Core는 사용자 공간 인터페이스를 공통으로 제공해 드라이버마다 같은 코드를 반복하지 않게 합니다.

따라서 개별 watchdog 드라이버는 하드웨어 타이머를 제어하는 operation만 제공하면 됩니다.

Watchdog core 역할 분담
사용자 공간 /dev/watchdog·ioctlWatchdog Driver Core의 공통 처리watchdog_ops 호출장치 고유 WDT 하드웨어 제어

공통 파일 API와 장치 고유 제어를 분리합니다.

Introduction
------------
This document does not describe what a WatchDog Timer (WDT) Driver or Device is.
It also does not describe the API which can be used by user space to communicate
with a WatchDog Timer. If you want to know this then please read the following
file: Documentation/watchdog/watchdog-api.rst .

So what does this document describe? It describes the API that can be used by
WatchDog Timer Drivers that want to use the WatchDog Timer Driver Core
Framework. This framework provides all interfacing towards user space so that
the same code does not have to be reproduced each time. This also means that
a watchdog timer driver then only needs to provide the different routines
(operations) that control the watchdog timer (WDT).

등록·해제 API와 지연 등록

23-44

Core를 쓰는 드라이버는 `linux/watchdog.h`를 포함해야 합니다. 이 헤더는 `watchdog_register_device(struct watchdog_device *)`와 `watchdog_unregister_device(struct watchdog_device *)`를 선언합니다.

`watchdog_register_device()`는 watchdog_device 포인터로 장치를 등록하고 성공 시 0, 실패 시 음수 errno를 반환합니다.

`watchdog_unregister_device()`는 이미 등록된 watchdog_device 포인터로 장치를 해제합니다.

Watchdog 서브시스템에는 등록 지연 메커니즘이 있어 부팅 과정의 매우 이른 시점에도 watchdog 등록을 요청할 수 있습니다.

Watchdog 등록 API
함수동작반환
watchdog_register_device(wdd)watchdog timer 장치 등록0 또는 음수 errno
watchdog_unregister_device(wdd)등록 장치 해제void

장치 생명주기의 두 핵심 함수입니다.

The API
-------
Each watchdog timer driver that wants to use the WatchDog Timer Driver Core
must #include <linux/watchdog.h> (you would have to do this anyway when
writing a watchdog device driver). This include file contains following
register/unregister routines::

        extern int watchdog_register_device(struct watchdog_device *);
        extern void watchdog_unregister_device(struct watchdog_device *);

The watchdog_register_device routine registers a watchdog timer device.
The parameter of this routine is a pointer to a watchdog_device structure.
This routine returns zero on success and a negative errno code for failure.

The watchdog_unregister_device routine deregisters a registered watchdog timer
device. The parameter of this routine is the pointer to the registered
watchdog_device structure.

The watchdog subsystem includes an registration deferral mechanism,
which allows you to register an watchdog as early as you wish during
the boot process.

watchdog_device 구조체

45-68

`struct watchdog_device`는 식별자·부모·sysfs 그룹, 정보·operation·governor, 부팅 상태와 timeout, 하드웨어 heartbeat 범위, reboot·restart notifier, private data와 core 내부 데이터, 상태 비트, 지연 등록 목록 항목을 한곳에 모읍니다.

원문의 전체 C 구조체 선언과 멤버 순서는 아래 원문 블록에 그대로 보존되어 있습니다.

watchdog_device 필드군
영역멤버
식별·계층id, parent, groups
기능info, ops, gov
시간timeout, pretimeout, min/max_timeout
하드웨어 heartbeatmin/max_hw_heartbeat_ms
알림reboot_nb, restart_nb
데이터·상태driver_data, wd_data, status
등록deferred

구조체 멤버를 역할별로 묶은 개요입니다.

The watchdog device structure looks like this::

  struct watchdog_device {
        int id;
        struct device *parent;
        const struct attribute_group **groups;
        const struct watchdog_info *info;
        const struct watchdog_ops *ops;
        const struct watchdog_governor *gov;
        unsigned int bootstatus;
        unsigned int timeout;
        unsigned int pretimeout;
        unsigned int min_timeout;
        unsigned int max_timeout;
        unsigned int min_hw_heartbeat_ms;
        unsigned int max_hw_heartbeat_ms;
        struct notifier_block reboot_nb;
        struct notifier_block restart_nb;
        void *driver_data;
        struct watchdog_core_data *wd_data;
        unsigned long status;
        struct list_head deferred;
  };

watchdog_device 멤버 의미

69-119

`id`는 등록 함수가 자동 설정합니다. ID 0은 동적 major·minor 0의 `/dev/watchdog0` cdev와 기존 `/dev/watchdog` miscdev를 모두 갖는 특별한 장치입니다. `parent`는 등록 전에 부모 장치 또는 NULL로 정하고 `groups`는 장치 생성 때 만들 sysfs attribute group 목록입니다.

`info`는 고유 이름 같은 추가 정보를 담은 `watchdog_info`, `ops`는 지원 operation 목록, `gov`는 배정된 pretimeout governor 또는 NULL입니다.

`timeout`은 초 단위 최종 timeout입니다. `WDOG_ACTIVE` 상태에서 사용자 공간이 heartbeat를 보내지 않으면 이 시간이 지난 뒤 시스템이 reboot합니다. `pretimeout`은 선행 알림 시간, `min_timeout`과 `max_timeout`은 사용자 공간에서 설정할 수 있는 최소·최대 초입니다. `max_hw_heartbeat_ms`가 0이 아니면 `max_timeout`은 사용하지 않습니다.

`min_hw_heartbeat_ms`는 heartbeat 사이의 최소 하드웨어 간격입니다. 보통 0이고 더 짧은 간격을 견디지 못하는 하드웨어에서만 설정합니다.

`max_hw_heartbeat_ms`는 하드웨어 최대 heartbeat 간격입니다. 사용자 timeout이 이 값보다 크면 core가 드라이버에 중간 heartbeat를 보냅니다. 단 `WDOG_ACTIVE`이고 사용자 공간이 timeout 이상 heartbeat를 보내지 않았다면 더 이상 연장하지 않습니다. `stop`을 구현하지 않은 드라이버는 이 값을 반드시 설정해야 합니다.

`reboot_nb`는 내부 reboot notifier입니다. 드라이버가 `watchdog_stop_on_reboot()`를 호출하면 core가 reboot 알림 때 watchdog을 멈춥니다. `restart_nb`는 machine restart notifier이고 기계를 restart할 수 있는 watchdog은 `ops->restart`를 정의합니다. `watchdog_set_restart_priority()`로 우선순위를 바꿀 수 있습니다.

`bootstatus`는 부팅 뒤 장치 상태를 `WDIOF_*` 비트로 기록합니다. `driver_data`는 드라이버 private data이며 반드시 `watchdog_set_drvdata()`와 `watchdog_get_drvdata()`로 접근합니다. `wd_data`는 core 내부 데이터 포인터입니다.

`status`는 활성 상태·nowayout 같은 추가 상태 비트를 담고, `deferred`는 일찍 초기화한 watchdog을 등록하는 `wtd_deferred_reg_list` 항목입니다.

watchdog_device 시간 필드
필드단위의미
timeout사용자 heartbeat 부재 뒤 reboot 시간
pretimeout최종 timeout 전 선행 알림
min_timeout사용자 설정 최소값
max_timeout사용자 설정 최대값
min_hw_heartbeat_msmsheartbeat 최소 간격
max_hw_heartbeat_msms하드웨어 최대 heartbeat 간격

사용자 timeout과 하드웨어 heartbeat 제한의 관계입니다.

긴 사용자 timeout 유지
사용자 timeout > max_hw_heartbeat_msCore가 max_hw_heartbeat_ms 안에 드라이버 heartbeat 전송사용자 heartbeat가 오면 계속 유지WDOG_ACTIVE 상태에서 사용자 timeout 초과 시 연장 중단

하드웨어 최대 간격보다 큰 timeout을 core가 중간 ping으로 유지합니다.

It contains following fields:

* id: set by watchdog_register_device, id 0 is special. It has both a
  /dev/watchdog0 cdev (dynamic major, minor 0) as well as the old
  /dev/watchdog miscdev. The id is set automatically when calling
  watchdog_register_device.
* parent: set this to the parent device (or NULL) before calling
  watchdog_register_device.
* groups: List of sysfs attribute groups to create when creating the watchdog
  device.
* info: a pointer to a watchdog_info structure. This structure gives some
  additional information about the watchdog timer itself. (Like its unique name)
* ops: a pointer to the list of watchdog operations that the watchdog supports.
* gov: a pointer to the assigned watchdog device pretimeout governor or NULL.
* timeout: the watchdog timer's timeout value (in seconds).
  This is the time after which the system will reboot if user space does
  not send a heartbeat request if WDOG_ACTIVE is set.
* pretimeout: the watchdog timer's pretimeout value (in seconds).
* min_timeout: the watchdog timer's minimum timeout value (in seconds).
  If set, the minimum configurable value for 'timeout'.
* max_timeout: the watchdog timer's maximum timeout value (in seconds),
  as seen from userspace. If set, the maximum configurable value for
  'timeout'. Not used if max_hw_heartbeat_ms is non-zero.
* min_hw_heartbeat_ms: Hardware limit for minimum time between heartbeats,
  in milli-seconds. This value is normally 0; it should only be provided
  if the hardware can not tolerate lower intervals between heartbeats.
* max_hw_heartbeat_ms: Maximum hardware heartbeat, in milli-seconds.
  If set, the infrastructure will send heartbeats to the watchdog driver
  if 'timeout' is larger than max_hw_heartbeat_ms, unless WDOG_ACTIVE
  is set and userspace failed to send a heartbeat for at least 'timeout'
  seconds. max_hw_heartbeat_ms must be set if a driver does not implement
  the stop function.
* reboot_nb: notifier block that is registered for reboot notifications, for
  internal use only. If the driver calls watchdog_stop_on_reboot, watchdog core
  will stop the watchdog on such notifications.
* restart_nb: notifier block that is registered for machine restart, for
  internal use only. If a watchdog is capable of restarting the machine, it
  should define ops->restart. Priority can be changed through
  watchdog_set_restart_priority.
* bootstatus: status of the device after booting (reported with watchdog
  WDIOF_* status bits).
* driver_data: a pointer to the drivers private data of a watchdog device.
  This data should only be accessed via the watchdog_set_drvdata and
  watchdog_get_drvdata routines.
* wd_data: a pointer to watchdog core internal data.
* status: this field contains a number of status bits that give extra
  information about the status of the device (Like: is the watchdog timer
  running/active, or is the nowayout bit set).
* deferred: entry in wtd_deferred_reg_list which is used to
  register early initialized watchdogs.

watchdog_ops와 필수 start

120-149

`struct watchdog_ops`는 module owner, 필수 `start`, 선택적인 `stop`, `ping`, `status`, `set_timeout`, `set_pretimeout`, `get_timeleft`, `restart`, `ioctl` 콜백을 정의합니다.

먼저 `owner`를 설정해야 합니다. Watchdog이 활성인 동안 module을 잠가 `/dev/watchdog`이 열린 상태에서 module을 unload해 system crash가 나는 일을 막습니다.

유일한 필수 operation인 `start`는 watchdog timer를 시작합니다. watchdog_device 포인터를 받고 성공 시 0, 실패 시 음수 errno를 반환합니다.

watchdog_ops
Operation필수역할
owner필수활성 중 module 고정
start필수watchdog 시작
stop선택watchdog 정지
ping선택keepalive
status선택WDIOF_* 상태
set_timeout선택timeout 설정
set_pretimeout선택pretimeout 설정
get_timeleft선택reset까지 남은 시간
restart선택기계 restart
ioctl선택장치 고유 ioctl

필수 여부와 함수 역할입니다.

The list of watchdog operations is defined as::

  struct watchdog_ops {
        struct module *owner;
        /* mandatory operations */
        int (*start)(struct watchdog_device *);
        /* optional operations */
        int (*stop)(struct watchdog_device *);
        int (*ping)(struct watchdog_device *);
        unsigned int (*status)(struct watchdog_device *);
        int (*set_timeout)(struct watchdog_device *, unsigned int);
        int (*set_pretimeout)(struct watchdog_device *, unsigned int);
        unsigned int (*get_timeleft)(struct watchdog_device *);
        int (*restart)(struct watchdog_device *);
        long (*ioctl)(struct watchdog_device *, unsigned int, unsigned long);
  };

It is important that you first define the module owner of the watchdog timer
driver's operations. This module owner will be used to lock the module when
the watchdog is active. (This to avoid a system crash when you unload the
module and /dev/watchdog is still open).

Some operations are mandatory and some are optional. The mandatory operations
are:

* start: this is a pointer to the routine that starts the watchdog timer
  device.
  The routine needs a pointer to the watchdog timer device structure as a
  parameter. It returns zero on success or a negative errno code for failure.

선택 operation 세부 규칙

150-238

`stop`은 watchdog을 멈추고 wdd 포인터를 받아 0 또는 음수 errno를 반환합니다. 시작만 가능하고 정지는 불가능한 하드웨어는 구현하지 않아도 됩니다. Stop이 없으면 core는 `WDOG_HW_RUNNING`을 설정하고 장치를 닫은 뒤에도 keepalive ping을 호출합니다. 이 경우 `max_hw_heartbeat_ms` 설정이 필수입니다.

`ping`은 하드웨어에 keepalive를 보냅니다. 별도 ping이 없는 하드웨어는 `start`로 timer를 다시 시작하며 core도 ping이 있으면 ping을, 없으면 start를 사용합니다. `WDIOC_KEEPALIVE`는 watchdog_info option에 `WDIOF_KEEPALIVEPING`이 있어야 활성화됩니다.

`status`는 장치 상태를 `WDIOF_*` 비트로 보고합니다. `WDIOF_MAGICCLOSE`와 `WDIOF_KEEPALIVEPING`은 core가 보고하므로 드라이버가 포함할 필요가 없습니다. Status 콜백이 없으면 core는 watchdog_device의 `bootstatus` 비트를 보고합니다.

`set_timeout`은 timeout을 검사·변경하고 성공 0, 범위 오류 `-EINVAL`, 하드웨어 기록 실패 `-EIO`를 반환합니다. 하드웨어 해상도 때문에 실제값이 요청과 다를 수 있으므로 성공 시 wdd의 timeout을 달성값으로 갱신합니다.

`max_hw_heartbeat_ms`를 쓰는 드라이버는 하드웨어 heartbeat를 `min(timeout, max_hw_heartbeat_ms)`로 설정합니다. 요청 timeout이 하드웨어 최대보다 크면 wdd timeout은 요청값으로, 그렇지 않으면 달성값으로 설정합니다. watchdog_info에 `WDIOF_SETTIMEOUT`도 필요합니다.

하드웨어 작업 없이 wdd timeout 값만 바꾸면 set_timeout 콜백을 생략할 수 있습니다. `WDIOF_SETTIMEOUT`은 있지만 콜백이 없으면 infrastructure가 요청값으로 내부 갱신합니다.

`WDIOF_PRETIMEOUT`을 쓰면 set_timeout이 변경 뒤 pretimeout이 여전히 유효한지 검사하고 timer를 다시 설정해야 합니다. Race 없이 core에서 처리할 수 없으므로 드라이버 책임입니다.

`set_pretimeout`은 최종 timeout 몇 초 전 알림을 검사·변경하며 0은 비활성화입니다. 성공 0, 범위 오류 `-EINVAL`, 기록 실패 `-EIO`를 반환하고 watchdog_info에 `WDIOF_PRETIMEOUT`이 필요합니다. 단순 wdd 값 갱신만 필요하면 콜백을 생략하고 core가 처리할 수 있습니다.

`get_timeleft`는 reset까지 남은 시간을 반환하고, `restart`는 기계를 restart해 0 또는 음수 errno를 반환합니다.

`ioctl`이 있으면 core 내부 ioctl 처리보다 먼저 호출합니다. 지원하지 않는 command는 `-ENOIOCTLCMD`를 반환해야 core 처리가 이어집니다. 인수는 watchdog_device, cmd, arg입니다.

선택 operation 반환 규칙
Operation성공대표 오류·대체
stop0음수 errno, 미구현 가능
ping0없으면 start 사용
set_timeout0-EINVAL / -EIO
set_pretimeout0-EINVAL / -EIO
restart0음수 errno
ioctl장치별미지원은 -ENOIOCTLCMD

주요 성공·오류 결과입니다.

Core keepalive 선택
Keepalive 요청ping 콜백이 있으면 ping(wdd)없으면 start(wdd)로 timer 재시작stop이 없으면 장치 close 뒤에도 core가 중간 ping

ping 콜백 유무와 stop 가능 여부에 따른 처리입니다.

Not all watchdog timer hardware supports the same functionality. That's why
all other routines/operations are optional. They only need to be provided if
they are supported. These optional routines/operations are:

* stop: with this routine the watchdog timer device is being stopped.

  The routine needs a pointer to the watchdog timer device structure as a
  parameter. It returns zero on success or a negative errno code for failure.
  Some watchdog timer hardware can only be started and not be stopped. A
  driver supporting such hardware does not have to implement the stop routine.

  If a driver has no stop function, the watchdog core will set WDOG_HW_RUNNING
  and start calling the driver's keepalive pings function after the watchdog
  device is closed.

  If a watchdog driver does not implement the stop function, it must set
  max_hw_heartbeat_ms.
* ping: this is the routine that sends a keepalive ping to the watchdog timer
  hardware.

  The routine needs a pointer to the watchdog timer device structure as a
  parameter. It returns zero on success or a negative errno code for failure.

  Most hardware that does not support this as a separate function uses the
  start function to restart the watchdog timer hardware. And that's also what
  the watchdog timer driver core does: to send a keepalive ping to the watchdog
  timer hardware it will either use the ping operation (when available) or the
  start operation (when the ping operation is not available).

  (Note: the WDIOC_KEEPALIVE ioctl call will only be active when the
  WDIOF_KEEPALIVEPING bit has been set in the option field on the watchdog's
  info structure).
* status: this routine checks the status of the watchdog timer device. The
  status of the device is reported with watchdog WDIOF_* status flags/bits.

  WDIOF_MAGICCLOSE and WDIOF_KEEPALIVEPING are reported by the watchdog core;
  it is not necessary to report those bits from the driver. Also, if no status
  function is provided by the driver, the watchdog core reports the status bits
  provided in the bootstatus variable of struct watchdog_device.

* set_timeout: this routine checks and changes the timeout of the watchdog
  timer device. It returns 0 on success, -EINVAL for "parameter out of range"
  and -EIO for "could not write value to the watchdog". On success this
  routine should set the timeout value of the watchdog_device to the
  achieved timeout value (which may be different from the requested one
  because the watchdog does not necessarily have a 1 second resolution).

  Drivers implementing max_hw_heartbeat_ms set the hardware watchdog heartbeat
  to the minimum of timeout and max_hw_heartbeat_ms. Those drivers set the
  timeout value of the watchdog_device either to the requested timeout value
  (if it is larger than max_hw_heartbeat_ms), or to the achieved timeout value.
  (Note: the WDIOF_SETTIMEOUT needs to be set in the options field of the
  watchdog's info structure).

  If the watchdog driver does not have to perform any action but setting the
  watchdog_device.timeout, this callback can be omitted.

  If set_timeout is not provided but, WDIOF_SETTIMEOUT is set, the watchdog
  infrastructure updates the timeout value of the watchdog_device internally
  to the requested value.

  If the pretimeout feature is used (WDIOF_PRETIMEOUT), then set_timeout must
  also take care of checking if pretimeout is still valid and set up the timer
  accordingly. This can't be done in the core without races, so it is the
  duty of the driver.
* set_pretimeout: this routine checks and changes the pretimeout value of
  the watchdog. It is optional because not all watchdogs support pretimeout
  notification. The timeout value is not an absolute time, but the number of
  seconds before the actual timeout would happen. It returns 0 on success,
  -EINVAL for "parameter out of range" and -EIO for "could not write value to
  the watchdog". A value of 0 disables pretimeout notification.

  (Note: the WDIOF_PRETIMEOUT needs to be set in the options field of the
  watchdog's info structure).

  If the watchdog driver does not have to perform any action but setting the
  watchdog_device.pretimeout, this callback can be omitted. That means if
  set_pretimeout is not provided but WDIOF_PRETIMEOUT is set, the watchdog
  infrastructure updates the pretimeout value of the watchdog_device internally
  to the requested value.

* get_timeleft: this routines returns the time that's left before a reset.
* restart: this routine restarts the machine. It returns 0 on success or a
  negative errno code for failure.
* ioctl: if this routine is present then it will be called first before we do
  our own internal ioctl call handling. This routine should return -ENOIOCTLCMD
  if a command is not supported. The parameters that are passed to the ioctl
  call are: watchdog_device, cmd and arg.

상태 비트와 nowayout

239-269

상태 비트는 가능하면 `set_bit`, `clear_bit` 계열 bit operation으로 변경합니다.

`WDOG_ACTIVE`는 사용자 관점에서 watchdog이 활성인지 나타내며 이 비트가 있는 동안 사용자 공간은 heartbeat를 보내야 합니다.

`WDOG_NO_WAY_OUT`은 nowayout 설정입니다. 비트가 있으면 watchdog timer를 멈출 수 없습니다.

`WDOG_HW_RUNNING`은 하드웨어 watchdog이 실제로 동작 중임을 뜻합니다. 멈출 수 없는 하드웨어는 반드시 설정하고, 장치를 열기 전 부팅 때부터 동작 중인 경우도 설정할 수 있습니다. 이 비트가 있고 `WDOG_ACTIVE`가 없으면 infrastructure가 하드웨어에 keepalive를 보냅니다.

`WDOG_HW_RUNNING` 상태로 등록한 장치를 `/dev/watchdog`에서 열면 start operation을 건너뛰고 keepalive를 보냅니다.

등록 전에 `WDOG_NO_WAY_OUT`을 설정하려면 watchdog_device의 `.status = WATCHDOG_NOWAYOUT_INIT_STATUS`로 `CONFIG_WATCHDOG_NOWAYOUT`과 같은 값을 쓰거나 `watchdog_set_nowayout(wdd, nowayout)` helper를 호출합니다.

Watchdog 상태 비트
비트의미
WDOG_ACTIVE사용자 공간 heartbeat가 필요한 활성 상태
WDOG_NO_WAY_OUTwatchdog 정지 불가
WDOG_HW_RUNNING하드웨어 timer가 실제 실행 중

Core가 추적하는 실행·정지 정책입니다.

The status bits should (preferably) be set with the set_bit and clear_bit alike
bit-operations. The status bits that are defined are:

* WDOG_ACTIVE: this status bit indicates whether or not a watchdog timer device
  is active or not from user perspective. User space is expected to send
  heartbeat requests to the driver while this flag is set.
* WDOG_NO_WAY_OUT: this bit stores the nowayout setting for the watchdog.
  If this bit is set then the watchdog timer will not be able to stop.
* WDOG_HW_RUNNING: Set by the watchdog driver if the hardware watchdog is
  running. The bit must be set if the watchdog timer hardware can not be
  stopped. The bit may also be set if the watchdog timer is running after
  booting, before the watchdog device is opened. If set, the watchdog
  infrastructure will send keepalives to the watchdog hardware while
  WDOG_ACTIVE is not set.
  Note: when you register the watchdog timer device with this bit set,
  then opening /dev/watchdog will skip the start operation but send a keepalive
  request instead.

  To set the WDOG_NO_WAY_OUT status bit (before registering your watchdog
  timer device) you can either:

  * set it statically in your watchdog_device struct with

        .status = WATCHDOG_NOWAYOUT_INIT_STATUS,

    (this will set the value the same as CONFIG_WATCHDOG_NOWAYOUT) or
  * use the following helper function::

        static inline void watchdog_set_nowayout(struct watchdog_device *wdd,
                                                 int nowayout)

Magic Close와 private data helper

270-291

Core는 Magic Close와 nowayout을 지원합니다. Magic Close를 쓰려면 watchdog_info options에 `WDIOF_MAGICCLOSE`를 설정합니다. Nowayout은 Magic Close보다 우선하므로 nowayout이면 magic close 조건을 만족해도 watchdog을 멈추지 않습니다.

드라이버별 private data는 `watchdog_set_drvdata(wdd, data)`로 저장하고 `watchdog_get_drvdata(wdd)`로 가져옵니다. 직접 `driver_data` 멤버를 다루지 않습니다.

Private data helper
함수역할
watchdog_set_drvdata(wdd, data)private data 저장
watchdog_get_drvdata(wdd)private data 포인터 조회

watchdog_device에 드라이버 데이터를 연결하는 API입니다.

Note:
   The WatchDog Timer Driver Core supports the magic close feature and
   the nowayout feature. To use the magic close feature you must set the
   WDIOF_MAGICCLOSE bit in the options field of the watchdog's info structure.

The nowayout feature will overrule the magic close feature.

To get or set driver specific data the following two helper functions should be
used::

  static inline void watchdog_set_drvdata(struct watchdog_device *wdd,
                                          void *data)
  static inline void *watchdog_get_drvdata(struct watchdog_device *wdd)

The watchdog_set_drvdata function allows you to add driver specific data. The
arguments of this function are the watchdog device where you want to add the
driver specific data to and a pointer to the data itself.

The watchdog_get_drvdata function allows you to retrieve driver specific data.
The argument of this function is the watchdog device where you want to retrieve
data from. The function returns the pointer to the driver specific data.

watchdog_init_timeout

292-304

`watchdog_init_timeout(wdd, timeout_parm, dev)`는 module timeout 매개변수로 wdd timeout을 초기화합니다. 매개변수가 유효하지 않으면 device tree의 `timeout-sec` property를 가져옵니다.

권장 방식은 watchdog_device에 기본 timeout을 먼저 넣고 이 함수로 사용자가 선호한 값을 적용하는 것입니다. 성공 시 0, 실패 시 음수 errno를 반환합니다.

Timeout 초기화 우선순위
watchdog_device에 드라이버 기본 timeout 설정유효한 module timeout_parm 확인유효하지 않으면 DT timeout-sec 조회선호값 적용 또는 기본값 유지

기본값 위에 사용자·firmware 설정을 적용합니다.

To initialize the timeout field, the following function can be used::

  extern int watchdog_init_timeout(struct watchdog_device *wdd,
                                   unsigned int timeout_parm,
                                   struct device *dev);

The watchdog_init_timeout function allows you to initialize the timeout field
using the module timeout parameter or by retrieving the timeout-sec property from
the device tree (if the module timeout parameter is invalid). Best practice is
to set the default timeout value as timeout value in the watchdog_device and
then use this function to set the user "preferred" timeout value.
This routine returns zero on success and a negative errno code for failure.

reboot·unregister 정지 helper

305-316

Reboot 때 watchdog을 끄려면 `watchdog_stop_on_reboot(wdd)`를 호출합니다.

장치 unregister 때 watchdog을 끄려면 `watchdog_stop_on_unregister(wdd)`를 호출합니다. 단 nowayout flag가 설정되어 있으면 이 helper도 watchdog을 멈추지 않습니다.

Watchdog 정지 helper
Helper시점nowayout
watchdog_stop_on_reboot시스템 reboot 알림드라이버 정책 적용
watchdog_stop_on_unregister장치 unregister설정 시 정지하지 않음

정지 시점과 nowayout 영향입니다.

To disable the watchdog on reboot, the user must call the following helper::

  static inline void watchdog_stop_on_reboot(struct watchdog_device *wdd);

To disable the watchdog when unregistering the watchdog, the user must call
the following helper. Note that this will only stop the watchdog if the
nowayout flag is not set.

::

  static inline void watchdog_stop_on_unregister(struct watchdog_device *wdd);

restart handler 우선순위

317-328

`watchdog_set_restart_priority(wdd, priority)`로 restart handler 우선순위를 바꿉니다.

우선순위 0은 제한된 restart 기능을 마지막 수단으로 호출할 때, 128은 다른 handler가 없을 것으로 예상하거나 전체 시스템 restart에 충분한 기본 handler에, 255는 다른 모든 handler를 선점해야 하는 최고 우선순위에 사용합니다.

Restart priority
의미
0마지막 수단, 제한된 restart
128기본 restart handler
255최고 우선순위, 다른 handler 선점

문서가 권장하는 세 기준값입니다.

To change the priority of the restart handler the following helper should be
used::

  void watchdog_set_restart_priority(struct watchdog_device *wdd, int priority);

User should follow the following guidelines for setting the priority:

* 0: should be called in last resort, has limited restart capabilities
* 128: default restart handler, use if no other handler is expected to be
  available, and/or if restart is sufficient to restart the entire system
* 255: highest priority, will preempt all other restart handlers

pretimeout 알림

329-339

Pretimeout을 알리려면 interrupt context에서도 호출 가능한 `watchdog_notify_pretimeout(wdd)`를 사용합니다.

`CONFIG_WATCHDOG_PRETIMEOUT_GOV`가 켜져 있으면 장치에 미리 배정한 pretimeout governor가 구성된 동작을 수행합니다. Governor framework가 없으면 kernel log buffer에 알림 메시지를 출력합니다.

Pretimeout 알림
하드웨어 pretimeout interruptwatchdog_notify_pretimeout(wdd)Pretimeout governor framework 확인있으면 지정 governor 동작없으면 kernel log 알림

빌드 설정에 따라 governor 또는 로그 경로로 나뉩니다.

To raise a pretimeout notification, the following function should be used::

  void watchdog_notify_pretimeout(struct watchdog_device *wdd)

The function can be called in the interrupt context. If watchdog pretimeout
governor framework (kbuild CONFIG_WATCHDOG_PRETIMEOUT_GOV symbol) is enabled,
an action is taken by a preconfigured pretimeout governor preassigned to
the watchdog device. If watchdog pretimeout governor framework is not
enabled, watchdog_notify_pretimeout() prints a notification message to
the kernel log buffer.

마지막 하드웨어 heartbeat 기록

340-350

`watchdog_set_last_hw_keepalive(wdd, last_ping_ms)`는 마지막으로 알려진 하드웨어 heartbeat가 현재 시각보다 `last_ping_ms` 전에 발생했다고 기록합니다.

Watchdog 등록 직후 즉시 호출해야 합니다. Probe 시 이미 watchdog이 실행 중이고 마지막 ping 뒤 `min_hw_heartbeat_ms`가 지나야 다음 ping이 가능한 경우에만 필요합니다.

To set the last known HW keepalive time for a watchdog, the following function
should be used::

  int watchdog_set_last_hw_keepalive(struct watchdog_device *wdd,
                                     unsigned int last_ping_ms)

This function must be called immediately after watchdog registration. It
sets the last known hardware heartbeat to have happened last_ping_ms before
current time. Calling this is only needed if the watchdog is already running
when probe is called, and the watchdog can only be pinged after the
min_hw_heartbeat_ms time has passed from the last ping.