← Documents Documentation/admin-guide/sysctl/kernel.rst GitHub 원문 ↗

Linux 6.18.37 · Administration

Documentation for /proc/sys/kernel/

`/proc/sys/kernel`의 계정·core dump·추적·lockup·panic·perf·printk·IPC·보안·watchdog 정책과 한도를 설명합니다.

Source pathDocumentation/admin-guide/sysctl/kernel.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

kernel.rst:1-1718

이 문서는 Linux 커널의 전역 동작을 런타임에 바꾸는 `kernel` sysctl을 다룹니다. 값 하나가 장애 진단, 보안 경계, 재부팅 정책, 성능 계측과 IPC 자원에 직접 영향을 줄 수 있으므로 각 항목의 단위와 되돌릴 수 있는지 여부를 함께 확인해야 합니다.

영역대표 항목
진단·복구`core_pattern`, `ftrace_dump_on_oops`, `panic*`, `printk*`
lockup 감시`hardlockup_*`, `softlockup_*`, `hung_task_*`, `watchdog*`
권한·보안`dmesg_restrict`, `kptr_restrict`, `modules_disabled`, `unprivileged_bpf_disabled`
성능·스케줄링`perf_*`, `numa_balancing*`, `sched_*`, `timer_migration`
IPC·프로세스`msg*`, `shm*`, `pid_max`, `threads-max`

`kexec_load_disabled`, `modules_disabled`, `unprivileged_bpf_disabled=1`처럼 한 번 강화하면 실행 중 되돌릴 수 없는 설정이 있습니다. 운영 변경 전에는 문서의 기본값, 커널 구성 조건, capability 요구 사항과 장애 시 복구 경로를 반드시 검토해야 합니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 ===================================
2 Documentation for /proc/sys/kernel/
3 ===================================
4
5 .. See scripts/check-sysctl-docs to keep this up to date
6
7
8 Copyright (c) 1998, 1999, Rik van Riel <[email protected]>
9
10 Copyright (c) 2009, Shen Feng<[email protected]>
11
12 For general info and legal blurb, please look in
13 Documentation/admin-guide/sysctl/index.rst.
14
15 ------------------------------------------------------------------------------
16
17 This file contains documentation for the sysctl files in
18 ``/proc/sys/kernel/``.
19
20 The files in this directory can be used to tune and monitor
21 miscellaneous and general things in the operation of the Linux
22 kernel. Since some of the files *can* be used to screw up your
23 system, it is advisable to read both documentation and source
24 before actually making adjustments.
25
26 Currently, these files might (depending on your configuration)
27 show up in ``/proc/sys/kernel``:
28
29 .. contents:: :local:
30
31
32 acct
33 ====
34
35 ::
36
37 highwater lowwater frequency
38
39 If BSD-style process accounting is enabled these values control
40 its behaviour. If free space on filesystem where the log lives
41 goes below ``lowwater``\ % accounting suspends. If free space gets
42 above ``highwater``\ % accounting resumes. ``frequency`` determines
43 how often do we check the amount of free space (value is in
44 seconds). Default:
45
46 ::
47
48 4 2 30
49
50 That is, suspend accounting if free space drops below 2%; resume it
51 if it increases to at least 4%; consider information about amount of
52 free space valid for 30 seconds.
53
54
55 acpi_video_flags
56 ================
57
58 See Documentation/power/video.rst. This allows the video resume mode to be set,
59 in a similar fashion to the ``acpi_sleep`` kernel parameter, by
60 combining the following values:
61
62 = =======
63 1 s3_bios
64 2 s3_mode
65 4 s3_beep
66 = =======
67
68 arch
69 ====
70
71 The machine hardware name, the same output as ``uname -m``
72 (e.g. ``x86_64`` or ``aarch64``).
73
74 auto_msgmni
75 ===========
76
77 This variable has no effect and may be removed in future kernel
78 releases. Reading it always returns 0.
79 Up to Linux 3.17, it enabled/disabled automatic recomputing of
80 `msgmni`_
81 upon memory add/remove or upon IPC namespace creation/removal.
82 Echoing "1" into this file enabled msgmni automatic recomputing.
83 Echoing "0" turned it off. The default value was 1.
84
85
86 bootloader_type (x86 only)
87 ==========================
88
89 This gives the bootloader type number as indicated by the bootloader,
90 shifted left by 4, and OR'd with the low four bits of the bootloader
91 version. The reason for this encoding is that this used to match the
92 ``type_of_loader`` field in the kernel header; the encoding is kept for
93 backwards compatibility. That is, if the full bootloader type number
94 is 0x15 and the full version number is 0x234, this file will contain
95 the value 340 = 0x154.
96
97 See the ``type_of_loader`` and ``ext_loader_type`` fields in
98 Documentation/arch/x86/boot.rst for additional information.
99
100
101 bootloader_version (x86 only)
102 =============================
103
104 The complete bootloader version number. In the example above, this
105 file will contain the value 564 = 0x234.
106
107 See the ``type_of_loader`` and ``ext_loader_ver`` fields in
108 Documentation/arch/x86/boot.rst for additional information.
109
110
111 bpf_stats_enabled
112 =================
113
114 Controls whether the kernel should collect statistics on BPF programs
115 (total time spent running, number of times run...). Enabling
116 statistics causes a slight reduction in performance on each program
117 run. The statistics can be seen using ``bpftool``.
118
119 = ===================================
120 0 Don't collect statistics (default).
121 1 Collect statistics.
122 = ===================================
123
124
125 cad_pid
126 =======
127
128 This is the pid which will be signalled on reboot (notably, by
129 Ctrl-Alt-Delete). Writing a value to this file which doesn't
130 correspond to a running process will result in ``-ESRCH``.
131
132 See also `ctrl-alt-del`_.
133
134
135 cap_last_cap
136 ============
137
138 Highest valid capability of the running kernel. Exports
139 ``CAP_LAST_CAP`` from the kernel.
140
141
142 .. _core_pattern:
143
144 core_pattern
145 ============
146
147 ``core_pattern`` is used to specify a core dumpfile pattern name.
148
149 * max length 127 characters; default value is "core"
150 * ``core_pattern`` is used as a pattern template for the output
151 filename; certain string patterns (beginning with '%') are
152 substituted with their actual values.
153 * backward compatibility with ``core_uses_pid``:
154
155 If ``core_pattern`` does not include "%p" (default does not)
156 and ``core_uses_pid`` is set, then .PID will be appended to
157 the filename.
158
159 * corename format specifiers
160
161 ======== ==========================================
162 %<NUL> '%' is dropped
163 %% output one '%'
164 %p pid
165 %P global pid (init PID namespace)
166 %i tid
167 %I global tid (init PID namespace)
168 %u uid (in initial user namespace)
169 %g gid (in initial user namespace)
170 %d dump mode, matches ``PR_SET_DUMPABLE`` and
171 ``/proc/sys/fs/suid_dumpable``
172 %s signal number
173 %t UNIX time of dump
174 %h hostname
175 %e executable filename (may be shortened, could be changed by prctl etc)
176 %f executable filename
177 %E executable path
178 %c maximum size of core file by resource limit RLIMIT_CORE
179 %C CPU the task ran on
180 %F pidfd number
181 %<OTHER> both are dropped
182 ======== ==========================================
183
184 * If the first character of the pattern is a '|', the kernel will treat
185 the rest of the pattern as a command to run. The core dump will be
186 written to the standard input of that program instead of to a file.
187
188
189 core_pipe_limit
190 ===============
191
192 This sysctl is only applicable when `core_pattern`_ is configured to
193 pipe core files to a user space helper (when the first character of
194 ``core_pattern`` is a '|', see above).
195 When collecting cores via a pipe to an application, it is occasionally
196 useful for the collecting application to gather data about the
197 crashing process from its ``/proc/pid`` directory.
198 In order to do this safely, the kernel must wait for the collecting
199 process to exit, so as not to remove the crashing processes proc files
200 prematurely.
201 This in turn creates the possibility that a misbehaving userspace
202 collecting process can block the reaping of a crashed process simply
203 by never exiting.
204 This sysctl defends against that.
205 It defines how many concurrent crashing processes may be piped to user
206 space applications in parallel.
207 If this value is exceeded, then those crashing processes above that
208 value are noted via the kernel log and their cores are skipped.
209 0 is a special value, indicating that unlimited processes may be
210 captured in parallel, but that no waiting will take place (i.e. the
211 collecting process is not guaranteed access to ``/proc/<crashing
212 pid>/``).
213 This value defaults to 0.
214
215
216 core_sort_vma
217 =============
218
219 The default coredump writes VMAs in address order. By setting
220 ``core_sort_vma`` to 1, VMAs will be written from smallest size
221 to largest size. This is known to break at least elfutils, but
222 can be handy when dealing with very large (and truncated)
223 coredumps where the more useful debugging details are included
224 in the smaller VMAs.
225
226
227 core_uses_pid
228 =============
229
230 The default coredump filename is "core". By setting
231 ``core_uses_pid`` to 1, the coredump filename becomes core.PID.
232 If `core_pattern`_ does not include "%p" (default does not)
233 and ``core_uses_pid`` is set, then .PID will be appended to
234 the filename.
235
236
237 ctrl-alt-del
238 ============
239
240 When the value in this file is 0, ctrl-alt-del is trapped and
241 sent to the ``init(1)`` program to handle a graceful restart.
242 When, however, the value is > 0, Linux's reaction to a Vulcan
243 Nerve Pinch (tm) will be an immediate reboot, without even
244 syncing its dirty buffers.
245
246 Note:
247 when a program (like dosemu) has the keyboard in 'raw'
248 mode, the ctrl-alt-del is intercepted by the program before it
249 ever reaches the kernel tty layer, and it's up to the program
250 to decide what to do with it.
251
252
253 dmesg_restrict
254 ==============
255
256 This toggle indicates whether unprivileged users are prevented
257 from using ``dmesg(8)`` to view messages from the kernel's log
258 buffer.
259 When ``dmesg_restrict`` is set to 0 there are no restrictions.
260 When ``dmesg_restrict`` is set to 1, users must have
261 ``CAP_SYSLOG`` to use ``dmesg(8)``.
262
263 The kernel config option ``CONFIG_SECURITY_DMESG_RESTRICT`` sets the
264 default value of ``dmesg_restrict``.
265
266
267 domainname & hostname
268 =====================
269
270 These files can be used to set the NIS/YP domainname and the
271 hostname of your box in exactly the same way as the commands
272 domainname and hostname, i.e.::
273
274 # echo "darkstar" > /proc/sys/kernel/hostname
275 # echo "mydomain" > /proc/sys/kernel/domainname
276
277 has the same effect as::
278
279 # hostname "darkstar"
280 # domainname "mydomain"
281
282 Note, however, that the classic darkstar.frop.org has the
283 hostname "darkstar" and DNS (Internet Domain Name Server)
284 domainname "frop.org", not to be confused with the NIS (Network
285 Information Service) or YP (Yellow Pages) domainname. These two
286 domain names are in general different. For a detailed discussion
287 see the ``hostname(1)`` man page.
288
289
290 firmware_config
291 ===============
292
293 See Documentation/driver-api/firmware/fallback-mechanisms.rst.
294
295 The entries in this directory allow the firmware loader helper
296 fallback to be controlled:
297
298 * ``force_sysfs_fallback``, when set to 1, forces the use of the
299 fallback;
300 * ``ignore_sysfs_fallback``, when set to 1, ignores any fallback.
301
302
303 ftrace_dump_on_oops
304 ===================
305
306 Determines whether ``ftrace_dump()`` should be called on an oops (or
307 kernel panic). This will output the contents of the ftrace buffers to
308 the console. This is very useful for capturing traces that lead to
309 crashes and outputting them to a serial console.
310
311 ======================= ===========================================
312 0 Disabled (default).
313 1 Dump buffers of all CPUs.
314 2(orig_cpu) Dump the buffer of the CPU that triggered the
315 oops.
316 <instance> Dump the specific instance buffer on all CPUs.
317 <instance>=2(orig_cpu) Dump the specific instance buffer on the CPU
318 that triggered the oops.
319 ======================= ===========================================
320
321 Multiple instance dump is also supported, and instances are separated
322 by commas. If global buffer also needs to be dumped, please specify
323 the dump mode (1/2/orig_cpu) first for global buffer.
324
325 So for example to dump "foo" and "bar" instance buffer on all CPUs,
326 user can::
327
328 echo "foo,bar" > /proc/sys/kernel/ftrace_dump_on_oops
329
330 To dump global buffer and "foo" instance buffer on all
331 CPUs along with the "bar" instance buffer on CPU that triggered the
332 oops, user can::
333
334 echo "1,foo,bar=2" > /proc/sys/kernel/ftrace_dump_on_oops
335
336 ftrace_enabled, stack_tracer_enabled
337 ====================================
338
339 See Documentation/trace/ftrace.rst.
340
341
342 hardlockup_all_cpu_backtrace
343 ============================
344
345 This value controls the hard lockup detector behavior when a hard
346 lockup condition is detected as to whether or not to gather further
347 debug information. If enabled, arch-specific all-CPU stack dumping
348 will be initiated.
349
350 = ============================================
351 0 Do nothing. This is the default behavior.
352 1 On detection capture more debug information.
353 = ============================================
354
355
356 hardlockup_panic
357 ================
358
359 This parameter can be used to control whether the kernel panics
360 when a hard lockup is detected.
361
362 = ===========================
363 0 Don't panic on hard lockup.
364 1 Panic on hard lockup.
365 = ===========================
366
367 See Documentation/admin-guide/lockup-watchdogs.rst for more information.
368 This can also be set using the nmi_watchdog kernel parameter.
369
370
371 hotplug
372 =======
373
374 Path for the hotplug policy agent.
375 Default value is ``CONFIG_UEVENT_HELPER_PATH``, which in turn defaults
376 to the empty string.
377
378 This file only exists when ``CONFIG_UEVENT_HELPER`` is enabled. Most
379 modern systems rely exclusively on the netlink-based uevent source and
380 don't need this.
381
382
383 hung_task_all_cpu_backtrace
384 ===========================
385
386 If this option is set, the kernel will send an NMI to all CPUs to dump
387 their backtraces when a hung task is detected. This file shows up if
388 CONFIG_DETECT_HUNG_TASK and CONFIG_SMP are enabled.
389
390 0: Won't show all CPUs backtraces when a hung task is detected.
391 This is the default behavior.
392
393 1: Will non-maskably interrupt all CPUs and dump their backtraces when
394 a hung task is detected.
395
396
397 hung_task_panic
398 ===============
399
400 Controls the kernel's behavior when a hung task is detected.
401 This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
402
403 = =================================================
404 0 Continue operation. This is the default behavior.
405 1 Panic immediately.
406 = =================================================
407
408
409 hung_task_check_count
410 =====================
411
412 The upper bound on the number of tasks that are checked.
413 This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
414
415
416 hung_task_detect_count
417 ======================
418
419 Indicates the total number of tasks that have been detected as hung since
420 the system boot.
421
422 This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
423
424
425 hung_task_timeout_secs
426 ======================
427
428 When a task in D state did not get scheduled
429 for more than this value report a warning.
430 This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
431
432 0 means infinite timeout, no checking is done.
433
434 Possible values to set are in range {0:``LONG_MAX``/``HZ``}.
435
436
437 hung_task_check_interval_secs
438 =============================
439
440 Hung task check interval. If hung task checking is enabled
441 (see `hung_task_timeout_secs`_), the check is done every
442 ``hung_task_check_interval_secs`` seconds.
443 This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
444
445 0 (default) means use ``hung_task_timeout_secs`` as checking
446 interval.
447
448 Possible values to set are in range {0:``LONG_MAX``/``HZ``}.
449
450
451 hung_task_warnings
452 ==================
453
454 The maximum number of warnings to report. During a check interval
455 if a hung task is detected, this value is decreased by 1.
456 When this value reaches 0, no more warnings will be reported.
457 This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
458
459 -1: report an infinite number of warnings.
460
461
462 hyperv_record_panic_msg
463 =======================
464
465 Controls whether the panic kmsg data should be reported to Hyper-V.
466
467 = =========================================================
468 0 Do not report panic kmsg data.
469 1 Report the panic kmsg data. This is the default behavior.
470 = =========================================================
471
472
473 ignore-unaligned-usertrap
474 =========================
475
476 On architectures where unaligned accesses cause traps, and where this
477 feature is supported (``CONFIG_SYSCTL_ARCH_UNALIGN_NO_WARN``;
478 currently, ``arc``, ``parisc`` and ``loongarch``), controls whether all
479 unaligned traps are logged.
480
481 = =============================================================
482 0 Log all unaligned accesses.
483 1 Only warn the first time a process traps. This is the default
484 setting.
485 = =============================================================
486
487 See also `unaligned-trap`_.
488
489 io_uring_disabled
490 =================
491
492 Prevents all processes from creating new io_uring instances. Enabling this
493 shrinks the kernel's attack surface.
494
495 = ======================================================================
496 0 All processes can create io_uring instances as normal. This is the
497 default setting.
498 1 io_uring creation is disabled (io_uring_setup() will fail with
499 -EPERM) for unprivileged processes not in the io_uring_group group.
500 Existing io_uring instances can still be used. See the
501 documentation for io_uring_group for more information.
502 2 io_uring creation is disabled for all processes. io_uring_setup()
503 always fails with -EPERM. Existing io_uring instances can still be
504 used.
505 = ======================================================================
506
507
508 io_uring_group
509 ==============
510
511 When io_uring_disabled is set to 1, a process must either be
512 privileged (CAP_SYS_ADMIN) or be in the io_uring_group group in order
513 to create an io_uring instance. If io_uring_group is set to -1 (the
514 default), only processes with the CAP_SYS_ADMIN capability may create
515 io_uring instances.
516
517
518 kexec_load_disabled
519 ===================
520
521 A toggle indicating if the syscalls ``kexec_load`` and
522 ``kexec_file_load`` have been disabled.
523 This value defaults to 0 (false: ``kexec_*load`` enabled), but can be
524 set to 1 (true: ``kexec_*load`` disabled).
525 Once true, kexec can no longer be used, and the toggle cannot be set
526 back to false.
527 This allows a kexec image to be loaded before disabling the syscall,
528 allowing a system to set up (and later use) an image without it being
529 altered.
530 Generally used together with the `modules_disabled`_ sysctl.
531
532 kexec_load_limit_panic
533 ======================
534
535 This parameter specifies a limit to the number of times the syscalls
536 ``kexec_load`` and ``kexec_file_load`` can be called with a crash
537 image. It can only be set with a more restrictive value than the
538 current one.
539
540 == ======================================================
541 -1 Unlimited calls to kexec. This is the default setting.
542 N Number of calls left.
543 == ======================================================
544
545 kexec_load_limit_reboot
546 =======================
547
548 Similar functionality as ``kexec_load_limit_panic``, but for a normal
549 image.
550
551 kptr_restrict
552 =============
553
554 This toggle indicates whether restrictions are placed on
555 exposing kernel addresses via ``/proc`` and other interfaces.
556
557 When ``kptr_restrict`` is set to 0 (the default) the address is hashed
558 before printing.
559 (This is the equivalent to %p.)
560
561 When ``kptr_restrict`` is set to 1, kernel pointers printed using the
562 %pK format specifier will be replaced with 0s unless the user has
563 ``CAP_SYSLOG`` and effective user and group ids are equal to the real
564 ids.
565 This is because %pK checks are done at read() time rather than open()
566 time, so if permissions are elevated between the open() and the read()
567 (e.g via a setuid binary) then %pK will not leak kernel pointers to
568 unprivileged users.
569 Note, this is a temporary solution only.
570 The correct long-term solution is to do the permission checks at
571 open() time.
572 Consider removing world read permissions from files that use %pK, and
573 using `dmesg_restrict`_ to protect against uses of %pK in ``dmesg(8)``
574 if leaking kernel pointer values to unprivileged users is a concern.
575
576 When ``kptr_restrict`` is set to 2, kernel pointers printed using
577 %pK will be replaced with 0s regardless of privileges.
578
579
580 modprobe
581 ========
582
583 The full path to the usermode helper for autoloading kernel modules,
584 by default ``CONFIG_MODPROBE_PATH``, which in turn defaults to
585 "/sbin/modprobe". This binary is executed when the kernel requests a
586 module. For example, if userspace passes an unknown filesystem type
587 to mount(), then the kernel will automatically request the
588 corresponding filesystem module by executing this usermode helper.
589 This usermode helper should insert the needed module into the kernel.
590
591 This sysctl only affects module autoloading. It has no effect on the
592 ability to explicitly insert modules.
593
594 This sysctl can be used to debug module loading requests::
595
596 echo '#! /bin/sh' > /tmp/modprobe
597 echo 'echo "$@" >> /tmp/modprobe.log' >> /tmp/modprobe
598 echo 'exec /sbin/modprobe "$@"' >> /tmp/modprobe
599 chmod a+x /tmp/modprobe
600 echo /tmp/modprobe > /proc/sys/kernel/modprobe
601
602 Alternatively, if this sysctl is set to the empty string, then module
603 autoloading is completely disabled. The kernel will not try to
604 execute a usermode helper at all, nor will it call the
605 kernel_module_request LSM hook.
606
607 If CONFIG_STATIC_USERMODEHELPER=y is set in the kernel configuration,
608 then the configured static usermode helper overrides this sysctl,
609 except that the empty string is still accepted to completely disable
610 module autoloading as described above.
611
612 modules_disabled
613 ================
614
615 A toggle value indicating if modules are allowed to be loaded
616 in an otherwise modular kernel. This toggle defaults to off
617 (0), but can be set true (1). Once true, modules can be
618 neither loaded nor unloaded, and the toggle cannot be set back
619 to false. Generally used with the `kexec_load_disabled`_ toggle.
620
621
622 .. _msgmni:
623
624 msgmax, msgmnb, and msgmni
625 ==========================
626
627 ``msgmax`` is the maximum size of an IPC message, in bytes. 8192 by
628 default (``MSGMAX``).
629
630 ``msgmnb`` is the maximum size of an IPC queue, in bytes. 16384 by
631 default (``MSGMNB``).
632
633 ``msgmni`` is the maximum number of IPC queues. 32000 by default
634 (``MSGMNI``).
635
636 All of these parameters are set per ipc namespace. The maximum number of bytes
637 in POSIX message queues is limited by ``RLIMIT_MSGQUEUE``. This limit is
638 respected hierarchically in the each user namespace.
639
640 msg_next_id, sem_next_id, and shm_next_id (System V IPC)
641 ========================================================
642
643 These three toggles allows to specify desired id for next allocated IPC
644 object: message, semaphore or shared memory respectively.
645
646 By default they are equal to -1, which means generic allocation logic.
647 Possible values to set are in range {0:``INT_MAX``}.
648
649 Notes:
650 1) kernel doesn't guarantee, that new object will have desired id. So,
651 it's up to userspace, how to handle an object with "wrong" id.
652 2) Toggle with non-default value will be set back to -1 by kernel after
653 successful IPC object allocation. If an IPC object allocation syscall
654 fails, it is undefined if the value remains unmodified or is reset to -1.
655
656
657 ngroups_max
658 ===========
659
660 Maximum number of supplementary groups, _i.e._ the maximum size which
661 ``setgroups`` will accept. Exports ``NGROUPS_MAX`` from the kernel.
662
663
664
665 nmi_watchdog
666 ============
667
668 This parameter can be used to control the NMI watchdog
669 (i.e. the hard lockup detector) on x86 systems.
670
671 = =================================
672 0 Disable the hard lockup detector.
673 1 Enable the hard lockup detector.
674 = =================================
675
676 The hard lockup detector monitors each CPU for its ability to respond to
677 timer interrupts. The mechanism utilizes CPU performance counter registers
678 that are programmed to generate Non-Maskable Interrupts (NMIs) periodically
679 while a CPU is busy. Hence, the alternative name 'NMI watchdog'.
680
681 The NMI watchdog is disabled by default if the kernel is running as a guest
682 in a KVM virtual machine. This default can be overridden by adding::
683
684 nmi_watchdog=1
685
686 to the guest kernel command line (see
687 Documentation/admin-guide/kernel-parameters.rst).
688
689
690 nmi_wd_lpm_factor (PPC only)
691 ============================
692
693 Factor to apply to the NMI watchdog timeout (only when ``nmi_watchdog`` is
694 set to 1). This factor represents the percentage added to
695 ``watchdog_thresh`` when calculating the NMI watchdog timeout during an
696 LPM. The soft lockup timeout is not impacted.
697
698 A value of 0 means no change. The default value is 200 meaning the NMI
699 watchdog is set to 30s (based on ``watchdog_thresh`` equal to 10).
700
701
702 numa_balancing
703 ==============
704
705 Enables/disables and configures automatic page fault based NUMA memory
706 balancing. Memory is moved automatically to nodes that access it often.
707 The value to set can be the result of ORing the following:
708
709 = =================================
710 0 NUMA_BALANCING_DISABLED
711 1 NUMA_BALANCING_NORMAL
712 2 NUMA_BALANCING_MEMORY_TIERING
713 = =================================
714
715 Or NUMA_BALANCING_NORMAL to optimize page placement among different
716 NUMA nodes to reduce remote accessing. On NUMA machines, there is a
717 performance penalty if remote memory is accessed by a CPU. When this
718 feature is enabled the kernel samples what task thread is accessing
719 memory by periodically unmapping pages and later trapping a page
720 fault. At the time of the page fault, it is determined if the data
721 being accessed should be migrated to a local memory node.
722
723 The unmapping of pages and trapping faults incur additional overhead that
724 ideally is offset by improved memory locality but there is no universal
725 guarantee. If the target workload is already bound to NUMA nodes then this
726 feature should be disabled.
727
728 Or NUMA_BALANCING_MEMORY_TIERING to optimize page placement among
729 different types of memory (represented as different NUMA nodes) to
730 place the hot pages in the fast memory. This is implemented based on
731 unmapping and page fault too.
732
733 numa_balancing_promote_rate_limit_MBps
734 ======================================
735
736 Too high promotion/demotion throughput between different memory types
737 may hurt application latency. This can be used to rate limit the
738 promotion throughput. The per-node max promotion throughput in MB/s
739 will be limited to be no more than the set value.
740
741 A rule of thumb is to set this to less than 1/10 of the PMEM node
742 write bandwidth.
743
744 oops_all_cpu_backtrace
745 ======================
746
747 If this option is set, the kernel will send an NMI to all CPUs to dump
748 their backtraces when an oops event occurs. It should be used as a last
749 resort in case a panic cannot be triggered (to protect VMs running, for
750 example) or kdump can't be collected. This file shows up if CONFIG_SMP
751 is enabled.
752
753 0: Won't show all CPUs backtraces when an oops is detected.
754 This is the default behavior.
755
756 1: Will non-maskably interrupt all CPUs and dump their backtraces when
757 an oops event is detected.
758
759
760 oops_limit
761 ==========
762
763 Number of kernel oopses after which the kernel should panic when
764 ``panic_on_oops`` is not set. Setting this to 0 disables checking
765 the count. Setting this to 1 has the same effect as setting
766 ``panic_on_oops=1``. The default value is 10000.
767
768
769 osrelease, ostype & version
770 ===========================
771
772 ::
773
774 # cat osrelease
775 2.1.88
776 # cat ostype
777 Linux
778 # cat version
779 #5 Wed Feb 25 21:49:24 MET 1998
780
781 The files ``osrelease`` and ``ostype`` should be clear enough.
782 ``version``
783 needs a little more clarification however. The '#5' means that
784 this is the fifth kernel built from this source base and the
785 date behind it indicates the time the kernel was built.
786 The only way to tune these values is to rebuild the kernel :-)
787
788
789 overflowgid & overflowuid
790 =========================
791
792 if your architecture did not always support 32-bit UIDs (i.e. arm,
793 i386, m68k, sh, and sparc32), a fixed UID and GID will be returned to
794 applications that use the old 16-bit UID/GID system calls, if the
795 actual UID or GID would exceed 65535.
796
797 These sysctls allow you to change the value of the fixed UID and GID.
798 The default is 65534.
799
800
801 panic
802 =====
803
804 The value in this file determines the behaviour of the kernel on a
805 panic:
806
807 * if zero, the kernel will loop forever;
808 * if negative, the kernel will reboot immediately;
809 * if positive, the kernel will reboot after the corresponding number
810 of seconds.
811
812 When you use the software watchdog, the recommended setting is 60.
813
814
815 panic_on_io_nmi
816 ===============
817
818 Controls the kernel's behavior when a CPU receives an NMI caused by
819 an IO error.
820
821 = ==================================================================
822 0 Try to continue operation (default).
823 1 Panic immediately. The IO error triggered an NMI. This indicates a
824 serious system condition which could result in IO data corruption.
825 Rather than continuing, panicking might be a better choice. Some
826 servers issue this sort of NMI when the dump button is pushed,
827 and you can use this option to take a crash dump.
828 = ==================================================================
829
830
831 panic_on_oops
832 =============
833
834 Controls the kernel's behaviour when an oops or BUG is encountered.
835
836 = ===================================================================
837 0 Try to continue operation.
838 1 Panic immediately. If the `panic` sysctl is also non-zero then the
839 machine will be rebooted.
840 = ===================================================================
841
842
843 panic_on_stackoverflow
844 ======================
845
846 Controls the kernel's behavior when detecting the overflows of
847 kernel, IRQ and exception stacks except a user stack.
848 This file shows up if ``CONFIG_DEBUG_STACKOVERFLOW`` is enabled.
849
850 = ==========================
851 0 Try to continue operation.
852 1 Panic immediately.
853 = ==========================
854
855
856 panic_on_unrecovered_nmi
857 ========================
858
859 The default Linux behaviour on an NMI of either memory or unknown is
860 to continue operation. For many environments such as scientific
861 computing it is preferable that the box is taken out and the error
862 dealt with than an uncorrected parity/ECC error get propagated.
863
864 A small number of systems do generate NMIs for bizarre random reasons
865 such as power management so the default is off. That sysctl works like
866 the existing panic controls already in that directory.
867
868
869 panic_on_warn
870 =============
871
872 Calls panic() in the WARN() path when set to 1. This is useful to avoid
873 a kernel rebuild when attempting to kdump at the location of a WARN().
874
875 = ================================================
876 0 Only WARN(), default behaviour.
877 1 Call panic() after printing out WARN() location.
878 = ================================================
879
880
881 panic_print
882 ===========
883
884 Bitmask for printing system info when panic happens. User can chose
885 combination of the following bits:
886
887 ===== ============================================
888 bit 0 print all tasks info
889 bit 1 print system memory info
890 bit 2 print timer info
891 bit 3 print locks info if ``CONFIG_LOCKDEP`` is on
892 bit 4 print ftrace buffer
893 bit 5 replay all kernel messages on consoles at the end of panic
894 bit 6 print all CPUs backtrace (if available in the arch)
895 bit 7 print only tasks in uninterruptible (blocked) state
896 ===== ============================================
897
898 So for example to print tasks and memory info on panic, user can::
899
900 echo 3 > /proc/sys/kernel/panic_print
901
902
903 panic_sys_info
904 ==============
905
906 A comma separated list of extra information to be dumped on panic,
907 for example, "tasks,mem,timers,...". It is a human readable alternative
908 to 'panic_print'. Possible values are:
909
910 ============= ===================================================
911 tasks print all tasks info
912 mem print system memory info
913 timer print timers info
914 lock print locks info if CONFIG_LOCKDEP is on
915 ftrace print ftrace buffer
916 all_bt print all CPUs backtrace (if available in the arch)
917 blocked_tasks print only tasks in uninterruptible (blocked) state
918 ============= ===================================================
919
920
921 panic_on_rcu_stall
922 ==================
923
924 When set to 1, calls panic() after RCU stall detection messages. This
925 is useful to define the root cause of RCU stalls using a vmcore.
926
927 = ============================================================
928 0 Do not panic() when RCU stall takes place, default behavior.
929 1 panic() after printing RCU stall messages.
930 = ============================================================
931
932 max_rcu_stall_to_panic
933 ======================
934
935 When ``panic_on_rcu_stall`` is set to 1, this value determines the
936 number of times that RCU can stall before panic() is called.
937
938 When ``panic_on_rcu_stall`` is set to 0, this value is has no effect.
939
940 perf_cpu_time_max_percent
941 =========================
942
943 Hints to the kernel how much CPU time it should be allowed to
944 use to handle perf sampling events. If the perf subsystem
945 is informed that its samples are exceeding this limit, it
946 will drop its sampling frequency to attempt to reduce its CPU
947 usage.
948
949 Some perf sampling happens in NMIs. If these samples
950 unexpectedly take too long to execute, the NMIs can become
951 stacked up next to each other so much that nothing else is
952 allowed to execute.
953
954 ===== ========================================================
955 0 Disable the mechanism. Do not monitor or correct perf's
956 sampling rate no matter how CPU time it takes.
957
958 1-100 Attempt to throttle perf's sample rate to this
959 percentage of CPU. Note: the kernel calculates an
960 "expected" length of each sample event. 100 here means
961 100% of that expected length. Even if this is set to
962 100, you may still see sample throttling if this
963 length is exceeded. Set to 0 if you truly do not care
964 how much CPU is consumed.
965 ===== ========================================================
966
967
968 perf_event_paranoid
969 ===================
970
971 Controls use of the performance events system by unprivileged
972 users (without CAP_PERFMON). The default value is 2.
973
974 For backward compatibility reasons access to system performance
975 monitoring and observability remains open for CAP_SYS_ADMIN
976 privileged processes but CAP_SYS_ADMIN usage for secure system
977 performance monitoring and observability operations is discouraged
978 with respect to CAP_PERFMON use cases.
979
980 === ==================================================================
981 -1 Allow use of (almost) all events by all users.
982
983 Ignore mlock limit after perf_event_mlock_kb without
984 ``CAP_IPC_LOCK``.
985
986 >=0 Disallow ftrace function tracepoint by users without
987 ``CAP_PERFMON``.
988
989 Disallow raw tracepoint access by users without ``CAP_PERFMON``.
990
991 >=1 Disallow CPU event access by users without ``CAP_PERFMON``.
992
993 >=2 Disallow kernel profiling by users without ``CAP_PERFMON``.
994 === ==================================================================
995
996
997 perf_event_max_stack
998 ====================
999
1000 Controls maximum number of stack frames to copy for (``attr.sample_type &
1001 PERF_SAMPLE_CALLCHAIN``) configured events, for instance, when using
1002 '``perf record -g``' or '``perf trace --call-graph fp``'.
1004 This can only be done when no events are in use that have callchains
1005 enabled, otherwise writing to this file will return ``-EBUSY``.
1007 The default value is 127.
1010 perf_event_mlock_kb
1011 ===================
1013 Control size of per-cpu ring buffer not counted against mlock limit.
1015 The default value is 512 + 1 page
1018 perf_event_max_contexts_per_stack
1019 =================================
1021 Controls maximum number of stack frame context entries for
1022 (``attr.sample_type & PERF_SAMPLE_CALLCHAIN``) configured events, for
1023 instance, when using '``perf record -g``' or '``perf trace --call-graph fp``'.
1025 This can only be done when no events are in use that have callchains
1026 enabled, otherwise writing to this file will return ``-EBUSY``.
1028 The default value is 8.
1031 perf_user_access (arm64 and riscv only)
1032 =======================================
1034 Controls user space access for reading perf event counters.
1036 * for arm64
1037 The default value is 0 (access disabled).
1039 When set to 1, user space can read performance monitor counter registers
1040 directly.
1042 See Documentation/arch/arm64/perf.rst for more information.
1044 * for riscv
1045 When set to 0, user space access is disabled.
1047 The default value is 1, user space can read performance monitor counter
1048 registers through perf, any direct access without perf intervention will trigger
1049 an illegal instruction.
1051 When set to 2, which enables legacy mode (user space has direct access to cycle
1052 and insret CSRs only). Note that this legacy value is deprecated and will be
1053 removed once all user space applications are fixed.
1055 Note that the time CSR is always directly accessible to all modes.
1057 pid_max
1058 =======
1060 PID allocation wrap value. When the kernel's next PID value
1061 reaches this value, it wraps back to a minimum PID value.
1062 PIDs of value ``pid_max`` or larger are not allocated.
1065 ns_last_pid
1066 ===========
1068 The last pid allocated in the current (the one task using this sysctl
1069 lives in) pid namespace. When selecting a pid for a next task on fork
1070 kernel tries to allocate a number starting from this one.
1073 powersave-nap (PPC only)
1074 ========================
1076 If set, Linux-PPC will use the 'nap' mode of powersaving,
1077 otherwise the 'doze' mode will be used.
1080 ==============================================================
1082 printk
1083 ======
1085 The four values in printk denote: ``console_loglevel``,
1086 ``default_message_loglevel``, ``minimum_console_loglevel`` and
1087 ``default_console_loglevel`` respectively.
1089 These values influence printk() behavior when printing or
1090 logging error messages. See '``man 2 syslog``' for more info on
1091 the different loglevels.
1093 ======================== =====================================
1094 console_loglevel messages with a higher priority than
1095 this will be printed to the console
1096 default_message_loglevel messages without an explicit priority
1097 will be printed with this priority
1098 minimum_console_loglevel minimum (highest) value to which
1099 console_loglevel can be set
1100 default_console_loglevel default value for console_loglevel
1101 ======================== =====================================
1104 printk_delay
1105 ============
1107 Delay each printk message in ``printk_delay`` milliseconds
1109 Value from 0 - 10000 is allowed.
1112 printk_ratelimit
1113 ================
1115 Some warning messages are rate limited. ``printk_ratelimit`` specifies
1116 the minimum length of time between these messages (in seconds).
1117 The default value is 5 seconds.
1119 A value of 0 will disable rate limiting.
1122 printk_ratelimit_burst
1123 ======================
1125 While long term we enforce one message per `printk_ratelimit`_
1126 seconds, we do allow a burst of messages to pass through.
1127 ``printk_ratelimit_burst`` specifies the number of messages we can
1128 send before ratelimiting kicks in. After `printk_ratelimit`_ seconds
1129 have elapsed, another burst of messages may be sent.
1131 The default value is 10 messages.
1134 printk_devkmsg
1135 ==============
1137 Control the logging to ``/dev/kmsg`` from userspace:
1139 ========= =============================================
1140 ratelimit default, ratelimited
1141 on unlimited logging to /dev/kmsg from userspace
1142 off logging to /dev/kmsg disabled
1143 ========= =============================================
1145 The kernel command line parameter ``printk.devkmsg=`` overrides this and is
1146 a one-time setting until next reboot: once set, it cannot be changed by
1147 this sysctl interface anymore.
1149 ==============================================================
1152 pty
1153 ===
1155 See Documentation/filesystems/devpts.rst.
1158 random
1159 ======
1161 This is a directory, with the following entries:
1163 * ``boot_id``: a UUID generated the first time this is retrieved, and
1164 unvarying after that;
1166 * ``uuid``: a UUID generated every time this is retrieved (this can
1167 thus be used to generate UUIDs at will);
1169 * ``entropy_avail``: the pool's entropy count, in bits;
1171 * ``poolsize``: the entropy pool size, in bits;
1173 * ``urandom_min_reseed_secs``: obsolete (used to determine the minimum
1174 number of seconds between urandom pool reseeding). This file is
1175 writable for compatibility purposes, but writing to it has no effect
1176 on any RNG behavior;
1178 * ``write_wakeup_threshold``: when the entropy count drops below this
1179 (as a number of bits), processes waiting to write to ``/dev/random``
1180 are woken up. This file is writable for compatibility purposes, but
1181 writing to it has no effect on any RNG behavior.
1184 randomize_va_space
1185 ==================
1187 This option can be used to select the type of process address
1188 space randomization that is used in the system, for architectures
1189 that support this feature.
1191 == ===========================================================================
1192 0 Turn the process address space randomization off. This is the
1193 default for architectures that do not support this feature anyways,
1194 and kernels that are booted with the "norandmaps" parameter.
1196 1 Make the addresses of mmap base, stack and VDSO page randomized.
1197 This, among other things, implies that shared libraries will be
1198 loaded to random addresses. Also for PIE-linked binaries, the
1199 location of code start is randomized. This is the default if the
1200 ``CONFIG_COMPAT_BRK`` option is enabled.
1202 2 Additionally enable heap randomization. This is the default if
1203 ``CONFIG_COMPAT_BRK`` is disabled.
1205 There are a few legacy applications out there (such as some ancient
1206 versions of libc.so.5 from 1996) that assume that brk area starts
1207 just after the end of the code+bss. These applications break when
1208 start of the brk area is randomized. There are however no known
1209 non-legacy applications that would be broken this way, so for most
1210 systems it is safe to choose full randomization.
1212 Systems with ancient and/or broken binaries should be configured
1213 with ``CONFIG_COMPAT_BRK`` enabled, which excludes the heap from process
1214 address space randomization.
1215 == ===========================================================================
1218 real-root-dev
1219 =============
1221 See Documentation/admin-guide/initrd.rst.
1224 reboot-cmd (SPARC only)
1225 =======================
1227 ??? This seems to be a way to give an argument to the Sparc
1228 ROM/Flash boot loader. Maybe to tell it what to do after
1229 rebooting. ???
1232 sched_energy_aware
1233 ==================
1235 Enables/disables Energy Aware Scheduling (EAS). EAS starts
1236 automatically on platforms where it can run (that is,
1237 platforms with asymmetric CPU topologies and having an Energy
1238 Model available). If your platform happens to meet the
1239 requirements for EAS but you do not want to use it, change
1240 this value to 0. On Non-EAS platforms, write operation fails and
1241 read doesn't return anything.
1243 task_delayacct
1244 ===============
1246 Enables/disables task delay accounting (see
1247 Documentation/accounting/delay-accounting.rst. Enabling this feature incurs
1248 a small amount of overhead in the scheduler but is useful for debugging
1249 and performance tuning. It is required by some tools such as iotop.
1251 sched_schedstats
1252 ================
1254 Enables/disables scheduler statistics. Enabling this feature
1255 incurs a small amount of overhead in the scheduler but is
1256 useful for debugging and performance tuning.
1258 sched_util_clamp_min
1259 ====================
1261 Max allowed *minimum* utilization.
1263 Default value is 1024, which is the maximum possible value.
1265 It means that any requested uclamp.min value cannot be greater than
1266 sched_util_clamp_min, i.e., it is restricted to the range
1267 [0:sched_util_clamp_min].
1269 sched_util_clamp_max
1270 ====================
1272 Max allowed *maximum* utilization.
1274 Default value is 1024, which is the maximum possible value.
1276 It means that any requested uclamp.max value cannot be greater than
1277 sched_util_clamp_max, i.e., it is restricted to the range
1278 [0:sched_util_clamp_max].
1280 sched_util_clamp_min_rt_default
1281 ===============================
1283 By default Linux is tuned for performance. Which means that RT tasks always run
1284 at the highest frequency and most capable (highest capacity) CPU (in
1285 heterogeneous systems).
1287 Uclamp achieves this by setting the requested uclamp.min of all RT tasks to
1288 1024 by default, which effectively boosts the tasks to run at the highest
1289 frequency and biases them to run on the biggest CPU.
1291 This knob allows admins to change the default behavior when uclamp is being
1292 used. In battery powered devices particularly, running at the maximum
1293 capacity and frequency will increase energy consumption and shorten the battery
1294 life.
1296 This knob is only effective for RT tasks which the user hasn't modified their
1297 requested uclamp.min value via sched_setattr() syscall.
1299 This knob will not escape the range constraint imposed by sched_util_clamp_min
1300 defined above.
1302 For example if
1304 sched_util_clamp_min_rt_default = 800
1305 sched_util_clamp_min = 600
1307 Then the boost will be clamped to 600 because 800 is outside of the permissible
1308 range of [0:600]. This could happen for instance if a powersave mode will
1309 restrict all boosts temporarily by modifying sched_util_clamp_min. As soon as
1310 this restriction is lifted, the requested sched_util_clamp_min_rt_default
1311 will take effect.
1313 seccomp
1314 =======
1316 See Documentation/userspace-api/seccomp_filter.rst.
1319 sg-big-buff
1320 ===========
1322 This file shows the size of the generic SCSI (sg) buffer.
1323 You can't tune it just yet, but you could change it on
1324 compile time by editing ``include/scsi/sg.h`` and changing
1325 the value of ``SG_BIG_BUFF``.
1327 There shouldn't be any reason to change this value. If
1328 you can come up with one, you probably know what you
1329 are doing anyway :)
1332 shmall
1333 ======
1335 This parameter sets the total amount of shared memory pages that can be used
1336 inside ipc namespace. The shared memory pages counting occurs for each ipc
1337 namespace separately and is not inherited. Hence, ``shmall`` should always be at
1338 least ``ceil(shmmax/PAGE_SIZE)``.
1340 If you are not sure what the default ``PAGE_SIZE`` is on your Linux
1341 system, you can run the following command::
1343 # getconf PAGE_SIZE
1345 To reduce or disable the ability to allocate shared memory, you must create a
1346 new ipc namespace, set this parameter to the required value and prohibit the
1347 creation of a new ipc namespace in the current user namespace or cgroups can
1348 be used.
1350 shmmax
1351 ======
1353 This value can be used to query and set the run time limit
1354 on the maximum shared memory segment size that can be created.
1355 Shared memory segments up to 1Gb are now supported in the
1356 kernel. This value defaults to ``SHMMAX``.
1359 shmmni
1360 ======
1362 This value determines the maximum number of shared memory segments.
1363 4096 by default (``SHMMNI``).
1366 shm_rmid_forced
1367 ===============
1369 Linux lets you set resource limits, including how much memory one
1370 process can consume, via ``setrlimit(2)``. Unfortunately, shared memory
1371 segments are allowed to exist without association with any process, and
1372 thus might not be counted against any resource limits. If enabled,
1373 shared memory segments are automatically destroyed when their attach
1374 count becomes zero after a detach or a process termination. It will
1375 also destroy segments that were created, but never attached to, on exit
1376 from the process. The only use left for ``IPC_RMID`` is to immediately
1377 destroy an unattached segment. Of course, this breaks the way things are
1378 defined, so some applications might stop working. Note that this
1379 feature will do you no good unless you also configure your resource
1380 limits (in particular, ``RLIMIT_AS`` and ``RLIMIT_NPROC``). Most systems don't
1381 need this.
1383 Note that if you change this from 0 to 1, already created segments
1384 without users and with a dead originative process will be destroyed.
1387 sysctl_writes_strict
1388 ====================
1390 Control how file position affects the behavior of updating sysctl values
1391 via the ``/proc/sys`` interface:
1393 == ======================================================================
1394 -1 Legacy per-write sysctl value handling, with no printk warnings.
1395 Each write syscall must fully contain the sysctl value to be
1396 written, and multiple writes on the same sysctl file descriptor
1397 will rewrite the sysctl value, regardless of file position.
1398 0 Same behavior as above, but warn about processes that perform writes
1399 to a sysctl file descriptor when the file position is not 0.
1400 1 (default) Respect file position when writing sysctl strings. Multiple
1401 writes will append to the sysctl value buffer. Anything past the max
1402 length of the sysctl value buffer will be ignored. Writes to numeric
1403 sysctl entries must always be at file position 0 and the value must
1404 be fully contained in the buffer sent in the write syscall.
1405 == ======================================================================
1408 softlockup_all_cpu_backtrace
1409 ============================
1411 This value controls the soft lockup detector thread's behavior
1412 when a soft lockup condition is detected as to whether or not
1413 to gather further debug information. If enabled, each cpu will
1414 be issued an NMI and instructed to capture stack trace.
1416 This feature is only applicable for architectures which support
1417 NMI.
1419 = ============================================
1420 0 Do nothing. This is the default behavior.
1421 1 On detection capture more debug information.
1422 = ============================================
1425 softlockup_panic
1426 =================
1428 This parameter can be used to control whether the kernel panics
1429 when a soft lockup is detected.
1431 = ============================================
1432 0 Don't panic on soft lockup.
1433 1 Panic on soft lockup.
1434 = ============================================
1436 This can also be set using the softlockup_panic kernel parameter.
1439 soft_watchdog
1440 =============
1442 This parameter can be used to control the soft lockup detector.
1444 = =================================
1445 0 Disable the soft lockup detector.
1446 1 Enable the soft lockup detector.
1447 = =================================
1449 The soft lockup detector monitors CPUs for threads that are hogging the CPUs
1450 without rescheduling voluntarily, and thus prevent the 'migration/N' threads
1451 from running, causing the watchdog work fail to execute. The mechanism depends
1452 on the CPUs ability to respond to timer interrupts which are needed for the
1453 watchdog work to be queued by the watchdog timer function, otherwise the NMI
1454 watchdog — if enabled — can detect a hard lockup condition.
1457 split_lock_mitigate (x86 only)
1458 ==============================
1460 On x86, each "split lock" imposes a system-wide performance penalty. On larger
1461 systems, large numbers of split locks from unprivileged users can result in
1462 denials of service to well-behaved and potentially more important users.
1464 The kernel mitigates these bad users by detecting split locks and imposing
1465 penalties: forcing them to wait and only allowing one core to execute split
1466 locks at a time.
1468 These mitigations can make those bad applications unbearably slow. Setting
1469 split_lock_mitigate=0 may restore some application performance, but will also
1470 increase system exposure to denial of service attacks from split lock users.
1472 = ===================================================================
1473 0 Disable the mitigation mode - just warns the split lock on kernel log
1474 and exposes the system to denials of service from the split lockers.
1475 1 Enable the mitigation mode (this is the default) - penalizes the split
1476 lockers with intentional performance degradation.
1477 = ===================================================================
1480 stack_erasing
1481 =============
1483 This parameter can be used to control kernel stack erasing at the end
1484 of syscalls for kernels built with ``CONFIG_KSTACK_ERASE``.
1486 That erasing reduces the information which kernel stack leak bugs
1487 can reveal and blocks some uninitialized stack variable attacks.
1488 The tradeoff is the performance impact: on a single CPU system kernel
1489 compilation sees a 1% slowdown, other systems and workloads may vary.
1491 = ====================================================================
1492 0 Kernel stack erasing is disabled, KSTACK_ERASE_METRICS are not updated.
1493 1 Kernel stack erasing is enabled (default), it is performed before
1494 returning to the userspace at the end of syscalls.
1495 = ====================================================================
1498 stop-a (SPARC only)
1499 ===================
1501 Controls Stop-A:
1503 = ====================================
1504 0 Stop-A has no effect.
1505 1 Stop-A breaks to the PROM (default).
1506 = ====================================
1508 Stop-A is always enabled on a panic, so that the user can return to
1509 the boot PROM.
1512 sysrq
1513 =====
1515 See Documentation/admin-guide/sysrq.rst.
1518 tainted
1519 =======
1521 Non-zero if the kernel has been tainted. Numeric values, which can be
1522 ORed together. The letters are seen in "Tainted" line of Oops reports.
1524 ====== ===== ==============================================================
1525 1 `(P)` proprietary module was loaded
1526 2 `(F)` module was force loaded
1527 4 `(S)` kernel running on an out of specification system
1528 8 `(R)` module was force unloaded
1529 16 `(M)` processor reported a Machine Check Exception (MCE)
1530 32 `(B)` bad page referenced or some unexpected page flags
1531 64 `(U)` taint requested by userspace application
1532 128 `(D)` kernel died recently, i.e. there was an OOPS or BUG
1533 256 `(A)` an ACPI table was overridden by user
1534 512 `(W)` kernel issued warning
1535 1024 `(C)` staging driver was loaded
1536 2048 `(I)` workaround for bug in platform firmware applied
1537 4096 `(O)` externally-built ("out-of-tree") module was loaded
1538 8192 `(E)` unsigned module was loaded
1539 16384 `(L)` soft lockup occurred
1540 32768 `(K)` kernel has been live patched
1541 65536 `(X)` Auxiliary taint, defined and used by for distros
1542 131072 `(T)` The kernel was built with the struct randomization plugin
1543 ====== ===== ==============================================================
1545 See Documentation/admin-guide/tainted-kernels.rst for more information.
1547 Note:
1548 writes to this sysctl interface will fail with ``EINVAL`` if the kernel is
1549 booted with the command line option ``panic_on_taint=<bitmask>,nousertaint``
1550 and any of the ORed together values being written to ``tainted`` match with
1551 the bitmask declared on panic_on_taint.
1552 See Documentation/admin-guide/kernel-parameters.rst for more details on
1553 that particular kernel command line option and its optional
1554 ``nousertaint`` switch.
1556 threads-max
1557 ===========
1559 This value controls the maximum number of threads that can be created
1560 using ``fork()``.
1562 During initialization the kernel sets this value such that even if the
1563 maximum number of threads is created, the thread structures occupy only
1564 a part (1/8th) of the available RAM pages.
1566 The minimum value that can be written to ``threads-max`` is 1.
1568 The maximum value that can be written to ``threads-max`` is given by the
1569 constant ``FUTEX_TID_MASK`` (0x3fffffff).
1571 If a value outside of this range is written to ``threads-max`` an
1572 ``EINVAL`` error occurs.
1574 timer_migration
1575 ===============
1577 When set to a non-zero value, attempt to migrate timers away from idle cpus to
1578 allow them to remain in low power states longer.
1580 Default is set (1).
1582 traceoff_on_warning
1583 ===================
1585 When set, disables tracing (see Documentation/trace/ftrace.rst) when a
1586 ``WARN()`` is hit.
1589 tracepoint_printk
1590 =================
1592 When tracepoints are sent to printk() (enabled by the ``tp_printk``
1593 boot parameter), this entry provides runtime control::
1595 echo 0 > /proc/sys/kernel/tracepoint_printk
1597 will stop tracepoints from being sent to printk(), and::
1599 echo 1 > /proc/sys/kernel/tracepoint_printk
1601 will send them to printk() again.
1603 This only works if the kernel was booted with ``tp_printk`` enabled.
1605 See Documentation/admin-guide/kernel-parameters.rst and
1606 Documentation/trace/boottime-trace.rst.
1609 unaligned-trap
1610 ==============
1612 On architectures where unaligned accesses cause traps, and where this
1613 feature is supported (``CONFIG_SYSCTL_ARCH_UNALIGN_ALLOW``; currently,
1614 ``arc``, ``parisc`` and ``loongarch``), controls whether unaligned traps
1615 are caught and emulated (instead of failing).
1617 = ========================================================
1618 0 Do not emulate unaligned accesses.
1619 1 Emulate unaligned accesses. This is the default setting.
1620 = ========================================================
1622 See also `ignore-unaligned-usertrap`_.
1625 unknown_nmi_panic
1626 =================
1628 The value in this file affects behavior of handling NMI. When the
1629 value is non-zero, unknown NMI is trapped and then panic occurs. At
1630 that time, kernel debugging information is displayed on console.
1632 NMI switch that most IA32 servers have fires unknown NMI up, for
1633 example. If a system hangs up, try pressing the NMI switch.
1636 unprivileged_bpf_disabled
1637 =========================
1639 Writing 1 to this entry will disable unprivileged calls to ``bpf()``;
1640 once disabled, calling ``bpf()`` without ``CAP_SYS_ADMIN`` or ``CAP_BPF``
1641 will return ``-EPERM``. Once set to 1, this can't be cleared from the
1642 running kernel anymore.
1644 Writing 2 to this entry will also disable unprivileged calls to ``bpf()``,
1645 however, an admin can still change this setting later on, if needed, by
1646 writing 0 or 1 to this entry.
1648 If ``BPF_UNPRIV_DEFAULT_OFF`` is enabled in the kernel config, then this
1649 entry will default to 2 instead of 0.
1651 = =============================================================
1652 0 Unprivileged calls to ``bpf()`` are enabled
1653 1 Unprivileged calls to ``bpf()`` are disabled without recovery
1654 2 Unprivileged calls to ``bpf()`` are disabled
1655 = =============================================================
1658 warn_limit
1659 ==========
1661 Number of kernel warnings after which the kernel should panic when
1662 ``panic_on_warn`` is not set. Setting this to 0 disables checking
1663 the warning count. Setting this to 1 has the same effect as setting
1664 ``panic_on_warn=1``. The default value is 0.
1667 watchdog
1668 ========
1670 This parameter can be used to disable or enable the soft lockup detector
1671 *and* the NMI watchdog (i.e. the hard lockup detector) at the same time.
1673 = ==============================
1674 0 Disable both lockup detectors.
1675 1 Enable both lockup detectors.
1676 = ==============================
1678 The soft lockup detector and the NMI watchdog can also be disabled or
1679 enabled individually, using the ``soft_watchdog`` and ``nmi_watchdog``
1680 parameters.
1681 If the ``watchdog`` parameter is read, for example by executing::
1683 cat /proc/sys/kernel/watchdog
1685 the output of this command (0 or 1) shows the logical OR of
1686 ``soft_watchdog`` and ``nmi_watchdog``.
1689 watchdog_cpumask
1690 ================
1692 This value can be used to control on which cpus the watchdog may run.
1693 The default cpumask is all possible cores, but if ``NO_HZ_FULL`` is
1694 enabled in the kernel config, and cores are specified with the
1695 ``nohz_full=`` boot argument, those cores are excluded by default.
1696 Offline cores can be included in this mask, and if the core is later
1697 brought online, the watchdog will be started based on the mask value.
1699 Typically this value would only be touched in the ``nohz_full`` case
1700 to re-enable cores that by default were not running the watchdog,
1701 if a kernel lockup was suspected on those cores.
1703 The argument value is the standard cpulist format for cpumasks,
1704 so for example to enable the watchdog on cores 0, 2, 3, and 4 you
1705 might say::
1707 echo 0,2-4 > /proc/sys/kernel/watchdog_cpumask
1710 watchdog_thresh
1711 ===============
1713 This value can be used to control the frequency of hrtimer and NMI
1714 events and the soft and hard lockup thresholds. The default threshold
1715 is 10 seconds.
1717 The softlockup threshold is (``2 * watchdog_thresh``). Setting this
1718 tunable to zero will disable lockup detection altogether.

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

문서 범위와 주의 사항

1-31

이 문서는 Rik van Riel과 Shen Feng이 작성한 `/proc/sys/kernel/` sysctl 설명서입니다. 일반 정보와 법적 안내는 `Documentation/admin-guide/sysctl/index.rst`를 참조합니다.

이 디렉터리의 파일은 Linux 커널의 여러 일반 동작을 감시하고 조정합니다. 일부 값은 시스템을 망가뜨릴 수 있으므로 실제 변경 전에 문서와 소스를 모두 읽어야 합니다. 나타나는 항목은 커널 구성에 따라 달라집니다.

acct

32-54
highwater lowwater frequency

BSD 방식 프로세스 회계가 활성화됐을 때 세 값이 동작을 제어합니다. 로그 파일 시스템의 여유 공간이 `lowwater`% 아래로 내려가면 회계를 중지하고, `highwater`% 위로 올라가면 재개합니다. `frequency`는 여유 공간을 다시 검사하는 간격이며 단위는 초입니다.

4 2 30

기본값은 여유 공간이 2% 미만이면 중지하고 4% 이상이면 재개하며, 여유 공간 정보를 30초 동안 유효한 것으로 간주한다는 뜻입니다.

acpi_video_flags

55-67

`Documentation/power/video.rst`를 참조하십시오. `acpi_sleep` 커널 매개변수와 비슷하게 다음 값을 조합해 비디오 resume 모드를 설정합니다.

플래그
`1``s3_bios`
`2``s3_mode`
`4``s3_beep`

arch

68-73

머신 하드웨어 이름입니다. `uname -m`과 같은 결과를 내며 예로 `x86_64`, `aarch64`가 있습니다.

auto_msgmni

74-85

현재는 아무 효과가 없고 앞으로 제거될 수 있으며 읽으면 항상 `0`을 반환합니다. Linux 3.17까지는 메모리 추가·제거 또는 IPC namespace 생성·제거 때 `msgmni`를 자동 재계산할지 제어했습니다. `1`은 활성화, `0`은 비활성화였고 기본값은 `1`이었습니다.

bootloader_type (x86 only)

86-100

bootloader가 알린 형식 번호를 4비트 왼쪽으로 이동한 뒤 bootloader 버전의 하위 4비트와 OR한 값입니다. 예전 커널 헤더의 `type_of_loader` 필드와 일치하던 인코딩을 하위 호환성 때문에 유지합니다. 전체 형식 번호가 `0x15`, 전체 버전이 `0x234`라면 값은 `340 = 0x154`입니다.

자세한 내용은 `Documentation/arch/x86/boot.rst`의 `type_of_loader`와 `ext_loader_type` 필드를 참조하십시오.

bootloader_version (x86 only)

101-110

완전한 bootloader 버전 번호입니다. 앞 예에서는 `564 = 0x234`입니다. 자세한 내용은 `Documentation/arch/x86/boot.rst`의 `type_of_loader`와 `ext_loader_ver` 필드를 참조하십시오.

bpf_stats_enabled

111-124

BPF 프로그램의 총 실행 시간과 실행 횟수 같은 통계를 커널이 수집할지 제어합니다. 수집을 켜면 프로그램을 실행할 때마다 성능이 약간 떨어지며, 통계는 `bpftool`로 볼 수 있습니다.

동작
`0`통계를 수집하지 않음(기본값)
`1`통계를 수집함

cad_pid

125-134

재부팅 때, 특히 Ctrl-Alt-Delete로 재부팅할 때 signal을 받을 PID입니다. 실행 중인 프로세스와 일치하지 않는 값을 쓰면 `-ESRCH`가 발생합니다. `ctrl-alt-del` 절도 참조하십시오.

cap_last_cap

135-143

실행 중인 커널에서 유효한 가장 높은 capability이며 커널의 `CAP_LAST_CAP`을 내보냅니다.

core_pattern

144-188

`core_pattern`은 core dump 파일 이름 패턴을 지정합니다. 최대 길이는 127자이고 기본값은 `core`입니다. `%`로 시작하는 형식 지정자는 실제 값으로 치환됩니다.

하위 호환성을 위해 패턴에 `%p`가 없고 `core_uses_pid`가 설정돼 있으면 파일 이름 뒤에 `.PID`를 붙입니다.

지정자치환 값
`%<NUL>``%`를 버림
`%%``%` 하나 출력
`%p`PID
`%P`init PID namespace의 전역 PID
`%i`TID
`%I`init PID namespace의 전역 TID
`%u`초기 user namespace의 UID
`%g`초기 user namespace의 GID
`%d``PR_SET_DUMPABLE` 및 `/proc/sys/fs/suid_dumpable`과 일치하는 dump 모드
`%s`signal 번호
`%t`dump의 UNIX 시간
`%h`hostname
`%e`실행 파일 이름. 줄어들거나 `prctl` 등으로 바뀔 수 있음
`%f`실행 파일 이름
`%E`실행 파일 경로
`%c``RLIMIT_CORE`가 정한 최대 core 파일 크기
`%C`task가 실행된 CPU
`%F`pidfd 번호
`%<OTHER>``%`와 뒤 문자를 모두 버림

패턴의 첫 문자가 `|`이면 나머지를 실행할 명령으로 취급하고, core dump를 파일 대신 그 프로그램의 표준 입력으로 씁니다.

core_pipe_limit

189-215

이 sysctl은 `core_pattern`이 `|`로 시작해 core를 사용자 공간 helper로 pipe하는 경우에만 적용됩니다. 수집기가 충돌 프로세스의 `/proc/pid` 정보를 읽을 수 있도록 커널은 수집 프로세스가 종료될 때까지 기다려야 합니다.

잘못된 수집기가 종료하지 않으면 충돌 프로세스의 회수를 막을 수 있으므로, 이 값은 동시에 사용자 공간으로 pipe할 수 있는 충돌 프로세스 수를 제한합니다. 초과한 프로세스는 커널 로그에 기록하고 core를 건너뜁니다.

`0`은 병렬 수집 수가 무제한이지만 기다리지 않는 특별한 값입니다. 따라서 수집기가 `/proc/<crashing pid>/`에 접근할 수 있다는 보장이 없습니다. 기본값은 `0`입니다.

core_sort_vma

216-226

기본 coredump는 VMA를 주소 순서로 씁니다. `1`이면 작은 VMA부터 큰 VMA 순으로 씁니다. 적어도 elfutils가 이 형식을 처리하지 못하는 것으로 알려졌지만, 매우 크고 잘린 coredump에서 유용한 디버깅 정보가 작은 VMA에 있을 때 도움이 될 수 있습니다.

core_uses_pid

227-236

기본 coredump 이름은 `core`입니다. `core_uses_pid=1`이면 `core.PID`가 됩니다. `core_pattern`에 `%p`가 없고 이 값이 설정돼 있으면 `.PID`를 붙입니다.

ctrl-alt-del

237-252

`0`이면 Ctrl-Alt-Delete를 가로채 `init(1)`에 보내 정상 재시작을 처리합니다. `0`보다 크면 dirty buffer를 동기화하지 않고 즉시 재부팅합니다.

dosemu 같은 프로그램이 키보드를 raw 모드로 사용하면 커널 tty 계층에 도달하기 전에 그 프로그램이 Ctrl-Alt-Delete를 가로채므로 처리 방식도 해당 프로그램이 결정합니다.

dmesg_restrict

253-266

권한 없는 사용자가 `dmesg(8)`로 커널 로그 버퍼를 볼 수 있는지 제어합니다. `0`은 제한 없음, `1`은 `CAP_SYSLOG`가 있어야 접근 가능하다는 뜻입니다. `CONFIG_SECURITY_DMESG_RESTRICT`가 기본값을 정합니다.

domainname & hostname

267-289

NIS/YP domainname과 hostname을 명령과 같은 방식으로 설정할 수 있습니다.

# echo "darkstar" > /proc/sys/kernel/hostname
# echo "mydomain" > /proc/sys/kernel/domainname

위 명령은 다음과 같은 효과를 냅니다.

# hostname "darkstar"
# domainname "mydomain"

`darkstar.frop.org`의 hostname은 `darkstar`, DNS domainname은 `frop.org`입니다. DNS domainname은 NIS 또는 YP domainname과 혼동하면 안 되며 일반적으로 서로 다릅니다. 자세한 설명은 `hostname(1)`을 참조하십시오.

firmware_config

290-302

`Documentation/driver-api/firmware/fallback-mechanisms.rst`를 참조하십시오. 이 디렉터리는 firmware loader helper fallback을 제어합니다.

항목값 `1`의 의미
`force_sysfs_fallback`fallback 사용을 강제함
`ignore_sysfs_fallback`모든 fallback을 무시함

ftrace_dump_on_oops

303-335

oops 또는 kernel panic 때 `ftrace_dump()`를 호출해 ftrace buffer를 console로 출력할지 정합니다. 충돌까지 이어진 trace를 serial console에 남길 때 유용합니다.

동작
`0`비활성화(기본값)
`1`모든 CPU buffer dump
`2` 또는 `orig_cpu`oops를 일으킨 CPU buffer dump
`<instance>`모든 CPU에서 지정 instance buffer dump
`<instance>=2` 또는 `<instance>=orig_cpu`oops를 일으킨 CPU에서 지정 instance buffer dump

쉼표로 여러 instance를 지정할 수 있습니다. 전역 buffer도 필요하면 전역 모드 `1`, `2`, `orig_cpu`를 맨 앞에 둡니다. 다음은 `foo`와 `bar` instance를 모든 CPU에서 dump합니다.

echo "foo,bar" > /proc/sys/kernel/ftrace_dump_on_oops

다음은 전역 buffer와 `foo`를 모든 CPU에서, `bar`를 oops 발생 CPU에서 dump합니다.

echo "1,foo,bar=2" > /proc/sys/kernel/ftrace_dump_on_oops

ftrace_enabled, stack_tracer_enabled

336-341

`Documentation/trace/ftrace.rst`를 참조하십시오.

hardlockup_all_cpu_backtrace

342-355

hard lockup 감지 때 추가 디버깅 정보를 수집할지 제어합니다. 활성화하면 아키텍처별 all-CPU stack dump를 시작합니다.

동작
`0`아무 작업도 하지 않음(기본값)
`1`감지 시 추가 디버깅 정보 수집

hardlockup_panic

356-370

hard lockup 감지 때 kernel panic을 일으킬지 제어합니다.

동작
`0`panic하지 않음
`1`panic함

자세한 내용은 `Documentation/admin-guide/lockup-watchdogs.rst`를 참조하십시오. `nmi_watchdog` 커널 매개변수로도 설정할 수 있습니다.

hotplug

371-382

hotplug 정책 agent의 경로입니다. 기본값은 `CONFIG_UEVENT_HELPER_PATH`이며, 그 기본값은 빈 문자열입니다. `CONFIG_UEVENT_HELPER`가 켜졌을 때만 존재합니다. 현대 시스템 대부분은 netlink 기반 uevent만 사용하므로 필요하지 않습니다.

hung_task_all_cpu_backtrace

383-396

hung task 감지 때 모든 CPU에 NMI를 보내 backtrace를 dump할지 정합니다. `CONFIG_DETECT_HUNG_TASK`와 `CONFIG_SMP`가 켜졌을 때 나타납니다.

동작
`0`모든 CPU backtrace를 표시하지 않음(기본값)
`1`모든 CPU를 non-maskable interrupt하고 backtrace를 dump함

hung_task_panic

397-408

hung task 감지 때의 커널 동작을 제어하며 `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다.

동작
`0`계속 실행(기본값)
`1`즉시 panic

hung_task_check_count

409-415

검사할 task 수의 상한입니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다.

hung_task_detect_count

416-424

부팅 이후 hung 상태로 감지된 task의 총수입니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다.

hung_task_timeout_secs

425-436

D 상태 task가 이 값보다 오래 schedule되지 않으면 경고합니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다. `0`은 무한 timeout이라 검사하지 않습니다. 설정 범위는 `0`부터 `LONG_MAX/HZ`까지입니다.

hung_task_check_interval_secs

437-450

hung task 검사가 활성화됐을 때 `hung_task_check_interval_secs`초마다 검사합니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다. 기본값 `0`은 검사 간격으로 `hung_task_timeout_secs`를 사용한다는 뜻이며 범위는 `0`부터 `LONG_MAX/HZ`까지입니다.

hung_task_warnings

451-461

보고할 최대 경고 수입니다. 검사 구간에 hung task를 찾을 때마다 1씩 줄고, `0`이면 더는 경고하지 않습니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타나며 `-1`은 무제한 경고입니다.

hyperv_record_panic_msg

462-472

panic kmsg 데이터를 Hyper-V에 보고할지 제어합니다.

동작
`0`보고하지 않음
`1`보고함(기본값)

ignore-unaligned-usertrap

473-488

unaligned 접근이 trap을 일으키고 `CONFIG_SYSCTL_ARCH_UNALIGN_NO_WARN`을 지원하는 아키텍처, 현재 `arc`, `parisc`, `loongarch`에서 모든 unaligned trap을 기록할지 제어합니다.

동작
`0`모든 unaligned 접근 기록
`1`프로세스가 처음 trap될 때만 경고(기본값)

`unaligned-trap` 절도 참조하십시오.

io_uring_disabled

489-507

새 `io_uring` instance 생성을 제한해 커널 공격 표면을 줄입니다.

동작
`0`모든 프로세스가 정상적으로 생성 가능(기본값)
`1``io_uring_group`에 속하지 않은 비특권 프로세스의 `io_uring_setup()`이 `-EPERM`으로 실패함. 기존 instance는 사용 가능
`2`모든 프로세스의 `io_uring_setup()`이 `-EPERM`으로 실패함. 기존 instance는 사용 가능

io_uring_group

508-517

`io_uring_disabled=1`일 때 새 instance를 만들려면 `CAP_SYS_ADMIN`이 있거나 `io_uring_group` 그룹에 속해야 합니다. 기본값 `-1`이면 `CAP_SYS_ADMIN`이 있는 프로세스만 만들 수 있습니다.

kexec_load_disabled

518-531

`kexec_load`와 `kexec_file_load` syscall을 비활성화했는지 나타냅니다. 기본값 `0`은 `kexec_*load` 활성, `1`은 비활성입니다. 한 번 `1`이 되면 되돌릴 수 없고 kexec를 더는 사용할 수 없습니다.

syscall을 막기 전에 kexec image를 적재하면 나중에 사용할 image를 변경하지 못하게 보호할 수 있습니다. 일반적으로 `modules_disabled`와 함께 사용합니다.

kexec_load_limit_panic

532-544

`kexec_load`와 `kexec_file_load`를 crash image와 함께 호출할 수 있는 횟수입니다. 현재보다 더 제한적인 값으로만 바꿀 수 있습니다.

동작
`-1`kexec 호출 무제한(기본값)
`N`남은 호출 횟수

kexec_load_limit_reboot

545-550

`kexec_load_limit_panic`과 같은 기능을 normal image에 적용합니다.

kptr_restrict

551-579

`/proc`와 다른 인터페이스에서 커널 주소를 노출할 때 적용할 제한을 정합니다.

동작
`0`출력 전에 주소를 hash함(기본값, `%p`와 같음)
`1``%pK` 포인터를 출력할 때 사용자가 `CAP_SYSLOG`를 갖고 effective UID/GID가 real UID/GID와 같지 않으면 0으로 바꿈
`2`권한과 무관하게 `%pK` 포인터를 0으로 바꿈

`%pK` 검사는 `open()`이 아니라 `read()` 때 수행하므로, open과 read 사이에 setuid binary 등으로 권한이 상승해도 비특권 사용자에게 포인터가 새지 않게 합니다. 이는 임시 해결책이며 장기적으로는 `open()` 때 검사해야 합니다.

포인터 노출이 우려되면 `%pK`를 사용하는 파일의 world-read 권한을 제거하고, `dmesg(8)`의 `%pK` 노출은 `dmesg_restrict`로 보호하는 방안을 고려하십시오.

modprobe

580-611

커널 module을 자동 적재하는 usermode helper의 전체 경로입니다. 기본값은 `CONFIG_MODPROBE_PATH`이며 그 기본값은 `/sbin/modprobe`입니다. 사용자 공간이 `mount()`에 알 수 없는 파일 시스템 형식을 넘기는 경우처럼 커널이 module을 요청하면 이 binary를 실행하며, helper는 필요한 module을 커널에 삽입해야 합니다.

이 sysctl은 module 자동 적재에만 영향을 주고 명시적인 module 삽입 능력에는 영향을 주지 않습니다. 다음처럼 module 적재 요청을 디버깅할 수 있습니다.

echo '#! /bin/sh' > /tmp/modprobe
echo 'echo "$@" >> /tmp/modprobe.log' >> /tmp/modprobe
echo 'exec /sbin/modprobe "$@"' >> /tmp/modprobe
chmod a+x /tmp/modprobe
echo /tmp/modprobe > /proc/sys/kernel/modprobe

빈 문자열로 설정하면 module 자동 적재를 완전히 끕니다. 커널은 usermode helper를 실행하지 않고 `kernel_module_request` LSM hook도 호출하지 않습니다.

커널이 `CONFIG_STATIC_USERMODEHELPER=y`로 구성됐다면 정적 helper가 이 sysctl보다 우선합니다. 다만 빈 문자열로 자동 적재를 완전히 끄는 동작은 그대로 허용됩니다.

modules_disabled

612-623

모듈식 커널에서 module 적재를 허용할지 나타내는 toggle입니다. 기본값 `0`은 허용이며 `1`로 바꿀 수 있습니다. 한 번 `1`이 되면 module을 적재하거나 제거할 수 없고 `0`으로 되돌릴 수도 없습니다. 일반적으로 `kexec_load_disabled`와 함께 사용합니다.

msgmax, msgmnb, and msgmni

624-639
항목의미와 기본값
`msgmax`IPC message 하나의 최대 크기(byte), 기본 `8192` (`MSGMAX`)
`msgmnb`IPC queue 하나의 최대 크기(byte), 기본 `16384` (`MSGMNB`)
`msgmni`IPC queue 최대 개수, 기본 `32000` (`MSGMNI`)

세 매개변수는 IPC namespace별로 설정됩니다. POSIX message queue의 최대 byte 수는 `RLIMIT_MSGQUEUE`가 제한하며, 각 user namespace에서 계층적으로 적용됩니다.

msg_next_id, sem_next_id, and shm_next_id (System V IPC)

640-656

다음에 할당할 message, semaphore, shared memory IPC 객체의 원하는 ID를 각각 지정합니다. 기본값 `-1`은 일반 할당 로직을 사용한다는 뜻이며 설정 범위는 `0`부터 `INT_MAX`까지입니다.

커널은 새 객체가 요청한 ID를 얻는다고 보장하지 않으므로 잘못된 ID를 처리하는 책임은 사용자 공간에 있습니다. 기본값이 아닌 toggle은 IPC 객체를 성공적으로 할당한 뒤 커널이 `-1`로 되돌립니다. 할당 syscall이 실패하면 값이 유지될지 `-1`로 초기화될지는 정의돼 있지 않습니다.

ngroups_max

657-664

`setgroups`가 허용하는 supplementary group 수의 최댓값이며 커널의 `NGROUPS_MAX`를 내보냅니다.

nmi_watchdog

665-689

x86 시스템에서 NMI watchdog, 즉 hard lockup detector를 제어합니다.

동작
`0`hard lockup detector 비활성화
`1`hard lockup detector 활성화

detector는 각 CPU가 timer interrupt에 응답하는지 감시합니다. CPU가 바쁜 동안에도 주기적으로 NMI를 발생하도록 CPU performance counter register를 설정하므로 NMI watchdog이라고도 부릅니다.

KVM virtual machine의 guest로 실행되는 커널에서는 기본적으로 꺼집니다. guest 커널 command line에 다음 값을 추가해 기본 동작을 덮어쓸 수 있습니다.

nmi_watchdog=1

자세한 내용은 `Documentation/admin-guide/kernel-parameters.rst`를 참조하십시오.

nmi_wd_lpm_factor (PPC only)

690-701

`nmi_watchdog=1`일 때 NMI watchdog timeout에 적용하는 계수입니다. LPM 중 timeout을 계산할 때 `watchdog_thresh`에 더할 백분율을 뜻하며 soft lockup timeout에는 영향을 주지 않습니다.

`0`은 변경 없음입니다. 기본값 `200`은 `watchdog_thresh=10`일 때 NMI watchdog을 30초로 설정합니다.

numa_balancing

702-732

page fault 기반 자동 NUMA memory balancing을 활성화·비활성화하고 모드를 정합니다. 자주 접근하는 node로 memory를 자동 이동하며 다음 값을 OR해 설정할 수 있습니다.

모드
`0``NUMA_BALANCING_DISABLED`
`1``NUMA_BALANCING_NORMAL`
`2``NUMA_BALANCING_MEMORY_TIERING`

`NUMA_BALANCING_NORMAL`은 원격 memory 접근 비용을 줄이기 위해 NUMA node 사이에서 page 배치를 최적화합니다. 커널이 주기적으로 page mapping을 제거하고 뒤따르는 page fault를 trap해 어떤 task thread가 memory에 접근하는지 표본화한 뒤, 해당 데이터를 local memory node로 옮길지 결정합니다.

unmapping과 fault trapping에는 추가 비용이 들며 memory locality 개선이 이를 항상 상쇄한다는 보장은 없습니다. workload가 이미 NUMA node에 bind돼 있다면 이 기능을 꺼야 합니다.

`NUMA_BALANCING_MEMORY_TIERING`은 서로 다른 NUMA node로 표현되는 memory 유형 사이에서 hot page를 빠른 memory에 두도록 최적화하며 역시 unmapping과 page fault를 사용합니다.

numa_balancing_promote_rate_limit_MBps

733-743

서로 다른 memory 유형 사이의 promotion·demotion 처리량이 지나치면 application latency가 나빠질 수 있습니다. 이 값은 node별 최대 promotion 처리량을 MB/s 단위로 제한합니다. 경험적으로 PMEM node 쓰기 bandwidth의 1/10보다 작게 설정합니다.

oops_all_cpu_backtrace

744-759

oops가 발생하면 모든 CPU에 NMI를 보내 backtrace를 dump할지 정합니다. VM 보호 때문에 panic을 일으킬 수 없거나 kdump를 수집할 수 없을 때 마지막 수단으로 사용합니다. `CONFIG_SMP`가 켜졌을 때 나타납니다.

동작
`0`모든 CPU backtrace를 표시하지 않음(기본값)
`1`모든 CPU를 non-maskable interrupt하고 backtrace를 dump함

oops_limit

760-768

`panic_on_oops`가 설정되지 않았을 때 몇 번째 kernel oops 뒤에 panic할지 정합니다. `0`은 횟수 검사를 끄고 `1`은 `panic_on_oops=1`과 같습니다. 기본값은 `10000`입니다.

osrelease, ostype & version

769-788
# cat osrelease
2.1.88
# cat ostype
Linux
# cat version
#5 Wed Feb 25 21:49:24 MET 1998

`osrelease`와 `ostype`의 의미는 명확합니다. `version`의 `#5`는 이 source base로 다섯 번째 빌드한 커널이라는 뜻이며 뒤 날짜는 빌드 시각입니다. 이 값은 커널을 다시 빌드해야만 바꿀 수 있습니다.

overflowgid & overflowuid

789-800

항상 32비트 UID를 지원하지 않았던 `arm`, `i386`, `m68k`, `sh`, `sparc32` 아키텍처에서 실제 UID/GID가 `65535`를 넘을 때 구형 16비트 UID/GID syscall을 쓰는 application에 반환할 고정 UID/GID를 정합니다. 기본값은 `65534`입니다.

panic

801-814

kernel panic 뒤의 동작을 정합니다.

동작
`0`영원히 loop
음수즉시 재부팅
양수해당 초가 지난 뒤 재부팅

software watchdog을 사용할 때 권장값은 `60`입니다.

panic_on_io_nmi

815-830

CPU가 IO error로 발생한 NMI를 받았을 때의 동작을 제어합니다.

동작
`0`계속 실행 시도(기본값)
`1`즉시 panic. 데이터 손상 가능성이 있는 심각한 IO 상태를 계속 실행하지 않고 crash dump를 수집할 수 있음

일부 server는 dump button을 누를 때 이런 NMI를 발생시키므로 이 옵션을 crash dump 수집에 이용할 수 있습니다.

panic_on_oops

831-842

oops 또는 BUG를 만났을 때의 커널 동작을 제어합니다.

동작
`0`계속 실행 시도
`1`즉시 panic. `panic` sysctl도 0이 아니면 machine 재부팅

panic_on_stackoverflow

843-855

user stack을 제외한 kernel, IRQ, exception stack overflow를 감지했을 때의 동작입니다. `CONFIG_DEBUG_STACKOVERFLOW`가 켜졌을 때 나타납니다.

동작
`0`계속 실행 시도
`1`즉시 panic

panic_on_unrecovered_nmi

856-868

memory 또는 알 수 없는 NMI에 대해 Linux 기본 동작은 계속 실행하는 것입니다. 과학 계산처럼 교정되지 않은 parity/ECC error의 전파보다 시스템을 격리하고 오류를 처리하는 편이 나은 환경에서는 panic이 적합할 수 있습니다.

일부 시스템은 전원 관리 같은 예상 밖의 이유로 NMI를 만들 수 있으므로 기본값은 off입니다. 이 sysctl은 같은 디렉터리의 다른 panic 제어와 같은 방식으로 동작합니다.

panic_on_warn

869-880

`1`이면 `WARN()` 경로에서 `panic()`을 호출합니다. WARN 위치에서 kdump를 얻기 위해 커널을 다시 빌드하지 않아도 됩니다.

동작
`0``WARN()`만 수행(기본값)
`1`WARN 위치를 출력한 뒤 `panic()` 호출

panic_print

881-902

panic 때 출력할 시스템 정보의 bitmask입니다. 다음 bit를 조합할 수 있습니다.

비트출력
bit 0모든 task 정보
bit 1시스템 memory 정보
bit 2timer 정보
bit 3`CONFIG_LOCKDEP`가 켜졌으면 lock 정보
bit 4ftrace buffer
bit 5panic 마지막에 모든 kernel message를 console에 재생
bit 6아키텍처가 지원하면 모든 CPU backtrace
bit 7uninterruptible(blocked) 상태 task만 출력

예를 들어 task와 memory 정보를 출력하려면 다음과 같이 설정합니다.

echo 3 > /proc/sys/kernel/panic_print

panic_sys_info

903-920

panic 때 추가로 dump할 정보를 쉼표로 구분한 목록입니다. `tasks,mem,timers,...`처럼 읽기 쉬운 `panic_print` 대안입니다.

출력
`tasks`모든 task 정보
`mem`시스템 memory 정보
`timer`timer 정보
`lock``CONFIG_LOCKDEP`가 켜졌으면 lock 정보
`ftrace`ftrace buffer
`all_bt`아키텍처가 지원하면 모든 CPU backtrace
`blocked_tasks`uninterruptible(blocked) 상태 task만 출력

panic_on_rcu_stall

921-931

`1`이면 RCU stall 감지 메시지를 출력한 뒤 `panic()`을 호출합니다. vmcore로 RCU stall의 root cause를 밝힐 때 유용합니다.

동작
`0`RCU stall 때 panic하지 않음(기본값)
`1`RCU stall 메시지 뒤 panic

max_rcu_stall_to_panic

932-939

`panic_on_rcu_stall=1`일 때 몇 번의 RCU stall 뒤에 `panic()`을 호출할지 정합니다. `panic_on_rcu_stall=0`이면 효과가 없습니다.

perf_cpu_time_max_percent

940-967

perf sampling event 처리에 허용할 CPU 시간의 비율을 커널에 알려 줍니다. sample이 한도를 넘으면 perf subsystem이 sampling frequency를 낮춰 CPU 사용량을 줄입니다.

일부 perf sample은 NMI에서 처리됩니다. 예상보다 오래 걸리면 NMI가 연속으로 쌓여 다른 작업이 실행되지 못할 수 있습니다.

동작
`0`감시와 sample rate 보정을 끔
`1-100`perf sample rate가 해당 CPU 비율을 넘지 않도록 throttle을 시도

커널은 sample event의 예상 길이를 계산하므로 `100`은 그 예상 길이의 100%라는 뜻입니다. 실제 길이가 넘으면 `100`에서도 throttle될 수 있으며 CPU 소비량을 전혀 제한하지 않으려면 `0`을 사용합니다.

perf_event_paranoid

968-996

`CAP_PERFMON`이 없는 비특권 사용자의 performance event system 사용을 제어합니다. 기본값은 `2`입니다. 하위 호환성을 위해 `CAP_SYS_ADMIN` 프로세스에도 시스템 performance monitoring 접근을 허용하지만, 보안 목적에는 `CAP_PERFMON` 사용이 권장됩니다.

제한
`-1`거의 모든 event를 모든 사용자에게 허용하고 `CAP_IPC_LOCK` 없이 `perf_event_mlock_kb` 이후 mlock 한도를 무시
`>=0``CAP_PERFMON` 없는 사용자의 ftrace function tracepoint와 raw tracepoint 접근 금지
`>=1``CAP_PERFMON` 없는 사용자의 CPU event 접근 금지
`>=2``CAP_PERFMON` 없는 사용자의 kernel profiling 금지

perf_event_max_stack

997-1009

`attr.sample_type & PERF_SAMPLE_CALLCHAIN` event에서 복사할 최대 stack frame 수를 정합니다. 예로 `perf record -g`, `perf trace --call-graph fp`가 있습니다.

callchain이 활성화된 event가 사용 중이면 바꿀 수 없고 쓰기가 `-EBUSY`를 반환합니다. 기본값은 `127`입니다.

perf_event_mlock_kb

1010-1017

mlock 한도에 계산하지 않는 CPU별 ring buffer 크기를 제어합니다. 기본값은 `512 + 1 page`입니다.

perf_event_max_contexts_per_stack

1018-1030

`attr.sample_type & PERF_SAMPLE_CALLCHAIN` event의 최대 stack frame context entry 수를 정합니다. 예로 `perf record -g`, `perf trace --call-graph fp`가 있습니다.

callchain이 활성화된 event가 사용 중이면 바꿀 수 없고 쓰기가 `-EBUSY`를 반환합니다. 기본값은 `8`입니다.

perf_user_access (arm64 and riscv only)

1031-1056

performance event counter를 읽는 사용자 공간 접근을 제어합니다.

arm64의 기본값은 `0`으로 접근 비활성화입니다. `1`이면 사용자 공간이 performance monitor counter register를 직접 읽을 수 있습니다. 자세한 내용은 `Documentation/arch/arm64/perf.rst`를 참조하십시오.

RISC-V에서 `0`은 사용자 공간 접근 비활성화입니다. 기본값 `1`은 perf를 통해 counter register를 읽게 하며 perf 개입 없는 직접 접근은 illegal instruction을 일으킵니다.

RISC-V 값 `2`는 legacy mode로 cycle과 `insret` CSR만 직접 접근할 수 있습니다. 이 값은 deprecated이며 사용자 공간 application이 모두 수정되면 제거됩니다. time CSR은 모든 모드에서 항상 직접 접근할 수 있습니다.

pid_max

1057-1064

PID 할당이 되감기는 값입니다. 다음 PID가 이 값에 이르면 최소 PID로 돌아가며 `pid_max` 이상인 PID는 할당하지 않습니다.

ns_last_pid

1065-1072

이 sysctl을 사용하는 task가 속한 현재 PID namespace에서 마지막으로 할당한 PID입니다. fork로 다음 task의 PID를 고를 때 커널은 이 번호부터 할당을 시도합니다.

powersave-nap (PPC only)

1073-1081

설정하면 Linux-PPC가 절전 `nap` 모드를 사용하고, 설정하지 않으면 `doze` 모드를 사용합니다.

printk

1082-1103

네 값은 순서대로 `console_loglevel`, `default_message_loglevel`, `minimum_console_loglevel`, `default_console_loglevel`입니다. 오류 message를 출력하거나 기록할 때 `printk()` 동작에 영향을 줍니다. loglevel은 `man 2 syslog`를 참조하십시오.

항목의미
`console_loglevel`이 값보다 우선순위가 높은 message를 console에 출력
`default_message_loglevel`명시적 우선순위가 없는 message에 적용할 우선순위
`minimum_console_loglevel``console_loglevel`에 설정할 수 있는 최소(가장 높은) 값
`default_console_loglevel``console_loglevel`의 기본값

printk_delay

1104-1111

각 printk message를 `printk_delay` millisecond만큼 지연합니다. 허용 범위는 `0-10000`입니다.

printk_ratelimit

1112-1121

일부 warning message에는 rate limit이 적용됩니다. 이 값은 message 사이의 최소 시간(초)을 정하며 기본값은 5초입니다. `0`은 rate limiting을 끕니다.

printk_ratelimit_burst

1122-1133

장기적으로 `printk_ratelimit`초마다 message 하나를 허용하지만 짧은 burst도 허용합니다. 이 값은 rate limiting이 시작되기 전에 보낼 수 있는 message 수이며, `printk_ratelimit`초 뒤 다시 같은 수의 burst를 보낼 수 있습니다. 기본값은 10개입니다.

printk_devkmsg

1134-1151

사용자 공간에서 `/dev/kmsg`로 기록하는 동작을 제어합니다.

동작
`ratelimit`기본값, rate limit 적용
`on`사용자 공간의 `/dev/kmsg` 기록 무제한
`off`사용자 공간의 `/dev/kmsg` 기록 비활성화

kernel command line의 `printk.devkmsg=`가 이 값을 덮어쓰며 다음 재부팅까지 한 번만 설정됩니다. 한 번 command line으로 설정하면 이 sysctl로 더는 바꿀 수 없습니다.

pty

1152-1157

`Documentation/filesystems/devpts.rst`를 참조하십시오.

random

1158-1183

다음 항목을 포함하는 디렉터리입니다.

항목의미
`boot_id`처음 읽을 때 생성되고 이후 바뀌지 않는 UUID
`uuid`읽을 때마다 생성되는 UUID. 필요할 때 UUID를 만드는 데 사용 가능
`entropy_avail`pool의 entropy 수(bit)
`poolsize`entropy pool 크기(bit)
`urandom_min_reseed_secs`obsolete. 예전에는 urandom pool reseed 사이의 최소 초를 정함. 호환성을 위해 쓸 수 있지만 RNG 동작에는 영향 없음
`write_wakeup_threshold`entropy가 이 bit 수 아래로 내려갈 때 `/dev/random`에 쓰려고 기다리는 프로세스를 깨움. 호환성을 위해 쓸 수 있지만 RNG 동작에는 영향 없음

randomize_va_space

1184-1217

기능을 지원하는 아키텍처에서 사용할 프로세스 주소 공간 randomization 유형을 고릅니다.

동작
`0`주소 공간 randomization을 끔. 기능 미지원 아키텍처와 `norandmaps`로 부팅한 커널의 기본값
`1`mmap base, stack, VDSO page 주소를 randomize함. shared library와 PIE binary의 code 시작 위치도 randomize함. `CONFIG_COMPAT_BRK`가 켜졌을 때 기본값
`2`heap randomization도 추가함. `CONFIG_COMPAT_BRK`가 꺼졌을 때 기본값

1996년의 일부 오래된 libc.so.5처럼 brk 영역이 code+bss 바로 뒤에서 시작한다고 가정하는 legacy application은 brk randomization으로 고장날 수 있습니다. 알려진 비legacy application 문제는 없으므로 대부분의 시스템에서는 full randomization이 안전합니다.

오래됐거나 잘못된 binary가 있는 시스템은 `CONFIG_COMPAT_BRK`를 켜서 heap을 프로세스 주소 공간 randomization에서 제외해야 합니다.

real-root-dev

1218-1223

`Documentation/admin-guide/initrd.rst`를 참조하십시오.

reboot-cmd (SPARC only)

1224-1231

문서 원문도 확정하지 못한 항목입니다. SPARC ROM/Flash boot loader에 인수를 전달해 재부팅 뒤 동작을 지시하는 방법으로 보입니다.

sched_energy_aware

1232-1242

Energy Aware Scheduling(EAS)을 켜거나 끕니다. asymmetric CPU topology와 Energy Model이 있어 EAS를 실행할 수 있는 platform에서는 자동으로 시작합니다. 요건을 충족하지만 사용하지 않으려면 `0`으로 바꿉니다. EAS를 지원하지 않는 platform에서는 쓰기가 실패하고 읽기는 아무것도 반환하지 않습니다.

task_delayacct

1243-1250

task delay accounting을 켜거나 끕니다. 자세한 내용은 `Documentation/accounting/delay-accounting.rst`를 참조하십시오. scheduler에 작은 overhead가 생기지만 디버깅과 성능 조정에 유용하고 `iotop` 같은 일부 도구가 필요로 합니다.

sched_schedstats

1251-1257

scheduler 통계를 켜거나 끕니다. scheduler에 작은 overhead가 생기지만 디버깅과 성능 조정에 유용합니다.

sched_util_clamp_min

1258-1268

허용할 minimum utilization의 최댓값입니다. 기본값 `1024`는 가능한 최대치입니다. 요청한 `uclamp.min`은 이 값보다 클 수 없으며 `[0:sched_util_clamp_min]` 범위로 제한됩니다.

sched_util_clamp_max

1269-1279

허용할 maximum utilization의 최댓값입니다. 기본값 `1024`는 가능한 최대치입니다. 요청한 `uclamp.max`는 이 값보다 클 수 없으며 `[0:sched_util_clamp_max]` 범위로 제한됩니다.

sched_util_clamp_min_rt_default

1280-1312

Linux는 기본적으로 성능에 맞춰 조정돼 있어 RT task가 항상 가장 높은 frequency와 heterogeneous system에서 가장 큰 capacity의 CPU에서 실행됩니다. uclamp는 모든 RT task의 요청 `uclamp.min`을 기본 `1024`로 설정해 이를 구현합니다.

이 knob는 uclamp를 사용할 때 관리자가 기본 동작을 바꿀 수 있게 합니다. 특히 battery 장치에서 최대 capacity와 frequency는 에너지 소비를 늘리고 battery 수명을 줄입니다.

사용자가 `sched_setattr()` syscall로 요청 `uclamp.min`을 바꾸지 않은 RT task에만 효과가 있으며, 위의 `sched_util_clamp_min` 범위 제약을 넘을 수 없습니다.

예를 들어 다음과 같이 설정했다고 가정합니다.

sched_util_clamp_min_rt_default = 800
sched_util_clamp_min = 600

`800`은 허용 범위 `[0:600]` 밖이므로 boost는 `600`으로 clamp됩니다. 절전 모드가 `sched_util_clamp_min`을 바꿔 모든 boost를 임시 제한할 때 이런 상황이 생길 수 있습니다. 제한을 해제하면 요청한 `sched_util_clamp_min_rt_default`가 다시 적용됩니다.

seccomp

1313-1318

`Documentation/userspace-api/seccomp_filter.rst`를 참조하십시오.

sg-big-buff

1319-1331

generic SCSI(sg) buffer 크기를 표시합니다. 아직 런타임에 조정할 수 없지만 빌드할 때 `include/scsi/sg.h`의 `SG_BIG_BUFF`를 바꿀 수 있습니다. 보통 바꿀 이유는 없습니다.

shmall

1332-1349

IPC namespace 안에서 사용할 수 있는 shared memory page 총량입니다. page 수는 namespace마다 따로 계산되고 상속되지 않으므로 `shmall`은 항상 `ceil(shmmax/PAGE_SIZE)` 이상이어야 합니다.

시스템의 기본 `PAGE_SIZE`를 모르면 다음 명령으로 확인합니다.

# getconf PAGE_SIZE

shared memory 할당 능력을 줄이거나 없애려면 새 IPC namespace를 만들고 이 값을 필요한 수준으로 설정한 뒤 현재 user namespace에서 새 IPC namespace 생성을 금지해야 합니다. cgroup을 사용할 수도 있습니다.

shmmax

1350-1358

생성 가능한 shared memory segment 최대 크기의 런타임 한도를 조회하고 설정합니다. 커널은 최대 1GiB segment를 지원하며 기본값은 `SHMMAX`입니다.

shmmni

1359-1365

shared memory segment의 최대 개수입니다. 기본값은 `4096` (`SHMMNI`)입니다.

shm_rmid_forced

1366-1386

`setrlimit(2)`은 한 프로세스의 memory 사용량 같은 자원 한도를 설정하지만, shared memory segment는 어떤 프로세스와도 연결되지 않은 채 존재할 수 있어 한도에 계산되지 않을 수 있습니다.

활성화하면 detach나 프로세스 종료 뒤 attach count가 0이 된 segment를 자동 파괴합니다. 생성했지만 한 번도 attach하지 않은 segment도 생성 프로세스가 종료하면 파괴합니다. `IPC_RMID`는 attach되지 않은 segment를 즉시 파괴하는 용도로만 남습니다.

정의된 System V 동작을 깨므로 일부 application이 멈출 수 있습니다. `RLIMIT_AS`, `RLIMIT_NPROC` 같은 자원 한도도 함께 구성하지 않으면 유용하지 않으며 대부분의 시스템에는 필요하지 않습니다.

`0`에서 `1`로 바꾸면 사용자가 없고 생성한 프로세스가 이미 죽은 기존 segment도 파괴됩니다.

sysctl_writes_strict

1387-1407

`/proc/sys`로 sysctl 값을 갱신할 때 file position이 쓰기 동작에 미치는 영향을 제어합니다.

동작
`-1`legacy per-write 처리. printk 경고 없음. 각 `write` syscall이 값 전체를 포함해야 하며 같은 descriptor의 여러 쓰기는 position과 무관하게 값을 다시 씀
`0``-1`과 같지만 file position이 0이 아닐 때 쓰는 프로세스를 경고
`1`기본값. 문자열 sysctl은 file position을 존중하고 여러 쓰기를 buffer에 append함. 최대 길이 뒤는 무시함. 숫자 sysctl은 항상 position 0에서 값 전체를 한 write buffer에 담아야 함

softlockup_all_cpu_backtrace

1408-1424

soft lockup 감지 때 detector thread가 추가 디버깅 정보를 수집할지 제어합니다. 활성화하면 각 CPU에 NMI를 보내 stack trace를 수집합니다. NMI를 지원하는 아키텍처에만 적용됩니다.

동작
`0`아무 작업도 하지 않음(기본값)
`1`감지 시 추가 디버깅 정보 수집

softlockup_panic

1425-1438

soft lockup 감지 때 kernel panic을 일으킬지 제어합니다.

동작
`0`panic하지 않음
`1`panic함

`softlockup_panic` kernel parameter로도 설정할 수 있습니다.

soft_watchdog

1439-1456

soft lockup detector를 제어합니다.

동작
`0`soft lockup detector 비활성화
`1`soft lockup detector 활성화

detector는 자발적으로 reschedule하지 않고 CPU를 독점해 `migration/N` thread 실행과 watchdog work를 막는 thread를 감시합니다. watchdog timer function이 work를 queue하려면 CPU가 timer interrupt에 응답해야 합니다. 그렇지 않으면 NMI watchdog이 활성화된 경우 hard lockup을 감지할 수 있습니다.

split_lock_mitigate (x86 only)

1457-1479

x86의 split lock 하나마다 시스템 전체 성능 비용이 생깁니다. 큰 시스템에서 비특권 사용자가 많은 split lock을 만들면 정상 사용자에게 denial of service를 일으킬 수 있습니다.

커널은 split lock을 감지해 기다리게 하고 한 번에 core 하나만 split lock을 실행하도록 벌점을 줍니다. 문제가 있는 application이 지나치게 느려질 수 있지만 완화를 끄면 split lock 사용자의 denial of service에 더 노출됩니다.

동작
`0`완화 비활성화. kernel log에 경고만 남기며 denial of service에 노출
`1`완화 활성화(기본값). 의도적 성능 저하로 split lock 사용자를 제약

stack_erasing

1480-1497

`CONFIG_KSTACK_ERASE`로 빌드한 커널에서 syscall 끝의 kernel stack 지우기를 제어합니다. stack leak bug가 드러낼 수 있는 정보를 줄이고 초기화되지 않은 stack 변수 공격 일부를 막습니다.

대가로 성능이 떨어집니다. single-CPU 시스템의 커널 컴파일에서는 약 1% 느려졌으며 다른 시스템과 workload는 달라질 수 있습니다.

동작
`0`kernel stack 지우기 비활성화, `KSTACK_ERASE_METRICS`도 갱신하지 않음
`1`활성화(기본값). syscall 끝에 사용자 공간으로 돌아가기 전에 수행

stop-a (SPARC only)

1498-1511

Stop-A 동작을 제어합니다.

동작
`0`Stop-A가 효과 없음
`1`PROM으로 진입(기본값)

panic 때는 사용자가 boot PROM으로 돌아갈 수 있도록 Stop-A가 항상 활성화됩니다.

sysrq

1512-1517

`Documentation/admin-guide/sysrq.rst`를 참조하십시오.

tainted

1518-1555

커널이 taint됐으면 0이 아닌 값입니다. 숫자는 OR할 수 있고 문자는 Oops 보고서의 `Tainted` 줄에 나타납니다.

문자의미
`1``P`proprietary module 적재
`2``F`module 강제 적재
`4``S`규격 밖 시스템에서 실행
`8``R`module 강제 제거
`16``M`processor가 Machine Check Exception(MCE) 보고
`32``B`잘못된 page 참조 또는 예상 밖 page flag
`64``U`사용자 공간 application이 taint 요청
`128``D`최근 OOPS 또는 BUG로 kernel 사망
`256``A`사용자가 ACPI table 덮어씀
`512``W`kernel warning 발생
`1024``C`staging driver 적재
`2048``I`platform firmware bug workaround 적용
`4096``O`외부 빌드(out-of-tree) module 적재
`8192``E`서명되지 않은 module 적재
`16384``L`soft lockup 발생
`32768``K`kernel live patch 적용
`65536``X`배포판이 정의하고 사용하는 auxiliary taint
`131072``T`struct randomization plugin으로 kernel 빌드

자세한 내용은 `Documentation/admin-guide/tainted-kernels.rst`를 참조하십시오.

command line에 `panic_on_taint=<bitmask>,nousertaint`를 지정해 부팅했고 `tainted`에 쓰려는 OR 값이 `panic_on_taint` bitmask와 겹치면 쓰기가 `EINVAL`로 실패합니다. 자세한 내용은 `Documentation/admin-guide/kernel-parameters.rst`의 해당 command line 옵션과 `nousertaint` switch를 참조하십시오.

threads-max

1556-1573

`fork()`로 만들 수 있는 thread의 최대 개수입니다. 초기화 때 커널은 최대 thread를 만들어도 thread 구조체가 사용 가능한 RAM page의 1/8만 차지하도록 값을 정합니다.

쓸 수 있는 최솟값은 `1`, 최댓값은 `FUTEX_TID_MASK` (`0x3fffffff`)입니다. 범위 밖 값을 쓰면 `EINVAL`이 발생합니다.

timer_migration

1574-1581

0이 아니면 idle CPU의 timer를 다른 CPU로 옮겨 저전력 상태를 더 오래 유지하도록 시도합니다. 기본값은 `1`입니다.

traceoff_on_warning

1582-1588

설정하면 `WARN()` 발생 때 tracing을 끕니다. tracing은 `Documentation/trace/ftrace.rst`를 참조하십시오.

tracepoint_printk

1589-1608

`tp_printk` boot parameter로 tracepoint를 `printk()`에 보내도록 켰을 때 런타임 동작을 제어합니다. 다음은 전송을 중지합니다.

echo 0 > /proc/sys/kernel/tracepoint_printk

다음은 다시 `printk()`로 보냅니다.

echo 1 > /proc/sys/kernel/tracepoint_printk

커널을 `tp_printk` 활성 상태로 부팅한 경우에만 동작합니다. `Documentation/admin-guide/kernel-parameters.rst`와 `Documentation/trace/boottime-trace.rst`를 참조하십시오.

unaligned-trap

1609-1624

unaligned 접근이 trap을 일으키고 `CONFIG_SYSCTL_ARCH_UNALIGN_ALLOW`를 지원하는 아키텍처, 현재 `arc`, `parisc`, `loongarch`에서 trap을 잡아 실패 대신 emulate할지 정합니다.

동작
`0`unaligned 접근을 emulate하지 않음
`1`unaligned 접근을 emulate함(기본값)

`ignore-unaligned-usertrap` 절도 참조하십시오.

unknown_nmi_panic

1625-1635

NMI 처리 동작에 영향을 줍니다. 0이 아니면 알 수 없는 NMI를 trap한 뒤 panic하고 kernel 디버깅 정보를 console에 표시합니다. 많은 IA32 server의 NMI switch가 이런 NMI를 발생시키므로 시스템이 멈췄을 때 switch를 눌러 볼 수 있습니다.

unprivileged_bpf_disabled

1636-1657

`1`을 쓰면 비특권 `bpf()` 호출을 비활성화합니다. 이후 `CAP_SYS_ADMIN` 또는 `CAP_BPF` 없는 호출은 `-EPERM`을 반환하며 실행 중인 커널에서 다시 해제할 수 없습니다.

`2`도 비특권 호출을 끄지만 관리자가 나중에 `0` 또는 `1`로 바꿀 수 있습니다. 커널 구성에 `BPF_UNPRIV_DEFAULT_OFF`가 켜졌다면 기본값은 `0` 대신 `2`입니다.

동작
`0`비특권 `bpf()` 호출 활성화
`1`비특권 `bpf()` 호출을 복구 불가능하게 비활성화
`2`비특권 `bpf()` 호출 비활성화, 관리자가 변경 가능

warn_limit

1658-1666

`panic_on_warn`이 설정되지 않았을 때 몇 번째 kernel warning 뒤에 panic할지 정합니다. `0`은 횟수 검사를 끄고 `1`은 `panic_on_warn=1`과 같습니다. 기본값은 `0`입니다.

watchdog

1667-1688

soft lockup detector와 NMI watchdog(hard lockup detector)을 동시에 켜거나 끕니다.

동작
`0`두 lockup detector 모두 비활성화
`1`두 lockup detector 모두 활성화

`soft_watchdog`와 `nmi_watchdog`로 각각 따로 제어할 수도 있습니다. 다음처럼 `watchdog`를 읽으면 두 값의 논리 OR 결과인 `0` 또는 `1`을 출력합니다.

cat /proc/sys/kernel/watchdog

watchdog_cpumask

1689-1709

watchdog가 실행될 CPU를 제어합니다. 기본 cpumask는 가능한 모든 core입니다. 커널에 `NO_HZ_FULL`이 켜졌고 `nohz_full=` boot argument로 core를 지정했다면 해당 core는 기본적으로 제외됩니다.

offline core도 mask에 넣을 수 있으며 나중에 online이 되면 mask에 따라 watchdog을 시작합니다. 보통 `nohz_full` 환경에서 watchdog이 기본적으로 없던 core의 kernel lockup이 의심될 때 다시 활성화하는 용도로만 바꿉니다.

값은 표준 cpumask cpulist 형식입니다. 예를 들어 core 0, 2, 3, 4에서 watchdog을 켜려면 다음과 같이 씁니다.

echo 0,2-4 > /proc/sys/kernel/watchdog_cpumask

watchdog_thresh

1710-1718

hrtimer와 NMI event frequency, soft·hard lockup threshold를 제어합니다. 기본 threshold는 10초이고 soft lockup threshold는 `2 * watchdog_thresh`입니다. `0`으로 설정하면 lockup detection을 모두 끕니다.