← Documents Documentation/trace/ftrace.rst GitHub 원문 ↗

Linux 6.18.37 · Tracing

ftrace 함수 추적 안내

Ftrace의 tracefs 제어 파일, tracer와 출력 옵션, latency·function·function graph 추적, 동적 함수 필터, trace command, 실시간 trace_pipe, buffer·snapshot·instance·stack tracer 운용을 Linux v6.18.37 원문 전체에 맞춰 설명합니다.

Source pathDocumentation/trace/ftrace.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

ftrace.rst:1-3801

Ftrace의 tracefs 제어 파일, tracer와 출력 옵션, latency·function·function graph 추적, 동적 함수 필터, trace command, 실시간 trace_pipe, buffer·snapshot·instance·stack tracer 운용을 Linux v6.18.37 원문 전체에 맞춰 설명합니다.

운용 순서는 tracefs를 확인한 뒤 tracer 또는 event를 고르고, 함수·PID·CPU 필터로 범위를 줄이며, buffer 크기와 overwrite 정책을 정하고, `tracing_on`으로 필요한 구간만 기록하는 방식이 기본이다. 장시간 수집은 `trace_pipe`, 특정 시점 보존은 `snapshot`, 서로 독립된 관찰은 `instances`, 최대 커널 스택 사용량 확인은 stack tracer를 사용한다.

동적 ftrace는 빌드 시 수집한 `mcount`·`__fentry__` 호출 지점을 부팅 때 NOP로 바꾸고 선택한 함수만 다시 ftrace 호출로 패치하여 비활성 상태의 비용을 줄인다. 전역 `ftrace_enabled`는 perf와 kprobes를 포함한 다른 사용자에게도 영향을 주므로 마지막 수단으로 주의해서 변경해야 한다.

ftrace 진단 흐름
목적주요 인터페이스권장 방식
함수 실행 경로`function`, `function_graph`함수·PID·graph 필터로 범위 제한
IRQ·preemption 지연`irqsoff`, `preemptoff`최대 지연 trace와 임계값 확인
실시간 장기 수집`trace_pipe`소비형 reader로 지속 저장
특정 시점 보존`snapshot`buffer swap 뒤 추적 계속
독립 관찰`instances`instance별 event와 buffer 분리
스택 한계 조사stack tracer`stack_max_size`와 `stack_trace` 확인

문제 성격에 맞는 tracer와 buffer 운용 방식을 선택한다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 ========================
2 ftrace - Function Tracer
3 ========================
4
5 Copyright 2008 Red Hat Inc.
6
7 :Author: Steven Rostedt <[email protected]>
8 :License: The GNU Free Documentation License, Version 1.2
9 (dual licensed under the GPL v2)
10 :Original Reviewers: Elias Oltmanns, Randy Dunlap, Andrew Morton,
11 John Kacur, and David Teigland.
12
13 - Written for: 2.6.28-rc2
14 - Updated for: 3.10
15 - Updated for: 4.13 - Copyright 2017 VMware Inc. Steven Rostedt
16 - Converted to rst format - Changbin Du <[email protected]>
17
18 Introduction
19 ------------
20
21 Ftrace is an internal tracer designed to help out developers and
22 designers of systems to find what is going on inside the kernel.
23 It can be used for debugging or analyzing latencies and
24 performance issues that take place outside of user-space.
25
26 Although ftrace is typically considered the function tracer, it
27 is really a framework of several assorted tracing utilities.
28 There's latency tracing to examine what occurs between interrupts
29 disabled and enabled, as well as for preemption and from a time
30 a task is woken to the task is actually scheduled in.
31
32 One of the most common uses of ftrace is the event tracing.
33 Throughout the kernel are hundreds of static event points that
34 can be enabled via the tracefs file system to see what is
35 going on in certain parts of the kernel.
36
37 See events.rst for more information.
38
39
40 Implementation Details
41 ----------------------
42
43 See Documentation/trace/ftrace-design.rst for details for arch porters and such.
44
45
46 The File System
47 ---------------
48
49 Ftrace uses the tracefs file system to hold the control files as
50 well as the files to display output.
51
52 When tracefs is configured into the kernel (which selecting any ftrace
53 option will do) the directory /sys/kernel/tracing will be created. To mount
54 this directory, you can add to your /etc/fstab file::
55
56 tracefs /sys/kernel/tracing tracefs defaults 0 0
57
58 Or you can mount it at run time with::
59
60 mount -t tracefs nodev /sys/kernel/tracing
61
62 For quicker access to that directory you may want to make a soft link to
63 it::
64
65 ln -s /sys/kernel/tracing /tracing
66
67 .. attention::
68
69 Before 4.1, all ftrace tracing control files were within the debugfs
70 file system, which is typically located at /sys/kernel/debug/tracing.
71 For backward compatibility, when mounting the debugfs file system,
72 the tracefs file system will be automatically mounted at:
73
74 /sys/kernel/debug/tracing
75
76 All files located in the tracefs file system will be located in that
77 debugfs file system directory as well.
78
79 .. attention::
80
81 Any selected ftrace option will also create the tracefs file system.
82 The rest of the document will assume that you are in the ftrace directory
83 (cd /sys/kernel/tracing) and will only concentrate on the files within that
84 directory and not distract from the content with the extended
85 "/sys/kernel/tracing" path name.
86
87 That's it! (assuming that you have ftrace configured into your kernel)
88
89 After mounting tracefs you will have access to the control and output files
90 of ftrace. Here is a list of some of the key files:
91
92
93 Note: all time values are in microseconds.
94
95 current_tracer:
96
97 This is used to set or display the current tracer
98 that is configured. Changing the current tracer clears
99 the ring buffer content as well as the "snapshot" buffer.
100
101 available_tracers:
102
103 This holds the different types of tracers that
104 have been compiled into the kernel. The
105 tracers listed here can be configured by
106 echoing their name into current_tracer.
107
108 tracing_on:
109
110 This sets or displays whether writing to the trace
111 ring buffer is enabled. Echo 0 into this file to disable
112 the tracer or 1 to enable it. Note, this only disables
113 writing to the ring buffer, the tracing overhead may
114 still be occurring.
115
116 The kernel function tracing_off() can be used within the
117 kernel to disable writing to the ring buffer, which will
118 set this file to "0". User space can re-enable tracing by
119 echoing "1" into the file.
120
121 Note, the function and event trigger "traceoff" will also
122 set this file to zero and stop tracing. Which can also
123 be re-enabled by user space using this file.
124
125 trace:
126
127 This file holds the output of the trace in a human
128 readable format (described below). Opening this file for
129 writing with the O_TRUNC flag clears the ring buffer content.
130 Note, this file is not a consumer. If tracing is off
131 (no tracer running, or tracing_on is zero), it will produce
132 the same output each time it is read. When tracing is on,
133 it may produce inconsistent results as it tries to read
134 the entire buffer without consuming it.
135
136 trace_pipe:
137
138 The output is the same as the "trace" file but this
139 file is meant to be streamed with live tracing.
140 Reads from this file will block until new data is
141 retrieved. Unlike the "trace" file, this file is a
142 consumer. This means reading from this file causes
143 sequential reads to display more current data. Once
144 data is read from this file, it is consumed, and
145 will not be read again with a sequential read. The
146 "trace" file is static, and if the tracer is not
147 adding more data, it will display the same
148 information every time it is read.
149
150 trace_options:
151
152 This file lets the user control the amount of data
153 that is displayed in one of the above output
154 files. Options also exist to modify how a tracer
155 or events work (stack traces, timestamps, etc).
156
157 options:
158
159 This is a directory that has a file for every available
160 trace option (also in trace_options). Options may also be set
161 or cleared by writing a "1" or "0" respectively into the
162 corresponding file with the option name.
163
164 tracing_max_latency:
165
166 Some of the tracers record the max latency.
167 For example, the maximum time that interrupts are disabled.
168 The maximum time is saved in this file. The max trace will also be
169 stored, and displayed by "trace". A new max trace will only be
170 recorded if the latency is greater than the value in this file
171 (in microseconds).
172
173 By echoing in a time into this file, no latency will be recorded
174 unless it is greater than the time in this file.
175
176 tracing_thresh:
177
178 Some latency tracers will record a trace whenever the
179 latency is greater than the number in this file.
180 Only active when the file contains a number greater than 0.
181 (in microseconds)
182
183 buffer_percent:
184
185 This is the watermark for how much the ring buffer needs to be filled
186 before a waiter is woken up. That is, if an application calls a
187 blocking read syscall on one of the per_cpu trace_pipe_raw files, it
188 will block until the given amount of data specified by buffer_percent
189 is in the ring buffer before it wakes the reader up. This also
190 controls how the splice system calls are blocked on this file::
191
192 0 - means to wake up as soon as there is any data in the ring buffer.
193 50 - means to wake up when roughly half of the ring buffer sub-buffers
194 are full.
195 100 - means to block until the ring buffer is totally full and is
196 about to start overwriting the older data.
197
198 buffer_size_kb:
199
200 This sets or displays the number of kilobytes each CPU
201 buffer holds. By default, the trace buffers are the same size
202 for each CPU. The displayed number is the size of the
203 CPU buffer and not total size of all buffers. The
204 trace buffers are allocated in pages (blocks of memory
205 that the kernel uses for allocation, usually 4 KB in size).
206 A few extra pages may be allocated to accommodate buffer management
207 meta-data. If the last page allocated has room for more bytes
208 than requested, the rest of the page will be used,
209 making the actual allocation bigger than requested or shown.
210 ( Note, the size may not be a multiple of the page size
211 due to buffer management meta-data. )
212
213 Buffer sizes for individual CPUs may vary
214 (see "per_cpu/cpu0/buffer_size_kb" below), and if they do
215 this file will show "X".
216
217 buffer_total_size_kb:
218
219 This displays the total combined size of all the trace buffers.
220
221 buffer_subbuf_size_kb:
222
223 This sets or displays the sub buffer size. The ring buffer is broken up
224 into several same size "sub buffers". An event can not be bigger than
225 the size of the sub buffer. Normally, the sub buffer is the size of the
226 architecture's page (4K on x86). The sub buffer also contains meta data
227 at the start which also limits the size of an event. That means when
228 the sub buffer is a page size, no event can be larger than the page
229 size minus the sub buffer meta data.
230
231 Note, the buffer_subbuf_size_kb is a way for the user to specify the
232 minimum size of the subbuffer. The kernel may make it bigger due to the
233 implementation details, or simply fail the operation if the kernel can
234 not handle the request.
235
236 Changing the sub buffer size allows for events to be larger than the
237 page size.
238
239 Note: When changing the sub-buffer size, tracing is stopped and any
240 data in the ring buffer and the snapshot buffer will be discarded.
241
242 free_buffer:
243
244 If a process is performing tracing, and the ring buffer should be
245 shrunk "freed" when the process is finished, even if it were to be
246 killed by a signal, this file can be used for that purpose. On close
247 of this file, the ring buffer will be resized to its minimum size.
248 Having a process that is tracing also open this file, when the process
249 exits its file descriptor for this file will be closed, and in doing so,
250 the ring buffer will be "freed".
251
252 It may also stop tracing if disable_on_free option is set.
253
254 tracing_cpumask:
255
256 This is a mask that lets the user only trace on specified CPUs.
257 The format is a hex string representing the CPUs.
258
259 set_ftrace_filter:
260
261 When dynamic ftrace is configured in (see the
262 section below "dynamic ftrace"), the code is dynamically
263 modified (code text rewrite) to disable calling of the
264 function profiler (mcount). This lets tracing be configured
265 in with practically no overhead in performance. This also
266 has a side effect of enabling or disabling specific functions
267 to be traced. Echoing names of functions into this file
268 will limit the trace to only those functions.
269 This influences the tracers "function" and "function_graph"
270 and thus also function profiling (see "function_profile_enabled").
271
272 The functions listed in "available_filter_functions" are what
273 can be written into this file.
274
275 This interface also allows for commands to be used. See the
276 "Filter commands" section for more details.
277
278 As a speed up, since processing strings can be quite expensive
279 and requires a check of all functions registered to tracing, instead
280 an index can be written into this file. A number (starting with "1")
281 written will instead select the same corresponding at the line position
282 of the "available_filter_functions" file.
283
284 set_ftrace_notrace:
285
286 This has an effect opposite to that of
287 set_ftrace_filter. Any function that is added here will not
288 be traced. If a function exists in both set_ftrace_filter
289 and set_ftrace_notrace, the function will _not_ be traced.
290
291 set_ftrace_pid:
292
293 Have the function tracer only trace the threads whose PID are
294 listed in this file.
295
296 If the "function-fork" option is set, then when a task whose
297 PID is listed in this file forks, the child's PID will
298 automatically be added to this file, and the child will be
299 traced by the function tracer as well. This option will also
300 cause PIDs of tasks that exit to be removed from the file.
301
302 set_ftrace_notrace_pid:
303
304 Have the function tracer ignore threads whose PID are listed in
305 this file.
306
307 If the "function-fork" option is set, then when a task whose
308 PID is listed in this file forks, the child's PID will
309 automatically be added to this file, and the child will not be
310 traced by the function tracer as well. This option will also
311 cause PIDs of tasks that exit to be removed from the file.
312
313 If a PID is in both this file and "set_ftrace_pid", then this
314 file takes precedence, and the thread will not be traced.
315
316 set_event_pid:
317
318 Have the events only trace a task with a PID listed in this file.
319 Note, sched_switch and sched_wake_up will also trace events
320 listed in this file.
321
322 To have the PIDs of children of tasks with their PID in this file
323 added on fork, enable the "event-fork" option. That option will also
324 cause the PIDs of tasks to be removed from this file when the task
325 exits.
326
327 set_event_notrace_pid:
328
329 Have the events not trace a task with a PID listed in this file.
330 Note, sched_switch and sched_wakeup will trace threads not listed
331 in this file, even if a thread's PID is in the file if the
332 sched_switch or sched_wakeup events also trace a thread that should
333 be traced.
334
335 To have the PIDs of children of tasks with their PID in this file
336 added on fork, enable the "event-fork" option. That option will also
337 cause the PIDs of tasks to be removed from this file when the task
338 exits.
339
340 set_graph_function:
341
342 Functions listed in this file will cause the function graph
343 tracer to only trace these functions and the functions that
344 they call. (See the section "dynamic ftrace" for more details).
345 Note, set_ftrace_filter and set_ftrace_notrace still affects
346 what functions are being traced.
347
348 set_graph_notrace:
349
350 Similar to set_graph_function, but will disable function graph
351 tracing when the function is hit until it exits the function.
352 This makes it possible to ignore tracing functions that are called
353 by a specific function.
354
355 available_filter_functions:
356
357 This lists the functions that ftrace has processed and can trace.
358 These are the function names that you can pass to
359 "set_ftrace_filter", "set_ftrace_notrace",
360 "set_graph_function", or "set_graph_notrace".
361 (See the section "dynamic ftrace" below for more details.)
362
363 available_filter_functions_addrs:
364
365 Similar to available_filter_functions, but with address displayed
366 for each function. The displayed address is the patch-site address
367 and can differ from /proc/kallsyms address.
368
369 dyn_ftrace_total_info:
370
371 This file is for debugging purposes. The number of functions that
372 have been converted to nops and are available to be traced.
373
374 enabled_functions:
375
376 This file is more for debugging ftrace, but can also be useful
377 in seeing if any function has a callback attached to it.
378 Not only does the trace infrastructure use ftrace function
379 trace utility, but other subsystems might too. This file
380 displays all functions that have a callback attached to them
381 as well as the number of callbacks that have been attached.
382 Note, a callback may also call multiple functions which will
383 not be listed in this count.
384
385 If the callback registered to be traced by a function with
386 the "save regs" attribute (thus even more overhead), an 'R'
387 will be displayed on the same line as the function that
388 is returning registers.
389
390 If the callback registered to be traced by a function with
391 the "ip modify" attribute (thus the regs->ip can be changed),
392 an 'I' will be displayed on the same line as the function that
393 can be overridden.
394
395 If a non-ftrace trampoline is attached (BPF) a 'D' will be displayed.
396 Note, normal ftrace trampolines can also be attached, but only one
397 "direct" trampoline can be attached to a given function at a time.
398
399 Some architectures can not call direct trampolines, but instead have
400 the ftrace ops function located above the function entry point. In
401 such cases an 'O' will be displayed.
402
403 If a function had either the "ip modify" or a "direct" call attached to
404 it in the past, a 'M' will be shown. This flag is never cleared. It is
405 used to know if a function was ever modified by the ftrace infrastructure,
406 and can be used for debugging.
407
408 If the architecture supports it, it will also show what callback
409 is being directly called by the function. If the count is greater
410 than 1 it most likely will be ftrace_ops_list_func().
411
412 If the callback of a function jumps to a trampoline that is
413 specific to the callback and which is not the standard trampoline,
414 its address will be printed as well as the function that the
415 trampoline calls.
416
417 touched_functions:
418
419 This file contains all the functions that ever had a function callback
420 to it via the ftrace infrastructure. It has the same format as
421 enabled_functions but shows all functions that have ever been
422 traced.
423
424 To see any function that has every been modified by "ip modify" or a
425 direct trampoline, one can perform the following command:
426
427 grep ' M ' /sys/kernel/tracing/touched_functions
428
429 function_profile_enabled:
430
431 When set it will enable all functions with either the function
432 tracer, or if configured, the function graph tracer. It will
433 keep a histogram of the number of functions that were called
434 and if the function graph tracer was configured, it will also keep
435 track of the time spent in those functions. The histogram
436 content can be displayed in the files:
437
438 trace_stat/function<cpu> ( function0, function1, etc).
439
440 trace_stat:
441
442 A directory that holds different tracing stats.
443
444 kprobe_events:
445
446 Enable dynamic trace points. See kprobetrace.rst.
447
448 kprobe_profile:
449
450 Dynamic trace points stats. See kprobetrace.rst.
451
452 max_graph_depth:
453
454 Used with the function graph tracer. This is the max depth
455 it will trace into a function. Setting this to a value of
456 one will show only the first kernel function that is called
457 from user space.
458
459 printk_formats:
460
461 This is for tools that read the raw format files. If an event in
462 the ring buffer references a string, only a pointer to the string
463 is recorded into the buffer and not the string itself. This prevents
464 tools from knowing what that string was. This file displays the string
465 and address for the string allowing tools to map the pointers to what
466 the strings were.
467
468 saved_cmdlines:
469
470 Only the pid of the task is recorded in a trace event unless
471 the event specifically saves the task comm as well. Ftrace
472 makes a cache of pid mappings to comms to try to display
473 comms for events. If a pid for a comm is not listed, then
474 "<...>" is displayed in the output.
475
476 If the option "record-cmd" is set to "0", then comms of tasks
477 will not be saved during recording. By default, it is enabled.
478
479 saved_cmdlines_size:
480
481 By default, 128 comms are saved (see "saved_cmdlines" above). To
482 increase or decrease the amount of comms that are cached, echo
483 the number of comms to cache into this file.
484
485 saved_tgids:
486
487 If the option "record-tgid" is set, on each scheduling context switch
488 the Task Group ID of a task is saved in a table mapping the PID of
489 the thread to its TGID. By default, the "record-tgid" option is
490 disabled.
491
492 snapshot:
493
494 This displays the "snapshot" buffer and also lets the user
495 take a snapshot of the current running trace.
496 See the "Snapshot" section below for more details.
497
498 stack_max_size:
499
500 When the stack tracer is activated, this will display the
501 maximum stack size it has encountered.
502 See the "Stack Trace" section below.
503
504 stack_trace:
505
506 This displays the stack back trace of the largest stack
507 that was encountered when the stack tracer is activated.
508 See the "Stack Trace" section below.
509
510 stack_trace_filter:
511
512 This is similar to "set_ftrace_filter" but it limits what
513 functions the stack tracer will check.
514
515 trace_clock:
516
517 Whenever an event is recorded into the ring buffer, a
518 "timestamp" is added. This stamp comes from a specified
519 clock. By default, ftrace uses the "local" clock. This
520 clock is very fast and strictly per CPU, but on some
521 systems it may not be monotonic with respect to other
522 CPUs. In other words, the local clocks may not be in sync
523 with local clocks on other CPUs.
524
525 Usual clocks for tracing::
526
527 # cat trace_clock
528 [local] global counter x86-tsc
529
530 The clock with the square brackets around it is the one in effect.
531
532 local:
533 Default clock, but may not be in sync across CPUs
534
535 global:
536 This clock is in sync with all CPUs but may
537 be a bit slower than the local clock.
538
539 counter:
540 This is not a clock at all, but literally an atomic
541 counter. It counts up one by one, but is in sync
542 with all CPUs. This is useful when you need to
543 know exactly the order events occurred with respect to
544 each other on different CPUs.
545
546 uptime:
547 This uses the jiffies counter and the time stamp
548 is relative to the time since boot up.
549
550 perf:
551 This makes ftrace use the same clock that perf uses.
552 Eventually perf will be able to read ftrace buffers
553 and this will help out in interleaving the data.
554
555 x86-tsc:
556 Architectures may define their own clocks. For
557 example, x86 uses its own TSC cycle clock here.
558
559 ppc-tb:
560 This uses the powerpc timebase register value.
561 This is in sync across CPUs and can also be used
562 to correlate events across hypervisor/guest if
563 tb_offset is known.
564
565 mono:
566 This uses the fast monotonic clock (CLOCK_MONOTONIC)
567 which is monotonic and is subject to NTP rate adjustments.
568
569 mono_raw:
570 This is the raw monotonic clock (CLOCK_MONOTONIC_RAW)
571 which is monotonic but is not subject to any rate adjustments
572 and ticks at the same rate as the hardware clocksource.
573
574 boot:
575 This is the boot clock (CLOCK_BOOTTIME) and is based on the
576 fast monotonic clock, but also accounts for time spent in
577 suspend. Since the clock access is designed for use in
578 tracing in the suspend path, some side effects are possible
579 if clock is accessed after the suspend time is accounted before
580 the fast mono clock is updated. In this case, the clock update
581 appears to happen slightly sooner than it normally would have.
582 Also on 32-bit systems, it's possible that the 64-bit boot offset
583 sees a partial update. These effects are rare and post
584 processing should be able to handle them. See comments in the
585 ktime_get_boot_fast_ns() function for more information.
586
587 tai:
588 This is the tai clock (CLOCK_TAI) and is derived from the wall-
589 clock time. However, this clock does not experience
590 discontinuities and backwards jumps caused by NTP inserting leap
591 seconds. Since the clock access is designed for use in tracing,
592 side effects are possible. The clock access may yield wrong
593 readouts in case the internal TAI offset is updated e.g., caused
594 by setting the system time or using adjtimex() with an offset.
595 These effects are rare and post processing should be able to
596 handle them. See comments in the ktime_get_tai_fast_ns()
597 function for more information.
598
599 To set a clock, simply echo the clock name into this file::
600
601 # echo global > trace_clock
602
603 Setting a clock clears the ring buffer content as well as the
604 "snapshot" buffer.
605
606 trace_marker:
607
608 This is a very useful file for synchronizing user space
609 with events happening in the kernel. Writing strings into
610 this file will be written into the ftrace buffer.
611
612 It is useful in applications to open this file at the start
613 of the application and just reference the file descriptor
614 for the file::
615
616 void trace_write(const char *fmt, ...)
617 {
618 va_list ap;
619 char buf[256];
620 int n;
621
622 if (trace_fd < 0)
623 return;
624
625 va_start(ap, fmt);
626 n = vsnprintf(buf, 256, fmt, ap);
627 va_end(ap);
628
629 write(trace_fd, buf, n);
630 }
631
632 start::
633
634 trace_fd = open("trace_marker", O_WRONLY);
635
636 Note: Writing into the trace_marker file can also initiate triggers
637 that are written into /sys/kernel/tracing/events/ftrace/print/trigger
638 See "Event triggers" in Documentation/trace/events.rst and an
639 example in Documentation/trace/histogram.rst (Section 3.)
640
641 trace_marker_raw:
642
643 This is similar to trace_marker above, but is meant for binary data
644 to be written to it, where a tool can be used to parse the data
645 from trace_pipe_raw.
646
647 uprobe_events:
648
649 Add dynamic tracepoints in programs.
650 See uprobetracer.rst
651
652 uprobe_profile:
653
654 Uprobe statistics. See uprobetrace.txt
655
656 instances:
657
658 This is a way to make multiple trace buffers where different
659 events can be recorded in different buffers.
660 See "Instances" section below.
661
662 events:
663
664 This is the trace event directory. It holds event tracepoints
665 (also known as static tracepoints) that have been compiled
666 into the kernel. It shows what event tracepoints exist
667 and how they are grouped by system. There are "enable"
668 files at various levels that can enable the tracepoints
669 when a "1" is written to them.
670
671 See events.rst for more information.
672
673 set_event:
674
675 By echoing in the event into this file, will enable that event.
676
677 See events.rst for more information.
678
679 available_events:
680
681 A list of events that can be enabled in tracing.
682
683 See events.rst for more information.
684
685 timestamp_mode:
686
687 Certain tracers may change the timestamp mode used when
688 logging trace events into the event buffer. Events with
689 different modes can coexist within a buffer but the mode in
690 effect when an event is logged determines which timestamp mode
691 is used for that event. The default timestamp mode is
692 'delta'.
693
694 Usual timestamp modes for tracing:
695
696 # cat timestamp_mode
697 [delta] absolute
698
699 The timestamp mode with the square brackets around it is the
700 one in effect.
701
702 delta: Default timestamp mode - timestamp is a delta against
703 a per-buffer timestamp.
704
705 absolute: The timestamp is a full timestamp, not a delta
706 against some other value. As such it takes up more
707 space and is less efficient.
708
709 hwlat_detector:
710
711 Directory for the Hardware Latency Detector.
712 See "Hardware Latency Detector" section below.
713
714 per_cpu:
715
716 This is a directory that contains the trace per_cpu information.
717
718 per_cpu/cpu0/buffer_size_kb:
719
720 The ftrace buffer is defined per_cpu. That is, there's a separate
721 buffer for each CPU to allow writes to be done atomically,
722 and free from cache bouncing. These buffers may have different
723 size buffers. This file is similar to the buffer_size_kb
724 file, but it only displays or sets the buffer size for the
725 specific CPU. (here cpu0).
726
727 per_cpu/cpu0/trace:
728
729 This is similar to the "trace" file, but it will only display
730 the data specific for the CPU. If written to, it only clears
731 the specific CPU buffer.
732
733 per_cpu/cpu0/trace_pipe
734
735 This is similar to the "trace_pipe" file, and is a consuming
736 read, but it will only display (and consume) the data specific
737 for the CPU.
738
739 per_cpu/cpu0/trace_pipe_raw
740
741 For tools that can parse the ftrace ring buffer binary format,
742 the trace_pipe_raw file can be used to extract the data
743 from the ring buffer directly. With the use of the splice()
744 system call, the buffer data can be quickly transferred to
745 a file or to the network where a server is collecting the
746 data.
747
748 Like trace_pipe, this is a consuming reader, where multiple
749 reads will always produce different data.
750
751 per_cpu/cpu0/snapshot:
752
753 This is similar to the main "snapshot" file, but will only
754 snapshot the current CPU (if supported). It only displays
755 the content of the snapshot for a given CPU, and if
756 written to, only clears this CPU buffer.
757
758 per_cpu/cpu0/snapshot_raw:
759
760 Similar to the trace_pipe_raw, but will read the binary format
761 from the snapshot buffer for the given CPU.
762
763 per_cpu/cpu0/stats:
764
765 This displays certain stats about the ring buffer:
766
767 entries:
768 The number of events that are still in the buffer.
769
770 overrun:
771 The number of lost events due to overwriting when
772 the buffer was full.
773
774 commit overrun:
775 Should always be zero.
776 This gets set if so many events happened within a nested
777 event (ring buffer is re-entrant), that it fills the
778 buffer and starts dropping events.
779
780 bytes:
781 Bytes actually read (not overwritten).
782
783 oldest event ts:
784 The oldest timestamp in the buffer
785
786 now ts:
787 The current timestamp
788
789 dropped events:
790 Events lost due to overwrite option being off.
791
792 read events:
793 The number of events read.
794
795 The Tracers
796 -----------
797
798 Here is the list of current tracers that may be configured.
799
800 "function"
801
802 Function call tracer to trace all kernel functions.
803
804 "function_graph"
805
806 Similar to the function tracer except that the
807 function tracer probes the functions on their entry
808 whereas the function graph tracer traces on both entry
809 and exit of the functions. It then provides the ability
810 to draw a graph of function calls similar to C code
811 source.
812
813 Note that the function graph calculates the timings of when the
814 function starts and returns internally and for each instance. If
815 there are two instances that run function graph tracer and traces
816 the same functions, the length of the timings may be slightly off as
817 each read the timestamp separately and not at the same time.
818
819 "blk"
820
821 The block tracer. The tracer used by the blktrace user
822 application.
823
824 "hwlat"
825
826 The Hardware Latency tracer is used to detect if the hardware
827 produces any latency. See "Hardware Latency Detector" section
828 below.
829
830 "irqsoff"
831
832 Traces the areas that disable interrupts and saves
833 the trace with the longest max latency.
834 See tracing_max_latency. When a new max is recorded,
835 it replaces the old trace. It is best to view this
836 trace with the latency-format option enabled, which
837 happens automatically when the tracer is selected.
838
839 "preemptoff"
840
841 Similar to irqsoff but traces and records the amount of
842 time for which preemption is disabled.
843
844 "preemptirqsoff"
845
846 Similar to irqsoff and preemptoff, but traces and
847 records the largest time for which irqs and/or preemption
848 is disabled.
849
850 "wakeup"
851
852 Traces and records the max latency that it takes for
853 the highest priority task to get scheduled after
854 it has been woken up.
855 Traces all tasks as an average developer would expect.
856
857 "wakeup_rt"
858
859 Traces and records the max latency that it takes for just
860 RT tasks (as the current "wakeup" does). This is useful
861 for those interested in wake up timings of RT tasks.
862
863 "wakeup_dl"
864
865 Traces and records the max latency that it takes for
866 a SCHED_DEADLINE task to be woken (as the "wakeup" and
867 "wakeup_rt" does).
868
869 "mmiotrace"
870
871 A special tracer that is used to trace binary modules.
872 It will trace all the calls that a module makes to the
873 hardware. Everything it writes and reads from the I/O
874 as well.
875
876 "branch"
877
878 This tracer can be configured when tracing likely/unlikely
879 calls within the kernel. It will trace when a likely and
880 unlikely branch is hit and if it was correct in its prediction
881 of being correct.
882
883 "nop"
884
885 This is the "trace nothing" tracer. To remove all
886 tracers from tracing simply echo "nop" into
887 current_tracer.
888
889 Error conditions
890 ----------------
891
892 For most ftrace commands, failure modes are obvious and communicated
893 using standard return codes.
894
895 For other more involved commands, extended error information may be
896 available via the tracing/error_log file. For the commands that
897 support it, reading the tracing/error_log file after an error will
898 display more detailed information about what went wrong, if
899 information is available. The tracing/error_log file is a circular
900 error log displaying a small number (currently, 8) of ftrace errors
901 for the last (8) failed commands.
902
903 The extended error information and usage takes the form shown in
904 this example::
905
906 # echo xxx > /sys/kernel/tracing/events/sched/sched_wakeup/trigger
907 echo: write error: Invalid argument
908
909 # cat /sys/kernel/tracing/error_log
910 [ 5348.887237] location: error: Couldn't yyy: zzz
911 Command: xxx
912 ^
913 [ 7517.023364] location: error: Bad rrr: sss
914 Command: ppp qqq
915 ^
916
917 To clear the error log, echo the empty string into it::
918
919 # echo > /sys/kernel/tracing/error_log
920
921 Examples of using the tracer
922 ----------------------------
923
924 Here are typical examples of using the tracers when controlling
925 them only with the tracefs interface (without using any
926 user-land utilities).
927
928 Output format:
929 --------------
930
931 Here is an example of the output format of the file "trace"::
932
933 # tracer: function
934 #
935 # entries-in-buffer/entries-written: 140080/250280 #P:4
936 #
937 # _-----=> irqs-off
938 # / _----=> need-resched
939 # | / _---=> hardirq/softirq
940 # || / _--=> preempt-depth
941 # ||| / delay
942 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
943 # | | | |||| | |
944 bash-1977 [000] .... 17284.993652: sys_close <-system_call_fastpath
945 bash-1977 [000] .... 17284.993653: __close_fd <-sys_close
946 bash-1977 [000] .... 17284.993653: _raw_spin_lock <-__close_fd
947 sshd-1974 [003] .... 17284.993653: __srcu_read_unlock <-fsnotify
948 bash-1977 [000] .... 17284.993654: add_preempt_count <-_raw_spin_lock
949 bash-1977 [000] ...1 17284.993655: _raw_spin_unlock <-__close_fd
950 bash-1977 [000] ...1 17284.993656: sub_preempt_count <-_raw_spin_unlock
951 bash-1977 [000] .... 17284.993657: filp_close <-__close_fd
952 bash-1977 [000] .... 17284.993657: dnotify_flush <-filp_close
953 sshd-1974 [003] .... 17284.993658: sys_select <-system_call_fastpath
954 ....
955
956 A header is printed with the tracer name that is represented by
957 the trace. In this case the tracer is "function". Then it shows the
958 number of events in the buffer as well as the total number of entries
959 that were written. The difference is the number of entries that were
960 lost due to the buffer filling up (250280 - 140080 = 110200 events
961 lost).
962
963 The header explains the content of the events. Task name "bash", the task
964 PID "1977", the CPU that it was running on "000", the latency format
965 (explained below), the timestamp in <secs>.<usecs> format, the
966 function name that was traced "sys_close" and the parent function that
967 called this function "system_call_fastpath". The timestamp is the time
968 at which the function was entered.
969
970 Latency trace format
971 --------------------
972
973 When the latency-format option is enabled or when one of the latency
974 tracers is set, the trace file gives somewhat more information to see
975 why a latency happened. Here is a typical trace::
976
977 # tracer: irqsoff
978 #
979 # irqsoff latency trace v1.1.5 on 3.8.0-test+
980 # --------------------------------------------------------------------
981 # latency: 259 us, #4/4, CPU#2 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
982 # -----------------
983 # | task: ps-6143 (uid:0 nice:0 policy:0 rt_prio:0)
984 # -----------------
985 # => started at: __lock_task_sighand
986 # => ended at: _raw_spin_unlock_irqrestore
987 #
988 #
989 # _------=> CPU#
990 # / _-----=> irqs-off
991 # | / _----=> need-resched
992 # || / _---=> hardirq/softirq
993 # ||| / _--=> preempt-depth
994 # |||| / delay
995 # cmd pid ||||| time | caller
996 # \ / ||||| \ | /
997 ps-6143 2d... 0us!: trace_hardirqs_off <-__lock_task_sighand
998 ps-6143 2d..1 259us+: trace_hardirqs_on <-_raw_spin_unlock_irqrestore
999 ps-6143 2d..1 263us+: time_hardirqs_on <-_raw_spin_unlock_irqrestore
1000 ps-6143 2d..1 306us : <stack trace>
1001 => trace_hardirqs_on_caller
1002 => trace_hardirqs_on
1003 => _raw_spin_unlock_irqrestore
1004 => do_task_stat
1005 => proc_tgid_stat
1006 => proc_single_show
1007 => seq_read
1008 => vfs_read
1009 => sys_read
1010 => system_call_fastpath
1013 This shows that the current tracer is "irqsoff" tracing the time
1014 for which interrupts were disabled. It gives the trace version (which
1015 never changes) and the version of the kernel upon which this was executed on
1016 (3.8). Then it displays the max latency in microseconds (259 us). The number
1017 of trace entries displayed and the total number (both are four: #4/4).
1018 VP, KP, SP, and HP are always zero and are reserved for later use.
1019 #P is the number of online CPUs (#P:4).
1021 The task is the process that was running when the latency
1022 occurred. (ps pid: 6143).
1024 The start and stop (the functions in which the interrupts were
1025 disabled and enabled respectively) that caused the latencies:
1027 - __lock_task_sighand is where the interrupts were disabled.
1028 - _raw_spin_unlock_irqrestore is where they were enabled again.
1030 The next lines after the header are the trace itself. The header
1031 explains which is which.
1033 cmd: The name of the process in the trace.
1035 pid: The PID of that process.
1037 CPU#: The CPU which the process was running on.
1039 irqs-off: 'd' interrupts are disabled. '.' otherwise.
1041 need-resched:
1042 - 'B' all, TIF_NEED_RESCHED, PREEMPT_NEED_RESCHED and TIF_RESCHED_LAZY is set,
1043 - 'N' both TIF_NEED_RESCHED and PREEMPT_NEED_RESCHED is set,
1044 - 'n' only TIF_NEED_RESCHED is set,
1045 - 'p' only PREEMPT_NEED_RESCHED is set,
1046 - 'L' both PREEMPT_NEED_RESCHED and TIF_RESCHED_LAZY is set,
1047 - 'b' both TIF_NEED_RESCHED and TIF_RESCHED_LAZY is set,
1048 - 'l' only TIF_RESCHED_LAZY is set
1049 - '.' otherwise.
1051 hardirq/softirq:
1052 - 'Z' - NMI occurred inside a hardirq
1053 - 'z' - NMI is running
1054 - 'H' - hard irq occurred inside a softirq.
1055 - 'h' - hard irq is running
1056 - 's' - soft irq is running
1057 - '.' - normal context.
1059 preempt-depth: The level of preempt_disabled
1061 The above is mostly meaningful for kernel developers.
1063 time:
1064 When the latency-format option is enabled, the trace file
1065 output includes a timestamp relative to the start of the
1066 trace. This differs from the output when latency-format
1067 is disabled, which includes an absolute timestamp.
1069 delay:
1070 This is just to help catch your eye a bit better. And
1071 needs to be fixed to be only relative to the same CPU.
1072 The marks are determined by the difference between this
1073 current trace and the next trace.
1075 - '$' - greater than 1 second
1076 - '@' - greater than 100 millisecond
1077 - '*' - greater than 10 millisecond
1078 - '#' - greater than 1000 microsecond
1079 - '!' - greater than 100 microsecond
1080 - '+' - greater than 10 microsecond
1081 - ' ' - less than or equal to 10 microsecond.
1083 The rest is the same as the 'trace' file.
1085 Note, the latency tracers will usually end with a back trace
1086 to easily find where the latency occurred.
1088 trace_options
1089 -------------
1091 The trace_options file (or the options directory) is used to control
1092 what gets printed in the trace output, or manipulate the tracers.
1093 To see what is available, simply cat the file::
1095 cat trace_options
1096 print-parent
1097 nosym-offset
1098 nosym-addr
1099 noverbose
1100 noraw
1101 nohex
1102 nobin
1103 noblock
1104 nofields
1105 trace_printk
1106 annotate
1107 nouserstacktrace
1108 nosym-userobj
1109 noprintk-msg-only
1110 context-info
1111 nolatency-format
1112 record-cmd
1113 norecord-tgid
1114 overwrite
1115 nodisable_on_free
1116 irq-info
1117 markers
1118 noevent-fork
1119 function-trace
1120 nofunction-fork
1121 nodisplay-graph
1122 nostacktrace
1123 nobranch
1125 To disable one of the options, echo in the option prepended with
1126 "no"::
1128 echo noprint-parent > trace_options
1130 To enable an option, leave off the "no"::
1132 echo sym-offset > trace_options
1134 Here are the available options:
1136 print-parent
1137 On function traces, display the calling (parent)
1138 function as well as the function being traced.
1139 ::
1141 print-parent:
1142 bash-4000 [01] 1477.606694: simple_strtoul <-kstrtoul
1144 noprint-parent:
1145 bash-4000 [01] 1477.606694: simple_strtoul
1148 sym-offset
1149 Display not only the function name, but also the
1150 offset in the function. For example, instead of
1151 seeing just "ktime_get", you will see
1152 "ktime_get+0xb/0x20".
1153 ::
1155 sym-offset:
1156 bash-4000 [01] 1477.606694: simple_strtoul+0x6/0xa0
1158 sym-addr
1159 This will also display the function address as well
1160 as the function name.
1161 ::
1163 sym-addr:
1164 bash-4000 [01] 1477.606694: simple_strtoul <c0339346>
1166 verbose
1167 This deals with the trace file when the
1168 latency-format option is enabled.
1169 ::
1171 bash 4000 1 0 00000000 00010a95 [58127d26] 1720.415ms \
1172 (+0.000ms): simple_strtoul (kstrtoul)
1174 raw
1175 This will display raw numbers. This option is best for
1176 use with user applications that can translate the raw
1177 numbers better than having it done in the kernel.
1179 hex
1180 Similar to raw, but the numbers will be in a hexadecimal format.
1182 bin
1183 This will print out the formats in raw binary.
1185 block
1186 When set, reading trace_pipe will not block when polled.
1188 fields
1189 Print the fields as described by their types. This is a better
1190 option than using hex, bin or raw, as it gives a better parsing
1191 of the content of the event.
1193 trace_printk
1194 Can disable trace_printk() from writing into the buffer.
1196 trace_printk_dest
1197 Set to have trace_printk() and similar internal tracing functions
1198 write into this instance. Note, only one trace instance can have
1199 this set. By setting this flag, it clears the trace_printk_dest flag
1200 of the instance that had it set previously. By default, the top
1201 level trace has this set, and will get it set again if another
1202 instance has it set then clears it.
1204 This flag cannot be cleared by the top level instance, as it is the
1205 default instance. The only way the top level instance has this flag
1206 cleared, is by it being set in another instance.
1208 copy_trace_marker
1209 If there are applications that hard code writing into the top level
1210 trace_marker file (/sys/kernel/tracing/trace_marker or trace_marker_raw),
1211 and the tooling would like it to go into an instance, this option can
1212 be used. Create an instance and set this option, and then all writes
1213 into the top level trace_marker file will also be redirected into this
1214 instance.
1216 Note, by default this option is set for the top level instance. If it
1217 is disabled, then writes to the trace_marker or trace_marker_raw files
1218 will not be written into the top level file. If no instance has this
1219 option set, then a write will error with the errno of ENODEV.
1221 annotate
1222 It is sometimes confusing when the CPU buffers are full
1223 and one CPU buffer had a lot of events recently, thus
1224 a shorter time frame, were another CPU may have only had
1225 a few events, which lets it have older events. When
1226 the trace is reported, it shows the oldest events first,
1227 and it may look like only one CPU ran (the one with the
1228 oldest events). When the annotate option is set, it will
1229 display when a new CPU buffer started::
1231 <idle>-0 [001] dNs4 21169.031481: wake_up_idle_cpu <-add_timer_on
1232 <idle>-0 [001] dNs4 21169.031482: _raw_spin_unlock_irqrestore <-add_timer_on
1233 <idle>-0 [001] .Ns4 21169.031484: sub_preempt_count <-_raw_spin_unlock_irqrestore
1234 ##### CPU 2 buffer started ####
1235 <idle>-0 [002] .N.1 21169.031484: rcu_idle_exit <-cpu_idle
1236 <idle>-0 [001] .Ns3 21169.031484: _raw_spin_unlock <-clocksource_watchdog
1237 <idle>-0 [001] .Ns3 21169.031485: sub_preempt_count <-_raw_spin_unlock
1239 userstacktrace
1240 This option changes the trace. It records a
1241 stacktrace of the current user space thread after
1242 each trace event.
1244 sym-userobj
1245 when user stacktrace are enabled, look up which
1246 object the address belongs to, and print a
1247 relative address. This is especially useful when
1248 ASLR is on, otherwise you don't get a chance to
1249 resolve the address to object/file/line after
1250 the app is no longer running
1252 The lookup is performed when you read
1253 trace,trace_pipe. Example::
1255 a.out-1623 [000] 40874.465068: /root/a.out[+0x480] <-/root/a.out[+0
1256 x494] <- /root/a.out[+0x4a8] <- /lib/libc-2.7.so[+0x1e1a6]
1259 printk-msg-only
1260 When set, trace_printk()s will only show the format
1261 and not their parameters (if trace_bprintk() or
1262 trace_bputs() was used to save the trace_printk()).
1264 context-info
1265 Show only the event data. Hides the comm, PID,
1266 timestamp, CPU, and other useful data.
1268 latency-format
1269 This option changes the trace output. When it is enabled,
1270 the trace displays additional information about the
1271 latency, as described in "Latency trace format".
1273 pause-on-trace
1274 When set, opening the trace file for read, will pause
1275 writing to the ring buffer (as if tracing_on was set to zero).
1276 This simulates the original behavior of the trace file.
1277 When the file is closed, tracing will be enabled again.
1279 hash-ptr
1280 When set, "%p" in the event printk format displays the
1281 hashed pointer value instead of real address.
1282 This will be useful if you want to find out which hashed
1283 value is corresponding to the real value in trace log.
1285 record-cmd
1286 When any event or tracer is enabled, a hook is enabled
1287 in the sched_switch trace point to fill comm cache
1288 with mapped pids and comms. But this may cause some
1289 overhead, and if you only care about pids, and not the
1290 name of the task, disabling this option can lower the
1291 impact of tracing. See "saved_cmdlines".
1293 record-tgid
1294 When any event or tracer is enabled, a hook is enabled
1295 in the sched_switch trace point to fill the cache of
1296 mapped Thread Group IDs (TGID) mapping to pids. See
1297 "saved_tgids".
1299 overwrite
1300 This controls what happens when the trace buffer is
1301 full. If "1" (default), the oldest events are
1302 discarded and overwritten. If "0", then the newest
1303 events are discarded.
1304 (see per_cpu/cpu0/stats for overrun and dropped)
1306 disable_on_free
1307 When the free_buffer is closed, tracing will
1308 stop (tracing_on set to 0).
1310 irq-info
1311 Shows the interrupt, preempt count, need resched data.
1312 When disabled, the trace looks like::
1314 # tracer: function
1315 #
1316 # entries-in-buffer/entries-written: 144405/9452052 #P:4
1317 #
1318 # TASK-PID CPU# TIMESTAMP FUNCTION
1319 # | | | | |
1320 <idle>-0 [002] 23636.756054: ttwu_do_activate.constprop.89 <-try_to_wake_up
1321 <idle>-0 [002] 23636.756054: activate_task <-ttwu_do_activate.constprop.89
1322 <idle>-0 [002] 23636.756055: enqueue_task <-activate_task
1325 markers
1326 When set, the trace_marker is writable (only by root).
1327 When disabled, the trace_marker will error with EINVAL
1328 on write.
1330 event-fork
1331 When set, tasks with PIDs listed in set_event_pid will have
1332 the PIDs of their children added to set_event_pid when those
1333 tasks fork. Also, when tasks with PIDs in set_event_pid exit,
1334 their PIDs will be removed from the file.
1336 This affects PIDs listed in set_event_notrace_pid as well.
1338 function-trace
1339 The latency tracers will enable function tracing
1340 if this option is enabled (default it is). When
1341 it is disabled, the latency tracers do not trace
1342 functions. This keeps the overhead of the tracer down
1343 when performing latency tests.
1345 function-fork
1346 When set, tasks with PIDs listed in set_ftrace_pid will
1347 have the PIDs of their children added to set_ftrace_pid
1348 when those tasks fork. Also, when tasks with PIDs in
1349 set_ftrace_pid exit, their PIDs will be removed from the
1350 file.
1352 This affects PIDs in set_ftrace_notrace_pid as well.
1354 display-graph
1355 When set, the latency tracers (irqsoff, wakeup, etc) will
1356 use function graph tracing instead of function tracing.
1358 stacktrace
1359 When set, a stack trace is recorded after any trace event
1360 is recorded.
1362 branch
1363 Enable branch tracing with the tracer. This enables branch
1364 tracer along with the currently set tracer. Enabling this
1365 with the "nop" tracer is the same as just enabling the
1366 "branch" tracer.
1368 .. tip:: Some tracers have their own options. They only appear in this
1369 file when the tracer is active. They always appear in the
1370 options directory.
1373 Here are the per tracer options:
1375 Options for function tracer:
1377 func_stack_trace
1378 When set, a stack trace is recorded after every
1379 function that is recorded. NOTE! Limit the functions
1380 that are recorded before enabling this, with
1381 "set_ftrace_filter" otherwise the system performance
1382 will be critically degraded. Remember to disable
1383 this option before clearing the function filter.
1385 Options for function_graph tracer:
1387 Since the function_graph tracer has a slightly different output
1388 it has its own options to control what is displayed.
1390 funcgraph-overrun
1391 When set, the "overrun" of the graph stack is
1392 displayed after each function traced. The
1393 overrun, is when the stack depth of the calls
1394 is greater than what is reserved for each task.
1395 Each task has a fixed array of functions to
1396 trace in the call graph. If the depth of the
1397 calls exceeds that, the function is not traced.
1398 The overrun is the number of functions missed
1399 due to exceeding this array.
1401 funcgraph-cpu
1402 When set, the CPU number of the CPU where the trace
1403 occurred is displayed.
1405 funcgraph-overhead
1406 When set, if the function takes longer than
1407 A certain amount, then a delay marker is
1408 displayed. See "delay" above, under the
1409 header description.
1411 funcgraph-proc
1412 Unlike other tracers, the process' command line
1413 is not displayed by default, but instead only
1414 when a task is traced in and out during a context
1415 switch. Enabling this options has the command
1416 of each process displayed at every line.
1418 funcgraph-duration
1419 At the end of each function (the return)
1420 the duration of the amount of time in the
1421 function is displayed in microseconds.
1423 funcgraph-abstime
1424 When set, the timestamp is displayed at each line.
1426 funcgraph-irqs
1427 When disabled, functions that happen inside an
1428 interrupt will not be traced.
1430 funcgraph-tail
1431 When set, the return event will include the function
1432 that it represents. By default this is off, and
1433 only a closing curly bracket "}" is displayed for
1434 the return of a function.
1436 funcgraph-retval
1437 When set, the return value of each traced function
1438 will be printed after an equal sign "=". By default
1439 this is off.
1441 funcgraph-retval-hex
1442 When set, the return value will always be printed
1443 in hexadecimal format. If the option is not set and
1444 the return value is an error code, it will be printed
1445 in signed decimal format; otherwise it will also be
1446 printed in hexadecimal format. By default, this option
1447 is off.
1449 sleep-time
1450 When running function graph tracer, to include
1451 the time a task schedules out in its function.
1452 When enabled, it will account time the task has been
1453 scheduled out as part of the function call.
1455 graph-time
1456 When running function profiler with function graph tracer,
1457 to include the time to call nested functions. When this is
1458 not set, the time reported for the function will only
1459 include the time the function itself executed for, not the
1460 time for functions that it called.
1462 Options for blk tracer:
1464 blk_classic
1465 Shows a more minimalistic output.
1468 irqsoff
1469 -------
1471 When interrupts are disabled, the CPU can not react to any other
1472 external event (besides NMIs and SMIs). This prevents the timer
1473 interrupt from triggering or the mouse interrupt from letting
1474 the kernel know of a new mouse event. The result is a latency
1475 with the reaction time.
1477 The irqsoff tracer tracks the time for which interrupts are
1478 disabled. When a new maximum latency is hit, the tracer saves
1479 the trace leading up to that latency point so that every time a
1480 new maximum is reached, the old saved trace is discarded and the
1481 new trace is saved.
1483 To reset the maximum, echo 0 into tracing_max_latency. Here is
1484 an example::
1486 # echo 0 > options/function-trace
1487 # echo irqsoff > current_tracer
1488 # echo 1 > tracing_on
1489 # echo 0 > tracing_max_latency
1490 # ls -ltr
1491 [...]
1492 # echo 0 > tracing_on
1493 # cat trace
1494 # tracer: irqsoff
1495 #
1496 # irqsoff latency trace v1.1.5 on 3.8.0-test+
1497 # --------------------------------------------------------------------
1498 # latency: 16 us, #4/4, CPU#0 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
1499 # -----------------
1500 # | task: swapper/0-0 (uid:0 nice:0 policy:0 rt_prio:0)
1501 # -----------------
1502 # => started at: run_timer_softirq
1503 # => ended at: run_timer_softirq
1504 #
1505 #
1506 # _------=> CPU#
1507 # / _-----=> irqs-off
1508 # | / _----=> need-resched
1509 # || / _---=> hardirq/softirq
1510 # ||| / _--=> preempt-depth
1511 # |||| / delay
1512 # cmd pid ||||| time | caller
1513 # \ / ||||| \ | /
1514 <idle>-0 0d.s2 0us+: _raw_spin_lock_irq <-run_timer_softirq
1515 <idle>-0 0dNs3 17us : _raw_spin_unlock_irq <-run_timer_softirq
1516 <idle>-0 0dNs3 17us+: trace_hardirqs_on <-run_timer_softirq
1517 <idle>-0 0dNs3 25us : <stack trace>
1518 => _raw_spin_unlock_irq
1519 => run_timer_softirq
1520 => __do_softirq
1521 => call_softirq
1522 => do_softirq
1523 => irq_exit
1524 => smp_apic_timer_interrupt
1525 => apic_timer_interrupt
1526 => rcu_idle_exit
1527 => cpu_idle
1528 => rest_init
1529 => start_kernel
1530 => x86_64_start_reservations
1531 => x86_64_start_kernel
1533 Here we see that we had a latency of 16 microseconds (which is
1534 very good). The _raw_spin_lock_irq in run_timer_softirq disabled
1535 interrupts. The difference between the 16 and the displayed
1536 timestamp 25us occurred because the clock was incremented
1537 between the time of recording the max latency and the time of
1538 recording the function that had that latency.
1540 Note the above example had function-trace not set. If we set
1541 function-trace, we get a much larger output::
1543 with echo 1 > options/function-trace
1545 # tracer: irqsoff
1546 #
1547 # irqsoff latency trace v1.1.5 on 3.8.0-test+
1548 # --------------------------------------------------------------------
1549 # latency: 71 us, #168/168, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
1550 # -----------------
1551 # | task: bash-2042 (uid:0 nice:0 policy:0 rt_prio:0)
1552 # -----------------
1553 # => started at: ata_scsi_queuecmd
1554 # => ended at: ata_scsi_queuecmd
1555 #
1556 #
1557 # _------=> CPU#
1558 # / _-----=> irqs-off
1559 # | / _----=> need-resched
1560 # || / _---=> hardirq/softirq
1561 # ||| / _--=> preempt-depth
1562 # |||| / delay
1563 # cmd pid ||||| time | caller
1564 # \ / ||||| \ | /
1565 bash-2042 3d... 0us : _raw_spin_lock_irqsave <-ata_scsi_queuecmd
1566 bash-2042 3d... 0us : add_preempt_count <-_raw_spin_lock_irqsave
1567 bash-2042 3d..1 1us : ata_scsi_find_dev <-ata_scsi_queuecmd
1568 bash-2042 3d..1 1us : __ata_scsi_find_dev <-ata_scsi_find_dev
1569 bash-2042 3d..1 2us : ata_find_dev.part.14 <-__ata_scsi_find_dev
1570 bash-2042 3d..1 2us : ata_qc_new_init <-__ata_scsi_queuecmd
1571 bash-2042 3d..1 3us : ata_sg_init <-__ata_scsi_queuecmd
1572 bash-2042 3d..1 4us : ata_scsi_rw_xlat <-__ata_scsi_queuecmd
1573 bash-2042 3d..1 4us : ata_build_rw_tf <-ata_scsi_rw_xlat
1574 [...]
1575 bash-2042 3d..1 67us : delay_tsc <-__delay
1576 bash-2042 3d..1 67us : add_preempt_count <-delay_tsc
1577 bash-2042 3d..2 67us : sub_preempt_count <-delay_tsc
1578 bash-2042 3d..1 67us : add_preempt_count <-delay_tsc
1579 bash-2042 3d..2 68us : sub_preempt_count <-delay_tsc
1580 bash-2042 3d..1 68us+: ata_bmdma_start <-ata_bmdma_qc_issue
1581 bash-2042 3d..1 71us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
1582 bash-2042 3d..1 71us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
1583 bash-2042 3d..1 72us+: trace_hardirqs_on <-ata_scsi_queuecmd
1584 bash-2042 3d..1 120us : <stack trace>
1585 => _raw_spin_unlock_irqrestore
1586 => ata_scsi_queuecmd
1587 => scsi_dispatch_cmd
1588 => scsi_request_fn
1589 => __blk_run_queue_uncond
1590 => __blk_run_queue
1591 => blk_queue_bio
1592 => submit_bio_noacct
1593 => submit_bio
1594 => submit_bh
1595 => __ext3_get_inode_loc
1596 => ext3_iget
1597 => ext3_lookup
1598 => lookup_real
1599 => __lookup_hash
1600 => walk_component
1601 => lookup_last
1602 => path_lookupat
1603 => filename_lookup
1604 => user_path_at_empty
1605 => user_path_at
1606 => vfs_fstatat
1607 => vfs_stat
1608 => sys_newstat
1609 => system_call_fastpath
1612 Here we traced a 71 microsecond latency. But we also see all the
1613 functions that were called during that time. Note that by
1614 enabling function tracing, we incur an added overhead. This
1615 overhead may extend the latency times. But nevertheless, this
1616 trace has provided some very helpful debugging information.
1618 If we prefer function graph output instead of function, we can set
1619 display-graph option::
1621 with echo 1 > options/display-graph
1623 # tracer: irqsoff
1624 #
1625 # irqsoff latency trace v1.1.5 on 4.20.0-rc6+
1626 # --------------------------------------------------------------------
1627 # latency: 3751 us, #274/274, CPU#0 | (M:desktop VP:0, KP:0, SP:0 HP:0 #P:4)
1628 # -----------------
1629 # | task: bash-1507 (uid:0 nice:0 policy:0 rt_prio:0)
1630 # -----------------
1631 # => started at: free_debug_processing
1632 # => ended at: return_to_handler
1633 #
1634 #
1635 # _-----=> irqs-off
1636 # / _----=> need-resched
1637 # | / _---=> hardirq/softirq
1638 # || / _--=> preempt-depth
1639 # ||| /
1640 # REL TIME CPU TASK/PID |||| DURATION FUNCTION CALLS
1641 # | | | | |||| | | | | | |
1642 0 us | 0) bash-1507 | d... | 0.000 us | _raw_spin_lock_irqsave();
1643 0 us | 0) bash-1507 | d..1 | 0.378 us | do_raw_spin_trylock();
1644 1 us | 0) bash-1507 | d..2 | | set_track() {
1645 2 us | 0) bash-1507 | d..2 | | save_stack_trace() {
1646 2 us | 0) bash-1507 | d..2 | | __save_stack_trace() {
1647 3 us | 0) bash-1507 | d..2 | | __unwind_start() {
1648 3 us | 0) bash-1507 | d..2 | | get_stack_info() {
1649 3 us | 0) bash-1507 | d..2 | 0.351 us | in_task_stack();
1650 4 us | 0) bash-1507 | d..2 | 1.107 us | }
1651 [...]
1652 3750 us | 0) bash-1507 | d..1 | 0.516 us | do_raw_spin_unlock();
1653 3750 us | 0) bash-1507 | d..1 | 0.000 us | _raw_spin_unlock_irqrestore();
1654 3764 us | 0) bash-1507 | d..1 | 0.000 us | tracer_hardirqs_on();
1655 bash-1507 0d..1 3792us : <stack trace>
1656 => free_debug_processing
1657 => __slab_free
1658 => kmem_cache_free
1659 => vm_area_free
1660 => remove_vma
1661 => exit_mmap
1662 => mmput
1663 => begin_new_exec
1664 => load_elf_binary
1665 => search_binary_handler
1666 => __do_execve_file.isra.32
1667 => __x64_sys_execve
1668 => do_syscall_64
1669 => entry_SYSCALL_64_after_hwframe
1671 preemptoff
1672 ----------
1674 When preemption is disabled, we may be able to receive
1675 interrupts but the task cannot be preempted and a higher
1676 priority task must wait for preemption to be enabled again
1677 before it can preempt a lower priority task.
1679 The preemptoff tracer traces the places that disable preemption.
1680 Like the irqsoff tracer, it records the maximum latency for
1681 which preemption was disabled. The control of preemptoff tracer
1682 is much like the irqsoff tracer.
1683 ::
1685 # echo 0 > options/function-trace
1686 # echo preemptoff > current_tracer
1687 # echo 1 > tracing_on
1688 # echo 0 > tracing_max_latency
1689 # ls -ltr
1690 [...]
1691 # echo 0 > tracing_on
1692 # cat trace
1693 # tracer: preemptoff
1694 #
1695 # preemptoff latency trace v1.1.5 on 3.8.0-test+
1696 # --------------------------------------------------------------------
1697 # latency: 46 us, #4/4, CPU#1 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
1698 # -----------------
1699 # | task: sshd-1991 (uid:0 nice:0 policy:0 rt_prio:0)
1700 # -----------------
1701 # => started at: do_IRQ
1702 # => ended at: do_IRQ
1703 #
1704 #
1705 # _------=> CPU#
1706 # / _-----=> irqs-off
1707 # | / _----=> need-resched
1708 # || / _---=> hardirq/softirq
1709 # ||| / _--=> preempt-depth
1710 # |||| / delay
1711 # cmd pid ||||| time | caller
1712 # \ / ||||| \ | /
1713 sshd-1991 1d.h. 0us+: irq_enter <-do_IRQ
1714 sshd-1991 1d..1 46us : irq_exit <-do_IRQ
1715 sshd-1991 1d..1 47us+: trace_preempt_on <-do_IRQ
1716 sshd-1991 1d..1 52us : <stack trace>
1717 => sub_preempt_count
1718 => irq_exit
1719 => do_IRQ
1720 => ret_from_intr
1723 This has some more changes. Preemption was disabled when an
1724 interrupt came in (notice the 'h'), and was enabled on exit.
1725 But we also see that interrupts have been disabled when entering
1726 the preempt off section and leaving it (the 'd'). We do not know if
1727 interrupts were enabled in the mean time or shortly after this
1728 was over.
1729 ::
1731 # tracer: preemptoff
1732 #
1733 # preemptoff latency trace v1.1.5 on 3.8.0-test+
1734 # --------------------------------------------------------------------
1735 # latency: 83 us, #241/241, CPU#1 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
1736 # -----------------
1737 # | task: bash-1994 (uid:0 nice:0 policy:0 rt_prio:0)
1738 # -----------------
1739 # => started at: wake_up_new_task
1740 # => ended at: task_rq_unlock
1741 #
1742 #
1743 # _------=> CPU#
1744 # / _-----=> irqs-off
1745 # | / _----=> need-resched
1746 # || / _---=> hardirq/softirq
1747 # ||| / _--=> preempt-depth
1748 # |||| / delay
1749 # cmd pid ||||| time | caller
1750 # \ / ||||| \ | /
1751 bash-1994 1d..1 0us : _raw_spin_lock_irqsave <-wake_up_new_task
1752 bash-1994 1d..1 0us : select_task_rq_fair <-select_task_rq
1753 bash-1994 1d..1 1us : __rcu_read_lock <-select_task_rq_fair
1754 bash-1994 1d..1 1us : source_load <-select_task_rq_fair
1755 bash-1994 1d..1 1us : source_load <-select_task_rq_fair
1756 [...]
1757 bash-1994 1d..1 12us : irq_enter <-smp_apic_timer_interrupt
1758 bash-1994 1d..1 12us : rcu_irq_enter <-irq_enter
1759 bash-1994 1d..1 13us : add_preempt_count <-irq_enter
1760 bash-1994 1d.h1 13us : exit_idle <-smp_apic_timer_interrupt
1761 bash-1994 1d.h1 13us : hrtimer_interrupt <-smp_apic_timer_interrupt
1762 bash-1994 1d.h1 13us : _raw_spin_lock <-hrtimer_interrupt
1763 bash-1994 1d.h1 14us : add_preempt_count <-_raw_spin_lock
1764 bash-1994 1d.h2 14us : ktime_get_update_offsets <-hrtimer_interrupt
1765 [...]
1766 bash-1994 1d.h1 35us : lapic_next_event <-clockevents_program_event
1767 bash-1994 1d.h1 35us : irq_exit <-smp_apic_timer_interrupt
1768 bash-1994 1d.h1 36us : sub_preempt_count <-irq_exit
1769 bash-1994 1d..2 36us : do_softirq <-irq_exit
1770 bash-1994 1d..2 36us : __do_softirq <-call_softirq
1771 bash-1994 1d..2 36us : __local_bh_disable <-__do_softirq
1772 bash-1994 1d.s2 37us : add_preempt_count <-_raw_spin_lock_irq
1773 bash-1994 1d.s3 38us : _raw_spin_unlock <-run_timer_softirq
1774 bash-1994 1d.s3 39us : sub_preempt_count <-_raw_spin_unlock
1775 bash-1994 1d.s2 39us : call_timer_fn <-run_timer_softirq
1776 [...]
1777 bash-1994 1dNs2 81us : cpu_needs_another_gp <-rcu_process_callbacks
1778 bash-1994 1dNs2 82us : __local_bh_enable <-__do_softirq
1779 bash-1994 1dNs2 82us : sub_preempt_count <-__local_bh_enable
1780 bash-1994 1dN.2 82us : idle_cpu <-irq_exit
1781 bash-1994 1dN.2 83us : rcu_irq_exit <-irq_exit
1782 bash-1994 1dN.2 83us : sub_preempt_count <-irq_exit
1783 bash-1994 1.N.1 84us : _raw_spin_unlock_irqrestore <-task_rq_unlock
1784 bash-1994 1.N.1 84us+: trace_preempt_on <-task_rq_unlock
1785 bash-1994 1.N.1 104us : <stack trace>
1786 => sub_preempt_count
1787 => _raw_spin_unlock_irqrestore
1788 => task_rq_unlock
1789 => wake_up_new_task
1790 => do_fork
1791 => sys_clone
1792 => stub_clone
1795 The above is an example of the preemptoff trace with
1796 function-trace set. Here we see that interrupts were not disabled
1797 the entire time. The irq_enter code lets us know that we entered
1798 an interrupt 'h'. Before that, the functions being traced still
1799 show that it is not in an interrupt, but we can see from the
1800 functions themselves that this is not the case.
1802 preemptirqsoff
1803 --------------
1805 Knowing the locations that have interrupts disabled or
1806 preemption disabled for the longest times is helpful. But
1807 sometimes we would like to know when either preemption and/or
1808 interrupts are disabled.
1810 Consider the following code::
1812 local_irq_disable();
1813 call_function_with_irqs_off();
1814 preempt_disable();
1815 call_function_with_irqs_and_preemption_off();
1816 local_irq_enable();
1817 call_function_with_preemption_off();
1818 preempt_enable();
1820 The irqsoff tracer will record the total length of
1821 call_function_with_irqs_off() and
1822 call_function_with_irqs_and_preemption_off().
1824 The preemptoff tracer will record the total length of
1825 call_function_with_irqs_and_preemption_off() and
1826 call_function_with_preemption_off().
1828 But neither will trace the time that interrupts and/or
1829 preemption is disabled. This total time is the time that we can
1830 not schedule. To record this time, use the preemptirqsoff
1831 tracer.
1833 Again, using this trace is much like the irqsoff and preemptoff
1834 tracers.
1835 ::
1837 # echo 0 > options/function-trace
1838 # echo preemptirqsoff > current_tracer
1839 # echo 1 > tracing_on
1840 # echo 0 > tracing_max_latency
1841 # ls -ltr
1842 [...]
1843 # echo 0 > tracing_on
1844 # cat trace
1845 # tracer: preemptirqsoff
1846 #
1847 # preemptirqsoff latency trace v1.1.5 on 3.8.0-test+
1848 # --------------------------------------------------------------------
1849 # latency: 100 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
1850 # -----------------
1851 # | task: ls-2230 (uid:0 nice:0 policy:0 rt_prio:0)
1852 # -----------------
1853 # => started at: ata_scsi_queuecmd
1854 # => ended at: ata_scsi_queuecmd
1855 #
1856 #
1857 # _------=> CPU#
1858 # / _-----=> irqs-off
1859 # | / _----=> need-resched
1860 # || / _---=> hardirq/softirq
1861 # ||| / _--=> preempt-depth
1862 # |||| / delay
1863 # cmd pid ||||| time | caller
1864 # \ / ||||| \ | /
1865 ls-2230 3d... 0us+: _raw_spin_lock_irqsave <-ata_scsi_queuecmd
1866 ls-2230 3...1 100us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
1867 ls-2230 3...1 101us+: trace_preempt_on <-ata_scsi_queuecmd
1868 ls-2230 3...1 111us : <stack trace>
1869 => sub_preempt_count
1870 => _raw_spin_unlock_irqrestore
1871 => ata_scsi_queuecmd
1872 => scsi_dispatch_cmd
1873 => scsi_request_fn
1874 => __blk_run_queue_uncond
1875 => __blk_run_queue
1876 => blk_queue_bio
1877 => submit_bio_noacct
1878 => submit_bio
1879 => submit_bh
1880 => ext3_bread
1881 => ext3_dir_bread
1882 => htree_dirblock_to_tree
1883 => ext3_htree_fill_tree
1884 => ext3_readdir
1885 => vfs_readdir
1886 => sys_getdents
1887 => system_call_fastpath
1890 The trace_hardirqs_off_thunk is called from assembly on x86 when
1891 interrupts are disabled in the assembly code. Without the
1892 function tracing, we do not know if interrupts were enabled
1893 within the preemption points. We do see that it started with
1894 preemption enabled.
1896 Here is a trace with function-trace set::
1898 # tracer: preemptirqsoff
1899 #
1900 # preemptirqsoff latency trace v1.1.5 on 3.8.0-test+
1901 # --------------------------------------------------------------------
1902 # latency: 161 us, #339/339, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
1903 # -----------------
1904 # | task: ls-2269 (uid:0 nice:0 policy:0 rt_prio:0)
1905 # -----------------
1906 # => started at: schedule
1907 # => ended at: mutex_unlock
1908 #
1909 #
1910 # _------=> CPU#
1911 # / _-----=> irqs-off
1912 # | / _----=> need-resched
1913 # || / _---=> hardirq/softirq
1914 # ||| / _--=> preempt-depth
1915 # |||| / delay
1916 # cmd pid ||||| time | caller
1917 # \ / ||||| \ | /
1918 kworker/-59 3...1 0us : __schedule <-schedule
1919 kworker/-59 3d..1 0us : rcu_preempt_qs <-rcu_note_context_switch
1920 kworker/-59 3d..1 1us : add_preempt_count <-_raw_spin_lock_irq
1921 kworker/-59 3d..2 1us : deactivate_task <-__schedule
1922 kworker/-59 3d..2 1us : dequeue_task <-deactivate_task
1923 kworker/-59 3d..2 2us : update_rq_clock <-dequeue_task
1924 kworker/-59 3d..2 2us : dequeue_task_fair <-dequeue_task
1925 kworker/-59 3d..2 2us : update_curr <-dequeue_task_fair
1926 kworker/-59 3d..2 2us : update_min_vruntime <-update_curr
1927 kworker/-59 3d..2 3us : cpuacct_charge <-update_curr
1928 kworker/-59 3d..2 3us : __rcu_read_lock <-cpuacct_charge
1929 kworker/-59 3d..2 3us : __rcu_read_unlock <-cpuacct_charge
1930 kworker/-59 3d..2 3us : update_cfs_rq_blocked_load <-dequeue_task_fair
1931 kworker/-59 3d..2 4us : clear_buddies <-dequeue_task_fair
1932 kworker/-59 3d..2 4us : account_entity_dequeue <-dequeue_task_fair
1933 kworker/-59 3d..2 4us : update_min_vruntime <-dequeue_task_fair
1934 kworker/-59 3d..2 4us : update_cfs_shares <-dequeue_task_fair
1935 kworker/-59 3d..2 5us : hrtick_update <-dequeue_task_fair
1936 kworker/-59 3d..2 5us : wq_worker_sleeping <-__schedule
1937 kworker/-59 3d..2 5us : kthread_data <-wq_worker_sleeping
1938 kworker/-59 3d..2 5us : put_prev_task_fair <-__schedule
1939 kworker/-59 3d..2 6us : pick_next_task_fair <-pick_next_task
1940 kworker/-59 3d..2 6us : clear_buddies <-pick_next_task_fair
1941 kworker/-59 3d..2 6us : set_next_entity <-pick_next_task_fair
1942 kworker/-59 3d..2 6us : update_stats_wait_end <-set_next_entity
1943 ls-2269 3d..2 7us : finish_task_switch <-__schedule
1944 ls-2269 3d..2 7us : _raw_spin_unlock_irq <-finish_task_switch
1945 ls-2269 3d..2 8us : do_IRQ <-ret_from_intr
1946 ls-2269 3d..2 8us : irq_enter <-do_IRQ
1947 ls-2269 3d..2 8us : rcu_irq_enter <-irq_enter
1948 ls-2269 3d..2 9us : add_preempt_count <-irq_enter
1949 ls-2269 3d.h2 9us : exit_idle <-do_IRQ
1950 [...]
1951 ls-2269 3d.h3 20us : sub_preempt_count <-_raw_spin_unlock
1952 ls-2269 3d.h2 20us : irq_exit <-do_IRQ
1953 ls-2269 3d.h2 21us : sub_preempt_count <-irq_exit
1954 ls-2269 3d..3 21us : do_softirq <-irq_exit
1955 ls-2269 3d..3 21us : __do_softirq <-call_softirq
1956 ls-2269 3d..3 21us+: __local_bh_disable <-__do_softirq
1957 ls-2269 3d.s4 29us : sub_preempt_count <-_local_bh_enable_ip
1958 ls-2269 3d.s5 29us : sub_preempt_count <-_local_bh_enable_ip
1959 ls-2269 3d.s5 31us : do_IRQ <-ret_from_intr
1960 ls-2269 3d.s5 31us : irq_enter <-do_IRQ
1961 ls-2269 3d.s5 31us : rcu_irq_enter <-irq_enter
1962 [...]
1963 ls-2269 3d.s5 31us : rcu_irq_enter <-irq_enter
1964 ls-2269 3d.s5 32us : add_preempt_count <-irq_enter
1965 ls-2269 3d.H5 32us : exit_idle <-do_IRQ
1966 ls-2269 3d.H5 32us : handle_irq <-do_IRQ
1967 ls-2269 3d.H5 32us : irq_to_desc <-handle_irq
1968 ls-2269 3d.H5 33us : handle_fasteoi_irq <-handle_irq
1969 [...]
1970 ls-2269 3d.s5 158us : _raw_spin_unlock_irqrestore <-rtl8139_poll
1971 ls-2269 3d.s3 158us : net_rps_action_and_irq_enable.isra.65 <-net_rx_action
1972 ls-2269 3d.s3 159us : __local_bh_enable <-__do_softirq
1973 ls-2269 3d.s3 159us : sub_preempt_count <-__local_bh_enable
1974 ls-2269 3d..3 159us : idle_cpu <-irq_exit
1975 ls-2269 3d..3 159us : rcu_irq_exit <-irq_exit
1976 ls-2269 3d..3 160us : sub_preempt_count <-irq_exit
1977 ls-2269 3d... 161us : __mutex_unlock_slowpath <-mutex_unlock
1978 ls-2269 3d... 162us+: trace_hardirqs_on <-mutex_unlock
1979 ls-2269 3d... 186us : <stack trace>
1980 => __mutex_unlock_slowpath
1981 => mutex_unlock
1982 => process_output
1983 => n_tty_write
1984 => tty_write
1985 => vfs_write
1986 => sys_write
1987 => system_call_fastpath
1989 This is an interesting trace. It started with kworker running and
1990 scheduling out and ls taking over. But as soon as ls released the
1991 rq lock and enabled interrupts (but not preemption) an interrupt
1992 triggered. When the interrupt finished, it started running softirqs.
1993 But while the softirq was running, another interrupt triggered.
1994 When an interrupt is running inside a softirq, the annotation is 'H'.
1997 wakeup
1998 ------
2000 One common case that people are interested in tracing is the
2001 time it takes for a task that is woken to actually wake up.
2002 Now for non Real-Time tasks, this can be arbitrary. But tracing
2003 it nonetheless can be interesting.
2005 Without function tracing::
2007 # echo 0 > options/function-trace
2008 # echo wakeup > current_tracer
2009 # echo 1 > tracing_on
2010 # echo 0 > tracing_max_latency
2011 # chrt -f 5 sleep 1
2012 # echo 0 > tracing_on
2013 # cat trace
2014 # tracer: wakeup
2015 #
2016 # wakeup latency trace v1.1.5 on 3.8.0-test+
2017 # --------------------------------------------------------------------
2018 # latency: 15 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
2019 # -----------------
2020 # | task: kworker/3:1H-312 (uid:0 nice:-20 policy:0 rt_prio:0)
2021 # -----------------
2022 #
2023 # _------=> CPU#
2024 # / _-----=> irqs-off
2025 # | / _----=> need-resched
2026 # || / _---=> hardirq/softirq
2027 # ||| / _--=> preempt-depth
2028 # |||| / delay
2029 # cmd pid ||||| time | caller
2030 # \ / ||||| \ | /
2031 <idle>-0 3dNs7 0us : 0:120:R + [003] 312:100:R kworker/3:1H
2032 <idle>-0 3dNs7 1us+: ttwu_do_activate.constprop.87 <-try_to_wake_up
2033 <idle>-0 3d..3 15us : __schedule <-schedule
2034 <idle>-0 3d..3 15us : 0:120:R ==> [003] 312:100:R kworker/3:1H
2036 The tracer only traces the highest priority task in the system
2037 to avoid tracing the normal circumstances. Here we see that
2038 the kworker with a nice priority of -20 (not very nice), took
2039 just 15 microseconds from the time it woke up, to the time it
2040 ran.
2042 Non Real-Time tasks are not that interesting. A more interesting
2043 trace is to concentrate only on Real-Time tasks.
2045 wakeup_rt
2046 ---------
2048 In a Real-Time environment it is very important to know the
2049 wakeup time it takes for the highest priority task that is woken
2050 up to the time that it executes. This is also known as "schedule
2051 latency". I stress the point that this is about RT tasks. It is
2052 also important to know the scheduling latency of non-RT tasks,
2053 but the average schedule latency is better for non-RT tasks.
2054 Tools like LatencyTop are more appropriate for such
2055 measurements.
2057 Real-Time environments are interested in the worst case latency.
2058 That is the longest latency it takes for something to happen,
2059 and not the average. We can have a very fast scheduler that may
2060 only have a large latency once in a while, but that would not
2061 work well with Real-Time tasks. The wakeup_rt tracer was designed
2062 to record the worst case wakeups of RT tasks. Non-RT tasks are
2063 not recorded because the tracer only records one worst case and
2064 tracing non-RT tasks that are unpredictable will overwrite the
2065 worst case latency of RT tasks (just run the normal wakeup
2066 tracer for a while to see that effect).
2068 Since this tracer only deals with RT tasks, we will run this
2069 slightly differently than we did with the previous tracers.
2070 Instead of performing an 'ls', we will run 'sleep 1' under
2071 'chrt' which changes the priority of the task.
2072 ::
2074 # echo 0 > options/function-trace
2075 # echo wakeup_rt > current_tracer
2076 # echo 1 > tracing_on
2077 # echo 0 > tracing_max_latency
2078 # chrt -f 5 sleep 1
2079 # echo 0 > tracing_on
2080 # cat trace
2081 # tracer: wakeup
2082 #
2083 # tracer: wakeup_rt
2084 #
2085 # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
2086 # --------------------------------------------------------------------
2087 # latency: 5 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
2088 # -----------------
2089 # | task: sleep-2389 (uid:0 nice:0 policy:1 rt_prio:5)
2090 # -----------------
2091 #
2092 # _------=> CPU#
2093 # / _-----=> irqs-off
2094 # | / _----=> need-resched
2095 # || / _---=> hardirq/softirq
2096 # ||| / _--=> preempt-depth
2097 # |||| / delay
2098 # cmd pid ||||| time | caller
2099 # \ / ||||| \ | /
2100 <idle>-0 3d.h4 0us : 0:120:R + [003] 2389: 94:R sleep
2101 <idle>-0 3d.h4 1us+: ttwu_do_activate.constprop.87 <-try_to_wake_up
2102 <idle>-0 3d..3 5us : __schedule <-schedule
2103 <idle>-0 3d..3 5us : 0:120:R ==> [003] 2389: 94:R sleep
2106 Running this on an idle system, we see that it only took 5 microseconds
2107 to perform the task switch. Note, since the trace point in the schedule
2108 is before the actual "switch", we stop the tracing when the recorded task
2109 is about to schedule in. This may change if we add a new marker at the
2110 end of the scheduler.
2112 Notice that the recorded task is 'sleep' with the PID of 2389
2113 and it has an rt_prio of 5. This priority is user-space priority
2114 and not the internal kernel priority. The policy is 1 for
2115 SCHED_FIFO and 2 for SCHED_RR.
2117 Note, that the trace data shows the internal priority (99 - rtprio).
2118 ::
2120 <idle>-0 3d..3 5us : 0:120:R ==> [003] 2389: 94:R sleep
2122 The 0:120:R means idle was running with a nice priority of 0 (120 - 120)
2123 and in the running state 'R'. The sleep task was scheduled in with
2124 2389: 94:R. That is the priority is the kernel rtprio (99 - 5 = 94)
2125 and it too is in the running state.
2127 Doing the same with chrt -r 5 and function-trace set.
2128 ::
2130 echo 1 > options/function-trace
2132 # tracer: wakeup_rt
2133 #
2134 # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
2135 # --------------------------------------------------------------------
2136 # latency: 29 us, #85/85, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
2137 # -----------------
2138 # | task: sleep-2448 (uid:0 nice:0 policy:1 rt_prio:5)
2139 # -----------------
2140 #
2141 # _------=> CPU#
2142 # / _-----=> irqs-off
2143 # | / _----=> need-resched
2144 # || / _---=> hardirq/softirq
2145 # ||| / _--=> preempt-depth
2146 # |||| / delay
2147 # cmd pid ||||| time | caller
2148 # \ / ||||| \ | /
2149 <idle>-0 3d.h4 1us+: 0:120:R + [003] 2448: 94:R sleep
2150 <idle>-0 3d.h4 2us : ttwu_do_activate.constprop.87 <-try_to_wake_up
2151 <idle>-0 3d.h3 3us : check_preempt_curr <-ttwu_do_wakeup
2152 <idle>-0 3d.h3 3us : resched_curr <-check_preempt_curr
2153 <idle>-0 3dNh3 4us : task_woken_rt <-ttwu_do_wakeup
2154 <idle>-0 3dNh3 4us : _raw_spin_unlock <-try_to_wake_up
2155 <idle>-0 3dNh3 4us : sub_preempt_count <-_raw_spin_unlock
2156 <idle>-0 3dNh2 5us : ttwu_stat <-try_to_wake_up
2157 <idle>-0 3dNh2 5us : _raw_spin_unlock_irqrestore <-try_to_wake_up
2158 <idle>-0 3dNh2 6us : sub_preempt_count <-_raw_spin_unlock_irqrestore
2159 <idle>-0 3dNh1 6us : _raw_spin_lock <-__run_hrtimer
2160 <idle>-0 3dNh1 6us : add_preempt_count <-_raw_spin_lock
2161 <idle>-0 3dNh2 7us : _raw_spin_unlock <-hrtimer_interrupt
2162 <idle>-0 3dNh2 7us : sub_preempt_count <-_raw_spin_unlock
2163 <idle>-0 3dNh1 7us : tick_program_event <-hrtimer_interrupt
2164 <idle>-0 3dNh1 7us : clockevents_program_event <-tick_program_event
2165 <idle>-0 3dNh1 8us : ktime_get <-clockevents_program_event
2166 <idle>-0 3dNh1 8us : lapic_next_event <-clockevents_program_event
2167 <idle>-0 3dNh1 8us : irq_exit <-smp_apic_timer_interrupt
2168 <idle>-0 3dNh1 9us : sub_preempt_count <-irq_exit
2169 <idle>-0 3dN.2 9us : idle_cpu <-irq_exit
2170 <idle>-0 3dN.2 9us : rcu_irq_exit <-irq_exit
2171 <idle>-0 3dN.2 10us : rcu_eqs_enter_common.isra.45 <-rcu_irq_exit
2172 <idle>-0 3dN.2 10us : sub_preempt_count <-irq_exit
2173 <idle>-0 3.N.1 11us : rcu_idle_exit <-cpu_idle
2174 <idle>-0 3dN.1 11us : rcu_eqs_exit_common.isra.43 <-rcu_idle_exit
2175 <idle>-0 3.N.1 11us : tick_nohz_idle_exit <-cpu_idle
2176 <idle>-0 3dN.1 12us : menu_hrtimer_cancel <-tick_nohz_idle_exit
2177 <idle>-0 3dN.1 12us : ktime_get <-tick_nohz_idle_exit
2178 <idle>-0 3dN.1 12us : tick_do_update_jiffies64 <-tick_nohz_idle_exit
2179 <idle>-0 3dN.1 13us : cpu_load_update_nohz <-tick_nohz_idle_exit
2180 <idle>-0 3dN.1 13us : _raw_spin_lock <-cpu_load_update_nohz
2181 <idle>-0 3dN.1 13us : add_preempt_count <-_raw_spin_lock
2182 <idle>-0 3dN.2 13us : __cpu_load_update <-cpu_load_update_nohz
2183 <idle>-0 3dN.2 14us : sched_avg_update <-__cpu_load_update
2184 <idle>-0 3dN.2 14us : _raw_spin_unlock <-cpu_load_update_nohz
2185 <idle>-0 3dN.2 14us : sub_preempt_count <-_raw_spin_unlock
2186 <idle>-0 3dN.1 15us : calc_load_nohz_stop <-tick_nohz_idle_exit
2187 <idle>-0 3dN.1 15us : touch_softlockup_watchdog <-tick_nohz_idle_exit
2188 <idle>-0 3dN.1 15us : hrtimer_cancel <-tick_nohz_idle_exit
2189 <idle>-0 3dN.1 15us : hrtimer_try_to_cancel <-hrtimer_cancel
2190 <idle>-0 3dN.1 16us : lock_hrtimer_base.isra.18 <-hrtimer_try_to_cancel
2191 <idle>-0 3dN.1 16us : _raw_spin_lock_irqsave <-lock_hrtimer_base.isra.18
2192 <idle>-0 3dN.1 16us : add_preempt_count <-_raw_spin_lock_irqsave
2193 <idle>-0 3dN.2 17us : __remove_hrtimer <-remove_hrtimer.part.16
2194 <idle>-0 3dN.2 17us : hrtimer_force_reprogram <-__remove_hrtimer
2195 <idle>-0 3dN.2 17us : tick_program_event <-hrtimer_force_reprogram
2196 <idle>-0 3dN.2 18us : clockevents_program_event <-tick_program_event
2197 <idle>-0 3dN.2 18us : ktime_get <-clockevents_program_event
2198 <idle>-0 3dN.2 18us : lapic_next_event <-clockevents_program_event
2199 <idle>-0 3dN.2 19us : _raw_spin_unlock_irqrestore <-hrtimer_try_to_cancel
2200 <idle>-0 3dN.2 19us : sub_preempt_count <-_raw_spin_unlock_irqrestore
2201 <idle>-0 3dN.1 19us : hrtimer_forward <-tick_nohz_idle_exit
2202 <idle>-0 3dN.1 20us : ktime_add_safe <-hrtimer_forward
2203 <idle>-0 3dN.1 20us : ktime_add_safe <-hrtimer_forward
2204 <idle>-0 3dN.1 20us : hrtimer_start_range_ns <-hrtimer_start_expires.constprop.11
2205 <idle>-0 3dN.1 20us : __hrtimer_start_range_ns <-hrtimer_start_range_ns
2206 <idle>-0 3dN.1 21us : lock_hrtimer_base.isra.18 <-__hrtimer_start_range_ns
2207 <idle>-0 3dN.1 21us : _raw_spin_lock_irqsave <-lock_hrtimer_base.isra.18
2208 <idle>-0 3dN.1 21us : add_preempt_count <-_raw_spin_lock_irqsave
2209 <idle>-0 3dN.2 22us : ktime_add_safe <-__hrtimer_start_range_ns
2210 <idle>-0 3dN.2 22us : enqueue_hrtimer <-__hrtimer_start_range_ns
2211 <idle>-0 3dN.2 22us : tick_program_event <-__hrtimer_start_range_ns
2212 <idle>-0 3dN.2 23us : clockevents_program_event <-tick_program_event
2213 <idle>-0 3dN.2 23us : ktime_get <-clockevents_program_event
2214 <idle>-0 3dN.2 23us : lapic_next_event <-clockevents_program_event
2215 <idle>-0 3dN.2 24us : _raw_spin_unlock_irqrestore <-__hrtimer_start_range_ns
2216 <idle>-0 3dN.2 24us : sub_preempt_count <-_raw_spin_unlock_irqrestore
2217 <idle>-0 3dN.1 24us : account_idle_ticks <-tick_nohz_idle_exit
2218 <idle>-0 3dN.1 24us : account_idle_time <-account_idle_ticks
2219 <idle>-0 3.N.1 25us : sub_preempt_count <-cpu_idle
2220 <idle>-0 3.N.. 25us : schedule <-cpu_idle
2221 <idle>-0 3.N.. 25us : __schedule <-preempt_schedule
2222 <idle>-0 3.N.. 26us : add_preempt_count <-__schedule
2223 <idle>-0 3.N.1 26us : rcu_note_context_switch <-__schedule
2224 <idle>-0 3.N.1 26us : rcu_sched_qs <-rcu_note_context_switch
2225 <idle>-0 3dN.1 27us : rcu_preempt_qs <-rcu_note_context_switch
2226 <idle>-0 3.N.1 27us : _raw_spin_lock_irq <-__schedule
2227 <idle>-0 3dN.1 27us : add_preempt_count <-_raw_spin_lock_irq
2228 <idle>-0 3dN.2 28us : put_prev_task_idle <-__schedule
2229 <idle>-0 3dN.2 28us : pick_next_task_stop <-pick_next_task
2230 <idle>-0 3dN.2 28us : pick_next_task_rt <-pick_next_task
2231 <idle>-0 3dN.2 29us : dequeue_pushable_task <-pick_next_task_rt
2232 <idle>-0 3d..3 29us : __schedule <-preempt_schedule
2233 <idle>-0 3d..3 30us : 0:120:R ==> [003] 2448: 94:R sleep
2235 This isn't that big of a trace, even with function tracing enabled,
2236 so I included the entire trace.
2238 The interrupt went off while when the system was idle. Somewhere
2239 before task_woken_rt() was called, the NEED_RESCHED flag was set,
2240 this is indicated by the first occurrence of the 'N' flag.
2242 Latency tracing and events
2243 --------------------------
2244 As function tracing can induce a much larger latency, but without
2245 seeing what happens within the latency it is hard to know what
2246 caused it. There is a middle ground, and that is with enabling
2247 events.
2248 ::
2250 # echo 0 > options/function-trace
2251 # echo wakeup_rt > current_tracer
2252 # echo 1 > events/enable
2253 # echo 1 > tracing_on
2254 # echo 0 > tracing_max_latency
2255 # chrt -f 5 sleep 1
2256 # echo 0 > tracing_on
2257 # cat trace
2258 # tracer: wakeup_rt
2259 #
2260 # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
2261 # --------------------------------------------------------------------
2262 # latency: 6 us, #12/12, CPU#2 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
2263 # -----------------
2264 # | task: sleep-5882 (uid:0 nice:0 policy:1 rt_prio:5)
2265 # -----------------
2266 #
2267 # _------=> CPU#
2268 # / _-----=> irqs-off
2269 # | / _----=> need-resched
2270 # || / _---=> hardirq/softirq
2271 # ||| / _--=> preempt-depth
2272 # |||| / delay
2273 # cmd pid ||||| time | caller
2274 # \ / ||||| \ | /
2275 <idle>-0 2d.h4 0us : 0:120:R + [002] 5882: 94:R sleep
2276 <idle>-0 2d.h4 0us : ttwu_do_activate.constprop.87 <-try_to_wake_up
2277 <idle>-0 2d.h4 1us : sched_wakeup: comm=sleep pid=5882 prio=94 success=1 target_cpu=002
2278 <idle>-0 2dNh2 1us : hrtimer_expire_exit: hrtimer=ffff88007796feb8
2279 <idle>-0 2.N.2 2us : power_end: cpu_id=2
2280 <idle>-0 2.N.2 3us : cpu_idle: state=4294967295 cpu_id=2
2281 <idle>-0 2dN.3 4us : hrtimer_cancel: hrtimer=ffff88007d50d5e0
2282 <idle>-0 2dN.3 4us : hrtimer_start: hrtimer=ffff88007d50d5e0 function=tick_sched_timer expires=34311211000000 softexpires=34311211000000
2283 <idle>-0 2.N.2 5us : rcu_utilization: Start context switch
2284 <idle>-0 2.N.2 5us : rcu_utilization: End context switch
2285 <idle>-0 2d..3 6us : __schedule <-schedule
2286 <idle>-0 2d..3 6us : 0:120:R ==> [002] 5882: 94:R sleep
2289 Hardware Latency Detector
2290 -------------------------
2292 The hardware latency detector is executed by enabling the "hwlat" tracer.
2294 NOTE, this tracer will affect the performance of the system as it will
2295 periodically make a CPU constantly busy with interrupts disabled.
2296 ::
2298 # echo hwlat > current_tracer
2299 # sleep 100
2300 # cat trace
2301 # tracer: hwlat
2302 #
2303 # entries-in-buffer/entries-written: 13/13 #P:8
2304 #
2305 # _-----=> irqs-off
2306 # / _----=> need-resched
2307 # | / _---=> hardirq/softirq
2308 # || / _--=> preempt-depth
2309 # ||| / delay
2310 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
2311 # | | | |||| | |
2312 <...>-1729 [001] d... 678.473449: #1 inner/outer(us): 11/12 ts:1581527483.343962693 count:6
2313 <...>-1729 [004] d... 689.556542: #2 inner/outer(us): 16/9 ts:1581527494.889008092 count:1
2314 <...>-1729 [005] d... 714.756290: #3 inner/outer(us): 16/16 ts:1581527519.678961629 count:5
2315 <...>-1729 [001] d... 718.788247: #4 inner/outer(us): 9/17 ts:1581527523.889012713 count:1
2316 <...>-1729 [002] d... 719.796341: #5 inner/outer(us): 13/9 ts:1581527524.912872606 count:1
2317 <...>-1729 [006] d... 844.787091: #6 inner/outer(us): 9/12 ts:1581527649.889048502 count:2
2318 <...>-1729 [003] d... 849.827033: #7 inner/outer(us): 18/9 ts:1581527654.889013793 count:1
2319 <...>-1729 [007] d... 853.859002: #8 inner/outer(us): 9/12 ts:1581527658.889065736 count:1
2320 <...>-1729 [001] d... 855.874978: #9 inner/outer(us): 9/11 ts:1581527660.861991877 count:1
2321 <...>-1729 [001] d... 863.938932: #10 inner/outer(us): 9/11 ts:1581527668.970010500 count:1 nmi-total:7 nmi-count:1
2322 <...>-1729 [007] d... 878.050780: #11 inner/outer(us): 9/12 ts:1581527683.385002600 count:1 nmi-total:5 nmi-count:1
2323 <...>-1729 [007] d... 886.114702: #12 inner/outer(us): 9/12 ts:1581527691.385001600 count:1
2326 The above output is somewhat the same in the header. All events will have
2327 interrupts disabled 'd'. Under the FUNCTION title there is:
2329 #1
2330 This is the count of events recorded that were greater than the
2331 tracing_threshold (See below).
2333 inner/outer(us): 11/11
2335 This shows two numbers as "inner latency" and "outer latency". The test
2336 runs in a loop checking a timestamp twice. The latency detected within
2337 the two timestamps is the "inner latency" and the latency detected
2338 after the previous timestamp and the next timestamp in the loop is
2339 the "outer latency".
2341 ts:1581527483.343962693
2343 The absolute timestamp that the first latency was recorded in the window.
2345 count:6
2347 The number of times a latency was detected during the window.
2349 nmi-total:7 nmi-count:1
2351 On architectures that support it, if an NMI comes in during the
2352 test, the time spent in NMI is reported in "nmi-total" (in
2353 microseconds).
2355 All architectures that have NMIs will show the "nmi-count" if an
2356 NMI comes in during the test.
2358 hwlat files:
2360 tracing_threshold
2361 This gets automatically set to "10" to represent 10
2362 microseconds. This is the threshold of latency that
2363 needs to be detected before the trace will be recorded.
2365 Note, when hwlat tracer is finished (another tracer is
2366 written into "current_tracer"), the original value for
2367 tracing_threshold is placed back into this file.
2369 hwlat_detector/width
2370 The length of time the test runs with interrupts disabled.
2372 hwlat_detector/window
2373 The length of time of the window which the test
2374 runs. That is, the test will run for "width"
2375 microseconds per "window" microseconds
2377 tracing_cpumask
2378 When the test is started. A kernel thread is created that
2379 runs the test. This thread will alternate between CPUs
2380 listed in the tracing_cpumask between each period
2381 (one "window"). To limit the test to specific CPUs
2382 set the mask in this file to only the CPUs that the test
2383 should run on.
2385 function
2386 --------
2388 This tracer is the function tracer. Enabling the function tracer
2389 can be done from the debug file system. Make sure the
2390 ftrace_enabled is set; otherwise this tracer is a nop.
2391 See the "ftrace_enabled" section below.
2392 ::
2394 # sysctl kernel.ftrace_enabled=1
2395 # echo function > current_tracer
2396 # echo 1 > tracing_on
2397 # usleep 1
2398 # echo 0 > tracing_on
2399 # cat trace
2400 # tracer: function
2401 #
2402 # entries-in-buffer/entries-written: 24799/24799 #P:4
2403 #
2404 # _-----=> irqs-off
2405 # / _----=> need-resched
2406 # | / _---=> hardirq/softirq
2407 # || / _--=> preempt-depth
2408 # ||| / delay
2409 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
2410 # | | | |||| | |
2411 bash-1994 [002] .... 3082.063030: mutex_unlock <-rb_simple_write
2412 bash-1994 [002] .... 3082.063031: __mutex_unlock_slowpath <-mutex_unlock
2413 bash-1994 [002] .... 3082.063031: __fsnotify_parent <-fsnotify_modify
2414 bash-1994 [002] .... 3082.063032: fsnotify <-fsnotify_modify
2415 bash-1994 [002] .... 3082.063032: __srcu_read_lock <-fsnotify
2416 bash-1994 [002] .... 3082.063032: add_preempt_count <-__srcu_read_lock
2417 bash-1994 [002] ...1 3082.063032: sub_preempt_count <-__srcu_read_lock
2418 bash-1994 [002] .... 3082.063033: __srcu_read_unlock <-fsnotify
2419 [...]
2422 Note: function tracer uses ring buffers to store the above
2423 entries. The newest data may overwrite the oldest data.
2424 Sometimes using echo to stop the trace is not sufficient because
2425 the tracing could have overwritten the data that you wanted to
2426 record. For this reason, it is sometimes better to disable
2427 tracing directly from a program. This allows you to stop the
2428 tracing at the point that you hit the part that you are
2429 interested in. To disable the tracing directly from a C program,
2430 something like following code snippet can be used::
2432 int trace_fd;
2433 [...]
2434 int main(int argc, char *argv[]) {
2435 [...]
2436 trace_fd = open(tracing_file("tracing_on"), O_WRONLY);
2437 [...]
2438 if (condition_hit()) {
2439 write(trace_fd, "0", 1);
2440 }
2441 [...]
2442 }
2445 Single thread tracing
2446 ---------------------
2448 By writing into set_ftrace_pid you can trace a
2449 single thread. For example::
2451 # cat set_ftrace_pid
2452 no pid
2453 # echo 3111 > set_ftrace_pid
2454 # cat set_ftrace_pid
2455 3111
2456 # echo function > current_tracer
2457 # cat trace | head
2458 # tracer: function
2459 #
2460 # TASK-PID CPU# TIMESTAMP FUNCTION
2461 # | | | | |
2462 yum-updatesd-3111 [003] 1637.254676: finish_task_switch <-thread_return
2463 yum-updatesd-3111 [003] 1637.254681: hrtimer_cancel <-schedule_hrtimeout_range
2464 yum-updatesd-3111 [003] 1637.254682: hrtimer_try_to_cancel <-hrtimer_cancel
2465 yum-updatesd-3111 [003] 1637.254683: lock_hrtimer_base <-hrtimer_try_to_cancel
2466 yum-updatesd-3111 [003] 1637.254685: fget_light <-do_sys_poll
2467 yum-updatesd-3111 [003] 1637.254686: pipe_poll <-do_sys_poll
2468 # echo > set_ftrace_pid
2469 # cat trace |head
2470 # tracer: function
2471 #
2472 # TASK-PID CPU# TIMESTAMP FUNCTION
2473 # | | | | |
2474 ##### CPU 3 buffer started ####
2475 yum-updatesd-3111 [003] 1701.957688: free_poll_entry <-poll_freewait
2476 yum-updatesd-3111 [003] 1701.957689: remove_wait_queue <-free_poll_entry
2477 yum-updatesd-3111 [003] 1701.957691: fput <-free_poll_entry
2478 yum-updatesd-3111 [003] 1701.957692: audit_syscall_exit <-sysret_audit
2479 yum-updatesd-3111 [003] 1701.957693: path_put <-audit_syscall_exit
2481 If you want to trace a function when executing, you could use
2482 something like this simple program.
2483 ::
2485 #include <stdio.h>
2486 #include <stdlib.h>
2487 #include <sys/types.h>
2488 #include <sys/stat.h>
2489 #include <fcntl.h>
2490 #include <unistd.h>
2491 #include <string.h>
2493 #define _STR(x) #x
2494 #define STR(x) _STR(x)
2495 #define MAX_PATH 256
2497 const char *find_tracefs(void)
2498 {
2499 static char tracefs[MAX_PATH+1];
2500 static int tracefs_found;
2501 char type[100];
2502 FILE *fp;
2504 if (tracefs_found)
2505 return tracefs;
2507 if ((fp = fopen("/proc/mounts","r")) == NULL) {
2508 perror("/proc/mounts");
2509 return NULL;
2510 }
2512 while (fscanf(fp, "%*s %"
2513 STR(MAX_PATH)
2514 "s %99s %*s %*d %*d\n",
2515 tracefs, type) == 2) {
2516 if (strcmp(type, "tracefs") == 0)
2517 break;
2518 }
2519 fclose(fp);
2521 if (strcmp(type, "tracefs") != 0) {
2522 fprintf(stderr, "tracefs not mounted");
2523 return NULL;
2524 }
2526 strcat(tracefs, "/tracing/");
2527 tracefs_found = 1;
2529 return tracefs;
2530 }
2532 const char *tracing_file(const char *file_name)
2533 {
2534 static char trace_file[MAX_PATH+1];
2535 snprintf(trace_file, MAX_PATH, "%s/%s", find_tracefs(), file_name);
2536 return trace_file;
2537 }
2539 int main (int argc, char **argv)
2540 {
2541 if (argc < 1)
2542 exit(-1);
2544 if (fork() > 0) {
2545 int fd, ffd;
2546 char line[64];
2547 int s;
2549 ffd = open(tracing_file("current_tracer"), O_WRONLY);
2550 if (ffd < 0)
2551 exit(-1);
2552 write(ffd, "nop", 3);
2554 fd = open(tracing_file("set_ftrace_pid"), O_WRONLY);
2555 s = sprintf(line, "%d\n", getpid());
2556 write(fd, line, s);
2558 write(ffd, "function", 8);
2560 close(fd);
2561 close(ffd);
2563 execvp(argv[1], argv+1);
2564 }
2566 return 0;
2567 }
2569 Or this simple script!
2570 ::
2572 #!/bin/bash
2574 tracefs=`sed -ne 's/^tracefs \(.*\) tracefs.*/\1/p' /proc/mounts`
2575 echo 0 > $tracefs/tracing_on
2576 echo $$ > $tracefs/set_ftrace_pid
2577 echo function > $tracefs/current_tracer
2578 echo 1 > $tracefs/tracing_on
2579 exec "$@"
2582 function graph tracer
2583 ---------------------------
2585 This tracer is similar to the function tracer except that it
2586 probes a function on its entry and its exit. This is done by
2587 using a dynamically allocated stack of return addresses in each
2588 task_struct. On function entry the tracer overwrites the return
2589 address of each function traced to set a custom probe. Thus the
2590 original return address is stored on the stack of return address
2591 in the task_struct.
2593 Probing on both ends of a function leads to special features
2594 such as:
2596 - measure of a function's time execution
2597 - having a reliable call stack to draw function calls graph
2599 This tracer is useful in several situations:
2601 - you want to find the reason of a strange kernel behavior and
2602 need to see what happens in detail on any areas (or specific
2603 ones).
2605 - you are experiencing weird latencies but it's difficult to
2606 find its origin.
2608 - you want to find quickly which path is taken by a specific
2609 function
2611 - you just want to peek inside a working kernel and want to see
2612 what happens there.
2614 ::
2616 # tracer: function_graph
2617 #
2618 # CPU DURATION FUNCTION CALLS
2619 # | | | | | | |
2621 0) | sys_open() {
2622 0) | do_sys_open() {
2623 0) | getname() {
2624 0) | kmem_cache_alloc() {
2625 0) 1.382 us | __might_sleep();
2626 0) 2.478 us | }
2627 0) | strncpy_from_user() {
2628 0) | might_fault() {
2629 0) 1.389 us | __might_sleep();
2630 0) 2.553 us | }
2631 0) 3.807 us | }
2632 0) 7.876 us | }
2633 0) | alloc_fd() {
2634 0) 0.668 us | _spin_lock();
2635 0) 0.570 us | expand_files();
2636 0) 0.586 us | _spin_unlock();
2639 There are several columns that can be dynamically
2640 enabled/disabled. You can use every combination of options you
2641 want, depending on your needs.
2643 - The cpu number on which the function executed is default
2644 enabled. It is sometimes better to only trace one cpu (see
2645 tracing_cpumask file) or you might sometimes see unordered
2646 function calls while cpu tracing switch.
2648 - hide: echo nofuncgraph-cpu > trace_options
2649 - show: echo funcgraph-cpu > trace_options
2651 - The duration (function's time of execution) is displayed on
2652 the closing bracket line of a function or on the same line
2653 than the current function in case of a leaf one. It is default
2654 enabled.
2656 - hide: echo nofuncgraph-duration > trace_options
2657 - show: echo funcgraph-duration > trace_options
2659 - The overhead field precedes the duration field in case of
2660 reached duration thresholds.
2662 - hide: echo nofuncgraph-overhead > trace_options
2663 - show: echo funcgraph-overhead > trace_options
2664 - depends on: funcgraph-duration
2666 ie::
2668 3) # 1837.709 us | } /* __switch_to */
2669 3) | finish_task_switch() {
2670 3) 0.313 us | _raw_spin_unlock_irq();
2671 3) 3.177 us | }
2672 3) # 1889.063 us | } /* __schedule */
2673 3) ! 140.417 us | } /* __schedule */
2674 3) # 2034.948 us | } /* schedule */
2675 3) * 33998.59 us | } /* schedule_preempt_disabled */
2677 [...]
2679 1) 0.260 us | msecs_to_jiffies();
2680 1) 0.313 us | __rcu_read_unlock();
2681 1) + 61.770 us | }
2682 1) + 64.479 us | }
2683 1) 0.313 us | rcu_bh_qs();
2684 1) 0.313 us | __local_bh_enable();
2685 1) ! 217.240 us | }
2686 1) 0.365 us | idle_cpu();
2687 1) | rcu_irq_exit() {
2688 1) 0.417 us | rcu_eqs_enter_common.isra.47();
2689 1) 3.125 us | }
2690 1) ! 227.812 us | }
2691 1) ! 457.395 us | }
2692 1) @ 119760.2 us | }
2694 [...]
2696 2) | handle_IPI() {
2697 1) 6.979 us | }
2698 2) 0.417 us | scheduler_ipi();
2699 1) 9.791 us | }
2700 1) + 12.917 us | }
2701 2) 3.490 us | }
2702 1) + 15.729 us | }
2703 1) + 18.542 us | }
2704 2) $ 3594274 us | }
2706 Flags::
2708 + means that the function exceeded 10 usecs.
2709 ! means that the function exceeded 100 usecs.
2710 # means that the function exceeded 1000 usecs.
2711 * means that the function exceeded 10 msecs.
2712 @ means that the function exceeded 100 msecs.
2713 $ means that the function exceeded 1 sec.
2716 - The task/pid field displays the thread cmdline and pid which
2717 executed the function. It is default disabled.
2719 - hide: echo nofuncgraph-proc > trace_options
2720 - show: echo funcgraph-proc > trace_options
2722 ie::
2724 # tracer: function_graph
2725 #
2726 # CPU TASK/PID DURATION FUNCTION CALLS
2727 # | | | | | | | | |
2728 0) sh-4802 | | d_free() {
2729 0) sh-4802 | | call_rcu() {
2730 0) sh-4802 | | __call_rcu() {
2731 0) sh-4802 | 0.616 us | rcu_process_gp_end();
2732 0) sh-4802 | 0.586 us | check_for_new_grace_period();
2733 0) sh-4802 | 2.899 us | }
2734 0) sh-4802 | 4.040 us | }
2735 0) sh-4802 | 5.151 us | }
2736 0) sh-4802 | + 49.370 us | }
2739 - The absolute time field is an absolute timestamp given by the
2740 system clock since it started. A snapshot of this time is
2741 given on each entry/exit of functions
2743 - hide: echo nofuncgraph-abstime > trace_options
2744 - show: echo funcgraph-abstime > trace_options
2746 ie::
2748 #
2749 # TIME CPU DURATION FUNCTION CALLS
2750 # | | | | | | | |
2751 360.774522 | 1) 0.541 us | }
2752 360.774522 | 1) 4.663 us | }
2753 360.774523 | 1) 0.541 us | __wake_up_bit();
2754 360.774524 | 1) 6.796 us | }
2755 360.774524 | 1) 7.952 us | }
2756 360.774525 | 1) 9.063 us | }
2757 360.774525 | 1) 0.615 us | journal_mark_dirty();
2758 360.774527 | 1) 0.578 us | __brelse();
2759 360.774528 | 1) | reiserfs_prepare_for_journal() {
2760 360.774528 | 1) | unlock_buffer() {
2761 360.774529 | 1) | wake_up_bit() {
2762 360.774529 | 1) | bit_waitqueue() {
2763 360.774530 | 1) 0.594 us | __phys_addr();
2766 The function name is always displayed after the closing bracket
2767 for a function if the start of that function is not in the
2768 trace buffer.
2770 Display of the function name after the closing bracket may be
2771 enabled for functions whose start is in the trace buffer,
2772 allowing easier searching with grep for function durations.
2773 It is default disabled.
2775 - hide: echo nofuncgraph-tail > trace_options
2776 - show: echo funcgraph-tail > trace_options
2778 Example with nofuncgraph-tail (default)::
2780 0) | putname() {
2781 0) | kmem_cache_free() {
2782 0) 0.518 us | __phys_addr();
2783 0) 1.757 us | }
2784 0) 2.861 us | }
2786 Example with funcgraph-tail::
2788 0) | putname() {
2789 0) | kmem_cache_free() {
2790 0) 0.518 us | __phys_addr();
2791 0) 1.757 us | } /* kmem_cache_free() */
2792 0) 2.861 us | } /* putname() */
2794 The return value of each traced function can be displayed after
2795 an equal sign "=". When encountering system call failures, it
2796 can be very helpful to quickly locate the function that first
2797 returns an error code.
2799 - hide: echo nofuncgraph-retval > trace_options
2800 - show: echo funcgraph-retval > trace_options
2802 Example with funcgraph-retval::
2804 1) | cgroup_migrate() {
2805 1) 0.651 us | cgroup_migrate_add_task(); /* = 0xffff93fcfd346c00 */
2806 1) | cgroup_migrate_execute() {
2807 1) | cpu_cgroup_can_attach() {
2808 1) | cgroup_taskset_first() {
2809 1) 0.732 us | cgroup_taskset_next(); /* = 0xffff93fc8fb20000 */
2810 1) 1.232 us | } /* cgroup_taskset_first = 0xffff93fc8fb20000 */
2811 1) 0.380 us | sched_rt_can_attach(); /* = 0x0 */
2812 1) 2.335 us | } /* cpu_cgroup_can_attach = -22 */
2813 1) 4.369 us | } /* cgroup_migrate_execute = -22 */
2814 1) 7.143 us | } /* cgroup_migrate = -22 */
2816 The above example shows that the function cpu_cgroup_can_attach
2817 returned the error code -22 firstly, then we can read the code
2818 of this function to get the root cause.
2820 When the option funcgraph-retval-hex is not set, the return value can
2821 be displayed in a smart way. Specifically, if it is an error code,
2822 it will be printed in signed decimal format, otherwise it will
2823 printed in hexadecimal format.
2825 - smart: echo nofuncgraph-retval-hex > trace_options
2826 - hexadecimal: echo funcgraph-retval-hex > trace_options
2828 Example with funcgraph-retval-hex::
2830 1) | cgroup_migrate() {
2831 1) 0.651 us | cgroup_migrate_add_task(); /* = 0xffff93fcfd346c00 */
2832 1) | cgroup_migrate_execute() {
2833 1) | cpu_cgroup_can_attach() {
2834 1) | cgroup_taskset_first() {
2835 1) 0.732 us | cgroup_taskset_next(); /* = 0xffff93fc8fb20000 */
2836 1) 1.232 us | } /* cgroup_taskset_first = 0xffff93fc8fb20000 */
2837 1) 0.380 us | sched_rt_can_attach(); /* = 0x0 */
2838 1) 2.335 us | } /* cpu_cgroup_can_attach = 0xffffffea */
2839 1) 4.369 us | } /* cgroup_migrate_execute = 0xffffffea */
2840 1) 7.143 us | } /* cgroup_migrate = 0xffffffea */
2842 At present, there are some limitations when using the funcgraph-retval
2843 option, and these limitations will be eliminated in the future:
2845 - Even if the function return type is void, a return value will still
2846 be printed, and you can just ignore it.
2848 - Even if return values are stored in multiple registers, only the
2849 value contained in the first register will be recorded and printed.
2850 To illustrate, in the x86 architecture, eax and edx are used to store
2851 a 64-bit return value, with the lower 32 bits saved in eax and the
2852 upper 32 bits saved in edx. However, only the value stored in eax
2853 will be recorded and printed.
2855 - In certain procedure call standards, such as arm64's AAPCS64, when a
2856 type is smaller than a GPR, it is the responsibility of the consumer
2857 to perform the narrowing, and the upper bits may contain UNKNOWN values.
2858 Therefore, it is advisable to check the code for such cases. For instance,
2859 when using a u8 in a 64-bit GPR, bits [63:8] may contain arbitrary values,
2860 especially when larger types are truncated, whether explicitly or implicitly.
2861 Here are some specific cases to illustrate this point:
2863 **Case One**:
2865 The function narrow_to_u8 is defined as follows::
2867 u8 narrow_to_u8(u64 val)
2868 {
2869 // implicitly truncated
2870 return val;
2871 }
2873 It may be compiled to::
2875 narrow_to_u8:
2876 < ... ftrace instrumentation ... >
2877 RET
2879 If you pass 0x123456789abcdef to this function and want to narrow it,
2880 it may be recorded as 0x123456789abcdef instead of 0xef.
2882 **Case Two**:
2884 The function error_if_not_4g_aligned is defined as follows::
2886 int error_if_not_4g_aligned(u64 val)
2887 {
2888 if (val & GENMASK(31, 0))
2889 return -EINVAL;
2891 return 0;
2892 }
2894 It could be compiled to::
2896 error_if_not_4g_aligned:
2897 CBNZ w0, .Lnot_aligned
2898 RET // bits [31:0] are zero, bits
2899 // [63:32] are UNKNOWN
2900 .Lnot_aligned:
2901 MOV x0, #-EINVAL
2902 RET
2904 When passing 0x2_0000_0000 to it, the return value may be recorded as
2905 0x2_0000_0000 instead of 0.
2907 You can put some comments on specific functions by using
2908 trace_printk() For example, if you want to put a comment inside
2909 the __might_sleep() function, you just have to include
2910 <linux/ftrace.h> and call trace_printk() inside __might_sleep()::
2912 trace_printk("I'm a comment!\n")
2914 will produce::
2916 1) | __might_sleep() {
2917 1) | /* I'm a comment! */
2918 1) 1.449 us | }
2921 You might find other useful features for this tracer in the
2922 following "dynamic ftrace" section such as tracing only specific
2923 functions or tasks.
2925 dynamic ftrace
2926 --------------
2928 If CONFIG_DYNAMIC_FTRACE is set, the system will run with
2929 virtually no overhead when function tracing is disabled. The way
2930 this works is the mcount function call (placed at the start of
2931 every kernel function, produced by the -pg switch in gcc),
2932 starts of pointing to a simple return. (Enabling FTRACE will
2933 include the -pg switch in the compiling of the kernel.)
2935 At compile time every C file object is run through the
2936 recordmcount program (located in the scripts directory). This
2937 program will parse the ELF headers in the C object to find all
2938 the locations in the .text section that call mcount. Starting
2939 with gcc version 4.6, the -mfentry has been added for x86, which
2940 calls "__fentry__" instead of "mcount". Which is called before
2941 the creation of the stack frame.
2943 Note, not all sections are traced. They may be prevented by either
2944 a notrace, or blocked another way and all inline functions are not
2945 traced. Check the "available_filter_functions" file to see what functions
2946 can be traced.
2948 A section called "__mcount_loc" is created that holds
2949 references to all the mcount/fentry call sites in the .text section.
2950 The recordmcount program re-links this section back into the
2951 original object. The final linking stage of the kernel will add all these
2952 references into a single table.
2954 On boot up, before SMP is initialized, the dynamic ftrace code
2955 scans this table and updates all the locations into nops. It
2956 also records the locations, which are added to the
2957 available_filter_functions list. Modules are processed as they
2958 are loaded and before they are executed. When a module is
2959 unloaded, it also removes its functions from the ftrace function
2960 list. This is automatic in the module unload code, and the
2961 module author does not need to worry about it.
2963 When tracing is enabled, the process of modifying the function
2964 tracepoints is dependent on architecture. The old method is to use
2965 kstop_machine to prevent races with the CPUs executing code being
2966 modified (which can cause the CPU to do undesirable things, especially
2967 if the modified code crosses cache (or page) boundaries), and the nops are
2968 patched back to calls. But this time, they do not call mcount
2969 (which is just a function stub). They now call into the ftrace
2970 infrastructure.
2972 The new method of modifying the function tracepoints is to place
2973 a breakpoint at the location to be modified, sync all CPUs, modify
2974 the rest of the instruction not covered by the breakpoint. Sync
2975 all CPUs again, and then remove the breakpoint with the finished
2976 version to the ftrace call site.
2978 Some archs do not even need to monkey around with the synchronization,
2979 and can just slap the new code on top of the old without any
2980 problems with other CPUs executing it at the same time.
2982 One special side-effect to the recording of the functions being
2983 traced is that we can now selectively choose which functions we
2984 wish to trace and which ones we want the mcount calls to remain
2985 as nops.
2987 Two files are used, one for enabling and one for disabling the
2988 tracing of specified functions. They are:
2990 set_ftrace_filter
2992 and
2994 set_ftrace_notrace
2996 A list of available functions that you can add to these files is
2997 listed in:
2999 available_filter_functions
3001 ::
3003 # cat available_filter_functions
3004 put_prev_task_idle
3005 kmem_cache_create
3006 pick_next_task_rt
3007 cpus_read_lock
3008 pick_next_task_fair
3009 mutex_lock
3010 [...]
3012 If I am only interested in sys_nanosleep and hrtimer_interrupt::
3014 # echo sys_nanosleep hrtimer_interrupt > set_ftrace_filter
3015 # echo function > current_tracer
3016 # echo 1 > tracing_on
3017 # usleep 1
3018 # echo 0 > tracing_on
3019 # cat trace
3020 # tracer: function
3021 #
3022 # entries-in-buffer/entries-written: 5/5 #P:4
3023 #
3024 # _-----=> irqs-off
3025 # / _----=> need-resched
3026 # | / _---=> hardirq/softirq
3027 # || / _--=> preempt-depth
3028 # ||| / delay
3029 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3030 # | | | |||| | |
3031 usleep-2665 [001] .... 4186.475355: sys_nanosleep <-system_call_fastpath
3032 <idle>-0 [001] d.h1 4186.475409: hrtimer_interrupt <-smp_apic_timer_interrupt
3033 usleep-2665 [001] d.h1 4186.475426: hrtimer_interrupt <-smp_apic_timer_interrupt
3034 <idle>-0 [003] d.h1 4186.475426: hrtimer_interrupt <-smp_apic_timer_interrupt
3035 <idle>-0 [002] d.h1 4186.475427: hrtimer_interrupt <-smp_apic_timer_interrupt
3037 To see which functions are being traced, you can cat the file:
3038 ::
3040 # cat set_ftrace_filter
3041 hrtimer_interrupt
3042 sys_nanosleep
3045 Perhaps this is not enough. The filters also allow glob(7) matching.
3047 ``<match>*``
3048 will match functions that begin with <match>
3049 ``*<match>``
3050 will match functions that end with <match>
3051 ``*<match>*``
3052 will match functions that have <match> in it
3053 ``<match1>*<match2>``
3054 will match functions that begin with <match1> and end with <match2>
3056 .. note::
3057 It is better to use quotes to enclose the wild cards,
3058 otherwise the shell may expand the parameters into names
3059 of files in the local directory.
3061 ::
3063 # echo 'hrtimer_*' > set_ftrace_filter
3065 Produces::
3067 # tracer: function
3068 #
3069 # entries-in-buffer/entries-written: 897/897 #P:4
3070 #
3071 # _-----=> irqs-off
3072 # / _----=> need-resched
3073 # | / _---=> hardirq/softirq
3074 # || / _--=> preempt-depth
3075 # ||| / delay
3076 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3077 # | | | |||| | |
3078 <idle>-0 [003] dN.1 4228.547803: hrtimer_cancel <-tick_nohz_idle_exit
3079 <idle>-0 [003] dN.1 4228.547804: hrtimer_try_to_cancel <-hrtimer_cancel
3080 <idle>-0 [003] dN.2 4228.547805: hrtimer_force_reprogram <-__remove_hrtimer
3081 <idle>-0 [003] dN.1 4228.547805: hrtimer_forward <-tick_nohz_idle_exit
3082 <idle>-0 [003] dN.1 4228.547805: hrtimer_start_range_ns <-hrtimer_start_expires.constprop.11
3083 <idle>-0 [003] d..1 4228.547858: hrtimer_get_next_event <-get_next_timer_interrupt
3084 <idle>-0 [003] d..1 4228.547859: hrtimer_start <-__tick_nohz_idle_enter
3085 <idle>-0 [003] d..2 4228.547860: hrtimer_force_reprogram <-__rem
3087 Notice that we lost the sys_nanosleep.
3088 ::
3090 # cat set_ftrace_filter
3091 hrtimer_run_queues
3092 hrtimer_run_pending
3093 hrtimer_setup
3094 hrtimer_cancel
3095 hrtimer_try_to_cancel
3096 hrtimer_forward
3097 hrtimer_start
3098 hrtimer_reprogram
3099 hrtimer_force_reprogram
3100 hrtimer_get_next_event
3101 hrtimer_interrupt
3102 hrtimer_nanosleep
3103 hrtimer_wakeup
3104 hrtimer_get_remaining
3105 hrtimer_get_res
3106 hrtimer_init_sleeper
3109 This is because the '>' and '>>' act just like they do in bash.
3110 To rewrite the filters, use '>'
3111 To append to the filters, use '>>'
3113 To clear out a filter so that all functions will be recorded
3114 again::
3116 # echo > set_ftrace_filter
3117 # cat set_ftrace_filter
3118 #
3120 Again, now we want to append.
3122 ::
3124 # echo sys_nanosleep > set_ftrace_filter
3125 # cat set_ftrace_filter
3126 sys_nanosleep
3127 # echo 'hrtimer_*' >> set_ftrace_filter
3128 # cat set_ftrace_filter
3129 hrtimer_run_queues
3130 hrtimer_run_pending
3131 hrtimer_setup
3132 hrtimer_cancel
3133 hrtimer_try_to_cancel
3134 hrtimer_forward
3135 hrtimer_start
3136 hrtimer_reprogram
3137 hrtimer_force_reprogram
3138 hrtimer_get_next_event
3139 hrtimer_interrupt
3140 sys_nanosleep
3141 hrtimer_nanosleep
3142 hrtimer_wakeup
3143 hrtimer_get_remaining
3144 hrtimer_get_res
3145 hrtimer_init_sleeper
3148 The set_ftrace_notrace prevents those functions from being
3149 traced.
3150 ::
3152 # echo '*preempt*' '*lock*' > set_ftrace_notrace
3154 Produces::
3156 # tracer: function
3157 #
3158 # entries-in-buffer/entries-written: 39608/39608 #P:4
3159 #
3160 # _-----=> irqs-off
3161 # / _----=> need-resched
3162 # | / _---=> hardirq/softirq
3163 # || / _--=> preempt-depth
3164 # ||| / delay
3165 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3166 # | | | |||| | |
3167 bash-1994 [000] .... 4342.324896: file_ra_state_init <-do_dentry_open
3168 bash-1994 [000] .... 4342.324897: open_check_o_direct <-do_last
3169 bash-1994 [000] .... 4342.324897: ima_file_check <-do_last
3170 bash-1994 [000] .... 4342.324898: process_measurement <-ima_file_check
3171 bash-1994 [000] .... 4342.324898: ima_get_action <-process_measurement
3172 bash-1994 [000] .... 4342.324898: ima_match_policy <-ima_get_action
3173 bash-1994 [000] .... 4342.324899: do_truncate <-do_last
3174 bash-1994 [000] .... 4342.324899: setattr_should_drop_suidgid <-do_truncate
3175 bash-1994 [000] .... 4342.324899: notify_change <-do_truncate
3176 bash-1994 [000] .... 4342.324900: current_fs_time <-notify_change
3177 bash-1994 [000] .... 4342.324900: current_kernel_time <-current_fs_time
3178 bash-1994 [000] .... 4342.324900: timespec_trunc <-current_fs_time
3180 We can see that there's no more lock or preempt tracing.
3182 Selecting function filters via index
3183 ------------------------------------
3185 Because processing of strings is expensive (the address of the function
3186 needs to be looked up before comparing to the string being passed in),
3187 an index can be used as well to enable functions. This is useful in the
3188 case of setting thousands of specific functions at a time. By passing
3189 in a list of numbers, no string processing will occur. Instead, the function
3190 at the specific location in the internal array (which corresponds to the
3191 functions in the "available_filter_functions" file), is selected.
3193 ::
3195 # echo 1 > set_ftrace_filter
3197 Will select the first function listed in "available_filter_functions"
3199 ::
3201 # head -1 available_filter_functions
3202 trace_initcall_finish_cb
3204 # cat set_ftrace_filter
3205 trace_initcall_finish_cb
3207 # head -50 available_filter_functions | tail -1
3208 x86_pmu_commit_txn
3210 # echo 1 50 > set_ftrace_filter
3211 # cat set_ftrace_filter
3212 trace_initcall_finish_cb
3213 x86_pmu_commit_txn
3215 Dynamic ftrace with the function graph tracer
3216 ---------------------------------------------
3218 Although what has been explained above concerns both the
3219 function tracer and the function-graph-tracer, there are some
3220 special features only available in the function-graph tracer.
3222 If you want to trace only one function and all of its children,
3223 you just have to echo its name into set_graph_function::
3225 echo __do_fault > set_graph_function
3227 will produce the following "expanded" trace of the __do_fault()
3228 function::
3230 0) | __do_fault() {
3231 0) | filemap_fault() {
3232 0) | find_lock_page() {
3233 0) 0.804 us | find_get_page();
3234 0) | __might_sleep() {
3235 0) 1.329 us | }
3236 0) 3.904 us | }
3237 0) 4.979 us | }
3238 0) 0.653 us | _spin_lock();
3239 0) 0.578 us | page_add_file_rmap();
3240 0) 0.525 us | native_set_pte_at();
3241 0) 0.585 us | _spin_unlock();
3242 0) | unlock_page() {
3243 0) 0.541 us | page_waitqueue();
3244 0) 0.639 us | __wake_up_bit();
3245 0) 2.786 us | }
3246 0) + 14.237 us | }
3247 0) | __do_fault() {
3248 0) | filemap_fault() {
3249 0) | find_lock_page() {
3250 0) 0.698 us | find_get_page();
3251 0) | __might_sleep() {
3252 0) 1.412 us | }
3253 0) 3.950 us | }
3254 0) 5.098 us | }
3255 0) 0.631 us | _spin_lock();
3256 0) 0.571 us | page_add_file_rmap();
3257 0) 0.526 us | native_set_pte_at();
3258 0) 0.586 us | _spin_unlock();
3259 0) | unlock_page() {
3260 0) 0.533 us | page_waitqueue();
3261 0) 0.638 us | __wake_up_bit();
3262 0) 2.793 us | }
3263 0) + 14.012 us | }
3265 You can also expand several functions at once::
3267 echo sys_open > set_graph_function
3268 echo sys_close >> set_graph_function
3270 Now if you want to go back to trace all functions you can clear
3271 this special filter via::
3273 echo > set_graph_function
3276 ftrace_enabled
3277 --------------
3279 Note, the proc sysctl ftrace_enable is a big on/off switch for the
3280 function tracer. By default it is enabled (when function tracing is
3281 enabled in the kernel). If it is disabled, all function tracing is
3282 disabled. This includes not only the function tracers for ftrace, but
3283 also for any other uses (perf, kprobes, stack tracing, profiling, etc). It
3284 cannot be disabled if there is a callback with FTRACE_OPS_FL_PERMANENT set
3285 registered.
3287 Please disable this with care.
3289 This can be disable (and enabled) with::
3291 sysctl kernel.ftrace_enabled=0
3292 sysctl kernel.ftrace_enabled=1
3294 or
3296 echo 0 > /proc/sys/kernel/ftrace_enabled
3297 echo 1 > /proc/sys/kernel/ftrace_enabled
3300 Filter commands
3301 ---------------
3303 A few commands are supported by the set_ftrace_filter interface.
3304 Trace commands have the following format::
3306 <function>:<command>:<parameter>
3308 The following commands are supported:
3310 - mod:
3311 This command enables function filtering per module. The
3312 parameter defines the module. For example, if only the write*
3313 functions in the ext3 module are desired, run:
3315 echo 'write*:mod:ext3' > set_ftrace_filter
3317 This command interacts with the filter in the same way as
3318 filtering based on function names. Thus, adding more functions
3319 in a different module is accomplished by appending (>>) to the
3320 filter file. Remove specific module functions by prepending
3321 '!'::
3323 echo '!writeback*:mod:ext3' >> set_ftrace_filter
3325 Mod command supports module globbing. Disable tracing for all
3326 functions except a specific module::
3328 echo '!*:mod:!ext3' >> set_ftrace_filter
3330 Disable tracing for all modules, but still trace kernel::
3332 echo '!*:mod:*' >> set_ftrace_filter
3334 Enable filter only for kernel::
3336 echo '*write*:mod:!*' >> set_ftrace_filter
3338 Enable filter for module globbing::
3340 echo '*write*:mod:*snd*' >> set_ftrace_filter
3342 - traceon/traceoff:
3343 These commands turn tracing on and off when the specified
3344 functions are hit. The parameter determines how many times the
3345 tracing system is turned on and off. If unspecified, there is
3346 no limit. For example, to disable tracing when a schedule bug
3347 is hit the first 5 times, run::
3349 echo '__schedule_bug:traceoff:5' > set_ftrace_filter
3351 To always disable tracing when __schedule_bug is hit::
3353 echo '__schedule_bug:traceoff' > set_ftrace_filter
3355 These commands are cumulative whether or not they are appended
3356 to set_ftrace_filter. To remove a command, prepend it by '!'
3357 and drop the parameter::
3359 echo '!__schedule_bug:traceoff:0' > set_ftrace_filter
3361 The above removes the traceoff command for __schedule_bug
3362 that have a counter. To remove commands without counters::
3364 echo '!__schedule_bug:traceoff' > set_ftrace_filter
3366 - snapshot:
3367 Will cause a snapshot to be triggered when the function is hit.
3368 ::
3370 echo 'native_flush_tlb_others:snapshot' > set_ftrace_filter
3372 To only snapshot once:
3373 ::
3375 echo 'native_flush_tlb_others:snapshot:1' > set_ftrace_filter
3377 To remove the above commands::
3379 echo '!native_flush_tlb_others:snapshot' > set_ftrace_filter
3380 echo '!native_flush_tlb_others:snapshot:0' > set_ftrace_filter
3382 - enable_event/disable_event:
3383 These commands can enable or disable a trace event. Note, because
3384 function tracing callbacks are very sensitive, when these commands
3385 are registered, the trace point is activated, but disabled in
3386 a "soft" mode. That is, the tracepoint will be called, but
3387 just will not be traced. The event tracepoint stays in this mode
3388 as long as there's a command that triggers it.
3389 ::
3391 echo 'try_to_wake_up:enable_event:sched:sched_switch:2' > \
3392 set_ftrace_filter
3394 The format is::
3396 <function>:enable_event:<system>:<event>[:count]
3397 <function>:disable_event:<system>:<event>[:count]
3399 To remove the events commands::
3401 echo '!try_to_wake_up:enable_event:sched:sched_switch:0' > \
3402 set_ftrace_filter
3403 echo '!schedule:disable_event:sched:sched_switch' > \
3404 set_ftrace_filter
3406 - dump:
3407 When the function is hit, it will dump the contents of the ftrace
3408 ring buffer to the console. This is useful if you need to debug
3409 something, and want to dump the trace when a certain function
3410 is hit. Perhaps it's a function that is called before a triple
3411 fault happens and does not allow you to get a regular dump.
3413 - cpudump:
3414 When the function is hit, it will dump the contents of the ftrace
3415 ring buffer for the current CPU to the console. Unlike the "dump"
3416 command, it only prints out the contents of the ring buffer for the
3417 CPU that executed the function that triggered the dump.
3419 - stacktrace:
3420 When the function is hit, a stack trace is recorded.
3422 trace_pipe
3423 ----------
3425 The trace_pipe outputs the same content as the trace file, but
3426 the effect on the tracing is different. Every read from
3427 trace_pipe is consumed. This means that subsequent reads will be
3428 different. The trace is live.
3429 ::
3431 # echo function > current_tracer
3432 # cat trace_pipe > /tmp/trace.out &
3433 [1] 4153
3434 # echo 1 > tracing_on
3435 # usleep 1
3436 # echo 0 > tracing_on
3437 # cat trace
3438 # tracer: function
3439 #
3440 # entries-in-buffer/entries-written: 0/0 #P:4
3441 #
3442 # _-----=> irqs-off
3443 # / _----=> need-resched
3444 # | / _---=> hardirq/softirq
3445 # || / _--=> preempt-depth
3446 # ||| / delay
3447 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3448 # | | | |||| | |
3450 #
3451 # cat /tmp/trace.out
3452 bash-1994 [000] .... 5281.568961: mutex_unlock <-rb_simple_write
3453 bash-1994 [000] .... 5281.568963: __mutex_unlock_slowpath <-mutex_unlock
3454 bash-1994 [000] .... 5281.568963: __fsnotify_parent <-fsnotify_modify
3455 bash-1994 [000] .... 5281.568964: fsnotify <-fsnotify_modify
3456 bash-1994 [000] .... 5281.568964: __srcu_read_lock <-fsnotify
3457 bash-1994 [000] .... 5281.568964: add_preempt_count <-__srcu_read_lock
3458 bash-1994 [000] ...1 5281.568965: sub_preempt_count <-__srcu_read_lock
3459 bash-1994 [000] .... 5281.568965: __srcu_read_unlock <-fsnotify
3460 bash-1994 [000] .... 5281.568967: sys_dup2 <-system_call_fastpath
3463 Note, reading the trace_pipe file will block until more input is
3464 added. This is contrary to the trace file. If any process opened
3465 the trace file for reading, it will actually disable tracing and
3466 prevent new entries from being added. The trace_pipe file does
3467 not have this limitation.
3469 trace entries
3470 -------------
3472 Having too much or not enough data can be troublesome in
3473 diagnosing an issue in the kernel. The file buffer_size_kb is
3474 used to modify the size of the internal trace buffers. The
3475 number listed is the number of entries that can be recorded per
3476 CPU. To know the full size, multiply the number of possible CPUs
3477 with the number of entries.
3478 ::
3480 # cat buffer_size_kb
3481 1408 (units kilobytes)
3483 Or simply read buffer_total_size_kb
3484 ::
3486 # cat buffer_total_size_kb
3487 5632
3489 To modify the buffer, simple echo in a number (in 1024 byte segments).
3490 ::
3492 # echo 10000 > buffer_size_kb
3493 # cat buffer_size_kb
3494 10000 (units kilobytes)
3496 It will try to allocate as much as possible. If you allocate too
3497 much, it can cause Out-Of-Memory to trigger.
3498 ::
3500 # echo 1000000000000 > buffer_size_kb
3501 -bash: echo: write error: Cannot allocate memory
3502 # cat buffer_size_kb
3503 85
3505 The per_cpu buffers can be changed individually as well:
3506 ::
3508 # echo 10000 > per_cpu/cpu0/buffer_size_kb
3509 # echo 100 > per_cpu/cpu1/buffer_size_kb
3511 When the per_cpu buffers are not the same, the buffer_size_kb
3512 at the top level will just show an X
3513 ::
3515 # cat buffer_size_kb
3516 X
3518 This is where the buffer_total_size_kb is useful:
3519 ::
3521 # cat buffer_total_size_kb
3522 12916
3524 Writing to the top level buffer_size_kb will reset all the buffers
3525 to be the same again.
3527 Snapshot
3528 --------
3529 CONFIG_TRACER_SNAPSHOT makes a generic snapshot feature
3530 available to all non latency tracers. (Latency tracers which
3531 record max latency, such as "irqsoff" or "wakeup", can't use
3532 this feature, since those are already using the snapshot
3533 mechanism internally.)
3535 Snapshot preserves a current trace buffer at a particular point
3536 in time without stopping tracing. Ftrace swaps the current
3537 buffer with a spare buffer, and tracing continues in the new
3538 current (=previous spare) buffer.
3540 The following tracefs files in "tracing" are related to this
3541 feature:
3543 snapshot:
3545 This is used to take a snapshot and to read the output
3546 of the snapshot. Echo 1 into this file to allocate a
3547 spare buffer and to take a snapshot (swap), then read
3548 the snapshot from this file in the same format as
3549 "trace" (described above in the section "The File
3550 System"). Both reads snapshot and tracing are executable
3551 in parallel. When the spare buffer is allocated, echoing
3552 0 frees it, and echoing else (positive) values clear the
3553 snapshot contents.
3554 More details are shown in the table below.
3556 +--------------+------------+------------+------------+
3557 |status\\input | 0 | 1 | else |
3558 +==============+============+============+============+
3559 |not allocated |(do nothing)| alloc+swap |(do nothing)|
3560 +--------------+------------+------------+------------+
3561 |allocated | free | swap | clear |
3562 +--------------+------------+------------+------------+
3564 Here is an example of using the snapshot feature.
3565 ::
3567 # echo 1 > events/sched/enable
3568 # echo 1 > snapshot
3569 # cat snapshot
3570 # tracer: nop
3571 #
3572 # entries-in-buffer/entries-written: 71/71 #P:8
3573 #
3574 # _-----=> irqs-off
3575 # / _----=> need-resched
3576 # | / _---=> hardirq/softirq
3577 # || / _--=> preempt-depth
3578 # ||| / delay
3579 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3580 # | | | |||| | |
3581 <idle>-0 [005] d... 2440.603828: sched_switch: prev_comm=swapper/5 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2242 next_prio=120
3582 sleep-2242 [005] d... 2440.603846: sched_switch: prev_comm=snapshot-test-2 prev_pid=2242 prev_prio=120 prev_state=R ==> next_comm=kworker/5:1 next_pid=60 next_prio=120
3583 [...]
3584 <idle>-0 [002] d... 2440.707230: sched_switch: prev_comm=swapper/2 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2229 next_prio=120
3586 # cat trace
3587 # tracer: nop
3588 #
3589 # entries-in-buffer/entries-written: 77/77 #P:8
3590 #
3591 # _-----=> irqs-off
3592 # / _----=> need-resched
3593 # | / _---=> hardirq/softirq
3594 # || / _--=> preempt-depth
3595 # ||| / delay
3596 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3597 # | | | |||| | |
3598 <idle>-0 [007] d... 2440.707395: sched_switch: prev_comm=swapper/7 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2243 next_prio=120
3599 snapshot-test-2-2229 [002] d... 2440.707438: sched_switch: prev_comm=snapshot-test-2 prev_pid=2229 prev_prio=120 prev_state=S ==> next_comm=swapper/2 next_pid=0 next_prio=120
3600 [...]
3603 If you try to use this snapshot feature when current tracer is
3604 one of the latency tracers, you will get the following results.
3605 ::
3607 # echo wakeup > current_tracer
3608 # echo 1 > snapshot
3609 bash: echo: write error: Device or resource busy
3610 # cat snapshot
3611 cat: snapshot: Device or resource busy
3614 Instances
3615 ---------
3616 In the tracefs tracing directory, there is a directory called "instances".
3617 This directory can have new directories created inside of it using
3618 mkdir, and removing directories with rmdir. The directory created
3619 with mkdir in this directory will already contain files and other
3620 directories after it is created.
3621 ::
3623 # mkdir instances/foo
3624 # ls instances/foo
3625 buffer_size_kb buffer_total_size_kb events free_buffer per_cpu
3626 set_event snapshot trace trace_clock trace_marker trace_options
3627 trace_pipe tracing_on
3629 As you can see, the new directory looks similar to the tracing directory
3630 itself. In fact, it is very similar, except that the buffer and
3631 events are agnostic from the main directory, or from any other
3632 instances that are created.
3634 The files in the new directory work just like the files with the
3635 same name in the tracing directory except the buffer that is used
3636 is a separate and new buffer. The files affect that buffer but do not
3637 affect the main buffer with the exception of trace_options. Currently,
3638 the trace_options affect all instances and the top level buffer
3639 the same, but this may change in future releases. That is, options
3640 may become specific to the instance they reside in.
3642 Notice that none of the function tracer files are there, nor is
3643 current_tracer and available_tracers. This is because the buffers
3644 can currently only have events enabled for them.
3645 ::
3647 # mkdir instances/foo
3648 # mkdir instances/bar
3649 # mkdir instances/zoot
3650 # echo 100000 > buffer_size_kb
3651 # echo 1000 > instances/foo/buffer_size_kb
3652 # echo 5000 > instances/bar/per_cpu/cpu1/buffer_size_kb
3653 # echo function > current_trace
3654 # echo 1 > instances/foo/events/sched/sched_wakeup/enable
3655 # echo 1 > instances/foo/events/sched/sched_wakeup_new/enable
3656 # echo 1 > instances/foo/events/sched/sched_switch/enable
3657 # echo 1 > instances/bar/events/irq/enable
3658 # echo 1 > instances/zoot/events/syscalls/enable
3659 # cat trace_pipe
3660 CPU:2 [LOST 11745 EVENTS]
3661 bash-2044 [002] .... 10594.481032: _raw_spin_lock_irqsave <-get_page_from_freelist
3662 bash-2044 [002] d... 10594.481032: add_preempt_count <-_raw_spin_lock_irqsave
3663 bash-2044 [002] d..1 10594.481032: __rmqueue <-get_page_from_freelist
3664 bash-2044 [002] d..1 10594.481033: _raw_spin_unlock <-get_page_from_freelist
3665 bash-2044 [002] d..1 10594.481033: sub_preempt_count <-_raw_spin_unlock
3666 bash-2044 [002] d... 10594.481033: get_pageblock_flags_group <-get_pageblock_migratetype
3667 bash-2044 [002] d... 10594.481034: __mod_zone_page_state <-get_page_from_freelist
3668 bash-2044 [002] d... 10594.481034: zone_statistics <-get_page_from_freelist
3669 bash-2044 [002] d... 10594.481034: __inc_zone_state <-zone_statistics
3670 bash-2044 [002] d... 10594.481034: __inc_zone_state <-zone_statistics
3671 bash-2044 [002] .... 10594.481035: arch_dup_task_struct <-copy_process
3672 [...]
3674 # cat instances/foo/trace_pipe
3675 bash-1998 [000] d..4 136.676759: sched_wakeup: comm=kworker/0:1 pid=59 prio=120 success=1 target_cpu=000
3676 bash-1998 [000] dN.4 136.676760: sched_wakeup: comm=bash pid=1998 prio=120 success=1 target_cpu=000
3677 <idle>-0 [003] d.h3 136.676906: sched_wakeup: comm=rcu_preempt pid=9 prio=120 success=1 target_cpu=003
3678 <idle>-0 [003] d..3 136.676909: sched_switch: prev_comm=swapper/3 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=rcu_preempt next_pid=9 next_prio=120
3679 rcu_preempt-9 [003] d..3 136.676916: sched_switch: prev_comm=rcu_preempt prev_pid=9 prev_prio=120 prev_state=S ==> next_comm=swapper/3 next_pid=0 next_prio=120
3680 bash-1998 [000] d..4 136.677014: sched_wakeup: comm=kworker/0:1 pid=59 prio=120 success=1 target_cpu=000
3681 bash-1998 [000] dN.4 136.677016: sched_wakeup: comm=bash pid=1998 prio=120 success=1 target_cpu=000
3682 bash-1998 [000] d..3 136.677018: sched_switch: prev_comm=bash prev_pid=1998 prev_prio=120 prev_state=R+ ==> next_comm=kworker/0:1 next_pid=59 next_prio=120
3683 kworker/0:1-59 [000] d..4 136.677022: sched_wakeup: comm=sshd pid=1995 prio=120 success=1 target_cpu=001
3684 kworker/0:1-59 [000] d..3 136.677025: sched_switch: prev_comm=kworker/0:1 prev_pid=59 prev_prio=120 prev_state=S ==> next_comm=bash next_pid=1998 next_prio=120
3685 [...]
3687 # cat instances/bar/trace_pipe
3688 migration/1-14 [001] d.h3 138.732674: softirq_raise: vec=3 [action=NET_RX]
3689 <idle>-0 [001] dNh3 138.732725: softirq_raise: vec=3 [action=NET_RX]
3690 bash-1998 [000] d.h1 138.733101: softirq_raise: vec=1 [action=TIMER]
3691 bash-1998 [000] d.h1 138.733102: softirq_raise: vec=9 [action=RCU]
3692 bash-1998 [000] ..s2 138.733105: softirq_entry: vec=1 [action=TIMER]
3693 bash-1998 [000] ..s2 138.733106: softirq_exit: vec=1 [action=TIMER]
3694 bash-1998 [000] ..s2 138.733106: softirq_entry: vec=9 [action=RCU]
3695 bash-1998 [000] ..s2 138.733109: softirq_exit: vec=9 [action=RCU]
3696 sshd-1995 [001] d.h1 138.733278: irq_handler_entry: irq=21 name=uhci_hcd:usb4
3697 sshd-1995 [001] d.h1 138.733280: irq_handler_exit: irq=21 ret=unhandled
3698 sshd-1995 [001] d.h1 138.733281: irq_handler_entry: irq=21 name=eth0
3699 sshd-1995 [001] d.h1 138.733283: irq_handler_exit: irq=21 ret=handled
3700 [...]
3702 # cat instances/zoot/trace
3703 # tracer: nop
3704 #
3705 # entries-in-buffer/entries-written: 18996/18996 #P:4
3706 #
3707 # _-----=> irqs-off
3708 # / _----=> need-resched
3709 # | / _---=> hardirq/softirq
3710 # || / _--=> preempt-depth
3711 # ||| / delay
3712 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
3713 # | | | |||| | |
3714 bash-1998 [000] d... 140.733501: sys_write -> 0x2
3715 bash-1998 [000] d... 140.733504: sys_dup2(oldfd: a, newfd: 1)
3716 bash-1998 [000] d... 140.733506: sys_dup2 -> 0x1
3717 bash-1998 [000] d... 140.733508: sys_fcntl(fd: a, cmd: 1, arg: 0)
3718 bash-1998 [000] d... 140.733509: sys_fcntl -> 0x1
3719 bash-1998 [000] d... 140.733510: sys_close(fd: a)
3720 bash-1998 [000] d... 140.733510: sys_close -> 0x0
3721 bash-1998 [000] d... 140.733514: sys_rt_sigprocmask(how: 0, nset: 0, oset: 6e2768, sigsetsize: 8)
3722 bash-1998 [000] d... 140.733515: sys_rt_sigprocmask -> 0x0
3723 bash-1998 [000] d... 140.733516: sys_rt_sigaction(sig: 2, act: 7fff718846f0, oact: 7fff71884650, sigsetsize: 8)
3724 bash-1998 [000] d... 140.733516: sys_rt_sigaction -> 0x0
3726 You can see that the trace of the top most trace buffer shows only
3727 the function tracing. The foo instance displays wakeups and task
3728 switches.
3730 To remove the instances, simply delete their directories:
3731 ::
3733 # rmdir instances/foo
3734 # rmdir instances/bar
3735 # rmdir instances/zoot
3737 Note, if a process has a trace file open in one of the instance
3738 directories, the rmdir will fail with EBUSY.
3741 Stack trace
3742 -----------
3743 Since the kernel has a fixed sized stack, it is important not to
3744 waste it in functions. A kernel developer must be conscious of
3745 what they allocate on the stack. If they add too much, the system
3746 can be in danger of a stack overflow, and corruption will occur,
3747 usually leading to a system panic.
3749 There are some tools that check this, usually with interrupts
3750 periodically checking usage. But if you can perform a check
3751 at every function call that will become very useful. As ftrace provides
3752 a function tracer, it makes it convenient to check the stack size
3753 at every function call. This is enabled via the stack tracer.
3755 CONFIG_STACK_TRACER enables the ftrace stack tracing functionality.
3756 To enable it, write a '1' into /proc/sys/kernel/stack_tracer_enabled.
3757 ::
3759 # echo 1 > /proc/sys/kernel/stack_tracer_enabled
3761 You can also enable it from the kernel command line to trace
3762 the stack size of the kernel during boot up, by adding "stacktrace"
3763 to the kernel command line parameter.
3765 After running it for a few minutes, the output looks like:
3766 ::
3768 # cat stack_max_size
3769 2928
3771 # cat stack_trace
3772 Depth Size Location (18 entries)
3773 ----- ---- --------
3774 0) 2928 224 update_sd_lb_stats+0xbc/0x4ac
3775 1) 2704 160 find_busiest_group+0x31/0x1f1
3776 2) 2544 256 load_balance+0xd9/0x662
3777 3) 2288 80 idle_balance+0xbb/0x130
3778 4) 2208 128 __schedule+0x26e/0x5b9
3779 5) 2080 16 schedule+0x64/0x66
3780 6) 2064 128 schedule_timeout+0x34/0xe0
3781 7) 1936 112 wait_for_common+0x97/0xf1
3782 8) 1824 16 wait_for_completion+0x1d/0x1f
3783 9) 1808 128 flush_work+0xfe/0x119
3784 10) 1680 16 tty_flush_to_ldisc+0x1e/0x20
3785 11) 1664 48 input_available_p+0x1d/0x5c
3786 12) 1616 48 n_tty_poll+0x6d/0x134
3787 13) 1568 64 tty_poll+0x64/0x7f
3788 14) 1504 880 do_select+0x31e/0x511
3789 15) 624 400 core_sys_select+0x177/0x216
3790 16) 224 96 sys_select+0x91/0xb9
3791 17) 128 128 system_call_fastpath+0x16/0x1b
3793 Note, if -mfentry is being used by gcc, functions get traced before
3794 they set up the stack frame. This means that leaf level functions
3795 are not tested by the stack tracer when -mfentry is used.
3797 Currently, -mfentry is used by gcc 4.6.0 and above on x86 only.
3799 More
3800 ----
3801 More details can be found in the source code, in the `kernel/trace/*.c` files.

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

문서 정보와 갱신 이력

1-18

`ftrace - Function Tracer` 문서는 Steven Rostedt가 작성한 커널 함수 추적기 사용 안내서다. 2008년 Red Hat 저작물이며 GNU Free Documentation License 1.2와 GPL v2로 이중 라이선스된다.

초기 검토자는 Elias Oltmanns, Randy Dunlap, Andrew Morton, John Kacur, David Teigland다. 문서는 Linux 2.6.28-rc2용으로 작성된 뒤 3.10과 4.13에 맞춰 갱신됐고, Changbin Du가 reStructuredText 형식으로 변환했다.

문서 이력
항목내용
최초 기준Linux 2.6.28-rc2
갱신Linux 3.10
갱신Linux 4.13, VMware Inc.와 Steven Rostedt
형식 변환Changbin Du, reStructuredText

원문에 기록된 기준 버전과 기여를 보존한다.

========================
ftrace - Function Tracer
========================

Copyright 2008 Red Hat Inc.

:Author:   Steven Rostedt <[email protected]>
:License:  The GNU Free Documentation License, Version 1.2
          (dual licensed under the GPL v2)
:Original Reviewers:  Elias Oltmanns, Randy Dunlap, Andrew Morton,
		      John Kacur, and David Teigland.

- Written for: 2.6.28-rc2
- Updated for: 3.10
- Updated for: 4.13 - Copyright 2017 VMware Inc. Steven Rostedt
- Converted to rst format - Changbin Du <[email protected]>

Introduction

소개

19-40

Ftrace는 개발자와 시스템 설계자가 커널 내부에서 일어나는 일을 파악하도록 만든 내부 추적기다. 사용자 공간 밖에서 발생하는 버그, 지연 시간, 성능 문제를 디버깅하고 분석하는 데 사용할 수 있다.

이름 때문에 흔히 함수 추적기 하나로 생각하지만, 실제로는 여러 추적 도구를 묶은 프레임워크다. 인터럽트 비활성화 구간, preemption 비활성화 구간, 태스크가 깨어난 뒤 실제로 스케줄될 때까지의 지연을 측정하는 추적기도 포함한다.

가장 널리 쓰이는 기능 가운데 하나는 이벤트 추적이다. 커널 전역의 수백 개 정적 이벤트 지점을 `tracefs`에서 선택적으로 활성화해 특정 하위 시스템의 동작을 관찰할 수 있다. 자세한 내용은 `events.rst`를 참고한다.

Ftrace 활용 범위
Ftrace함수 호출 흐름
FtraceIRQ/preemption 지연
Ftracewakeup-to-schedule 지연
Ftrace정적 이벤트 추적

하나의 프레임워크가 함수 흐름과 여러 지연 및 이벤트를 관찰한다.

------------

Ftrace is an internal tracer designed to help out developers and
designers of systems to find what is going on inside the kernel.
It can be used for debugging or analyzing latencies and
performance issues that take place outside of user-space.

Although ftrace is typically considered the function tracer, it
is really a framework of several assorted tracing utilities.
There's latency tracing to examine what occurs between interrupts
disabled and enabled, as well as for preemption and from a time
a task is woken to the task is actually scheduled in.

One of the most common uses of ftrace is the event tracing.
Throughout the kernel are hundreds of static event points that
can be enabled via the tracefs file system to see what is
going on in certain parts of the kernel.

See events.rst for more information.


Implementation Details

구현 세부 사항

41-46

아키텍처 포팅과 내부 구현 세부 사항은 `Documentation/trace/ftrace-design.rst`에서 다룬다. 이 문서는 사용자 관점의 `tracefs` 인터페이스와 추적기 운용에 집중한다.

----------------------

See Documentation/trace/ftrace-design.rst for details for arch porters and such.


The File System

tracefs 파일 시스템

47-94

Ftrace는 제어 파일과 출력 파일을 `tracefs` 파일 시스템에 둔다. 커널에서 ftrace 선택 항목을 하나라도 켜면 `/sys/kernel/tracing` 디렉터리가 만들어진다.

부팅 때 자동으로 마운트하려면 `/etc/fstab`에 다음 항목을 추가한다.

 tracefs       /sys/kernel/tracing       tracefs defaults        0       0

실행 중 직접 마운트하려면 다음 명령을 사용한다.

 mount -t tracefs nodev /sys/kernel/tracing

짧은 경로로 접근하고 싶다면 `/tracing` 심볼릭 링크를 만들 수 있다.

 ln -s /sys/kernel/tracing /tracing

Linux 4.1 이전에는 모든 ftrace 제어 파일이 보통 `/sys/kernel/debug/tracing`에 있는 `debugfs` 아래에 있었다. 호환성을 위해 `debugfs`를 마운트하면 `tracefs`도 이 경로에 자동 마운트되며, `tracefs`의 모든 파일을 같은 `debugfs` 디렉터리에서도 볼 수 있다.

이 문서의 나머지 부분은 `cd /sys/kernel/tracing`을 실행한 상태를 가정하고 긴 경로 접두사를 생략한다. 문서에서 사용하는 모든 시간 값의 단위는 마이크로초다.

tracefs 접근 경로
ftrace 커널 설정/sys/kernel/tracing 생성
tracefs 마운트/sys/kernel/tracing
debugfs 호환 마운트/sys/kernel/debug/tracing
선택적 심볼릭 링크/tracing

독립 tracefs 경로와 이전 debugfs 호환 경로가 같은 제어 파일을 제공한다.

---------------

Ftrace uses the tracefs file system to hold the control files as
well as the files to display output.

When tracefs is configured into the kernel (which selecting any ftrace
option will do) the directory /sys/kernel/tracing will be created. To mount
this directory, you can add to your /etc/fstab file::

 tracefs       /sys/kernel/tracing       tracefs defaults        0       0

Or you can mount it at run time with::

 mount -t tracefs nodev /sys/kernel/tracing

For quicker access to that directory you may want to make a soft link to
it::

 ln -s /sys/kernel/tracing /tracing

.. attention::

  Before 4.1, all ftrace tracing control files were within the debugfs
  file system, which is typically located at /sys/kernel/debug/tracing.
  For backward compatibility, when mounting the debugfs file system,
  the tracefs file system will be automatically mounted at:

  /sys/kernel/debug/tracing

  All files located in the tracefs file system will be located in that
  debugfs file system directory as well.

.. attention::

  Any selected ftrace option will also create the tracefs file system.
  The rest of the document will assume that you are in the ftrace directory
  (cd /sys/kernel/tracing) and will only concentrate on the files within that
  directory and not distract from the content with the extended
  "/sys/kernel/tracing" path name.

That's it! (assuming that you have ftrace configured into your kernel)

After mounting tracefs you will have access to the control and output files
of ftrace. Here is a list of some of the key files:


 Note: all time values are in microseconds.

핵심 추적기와 출력 제어 파일

95-163

`current_tracer`는 현재 설정된 추적기를 표시하거나 바꾼다. 추적기를 바꾸면 기본 링 버퍼와 `snapshot` 버퍼의 내용이 모두 지워진다.

`available_tracers`는 현재 커널에 컴파일된 추적기 종류를 나열한다. 이 파일에 보이는 이름을 `current_tracer`에 쓰면 해당 추적기를 선택할 수 있다.

`tracing_on`은 추적 링 버퍼에 기록할지를 표시하고 제어한다. 0을 쓰면 기록을 멈추고 1을 쓰면 다시 켠다. 기록을 막을 뿐 계측이나 추적기 자체의 오버헤드는 남을 수 있다.

커널 안에서는 `tracing_off()`로 기록을 끌 수 있으며 이때 `tracing_on`도 0이 된다. 사용자 공간은 파일에 1을 써서 다시 활성화할 수 있다. 함수 또는 이벤트의 `traceoff` 트리거도 이 파일을 0으로 만들며 같은 방식으로 재활성화할 수 있다.

`trace`는 사람이 읽을 수 있는 추적 결과를 보여 준다. `O_TRUNC`로 쓰기 열면 링 버퍼가 지워진다. 이 파일은 소비자가 아니므로 추적이 멈춘 상태에서는 반복해서 같은 내용을 읽는다. 추적 중에는 버퍼 전체를 소비하지 않고 읽기 때문에 결과가 일관되지 않을 수 있다.

`trace_pipe`는 `trace`와 같은 형식의 실시간 스트림이다. 새 데이터가 생길 때까지 읽기가 블록되며, 읽은 데이터를 소비하므로 연속 읽기는 매번 더 최신 내용을 반환하고 이미 읽은 레코드는 다시 나오지 않는다.

`trace_options`는 출력에 표시할 데이터의 양과 스택 추적, 타임스탬프 등 추적기와 이벤트의 동작 옵션을 제어한다. `options/` 디렉터리는 각 옵션을 개별 파일로 제공하며 1을 쓰면 설정하고 0을 쓰면 해제한다.

trace와 trace_pipe
파일읽기 방식데이터 소비주 용도
trace현재 전체 버퍼 표시소비하지 않음정지 후 결과 검토
trace_pipe새 데이터까지 블록읽은 레코드 소비실시간 스트리밍

정적 검사와 실시간 소비 방식의 차이를 구분한다.

기록 활성 상태
tracing_on = 1링 버퍼 기록
tracing_off() 또는 traceoff 트리거tracing_on = 0
tracing_on = 0기록 중지
echo 1링 버퍼 기록

여러 경로가 링 버퍼 기록을 멈추고 사용자 공간에서 다시 켤 수 있다.

  current_tracer:

	This is used to set or display the current tracer
	that is configured. Changing the current tracer clears
	the ring buffer content as well as the "snapshot" buffer.

  available_tracers:

	This holds the different types of tracers that
	have been compiled into the kernel. The
	tracers listed here can be configured by
	echoing their name into current_tracer.

  tracing_on:

	This sets or displays whether writing to the trace
	ring buffer is enabled. Echo 0 into this file to disable
	the tracer or 1 to enable it. Note, this only disables
	writing to the ring buffer, the tracing overhead may
	still be occurring.

	The kernel function tracing_off() can be used within the
	kernel to disable writing to the ring buffer, which will
	set this file to "0". User space can re-enable tracing by
	echoing "1" into the file.

	Note, the function and event trigger "traceoff" will also
	set this file to zero and stop tracing. Which can also
	be re-enabled by user space using this file.

  trace:

	This file holds the output of the trace in a human
	readable format (described below). Opening this file for
	writing with the O_TRUNC flag clears the ring buffer content.
        Note, this file is not a consumer. If tracing is off
        (no tracer running, or tracing_on is zero), it will produce
        the same output each time it is read. When tracing is on,
        it may produce inconsistent results as it tries to read
        the entire buffer without consuming it.

  trace_pipe:

	The output is the same as the "trace" file but this
	file is meant to be streamed with live tracing.
	Reads from this file will block until new data is
	retrieved.  Unlike the "trace" file, this file is a
	consumer. This means reading from this file causes
	sequential reads to display more current data. Once
	data is read from this file, it is consumed, and
	will not be read again with a sequential read. The
	"trace" file is static, and if the tracer is not
	adding more data, it will display the same
	information every time it is read.

  trace_options:

	This file lets the user control the amount of data
	that is displayed in one of the above output
	files. Options also exist to modify how a tracer
	or events work (stack traces, timestamps, etc).

  options:

	This is a directory that has a file for every available
	trace option (also in trace_options). Options may also be set
	or cleared by writing a "1" or "0" respectively into the
	corresponding file with the option name.

지연 시간과 버퍼 제어

164-253

`tracing_max_latency`는 irqsoff 같은 일부 추적기가 관찰한 최대 지연 시간과 그때의 추적을 저장한다. 새 지연이 파일의 현재 값보다 클 때만 최대 추적을 교체한다. 사용자가 임계 시간을 쓰면 그 값보다 큰 지연만 기록된다.

`tracing_thresh`는 지원하는 지연 추적기가 이 파일의 값보다 큰 지연을 만날 때마다 추적을 기록하도록 한다. 0보다 큰 값일 때만 활성화되며 단위는 마이크로초다.

`buffer_percent`는 CPU별 `trace_pipe_raw`를 블로킹 읽기하거나 splice할 때 대기자를 깨울 링 버퍼 채움 비율이다. 0은 데이터가 조금이라도 생기면 깨우고, 50은 하위 버퍼가 대략 절반 찰 때, 100은 링 버퍼가 완전히 차서 이전 데이터를 덮기 직전에 깨운다.

`buffer_size_kb`는 CPU 하나의 추적 버퍼 크기를 KiB 단위로 설정하거나 표시한다. 기본적으로 CPU별 크기가 같으며 표시값은 모든 CPU 합계가 아니다. 버퍼는 페이지 단위로 할당되고 관리 메타데이터용 페이지가 추가될 수 있으며 마지막 페이지의 남는 공간도 사용하므로 실제 할당량은 요청값보다 클 수 있다.

CPU별 버퍼 크기를 `per_cpu/cpu0/buffer_size_kb` 등에서 다르게 설정하면 전역 `buffer_size_kb`는 숫자 대신 `X`를 표시한다. `buffer_total_size_kb`는 모든 CPU 추적 버퍼의 합산 크기를 보여 준다.

`buffer_subbuf_size_kb`는 링 버퍼를 구성하는 동일 크기 하위 버퍼의 최소 크기를 설정하거나 표시한다. 이벤트 하나는 하위 버퍼보다 클 수 없고 시작 부분의 메타데이터 공간도 차감되므로, 기본 페이지 크기 하위 버퍼에서는 이벤트 최대 크기가 페이지보다 작다.

커널 구현상 요청보다 더 큰 하위 버퍼를 만들거나 요청을 처리할 수 없어 실패할 수 있다. 하위 버퍼 크기를 키우면 페이지보다 큰 이벤트를 저장할 수 있지만, 변경할 때 추적이 중지되고 링 버퍼와 snapshot 버퍼의 데이터가 모두 폐기된다.

`free_buffer`는 추적 프로세스가 정상 종료되거나 신호로 죽을 때도 버퍼를 최소 크기로 줄이는 수명 주기 장치다. 프로세스가 이 파일을 열어 둔 채 추적하면 종료 시 파일 디스크립터가 닫히면서 버퍼가 해제된다. `disable_on_free` 옵션이 설정돼 있으면 추적도 멈출 수 있다.

buffer_percent 깨움 조건
깨움 조건
0링 버퍼에 데이터가 하나라도 생김
50하위 버퍼가 대략 절반 채워짐
100버퍼가 가득 차 기존 데이터를 덮기 직전

블로킹 raw reader와 splice의 워터마크다.

버퍼 크기 파일
파일의미변경 영향
buffer_size_kbCPU 하나의 버퍼 크기CPU별 값이 다르면 X 표시
buffer_total_size_kb모든 CPU 버퍼의 합읽기 전용 합계
buffer_subbuf_size_kb하위 버퍼 최소 크기추적 정지 및 버퍼 폐기

CPU별 크기, 전체 크기, 이벤트 상한을 각각 제어한다.

  tracing_max_latency:

	Some of the tracers record the max latency.
	For example, the maximum time that interrupts are disabled.
	The maximum time is saved in this file. The max trace will also be
	stored,	and displayed by "trace". A new max trace will only be
	recorded if the latency is greater than the value in this file
	(in microseconds).

	By echoing in a time into this file, no latency will be recorded
	unless it is greater than the time in this file.

  tracing_thresh:

	Some latency tracers will record a trace whenever the
	latency is greater than the number in this file.
	Only active when the file contains a number greater than 0.
	(in microseconds)

  buffer_percent:

	This is the watermark for how much the ring buffer needs to be filled
	before a waiter is woken up. That is, if an application calls a
	blocking read syscall on one of the per_cpu trace_pipe_raw files, it
	will block until the given amount of data specified by buffer_percent
	is in the ring buffer before it wakes the reader up. This also
	controls how the splice system calls are blocked on this file::

	  0   - means to wake up as soon as there is any data in the ring buffer.
	  50  - means to wake up when roughly half of the ring buffer sub-buffers
	        are full.
	  100 - means to block until the ring buffer is totally full and is
	        about to start overwriting the older data.

  buffer_size_kb:

	This sets or displays the number of kilobytes each CPU
	buffer holds. By default, the trace buffers are the same size
	for each CPU. The displayed number is the size of the
	CPU buffer and not total size of all buffers. The
	trace buffers are allocated in pages (blocks of memory
	that the kernel uses for allocation, usually 4 KB in size).
	A few extra pages may be allocated to accommodate buffer management
	meta-data. If the last page allocated has room for more bytes
	than requested, the rest of the page will be used,
	making the actual allocation bigger than requested or shown.
	( Note, the size may not be a multiple of the page size
	due to buffer management meta-data. )

	Buffer sizes for individual CPUs may vary
	(see "per_cpu/cpu0/buffer_size_kb" below), and if they do
	this file will show "X".

  buffer_total_size_kb:

	This displays the total combined size of all the trace buffers.

  buffer_subbuf_size_kb:

	This sets or displays the sub buffer size. The ring buffer is broken up
	into several same size "sub buffers". An event can not be bigger than
	the size of the sub buffer. Normally, the sub buffer is the size of the
	architecture's page (4K on x86). The sub buffer also contains meta data
	at the start which also limits the size of an event.  That means when
	the sub buffer is a page size, no event can be larger than the page
	size minus the sub buffer meta data.

	Note, the buffer_subbuf_size_kb is a way for the user to specify the
	minimum size of the subbuffer. The kernel may make it bigger due to the
	implementation details, or simply fail the operation if the kernel can
	not handle the request.

	Changing the sub buffer size allows for events to be larger than the
	page size.

	Note: When changing the sub-buffer size, tracing is stopped and any
	data in the ring buffer and the snapshot buffer will be discarded.

  free_buffer:

	If a process is performing tracing, and the ring buffer	should be
	shrunk "freed" when the process is finished, even if it were to be
	killed by a signal, this file can be used for that purpose. On close
	of this file, the ring buffer will be resized to its minimum size.
	Having a process that is tracing also open this file, when the process
	exits its file descriptor for this file will be closed, and in doing so,
	the ring buffer will be "freed".

	It may also stop tracing if disable_on_free option is set.

CPU, 함수와 PID 선택

254-339

`tracing_cpumask`는 추적할 CPU만 선택하는 16진수 비트 마스크다.

동적 ftrace를 구성하면 `set_ftrace_filter`가 함수 이름을 받아 `function`과 `function_graph` 추적기 및 함수 프로파일링 대상을 제한한다. 동적 코드 재작성은 `mcount` 호출을 비활성화해 추적을 사용하지 않을 때 성능 오버헤드를 거의 없애며, `available_filter_functions`에 나오는 이름만 지정할 수 있다.

`set_ftrace_filter`는 필터 명령도 지원한다. 문자열 처리와 모든 등록 함수 검사는 비용이 크므로, `available_filter_functions`의 1부터 시작하는 줄 번호를 숫자로 써서 같은 함수를 더 빠르게 선택할 수도 있다.

`set_ftrace_notrace`는 반대 의미의 함수 제외 목록이다. 같은 함수가 포함 필터와 제외 필터에 모두 있으면 제외가 우선해 추적하지 않는다.

`set_ftrace_pid`는 나열한 PID의 스레드만 함수 추적한다. `function-fork` 옵션을 켜면 이 태스크가 fork한 자식 PID를 자동으로 추가하고, 태스크가 종료되면 목록에서 제거한다.

`set_ftrace_notrace_pid`는 나열한 PID를 함수 추적에서 제외한다. `function-fork`가 켜져 있으면 자식도 제외 목록에 추가되고 종료 시 제거된다. 같은 PID가 포함 목록에도 있으면 제외 목록이 우선한다.

`set_event_pid`는 나열한 PID에 대한 이벤트만 기록한다. `sched_switch`와 `sched_wake_up`은 해당 PID가 관계된 이벤트도 추적한다. `event-fork` 옵션은 자식 PID를 자동 추가하고 종료 PID를 제거한다.

`set_event_notrace_pid`는 나열한 PID의 이벤트를 제외한다. 다만 `sched_switch`와 `sched_wakeup`이 추적해야 하는 다른 스레드도 함께 담으면 제외 PID가 관계돼도 이벤트가 기록될 수 있다. `event-fork`의 상속과 제거 규칙은 포함 목록과 같다.

포함과 제외 선택
포함제외범위상속 옵션
set_ftrace_filterset_ftrace_notrace함수 이름/명령해당 없음
set_ftrace_pidset_ftrace_notrace_pid함수 추적 PIDfunction-fork
set_event_pidset_event_notrace_pid이벤트 PIDevent-fork

함수 추적과 이벤트 추적의 이름 및 PID 필터를 구분한다.

필터 우선순위
함수 또는 PID포함 목록 검사
포함됨제외 목록 검사
제외됨추적하지 않음
제외되지 않음추적

포함 목록과 제외 목록이 겹치면 제외가 이긴다.

  tracing_cpumask:

	This is a mask that lets the user only trace on specified CPUs.
	The format is a hex string representing the CPUs.

  set_ftrace_filter:

	When dynamic ftrace is configured in (see the
	section below "dynamic ftrace"), the code is dynamically
	modified (code text rewrite) to disable calling of the
	function profiler (mcount). This lets tracing be configured
	in with practically no overhead in performance.  This also
	has a side effect of enabling or disabling specific functions
	to be traced. Echoing names of functions into this file
	will limit the trace to only those functions.
	This influences the tracers "function" and "function_graph"
	and thus also function profiling (see "function_profile_enabled").

	The functions listed in "available_filter_functions" are what
	can be written into this file.

	This interface also allows for commands to be used. See the
	"Filter commands" section for more details.

	As a speed up, since processing strings can be quite expensive
	and requires a check of all functions registered to tracing, instead
	an index can be written into this file. A number (starting with "1")
	written will instead select the same corresponding at the line position
	of the "available_filter_functions" file.

  set_ftrace_notrace:

	This has an effect opposite to that of
	set_ftrace_filter. Any function that is added here will not
	be traced. If a function exists in both set_ftrace_filter
	and set_ftrace_notrace,	the function will _not_ be traced.

  set_ftrace_pid:

	Have the function tracer only trace the threads whose PID are
	listed in this file.

	If the "function-fork" option is set, then when a task whose
	PID is listed in this file forks, the child's PID will
	automatically be added to this file, and the child will be
	traced by the function tracer as well. This option will also
	cause PIDs of tasks that exit to be removed from the file.

  set_ftrace_notrace_pid:

        Have the function tracer ignore threads whose PID are listed in
        this file.

        If the "function-fork" option is set, then when a task whose
	PID is listed in this file forks, the child's PID will
	automatically be added to this file, and the child will not be
	traced by the function tracer as well. This option will also
	cause PIDs of tasks that exit to be removed from the file.

        If a PID is in both this file and "set_ftrace_pid", then this
        file takes precedence, and the thread will not be traced.

  set_event_pid:

	Have the events only trace a task with a PID listed in this file.
	Note, sched_switch and sched_wake_up will also trace events
	listed in this file.

	To have the PIDs of children of tasks with their PID in this file
	added on fork, enable the "event-fork" option. That option will also
	cause the PIDs of tasks to be removed from this file when the task
	exits.

  set_event_notrace_pid:

	Have the events not trace a task with a PID listed in this file.
	Note, sched_switch and sched_wakeup will trace threads not listed
	in this file, even if a thread's PID is in the file if the
        sched_switch or sched_wakeup events also trace a thread that should
        be traced.

	To have the PIDs of children of tasks with their PID in this file
	added on fork, enable the "event-fork" option. That option will also
	cause the PIDs of tasks to be removed from this file when the task
	exits.

함수 그래프 필터와 콜백 진단

340-428

`set_graph_function`에 함수를 나열하면 함수 그래프 추적기는 그 함수와 그 함수가 호출한 하위 함수만 추적한다. `set_ftrace_filter`와 `set_ftrace_notrace`도 여전히 실제 추적 대상에 영향을 준다.

`set_graph_notrace`는 지정 함수에 진입한 순간부터 그 함수가 끝날 때까지 함수 그래프 추적을 끈다. 특정 함수가 호출하는 전체 하위 트리를 무시할 때 유용하다.

`available_filter_functions`는 ftrace가 처리해 추적할 수 있는 함수 이름을 나열하며, 함수 필터와 그래프 필터 파일에 전달할 수 있다. `available_filter_functions_addrs`는 같은 목록에 각 함수의 패치 지점 주소를 함께 표시한다. 이 주소는 `/proc/kallsyms`의 함수 주소와 다를 수 있다.

`dyn_ftrace_total_info`는 디버깅용으로, nop으로 변환되어 추적 가능해진 함수 수를 보여 준다.

`enabled_functions`는 현재 함수 콜백이 연결된 모든 함수와 연결된 콜백 수를 표시한다. 추적 기반 구조 외의 하위 시스템도 함수 추적 기능을 사용할 수 있으며, 콜백 하나가 여러 함수를 부르는 사실은 이 수에 반영되지 않는다.

콜백이 레지스터 저장 속성을 사용하면 해당 함수 줄에 `R`, `regs->ip`를 바꿀 수 있는 IP modify 속성이면 `I`가 표시된다. BPF 같은 비-ftrace 직접 트램펄린은 `D`, 직접 호출을 지원하지 않아 ops 함수가 진입점 위에 배치되는 아키텍처는 `O`로 표시된다.

함수가 과거에 IP modify 또는 direct call로 수정된 적이 있으면 `M`이 표시되며 이 플래그는 지워지지 않는다. 아키텍처가 지원하면 직접 호출되는 콜백도 보여 준다. 콜백 수가 1보다 크면 보통 `ftrace_ops_list_func()`이고, 표준이 아닌 콜백 전용 트램펄린이면 트램펄린 주소와 그 대상 함수도 출력한다.

`touched_functions`는 ftrace 기반 구조를 통해 한 번이라도 함수 콜백이 연결된 모든 함수를 `enabled_functions`와 같은 형식으로 보존한다. IP modify 또는 direct trampoline으로 수정된 적 있는 함수는 다음 명령으로 찾는다.

	grep ' M ' /sys/kernel/tracing/touched_functions
enabled_functions 상태 문자
문자의미
Rsave regs 콜백
IIP modify 콜백
DBPF 등 direct trampoline
O아키텍처의 ops 함수 배치 방식
M과거 IP modify 또는 direct call 이력, 지워지지 않음

콜백이 요구하거나 과거에 사용한 특별한 동작을 표시한다.

함수 그래프 제외
set_graph_function 대상그래프 추적 시작
set_graph_notrace 함수 진입그래프 추적 일시 중지
제외 함수와 하위 호출기록하지 않음
제외 함수 반환그래프 추적 재개

지정 함수의 전체 호출 하위 트리를 잠시 숨긴다.

  set_graph_function:

	Functions listed in this file will cause the function graph
	tracer to only trace these functions and the functions that
	they call. (See the section "dynamic ftrace" for more details).
	Note, set_ftrace_filter and set_ftrace_notrace still affects
	what functions are being traced.

  set_graph_notrace:

	Similar to set_graph_function, but will disable function graph
	tracing when the function is hit until it exits the function.
	This makes it possible to ignore tracing functions that are called
	by a specific function.

  available_filter_functions:

	This lists the functions that ftrace has processed and can trace.
	These are the function names that you can pass to
	"set_ftrace_filter", "set_ftrace_notrace",
	"set_graph_function", or "set_graph_notrace".
	(See the section "dynamic ftrace" below for more details.)

  available_filter_functions_addrs:

	Similar to available_filter_functions, but with address displayed
	for each function. The displayed address is the patch-site address
	and can differ from /proc/kallsyms address.

  dyn_ftrace_total_info:

	This file is for debugging purposes. The number of functions that
	have been converted to nops and are available to be traced.

  enabled_functions:

	This file is more for debugging ftrace, but can also be useful
	in seeing if any function has a callback attached to it.
	Not only does the trace infrastructure use ftrace function
	trace utility, but other subsystems might too. This file
	displays all functions that have a callback attached to them
	as well as the number of callbacks that have been attached.
	Note, a callback may also call multiple functions which will
	not be listed in this count.

	If the callback registered to be traced by a function with
	the "save regs" attribute (thus even more overhead), an 'R'
	will be displayed on the same line as the function that
	is returning registers.

	If the callback registered to be traced by a function with
	the "ip modify" attribute (thus the regs->ip can be changed),
	an 'I' will be displayed on the same line as the function that
	can be overridden.

	If a non-ftrace trampoline is attached (BPF) a 'D' will be displayed.
	Note, normal ftrace trampolines can also be attached, but only one
	"direct" trampoline can be attached to a given function at a time.

	Some architectures can not call direct trampolines, but instead have
	the ftrace ops function located above the function entry point. In
	such cases an 'O' will be displayed.

	If a function had either the "ip modify" or a "direct" call attached to
	it in the past, a 'M' will be shown. This flag is never cleared. It is
	used to know if a function was ever modified by the ftrace infrastructure,
	and can be used for debugging.

	If the architecture supports it, it will also show what callback
	is being directly called by the function. If the count is greater
	than 1 it most likely will be ftrace_ops_list_func().

	If the callback of a function jumps to a trampoline that is
	specific to the callback and which is not the standard trampoline,
	its address will be printed as well as the function that the
	trampoline calls.

  touched_functions:

	This file contains all the functions that ever had a function callback
	to it via the ftrace infrastructure. It has the same format as
	enabled_functions but shows all functions that have ever been
	traced.

	To see any function that has every been modified by "ip modify" or a
	direct trampoline, one can perform the following command:

	grep ' M ' /sys/kernel/tracing/touched_functions

프로파일, 동적 probe와 메타데이터

429-514

`function_profile_enabled`를 설정하면 function 추적기 또는 구성된 경우 function_graph 추적기로 모든 함수를 프로파일링한다. 함수 호출 횟수 히스토그램을 유지하며 함수 그래프 추적기가 있으면 함수에서 소비한 시간도 집계한다. CPU별 결과는 `trace_stat/function0`, `function1` 같은 파일에서 확인한다.

`trace_stat/` 디렉터리는 여러 추적 통계를 보관한다. `kprobe_events`는 동적 추적 지점을 만들며 자세한 내용은 `kprobetrace.rst`, `kprobe_profile`은 그 동적 지점의 통계를 제공한다.

`max_graph_depth`는 함수 그래프 추적기가 내려갈 최대 호출 깊이다. 1로 설정하면 사용자 공간에서 들어온 첫 번째 커널 함수만 보인다.

`printk_formats`는 raw format을 읽는 도구를 위한 포인터-문자열 매핑이다. 링 버퍼 이벤트는 문자열 자체가 아니라 포인터만 저장할 수 있으므로, 이 파일이 문자열과 주소를 제공해 외부 도구가 포인터를 원문 문자열로 해석하게 한다.

`saved_cmdlines`는 이벤트에 기록된 PID를 태스크 명령 이름 `comm`으로 표시하기 위한 캐시다. 매핑이 없으면 출력에 `<...>`가 나타난다. `record-cmd` 옵션이 0이면 기록 중 comm 저장을 중지하며 기본값은 활성화다.

`saved_cmdlines_size`는 기본 128개인 comm 캐시 수를 조절한다. `saved_tgids`는 `record-tgid` 옵션을 켰을 때 문맥 전환마다 스레드 PID와 태스크 그룹 ID의 매핑을 저장하며 기본적으로 꺼져 있다.

`snapshot`은 별도 snapshot 버퍼의 내용을 표시하고 현재 실행 중인 추적의 스냅샷을 만들 수 있게 한다. 자세한 운용은 뒤의 Snapshot 절에서 설명한다.

스택 추적기를 켜면 `stack_max_size`가 관찰된 최대 스택 크기를, `stack_trace`가 그 최대 사용 시점의 역추적을 표시한다. `stack_trace_filter`는 스택 추적기가 검사할 함수를 `set_ftrace_filter`와 비슷한 방식으로 제한한다.

프로파일과 메타데이터 파일
파일제공 정보
function_profile_enabled함수 횟수와 선택적 실행 시간
printk_formats문자열 포인터와 실제 문자열 매핑
saved_cmdlinesPID와 comm 매핑
saved_tgidsPID와 TGID 매핑
snapshot별도 시점 보존 버퍼
stack_trace최대 스택 사용 시 역추적

수집 데이터가 무엇을 보완하는지 정리한다.

  function_profile_enabled:

	When set it will enable all functions with either the function
	tracer, or if configured, the function graph tracer. It will
	keep a histogram of the number of functions that were called
	and if the function graph tracer was configured, it will also keep
	track of the time spent in those functions. The histogram
	content can be displayed in the files:

	trace_stat/function<cpu> ( function0, function1, etc).

  trace_stat:

	A directory that holds different tracing stats.

  kprobe_events:

	Enable dynamic trace points. See kprobetrace.rst.

  kprobe_profile:

	Dynamic trace points stats. See kprobetrace.rst.

  max_graph_depth:

	Used with the function graph tracer. This is the max depth
	it will trace into a function. Setting this to a value of
	one will show only the first kernel function that is called
	from user space.

  printk_formats:

	This is for tools that read the raw format files. If an event in
	the ring buffer references a string, only a pointer to the string
	is recorded into the buffer and not the string itself. This prevents
	tools from knowing what that string was. This file displays the string
	and address for	the string allowing tools to map the pointers to what
	the strings were.

  saved_cmdlines:

	Only the pid of the task is recorded in a trace event unless
	the event specifically saves the task comm as well. Ftrace
	makes a cache of pid mappings to comms to try to display
	comms for events. If a pid for a comm is not listed, then
	"<...>" is displayed in the output.

	If the option "record-cmd" is set to "0", then comms of tasks
	will not be saved during recording. By default, it is enabled.

  saved_cmdlines_size:

	By default, 128 comms are saved (see "saved_cmdlines" above). To
	increase or decrease the amount of comms that are cached, echo
	the number of comms to cache into this file.

  saved_tgids:

	If the option "record-tgid" is set, on each scheduling context switch
	the Task Group ID of a task is saved in a table mapping the PID of
	the thread to its TGID. By default, the "record-tgid" option is
	disabled.

  snapshot:

	This displays the "snapshot" buffer and also lets the user
	take a snapshot of the current running trace.
	See the "Snapshot" section below for more details.

  stack_max_size:

	When the stack tracer is activated, this will display the
	maximum stack size it has encountered.
	See the "Stack Trace" section below.

  stack_trace:

	This displays the stack back trace of the largest stack
	that was encountered when the stack tracer is activated.
	See the "Stack Trace" section below.

  stack_trace_filter:

	This is similar to "set_ftrace_filter" but it limits what
	functions the stack tracer will check.

추적 시계 선택

515-605

`trace_clock`은 링 버퍼 이벤트에 붙일 타임스탬프의 시계를 선택한다. 기본 `local`은 매우 빠른 CPU별 시계지만 시스템에 따라 CPU 간 동기화가 되지 않아 다른 CPU의 이벤트와 비교할 때 단조 증가하지 않을 수 있다.

현재 사용할 수 있는 시계와 선택 상태는 다음처럼 확인한다. 대괄호로 둘러싼 이름이 현재 시계다.

	  # cat trace_clock
	  [local] global counter x86-tsc

`local`은 빠른 기본 시계지만 CPU 간 동기화를 보장하지 않는다. `global`은 모든 CPU에서 동기화되지만 조금 느릴 수 있다. `counter`는 실제 시계가 아니라 모든 CPU가 공유하는 원자 카운터로, 시간보다 CPU 간 이벤트의 정확한 순서를 알아야 할 때 유용하다.

`uptime`은 jiffies 카운터를 사용해 부팅 이후 상대 시간을 기록한다. `perf`는 perf와 같은 시계를 사용해 두 데이터 흐름을 서로 끼워 맞추기 쉽게 한다.

아키텍처는 자체 시계를 제공할 수 있다. `x86-tsc`는 x86 TSC 사이클 시계이고, `ppc-tb`는 CPU 간 동기화된 PowerPC timebase 레지스터를 사용한다. `tb_offset`을 알면 hypervisor와 guest 이벤트도 상관 분석할 수 있다.

`mono`는 NTP 속도 조정의 영향을 받는 빠른 `CLOCK_MONOTONIC`, `mono_raw`는 속도 조정을 받지 않고 하드웨어 clocksource와 같은 속도로 흐르는 `CLOCK_MONOTONIC_RAW`다.

`boot`는 `CLOCK_BOOTTIME`으로 fast monotonic에 suspend 시간을 더한다. suspend 경로에서 빠르게 읽도록 설계돼 suspend 시간 반영과 fast mono 갱신 사이에 읽으면 갱신이 약간 일찍 보일 수 있다. 32비트 시스템에서는 64비트 boot offset의 부분 갱신을 볼 가능성도 있다. 드문 현상이며 후처리로 처리할 수 있고 자세한 내용은 `ktime_get_boot_fast_ns()` 주석에 있다.

`tai`는 벽시계에서 유도한 `CLOCK_TAI`이며 NTP 윤초 삽입 때문에 불연속 또는 역행하지 않는다. 다만 시스템 시간 설정이나 offset을 지정한 `adjtimex()`로 내부 TAI offset이 갱신되는 순간에는 빠른 추적용 읽기가 잘못된 값을 낼 수 있다. 자세한 내용은 `ktime_get_tai_fast_ns()` 주석을 참고한다.

시계를 선택하려면 이름을 `trace_clock`에 쓴다. 변경하면 링 버퍼와 snapshot 버퍼 내용이 모두 지워진다.

	  # echo global > trace_clock
주요 trace_clock
시계특성용도/주의
local빠른 CPU별 시계CPU 간 비동기 가능
globalCPU 간 동기화local보다 느릴 수 있음
counter원자 순번CPU 간 사건 순서
uptimejiffies 기반부팅 이후 상대 시간
perfperf와 동일 시계데이터 상관 분석
mono / mono_raw단조 시계NTP 속도 조정 적용 여부가 다름
bootsuspend 포함빠른 읽기의 드문 경계 효과
tai윤초 역행 없음TAI offset 갱신 경계 주의

속도, CPU 간 비교, 시간 의미에 따라 선택한다.

  trace_clock:

	Whenever an event is recorded into the ring buffer, a
	"timestamp" is added. This stamp comes from a specified
	clock. By default, ftrace uses the "local" clock. This
	clock is very fast and strictly per CPU, but on some
	systems it may not be monotonic with respect to other
	CPUs. In other words, the local clocks may not be in sync
	with local clocks on other CPUs.

	Usual clocks for tracing::

	  # cat trace_clock
	  [local] global counter x86-tsc

	The clock with the square brackets around it is the one in effect.

	local:
		Default clock, but may not be in sync across CPUs

	global:
		This clock is in sync with all CPUs but may
		be a bit slower than the local clock.

	counter:
		This is not a clock at all, but literally an atomic
		counter. It counts up one by one, but is in sync
		with all CPUs. This is useful when you need to
		know exactly the order events occurred with respect to
		each other on different CPUs.

	uptime:
		This uses the jiffies counter and the time stamp
		is relative to the time since boot up.

	perf:
		This makes ftrace use the same clock that perf uses.
		Eventually perf will be able to read ftrace buffers
		and this will help out in interleaving the data.

	x86-tsc:
		Architectures may define their own clocks. For
		example, x86 uses its own TSC cycle clock here.

	ppc-tb:
		This uses the powerpc timebase register value.
		This is in sync across CPUs and can also be used
		to correlate events across hypervisor/guest if
		tb_offset is known.

	mono:
		This uses the fast monotonic clock (CLOCK_MONOTONIC)
		which is monotonic and is subject to NTP rate adjustments.

	mono_raw:
		This is the raw monotonic clock (CLOCK_MONOTONIC_RAW)
		which is monotonic but is not subject to any rate adjustments
		and ticks at the same rate as the hardware clocksource.

	boot:
		This is the boot clock (CLOCK_BOOTTIME) and is based on the
		fast monotonic clock, but also accounts for time spent in
		suspend. Since the clock access is designed for use in
		tracing in the suspend path, some side effects are possible
		if clock is accessed after the suspend time is accounted before
		the fast mono clock is updated. In this case, the clock update
		appears to happen slightly sooner than it normally would have.
		Also on 32-bit systems, it's possible that the 64-bit boot offset
		sees a partial update. These effects are rare and post
		processing should be able to handle them. See comments in the
		ktime_get_boot_fast_ns() function for more information.

	tai:
		This is the tai clock (CLOCK_TAI) and is derived from the wall-
		clock time. However, this clock does not experience
		discontinuities and backwards jumps caused by NTP inserting leap
		seconds. Since the clock access is designed for use in tracing,
		side effects are possible. The clock access may yield wrong
		readouts in case the internal TAI offset is updated e.g., caused
		by setting the system time or using adjtimex() with an offset.
		These effects are rare and post processing should be able to
		handle them. See comments in the ktime_get_tai_fast_ns()
		function for more information.

	To set a clock, simply echo the clock name into this file::

	  # echo global > trace_clock

	Setting a clock clears the ring buffer content as well as the
	"snapshot" buffer.

사용자 공간 마커와 uprobe

606-655

`trace_marker`는 사용자 공간 동작과 커널 이벤트를 동기화하는 파일이다. 문자열을 쓰면 ftrace 버퍼에 레코드로 들어간다. 애플리케이션 시작 때 한 번 열어 파일 디스크립터를 보관하고 필요한 시점마다 포맷된 문자열을 쓰는 방식이 효율적이다.

		void trace_write(const char *fmt, ...)
		{
			va_list ap;
			char buf[256];
			int n;

			if (trace_fd < 0)
				return;

			va_start(ap, fmt);
			n = vsnprintf(buf, 256, fmt, ap);
			va_end(ap);

			write(trace_fd, buf, n);
		}

애플리케이션 초기화에서는 다음처럼 쓰기 전용으로 파일을 연다.

		trace_fd = open("trace_marker", O_WRONLY);

`trace_marker` 쓰기는 `/sys/kernel/tracing/events/ftrace/print/trigger`에 설정한 트리거를 시작할 수도 있다. 이벤트 트리거는 `Documentation/trace/events.rst`, 예제는 `Documentation/trace/histogram.rst`의 3절을 참고한다.

`trace_marker_raw`는 같은 목적의 바이너리 데이터 입력 파일이다. 도구는 `trace_pipe_raw`에서 데이터를 읽어 자체 형식으로 해석할 수 있다.

`uprobe_events`는 프로그램에 동적 추적 지점을 추가하며 `uprobetracer.rst`를 참고한다. `uprobe_profile`은 uprobe 통계를 제공한다.

사용자 공간 마커
애플리케이션 시작trace_marker 열기
사용자 공간 상태 변화문자열 쓰기
trace_markerftrace 링 버퍼
ftrace 링 버퍼커널 이벤트와 시간순 분석

애플리케이션의 의미 있는 시점을 커널 추적 스트림과 합친다.

  trace_marker:

	This is a very useful file for synchronizing user space
	with events happening in the kernel. Writing strings into
	this file will be written into the ftrace buffer.

	It is useful in applications to open this file at the start
	of the application and just reference the file descriptor
	for the file::

		void trace_write(const char *fmt, ...)
		{
			va_list ap;
			char buf[256];
			int n;

			if (trace_fd < 0)
				return;

			va_start(ap, fmt);
			n = vsnprintf(buf, 256, fmt, ap);
			va_end(ap);

			write(trace_fd, buf, n);
		}

	start::

		trace_fd = open("trace_marker", O_WRONLY);

	Note: Writing into the trace_marker file can also initiate triggers
	      that are written into /sys/kernel/tracing/events/ftrace/print/trigger
	      See "Event triggers" in Documentation/trace/events.rst and an
              example in Documentation/trace/histogram.rst (Section 3.)

  trace_marker_raw:

	This is similar to trace_marker above, but is meant for binary data
	to be written to it, where a tool can be used to parse the data
	from trace_pipe_raw.

  uprobe_events:

	Add dynamic tracepoints in programs.
	See uprobetracer.rst

  uprobe_profile:

	Uprobe statistics. See uprobetrace.txt

인스턴스, 이벤트와 타임스탬프 모드

656-713

`instances/`는 여러 독립 추적 버퍼를 만들어 서로 다른 이벤트를 각 버퍼에 기록하는 기능이다. 자세한 내용은 뒤의 Instances 절에서 설명한다.

`events/`는 커널에 컴파일된 정적 tracepoint를 시스템별로 묶어 보여 주는 이벤트 디렉터리다. 여러 계층의 `enable` 파일에 1을 써서 범위별로 tracepoint를 켤 수 있으며 자세한 내용은 `events.rst`에 있다.

`set_event`에 이벤트 이름을 쓰면 해당 이벤트를 활성화한다. `available_events`는 활성화할 수 있는 전체 이벤트 목록을 제공한다.

`timestamp_mode`는 이벤트 버퍼에 기록할 타임스탬프 표현을 제어한다. 서로 다른 모드의 이벤트가 한 버퍼에 함께 있을 수 있지만, 각 이벤트는 기록 시점에 활성화된 모드를 사용한다. 기본값은 `delta`다.

현재 모드는 다음처럼 확인하며 대괄호가 선택 상태를 표시한다.

	  # cat timestamp_mode
	  [delta] absolute

`delta`는 버퍼별 기준 타임스탬프에 대한 차이를 저장하는 효율적인 기본 모드다. `absolute`는 다른 값에 대한 차이가 아닌 전체 타임스탬프를 저장하므로 더 많은 공간을 사용하고 효율이 낮다.

`hwlat_detector/`는 Hardware Latency Detector의 제어 디렉터리며 자세한 사용법은 뒤의 해당 절에 있다.

이벤트 타임스탬프 모드
모드기록값비용
delta버퍼 기준에 대한 차이기본값, 공간 효율적
absolute완전한 타임스탬프더 큰 레코드

저장 공간과 독립적인 시간값 사이의 선택이다.

  instances:

	This is a way to make multiple trace buffers where different
	events can be recorded in different buffers.
	See "Instances" section below.

  events:

	This is the trace event directory. It holds event tracepoints
	(also known as static tracepoints) that have been compiled
	into the kernel. It shows what event tracepoints exist
	and how they are grouped by system. There are "enable"
	files at various levels that can enable the tracepoints
	when a "1" is written to them.

	See events.rst for more information.

  set_event:

	By echoing in the event into this file, will enable that event.

	See events.rst for more information.

  available_events:

	A list of events that can be enabled in tracing.

	See events.rst for more information.

  timestamp_mode:

	Certain tracers may change the timestamp mode used when
	logging trace events into the event buffer.  Events with
	different modes can coexist within a buffer but the mode in
	effect when an event is logged determines which timestamp mode
	is used for that event.  The default timestamp mode is
	'delta'.

	Usual timestamp modes for tracing:

	  # cat timestamp_mode
	  [delta] absolute

	  The timestamp mode with the square brackets around it is the
	  one in effect.

	  delta: Default timestamp mode - timestamp is a delta against
	         a per-buffer timestamp.

	  absolute: The timestamp is a full timestamp, not a delta
                 against some other value.  As such it takes up more
                 space and is less efficient.

  hwlat_detector:

	Directory for the Hardware Latency Detector.
	See "Hardware Latency Detector" section below.

CPU별 버퍼와 통계

714-794

`per_cpu/`는 CPU별 추적 정보를 담는 디렉터리다. ftrace 링 버퍼는 CPU마다 분리되어 쓰기 작업을 원자적으로 수행하고 CPU 사이 캐시 라인 이동을 피한다.

`per_cpu/cpu0/buffer_size_kb`는 CPU 0 버퍼의 크기만 표시하거나 설정한다. CPU별 버퍼는 서로 다른 크기를 가질 수 있으며 전역 `buffer_size_kb`와 달리 해당 CPU만 바꾼다.

`per_cpu/cpu0/trace`는 CPU 0의 데이터만 보여 주고, 이 파일에 쓰면 CPU 0 버퍼만 지운다.

`per_cpu/cpu0/trace_pipe`는 CPU 0 전용 소비형 실시간 읽기다. `per_cpu/cpu0/trace_pipe_raw`는 링 버퍼의 바이너리 형식을 직접 추출하며, `splice()`로 파일이나 수집 서버가 있는 네트워크에 빠르게 전송할 수 있다. `trace_pipe`처럼 읽은 데이터는 소비되어 다음 읽기에 다시 나오지 않는다.

`per_cpu/cpu0/snapshot`은 지원되는 경우 CPU 0만 스냅샷하고 그 내용을 표시하며, 쓰면 CPU 0 snapshot 버퍼만 지운다. `per_cpu/cpu0/snapshot_raw`는 같은 버퍼를 바이너리 형식으로 읽는다.

`per_cpu/cpu0/stats`는 CPU 0 링 버퍼 상태를 보여 준다. `entries`는 아직 버퍼에 남은 이벤트 수, `overrun`은 버퍼가 가득 차 이전 이벤트를 덮어 잃은 수, `bytes`는 덮이지 않고 실제 읽은 바이트 수다.

`commit overrun`은 항상 0이어야 한다. 재진입 가능한 링 버퍼 안에서 중첩 이벤트가 너무 많이 발생해 버퍼를 채우고 이벤트를 버리기 시작하면 증가한다.

`oldest event ts`는 버퍼의 가장 오래된 타임스탬프, `now ts`는 현재 타임스탬프다. `dropped events`는 overwrite 옵션이 꺼져 있어서 잃은 이벤트 수, `read events`는 읽은 이벤트 수다.

CPU별 파일
파일역할소비 여부
buffer_size_kb해당 CPU 버퍼 크기해당 없음
trace해당 CPU 텍스트 출력/지우기비소비
trace_pipe해당 CPU 실시간 텍스트소비
trace_pipe_raw해당 CPU 바이너리 스트림소비
snapshot해당 CPU snapshot비소비
snapshot_raw해당 CPU snapshot 바이너리비소비
stats해당 CPU 링 버퍼 통계해당 없음

전역 버퍼와 분리해 한 CPU의 데이터만 제어하거나 소비한다.

per_cpu stats 필드
필드의미
entries버퍼에 남은 이벤트
overrunoverwrite로 잃은 이벤트
commit overrun중첩 이벤트로 commit 중 손실, 정상은 0
bytes실제로 읽은 바이트
oldest event ts / now ts가장 오래된 시각 / 현재 시각
dropped eventsoverwrite 비활성 때문에 버린 이벤트
read events읽은 이벤트 수

손실, 읽기, 시간 상태를 진단한다.

CPU별 raw 데이터 수집
CPU별 링 버퍼trace_pipe_raw
trace_pipe_rawsplice()
splice()파일 또는 네트워크
읽은 레코드버퍼에서 소비

소비형 raw 스트림을 복사 없이 수집 대상으로 전송할 수 있다.

  per_cpu:

	This is a directory that contains the trace per_cpu information.

  per_cpu/cpu0/buffer_size_kb:

	The ftrace buffer is defined per_cpu. That is, there's a separate
	buffer for each CPU to allow writes to be done atomically,
	and free from cache bouncing. These buffers may have different
	size buffers. This file is similar to the buffer_size_kb
	file, but it only displays or sets the buffer size for the
	specific CPU. (here cpu0).

  per_cpu/cpu0/trace:

	This is similar to the "trace" file, but it will only display
	the data specific for the CPU. If written to, it only clears
	the specific CPU buffer.

  per_cpu/cpu0/trace_pipe

	This is similar to the "trace_pipe" file, and is a consuming
	read, but it will only display (and consume) the data specific
	for the CPU.

  per_cpu/cpu0/trace_pipe_raw

	For tools that can parse the ftrace ring buffer binary format,
	the trace_pipe_raw file can be used to extract the data
	from the ring buffer directly. With the use of the splice()
	system call, the buffer data can be quickly transferred to
	a file or to the network where a server is collecting the
	data.

	Like trace_pipe, this is a consuming reader, where multiple
	reads will always produce different data.

  per_cpu/cpu0/snapshot:

	This is similar to the main "snapshot" file, but will only
	snapshot the current CPU (if supported). It only displays
	the content of the snapshot for a given CPU, and if
	written to, only clears this CPU buffer.

  per_cpu/cpu0/snapshot_raw:

	Similar to the trace_pipe_raw, but will read the binary format
	from the snapshot buffer for the given CPU.

  per_cpu/cpu0/stats:

	This displays certain stats about the ring buffer:

	entries:
		The number of events that are still in the buffer.

	overrun:
		The number of lost events due to overwriting when
		the buffer was full.

	commit overrun:
		Should always be zero.
		This gets set if so many events happened within a nested
		event (ring buffer is re-entrant), that it fills the
		buffer and starts dropping events.

	bytes:
		Bytes actually read (not overwritten).

	oldest event ts:
		The oldest timestamp in the buffer

	now ts:
		The current timestamp

	dropped events:
		Events lost due to overwrite option being off.

	read events:
		The number of events read.

사용 가능한 추적기

795-888

`The Tracers` 절은 `current_tracer`에 설정할 수 있는 주요 추적기를 설명한다. 실제 사용 가능한 목록은 커널 구성에 따라 달라지며 `available_tracers`에서 확인한다.

`function`은 모든 커널 함수의 호출 진입을 추적한다. `function_graph`는 진입과 종료를 모두 추적해 C 소스와 비슷한 호출 그래프를 그리며, 함수가 시작하고 반환한 시각을 인스턴스별로 내부 계산한다.

두 인스턴스가 같은 함수에 `function_graph`를 동시에 사용하면 각 인스턴스가 타임스탬프를 서로 다른 순간에 읽으므로 계산된 실행 시간이 약간 다를 수 있다.

`blk`는 `blktrace` 사용자 애플리케이션이 사용하는 블록 추적기다. `hwlat`는 하드웨어가 만든 지연을 탐지하며 뒤의 Hardware Latency Detector 절에서 자세히 다룬다.

`irqsoff`는 인터럽트를 비활성화한 구간을 추적하고 가장 긴 최대 지연의 추적을 저장한다. 새 최대값은 이전 추적을 교체한다. 이 추적기를 선택하면 `latency-format`이 자동 활성화되며 `tracing_max_latency`와 함께 사용한다.

`preemptoff`는 preemption 비활성 시간을, `preemptirqsoff`는 irq 또는 preemption 가운데 하나 이상이 비활성인 가장 긴 시간을 기록한다.

`wakeup`은 깨어난 최고 우선순위 태스크가 실제로 스케줄될 때까지의 최대 지연을 추적하며 일반 개발자가 기대하는 모든 태스크를 다룬다. `wakeup_rt`는 RT 태스크만, `wakeup_dl`은 `SCHED_DEADLINE` 태스크만 같은 지연을 측정한다.

`mmiotrace`는 바이너리 모듈이 하드웨어에 수행하는 모든 I/O 읽기와 쓰기 호출을 추적하는 특수 추적기다. `branch`는 커널의 likely/unlikely 분기가 실행됐는지와 예측이 맞았는지를 기록한다.

`nop`는 아무것도 추적하지 않는다. 활성 추적기를 모두 제거하려면 `current_tracer`에 `nop`을 쓴다.

주요 ftrace 추적기
추적기관찰 대상특징
function커널 함수 진입호출 흐름
function_graph함수 진입과 종료호출 그래프와 실행 시간
blk블록 I/Oblktrace 백엔드
hwlat하드웨어 유발 지연Hardware Latency Detector
irqsoffIRQ 비활성 구간가장 긴 지연 저장
preemptoffpreemption 비활성 구간가장 긴 지연 저장
preemptirqsoffIRQ 또는 preemption 비활성가장 긴 지연 저장
wakeup / wakeup_rt / wakeup_dl태스크 wakeup 지연태스크 클래스별 측정
mmiotrace모듈의 하드웨어 I/O읽기와 쓰기 호출
branchlikely/unlikely 분기예측 적중 여부
nop없음활성 추적기 해제

관찰 대상과 최대 지연 보존 여부를 구분한다.

지연 추적기 최대값 교체
지연 구간 관찰tracing_max_latency와 비교
기존 값 이하현재 최대 추적 유지
새 최대값이전 추적 교체
이전 추적 교체trace에서 latency-format으로 표시

더 긴 지연이 관찰될 때만 보존된 추적을 바꾼다.

The Tracers
-----------

Here is the list of current tracers that may be configured.

  "function"

	Function call tracer to trace all kernel functions.

  "function_graph"

	Similar to the function tracer except that the
	function tracer probes the functions on their entry
	whereas the function graph tracer traces on both entry
	and exit of the functions. It then provides the ability
	to draw a graph of function calls similar to C code
	source.

	Note that the function graph calculates the timings of when the
	function starts and returns internally and for each instance. If
	there are two instances that run function graph tracer and traces
	the same functions, the length of the timings may be slightly off as
	each read the timestamp separately and not at the same time.

  "blk"

	The block tracer. The tracer used by the blktrace user
	application.

  "hwlat"

	The Hardware Latency tracer is used to detect if the hardware
	produces any latency. See "Hardware Latency Detector" section
	below.

  "irqsoff"

	Traces the areas that disable interrupts and saves
	the trace with the longest max latency.
	See tracing_max_latency. When a new max is recorded,
	it replaces the old trace. It is best to view this
	trace with the latency-format option enabled, which
	happens automatically when the tracer is selected.

  "preemptoff"

	Similar to irqsoff but traces and records the amount of
	time for which preemption is disabled.

  "preemptirqsoff"

	Similar to irqsoff and preemptoff, but traces and
	records the largest time for which irqs and/or preemption
	is disabled.

  "wakeup"

	Traces and records the max latency that it takes for
	the highest priority task to get scheduled after
	it has been woken up.
        Traces all tasks as an average developer would expect.

  "wakeup_rt"

        Traces and records the max latency that it takes for just
        RT tasks (as the current "wakeup" does). This is useful
        for those interested in wake up timings of RT tasks.

  "wakeup_dl"

	Traces and records the max latency that it takes for
	a SCHED_DEADLINE task to be woken (as the "wakeup" and
	"wakeup_rt" does).

  "mmiotrace"

	A special tracer that is used to trace binary modules.
	It will trace all the calls that a module makes to the
	hardware. Everything it writes and reads from the I/O
	as well.

  "branch"

	This tracer can be configured when tracing likely/unlikely
	calls within the kernel. It will trace when a likely and
	unlikely branch is hit and if it was correct in its prediction
	of being correct.

  "nop"

	This is the "trace nothing" tracer. To remove all
	tracers from tracing simply echo "nop" into
	current_tracer.

오류 조건과 error_log

889-920

대부분의 ftrace 명령 오류는 표준 반환 코드로 분명하게 전달된다. 복잡한 명령 가운데 일부는 `tracing/error_log`에 더 자세한 원인과 명령 안의 오류 위치를 남긴다.

`tracing/error_log`는 순환 오류 로그이며 현재는 최근 실패 명령 8개의 ftrace 오류를 보관한다. 지원 명령이 실패한 뒤 파일을 읽으면 위치, 오류 설명, 원래 명령과 caret 위치를 확인할 수 있다.

    # echo xxx > /sys/kernel/tracing/events/sched/sched_wakeup/trigger
    echo: write error: Invalid argument

    # cat /sys/kernel/tracing/error_log
    [ 5348.887237] location: error: Couldn't yyy: zzz
      Command: xxx
               ^
    [ 7517.023364] location: error: Bad rrr: sss
      Command: ppp qqq
                   ^

빈 문자열을 쓰면 오류 로그를 지운다.

    # echo > /sys/kernel/tracing/error_log
확장 오류 진단
tracefs 명령 실패표준 errno 반환
확장 오류 지원 명령tracing/error_log 읽기
tracing/error_log위치와 실패 명령 표시
빈 문자열 쓰기오류 로그 초기화

표준 쓰기 오류 뒤 error_log에서 명령의 상세 실패 위치를 찾는다.

Error conditions
----------------

  For most ftrace commands, failure modes are obvious and communicated
  using standard return codes.

  For other more involved commands, extended error information may be
  available via the tracing/error_log file.  For the commands that
  support it, reading the tracing/error_log file after an error will
  display more detailed information about what went wrong, if
  information is available.  The tracing/error_log file is a circular
  error log displaying a small number (currently, 8) of ftrace errors
  for the last (8) failed commands.

  The extended error information and usage takes the form shown in
  this example::

    # echo xxx > /sys/kernel/tracing/events/sched/sched_wakeup/trigger
    echo: write error: Invalid argument

    # cat /sys/kernel/tracing/error_log
    [ 5348.887237] location: error: Couldn't yyy: zzz
      Command: xxx
               ^
    [ 7517.023364] location: error: Bad rrr: sss
      Command: ppp qqq
                   ^

  To clear the error log, echo the empty string into it::

    # echo > /sys/kernel/tracing/error_log

tracefs만 사용하는 예제

921-927

뒤의 예제는 별도 사용자 공간 유틸리티 없이 `tracefs` 인터페이스만으로 추적기를 제어하는 전형적인 절차를 보여 준다.

Examples of using the tracer
----------------------------

Here are typical examples of using the tracers when controlling
them only with the tracefs interface (without using any
user-land utilities).

일반 trace 출력 형식

928-969

다음은 `trace` 파일에서 `function` 추적기를 읽은 출력 예다. 헤더와 열 배치를 포함한 원문을 그대로 보존한다.

  # tracer: function
  #
  # entries-in-buffer/entries-written: 140080/250280   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1977  [000] .... 17284.993652: sys_close <-system_call_fastpath
              bash-1977  [000] .... 17284.993653: __close_fd <-sys_close
              bash-1977  [000] .... 17284.993653: _raw_spin_lock <-__close_fd
              sshd-1974  [003] .... 17284.993653: __srcu_read_unlock <-fsnotify
              bash-1977  [000] .... 17284.993654: add_preempt_count <-_raw_spin_lock
              bash-1977  [000] ...1 17284.993655: _raw_spin_unlock <-__close_fd
              bash-1977  [000] ...1 17284.993656: sub_preempt_count <-_raw_spin_unlock
              bash-1977  [000] .... 17284.993657: filp_close <-__close_fd
              bash-1977  [000] .... 17284.993657: dnotify_flush <-filp_close
              sshd-1974  [003] .... 17284.993658: sys_select <-system_call_fastpath
              ....

`# tracer: function`은 결과를 만든 추적기 이름이다. `entries-in-buffer/entries-written`은 현재 버퍼에 남은 이벤트 수와 지금까지 쓴 총 이벤트 수다. 예제에서는 250,280개를 썼지만 140,080개만 남았으므로 버퍼가 차면서 110,200개를 잃었다.

각 레코드는 태스크 이름 `bash`, PID `1977`, 실행 CPU `000`, 지연 상태 문자, `<secs>.<usecs>` 형식 타임스탬프, 추적 함수 `sys_close`, 부모 호출 함수 `system_call_fastpath`를 표시한다. 타임스탬프는 함수 진입 시각이다.

function trace 열
의미
TASK-PIDbash-1977태스크 이름과 PID
CPU#[000]실행 CPU
상태....IRQ, resched, interrupt, preempt 상태
TIMESTAMP17284.993652함수 진입 시각
FUNCTIONsys_close <-system_call_fastpath추적 함수와 호출자

한 줄에서 태스크 문맥과 함수 호출 관계를 읽는다.

버퍼 손실 계산
entries-written 250280entries-in-buffer 140080 차감
entries-in-buffer 140080 차감lost 110200

누적 쓰기와 현재 보존 수의 차이가 덮어쓴 이벤트 수다.

Output format:
--------------

Here is an example of the output format of the file "trace"::

  # tracer: function
  #
  # entries-in-buffer/entries-written: 140080/250280   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1977  [000] .... 17284.993652: sys_close <-system_call_fastpath
              bash-1977  [000] .... 17284.993653: __close_fd <-sys_close
              bash-1977  [000] .... 17284.993653: _raw_spin_lock <-__close_fd
              sshd-1974  [003] .... 17284.993653: __srcu_read_unlock <-fsnotify
              bash-1977  [000] .... 17284.993654: add_preempt_count <-_raw_spin_lock
              bash-1977  [000] ...1 17284.993655: _raw_spin_unlock <-__close_fd
              bash-1977  [000] ...1 17284.993656: sub_preempt_count <-_raw_spin_unlock
              bash-1977  [000] .... 17284.993657: filp_close <-__close_fd
              bash-1977  [000] .... 17284.993657: dnotify_flush <-filp_close
              sshd-1974  [003] .... 17284.993658: sys_select <-system_call_fastpath
              ....

A header is printed with the tracer name that is represented by
the trace. In this case the tracer is "function". Then it shows the
number of events in the buffer as well as the total number of entries
that were written. The difference is the number of entries that were
lost due to the buffer filling up (250280 - 140080 = 110200 events
lost).

The header explains the content of the events. Task name "bash", the task
PID "1977", the CPU that it was running on "000", the latency format
(explained below), the timestamp in <secs>.<usecs> format, the
function name that was traced "sys_close" and the parent function that
called this function "system_call_fastpath". The timestamp is the time
at which the function was entered.

지연 추적 출력 형식

970-1087

`latency-format` 옵션을 켜거나 지연 추적기를 선택하면 `trace`는 지연 원인을 분석할 수 있도록 더 많은 정보를 표시한다. 다음 `irqsoff` 예제는 원문 헤더, 레코드, 역추적을 그대로 보존한다.

  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 259 us, #4/4, CPU#2 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: ps-6143 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: __lock_task_sighand
  #  => ended at:   _raw_spin_unlock_irqrestore
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
        ps-6143    2d...    0us!: trace_hardirqs_off <-__lock_task_sighand
        ps-6143    2d..1  259us+: trace_hardirqs_on <-_raw_spin_unlock_irqrestore
        ps-6143    2d..1  263us+: time_hardirqs_on <-_raw_spin_unlock_irqrestore
        ps-6143    2d..1  306us : <stack trace>
   => trace_hardirqs_on_caller
   => trace_hardirqs_on
   => _raw_spin_unlock_irqrestore
   => do_task_stat
   => proc_tgid_stat
   => proc_single_show
   => seq_read
   => vfs_read
   => sys_read
   => system_call_fastpath

예제는 인터럽트 비활성 시간을 재는 `irqsoff` 추적기와 고정 추적 형식 버전 1.1.5, 실행 커널 3.8을 보여 준다. 최대 지연은 259마이크로초이고 `#4/4`는 표시 레코드 수와 전체 레코드 수가 각각 4개임을 뜻한다.

`VP`, `KP`, `SP`, `HP`는 향후 사용을 위해 예약되어 항상 0이다. `#P:4`는 온라인 CPU가 4개라는 뜻이다. 지연 당시 실행 중인 태스크는 `ps`, PID는 6143이다.

`started at`의 `__lock_task_sighand`에서 인터럽트가 비활성화됐고, `ended at`의 `_raw_spin_unlock_irqrestore`에서 다시 활성화됐다. 뒤 레코드와 스택 역추적은 이 구간의 호출 경로를 보여 준다.

`cmd`는 프로세스 이름, `pid`는 프로세스 ID, `CPU#`은 실행 CPU다. `irqs-off` 열에서 `d`는 인터럽트 비활성, `.`은 그 밖의 상태다. `preempt-depth`는 `preempt_disabled` 중첩 깊이다.

`need-resched`의 `B`는 `TIF_NEED_RESCHED`, `PREEMPT_NEED_RESCHED`, `TIF_RESCHED_LAZY`가 모두 설정된 상태다. `N`은 앞의 두 플래그, `n`은 `TIF_NEED_RESCHED`만, `p`는 `PREEMPT_NEED_RESCHED`만 설정된 상태다.

같은 열의 `L`은 `PREEMPT_NEED_RESCHED`와 `TIF_RESCHED_LAZY`, `b`는 `TIF_NEED_RESCHED`와 `TIF_RESCHED_LAZY`, `l`은 `TIF_RESCHED_LAZY`만 설정됐음을 뜻하며 `.`은 어느 것도 아님을 뜻한다.

`hardirq/softirq` 열에서 `Z`는 hardirq 안에서 NMI 발생, `z`는 NMI 실행 중, `H`는 softirq 안에서 hardirq 발생, `h`는 hardirq 실행 중, `s`는 softirq 실행 중, `.`은 일반 문맥이다. 이 상태 정보는 주로 커널 개발자에게 의미가 있다.

`time`은 `latency-format`에서 추적 시작점에 대한 상대 시간이다. 옵션을 끈 일반 출력은 절대 타임스탬프를 사용한다.

`delay` 문자는 현재 레코드와 다음 레코드의 시간 차이를 눈에 띄게 표시한다. 현재 구현은 같은 CPU에 대한 차이만 보도록 개선할 필요가 있다. 지연 추적은 보통 마지막에 역추적을 붙여 지연 발생 위치를 쉽게 찾게 한다.

need-resched 문자
문자설정된 플래그
BTIF_NEED_RESCHED + PREEMPT_NEED_RESCHED + TIF_RESCHED_LAZY
NTIF_NEED_RESCHED + PREEMPT_NEED_RESCHED
nTIF_NEED_RESCHED
pPREEMPT_NEED_RESCHED
LPREEMPT_NEED_RESCHED + TIF_RESCHED_LAZY
bTIF_NEED_RESCHED + TIF_RESCHED_LAZY
lTIF_RESCHED_LAZY
.설정 없음

세 reschedule 플래그의 조합을 한 문자로 표시한다.

interrupt 문맥 문자
문자상태
Zhardirq 안에서 NMI 발생
zNMI 실행 중
Hsoftirq 안에서 hardirq 발생
hhardirq 실행 중
ssoftirq 실행 중
.일반 문맥

NMI, hardirq, softirq의 현재 또는 중첩 상태를 나타낸다.

delay 표식
문자시간 차이
$1초 초과
@100밀리초 초과
*10밀리초 초과
#1,000마이크로초 초과
!100마이크로초 초과
+10마이크로초 초과
공백10마이크로초 이하

인접 레코드 사이 지연 크기를 문자로 강조한다.

Latency trace format
--------------------

When the latency-format option is enabled or when one of the latency
tracers is set, the trace file gives somewhat more information to see
why a latency happened. Here is a typical trace::

  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 259 us, #4/4, CPU#2 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: ps-6143 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: __lock_task_sighand
  #  => ended at:   _raw_spin_unlock_irqrestore
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
        ps-6143    2d...    0us!: trace_hardirqs_off <-__lock_task_sighand
        ps-6143    2d..1  259us+: trace_hardirqs_on <-_raw_spin_unlock_irqrestore
        ps-6143    2d..1  263us+: time_hardirqs_on <-_raw_spin_unlock_irqrestore
        ps-6143    2d..1  306us : <stack trace>
   => trace_hardirqs_on_caller
   => trace_hardirqs_on
   => _raw_spin_unlock_irqrestore
   => do_task_stat
   => proc_tgid_stat
   => proc_single_show
   => seq_read
   => vfs_read
   => sys_read
   => system_call_fastpath


This shows that the current tracer is "irqsoff" tracing the time
for which interrupts were disabled. It gives the trace version (which
never changes) and the version of the kernel upon which this was executed on
(3.8). Then it displays the max latency in microseconds (259 us). The number
of trace entries displayed and the total number (both are four: #4/4).
VP, KP, SP, and HP are always zero and are reserved for later use.
#P is the number of online CPUs (#P:4).

The task is the process that was running when the latency
occurred. (ps pid: 6143).

The start and stop (the functions in which the interrupts were
disabled and enabled respectively) that caused the latencies:

  - __lock_task_sighand is where the interrupts were disabled.
  - _raw_spin_unlock_irqrestore is where they were enabled again.

The next lines after the header are the trace itself. The header
explains which is which.

  cmd: The name of the process in the trace.

  pid: The PID of that process.

  CPU#: The CPU which the process was running on.

  irqs-off: 'd' interrupts are disabled. '.' otherwise.

  need-resched:
	- 'B' all, TIF_NEED_RESCHED, PREEMPT_NEED_RESCHED and TIF_RESCHED_LAZY is set,
	- 'N' both TIF_NEED_RESCHED and PREEMPT_NEED_RESCHED is set,
	- 'n' only TIF_NEED_RESCHED is set,
	- 'p' only PREEMPT_NEED_RESCHED is set,
	- 'L' both PREEMPT_NEED_RESCHED and TIF_RESCHED_LAZY is set,
	- 'b' both TIF_NEED_RESCHED and TIF_RESCHED_LAZY is set,
	- 'l' only TIF_RESCHED_LAZY is set
	- '.' otherwise.

  hardirq/softirq:
	- 'Z' - NMI occurred inside a hardirq
	- 'z' - NMI is running
	- 'H' - hard irq occurred inside a softirq.
	- 'h' - hard irq is running
	- 's' - soft irq is running
	- '.' - normal context.

  preempt-depth: The level of preempt_disabled

The above is mostly meaningful for kernel developers.

  time:
	When the latency-format option is enabled, the trace file
	output includes a timestamp relative to the start of the
	trace. This differs from the output when latency-format
	is disabled, which includes an absolute timestamp.

  delay:
	This is just to help catch your eye a bit better. And
	needs to be fixed to be only relative to the same CPU.
	The marks are determined by the difference between this
	current trace and the next trace.

	  - '$' - greater than 1 second
	  - '@' - greater than 100 millisecond
	  - '*' - greater than 10 millisecond
	  - '#' - greater than 1000 microsecond
	  - '!' - greater than 100 microsecond
	  - '+' - greater than 10 microsecond
	  - ' ' - less than or equal to 10 microsecond.

  The rest is the same as the 'trace' file.

  Note, the latency tracers will usually end with a back trace
  to easily find where the latency occurred.

공통 trace_options

1088-1372

`trace_options` 파일과 `options/` 디렉터리는 추적 출력에 표시할 정보와 추적기 동작을 제어한다. 다음 출력에서 `no` 접두사가 붙은 항목은 비활성 상태다.

  cat trace_options
	print-parent
	nosym-offset
	nosym-addr
	noverbose
	noraw
	nohex
	nobin
	noblock
	nofields
	trace_printk
	annotate
	nouserstacktrace
	nosym-userobj
	noprintk-msg-only
	context-info
	nolatency-format
	record-cmd
	norecord-tgid
	overwrite
	nodisable_on_free
	irq-info
	markers
	noevent-fork
	function-trace
	nofunction-fork
	nodisplay-graph
	nostacktrace
	nobranch

옵션을 끌 때는 이름 앞에 `no`를 붙여 쓰고, 켤 때는 `no`를 뺀 이름을 쓴다.

  echo noprint-parent > trace_options
  echo sym-offset > trace_options

`print-parent`는 함수와 함께 호출자 함수를 표시한다. 끄면 현재 함수만 보인다.

	  print-parent:
	   bash-4000  [01]  1477.606694: simple_strtoul <-kstrtoul

	  noprint-parent:
	   bash-4000  [01]  1477.606694: simple_strtoul

`sym-offset`은 함수 이름 뒤에 함수 내부 오프셋과 전체 크기를 표시하고, `sym-addr`은 함수 주소도 표시한다.

	  sym-offset:
	   bash-4000  [01]  1477.606694: simple_strtoul+0x6/0xa0
	  sym-addr:
	   bash-4000  [01]  1477.606694: simple_strtoul <c0339346>

`verbose`는 `latency-format` 출력에 더 상세한 레코드 형식을 사용한다.

	    bash  4000 1 0 00000000 00010a95 [58127d26] 1720.415ms \
	    (+0.000ms): simple_strtoul (kstrtoul)

`raw`는 사용자 애플리케이션이 해석하기 좋은 원시 숫자를, `hex`는 같은 값을 16진수로, `bin`은 원시 바이너리 형식으로 출력한다. `fields`는 이벤트 필드의 선언된 타입에 따라 출력하므로 `hex`, `bin`, `raw`보다 내용을 파싱하기 좋다.

`block`을 설정하면 poll된 `trace_pipe` 읽기가 블록되지 않는다. `trace_printk`는 `trace_printk()`의 버퍼 기록을 끌 수 있다.

`trace_printk_dest`는 `trace_printk()`와 유사한 내부 추적 함수가 기록할 인스턴스를 선택한다. 동시에 인스턴스 하나만 가질 수 있어 새 인스턴스에 설정하면 이전 인스턴스의 플래그가 지워진다. 기본은 최상위 인스턴스이며 다른 인스턴스가 플래그를 지우면 최상위가 다시 받는다.

최상위 인스턴스는 기본 대상이므로 자신의 `trace_printk_dest`를 직접 지울 수 없다. 다른 인스턴스에 플래그를 설정할 때만 최상위에서 제거된다.

`copy_trace_marker`는 최상위 `/sys/kernel/tracing/trace_marker` 또는 `trace_marker_raw` 경로에 하드코딩해 쓰는 애플리케이션의 마커를 선택한 인스턴스에도 복사한다. 최상위 인스턴스에는 기본 설정돼 있으며, 이를 끄면 최상위 마커에 기록되지 않는다. 어느 인스턴스에도 설정하지 않으면 쓰기가 `ENODEV`로 실패한다.

`annotate`는 CPU별 버퍼가 서로 다른 시간 범위를 담아 출력에서 한 CPU만 실행된 것처럼 보이는 혼동을 줄인다. 새 CPU 버퍼의 시작 지점에 구분선을 표시한다.

			  <idle>-0     [001] dNs4 21169.031481: wake_up_idle_cpu <-add_timer_on
			  <idle>-0     [001] dNs4 21169.031482: _raw_spin_unlock_irqrestore <-add_timer_on
			  <idle>-0     [001] .Ns4 21169.031484: sub_preempt_count <-_raw_spin_unlock_irqrestore
		##### CPU 2 buffer started ####
			  <idle>-0     [002] .N.1 21169.031484: rcu_idle_exit <-cpu_idle
			  <idle>-0     [001] .Ns3 21169.031484: _raw_spin_unlock <-clocksource_watchdog
			  <idle>-0     [001] .Ns3 21169.031485: sub_preempt_count <-_raw_spin_unlock

`userstacktrace`는 각 추적 이벤트 뒤에 현재 사용자 공간 스레드의 스택을 기록한다. `sym-userobj`는 사용자 스택 주소가 속한 객체와 상대 주소를 찾아 표시하며 ASLR 사용 시 프로세스 종료 뒤에도 객체·파일·줄을 해석하는 데 유용하다. 조회는 `trace` 또는 `trace_pipe`를 읽을 때 수행된다.

		  a.out-1623  [000] 40874.465068: /root/a.out[+0x480] <-/root/a.out[+0
		  x494] <- /root/a.out[+0x4a8] <- /lib/libc-2.7.so[+0x1e1a6]

`printk-msg-only`는 `trace_bprintk()`나 `trace_bputs()`가 저장한 `trace_printk()` 레코드에서 인자 없이 형식 문자열만 표시한다. `context-info`는 comm, PID, 타임스탬프, CPU 등 문맥을 숨기고 이벤트 데이터만 보여 준다.

`latency-format`은 앞 절의 추가 지연 정보를 표시한다. `pause-on-trace`는 `trace` 파일을 읽기 위해 여는 동안 `tracing_on=0`처럼 링 버퍼 기록을 멈추고 파일을 닫으면 다시 켜 예전 `trace` 동작을 모사한다.

`hash-ptr`은 이벤트 printk 형식의 `%p`가 실제 주소 대신 해시된 포인터를 표시하게 한다. 추적 로그의 해시 값이 실제 값과 어떻게 대응되는지 조사할 때 사용할 수 있다.

`record-cmd`는 이벤트나 추적기를 켰을 때 `sched_switch` 훅으로 PID와 comm 캐시를 채운다. 이름이 필요 없고 PID만 중요하다면 끄면 오버헤드를 줄일 수 있다. `record-tgid`는 같은 방식으로 PID와 TGID 매핑 캐시를 채운다.

`overwrite`가 기본값 1이면 버퍼가 찼을 때 가장 오래된 이벤트를 버리고 새 이벤트로 덮는다. 0이면 새 이벤트를 버린다. 손실 수는 `per_cpu/cpu0/stats`의 `overrun`과 `dropped events`에서 확인한다.

`disable_on_free`는 `free_buffer`가 닫힐 때 `tracing_on`을 0으로 만들어 추적을 멈춘다. `irq-info`는 인터럽트 상태, preempt count, need-resched 정보를 표시한다. 끄면 다음처럼 상태 열이 없는 간결한 출력이 된다.

		# tracer: function
		#
		# entries-in-buffer/entries-written: 144405/9452052   #P:4
		#
		#           TASK-PID   CPU#      TIMESTAMP  FUNCTION
		#              | |       |          |         |
			  <idle>-0     [002]  23636.756054: ttwu_do_activate.constprop.89 <-try_to_wake_up
			  <idle>-0     [002]  23636.756054: activate_task <-ttwu_do_activate.constprop.89
			  <idle>-0     [002]  23636.756055: enqueue_task <-activate_task

`markers`를 설정하면 root가 `trace_marker`에 쓸 수 있고, 끄면 쓰기가 `EINVAL`로 실패한다.

`event-fork`는 `set_event_pid`의 태스크가 fork할 때 자식 PID를 자동 추가하고 태스크 종료 시 PID를 제거한다. `set_event_notrace_pid`에도 같은 규칙이 적용된다.

`function-trace`는 지연 추적기가 함수 추적을 함께 수행할지 제어하며 기본값은 활성이다. 끄면 지연 시험 중 함수 추적 오버헤드를 줄인다. `function-fork`는 `set_ftrace_pid`와 `set_ftrace_notrace_pid` 목록에 자식 PID 추가와 종료 PID 제거를 적용한다.

`display-graph`는 irqsoff, wakeup 같은 지연 추적기가 일반 함수 추적 대신 함수 그래프를 사용하게 한다. `stacktrace`는 모든 추적 이벤트 뒤에 스택을 기록한다. `branch`는 현재 추적기와 branch 추적을 동시에 켜며, `nop`과 함께 켜면 branch 추적기만 선택한 것과 같다.

출력 표현 옵션
옵션효과
print-parent호출자 함수 표시
sym-offset / sym-addr함수 오프셋 / 주소 표시
verbose상세 latency 레코드
raw / hex / bin원시 숫자 / 16진수 / 바이너리
fields선언 타입에 맞춘 이벤트 필드
context-info문맥을 숨기고 이벤트 데이터만 표시
irq-infoIRQ, preempt, resched 상태 표시

같은 레코드를 어떤 수준과 형식으로 표시할지 결정한다.

기록과 수명 주기 옵션
옵션효과
trace_printk / trace_printk_dest내부 printk 기록과 대상 인스턴스
copy_trace_marker최상위 마커 쓰기를 인스턴스에 복사
pause-on-tracetrace 파일을 연 동안 기록 중지
overwrite오래된 이벤트 또는 새 이벤트 폐기 선택
disable_on_freefree_buffer 닫힘 때 추적 중지
markerstrace_marker 쓰기 허용

링 버퍼 기록 대상, 가득 찬 버퍼, 파일 수명을 제어한다.

PID 상속 옵션
event-forkset_event_pid와 notrace 목록
function-forkset_ftrace_pid와 notrace 목록
부모 fork자식 PID 자동 추가
태스크 exitPID 자동 제거

fork와 exit에 맞춰 함수 또는 이벤트 PID 목록을 유지한다.

trace_options
-------------

The trace_options file (or the options directory) is used to control
what gets printed in the trace output, or manipulate the tracers.
To see what is available, simply cat the file::

  cat trace_options
	print-parent
	nosym-offset
	nosym-addr
	noverbose
	noraw
	nohex
	nobin
	noblock
	nofields
	trace_printk
	annotate
	nouserstacktrace
	nosym-userobj
	noprintk-msg-only
	context-info
	nolatency-format
	record-cmd
	norecord-tgid
	overwrite
	nodisable_on_free
	irq-info
	markers
	noevent-fork
	function-trace
	nofunction-fork
	nodisplay-graph
	nostacktrace
	nobranch

To disable one of the options, echo in the option prepended with
"no"::

  echo noprint-parent > trace_options

To enable an option, leave off the "no"::

  echo sym-offset > trace_options

Here are the available options:

  print-parent
	On function traces, display the calling (parent)
	function as well as the function being traced.
	::

	  print-parent:
	   bash-4000  [01]  1477.606694: simple_strtoul <-kstrtoul

	  noprint-parent:
	   bash-4000  [01]  1477.606694: simple_strtoul


  sym-offset
	Display not only the function name, but also the
	offset in the function. For example, instead of
	seeing just "ktime_get", you will see
	"ktime_get+0xb/0x20".
	::

	  sym-offset:
	   bash-4000  [01]  1477.606694: simple_strtoul+0x6/0xa0

  sym-addr
	This will also display the function address as well
	as the function name.
	::

	  sym-addr:
	   bash-4000  [01]  1477.606694: simple_strtoul <c0339346>

  verbose
	This deals with the trace file when the
        latency-format option is enabled.
	::

	    bash  4000 1 0 00000000 00010a95 [58127d26] 1720.415ms \
	    (+0.000ms): simple_strtoul (kstrtoul)

  raw
	This will display raw numbers. This option is best for
	use with user applications that can translate the raw
	numbers better than having it done in the kernel.

  hex
	Similar to raw, but the numbers will be in a hexadecimal format.

  bin
	This will print out the formats in raw binary.

  block
	When set, reading trace_pipe will not block when polled.

  fields
	Print the fields as described by their types. This is a better
	option than using hex, bin or raw, as it gives a better parsing
	of the content of the event.

  trace_printk
	Can disable trace_printk() from writing into the buffer.

  trace_printk_dest
	Set to have trace_printk() and similar internal tracing functions
	write into this instance. Note, only one trace instance can have
	this set. By setting this flag, it clears the trace_printk_dest flag
	of the instance that had it set previously. By default, the top
	level trace has this set, and will get it set again if another
	instance has it set then clears it.

	This flag cannot be cleared by the top level instance, as it is the
	default instance. The only way the top level instance has this flag
	cleared, is by it being set in another instance.

  copy_trace_marker
	If there are applications that hard code writing into the top level
	trace_marker file (/sys/kernel/tracing/trace_marker or trace_marker_raw),
	and the tooling would like it to go into an instance, this option can
	be used. Create an instance and set this option, and then all writes
	into the top level trace_marker file will also be redirected into this
	instance.

	Note, by default this option is set for the top level instance. If it
	is disabled, then writes to the trace_marker or trace_marker_raw files
	will not be written into the top level file. If no instance has this
	option set, then a write will error with the errno of ENODEV.

  annotate
	It is sometimes confusing when the CPU buffers are full
	and one CPU buffer had a lot of events recently, thus
	a shorter time frame, were another CPU may have only had
	a few events, which lets it have older events. When
	the trace is reported, it shows the oldest events first,
	and it may look like only one CPU ran (the one with the
	oldest events). When the annotate option is set, it will
	display when a new CPU buffer started::

			  <idle>-0     [001] dNs4 21169.031481: wake_up_idle_cpu <-add_timer_on
			  <idle>-0     [001] dNs4 21169.031482: _raw_spin_unlock_irqrestore <-add_timer_on
			  <idle>-0     [001] .Ns4 21169.031484: sub_preempt_count <-_raw_spin_unlock_irqrestore
		##### CPU 2 buffer started ####
			  <idle>-0     [002] .N.1 21169.031484: rcu_idle_exit <-cpu_idle
			  <idle>-0     [001] .Ns3 21169.031484: _raw_spin_unlock <-clocksource_watchdog
			  <idle>-0     [001] .Ns3 21169.031485: sub_preempt_count <-_raw_spin_unlock

  userstacktrace
	This option changes the trace. It records a
	stacktrace of the current user space thread after
	each trace event.

  sym-userobj
	when user stacktrace are enabled, look up which
	object the address belongs to, and print a
	relative address. This is especially useful when
	ASLR is on, otherwise you don't get a chance to
	resolve the address to object/file/line after
	the app is no longer running

	The lookup is performed when you read
	trace,trace_pipe. Example::

		  a.out-1623  [000] 40874.465068: /root/a.out[+0x480] <-/root/a.out[+0
		  x494] <- /root/a.out[+0x4a8] <- /lib/libc-2.7.so[+0x1e1a6]


  printk-msg-only
	When set, trace_printk()s will only show the format
	and not their parameters (if trace_bprintk() or
	trace_bputs() was used to save the trace_printk()).

  context-info
	Show only the event data. Hides the comm, PID,
	timestamp, CPU, and other useful data.

  latency-format
	This option changes the trace output. When it is enabled,
	the trace displays additional information about the
	latency, as described in "Latency trace format".

  pause-on-trace
	When set, opening the trace file for read, will pause
	writing to the ring buffer (as if tracing_on was set to zero).
	This simulates the original behavior of the trace file.
	When the file is closed, tracing will be enabled again.

  hash-ptr
        When set, "%p" in the event printk format displays the
        hashed pointer value instead of real address.
        This will be useful if you want to find out which hashed
        value is corresponding to the real value in trace log.

  record-cmd
	When any event or tracer is enabled, a hook is enabled
	in the sched_switch trace point to fill comm cache
	with mapped pids and comms. But this may cause some
	overhead, and if you only care about pids, and not the
	name of the task, disabling this option can lower the
	impact of tracing. See "saved_cmdlines".

  record-tgid
	When any event or tracer is enabled, a hook is enabled
	in the sched_switch trace point to fill the cache of
	mapped Thread Group IDs (TGID) mapping to pids. See
	"saved_tgids".

  overwrite
	This controls what happens when the trace buffer is
	full. If "1" (default), the oldest events are
	discarded and overwritten. If "0", then the newest
	events are discarded.
	(see per_cpu/cpu0/stats for overrun and dropped)

  disable_on_free
	When the free_buffer is closed, tracing will
	stop (tracing_on set to 0).

  irq-info
	Shows the interrupt, preempt count, need resched data.
	When disabled, the trace looks like::

		# tracer: function
		#
		# entries-in-buffer/entries-written: 144405/9452052   #P:4
		#
		#           TASK-PID   CPU#      TIMESTAMP  FUNCTION
		#              | |       |          |         |
			  <idle>-0     [002]  23636.756054: ttwu_do_activate.constprop.89 <-try_to_wake_up
			  <idle>-0     [002]  23636.756054: activate_task <-ttwu_do_activate.constprop.89
			  <idle>-0     [002]  23636.756055: enqueue_task <-activate_task


  markers
	When set, the trace_marker is writable (only by root).
	When disabled, the trace_marker will error with EINVAL
	on write.

  event-fork
	When set, tasks with PIDs listed in set_event_pid will have
	the PIDs of their children added to set_event_pid when those
	tasks fork. Also, when tasks with PIDs in set_event_pid exit,
	their PIDs will be removed from the file.

        This affects PIDs listed in set_event_notrace_pid as well.

  function-trace
	The latency tracers will enable function tracing
	if this option is enabled (default it is). When
	it is disabled, the latency tracers do not trace
	functions. This keeps the overhead of the tracer down
	when performing latency tests.

  function-fork
	When set, tasks with PIDs listed in set_ftrace_pid will
	have the PIDs of their children added to set_ftrace_pid
	when those tasks fork. Also, when tasks with PIDs in
	set_ftrace_pid exit, their PIDs will be removed from the
	file.

        This affects PIDs in set_ftrace_notrace_pid as well.

  display-graph
	When set, the latency tracers (irqsoff, wakeup, etc) will
	use function graph tracing instead of function tracing.

  stacktrace
	When set, a stack trace is recorded after any trace event
	is recorded.

  branch
	Enable branch tracing with the tracer. This enables branch
	tracer along with the currently set tracer. Enabling this
	with the "nop" tracer is the same as just enabling the
	"branch" tracer.

.. tip:: Some tracers have their own options. They only appear in this
       file when the tracer is active. They always appear in the
       options directory.

추적기별 옵션

1373-1467

일부 추적기는 자체 옵션을 제공한다. 해당 추적기가 활성일 때만 `trace_options`에 나타나지만 `options/` 디렉터리에는 항상 보인다.

function 추적기의 `func_stack_trace`는 기록한 모든 함수 뒤에 스택을 추가한다. 시스템 성능을 심각하게 떨어뜨릴 수 있으므로 먼저 `set_ftrace_filter`로 함수를 제한해야 하며, 함수 필터를 지우기 전에 반드시 이 옵션을 꺼야 한다.

function_graph 추적기는 출력 형식이 달라 별도 옵션을 갖는다. `funcgraph-overrun`은 태스크마다 고정된 함수 그래프 스택보다 호출 깊이가 커져 추적하지 못한 함수 수를 각 함수 뒤에 표시한다.

`funcgraph-cpu`는 이벤트가 발생한 CPU 번호를, `funcgraph-overhead`는 함수 시간이 일정 임계값을 넘을 때 앞에서 설명한 delay 표식을 표시한다.

`funcgraph-proc`는 각 줄에 프로세스 명령을 표시한다. 함수 그래프는 기본적으로 문맥 전환에서 태스크가 들어오고 나갈 때만 명령을 표시하므로 이 옵션이 줄별 명령 표시를 추가한다.

`funcgraph-duration`은 함수 반환 지점에 함수 안에서 보낸 시간을 마이크로초로 표시하고, `funcgraph-abstime`은 각 줄에 타임스탬프를 표시한다. `funcgraph-irqs`를 끄면 인터럽트 안에서 실행된 함수를 추적하지 않는다.

`funcgraph-tail`은 반환 이벤트에 대응 함수 이름을 표시한다. 기본은 꺼져 있어 함수 반환을 닫는 중괄호 `}`만 출력한다.

`funcgraph-retval`은 각 함수 반환값을 등호 `=` 뒤에 표시한다. `funcgraph-retval-hex`를 켜면 항상 16진수로 출력한다. 끈 상태에서는 오류 코드만 signed decimal로, 나머지는 16진수로 표시하며 두 옵션 모두 기본적으로 꺼져 있다.

`sleep-time`은 태스크가 함수 실행 중 스케줄 아웃된 시간도 함수 호출 시간에 포함한다. `graph-time`은 함수 그래프 기반 함수 프로파일러에서 중첩 호출 함수에 소비한 시간을 포함한다. 끄면 해당 함수 자체가 실행한 시간만 보고한다.

blk 추적기의 `blk_classic`은 더 간결한 고전 출력 형식을 사용한다.

function_graph 표시 옵션
옵션표시/계산 효과
funcgraph-overrun고정 그래프 스택 초과로 놓친 함수 수
funcgraph-cpuCPU 번호
funcgraph-overhead긴 함수의 delay 표식
funcgraph-proc모든 줄의 프로세스 명령
funcgraph-duration함수 실행 시간
funcgraph-abstime각 줄의 절대 시각
funcgraph-irqs인터럽트 내부 함수 포함 여부
funcgraph-tail반환 이벤트의 함수 이름
funcgraph-retval함수 반환값
funcgraph-retval-hex반환값을 항상 16진수로 표시

호출 그래프의 문맥과 반환 정보를 선택한다.

함수 시간 계산 옵션
옵션활성 시 포함비활성 시
sleep-time태스크가 스케줄 아웃된 시간실행 가능/실행 시간 중심
graph-time호출한 하위 함수 시간함수 자체 실행 시간만

스케줄 아웃 시간과 중첩 호출 시간을 포함할지 구분한다.

Here are the per tracer options:

Options for function tracer:

  func_stack_trace
	When set, a stack trace is recorded after every
	function that is recorded. NOTE! Limit the functions
	that are recorded before enabling this, with
	"set_ftrace_filter" otherwise the system performance
	will be critically degraded. Remember to disable
	this option before clearing the function filter.

Options for function_graph tracer:

 Since the function_graph tracer has a slightly different output
 it has its own options to control what is displayed.

  funcgraph-overrun
	When set, the "overrun" of the graph stack is
	displayed after each function traced. The
	overrun, is when the stack depth of the calls
	is greater than what is reserved for each task.
	Each task has a fixed array of functions to
	trace in the call graph. If the depth of the
	calls exceeds that, the function is not traced.
	The overrun is the number of functions missed
	due to exceeding this array.

  funcgraph-cpu
	When set, the CPU number of the CPU where the trace
	occurred is displayed.

  funcgraph-overhead
	When set, if the function takes longer than
	A certain amount, then a delay marker is
	displayed. See "delay" above, under the
	header description.

  funcgraph-proc
	Unlike other tracers, the process' command line
	is not displayed by default, but instead only
	when a task is traced in and out during a context
	switch. Enabling this options has the command
	of each process displayed at every line.

  funcgraph-duration
	At the end of each function (the return)
	the duration of the amount of time in the
	function is displayed in microseconds.

  funcgraph-abstime
	When set, the timestamp is displayed at each line.

  funcgraph-irqs
	When disabled, functions that happen inside an
	interrupt will not be traced.

  funcgraph-tail
	When set, the return event will include the function
	that it represents. By default this is off, and
	only a closing curly bracket "}" is displayed for
	the return of a function.

  funcgraph-retval
	When set, the return value of each traced function
	will be printed after an equal sign "=". By default
	this is off.

  funcgraph-retval-hex
	When set, the return value will always be printed
	in hexadecimal format. If the option is not set and
	the return value is an error code, it will be printed
	in signed decimal format; otherwise it will also be
	printed in hexadecimal format. By default, this option
	is off.

  sleep-time
	When running function graph tracer, to include
	the time a task schedules out in its function.
	When enabled, it will account time the task has been
	scheduled out as part of the function call.

  graph-time
	When running function profiler with function graph tracer,
	to include the time to call nested functions. When this is
	not set, the time reported for the function will only
	include the time the function itself executed for, not the
	time for functions that it called.

Options for blk tracer:

  blk_classic
	Shows a more minimalistic output.

irqsoff 추적기

1468-1670

인터럽트가 비활성화된 동안 CPU는 NMI와 SMI를 제외한 외부 사건에 반응할 수 없다. 타이머 인터럽트가 발생하지 못하고 새 마우스 사건도 커널에 전달되지 않으므로, 시스템의 반응 시간이 그만큼 늦어진다.

`irqsoff` 추적기는 인터럽트가 비활성화된 시간을 측정한다. 새 최장 지연이 발견되면 그 지연 지점까지의 추적을 저장하고, 이전에 저장한 최장 지연 추적은 새 기록으로 교체한다.

최대 지연 기록은 `echo 0 > tracing_max_latency`로 초기화한다. 다음 예는 `options/function-trace`를 끈 상태에서 `irqsoff`를 선택하고 추적을 실행한 결과다.

  # echo 0 > options/function-trace
  # echo irqsoff > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # ls -ltr
  [...]
  # echo 0 > tracing_on
  # cat trace
  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 16 us, #4/4, CPU#0 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: swapper/0-0 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: run_timer_softirq
  #  => ended at:   run_timer_softirq
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       0d.s2    0us+: _raw_spin_lock_irq <-run_timer_softirq
    <idle>-0       0dNs3   17us : _raw_spin_unlock_irq <-run_timer_softirq
    <idle>-0       0dNs3   17us+: trace_hardirqs_on <-run_timer_softirq
    <idle>-0       0dNs3   25us : <stack trace>
   => _raw_spin_unlock_irq
   => run_timer_softirq
   => __do_softirq
   => call_softirq
   => do_softirq
   => irq_exit
   => smp_apic_timer_interrupt
   => apic_timer_interrupt
   => rcu_idle_exit
   => cpu_idle
   => rest_init
   => start_kernel
   => x86_64_start_reservations
   => x86_64_start_kernel

이 예에서는 `run_timer_softirq` 안의 `_raw_spin_lock_irq`가 인터럽트를 비활성화했고, 측정된 최장 지연은 16마이크로초다. 표시된 스택의 타임스탬프가 25us인 이유는 최대 지연을 기록한 시점과 그 지연을 만든 함수를 기록한 시점 사이에 시계 값이 증가했기 때문이다.

위 예는 함수 추적을 사용하지 않았다. `options/function-trace`를 켜면 인터럽트 비활성화 구간에서 호출된 함수까지 모두 기록되므로 출력이 훨씬 길어진다.

 with echo 1 > options/function-trace

  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 71 us, #168/168, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: bash-2042 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: ata_scsi_queuecmd
  #  => ended at:   ata_scsi_queuecmd
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
      bash-2042    3d...    0us : _raw_spin_lock_irqsave <-ata_scsi_queuecmd
      bash-2042    3d...    0us : add_preempt_count <-_raw_spin_lock_irqsave
      bash-2042    3d..1    1us : ata_scsi_find_dev <-ata_scsi_queuecmd
      bash-2042    3d..1    1us : __ata_scsi_find_dev <-ata_scsi_find_dev
      bash-2042    3d..1    2us : ata_find_dev.part.14 <-__ata_scsi_find_dev
      bash-2042    3d..1    2us : ata_qc_new_init <-__ata_scsi_queuecmd
      bash-2042    3d..1    3us : ata_sg_init <-__ata_scsi_queuecmd
      bash-2042    3d..1    4us : ata_scsi_rw_xlat <-__ata_scsi_queuecmd
      bash-2042    3d..1    4us : ata_build_rw_tf <-ata_scsi_rw_xlat
  [...]
      bash-2042    3d..1   67us : delay_tsc <-__delay
      bash-2042    3d..1   67us : add_preempt_count <-delay_tsc
      bash-2042    3d..2   67us : sub_preempt_count <-delay_tsc
      bash-2042    3d..1   67us : add_preempt_count <-delay_tsc
      bash-2042    3d..2   68us : sub_preempt_count <-delay_tsc
      bash-2042    3d..1   68us+: ata_bmdma_start <-ata_bmdma_qc_issue
      bash-2042    3d..1   71us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
      bash-2042    3d..1   71us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
      bash-2042    3d..1   72us+: trace_hardirqs_on <-ata_scsi_queuecmd
      bash-2042    3d..1  120us : <stack trace>
   => _raw_spin_unlock_irqrestore
   => ata_scsi_queuecmd
   => scsi_dispatch_cmd
   => scsi_request_fn
   => __blk_run_queue_uncond
   => __blk_run_queue
   => blk_queue_bio
   => submit_bio_noacct
   => submit_bio
   => submit_bh
   => __ext3_get_inode_loc
   => ext3_iget
   => ext3_lookup
   => lookup_real
   => __lookup_hash
   => walk_component
   => lookup_last
   => path_lookupat
   => filename_lookup
   => user_path_at_empty
   => user_path_at
   => vfs_fstatat
   => vfs_stat
   => sys_newstat
   => system_call_fastpath

두 번째 예의 지연은 71마이크로초이며 `ata_scsi_queuecmd` 안에서 호출된 함수들을 함께 보여 준다. 함수 추적 자체의 오버헤드가 실제 지연 시간을 늘릴 수 있다는 점은 감안해야 하지만, 어느 호출 경로가 지연을 만들었는지 조사할 때 매우 유용하다.

함수 목록 대신 함수 그래프 형식을 원하면 `options/display-graph`를 켠다. 그러면 호출의 중첩과 각 함수의 실행 시간이 C 코드와 비슷한 구조로 나타난다.

 with echo 1 > options/display-graph

  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 4.20.0-rc6+
  # --------------------------------------------------------------------
  # latency: 3751 us, #274/274, CPU#0 | (M:desktop VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: bash-1507 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: free_debug_processing
  #  => ended at:   return_to_handler
  #
  #
  #                                       _-----=> irqs-off
  #                                      / _----=> need-resched
  #                                     | / _---=> hardirq/softirq
  #                                     || / _--=> preempt-depth
  #                                     ||| /
  #   REL TIME      CPU  TASK/PID       ||||     DURATION                  FUNCTION CALLS
  #      |          |     |    |        ||||      |   |                     |   |   |   |
          0 us |   0)   bash-1507    |  d... |   0.000 us    |  _raw_spin_lock_irqsave();
          0 us |   0)   bash-1507    |  d..1 |   0.378 us    |    do_raw_spin_trylock();
          1 us |   0)   bash-1507    |  d..2 |               |    set_track() {
          2 us |   0)   bash-1507    |  d..2 |               |      save_stack_trace() {
          2 us |   0)   bash-1507    |  d..2 |               |        __save_stack_trace() {
          3 us |   0)   bash-1507    |  d..2 |               |          __unwind_start() {
          3 us |   0)   bash-1507    |  d..2 |               |            get_stack_info() {
          3 us |   0)   bash-1507    |  d..2 |   0.351 us    |              in_task_stack();
          4 us |   0)   bash-1507    |  d..2 |   1.107 us    |            }
  [...]
       3750 us |   0)   bash-1507    |  d..1 |   0.516 us    |      do_raw_spin_unlock();
       3750 us |   0)   bash-1507    |  d..1 |   0.000 us    |  _raw_spin_unlock_irqrestore();
       3764 us |   0)   bash-1507    |  d..1 |   0.000 us    |  tracer_hardirqs_on();
      bash-1507    0d..1 3792us : <stack trace>
   => free_debug_processing
   => __slab_free
   => kmem_cache_free
   => vm_area_free
   => remove_vma
   => exit_mmap
   => mmput
   => begin_new_exec
   => load_elf_binary
   => search_binary_handler
   => __do_execve_file.isra.32
   => __x64_sys_execve
   => do_syscall_64
   => entry_SYSCALL_64_after_hwframe

함수 그래프 예는 `free_debug_processing`에서 시작해 `return_to_handler`에서 끝난 3751마이크로초 구간을 보여 준다. 중첩된 호출과 마지막 스택 추적을 함께 읽으면 IRQ를 막은 호출 경로를 계층적으로 좁힐 수 있다.

irqsoff 예제 비교
설정예제 지연주요 정보주의점
function-trace 끔16 usIRQ 비활성화 시작·종료와 스택호출된 모든 함수는 보이지 않음
function-trace 켬71 us구간 안의 전체 함수 호출추적 오버헤드가 지연을 늘릴 수 있음
display-graph 켬3751 us중첩 호출과 함수별 실행 시간출력이 길지만 호출 구조가 선명함

함수 추적 옵션에 따라 얻는 정보와 측정 오버헤드가 달라진다.

irqsoff 최장 지연 보존
IRQ 비활성화경과 시간 측정
경과 시간 측정tracing_max_latency와 비교
기존 최대값 이하기존 추적 유지
새 최대값이전 추적 폐기
이전 추적 폐기새 지연 경로 저장

새 최대값이 나올 때마다 보존 추적을 교체한다.

16 us 예제 판독
관찰값의미
0us+ `_raw_spin_lock_irq``run_timer_softirq`에서 IRQ 비활성화 시작
17us `_raw_spin_unlock_irq`IRQ 비활성화 구간 종료
latency: 16 us최대 지연을 판정한 시점의 측정값
25us `<stack trace>`함수·스택 기록 시점까지 시계가 더 진행된 값

측정값과 마지막 스택 타임스탬프가 다른 이유를 구분한다.

irqsoff
-------

When interrupts are disabled, the CPU can not react to any other
external event (besides NMIs and SMIs). This prevents the timer
interrupt from triggering or the mouse interrupt from letting
the kernel know of a new mouse event. The result is a latency
with the reaction time.

The irqsoff tracer tracks the time for which interrupts are
disabled. When a new maximum latency is hit, the tracer saves
the trace leading up to that latency point so that every time a
new maximum is reached, the old saved trace is discarded and the
new trace is saved.

To reset the maximum, echo 0 into tracing_max_latency. Here is
an example::

  # echo 0 > options/function-trace
  # echo irqsoff > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # ls -ltr
  [...]
  # echo 0 > tracing_on
  # cat trace
  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 16 us, #4/4, CPU#0 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: swapper/0-0 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: run_timer_softirq
  #  => ended at:   run_timer_softirq
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       0d.s2    0us+: _raw_spin_lock_irq <-run_timer_softirq
    <idle>-0       0dNs3   17us : _raw_spin_unlock_irq <-run_timer_softirq
    <idle>-0       0dNs3   17us+: trace_hardirqs_on <-run_timer_softirq
    <idle>-0       0dNs3   25us : <stack trace>
   => _raw_spin_unlock_irq
   => run_timer_softirq
   => __do_softirq
   => call_softirq
   => do_softirq
   => irq_exit
   => smp_apic_timer_interrupt
   => apic_timer_interrupt
   => rcu_idle_exit
   => cpu_idle
   => rest_init
   => start_kernel
   => x86_64_start_reservations
   => x86_64_start_kernel

Here we see that we had a latency of 16 microseconds (which is
very good). The _raw_spin_lock_irq in run_timer_softirq disabled
interrupts. The difference between the 16 and the displayed
timestamp 25us occurred because the clock was incremented
between the time of recording the max latency and the time of
recording the function that had that latency.

Note the above example had function-trace not set. If we set
function-trace, we get a much larger output::

 with echo 1 > options/function-trace

  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 71 us, #168/168, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: bash-2042 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: ata_scsi_queuecmd
  #  => ended at:   ata_scsi_queuecmd
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
      bash-2042    3d...    0us : _raw_spin_lock_irqsave <-ata_scsi_queuecmd
      bash-2042    3d...    0us : add_preempt_count <-_raw_spin_lock_irqsave
      bash-2042    3d..1    1us : ata_scsi_find_dev <-ata_scsi_queuecmd
      bash-2042    3d..1    1us : __ata_scsi_find_dev <-ata_scsi_find_dev
      bash-2042    3d..1    2us : ata_find_dev.part.14 <-__ata_scsi_find_dev
      bash-2042    3d..1    2us : ata_qc_new_init <-__ata_scsi_queuecmd
      bash-2042    3d..1    3us : ata_sg_init <-__ata_scsi_queuecmd
      bash-2042    3d..1    4us : ata_scsi_rw_xlat <-__ata_scsi_queuecmd
      bash-2042    3d..1    4us : ata_build_rw_tf <-ata_scsi_rw_xlat
  [...]
      bash-2042    3d..1   67us : delay_tsc <-__delay
      bash-2042    3d..1   67us : add_preempt_count <-delay_tsc
      bash-2042    3d..2   67us : sub_preempt_count <-delay_tsc
      bash-2042    3d..1   67us : add_preempt_count <-delay_tsc
      bash-2042    3d..2   68us : sub_preempt_count <-delay_tsc
      bash-2042    3d..1   68us+: ata_bmdma_start <-ata_bmdma_qc_issue
      bash-2042    3d..1   71us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
      bash-2042    3d..1   71us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
      bash-2042    3d..1   72us+: trace_hardirqs_on <-ata_scsi_queuecmd
      bash-2042    3d..1  120us : <stack trace>
   => _raw_spin_unlock_irqrestore
   => ata_scsi_queuecmd
   => scsi_dispatch_cmd
   => scsi_request_fn
   => __blk_run_queue_uncond
   => __blk_run_queue
   => blk_queue_bio
   => submit_bio_noacct
   => submit_bio
   => submit_bh
   => __ext3_get_inode_loc
   => ext3_iget
   => ext3_lookup
   => lookup_real
   => __lookup_hash
   => walk_component
   => lookup_last
   => path_lookupat
   => filename_lookup
   => user_path_at_empty
   => user_path_at
   => vfs_fstatat
   => vfs_stat
   => sys_newstat
   => system_call_fastpath


Here we traced a 71 microsecond latency. But we also see all the
functions that were called during that time. Note that by
enabling function tracing, we incur an added overhead. This
overhead may extend the latency times. But nevertheless, this
trace has provided some very helpful debugging information.

If we prefer function graph output instead of function, we can set
display-graph option::

 with echo 1 > options/display-graph

  # tracer: irqsoff
  #
  # irqsoff latency trace v1.1.5 on 4.20.0-rc6+
  # --------------------------------------------------------------------
  # latency: 3751 us, #274/274, CPU#0 | (M:desktop VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: bash-1507 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: free_debug_processing
  #  => ended at:   return_to_handler
  #
  #
  #                                       _-----=> irqs-off
  #                                      / _----=> need-resched
  #                                     | / _---=> hardirq/softirq
  #                                     || / _--=> preempt-depth
  #                                     ||| /
  #   REL TIME      CPU  TASK/PID       ||||     DURATION                  FUNCTION CALLS
  #      |          |     |    |        ||||      |   |                     |   |   |   |
          0 us |   0)   bash-1507    |  d... |   0.000 us    |  _raw_spin_lock_irqsave();
          0 us |   0)   bash-1507    |  d..1 |   0.378 us    |    do_raw_spin_trylock();
          1 us |   0)   bash-1507    |  d..2 |               |    set_track() {
          2 us |   0)   bash-1507    |  d..2 |               |      save_stack_trace() {
          2 us |   0)   bash-1507    |  d..2 |               |        __save_stack_trace() {
          3 us |   0)   bash-1507    |  d..2 |               |          __unwind_start() {
          3 us |   0)   bash-1507    |  d..2 |               |            get_stack_info() {
          3 us |   0)   bash-1507    |  d..2 |   0.351 us    |              in_task_stack();
          4 us |   0)   bash-1507    |  d..2 |   1.107 us    |            }
  [...]
       3750 us |   0)   bash-1507    |  d..1 |   0.516 us    |      do_raw_spin_unlock();
       3750 us |   0)   bash-1507    |  d..1 |   0.000 us    |  _raw_spin_unlock_irqrestore();
       3764 us |   0)   bash-1507    |  d..1 |   0.000 us    |  tracer_hardirqs_on();
      bash-1507    0d..1 3792us : <stack trace>
   => free_debug_processing
   => __slab_free
   => kmem_cache_free
   => vm_area_free
   => remove_vma
   => exit_mmap
   => mmput
   => begin_new_exec
   => load_elf_binary
   => search_binary_handler
   => __do_execve_file.isra.32
   => __x64_sys_execve
   => do_syscall_64
   => entry_SYSCALL_64_after_hwframe

preemptoff 추적기

1671-1801

preemption이 비활성화되어도 인터럽트는 받을 수 있지만 현재 태스크를 선점할 수는 없다. 따라서 더 높은 우선순위 태스크가 준비되어도 preemption이 다시 활성화될 때까지 낮은 우선순위 태스크를 밀어내지 못하고 기다린다.

`preemptoff` 추적기는 preemption을 비활성화한 위치와 그 상태가 지속된 최대 시간을 기록한다. 선택, 시작·정지, `tracing_max_latency` 초기화 방식은 `irqsoff`와 거의 같다.

  # echo 0 > options/function-trace
  # echo preemptoff > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # ls -ltr
  [...]
  # echo 0 > tracing_on
  # cat trace
  # tracer: preemptoff
  #
  # preemptoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 46 us, #4/4, CPU#1 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sshd-1991 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: do_IRQ
  #  => ended at:   do_IRQ
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
      sshd-1991    1d.h.    0us+: irq_enter <-do_IRQ
      sshd-1991    1d..1   46us : irq_exit <-do_IRQ
      sshd-1991    1d..1   47us+: trace_preempt_on <-do_IRQ
      sshd-1991    1d..1   52us : <stack trace>
   => sub_preempt_count
   => irq_exit
   => do_IRQ
   => ret_from_intr

첫 예는 `do_IRQ` 구간에서 46마이크로초 동안 preemption이 비활성화되었음을 보여 준다. 상태 문자의 `h`는 하드 인터럽트 문맥에 들어왔음을 뜻하고, 종료 지점의 `trace_preempt_on`은 preemption이 다시 활성화되었음을 나타낸다.

진입과 이탈 기록에 보이는 `d`는 그 두 기록 시점에 IRQ도 비활성화되어 있었다는 뜻이다. 이 정보만으로는 그 사이 전체 구간에서도 IRQ가 계속 꺼져 있었는지, 또는 구간 종료 직후 언제 켜졌는지 확정할 수 없다.

  # tracer: preemptoff
  #
  # preemptoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 83 us, #241/241, CPU#1 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: bash-1994 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: wake_up_new_task
  #  => ended at:   task_rq_unlock
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
      bash-1994    1d..1    0us : _raw_spin_lock_irqsave <-wake_up_new_task
      bash-1994    1d..1    0us : select_task_rq_fair <-select_task_rq
      bash-1994    1d..1    1us : __rcu_read_lock <-select_task_rq_fair
      bash-1994    1d..1    1us : source_load <-select_task_rq_fair
      bash-1994    1d..1    1us : source_load <-select_task_rq_fair
  [...]
      bash-1994    1d..1   12us : irq_enter <-smp_apic_timer_interrupt
      bash-1994    1d..1   12us : rcu_irq_enter <-irq_enter
      bash-1994    1d..1   13us : add_preempt_count <-irq_enter
      bash-1994    1d.h1   13us : exit_idle <-smp_apic_timer_interrupt
      bash-1994    1d.h1   13us : hrtimer_interrupt <-smp_apic_timer_interrupt
      bash-1994    1d.h1   13us : _raw_spin_lock <-hrtimer_interrupt
      bash-1994    1d.h1   14us : add_preempt_count <-_raw_spin_lock
      bash-1994    1d.h2   14us : ktime_get_update_offsets <-hrtimer_interrupt
  [...]
      bash-1994    1d.h1   35us : lapic_next_event <-clockevents_program_event
      bash-1994    1d.h1   35us : irq_exit <-smp_apic_timer_interrupt
      bash-1994    1d.h1   36us : sub_preempt_count <-irq_exit
      bash-1994    1d..2   36us : do_softirq <-irq_exit
      bash-1994    1d..2   36us : __do_softirq <-call_softirq
      bash-1994    1d..2   36us : __local_bh_disable <-__do_softirq
      bash-1994    1d.s2   37us : add_preempt_count <-_raw_spin_lock_irq
      bash-1994    1d.s3   38us : _raw_spin_unlock <-run_timer_softirq
      bash-1994    1d.s3   39us : sub_preempt_count <-_raw_spin_unlock
      bash-1994    1d.s2   39us : call_timer_fn <-run_timer_softirq
  [...]
      bash-1994    1dNs2   81us : cpu_needs_another_gp <-rcu_process_callbacks
      bash-1994    1dNs2   82us : __local_bh_enable <-__do_softirq
      bash-1994    1dNs2   82us : sub_preempt_count <-__local_bh_enable
      bash-1994    1dN.2   82us : idle_cpu <-irq_exit
      bash-1994    1dN.2   83us : rcu_irq_exit <-irq_exit
      bash-1994    1dN.2   83us : sub_preempt_count <-irq_exit
      bash-1994    1.N.1   84us : _raw_spin_unlock_irqrestore <-task_rq_unlock
      bash-1994    1.N.1   84us+: trace_preempt_on <-task_rq_unlock
      bash-1994    1.N.1  104us : <stack trace>
   => sub_preempt_count
   => _raw_spin_unlock_irqrestore
   => task_rq_unlock
   => wake_up_new_task
   => do_fork
   => sys_clone
   => stub_clone

두 번째 예는 `function-trace`를 켠 `preemptoff` 결과다. `wake_up_new_task`에서 `task_rq_unlock`까지 83마이크로초가 걸렸고, 그 사이 타이머 인터럽트에 진입한 뒤 softirq까지 실행한 호출 경로가 드러난다.

`irq_enter`와 상태 문자 `h`는 실제 인터럽트 진입을 권위 있게 표시한다. 그 앞의 함수 이름만 보면 인터럽트 문맥이 아닌 것처럼 오해할 수 있으므로, 함수 이름의 정황보다 각 행의 IRQ·softirq 상태 표기를 우선해 해석해야 한다.

preemptoff 상태 문자 해석
표기예제에서의 의미판독 한계
h하드 인터럽트 문맥인터럽트 진입·이탈 구간과 함께 읽어야 함
ssoftirq 문맥preemption은 계속 비활성화될 수 있음
d해당 기록 시점에 IRQ 비활성화두 기록 사이 전체 상태를 단독으로 증명하지 않음
preempt-depthpreemption 중첩 깊이0으로 돌아가는 지점이 구간 종료

preemption 비활성화 구간 안에서도 IRQ 문맥은 바뀔 수 있다.

preemptoff 83 us 구간
wake_up_new_taskpreemption 비활성화
preemption 비활성화타이머 hardirq 진입
타이머 hardirq 진입irq_exit
irq_exitsoftirq 실행
softirq 실행task_rq_unlock
task_rq_unlocktrace_preempt_on

preemption을 끈 채 인터럽트와 softirq가 끼어드는 흐름이다.

preemptoff
----------

When preemption is disabled, we may be able to receive
interrupts but the task cannot be preempted and a higher
priority task must wait for preemption to be enabled again
before it can preempt a lower priority task.

The preemptoff tracer traces the places that disable preemption.
Like the irqsoff tracer, it records the maximum latency for
which preemption was disabled. The control of preemptoff tracer
is much like the irqsoff tracer.
::

  # echo 0 > options/function-trace
  # echo preemptoff > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # ls -ltr
  [...]
  # echo 0 > tracing_on
  # cat trace
  # tracer: preemptoff
  #
  # preemptoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 46 us, #4/4, CPU#1 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sshd-1991 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: do_IRQ
  #  => ended at:   do_IRQ
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
      sshd-1991    1d.h.    0us+: irq_enter <-do_IRQ
      sshd-1991    1d..1   46us : irq_exit <-do_IRQ
      sshd-1991    1d..1   47us+: trace_preempt_on <-do_IRQ
      sshd-1991    1d..1   52us : <stack trace>
   => sub_preempt_count
   => irq_exit
   => do_IRQ
   => ret_from_intr


This has some more changes. Preemption was disabled when an
interrupt came in (notice the 'h'), and was enabled on exit.
But we also see that interrupts have been disabled when entering
the preempt off section and leaving it (the 'd'). We do not know if
interrupts were enabled in the mean time or shortly after this
was over.
::

  # tracer: preemptoff
  #
  # preemptoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 83 us, #241/241, CPU#1 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: bash-1994 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: wake_up_new_task
  #  => ended at:   task_rq_unlock
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
      bash-1994    1d..1    0us : _raw_spin_lock_irqsave <-wake_up_new_task
      bash-1994    1d..1    0us : select_task_rq_fair <-select_task_rq
      bash-1994    1d..1    1us : __rcu_read_lock <-select_task_rq_fair
      bash-1994    1d..1    1us : source_load <-select_task_rq_fair
      bash-1994    1d..1    1us : source_load <-select_task_rq_fair
  [...]
      bash-1994    1d..1   12us : irq_enter <-smp_apic_timer_interrupt
      bash-1994    1d..1   12us : rcu_irq_enter <-irq_enter
      bash-1994    1d..1   13us : add_preempt_count <-irq_enter
      bash-1994    1d.h1   13us : exit_idle <-smp_apic_timer_interrupt
      bash-1994    1d.h1   13us : hrtimer_interrupt <-smp_apic_timer_interrupt
      bash-1994    1d.h1   13us : _raw_spin_lock <-hrtimer_interrupt
      bash-1994    1d.h1   14us : add_preempt_count <-_raw_spin_lock
      bash-1994    1d.h2   14us : ktime_get_update_offsets <-hrtimer_interrupt
  [...]
      bash-1994    1d.h1   35us : lapic_next_event <-clockevents_program_event
      bash-1994    1d.h1   35us : irq_exit <-smp_apic_timer_interrupt
      bash-1994    1d.h1   36us : sub_preempt_count <-irq_exit
      bash-1994    1d..2   36us : do_softirq <-irq_exit
      bash-1994    1d..2   36us : __do_softirq <-call_softirq
      bash-1994    1d..2   36us : __local_bh_disable <-__do_softirq
      bash-1994    1d.s2   37us : add_preempt_count <-_raw_spin_lock_irq
      bash-1994    1d.s3   38us : _raw_spin_unlock <-run_timer_softirq
      bash-1994    1d.s3   39us : sub_preempt_count <-_raw_spin_unlock
      bash-1994    1d.s2   39us : call_timer_fn <-run_timer_softirq
  [...]
      bash-1994    1dNs2   81us : cpu_needs_another_gp <-rcu_process_callbacks
      bash-1994    1dNs2   82us : __local_bh_enable <-__do_softirq
      bash-1994    1dNs2   82us : sub_preempt_count <-__local_bh_enable
      bash-1994    1dN.2   82us : idle_cpu <-irq_exit
      bash-1994    1dN.2   83us : rcu_irq_exit <-irq_exit
      bash-1994    1dN.2   83us : sub_preempt_count <-irq_exit
      bash-1994    1.N.1   84us : _raw_spin_unlock_irqrestore <-task_rq_unlock
      bash-1994    1.N.1   84us+: trace_preempt_on <-task_rq_unlock
      bash-1994    1.N.1  104us : <stack trace>
   => sub_preempt_count
   => _raw_spin_unlock_irqrestore
   => task_rq_unlock
   => wake_up_new_task
   => do_fork
   => sys_clone
   => stub_clone


The above is an example of the preemptoff trace with
function-trace set. Here we see that interrupts were not disabled
the entire time. The irq_enter code lets us know that we entered
an interrupt 'h'. Before that, the functions being traced still
show that it is not in an interrupt, but we can see from the
functions themselves that this is not the case.

preemptirqsoff 추적기

1802-1996

IRQ 비활성화가 가장 긴 위치와 preemption 비활성화가 가장 긴 위치를 각각 찾는 것만으로는 부족할 때가 있다. 실제 스케줄 불가능 시간은 두 상태가 겹치거나 차례로 이어지는 전체 구간이기 때문이다.

다음 의사 코드는 IRQ만 꺼진 구간, IRQ와 preemption이 모두 꺼진 구간, preemption만 꺼진 구간이 연속되는 상황을 보여 준다.

    local_irq_disable();
    call_function_with_irqs_off();
    preempt_disable();
    call_function_with_irqs_and_preemption_off();
    local_irq_enable();
    call_function_with_preemption_off();
    preempt_enable();

`irqsoff`는 `call_function_with_irqs_off()`와 `call_function_with_irqs_and_preemption_off()`의 합계만 기록한다. 반대로 `preemptoff`는 `call_function_with_irqs_and_preemption_off()`와 `call_function_with_preemption_off()`의 합계만 기록한다.

어느 하나만으로는 IRQ 또는 preemption 중 적어도 하나가 비활성화되어 스케줄할 수 없는 전체 시간을 얻지 못한다. `preemptirqsoff`는 첫 `local_irq_disable()`부터 마지막 `preempt_enable()`까지 이어지는 이 합성 구간을 측정한다.

사용 절차는 `irqsoff`, `preemptoff`와 같다. 다음 예는 함수 추적을 끄고 `preemptirqsoff`를 선택한 뒤 최대 지연을 초기화해 얻은 결과다.

  # echo 0 > options/function-trace
  # echo preemptirqsoff > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # ls -ltr
  [...]
  # echo 0 > tracing_on
  # cat trace
  # tracer: preemptirqsoff
  #
  # preemptirqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 100 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: ls-2230 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: ata_scsi_queuecmd
  #  => ended at:   ata_scsi_queuecmd
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
        ls-2230    3d...    0us+: _raw_spin_lock_irqsave <-ata_scsi_queuecmd
        ls-2230    3...1  100us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
        ls-2230    3...1  101us+: trace_preempt_on <-ata_scsi_queuecmd
        ls-2230    3...1  111us : <stack trace>
   => sub_preempt_count
   => _raw_spin_unlock_irqrestore
   => ata_scsi_queuecmd
   => scsi_dispatch_cmd
   => scsi_request_fn
   => __blk_run_queue_uncond
   => __blk_run_queue
   => blk_queue_bio
   => submit_bio_noacct
   => submit_bio
   => submit_bh
   => ext3_bread
   => ext3_dir_bread
   => htree_dirblock_to_tree
   => ext3_htree_fill_tree
   => ext3_readdir
   => vfs_readdir
   => sys_getdents
   => system_call_fastpath

첫 예는 `ata_scsi_queuecmd` 안에서 스케줄할 수 없었던 100마이크로초 구간을 보여 준다. x86에서는 어셈블리 코드가 IRQ를 비활성화할 때 `trace_hardirqs_off_thunk`가 호출된다.

함수 추적이 꺼져 있으므로 두 preemption 지점 사이에서 IRQ가 잠시 활성화되었는지는 알 수 없다. 다만 추적 시작 시점에는 preemption이 활성화되어 있었다는 사실은 상태 정보에서 확인할 수 있다.

다음은 `function-trace`를 켠 결과로, 호출 경로와 중첩 인터럽트를 자세히 보여 준다.

  # tracer: preemptirqsoff
  #
  # preemptirqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 161 us, #339/339, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: ls-2269 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: schedule
  #  => ended at:   mutex_unlock
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
  kworker/-59      3...1    0us : __schedule <-schedule
  kworker/-59      3d..1    0us : rcu_preempt_qs <-rcu_note_context_switch
  kworker/-59      3d..1    1us : add_preempt_count <-_raw_spin_lock_irq
  kworker/-59      3d..2    1us : deactivate_task <-__schedule
  kworker/-59      3d..2    1us : dequeue_task <-deactivate_task
  kworker/-59      3d..2    2us : update_rq_clock <-dequeue_task
  kworker/-59      3d..2    2us : dequeue_task_fair <-dequeue_task
  kworker/-59      3d..2    2us : update_curr <-dequeue_task_fair
  kworker/-59      3d..2    2us : update_min_vruntime <-update_curr
  kworker/-59      3d..2    3us : cpuacct_charge <-update_curr
  kworker/-59      3d..2    3us : __rcu_read_lock <-cpuacct_charge
  kworker/-59      3d..2    3us : __rcu_read_unlock <-cpuacct_charge
  kworker/-59      3d..2    3us : update_cfs_rq_blocked_load <-dequeue_task_fair
  kworker/-59      3d..2    4us : clear_buddies <-dequeue_task_fair
  kworker/-59      3d..2    4us : account_entity_dequeue <-dequeue_task_fair
  kworker/-59      3d..2    4us : update_min_vruntime <-dequeue_task_fair
  kworker/-59      3d..2    4us : update_cfs_shares <-dequeue_task_fair
  kworker/-59      3d..2    5us : hrtick_update <-dequeue_task_fair
  kworker/-59      3d..2    5us : wq_worker_sleeping <-__schedule
  kworker/-59      3d..2    5us : kthread_data <-wq_worker_sleeping
  kworker/-59      3d..2    5us : put_prev_task_fair <-__schedule
  kworker/-59      3d..2    6us : pick_next_task_fair <-pick_next_task
  kworker/-59      3d..2    6us : clear_buddies <-pick_next_task_fair
  kworker/-59      3d..2    6us : set_next_entity <-pick_next_task_fair
  kworker/-59      3d..2    6us : update_stats_wait_end <-set_next_entity
        ls-2269    3d..2    7us : finish_task_switch <-__schedule
        ls-2269    3d..2    7us : _raw_spin_unlock_irq <-finish_task_switch
        ls-2269    3d..2    8us : do_IRQ <-ret_from_intr
        ls-2269    3d..2    8us : irq_enter <-do_IRQ
        ls-2269    3d..2    8us : rcu_irq_enter <-irq_enter
        ls-2269    3d..2    9us : add_preempt_count <-irq_enter
        ls-2269    3d.h2    9us : exit_idle <-do_IRQ
  [...]
        ls-2269    3d.h3   20us : sub_preempt_count <-_raw_spin_unlock
        ls-2269    3d.h2   20us : irq_exit <-do_IRQ
        ls-2269    3d.h2   21us : sub_preempt_count <-irq_exit
        ls-2269    3d..3   21us : do_softirq <-irq_exit
        ls-2269    3d..3   21us : __do_softirq <-call_softirq
        ls-2269    3d..3   21us+: __local_bh_disable <-__do_softirq
        ls-2269    3d.s4   29us : sub_preempt_count <-_local_bh_enable_ip
        ls-2269    3d.s5   29us : sub_preempt_count <-_local_bh_enable_ip
        ls-2269    3d.s5   31us : do_IRQ <-ret_from_intr
        ls-2269    3d.s5   31us : irq_enter <-do_IRQ
        ls-2269    3d.s5   31us : rcu_irq_enter <-irq_enter
  [...]
        ls-2269    3d.s5   31us : rcu_irq_enter <-irq_enter
        ls-2269    3d.s5   32us : add_preempt_count <-irq_enter
        ls-2269    3d.H5   32us : exit_idle <-do_IRQ
        ls-2269    3d.H5   32us : handle_irq <-do_IRQ
        ls-2269    3d.H5   32us : irq_to_desc <-handle_irq
        ls-2269    3d.H5   33us : handle_fasteoi_irq <-handle_irq
  [...]
        ls-2269    3d.s5  158us : _raw_spin_unlock_irqrestore <-rtl8139_poll
        ls-2269    3d.s3  158us : net_rps_action_and_irq_enable.isra.65 <-net_rx_action
        ls-2269    3d.s3  159us : __local_bh_enable <-__do_softirq
        ls-2269    3d.s3  159us : sub_preempt_count <-__local_bh_enable
        ls-2269    3d..3  159us : idle_cpu <-irq_exit
        ls-2269    3d..3  159us : rcu_irq_exit <-irq_exit
        ls-2269    3d..3  160us : sub_preempt_count <-irq_exit
        ls-2269    3d...  161us : __mutex_unlock_slowpath <-mutex_unlock
        ls-2269    3d...  162us+: trace_hardirqs_on <-mutex_unlock
        ls-2269    3d...  186us : <stack trace>
   => __mutex_unlock_slowpath
   => mutex_unlock
   => process_output
   => n_tty_write
   => tty_write
   => vfs_write
   => sys_write
   => system_call_fastpath

161마이크로초 예는 `kworker`가 `schedule`에서 스케줄 아웃되고 `ls`가 실행을 이어받는 것으로 시작한다. `ls`가 runqueue 잠금을 풀며 IRQ를 활성화했지만 preemption은 아직 비활성화된 상태였고, 바로 그때 인터럽트가 발생했다.

첫 인터럽트가 끝난 뒤 softirq가 실행되었고, 그 softirq 도중 또 다른 하드 인터럽트가 들어왔다. softirq 안에서 하드 인터럽트가 실행되는 중첩 상태는 상태 문자 `H`로 표시된다.

세 지연 추적기의 측정 구간
구간IRQ 상태preemption 상태irqsoffpreemptoffpreemptirqsoff
`call_function_with_irqs_off()`포함제외포함
`call_function_with_irqs_and_preemption_off()`포함포함포함
`call_function_with_preemption_off()`제외포함포함

IRQ와 preemption 상태가 이어질 때 각 추적기가 포함하는 범위를 비교한다.

스케줄 불가능 전체 구간
local_irq_disable()IRQ만 비활성화
IRQ만 비활성화preempt_disable()
preempt_disable()IRQ와 preemption 모두 비활성화
IRQ와 preemption 모두 비활성화local_irq_enable()
local_irq_enable()preemption만 비활성화
preemption만 비활성화preempt_enable()
preempt_enable()측정 종료·스케줄 가능

둘 중 하나라도 비활성화된 동안 `preemptirqsoff` 측정은 계속된다.

161 us 예제의 문맥 전환
단계대표 함수/표기상태
태스크 교체`schedule` → `finish_task_switch`kworker에서 ls로 전환
잠금 해제`_raw_spin_unlock_irq`IRQ는 켜지지만 preemption은 계속 꺼짐
첫 hardirq`do_IRQ` / `h`인터럽트 처리
softirq`__do_softirq` / `s`첫 인터럽트 종료 뒤 실행
중첩 hardirq`do_IRQ` / `H`softirq 안에서 다시 인터럽트 발생
구간 종료`mutex_unlock`preemption·IRQ 제약 해제

함수 추적 출력에서 태스크·IRQ·softirq의 중첩 순서를 읽는다.

preemptirqsoff
--------------

Knowing the locations that have interrupts disabled or
preemption disabled for the longest times is helpful. But
sometimes we would like to know when either preemption and/or
interrupts are disabled.

Consider the following code::

    local_irq_disable();
    call_function_with_irqs_off();
    preempt_disable();
    call_function_with_irqs_and_preemption_off();
    local_irq_enable();
    call_function_with_preemption_off();
    preempt_enable();

The irqsoff tracer will record the total length of
call_function_with_irqs_off() and
call_function_with_irqs_and_preemption_off().

The preemptoff tracer will record the total length of
call_function_with_irqs_and_preemption_off() and
call_function_with_preemption_off().

But neither will trace the time that interrupts and/or
preemption is disabled. This total time is the time that we can
not schedule. To record this time, use the preemptirqsoff
tracer.

Again, using this trace is much like the irqsoff and preemptoff
tracers.
::

  # echo 0 > options/function-trace
  # echo preemptirqsoff > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # ls -ltr
  [...]
  # echo 0 > tracing_on
  # cat trace
  # tracer: preemptirqsoff
  #
  # preemptirqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 100 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: ls-2230 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: ata_scsi_queuecmd
  #  => ended at:   ata_scsi_queuecmd
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
        ls-2230    3d...    0us+: _raw_spin_lock_irqsave <-ata_scsi_queuecmd
        ls-2230    3...1  100us : _raw_spin_unlock_irqrestore <-ata_scsi_queuecmd
        ls-2230    3...1  101us+: trace_preempt_on <-ata_scsi_queuecmd
        ls-2230    3...1  111us : <stack trace>
   => sub_preempt_count
   => _raw_spin_unlock_irqrestore
   => ata_scsi_queuecmd
   => scsi_dispatch_cmd
   => scsi_request_fn
   => __blk_run_queue_uncond
   => __blk_run_queue
   => blk_queue_bio
   => submit_bio_noacct
   => submit_bio
   => submit_bh
   => ext3_bread
   => ext3_dir_bread
   => htree_dirblock_to_tree
   => ext3_htree_fill_tree
   => ext3_readdir
   => vfs_readdir
   => sys_getdents
   => system_call_fastpath


The trace_hardirqs_off_thunk is called from assembly on x86 when
interrupts are disabled in the assembly code. Without the
function tracing, we do not know if interrupts were enabled
within the preemption points. We do see that it started with
preemption enabled.

Here is a trace with function-trace set::

  # tracer: preemptirqsoff
  #
  # preemptirqsoff latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 161 us, #339/339, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: ls-2269 (uid:0 nice:0 policy:0 rt_prio:0)
  #    -----------------
  #  => started at: schedule
  #  => ended at:   mutex_unlock
  #
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
  kworker/-59      3...1    0us : __schedule <-schedule
  kworker/-59      3d..1    0us : rcu_preempt_qs <-rcu_note_context_switch
  kworker/-59      3d..1    1us : add_preempt_count <-_raw_spin_lock_irq
  kworker/-59      3d..2    1us : deactivate_task <-__schedule
  kworker/-59      3d..2    1us : dequeue_task <-deactivate_task
  kworker/-59      3d..2    2us : update_rq_clock <-dequeue_task
  kworker/-59      3d..2    2us : dequeue_task_fair <-dequeue_task
  kworker/-59      3d..2    2us : update_curr <-dequeue_task_fair
  kworker/-59      3d..2    2us : update_min_vruntime <-update_curr
  kworker/-59      3d..2    3us : cpuacct_charge <-update_curr
  kworker/-59      3d..2    3us : __rcu_read_lock <-cpuacct_charge
  kworker/-59      3d..2    3us : __rcu_read_unlock <-cpuacct_charge
  kworker/-59      3d..2    3us : update_cfs_rq_blocked_load <-dequeue_task_fair
  kworker/-59      3d..2    4us : clear_buddies <-dequeue_task_fair
  kworker/-59      3d..2    4us : account_entity_dequeue <-dequeue_task_fair
  kworker/-59      3d..2    4us : update_min_vruntime <-dequeue_task_fair
  kworker/-59      3d..2    4us : update_cfs_shares <-dequeue_task_fair
  kworker/-59      3d..2    5us : hrtick_update <-dequeue_task_fair
  kworker/-59      3d..2    5us : wq_worker_sleeping <-__schedule
  kworker/-59      3d..2    5us : kthread_data <-wq_worker_sleeping
  kworker/-59      3d..2    5us : put_prev_task_fair <-__schedule
  kworker/-59      3d..2    6us : pick_next_task_fair <-pick_next_task
  kworker/-59      3d..2    6us : clear_buddies <-pick_next_task_fair
  kworker/-59      3d..2    6us : set_next_entity <-pick_next_task_fair
  kworker/-59      3d..2    6us : update_stats_wait_end <-set_next_entity
        ls-2269    3d..2    7us : finish_task_switch <-__schedule
        ls-2269    3d..2    7us : _raw_spin_unlock_irq <-finish_task_switch
        ls-2269    3d..2    8us : do_IRQ <-ret_from_intr
        ls-2269    3d..2    8us : irq_enter <-do_IRQ
        ls-2269    3d..2    8us : rcu_irq_enter <-irq_enter
        ls-2269    3d..2    9us : add_preempt_count <-irq_enter
        ls-2269    3d.h2    9us : exit_idle <-do_IRQ
  [...]
        ls-2269    3d.h3   20us : sub_preempt_count <-_raw_spin_unlock
        ls-2269    3d.h2   20us : irq_exit <-do_IRQ
        ls-2269    3d.h2   21us : sub_preempt_count <-irq_exit
        ls-2269    3d..3   21us : do_softirq <-irq_exit
        ls-2269    3d..3   21us : __do_softirq <-call_softirq
        ls-2269    3d..3   21us+: __local_bh_disable <-__do_softirq
        ls-2269    3d.s4   29us : sub_preempt_count <-_local_bh_enable_ip
        ls-2269    3d.s5   29us : sub_preempt_count <-_local_bh_enable_ip
        ls-2269    3d.s5   31us : do_IRQ <-ret_from_intr
        ls-2269    3d.s5   31us : irq_enter <-do_IRQ
        ls-2269    3d.s5   31us : rcu_irq_enter <-irq_enter
  [...]
        ls-2269    3d.s5   31us : rcu_irq_enter <-irq_enter
        ls-2269    3d.s5   32us : add_preempt_count <-irq_enter
        ls-2269    3d.H5   32us : exit_idle <-do_IRQ
        ls-2269    3d.H5   32us : handle_irq <-do_IRQ
        ls-2269    3d.H5   32us : irq_to_desc <-handle_irq
        ls-2269    3d.H5   33us : handle_fasteoi_irq <-handle_irq
  [...]
        ls-2269    3d.s5  158us : _raw_spin_unlock_irqrestore <-rtl8139_poll
        ls-2269    3d.s3  158us : net_rps_action_and_irq_enable.isra.65 <-net_rx_action
        ls-2269    3d.s3  159us : __local_bh_enable <-__do_softirq
        ls-2269    3d.s3  159us : sub_preempt_count <-__local_bh_enable
        ls-2269    3d..3  159us : idle_cpu <-irq_exit
        ls-2269    3d..3  159us : rcu_irq_exit <-irq_exit
        ls-2269    3d..3  160us : sub_preempt_count <-irq_exit
        ls-2269    3d...  161us : __mutex_unlock_slowpath <-mutex_unlock
        ls-2269    3d...  162us+: trace_hardirqs_on <-mutex_unlock
        ls-2269    3d...  186us : <stack trace>
   => __mutex_unlock_slowpath
   => mutex_unlock
   => process_output
   => n_tty_write
   => tty_write
   => vfs_write
   => sys_write
   => system_call_fastpath

This is an interesting trace. It started with kworker running and
scheduling out and ls taking over. But as soon as ls released the
rq lock and enabled interrupts (but not preemption) an interrupt
triggered. When the interrupt finished, it started running softirqs.
But while the softirq was running, another interrupt triggered.
When an interrupt is running inside a softirq, the annotation is 'H'.

wakeup 추적기

1997-2044

자주 조사하는 지연 가운데 하나는 잠에서 깨어난 태스크가 실제로 CPU에서 실행되기까지 걸리는 시간이다. 비실시간 태스크의 이 시간은 임의로 길어질 수 있지만, 시스템 스케줄링 동작을 이해하는 데 여전히 흥미로운 지표다.

다음 예는 함수 추적을 끄고 `wakeup` 추적기를 실행한다. `chrt -f 5 sleep 1`로 태스크를 깨운 뒤 최대 지연 추적을 읽는다.

  # echo 0 > options/function-trace
  # echo wakeup > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # chrt -f 5 sleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: wakeup
  #
  # wakeup latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 15 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: kworker/3:1H-312 (uid:0 nice:-20 policy:0 rt_prio:0)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       3dNs7    0us :      0:120:R   + [003]   312:100:R kworker/3:1H
    <idle>-0       3dNs7    1us+: ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       3d..3   15us : __schedule <-schedule
    <idle>-0       3d..3   15us :      0:120:R ==> [003]   312:100:R kworker/3:1H

`wakeup`은 일상적인 모든 wakeup을 기록하지 않고 시스템에서 가장 우선순위가 높은 태스크만 추적한다. 예에서는 nice 값 -20인 `kworker/3:1H`가 깨어난 시점부터 실제로 실행될 때까지 15마이크로초가 걸렸다.

일반 태스크보다 실시간 태스크의 최악 wakeup 지연이 더 중요한 경우에는 다음 절의 `wakeup_rt` 추적기를 사용한다.

wakeup 예제의 네 기록
시각기록의미
0us`0:120:R + [003] 312:100:R`CPU 3의 kworker를 깨우기 시작
1us+`ttwu_do_activate`깨어난 태스크를 활성화
15us`__schedule`스케줄러가 태스크 전환 직전까지 도달
15us`0:120:R ==> [003] 312:100:R`kworker가 CPU 3에서 실행 대상으로 선택

태스크가 준비 큐에 들어가 실제로 선택되기까지의 핵심 지점을 읽는다.

wakeup 지연 구간
태스크 wakeuptry_to_wake_up
try_to_wake_upttwu_do_activate
ttwu_do_activate스케줄러 진입
스케줄러 진입깨어난 태스크 선택

깨우기 요청부터 스케줄 인 직전까지를 측정한다.

wakeup
------

One common case that people are interested in tracing is the
time it takes for a task that is woken to actually wake up.
Now for non Real-Time tasks, this can be arbitrary. But tracing
it nonetheless can be interesting. 

Without function tracing::

  # echo 0 > options/function-trace
  # echo wakeup > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # chrt -f 5 sleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: wakeup
  #
  # wakeup latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 15 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: kworker/3:1H-312 (uid:0 nice:-20 policy:0 rt_prio:0)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       3dNs7    0us :      0:120:R   + [003]   312:100:R kworker/3:1H
    <idle>-0       3dNs7    1us+: ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       3d..3   15us : __schedule <-schedule
    <idle>-0       3d..3   15us :      0:120:R ==> [003]   312:100:R kworker/3:1H

The tracer only traces the highest priority task in the system
to avoid tracing the normal circumstances. Here we see that
the kworker with a nice priority of -20 (not very nice), took
just 15 microseconds from the time it woke up, to the time it
ran.

Non Real-Time tasks are not that interesting. A more interesting
trace is to concentrate only on Real-Time tasks.

wakeup_rt 추적기

2045-2241

실시간 환경에서는 깨어난 최우선순위 RT 태스크가 실제로 실행될 때까지의 wakeup 시간, 즉 스케줄 지연을 아는 것이 매우 중요하다. 비RT 태스크의 스케줄 지연도 중요하지만 그 경우에는 평균값이 더 유용하고, LatencyTop 같은 도구가 이런 측정에 더 적합하다.

실시간 시스템은 평균보다 최악 지연에 관심을 둔다. 대부분 매우 빠르더라도 드물게 큰 지연이 생기는 스케줄러는 RT 태스크에 적합하지 않다. `wakeup_rt`는 RT 태스크 wakeup의 최악값만 보존하도록 설계되었다.

비RT 태스크는 예측하기 어려운 wakeup으로 단 하나뿐인 최악값 기록을 덮어쓸 수 있으므로 이 추적기에서 제외한다. 일반 태스크까지 보고 싶다면 `wakeup` 추적기를 별도로 사용한다.

RT 태스크를 대상으로 하므로 앞선 예의 `ls` 대신 `chrt`로 우선순위를 바꾼 `sleep 1`을 실행한다. 다음 명령은 함수 추적을 끄고 FIFO 우선순위 5인 태스크의 wakeup을 측정한다.

  # echo 0 > options/function-trace
  # echo wakeup_rt > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # chrt -f 5 sleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: wakeup
  #
  # tracer: wakeup_rt
  #
  # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 5 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sleep-2389 (uid:0 nice:0 policy:1 rt_prio:5)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       3d.h4    0us :      0:120:R   + [003]  2389: 94:R sleep
    <idle>-0       3d.h4    1us+: ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       3d..3    5us : __schedule <-schedule
    <idle>-0       3d..3    5us :      0:120:R ==> [003]  2389: 94:R sleep

유휴 시스템의 예에서는 태스크 전환까지 5마이크로초가 걸렸다. 스케줄러의 tracepoint가 실제 문맥 전환보다 앞에 있으므로, 기록 대상 태스크가 스케줄 인되기 직전에 추적을 멈춘다. 스케줄러 끝에 새 표식이 추가되면 이 종료 기준은 달라질 수 있다.

기록 대상은 PID 2389의 `sleep`이며 `rt_prio`는 5다. 이 값은 사용자 공간 RT 우선순위다. 정책 값 1은 `SCHED_FIFO`, 2는 `SCHED_RR`을 뜻한다.

추적 데이터의 우선순위는 사용자 값이 아니라 커널 내부 값 `99 - rtprio`를 표시한다.

  <idle>-0       3d..3    5us :      0:120:R ==> [003]  2389: 94:R sleep

`0:120:R`은 nice 0인 idle 태스크가 실행 상태 `R`에 있었음을 뜻한다. 일반 우선순위는 `120 - nice` 관계로 읽는다. 스케줄 인된 `sleep`은 `2389:94:R`로 표시되며, 내부 RT 우선순위 94는 `99 - 5`의 결과다.

같은 측정을 `chrt -r 5`와 함수 추적을 켠 상태로 수행하려면 먼저 `options/function-trace`에 1을 쓴다. 다음 출력은 문서가 제시하는 전체 호출 기록이다.

  echo 1 > options/function-trace

  # tracer: wakeup_rt
  #
  # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 29 us, #85/85, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sleep-2448 (uid:0 nice:0 policy:1 rt_prio:5)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       3d.h4    1us+:      0:120:R   + [003]  2448: 94:R sleep
    <idle>-0       3d.h4    2us : ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       3d.h3    3us : check_preempt_curr <-ttwu_do_wakeup
    <idle>-0       3d.h3    3us : resched_curr <-check_preempt_curr
    <idle>-0       3dNh3    4us : task_woken_rt <-ttwu_do_wakeup
    <idle>-0       3dNh3    4us : _raw_spin_unlock <-try_to_wake_up
    <idle>-0       3dNh3    4us : sub_preempt_count <-_raw_spin_unlock
    <idle>-0       3dNh2    5us : ttwu_stat <-try_to_wake_up
    <idle>-0       3dNh2    5us : _raw_spin_unlock_irqrestore <-try_to_wake_up
    <idle>-0       3dNh2    6us : sub_preempt_count <-_raw_spin_unlock_irqrestore
    <idle>-0       3dNh1    6us : _raw_spin_lock <-__run_hrtimer
    <idle>-0       3dNh1    6us : add_preempt_count <-_raw_spin_lock
    <idle>-0       3dNh2    7us : _raw_spin_unlock <-hrtimer_interrupt
    <idle>-0       3dNh2    7us : sub_preempt_count <-_raw_spin_unlock
    <idle>-0       3dNh1    7us : tick_program_event <-hrtimer_interrupt
    <idle>-0       3dNh1    7us : clockevents_program_event <-tick_program_event
    <idle>-0       3dNh1    8us : ktime_get <-clockevents_program_event
    <idle>-0       3dNh1    8us : lapic_next_event <-clockevents_program_event
    <idle>-0       3dNh1    8us : irq_exit <-smp_apic_timer_interrupt
    <idle>-0       3dNh1    9us : sub_preempt_count <-irq_exit
    <idle>-0       3dN.2    9us : idle_cpu <-irq_exit
    <idle>-0       3dN.2    9us : rcu_irq_exit <-irq_exit
    <idle>-0       3dN.2   10us : rcu_eqs_enter_common.isra.45 <-rcu_irq_exit
    <idle>-0       3dN.2   10us : sub_preempt_count <-irq_exit
    <idle>-0       3.N.1   11us : rcu_idle_exit <-cpu_idle
    <idle>-0       3dN.1   11us : rcu_eqs_exit_common.isra.43 <-rcu_idle_exit
    <idle>-0       3.N.1   11us : tick_nohz_idle_exit <-cpu_idle
    <idle>-0       3dN.1   12us : menu_hrtimer_cancel <-tick_nohz_idle_exit
    <idle>-0       3dN.1   12us : ktime_get <-tick_nohz_idle_exit
    <idle>-0       3dN.1   12us : tick_do_update_jiffies64 <-tick_nohz_idle_exit
    <idle>-0       3dN.1   13us : cpu_load_update_nohz <-tick_nohz_idle_exit
    <idle>-0       3dN.1   13us : _raw_spin_lock <-cpu_load_update_nohz
    <idle>-0       3dN.1   13us : add_preempt_count <-_raw_spin_lock
    <idle>-0       3dN.2   13us : __cpu_load_update <-cpu_load_update_nohz
    <idle>-0       3dN.2   14us : sched_avg_update <-__cpu_load_update
    <idle>-0       3dN.2   14us : _raw_spin_unlock <-cpu_load_update_nohz
    <idle>-0       3dN.2   14us : sub_preempt_count <-_raw_spin_unlock
    <idle>-0       3dN.1   15us : calc_load_nohz_stop <-tick_nohz_idle_exit
    <idle>-0       3dN.1   15us : touch_softlockup_watchdog <-tick_nohz_idle_exit
    <idle>-0       3dN.1   15us : hrtimer_cancel <-tick_nohz_idle_exit
    <idle>-0       3dN.1   15us : hrtimer_try_to_cancel <-hrtimer_cancel
    <idle>-0       3dN.1   16us : lock_hrtimer_base.isra.18 <-hrtimer_try_to_cancel
    <idle>-0       3dN.1   16us : _raw_spin_lock_irqsave <-lock_hrtimer_base.isra.18
    <idle>-0       3dN.1   16us : add_preempt_count <-_raw_spin_lock_irqsave
    <idle>-0       3dN.2   17us : __remove_hrtimer <-remove_hrtimer.part.16
    <idle>-0       3dN.2   17us : hrtimer_force_reprogram <-__remove_hrtimer
    <idle>-0       3dN.2   17us : tick_program_event <-hrtimer_force_reprogram
    <idle>-0       3dN.2   18us : clockevents_program_event <-tick_program_event
    <idle>-0       3dN.2   18us : ktime_get <-clockevents_program_event
    <idle>-0       3dN.2   18us : lapic_next_event <-clockevents_program_event
    <idle>-0       3dN.2   19us : _raw_spin_unlock_irqrestore <-hrtimer_try_to_cancel
    <idle>-0       3dN.2   19us : sub_preempt_count <-_raw_spin_unlock_irqrestore
    <idle>-0       3dN.1   19us : hrtimer_forward <-tick_nohz_idle_exit
    <idle>-0       3dN.1   20us : ktime_add_safe <-hrtimer_forward
    <idle>-0       3dN.1   20us : ktime_add_safe <-hrtimer_forward
    <idle>-0       3dN.1   20us : hrtimer_start_range_ns <-hrtimer_start_expires.constprop.11
    <idle>-0       3dN.1   20us : __hrtimer_start_range_ns <-hrtimer_start_range_ns
    <idle>-0       3dN.1   21us : lock_hrtimer_base.isra.18 <-__hrtimer_start_range_ns
    <idle>-0       3dN.1   21us : _raw_spin_lock_irqsave <-lock_hrtimer_base.isra.18
    <idle>-0       3dN.1   21us : add_preempt_count <-_raw_spin_lock_irqsave
    <idle>-0       3dN.2   22us : ktime_add_safe <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   22us : enqueue_hrtimer <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   22us : tick_program_event <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   23us : clockevents_program_event <-tick_program_event
    <idle>-0       3dN.2   23us : ktime_get <-clockevents_program_event
    <idle>-0       3dN.2   23us : lapic_next_event <-clockevents_program_event
    <idle>-0       3dN.2   24us : _raw_spin_unlock_irqrestore <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   24us : sub_preempt_count <-_raw_spin_unlock_irqrestore
    <idle>-0       3dN.1   24us : account_idle_ticks <-tick_nohz_idle_exit
    <idle>-0       3dN.1   24us : account_idle_time <-account_idle_ticks
    <idle>-0       3.N.1   25us : sub_preempt_count <-cpu_idle
    <idle>-0       3.N..   25us : schedule <-cpu_idle
    <idle>-0       3.N..   25us : __schedule <-preempt_schedule
    <idle>-0       3.N..   26us : add_preempt_count <-__schedule
    <idle>-0       3.N.1   26us : rcu_note_context_switch <-__schedule
    <idle>-0       3.N.1   26us : rcu_sched_qs <-rcu_note_context_switch
    <idle>-0       3dN.1   27us : rcu_preempt_qs <-rcu_note_context_switch
    <idle>-0       3.N.1   27us : _raw_spin_lock_irq <-__schedule
    <idle>-0       3dN.1   27us : add_preempt_count <-_raw_spin_lock_irq
    <idle>-0       3dN.2   28us : put_prev_task_idle <-__schedule
    <idle>-0       3dN.2   28us : pick_next_task_stop <-pick_next_task
    <idle>-0       3dN.2   28us : pick_next_task_rt <-pick_next_task
    <idle>-0       3dN.2   29us : dequeue_pushable_task <-pick_next_task_rt
    <idle>-0       3d..3   29us : __schedule <-preempt_schedule
    <idle>-0       3d..3   30us :      0:120:R ==> [003]  2448: 94:R sleep

함수 추적을 켰는데도 이 예의 전체 기록은 비교적 짧아 문서에 생략 없이 실려 있다. 측정 지연은 29마이크로초이며, 타이머 인터럽트에서 태스크를 깨운 뒤 idle을 빠져나와 `__schedule`이 RT 태스크를 고르는 전체 호출 경로를 보여 준다.

시스템이 idle인 동안 인터럽트가 발생했다. `task_woken_rt()`가 호출되기 전 어느 지점에서 `NEED_RESCHED` 플래그가 설정되었고, 추적 행에서 처음 나타나는 상태 문자 `N`이 그 순간 이후를 표시한다.

RT 우선순위 표기
항목예제 값해석
사용자 `rt_prio`5`chrt`에 지정한 RT 우선순위
커널 내부 우선순위94`99 - 5`
idle 일반 우선순위120nice 0에 대한 내부 값
정책 1SCHED_FIFO선입선출 RT 정책
정책 2SCHED_RR라운드로빈 RT 정책

사용자 공간의 RT 우선순위와 추적에 표시되는 내부 값을 구분한다.

wakeup_rt 예제 비교
설정지연기록 수얻는 정보
function-trace 끔5 us4/4wakeup과 스케줄 인 경계
function-trace 켬29 us85/85인터럽트·idle 이탈·스케줄러 호출 경로

함수 추적의 정보량과 오버헤드를 함께 본다.

RT wakeup에서 실행까지
타이머 hardirqRT 태스크 wakeup
RT 태스크 wakeupNEED_RESCHED 설정·N 표시
NEED_RESCHED 설정·N 표시idle 이탈
idle 이탈preempt_schedule
preempt_schedulepick_next_task_rt
pick_next_task_rtRT 태스크 스케줄 인 직전

idle CPU가 RT 태스크를 선택하는 동안 상태 문자 N이 재스케줄 필요를 드러낸다.

wakeup_rt
---------

In a Real-Time environment it is very important to know the
wakeup time it takes for the highest priority task that is woken
up to the time that it executes. This is also known as "schedule
latency". I stress the point that this is about RT tasks. It is
also important to know the scheduling latency of non-RT tasks,
but the average schedule latency is better for non-RT tasks.
Tools like LatencyTop are more appropriate for such
measurements.

Real-Time environments are interested in the worst case latency.
That is the longest latency it takes for something to happen,
and not the average. We can have a very fast scheduler that may
only have a large latency once in a while, but that would not
work well with Real-Time tasks.  The wakeup_rt tracer was designed
to record the worst case wakeups of RT tasks. Non-RT tasks are
not recorded because the tracer only records one worst case and
tracing non-RT tasks that are unpredictable will overwrite the
worst case latency of RT tasks (just run the normal wakeup
tracer for a while to see that effect).

Since this tracer only deals with RT tasks, we will run this
slightly differently than we did with the previous tracers.
Instead of performing an 'ls', we will run 'sleep 1' under
'chrt' which changes the priority of the task.
::

  # echo 0 > options/function-trace
  # echo wakeup_rt > current_tracer
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # chrt -f 5 sleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: wakeup
  #
  # tracer: wakeup_rt
  #
  # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 5 us, #4/4, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sleep-2389 (uid:0 nice:0 policy:1 rt_prio:5)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       3d.h4    0us :      0:120:R   + [003]  2389: 94:R sleep
    <idle>-0       3d.h4    1us+: ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       3d..3    5us : __schedule <-schedule
    <idle>-0       3d..3    5us :      0:120:R ==> [003]  2389: 94:R sleep


Running this on an idle system, we see that it only took 5 microseconds
to perform the task switch.  Note, since the trace point in the schedule
is before the actual "switch", we stop the tracing when the recorded task
is about to schedule in. This may change if we add a new marker at the
end of the scheduler.

Notice that the recorded task is 'sleep' with the PID of 2389
and it has an rt_prio of 5. This priority is user-space priority
and not the internal kernel priority. The policy is 1 for
SCHED_FIFO and 2 for SCHED_RR.

Note, that the trace data shows the internal priority (99 - rtprio).
::

  <idle>-0       3d..3    5us :      0:120:R ==> [003]  2389: 94:R sleep

The 0:120:R means idle was running with a nice priority of 0 (120 - 120)
and in the running state 'R'. The sleep task was scheduled in with
2389: 94:R. That is the priority is the kernel rtprio (99 - 5 = 94)
and it too is in the running state.

Doing the same with chrt -r 5 and function-trace set.
::

  echo 1 > options/function-trace

  # tracer: wakeup_rt
  #
  # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 29 us, #85/85, CPU#3 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sleep-2448 (uid:0 nice:0 policy:1 rt_prio:5)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       3d.h4    1us+:      0:120:R   + [003]  2448: 94:R sleep
    <idle>-0       3d.h4    2us : ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       3d.h3    3us : check_preempt_curr <-ttwu_do_wakeup
    <idle>-0       3d.h3    3us : resched_curr <-check_preempt_curr
    <idle>-0       3dNh3    4us : task_woken_rt <-ttwu_do_wakeup
    <idle>-0       3dNh3    4us : _raw_spin_unlock <-try_to_wake_up
    <idle>-0       3dNh3    4us : sub_preempt_count <-_raw_spin_unlock
    <idle>-0       3dNh2    5us : ttwu_stat <-try_to_wake_up
    <idle>-0       3dNh2    5us : _raw_spin_unlock_irqrestore <-try_to_wake_up
    <idle>-0       3dNh2    6us : sub_preempt_count <-_raw_spin_unlock_irqrestore
    <idle>-0       3dNh1    6us : _raw_spin_lock <-__run_hrtimer
    <idle>-0       3dNh1    6us : add_preempt_count <-_raw_spin_lock
    <idle>-0       3dNh2    7us : _raw_spin_unlock <-hrtimer_interrupt
    <idle>-0       3dNh2    7us : sub_preempt_count <-_raw_spin_unlock
    <idle>-0       3dNh1    7us : tick_program_event <-hrtimer_interrupt
    <idle>-0       3dNh1    7us : clockevents_program_event <-tick_program_event
    <idle>-0       3dNh1    8us : ktime_get <-clockevents_program_event
    <idle>-0       3dNh1    8us : lapic_next_event <-clockevents_program_event
    <idle>-0       3dNh1    8us : irq_exit <-smp_apic_timer_interrupt
    <idle>-0       3dNh1    9us : sub_preempt_count <-irq_exit
    <idle>-0       3dN.2    9us : idle_cpu <-irq_exit
    <idle>-0       3dN.2    9us : rcu_irq_exit <-irq_exit
    <idle>-0       3dN.2   10us : rcu_eqs_enter_common.isra.45 <-rcu_irq_exit
    <idle>-0       3dN.2   10us : sub_preempt_count <-irq_exit
    <idle>-0       3.N.1   11us : rcu_idle_exit <-cpu_idle
    <idle>-0       3dN.1   11us : rcu_eqs_exit_common.isra.43 <-rcu_idle_exit
    <idle>-0       3.N.1   11us : tick_nohz_idle_exit <-cpu_idle
    <idle>-0       3dN.1   12us : menu_hrtimer_cancel <-tick_nohz_idle_exit
    <idle>-0       3dN.1   12us : ktime_get <-tick_nohz_idle_exit
    <idle>-0       3dN.1   12us : tick_do_update_jiffies64 <-tick_nohz_idle_exit
    <idle>-0       3dN.1   13us : cpu_load_update_nohz <-tick_nohz_idle_exit
    <idle>-0       3dN.1   13us : _raw_spin_lock <-cpu_load_update_nohz
    <idle>-0       3dN.1   13us : add_preempt_count <-_raw_spin_lock
    <idle>-0       3dN.2   13us : __cpu_load_update <-cpu_load_update_nohz
    <idle>-0       3dN.2   14us : sched_avg_update <-__cpu_load_update
    <idle>-0       3dN.2   14us : _raw_spin_unlock <-cpu_load_update_nohz
    <idle>-0       3dN.2   14us : sub_preempt_count <-_raw_spin_unlock
    <idle>-0       3dN.1   15us : calc_load_nohz_stop <-tick_nohz_idle_exit
    <idle>-0       3dN.1   15us : touch_softlockup_watchdog <-tick_nohz_idle_exit
    <idle>-0       3dN.1   15us : hrtimer_cancel <-tick_nohz_idle_exit
    <idle>-0       3dN.1   15us : hrtimer_try_to_cancel <-hrtimer_cancel
    <idle>-0       3dN.1   16us : lock_hrtimer_base.isra.18 <-hrtimer_try_to_cancel
    <idle>-0       3dN.1   16us : _raw_spin_lock_irqsave <-lock_hrtimer_base.isra.18
    <idle>-0       3dN.1   16us : add_preempt_count <-_raw_spin_lock_irqsave
    <idle>-0       3dN.2   17us : __remove_hrtimer <-remove_hrtimer.part.16
    <idle>-0       3dN.2   17us : hrtimer_force_reprogram <-__remove_hrtimer
    <idle>-0       3dN.2   17us : tick_program_event <-hrtimer_force_reprogram
    <idle>-0       3dN.2   18us : clockevents_program_event <-tick_program_event
    <idle>-0       3dN.2   18us : ktime_get <-clockevents_program_event
    <idle>-0       3dN.2   18us : lapic_next_event <-clockevents_program_event
    <idle>-0       3dN.2   19us : _raw_spin_unlock_irqrestore <-hrtimer_try_to_cancel
    <idle>-0       3dN.2   19us : sub_preempt_count <-_raw_spin_unlock_irqrestore
    <idle>-0       3dN.1   19us : hrtimer_forward <-tick_nohz_idle_exit
    <idle>-0       3dN.1   20us : ktime_add_safe <-hrtimer_forward
    <idle>-0       3dN.1   20us : ktime_add_safe <-hrtimer_forward
    <idle>-0       3dN.1   20us : hrtimer_start_range_ns <-hrtimer_start_expires.constprop.11
    <idle>-0       3dN.1   20us : __hrtimer_start_range_ns <-hrtimer_start_range_ns
    <idle>-0       3dN.1   21us : lock_hrtimer_base.isra.18 <-__hrtimer_start_range_ns
    <idle>-0       3dN.1   21us : _raw_spin_lock_irqsave <-lock_hrtimer_base.isra.18
    <idle>-0       3dN.1   21us : add_preempt_count <-_raw_spin_lock_irqsave
    <idle>-0       3dN.2   22us : ktime_add_safe <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   22us : enqueue_hrtimer <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   22us : tick_program_event <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   23us : clockevents_program_event <-tick_program_event
    <idle>-0       3dN.2   23us : ktime_get <-clockevents_program_event
    <idle>-0       3dN.2   23us : lapic_next_event <-clockevents_program_event
    <idle>-0       3dN.2   24us : _raw_spin_unlock_irqrestore <-__hrtimer_start_range_ns
    <idle>-0       3dN.2   24us : sub_preempt_count <-_raw_spin_unlock_irqrestore
    <idle>-0       3dN.1   24us : account_idle_ticks <-tick_nohz_idle_exit
    <idle>-0       3dN.1   24us : account_idle_time <-account_idle_ticks
    <idle>-0       3.N.1   25us : sub_preempt_count <-cpu_idle
    <idle>-0       3.N..   25us : schedule <-cpu_idle
    <idle>-0       3.N..   25us : __schedule <-preempt_schedule
    <idle>-0       3.N..   26us : add_preempt_count <-__schedule
    <idle>-0       3.N.1   26us : rcu_note_context_switch <-__schedule
    <idle>-0       3.N.1   26us : rcu_sched_qs <-rcu_note_context_switch
    <idle>-0       3dN.1   27us : rcu_preempt_qs <-rcu_note_context_switch
    <idle>-0       3.N.1   27us : _raw_spin_lock_irq <-__schedule
    <idle>-0       3dN.1   27us : add_preempt_count <-_raw_spin_lock_irq
    <idle>-0       3dN.2   28us : put_prev_task_idle <-__schedule
    <idle>-0       3dN.2   28us : pick_next_task_stop <-pick_next_task
    <idle>-0       3dN.2   28us : pick_next_task_rt <-pick_next_task
    <idle>-0       3dN.2   29us : dequeue_pushable_task <-pick_next_task_rt
    <idle>-0       3d..3   29us : __schedule <-preempt_schedule
    <idle>-0       3d..3   30us :      0:120:R ==> [003]  2448: 94:R sleep

This isn't that big of a trace, even with function tracing enabled,
so I included the entire trace.

The interrupt went off while when the system was idle. Somewhere
before task_woken_rt() was called, the NEED_RESCHED flag was set,
this is indicated by the first occurrence of the 'N' flag.

지연 추적과 이벤트

2242-2288

함수 추적은 지연 구간 내부를 자세히 보여 주지만 측정 지연 자체를 크게 늘릴 수 있다. 반대로 함수 추적을 완전히 끄면 원인을 알기 어렵다. 두 극단 사이의 절충안은 tracepoint 이벤트를 활성화하는 것이다.

다음 예는 `function-trace`를 끈 채 `wakeup_rt`를 선택하고 `events/enable`로 모든 이벤트를 켠다.

  # echo 0 > options/function-trace
  # echo wakeup_rt > current_tracer
  # echo 1 > events/enable
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # chrt -f 5 sleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: wakeup_rt
  #
  # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 6 us, #12/12, CPU#2 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sleep-5882 (uid:0 nice:0 policy:1 rt_prio:5)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       2d.h4    0us :      0:120:R   + [002]  5882: 94:R sleep
    <idle>-0       2d.h4    0us : ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       2d.h4    1us : sched_wakeup: comm=sleep pid=5882 prio=94 success=1 target_cpu=002
    <idle>-0       2dNh2    1us : hrtimer_expire_exit: hrtimer=ffff88007796feb8
    <idle>-0       2.N.2    2us : power_end: cpu_id=2
    <idle>-0       2.N.2    3us : cpu_idle: state=4294967295 cpu_id=2
    <idle>-0       2dN.3    4us : hrtimer_cancel: hrtimer=ffff88007d50d5e0
    <idle>-0       2dN.3    4us : hrtimer_start: hrtimer=ffff88007d50d5e0 function=tick_sched_timer expires=34311211000000 softexpires=34311211000000
    <idle>-0       2.N.2    5us : rcu_utilization: Start context switch
    <idle>-0       2.N.2    5us : rcu_utilization: End context switch
    <idle>-0       2d..3    6us : __schedule <-schedule
    <idle>-0       2d..3    6us :      0:120:R ==> [002]  5882: 94:R sleep

결과는 6마이크로초 지연 동안 12개 기록을 남긴다. 함수 호출 전체를 기록하지 않아도 `sched_wakeup`, hrtimer, CPU idle, RCU 문맥 전환 이벤트를 통해 wakeup에서 스케줄 인까지의 주요 상태 변화를 확인할 수 있다.

지연 내부 관찰 방식
방식정보량오버헤드적합한 용도
함수 추적 끔최소낮음순수 지연값과 경계 확인
이벤트 활성화중간중간주요 커널 상태 변화와 원인 단서
함수 추적 켬최대높음세부 호출 경로 분석

원인 정보와 추적 오버헤드 사이에서 알맞은 수준을 고른다.

이벤트 기반 wakeup 관찰
sched_wakeuphrtimer_expire_exit
hrtimer_expire_exitcpu_idle 이탈
cpu_idle 이탈RCU context switch
RCU context switch__schedule
__schedulesleep 스케줄 인

선택된 tracepoint가 함수 전체 대신 핵심 전환을 연결한다.

Latency tracing and events
--------------------------
As function tracing can induce a much larger latency, but without
seeing what happens within the latency it is hard to know what
caused it. There is a middle ground, and that is with enabling
events.
::

  # echo 0 > options/function-trace
  # echo wakeup_rt > current_tracer
  # echo 1 > events/enable
  # echo 1 > tracing_on
  # echo 0 > tracing_max_latency
  # chrt -f 5 sleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: wakeup_rt
  #
  # wakeup_rt latency trace v1.1.5 on 3.8.0-test+
  # --------------------------------------------------------------------
  # latency: 6 us, #12/12, CPU#2 | (M:preempt VP:0, KP:0, SP:0 HP:0 #P:4)
  #    -----------------
  #    | task: sleep-5882 (uid:0 nice:0 policy:1 rt_prio:5)
  #    -----------------
  #
  #                  _------=> CPU#            
  #                 / _-----=> irqs-off        
  #                | / _----=> need-resched    
  #                || / _---=> hardirq/softirq 
  #                ||| / _--=> preempt-depth   
  #                |||| /     delay             
  #  cmd     pid   ||||| time  |   caller      
  #     \   /      |||||  \    |   /           
    <idle>-0       2d.h4    0us :      0:120:R   + [002]  5882: 94:R sleep
    <idle>-0       2d.h4    0us : ttwu_do_activate.constprop.87 <-try_to_wake_up
    <idle>-0       2d.h4    1us : sched_wakeup: comm=sleep pid=5882 prio=94 success=1 target_cpu=002
    <idle>-0       2dNh2    1us : hrtimer_expire_exit: hrtimer=ffff88007796feb8
    <idle>-0       2.N.2    2us : power_end: cpu_id=2
    <idle>-0       2.N.2    3us : cpu_idle: state=4294967295 cpu_id=2
    <idle>-0       2dN.3    4us : hrtimer_cancel: hrtimer=ffff88007d50d5e0
    <idle>-0       2dN.3    4us : hrtimer_start: hrtimer=ffff88007d50d5e0 function=tick_sched_timer expires=34311211000000 softexpires=34311211000000
    <idle>-0       2.N.2    5us : rcu_utilization: Start context switch
    <idle>-0       2.N.2    5us : rcu_utilization: End context switch
    <idle>-0       2d..3    6us : __schedule <-schedule
    <idle>-0       2d..3    6us :      0:120:R ==> [002]  5882: 94:R sleep

하드웨어 지연 탐지기

2289-2384

하드웨어 지연 탐지기는 `current_tracer`에 `hwlat`를 써서 실행한다.

이 추적기는 주기적으로 한 CPU의 인터럽트를 끄고 계속 바쁘게 돌리므로 시스템 성능에 직접 영향을 준다. 운영 부하에서 사용할 때는 측정 자체가 만드는 정지 시간과 CPU 점유를 반드시 고려해야 한다.

  # echo hwlat > current_tracer
  # sleep 100
  # cat trace
  # tracer: hwlat
  #
  # entries-in-buffer/entries-written: 13/13   #P:8
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
             <...>-1729  [001] d...   678.473449: #1     inner/outer(us):   11/12    ts:1581527483.343962693 count:6
             <...>-1729  [004] d...   689.556542: #2     inner/outer(us):   16/9     ts:1581527494.889008092 count:1
             <...>-1729  [005] d...   714.756290: #3     inner/outer(us):   16/16    ts:1581527519.678961629 count:5
             <...>-1729  [001] d...   718.788247: #4     inner/outer(us):    9/17    ts:1581527523.889012713 count:1
             <...>-1729  [002] d...   719.796341: #5     inner/outer(us):   13/9     ts:1581527524.912872606 count:1
             <...>-1729  [006] d...   844.787091: #6     inner/outer(us):    9/12    ts:1581527649.889048502 count:2
             <...>-1729  [003] d...   849.827033: #7     inner/outer(us):   18/9     ts:1581527654.889013793 count:1
             <...>-1729  [007] d...   853.859002: #8     inner/outer(us):    9/12    ts:1581527658.889065736 count:1
             <...>-1729  [001] d...   855.874978: #9     inner/outer(us):    9/11    ts:1581527660.861991877 count:1
             <...>-1729  [001] d...   863.938932: #10    inner/outer(us):    9/11    ts:1581527668.970010500 count:1 nmi-total:7 nmi-count:1
             <...>-1729  [007] d...   878.050780: #11    inner/outer(us):    9/12    ts:1581527683.385002600 count:1 nmi-total:5 nmi-count:1
             <...>-1729  [007] d...   886.114702: #12    inner/outer(us):    9/12    ts:1581527691.385001600 count:1

헤더 형식은 다른 추적기와 비슷하며 모든 사건은 IRQ가 비활성화된 상태 문자 `d`로 표시된다. `FUNCTION` 열에는 탐지 순번, inner/outer 지연, 절대 시각, 창 안의 탐지 횟수와 선택적인 NMI 통계가 기록된다.

`#1`은 `tracing_threshold`보다 큰 지연 사건이 기록된 순번이다.

`inner/outer(us): 11/11`의 inner latency는 루프 안에서 타임스탬프를 연속 두 번 확인하는 사이에 탐지된 지연이다. outer latency는 이전 반복의 마지막 타임스탬프와 다음 반복의 첫 타임스탬프 사이에서 탐지된 지연이다.

`ts:1581527483.343962693`은 해당 측정 창에서 첫 지연을 기록한 절대 타임스탬프다. `count:6`은 그 창 안에서 지연을 탐지한 횟수다.

아키텍처가 지원하고 테스트 도중 NMI가 들어오면 `nmi-total`에 NMI에서 보낸 총 시간을 마이크로초 단위로 기록한다. NMI가 있는 모든 아키텍처는 테스트 중 NMI 발생 시 `nmi-count`에 발생 횟수를 표시한다.

`tracing_threshold`는 `hwlat` 시작 시 자동으로 10, 즉 10마이크로초로 설정된다. 이 값보다 큰 지연만 추적에 기록한다. 다른 추적기를 `current_tracer`에 써서 `hwlat`를 끝내면 파일의 원래 값이 복원된다.

`hwlat_detector/width`는 인터럽트를 비활성화하고 실제 테스트를 수행하는 시간이다. `hwlat_detector/window`는 한 측정 주기의 전체 길이이며, 매 window 마이크로초마다 width 마이크로초 동안 테스트한다.

테스트를 시작하면 커널 스레드가 만들어진다. 이 스레드는 각 window가 끝날 때마다 `tracing_cpumask`에 포함된 CPU 사이를 번갈아 이동한다. 특정 CPU에서만 측정하려면 이 마스크에 그 CPU들만 지정한다.

hwlat 출력 필드
필드의미단위/조건
#N임계값을 넘은 기록 순번사건 번호
inner연속한 두 내부 타임스탬프 사이 지연us
outer반복 경계의 두 타임스탬프 사이 지연us
ts창에서 첫 지연이 기록된 절대 시각초 단위 절대 타임스탬프
count한 window 안의 지연 탐지 횟수횟수
nmi-totalNMI 안에서 보낸 총 시간지원 아키텍처, us
nmi-count테스트 중 들어온 NMI 수NMI 발생 시

한 측정 창에서 탐지한 하드웨어 지연과 NMI 영향을 구분한다.

hwlat 제어 파일
파일역할
tracing_threshold기록할 최소 지연, 시작 시 10us로 임시 설정
hwlat_detector/widthIRQ를 끄고 바쁜 루프를 실행하는 시간
hwlat_detector/windowwidth 테스트를 한 번 수행하는 전체 주기
tracing_cpumask측정 커널 스레드가 번갈아 실행될 CPU 집합

임계값·측정 폭·주기·대상 CPU를 각각 제어한다.

hwlat 측정 창
window 시작대상 CPU 선택
대상 CPU 선택IRQ 비활성화
IRQ 비활성화width 동안 타임스탬프 반복 검사
width 동안 타임스탬프 반복 검사inner/outer 지연 계산
inner/outer 지연 계산tracing_threshold 초과 기록
tracing_threshold 초과 기록다음 window에서 CPU 교대

각 window 안에서 width만큼 IRQ를 끄고 타임스탬프 간격을 반복 검사한다.

Hardware Latency Detector
-------------------------

The hardware latency detector is executed by enabling the "hwlat" tracer.

NOTE, this tracer will affect the performance of the system as it will
periodically make a CPU constantly busy with interrupts disabled.
::

  # echo hwlat > current_tracer
  # sleep 100
  # cat trace
  # tracer: hwlat
  #
  # entries-in-buffer/entries-written: 13/13   #P:8
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
             <...>-1729  [001] d...   678.473449: #1     inner/outer(us):   11/12    ts:1581527483.343962693 count:6
             <...>-1729  [004] d...   689.556542: #2     inner/outer(us):   16/9     ts:1581527494.889008092 count:1
             <...>-1729  [005] d...   714.756290: #3     inner/outer(us):   16/16    ts:1581527519.678961629 count:5
             <...>-1729  [001] d...   718.788247: #4     inner/outer(us):    9/17    ts:1581527523.889012713 count:1
             <...>-1729  [002] d...   719.796341: #5     inner/outer(us):   13/9     ts:1581527524.912872606 count:1
             <...>-1729  [006] d...   844.787091: #6     inner/outer(us):    9/12    ts:1581527649.889048502 count:2
             <...>-1729  [003] d...   849.827033: #7     inner/outer(us):   18/9     ts:1581527654.889013793 count:1
             <...>-1729  [007] d...   853.859002: #8     inner/outer(us):    9/12    ts:1581527658.889065736 count:1
             <...>-1729  [001] d...   855.874978: #9     inner/outer(us):    9/11    ts:1581527660.861991877 count:1
             <...>-1729  [001] d...   863.938932: #10    inner/outer(us):    9/11    ts:1581527668.970010500 count:1 nmi-total:7 nmi-count:1
             <...>-1729  [007] d...   878.050780: #11    inner/outer(us):    9/12    ts:1581527683.385002600 count:1 nmi-total:5 nmi-count:1
             <...>-1729  [007] d...   886.114702: #12    inner/outer(us):    9/12    ts:1581527691.385001600 count:1


The above output is somewhat the same in the header. All events will have
interrupts disabled 'd'. Under the FUNCTION title there is:

 #1
	This is the count of events recorded that were greater than the
	tracing_threshold (See below).

 inner/outer(us):   11/11

      This shows two numbers as "inner latency" and "outer latency". The test
      runs in a loop checking a timestamp twice. The latency detected within
      the two timestamps is the "inner latency" and the latency detected
      after the previous timestamp and the next timestamp in the loop is
      the "outer latency".

 ts:1581527483.343962693

      The absolute timestamp that the first latency was recorded in the window.

 count:6

      The number of times a latency was detected during the window.

 nmi-total:7 nmi-count:1

      On architectures that support it, if an NMI comes in during the
      test, the time spent in NMI is reported in "nmi-total" (in
      microseconds).

      All architectures that have NMIs will show the "nmi-count" if an
      NMI comes in during the test.

hwlat files:

  tracing_threshold
	This gets automatically set to "10" to represent 10
	microseconds. This is the threshold of latency that
	needs to be detected before the trace will be recorded.

	Note, when hwlat tracer is finished (another tracer is
	written into "current_tracer"), the original value for
	tracing_threshold is placed back into this file.

  hwlat_detector/width
	The length of time the test runs with interrupts disabled.

  hwlat_detector/window
	The length of time of the window which the test
	runs. That is, the test will run for "width"
	microseconds per "window" microseconds

  tracing_cpumask
	When the test is started. A kernel thread is created that
	runs the test. This thread will alternate between CPUs
	listed in the tracing_cpumask between each period
	(one "window"). To limit the test to specific CPUs
	set the mask in this file to only the CPUs that the test
	should run on.

function 추적기

2385-2444

`function`은 커널 함수 진입을 기록하는 함수 추적기다. tracefs 제어 파일에서 활성화할 수 있지만 먼저 `ftrace_enabled`가 켜져 있어야 한다. 이 전역 설정이 꺼져 있으면 추적기를 선택해도 `nop`처럼 아무것도 기록하지 않는다.

다음 예는 sysctl로 `kernel.ftrace_enabled=1`을 설정하고 `function` 추적기를 잠깐 실행한 뒤 결과를 읽는다.

  # sysctl kernel.ftrace_enabled=1
  # echo function > current_tracer
  # echo 1 > tracing_on
  # usleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: function
  #
  # entries-in-buffer/entries-written: 24799/24799   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1994  [002] ....  3082.063030: mutex_unlock <-rb_simple_write
              bash-1994  [002] ....  3082.063031: __mutex_unlock_slowpath <-mutex_unlock
              bash-1994  [002] ....  3082.063031: __fsnotify_parent <-fsnotify_modify
              bash-1994  [002] ....  3082.063032: fsnotify <-fsnotify_modify
              bash-1994  [002] ....  3082.063032: __srcu_read_lock <-fsnotify
              bash-1994  [002] ....  3082.063032: add_preempt_count <-__srcu_read_lock
              bash-1994  [002] ...1  3082.063032: sub_preempt_count <-__srcu_read_lock
              bash-1994  [002] ....  3082.063033: __srcu_read_unlock <-fsnotify
  [...]

각 행은 태스크와 PID, CPU, IRQ·재스케줄·문맥 상태, 타임스탬프, 추적된 함수와 호출자를 보여 준다. 예에서는 `bash-1994`가 CPU 2에서 `mutex_unlock`, fsnotify, SRCU 관련 함수를 차례로 호출한다.

함수 추적기는 링 버퍼에 기록하므로 새 데이터가 오래된 데이터를 덮어쓸 수 있다. 셸에서 `echo 0 > tracing_on`을 실행해 멈추는 사이에도 관심 있던 기록이 이미 사라질 수 있다.

관심 조건을 만난 정확한 지점에서 추적을 멈추려면 프로그램이 `tracing_on` 파일을 미리 열어 두었다가 조건이 참일 때 직접 0을 쓰는 편이 낫다. 다음 C 조각은 그 방식을 보여 준다.

	int trace_fd;
	[...]
	int main(int argc, char *argv[]) {
		[...]
		trace_fd = open(tracing_file("tracing_on"), O_WRONLY);
		[...]
		if (condition_hit()) {
			write(trace_fd, "0", 1);
		}
		[...]
	}
함수 추적 중지 방식
방식중지 시점위험/장점
셸에서 echo 0사용자가 명령을 실행한 뒤반응 지연 동안 원하는 기록이 덮일 수 있음
프로그램에서 write`condition_hit()` 직후관심 사건과 가까운 시점에 즉시 정지

링 버퍼 덮어쓰기를 피하려면 중지 지점을 관심 조건 가까이에 둔다.

프로그램 내부 추적 정지
tracing_on 열기대상 코드 실행
대상 코드 실행condition_hit() 검사
조건 거짓대상 코드 계속 실행
조건 참write(trace_fd, 0, 1)
write(trace_fd, 0, 1)링 버퍼 상태 고정

`tracing_on` 파일 설명자를 미리 열어 조건 충족 즉시 기록을 멈춘다.

function
--------

This tracer is the function tracer. Enabling the function tracer
can be done from the debug file system. Make sure the
ftrace_enabled is set; otherwise this tracer is a nop.
See the "ftrace_enabled" section below.
::

  # sysctl kernel.ftrace_enabled=1
  # echo function > current_tracer
  # echo 1 > tracing_on
  # usleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: function
  #
  # entries-in-buffer/entries-written: 24799/24799   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1994  [002] ....  3082.063030: mutex_unlock <-rb_simple_write
              bash-1994  [002] ....  3082.063031: __mutex_unlock_slowpath <-mutex_unlock
              bash-1994  [002] ....  3082.063031: __fsnotify_parent <-fsnotify_modify
              bash-1994  [002] ....  3082.063032: fsnotify <-fsnotify_modify
              bash-1994  [002] ....  3082.063032: __srcu_read_lock <-fsnotify
              bash-1994  [002] ....  3082.063032: add_preempt_count <-__srcu_read_lock
              bash-1994  [002] ...1  3082.063032: sub_preempt_count <-__srcu_read_lock
              bash-1994  [002] ....  3082.063033: __srcu_read_unlock <-fsnotify
  [...]


Note: function tracer uses ring buffers to store the above
entries. The newest data may overwrite the oldest data.
Sometimes using echo to stop the trace is not sufficient because
the tracing could have overwritten the data that you wanted to
record. For this reason, it is sometimes better to disable
tracing directly from a program. This allows you to stop the
tracing at the point that you hit the part that you are
interested in. To disable the tracing directly from a C program,
something like following code snippet can be used::

	int trace_fd;
	[...]
	int main(int argc, char *argv[]) {
		[...]
		trace_fd = open(tracing_file("tracing_on"), O_WRONLY);
		[...]
		if (condition_hit()) {
			write(trace_fd, "0", 1);
		}
		[...]
	}

단일 스레드 추적

2445-2581

`set_ftrace_pid`에 PID를 쓰면 함수 추적 대상을 단일 스레드로 제한할 수 있다. 파일이 비어 있을 때는 `no pid`가 표시된다.

  # cat set_ftrace_pid
  no pid
  # echo 3111 > set_ftrace_pid
  # cat set_ftrace_pid
  3111
  # echo function > current_tracer
  # cat trace | head
  # tracer: function
  #
  #           TASK-PID    CPU#    TIMESTAMP  FUNCTION
  #              | |       |          |         |
      yum-updatesd-3111  [003]  1637.254676: finish_task_switch <-thread_return
      yum-updatesd-3111  [003]  1637.254681: hrtimer_cancel <-schedule_hrtimeout_range
      yum-updatesd-3111  [003]  1637.254682: hrtimer_try_to_cancel <-hrtimer_cancel
      yum-updatesd-3111  [003]  1637.254683: lock_hrtimer_base <-hrtimer_try_to_cancel
      yum-updatesd-3111  [003]  1637.254685: fget_light <-do_sys_poll
      yum-updatesd-3111  [003]  1637.254686: pipe_poll <-do_sys_poll
  # echo > set_ftrace_pid
  # cat trace |head
  # tracer: function
  #
  #           TASK-PID    CPU#    TIMESTAMP  FUNCTION
  #              | |       |          |         |
  ##### CPU 3 buffer started ####
      yum-updatesd-3111  [003]  1701.957688: free_poll_entry <-poll_freewait
      yum-updatesd-3111  [003]  1701.957689: remove_wait_queue <-free_poll_entry
      yum-updatesd-3111  [003]  1701.957691: fput <-free_poll_entry
      yum-updatesd-3111  [003]  1701.957692: audit_syscall_exit <-sysret_audit
      yum-updatesd-3111  [003]  1701.957693: path_put <-audit_syscall_exit

예에서는 PID 3111을 설정한 뒤 `function` 추적기를 선택해 `yum-updatesd-3111`의 함수만 기록한다. 빈 값을 써서 `set_ftrace_pid`를 지우면 PID 필터가 해제되고 이후에는 일반 함수 추적으로 돌아간다.

특정 명령이 실행되는 동안 그 프로세스만 추적하려면 작은 실행 래퍼를 사용할 수 있다. 다음 C 프로그램은 tracefs 마운트 지점을 찾고, `current_tracer`와 `set_ftrace_pid`를 설정한 뒤 대상 명령을 `execvp()`로 실행한다.

	#include <stdio.h>
	#include <stdlib.h>
	#include <sys/types.h>
	#include <sys/stat.h>
	#include <fcntl.h>
	#include <unistd.h>
	#include <string.h>

	#define _STR(x) #x
	#define STR(x) _STR(x)
	#define MAX_PATH 256

	const char *find_tracefs(void)
	{
	       static char tracefs[MAX_PATH+1];
	       static int tracefs_found;
	       char type[100];
	       FILE *fp;

	       if (tracefs_found)
		       return tracefs;

	       if ((fp = fopen("/proc/mounts","r")) == NULL) {
		       perror("/proc/mounts");
		       return NULL;
	       }

	       while (fscanf(fp, "%*s %"
		             STR(MAX_PATH)
		             "s %99s %*s %*d %*d\n",
		             tracefs, type) == 2) {
		       if (strcmp(type, "tracefs") == 0)
		               break;
	       }
	       fclose(fp);

	       if (strcmp(type, "tracefs") != 0) {
		       fprintf(stderr, "tracefs not mounted");
		       return NULL;
	       }

	       strcat(tracefs, "/tracing/");
	       tracefs_found = 1;

	       return tracefs;
	}

	const char *tracing_file(const char *file_name)
	{
	       static char trace_file[MAX_PATH+1];
	       snprintf(trace_file, MAX_PATH, "%s/%s", find_tracefs(), file_name);
	       return trace_file;
	}

	int main (int argc, char **argv)
	{
		if (argc < 1)
		        exit(-1);

		if (fork() > 0) {
		        int fd, ffd;
		        char line[64];
		        int s;

		        ffd = open(tracing_file("current_tracer"), O_WRONLY);
		        if (ffd < 0)
		                exit(-1);
		        write(ffd, "nop", 3);

		        fd = open(tracing_file("set_ftrace_pid"), O_WRONLY);
		        s = sprintf(line, "%d\n", getpid());
		        write(fd, line, s);

		        write(ffd, "function", 8);

		        close(fd);
		        close(ffd);

		        execvp(argv[1], argv+1);
		}

		return 0;
	}

`find_tracefs()`는 `/proc/mounts`를 읽어 파일 시스템 유형이 `tracefs`인 마운트를 찾고 제어 파일 경로를 구성한다. `tracing_file()`은 그 기준 경로에 요청한 파일 이름을 붙인다.

래퍼는 먼저 `current_tracer`를 `nop`으로 바꾸고 자신의 PID를 `set_ftrace_pid`에 쓴 다음 `function`을 선택한다. 마지막 `execvp(argv[1], argv+1)`은 같은 PID로 대상 프로그램을 실행하므로 필터가 그 명령의 함수 호출을 따라간다.

같은 절차는 다음과 같은 간단한 셸 스크립트로도 수행할 수 있다.

  #!/bin/bash

  tracefs=`sed -ne 's/^tracefs \(.*\) tracefs.*/\1/p' /proc/mounts`
  echo 0 > $tracefs/tracing_on
  echo $$ > $tracefs/set_ftrace_pid
  echo function > $tracefs/current_tracer
  echo 1 > $tracefs/tracing_on
  exec "$@"

스크립트는 `/proc/mounts`에서 tracefs 경로를 찾고 추적을 잠시 끈 뒤 `$$`를 `set_ftrace_pid`에 쓴다. `function` 추적기를 선택하고 기록을 다시 켠 다음 `exec "$@"`로 대상 명령을 같은 셸 PID에 덮어씌운다.

단일 스레드 추적 단계
단계제어 파일/호출효과
1tracing_on = 0설정 중 기록 중지
2set_ftrace_pid = 대상 PID추적 스레드 제한
3current_tracer = function함수 진입 추적 선택
4tracing_on = 1기록 시작
5execvp 또는 exec같은 PID로 대상 명령 실행
6set_ftrace_pid 비우기PID 필터 해제

필터를 먼저 설정한 뒤 함수 추적기를 켜 대상 PID의 호출만 수집한다.

실행 래퍼 비교
구현tracefs 검색실행 전 제어대상 실행
C 프로그램`/proc/mounts`를 `fscanf`로 검색nop·PID·function 설정`execvp()`
셸 스크립트`sed`로 tracefs 마운트 추출정지·PID·function·시작`exec "$@"`

C와 셸 구현은 같은 제어 순서를 서로 다른 수준으로 제공한다.

PID를 유지한 명령 추적
래퍼 PID를 set_ftrace_pid에 기록function 추적 활성화
function 추적 활성화exec로 대상 명령 실행
exec로 대상 명령 실행PID 유지
PID 유지대상 명령의 함수만 링 버퍼에 기록

exec가 프로세스 이미지만 바꾸고 PID는 유지하므로 미리 설정한 필터가 계속 적용된다.

Single thread tracing
---------------------

By writing into set_ftrace_pid you can trace a
single thread. For example::

  # cat set_ftrace_pid
  no pid
  # echo 3111 > set_ftrace_pid
  # cat set_ftrace_pid
  3111
  # echo function > current_tracer
  # cat trace | head
  # tracer: function
  #
  #           TASK-PID    CPU#    TIMESTAMP  FUNCTION
  #              | |       |          |         |
      yum-updatesd-3111  [003]  1637.254676: finish_task_switch <-thread_return
      yum-updatesd-3111  [003]  1637.254681: hrtimer_cancel <-schedule_hrtimeout_range
      yum-updatesd-3111  [003]  1637.254682: hrtimer_try_to_cancel <-hrtimer_cancel
      yum-updatesd-3111  [003]  1637.254683: lock_hrtimer_base <-hrtimer_try_to_cancel
      yum-updatesd-3111  [003]  1637.254685: fget_light <-do_sys_poll
      yum-updatesd-3111  [003]  1637.254686: pipe_poll <-do_sys_poll
  # echo > set_ftrace_pid
  # cat trace |head
  # tracer: function
  #
  #           TASK-PID    CPU#    TIMESTAMP  FUNCTION
  #              | |       |          |         |
  ##### CPU 3 buffer started ####
      yum-updatesd-3111  [003]  1701.957688: free_poll_entry <-poll_freewait
      yum-updatesd-3111  [003]  1701.957689: remove_wait_queue <-free_poll_entry
      yum-updatesd-3111  [003]  1701.957691: fput <-free_poll_entry
      yum-updatesd-3111  [003]  1701.957692: audit_syscall_exit <-sysret_audit
      yum-updatesd-3111  [003]  1701.957693: path_put <-audit_syscall_exit

If you want to trace a function when executing, you could use
something like this simple program.
::

	#include <stdio.h>
	#include <stdlib.h>
	#include <sys/types.h>
	#include <sys/stat.h>
	#include <fcntl.h>
	#include <unistd.h>
	#include <string.h>

	#define _STR(x) #x
	#define STR(x) _STR(x)
	#define MAX_PATH 256

	const char *find_tracefs(void)
	{
	       static char tracefs[MAX_PATH+1];
	       static int tracefs_found;
	       char type[100];
	       FILE *fp;

	       if (tracefs_found)
		       return tracefs;

	       if ((fp = fopen("/proc/mounts","r")) == NULL) {
		       perror("/proc/mounts");
		       return NULL;
	       }

	       while (fscanf(fp, "%*s %"
		             STR(MAX_PATH)
		             "s %99s %*s %*d %*d\n",
		             tracefs, type) == 2) {
		       if (strcmp(type, "tracefs") == 0)
		               break;
	       }
	       fclose(fp);

	       if (strcmp(type, "tracefs") != 0) {
		       fprintf(stderr, "tracefs not mounted");
		       return NULL;
	       }

	       strcat(tracefs, "/tracing/");
	       tracefs_found = 1;

	       return tracefs;
	}

	const char *tracing_file(const char *file_name)
	{
	       static char trace_file[MAX_PATH+1];
	       snprintf(trace_file, MAX_PATH, "%s/%s", find_tracefs(), file_name);
	       return trace_file;
	}

	int main (int argc, char **argv)
	{
		if (argc < 1)
		        exit(-1);

		if (fork() > 0) {
		        int fd, ffd;
		        char line[64];
		        int s;

		        ffd = open(tracing_file("current_tracer"), O_WRONLY);
		        if (ffd < 0)
		                exit(-1);
		        write(ffd, "nop", 3);

		        fd = open(tracing_file("set_ftrace_pid"), O_WRONLY);
		        s = sprintf(line, "%d\n", getpid());
		        write(fd, line, s);

		        write(ffd, "function", 8);

		        close(fd);
		        close(ffd);

		        execvp(argv[1], argv+1);
		}

		return 0;
	}

Or this simple script!
::

  #!/bin/bash

  tracefs=`sed -ne 's/^tracefs \(.*\) tracefs.*/\1/p' /proc/mounts`
  echo 0 > $tracefs/tracing_on
  echo $$ > $tracefs/set_ftrace_pid
  echo function > $tracefs/current_tracer
  echo 1 > $tracefs/tracing_on
  exec "$@"

function graph 추적기

2582-2924

`function_graph` 추적기는 함수 추적기와 비슷하지만 함수 진입과 반환을 모두 탐침한다. 각 `task_struct`에 동적으로 할당한 반환 주소 스택을 사용하며, 함수 진입 때 추적 대상 함수의 원래 반환 주소를 사용자 정의 탐침 주소로 바꾸고 원래 주소는 태스크의 반환 주소 스택에 저장한다.

함수의 양 끝을 모두 관찰하므로 함수 실행 시간을 측정할 수 있고, 신뢰할 수 있는 호출 스택을 바탕으로 함수 호출 그래프를 그릴 수 있다.

이 추적기는 이상한 커널 동작의 상세 원인을 찾을 때, 출처가 불분명한 지연을 조사할 때, 특정 함수가 택하는 경로를 빠르게 확인할 때, 또는 동작 중인 커널 내부를 살펴볼 때 유용하다.

  # tracer: function_graph
  #
  # CPU  DURATION                  FUNCTION CALLS
  # |     |   |                     |   |   |   |

   0)               |  sys_open() {
   0)               |    do_sys_open() {
   0)               |      getname() {
   0)               |        kmem_cache_alloc() {
   0)   1.382 us    |          __might_sleep();
   0)   2.478 us    |        }
   0)               |        strncpy_from_user() {
   0)               |          might_fault() {
   0)   1.389 us    |            __might_sleep();
   0)   2.553 us    |          }
   0)   3.807 us    |        }
   0)   7.876 us    |      }
   0)               |      alloc_fd() {
   0)   0.668 us    |        _spin_lock();
   0)   0.570 us    |        expand_files();
   0)   0.586 us    |        _spin_unlock();

기본 출력은 CPU 번호, 실행 시간, 중첩된 함수 호출을 보여 준다. 자식 호출이 있는 함수는 여는 중괄호와 닫는 중괄호로 범위를 나타내고, leaf 함수는 한 줄에 함수 이름과 실행 시간을 함께 표시한다.

출력 열은 필요에 따라 동적으로 켜고 끌 수 있으며 여러 옵션을 원하는 조합으로 사용할 수 있다.

`funcgraph-cpu`는 함수가 실행된 CPU 번호를 표시하며 기본적으로 켜져 있다. CPU가 바뀌며 추적되면 호출 순서가 뒤섞여 보일 수 있으므로 `tracing_cpumask`로 한 CPU만 추적하는 편이 나을 때가 있다. `nofuncgraph-cpu`로 숨기고 `funcgraph-cpu`로 다시 표시한다.

`funcgraph-duration`은 함수 실행 시간을 표시하며 기본적으로 켜져 있다. 자식 호출이 있는 함수는 닫는 중괄호 행에, leaf 함수는 같은 행에 시간을 표시한다. `nofuncgraph-duration`과 `funcgraph-duration`으로 제어한다.

`funcgraph-overhead`는 실행 시간이 정해진 임계값을 넘었을 때 duration 앞에 강조 문자를 붙인다. 이 옵션은 `funcgraph-duration`에 의존하며 `nofuncgraph-overhead`와 `funcgraph-overhead`로 제어한다.

    3) # 1837.709 us |          } /* __switch_to */
    3)               |          finish_task_switch() {
    3)   0.313 us    |            _raw_spin_unlock_irq();
    3)   3.177 us    |          }
    3) # 1889.063 us |        } /* __schedule */
    3) ! 140.417 us  |      } /* __schedule */
    3) # 2034.948 us |    } /* schedule */
    3) * 33998.59 us |  } /* schedule_preempt_disabled */

    [...]

    1)   0.260 us    |              msecs_to_jiffies();
    1)   0.313 us    |              __rcu_read_unlock();
    1) + 61.770 us   |            }
    1) + 64.479 us   |          }
    1)   0.313 us    |          rcu_bh_qs();
    1)   0.313 us    |          __local_bh_enable();
    1) ! 217.240 us  |        }
    1)   0.365 us    |        idle_cpu();
    1)               |        rcu_irq_exit() {
    1)   0.417 us    |          rcu_eqs_enter_common.isra.47();
    1)   3.125 us    |        }
    1) ! 227.812 us  |      }
    1) ! 457.395 us  |    }
    1) @ 119760.2 us |  }

    [...]

    2)               |    handle_IPI() {
    1)   6.979 us    |                  }
    2)   0.417 us    |      scheduler_ipi();
    1)   9.791 us    |                }
    1) + 12.917 us   |              }
    2)   3.490 us    |    }
    1) + 15.729 us   |            }
    1) + 18.542 us   |          }
    2) $ 3594274 us  |  }

강조 문자는 실행 시간이 10마이크로초, 100마이크로초, 1밀리초, 10밀리초, 100밀리초, 1초를 넘는 구간을 단계별로 구분한다.

  + means that the function exceeded 10 usecs.
  ! means that the function exceeded 100 usecs.
  # means that the function exceeded 1000 usecs.
  * means that the function exceeded 10 msecs.
  @ means that the function exceeded 100 msecs.
  $ means that the function exceeded 1 sec.
function graph 지연 강조 문자
문자임계값단위
+10 초과us
!100 초과us
#1000 초과us
*10 초과ms
@100 초과ms
$1 초과s

duration 임계값을 넘긴 함수를 한 글자로 빠르게 찾는다.

`funcgraph-proc`는 함수를 실행한 스레드의 명령 이름과 PID를 표시한다. 기본값은 꺼짐이며 `nofuncgraph-proc`로 숨기고 `funcgraph-proc`로 표시한다.

    # tracer: function_graph
    #
    # CPU  TASK/PID        DURATION                  FUNCTION CALLS
    # |    |    |           |   |                     |   |   |   |
    0)    sh-4802     |               |                  d_free() {
    0)    sh-4802     |               |                    call_rcu() {
    0)    sh-4802     |               |                      __call_rcu() {
    0)    sh-4802     |   0.616 us    |                        rcu_process_gp_end();
    0)    sh-4802     |   0.586 us    |                        check_for_new_grace_period();
    0)    sh-4802     |   2.899 us    |                      }
    0)    sh-4802     |   4.040 us    |                    }
    0)    sh-4802     |   5.151 us    |                  }
    0)    sh-4802     | + 49.370 us   |                }

`funcgraph-abstime`은 시스템 시작 이후의 절대 시각을 함수 진입·반환마다 표시한다. 기본값은 꺼짐이며 `nofuncgraph-abstime`과 `funcgraph-abstime`으로 제어한다.

    #
    #      TIME       CPU  DURATION                  FUNCTION CALLS
    #       |         |     |   |                     |   |   |   |
    360.774522 |   1)   0.541 us    |                                          }
    360.774522 |   1)   4.663 us    |                                        }
    360.774523 |   1)   0.541 us    |                                        __wake_up_bit();
    360.774524 |   1)   6.796 us    |                                      }
    360.774524 |   1)   7.952 us    |                                    }
    360.774525 |   1)   9.063 us    |                                  }
    360.774525 |   1)   0.615 us    |                                  journal_mark_dirty();
    360.774527 |   1)   0.578 us    |                                  __brelse();
    360.774528 |   1)               |                                  reiserfs_prepare_for_journal() {
    360.774528 |   1)               |                                    unlock_buffer() {
    360.774529 |   1)               |                                      wake_up_bit() {
    360.774529 |   1)               |                                        bit_waitqueue() {
    360.774530 |   1)   0.594 us    |                                          __phys_addr();

함수 시작 기록이 이미 링 버퍼 밖으로 밀려났다면 닫는 중괄호 뒤에 함수 이름을 항상 표시한다. 시작 기록이 버퍼 안에 있는 함수도 닫는 행에 이름을 보이게 하면 `grep`으로 함수 실행 시간을 찾기 쉬워진다.

이 동작은 `funcgraph-tail`로 켜며 기본값은 꺼짐이다. 기본 `nofuncgraph-tail`에서는 닫는 중괄호만 보인다.

    0)               |      putname() {
    0)               |        kmem_cache_free() {
    0)   0.518 us    |          __phys_addr();
    0)   1.757 us    |        }
    0)   2.861 us    |      }

`funcgraph-tail`을 켜면 닫는 중괄호 뒤에 주석 형식으로 해당 함수 이름이 붙는다.

    0)               |      putname() {
    0)               |        kmem_cache_free() {
    0)   0.518 us    |          __phys_addr();
    0)   1.757 us    |        } /* kmem_cache_free() */
    0)   2.861 us    |      } /* putname() */

`funcgraph-retval`은 각 추적 함수의 반환값을 등호 `=` 뒤에 표시한다. 시스템 호출 실패를 조사할 때 처음 오류 코드를 반환한 함수를 빠르게 찾는 데 특히 유용하다. `nofuncgraph-retval`로 숨기고 `funcgraph-retval`로 표시한다.

    1)               |    cgroup_migrate() {
    1)   0.651 us    |      cgroup_migrate_add_task(); /* = 0xffff93fcfd346c00 */
    1)               |      cgroup_migrate_execute() {
    1)               |        cpu_cgroup_can_attach() {
    1)               |          cgroup_taskset_first() {
    1)   0.732 us    |            cgroup_taskset_next(); /* = 0xffff93fc8fb20000 */
    1)   1.232 us    |          } /* cgroup_taskset_first = 0xffff93fc8fb20000 */
    1)   0.380 us    |          sched_rt_can_attach(); /* = 0x0 */
    1)   2.335 us    |        } /* cpu_cgroup_can_attach = -22 */
    1)   4.369 us    |      } /* cgroup_migrate_execute = -22 */
    1)   7.143 us    |    } /* cgroup_migrate = -22 */

예에서는 `cpu_cgroup_can_attach()`가 먼저 오류 코드 -22를 반환했고 뒤의 호출자들이 같은 오류를 전달한다. 이 최초 반환 함수를 읽으면 근본 원인을 좁힐 수 있다.

`funcgraph-retval-hex`가 꺼진 스마트 모드에서는 오류 코드는 부호 있는 10진수로, 그 밖의 값은 16진수로 표시한다. 옵션을 켜면 모든 반환값을 16진수로 표시한다.

    1)               |      cgroup_migrate() {
    1)   0.651 us    |        cgroup_migrate_add_task(); /* = 0xffff93fcfd346c00 */
    1)               |        cgroup_migrate_execute() {
    1)               |          cpu_cgroup_can_attach() {
    1)               |            cgroup_taskset_first() {
    1)   0.732 us    |              cgroup_taskset_next(); /* = 0xffff93fc8fb20000 */
    1)   1.232 us    |            } /* cgroup_taskset_first = 0xffff93fc8fb20000 */
    1)   0.380 us    |            sched_rt_can_attach(); /* = 0x0 */
    1)   2.335 us    |          } /* cpu_cgroup_can_attach = 0xffffffea */
    1)   4.369 us    |        } /* cgroup_migrate_execute = 0xffffffea */
    1)   7.143 us    |      } /* cgroup_migrate = 0xffffffea */

반환값 추적에는 현재 몇 가지 한계가 있다. 반환형이 `void`여도 값이 출력되므로 무시해야 하며, 반환값이 여러 레지스터에 나뉘어 저장되어도 첫 레지스터 값만 기록한다.

예를 들어 x86에서 64비트 반환값을 `eax`와 `edx`에 나누어 저장하면 낮은 32비트가 든 `eax`만 기록되고 높은 32비트가 든 `edx`는 빠진다.

arm64 AAPCS64처럼 반환형이 GPR보다 작을 때 narrowing을 소비자가 담당하는 호출 규약에서는 상위 비트가 UNKNOWN 값을 가질 수 있다. 특히 큰 타입을 명시적 또는 암묵적으로 잘라 64비트 GPR에 `u8`을 반환하면 비트 [63:8]이 임의 값일 수 있으므로 코드를 함께 확인해야 한다.

첫 번째 사례의 `narrow_to_u8()`은 `u64`를 `u8`로 암묵적으로 잘라 반환한다.

	u8 narrow_to_u8(u64 val)
	{
		// implicitly truncated
		return val;
	}

컴파일 결과는 반환 레지스터를 별도로 좁히지 않고 바로 `RET`할 수 있다.

	narrow_to_u8:
		< ... ftrace instrumentation ... >
		RET

따라서 입력 `0x123456789abcdef`를 `0xef`로 좁히려 해도 추적에는 상위 비트가 남은 `0x123456789abcdef`가 기록될 수 있다.

두 번째 사례의 `error_if_not_4g_aligned()`는 낮은 32비트가 0이 아니면 `-EINVAL`, 그렇지 않으면 0을 반환한다.

	int error_if_not_4g_aligned(u64 val)
	{
		if (val & GENMASK(31, 0))
			return -EINVAL;

		return 0;
	}

컴파일된 정렬 성공 경로는 `w0`만 검사한 뒤 `x0` 상위 비트를 정리하지 않고 반환할 수 있다.

	error_if_not_4g_aligned:
		CBNZ    w0, .Lnot_aligned
		RET			// bits [31:0] are zero, bits
					// [63:32] are UNKNOWN
	.Lnot_aligned:
		MOV    x0, #-EINVAL
		RET

이 경우 `0x2_0000_0000`을 전달하면 함수의 논리적 반환값은 0이지만 상위 비트가 남아 추적에는 `0x2_0000_0000`이 기록될 수 있다.

특정 함수에 설명을 남기려면 `trace_printk()`를 사용할 수 있다. 예를 들어 `__might_sleep()` 안에서 `<linux/ftrace.h>`를 포함하고 다음 호출을 추가한다.

	trace_printk("I'm a comment!\n")

그러면 함수 그래프 안에 주석 행이 삽입된다.

   1)               |             __might_sleep() {
   1)               |                /* I'm a comment! */
   1)   1.449 us    |             }

특정 함수나 태스크만 추적하는 기능 등 이 추적기의 추가 활용법은 이어지는 `dynamic ftrace` 절에서 설명한다.

주요 function_graph 표시 옵션
옵션표시 내용기본값
funcgraph-cpu실행 CPU
funcgraph-duration함수 실행 시간
funcgraph-overhead지연 임계값 강조 문자duration에 의존
funcgraph-proc태스크 명령과 PID
funcgraph-abstime시스템 시작 이후 절대 시각
funcgraph-tail닫는 중괄호 뒤 함수 이름
funcgraph-retval함수 반환값
funcgraph-retval-hex반환값을 항상 16진수로 표시

그래프의 열과 닫는 행, 반환값 표현을 목적에 맞게 조합한다.

반환값 추적 한계
상황기록되는 값주의
void 반환레지스터의 값이 표시될 수 있음무시
여러 레지스터 반환첫 레지스터만나머지 비트 누락
GPR보다 작은 타입상위 UNKNOWN 비트 포함 가능소비자의 narrowing 확인
오류 코드 스마트 모드부호 있는 10진수정상 값은 16진수

출력값을 함수의 선언형과 호출 규약에 맞춰 해석해야 한다.

함수 진입과 반환 탐침
추적 함수 진입원래 반환 주소 저장
원래 반환 주소 저장반환 주소를 사용자 탐침으로 교체
반환 주소를 사용자 탐침으로 교체함수 본문과 자식 호출 실행
함수 본문과 자식 호출 실행사용자 반환 탐침
사용자 반환 탐침실행 시간·반환값 기록
실행 시간·반환값 기록원래 반환 주소로 복귀

원래 반환 주소를 태스크별 스택에 보관해 안정적인 호출 그래프를 만든다.

반환 오류 원인 추적
leaf 또는 하위 함수 반환funcgraph-retval 확인
첫 음수 오류 코드해당 함수 구현 조사
해당 함수 구현 조사호출자에게 전달되는 같은 오류 확인

중첩 그래프에서 오류를 처음 만든 함수를 위쪽 호출자로 거슬러 읽는다.

function graph tracer
---------------------------

This tracer is similar to the function tracer except that it
probes a function on its entry and its exit. This is done by
using a dynamically allocated stack of return addresses in each
task_struct. On function entry the tracer overwrites the return
address of each function traced to set a custom probe. Thus the
original return address is stored on the stack of return address
in the task_struct.

Probing on both ends of a function leads to special features
such as:

- measure of a function's time execution
- having a reliable call stack to draw function calls graph

This tracer is useful in several situations:

- you want to find the reason of a strange kernel behavior and
  need to see what happens in detail on any areas (or specific
  ones).

- you are experiencing weird latencies but it's difficult to
  find its origin.

- you want to find quickly which path is taken by a specific
  function

- you just want to peek inside a working kernel and want to see
  what happens there.

::

  # tracer: function_graph
  #
  # CPU  DURATION                  FUNCTION CALLS
  # |     |   |                     |   |   |   |

   0)               |  sys_open() {
   0)               |    do_sys_open() {
   0)               |      getname() {
   0)               |        kmem_cache_alloc() {
   0)   1.382 us    |          __might_sleep();
   0)   2.478 us    |        }
   0)               |        strncpy_from_user() {
   0)               |          might_fault() {
   0)   1.389 us    |            __might_sleep();
   0)   2.553 us    |          }
   0)   3.807 us    |        }
   0)   7.876 us    |      }
   0)               |      alloc_fd() {
   0)   0.668 us    |        _spin_lock();
   0)   0.570 us    |        expand_files();
   0)   0.586 us    |        _spin_unlock();


There are several columns that can be dynamically
enabled/disabled. You can use every combination of options you
want, depending on your needs.

- The cpu number on which the function executed is default
  enabled.  It is sometimes better to only trace one cpu (see
  tracing_cpumask file) or you might sometimes see unordered
  function calls while cpu tracing switch.

	- hide: echo nofuncgraph-cpu > trace_options
	- show: echo funcgraph-cpu > trace_options

- The duration (function's time of execution) is displayed on
  the closing bracket line of a function or on the same line
  than the current function in case of a leaf one. It is default
  enabled.

	- hide: echo nofuncgraph-duration > trace_options
	- show: echo funcgraph-duration > trace_options

- The overhead field precedes the duration field in case of
  reached duration thresholds.

	- hide: echo nofuncgraph-overhead > trace_options
	- show: echo funcgraph-overhead > trace_options
	- depends on: funcgraph-duration

  ie::

    3) # 1837.709 us |          } /* __switch_to */
    3)               |          finish_task_switch() {
    3)   0.313 us    |            _raw_spin_unlock_irq();
    3)   3.177 us    |          }
    3) # 1889.063 us |        } /* __schedule */
    3) ! 140.417 us  |      } /* __schedule */
    3) # 2034.948 us |    } /* schedule */
    3) * 33998.59 us |  } /* schedule_preempt_disabled */

    [...]

    1)   0.260 us    |              msecs_to_jiffies();
    1)   0.313 us    |              __rcu_read_unlock();
    1) + 61.770 us   |            }
    1) + 64.479 us   |          }
    1)   0.313 us    |          rcu_bh_qs();
    1)   0.313 us    |          __local_bh_enable();
    1) ! 217.240 us  |        }
    1)   0.365 us    |        idle_cpu();
    1)               |        rcu_irq_exit() {
    1)   0.417 us    |          rcu_eqs_enter_common.isra.47();
    1)   3.125 us    |        }
    1) ! 227.812 us  |      }
    1) ! 457.395 us  |    }
    1) @ 119760.2 us |  }

    [...]

    2)               |    handle_IPI() {
    1)   6.979 us    |                  }
    2)   0.417 us    |      scheduler_ipi();
    1)   9.791 us    |                }
    1) + 12.917 us   |              }
    2)   3.490 us    |    }
    1) + 15.729 us   |            }
    1) + 18.542 us   |          }
    2) $ 3594274 us  |  }

Flags::

  + means that the function exceeded 10 usecs.
  ! means that the function exceeded 100 usecs.
  # means that the function exceeded 1000 usecs.
  * means that the function exceeded 10 msecs.
  @ means that the function exceeded 100 msecs.
  $ means that the function exceeded 1 sec.


- The task/pid field displays the thread cmdline and pid which
  executed the function. It is default disabled.

	- hide: echo nofuncgraph-proc > trace_options
	- show: echo funcgraph-proc > trace_options

  ie::

    # tracer: function_graph
    #
    # CPU  TASK/PID        DURATION                  FUNCTION CALLS
    # |    |    |           |   |                     |   |   |   |
    0)    sh-4802     |               |                  d_free() {
    0)    sh-4802     |               |                    call_rcu() {
    0)    sh-4802     |               |                      __call_rcu() {
    0)    sh-4802     |   0.616 us    |                        rcu_process_gp_end();
    0)    sh-4802     |   0.586 us    |                        check_for_new_grace_period();
    0)    sh-4802     |   2.899 us    |                      }
    0)    sh-4802     |   4.040 us    |                    }
    0)    sh-4802     |   5.151 us    |                  }
    0)    sh-4802     | + 49.370 us   |                }


- The absolute time field is an absolute timestamp given by the
  system clock since it started. A snapshot of this time is
  given on each entry/exit of functions

	- hide: echo nofuncgraph-abstime > trace_options
	- show: echo funcgraph-abstime > trace_options

  ie::

    #
    #      TIME       CPU  DURATION                  FUNCTION CALLS
    #       |         |     |   |                     |   |   |   |
    360.774522 |   1)   0.541 us    |                                          }
    360.774522 |   1)   4.663 us    |                                        }
    360.774523 |   1)   0.541 us    |                                        __wake_up_bit();
    360.774524 |   1)   6.796 us    |                                      }
    360.774524 |   1)   7.952 us    |                                    }
    360.774525 |   1)   9.063 us    |                                  }
    360.774525 |   1)   0.615 us    |                                  journal_mark_dirty();
    360.774527 |   1)   0.578 us    |                                  __brelse();
    360.774528 |   1)               |                                  reiserfs_prepare_for_journal() {
    360.774528 |   1)               |                                    unlock_buffer() {
    360.774529 |   1)               |                                      wake_up_bit() {
    360.774529 |   1)               |                                        bit_waitqueue() {
    360.774530 |   1)   0.594 us    |                                          __phys_addr();


The function name is always displayed after the closing bracket
for a function if the start of that function is not in the
trace buffer.

Display of the function name after the closing bracket may be
enabled for functions whose start is in the trace buffer,
allowing easier searching with grep for function durations.
It is default disabled.

	- hide: echo nofuncgraph-tail > trace_options
	- show: echo funcgraph-tail > trace_options

  Example with nofuncgraph-tail (default)::

    0)               |      putname() {
    0)               |        kmem_cache_free() {
    0)   0.518 us    |          __phys_addr();
    0)   1.757 us    |        }
    0)   2.861 us    |      }

  Example with funcgraph-tail::

    0)               |      putname() {
    0)               |        kmem_cache_free() {
    0)   0.518 us    |          __phys_addr();
    0)   1.757 us    |        } /* kmem_cache_free() */
    0)   2.861 us    |      } /* putname() */

The return value of each traced function can be displayed after
an equal sign "=". When encountering system call failures, it
can be very helpful to quickly locate the function that first
returns an error code.

	- hide: echo nofuncgraph-retval > trace_options
	- show: echo funcgraph-retval > trace_options

  Example with funcgraph-retval::

    1)               |    cgroup_migrate() {
    1)   0.651 us    |      cgroup_migrate_add_task(); /* = 0xffff93fcfd346c00 */
    1)               |      cgroup_migrate_execute() {
    1)               |        cpu_cgroup_can_attach() {
    1)               |          cgroup_taskset_first() {
    1)   0.732 us    |            cgroup_taskset_next(); /* = 0xffff93fc8fb20000 */
    1)   1.232 us    |          } /* cgroup_taskset_first = 0xffff93fc8fb20000 */
    1)   0.380 us    |          sched_rt_can_attach(); /* = 0x0 */
    1)   2.335 us    |        } /* cpu_cgroup_can_attach = -22 */
    1)   4.369 us    |      } /* cgroup_migrate_execute = -22 */
    1)   7.143 us    |    } /* cgroup_migrate = -22 */

The above example shows that the function cpu_cgroup_can_attach
returned the error code -22 firstly, then we can read the code
of this function to get the root cause.

When the option funcgraph-retval-hex is not set, the return value can
be displayed in a smart way. Specifically, if it is an error code,
it will be printed in signed decimal format, otherwise it will
printed in hexadecimal format.

	- smart: echo nofuncgraph-retval-hex > trace_options
	- hexadecimal: echo funcgraph-retval-hex > trace_options

  Example with funcgraph-retval-hex::

    1)               |      cgroup_migrate() {
    1)   0.651 us    |        cgroup_migrate_add_task(); /* = 0xffff93fcfd346c00 */
    1)               |        cgroup_migrate_execute() {
    1)               |          cpu_cgroup_can_attach() {
    1)               |            cgroup_taskset_first() {
    1)   0.732 us    |              cgroup_taskset_next(); /* = 0xffff93fc8fb20000 */
    1)   1.232 us    |            } /* cgroup_taskset_first = 0xffff93fc8fb20000 */
    1)   0.380 us    |            sched_rt_can_attach(); /* = 0x0 */
    1)   2.335 us    |          } /* cpu_cgroup_can_attach = 0xffffffea */
    1)   4.369 us    |        } /* cgroup_migrate_execute = 0xffffffea */
    1)   7.143 us    |      } /* cgroup_migrate = 0xffffffea */

At present, there are some limitations when using the funcgraph-retval
option, and these limitations will be eliminated in the future:

- Even if the function return type is void, a return value will still
  be printed, and you can just ignore it.

- Even if return values are stored in multiple registers, only the
  value contained in the first register will be recorded and printed.
  To illustrate, in the x86 architecture, eax and edx are used to store
  a 64-bit return value, with the lower 32 bits saved in eax and the
  upper 32 bits saved in edx. However, only the value stored in eax
  will be recorded and printed.

- In certain procedure call standards, such as arm64's AAPCS64, when a
  type is smaller than a GPR, it is the responsibility of the consumer
  to perform the narrowing, and the upper bits may contain UNKNOWN values.
  Therefore, it is advisable to check the code for such cases. For instance,
  when using a u8 in a 64-bit GPR, bits [63:8] may contain arbitrary values,
  especially when larger types are truncated, whether explicitly or implicitly.
  Here are some specific cases to illustrate this point:

  **Case One**:

  The function narrow_to_u8 is defined as follows::

	u8 narrow_to_u8(u64 val)
	{
		// implicitly truncated
		return val;
	}

  It may be compiled to::

	narrow_to_u8:
		< ... ftrace instrumentation ... >
		RET

  If you pass 0x123456789abcdef to this function and want to narrow it,
  it may be recorded as 0x123456789abcdef instead of 0xef.

  **Case Two**:

  The function error_if_not_4g_aligned is defined as follows::

	int error_if_not_4g_aligned(u64 val)
	{
		if (val & GENMASK(31, 0))
			return -EINVAL;

		return 0;
	}

  It could be compiled to::

	error_if_not_4g_aligned:
		CBNZ    w0, .Lnot_aligned
		RET			// bits [31:0] are zero, bits
					// [63:32] are UNKNOWN
	.Lnot_aligned:
		MOV    x0, #-EINVAL
		RET

  When passing 0x2_0000_0000 to it, the return value may be recorded as
  0x2_0000_0000 instead of 0.

You can put some comments on specific functions by using
trace_printk() For example, if you want to put a comment inside
the __might_sleep() function, you just have to include
<linux/ftrace.h> and call trace_printk() inside __might_sleep()::

	trace_printk("I'm a comment!\n")

will produce::

   1)               |             __might_sleep() {
   1)               |                /* I'm a comment! */
   1)   1.449 us    |             }


You might find other useful features for this tracer in the
following "dynamic ftrace" section such as tracing only specific
functions or tasks.

동적 ftrace와 함수 필터

2925-3214

`CONFIG_DYNAMIC_FTRACE`가 설정되면 함수 추적을 꺼 둔 동안의 실행 오버헤드는 사실상 없다. GCC의 `-pg` 옵션은 모든 커널 함수 시작점에 `mcount` 호출을 넣지만, 추적이 비활성화된 초기 상태의 `mcount`는 곧바로 반환하는 단순 스텁으로 동작한다. Ftrace를 켜는 커널 빌드에는 이 `-pg` 옵션이 포함된다.

컴파일할 때 `scripts/` 디렉터리의 `recordmcount` 프로그램이 각 C 오브젝트의 ELF 헤더를 분석하여 `.text` 섹션에서 `mcount`를 호출하는 위치를 찾는다. GCC 4.6부터 x86에서는 `-mfentry`를 사용할 수 있으며, 이 경우 스택 프레임을 만들기 전에 `mcount` 대신 `__fentry__`를 호출한다.

모든 섹션과 함수가 추적 대상인 것은 아니다. `notrace`로 지정했거나 다른 방식으로 차단된 영역과 모든 인라인 함수는 추적하지 않는다. 실제로 선택할 수 있는 함수는 `available_filter_functions`에서 확인한다.

`recordmcount`는 `.text`의 모든 `mcount` 또는 `fentry` 호출 지점 참조를 담는 `__mcount_loc` 섹션을 만들고 원래 오브젝트에 다시 링크한다. 마지막 커널 링크 단계에서는 각 오브젝트의 참조를 하나의 테이블로 합친다.

부팅할 때 SMP를 초기화하기 전에 동적 ftrace 코드가 이 테이블을 훑어 모든 호출 지점을 NOP로 바꾸고 위치를 기록하여 `available_filter_functions` 목록에 추가한다. 모듈도 실행되기 전 로드 과정에서 같은 처리를 받으며, 언로드할 때는 해당 함수를 ftrace 함수 목록에서 자동으로 제거하므로 모듈 작성자가 별도로 처리할 필요가 없다.

추적을 켤 때 함수 추적 지점을 수정하는 방식은 아키텍처에 따라 다르다. 예전 방식은 `kstop_machine`으로 다른 CPU가 수정 중인 코드를 실행하지 못하게 한 뒤 NOP를 ftrace 인프라 호출로 다시 패치한다. 수정 명령이 캐시나 페이지 경계를 가로지르면 경쟁 상태가 특히 위험하기 때문이다. 이때 원래의 단순 `mcount` 스텁을 호출하는 것이 아니라 ftrace 내부로 진입한다.

새 방식은 수정할 위치에 breakpoint를 놓고 모든 CPU를 동기화한 다음 breakpoint가 덮지 않은 나머지 명령을 바꾼다. 다시 모든 CPU를 동기화하고 완성된 ftrace 호출 지점 명령으로 breakpoint를 제거한다. 일부 아키텍처는 별도 동기화 없이도 다른 CPU가 동시에 실행하는 가운데 새 코드를 안전하게 덮어쓸 수 있다.

함수 호출 지점을 모두 기록해 두었기 때문에 추적할 함수만 선택하고 나머지 `mcount` 호출 지점은 NOP 상태로 유지할 수 있다. 포함 필터와 제외 필터, 선택 가능한 함수 목록은 다음 파일을 사용한다.

동적 함수 필터 파일
파일역할
`set_ftrace_filter`지정한 함수의 추적을 활성화
`set_ftrace_notrace`지정한 함수를 추적 대상에서 제외
`available_filter_functions`필터에 넣을 수 있는 함수 목록

추적할 함수 집합과 제외할 함수 집합을 별도로 관리한다.

`available_filter_functions`를 읽으면 현재 선택 가능한 함수가 나열된다.

  # cat available_filter_functions
  put_prev_task_idle
  kmem_cache_create
  pick_next_task_rt
  cpus_read_lock
  pick_next_task_fair
  mutex_lock
  [...]

`sys_nanosleep`과 `hrtimer_interrupt`만 관심 대상이라면 두 이름을 `set_ftrace_filter`에 쓰고 function tracer를 시작한다. 아래 결과에는 요청한 두 함수만 기록된다.

  # echo sys_nanosleep hrtimer_interrupt > set_ftrace_filter
  # echo function > current_tracer
  # echo 1 > tracing_on
  # usleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: function
  #
  # entries-in-buffer/entries-written: 5/5   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            usleep-2665  [001] ....  4186.475355: sys_nanosleep <-system_call_fastpath
            <idle>-0     [001] d.h1  4186.475409: hrtimer_interrupt <-smp_apic_timer_interrupt
            usleep-2665  [001] d.h1  4186.475426: hrtimer_interrupt <-smp_apic_timer_interrupt
            <idle>-0     [003] d.h1  4186.475426: hrtimer_interrupt <-smp_apic_timer_interrupt
            <idle>-0     [002] d.h1  4186.475427: hrtimer_interrupt <-smp_apic_timer_interrupt

현재 실제로 추적하는 함수 집합은 `set_ftrace_filter`를 읽어 확인할 수 있다.

  # cat set_ftrace_filter
  hrtimer_interrupt
  sys_nanosleep

필터는 `glob(7)` 방식의 와일드카드 일치를 지원한다. 셸이 현재 디렉터리의 파일 이름으로 와일드카드를 먼저 확장하지 않도록 패턴을 따옴표로 감싸는 편이 안전하다.

함수 이름 glob 규칙
패턴일치 조건
`<match>*``<match>`로 시작
`*<match>``<match>`로 끝남
`*<match>*`이름 안에 `<match>`가 포함됨
`<match1>*<match2>``<match1>`로 시작하고 `<match2>`로 끝남

와일드카드 위치에 따라 함수 이름의 시작, 끝 또는 중간을 일치시킨다.

다음 명령은 이름이 `hrtimer_`로 시작하는 함수를 필터에 설정한다.

  # echo 'hrtimer_*' > set_ftrace_filter

그 결과 여러 `hrtimer_*` 함수가 기록된다.

  # tracer: function
  #
  # entries-in-buffer/entries-written: 897/897   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            <idle>-0     [003] dN.1  4228.547803: hrtimer_cancel <-tick_nohz_idle_exit
            <idle>-0     [003] dN.1  4228.547804: hrtimer_try_to_cancel <-hrtimer_cancel
            <idle>-0     [003] dN.2  4228.547805: hrtimer_force_reprogram <-__remove_hrtimer
            <idle>-0     [003] dN.1  4228.547805: hrtimer_forward <-tick_nohz_idle_exit
            <idle>-0     [003] dN.1  4228.547805: hrtimer_start_range_ns <-hrtimer_start_expires.constprop.11
            <idle>-0     [003] d..1  4228.547858: hrtimer_get_next_event <-get_next_timer_interrupt
            <idle>-0     [003] d..1  4228.547859: hrtimer_start <-__tick_nohz_idle_enter
            <idle>-0     [003] d..2  4228.547860: hrtimer_force_reprogram <-__rem

이때 `sys_nanosleep`이 사라진 이유는 셸의 `>`와 `>>`가 평소와 똑같이 동작하기 때문이다. `>`는 필터를 새로 쓰고 `>>`는 기존 필터에 추가한다. 실제 필터 목록도 `hrtimer_*`로 다시 작성되어 있다.

  # cat set_ftrace_filter
  hrtimer_run_queues
  hrtimer_run_pending
  hrtimer_setup
  hrtimer_cancel
  hrtimer_try_to_cancel
  hrtimer_forward
  hrtimer_start
  hrtimer_reprogram
  hrtimer_force_reprogram
  hrtimer_get_next_event
  hrtimer_interrupt
  hrtimer_nanosleep
  hrtimer_wakeup
  hrtimer_get_remaining
  hrtimer_get_res
  hrtimer_init_sleeper

모든 함수를 다시 기록하도록 포함 필터를 비우려면 빈 값을 `set_ftrace_filter`에 새로 쓴다.

 # echo > set_ftrace_filter
 # cat set_ftrace_filter
 #

여러 조건을 누적하려면 첫 조건은 `>`로 설정하고 뒤 조건은 `>>`로 추가한다. 아래에서는 `sys_nanosleep`을 먼저 넣고 `hrtimer_*` 함수들을 덧붙인다.

  # echo sys_nanosleep > set_ftrace_filter
  # cat set_ftrace_filter
  sys_nanosleep
  # echo 'hrtimer_*' >> set_ftrace_filter
  # cat set_ftrace_filter
  hrtimer_run_queues
  hrtimer_run_pending
  hrtimer_setup
  hrtimer_cancel
  hrtimer_try_to_cancel
  hrtimer_forward
  hrtimer_start
  hrtimer_reprogram
  hrtimer_force_reprogram
  hrtimer_get_next_event
  hrtimer_interrupt
  sys_nanosleep
  hrtimer_nanosleep
  hrtimer_wakeup
  hrtimer_get_remaining
  hrtimer_get_res
  hrtimer_init_sleeper
필터 쓰기 연산
연산결과
`>`기존 필터를 지우고 새 조건으로 교체
`>>`기존 필터를 유지하고 새 조건을 추가
빈 값을 `>`로 쓰기포함 필터를 지워 모든 함수가 다시 대상이 됨

리디렉션 연산자가 기존 선택 집합을 유지할지 결정한다.

`set_ftrace_notrace`는 일치하는 함수를 추적하지 못하게 한다. 다음 예는 이름에 `preempt` 또는 `lock`이 들어가는 함수를 제외한다.

  # echo '*preempt*' '*lock*' > set_ftrace_notrace

출력에는 더 이상 lock 또는 preempt 관련 함수 추적이 나타나지 않는다.

  # tracer: function
  #
  # entries-in-buffer/entries-written: 39608/39608   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1994  [000] ....  4342.324896: file_ra_state_init <-do_dentry_open
              bash-1994  [000] ....  4342.324897: open_check_o_direct <-do_last
              bash-1994  [000] ....  4342.324897: ima_file_check <-do_last
              bash-1994  [000] ....  4342.324898: process_measurement <-ima_file_check
              bash-1994  [000] ....  4342.324898: ima_get_action <-process_measurement
              bash-1994  [000] ....  4342.324898: ima_match_policy <-ima_get_action
              bash-1994  [000] ....  4342.324899: do_truncate <-do_last
              bash-1994  [000] ....  4342.324899: setattr_should_drop_suidgid <-do_truncate
              bash-1994  [000] ....  4342.324899: notify_change <-do_truncate
              bash-1994  [000] ....  4342.324900: current_fs_time <-notify_change
              bash-1994  [000] ....  4342.324900: current_kernel_time <-current_fs_time
              bash-1994  [000] ....  4342.324900: timespec_trunc <-current_fs_time

문자열 필터는 전달된 이름과 비교하기 전에 함수 주소를 찾아야 하므로 대량 설정 비용이 크다. 수천 개의 특정 함수를 한꺼번에 고를 때는 숫자 인덱스를 대신 사용할 수 있다. 인덱스는 `available_filter_functions` 순서와 대응하는 내부 함수 배열의 위치이며 문자열 처리를 거치지 않는다.

숫자 `1`을 쓰면 `available_filter_functions`의 첫 번째 함수를 선택한다.

  # echo 1 > set_ftrace_filter

여러 숫자를 한 번에 쓰면 해당 위치의 함수들이 함께 선택된다. 다음 예는 첫 번째와 50번째 함수를 고른다.

  # head -1 available_filter_functions
  trace_initcall_finish_cb

  # cat set_ftrace_filter
  trace_initcall_finish_cb

  # head -50 available_filter_functions | tail -1
  x86_pmu_commit_txn

  # echo 1 50 > set_ftrace_filter
  # cat set_ftrace_filter
  trace_initcall_finish_cb
  x86_pmu_commit_txn
동적 ftrace 호출 지점 구성
`-pg` 또는 `-mfentry`로 호출 지점 생성`recordmcount`가 ELF `.text` 분석
`recordmcount`가 ELF `.text` 분석`__mcount_loc` 테이블 링크
`__mcount_loc` 테이블 링크부팅 전 호출 지점을 NOP로 패치
부팅 전 호출 지점을 NOP로 패치선택한 함수만 ftrace 호출로 패치
선택한 함수만 ftrace 호출로 패치나머지 함수는 NOP 유지

빌드 시 수집한 호출 지점을 부팅 때 NOP로 만들고 필요한 함수만 ftrace 호출로 전환한다.

dynamic ftrace
--------------

If CONFIG_DYNAMIC_FTRACE is set, the system will run with
virtually no overhead when function tracing is disabled. The way
this works is the mcount function call (placed at the start of
every kernel function, produced by the -pg switch in gcc),
starts of pointing to a simple return. (Enabling FTRACE will
include the -pg switch in the compiling of the kernel.)

At compile time every C file object is run through the
recordmcount program (located in the scripts directory). This
program will parse the ELF headers in the C object to find all
the locations in the .text section that call mcount. Starting
with gcc version 4.6, the -mfentry has been added for x86, which
calls "__fentry__" instead of "mcount". Which is called before
the creation of the stack frame.

Note, not all sections are traced. They may be prevented by either
a notrace, or blocked another way and all inline functions are not
traced. Check the "available_filter_functions" file to see what functions
can be traced.

A section called "__mcount_loc" is created that holds
references to all the mcount/fentry call sites in the .text section.
The recordmcount program re-links this section back into the
original object. The final linking stage of the kernel will add all these
references into a single table.

On boot up, before SMP is initialized, the dynamic ftrace code
scans this table and updates all the locations into nops. It
also records the locations, which are added to the
available_filter_functions list.  Modules are processed as they
are loaded and before they are executed.  When a module is
unloaded, it also removes its functions from the ftrace function
list. This is automatic in the module unload code, and the
module author does not need to worry about it.

When tracing is enabled, the process of modifying the function
tracepoints is dependent on architecture. The old method is to use
kstop_machine to prevent races with the CPUs executing code being
modified (which can cause the CPU to do undesirable things, especially
if the modified code crosses cache (or page) boundaries), and the nops are
patched back to calls. But this time, they do not call mcount
(which is just a function stub). They now call into the ftrace
infrastructure.

The new method of modifying the function tracepoints is to place
a breakpoint at the location to be modified, sync all CPUs, modify
the rest of the instruction not covered by the breakpoint. Sync
all CPUs again, and then remove the breakpoint with the finished
version to the ftrace call site.

Some archs do not even need to monkey around with the synchronization,
and can just slap the new code on top of the old without any
problems with other CPUs executing it at the same time.

One special side-effect to the recording of the functions being
traced is that we can now selectively choose which functions we
wish to trace and which ones we want the mcount calls to remain
as nops.

Two files are used, one for enabling and one for disabling the
tracing of specified functions. They are:

  set_ftrace_filter

and

  set_ftrace_notrace

A list of available functions that you can add to these files is
listed in:

   available_filter_functions

::

  # cat available_filter_functions
  put_prev_task_idle
  kmem_cache_create
  pick_next_task_rt
  cpus_read_lock
  pick_next_task_fair
  mutex_lock
  [...]

If I am only interested in sys_nanosleep and hrtimer_interrupt::

  # echo sys_nanosleep hrtimer_interrupt > set_ftrace_filter
  # echo function > current_tracer
  # echo 1 > tracing_on
  # usleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: function
  #
  # entries-in-buffer/entries-written: 5/5   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            usleep-2665  [001] ....  4186.475355: sys_nanosleep <-system_call_fastpath
            <idle>-0     [001] d.h1  4186.475409: hrtimer_interrupt <-smp_apic_timer_interrupt
            usleep-2665  [001] d.h1  4186.475426: hrtimer_interrupt <-smp_apic_timer_interrupt
            <idle>-0     [003] d.h1  4186.475426: hrtimer_interrupt <-smp_apic_timer_interrupt
            <idle>-0     [002] d.h1  4186.475427: hrtimer_interrupt <-smp_apic_timer_interrupt

To see which functions are being traced, you can cat the file:
::

  # cat set_ftrace_filter
  hrtimer_interrupt
  sys_nanosleep


Perhaps this is not enough. The filters also allow glob(7) matching.

  ``<match>*``
	will match functions that begin with <match>
  ``*<match>``
	will match functions that end with <match>
  ``*<match>*``
	will match functions that have <match> in it
  ``<match1>*<match2>``
	will match functions that begin with <match1> and end with <match2>

.. note::
      It is better to use quotes to enclose the wild cards,
      otherwise the shell may expand the parameters into names
      of files in the local directory.

::

  # echo 'hrtimer_*' > set_ftrace_filter

Produces::

  # tracer: function
  #
  # entries-in-buffer/entries-written: 897/897   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            <idle>-0     [003] dN.1  4228.547803: hrtimer_cancel <-tick_nohz_idle_exit
            <idle>-0     [003] dN.1  4228.547804: hrtimer_try_to_cancel <-hrtimer_cancel
            <idle>-0     [003] dN.2  4228.547805: hrtimer_force_reprogram <-__remove_hrtimer
            <idle>-0     [003] dN.1  4228.547805: hrtimer_forward <-tick_nohz_idle_exit
            <idle>-0     [003] dN.1  4228.547805: hrtimer_start_range_ns <-hrtimer_start_expires.constprop.11
            <idle>-0     [003] d..1  4228.547858: hrtimer_get_next_event <-get_next_timer_interrupt
            <idle>-0     [003] d..1  4228.547859: hrtimer_start <-__tick_nohz_idle_enter
            <idle>-0     [003] d..2  4228.547860: hrtimer_force_reprogram <-__rem

Notice that we lost the sys_nanosleep.
::

  # cat set_ftrace_filter
  hrtimer_run_queues
  hrtimer_run_pending
  hrtimer_setup
  hrtimer_cancel
  hrtimer_try_to_cancel
  hrtimer_forward
  hrtimer_start
  hrtimer_reprogram
  hrtimer_force_reprogram
  hrtimer_get_next_event
  hrtimer_interrupt
  hrtimer_nanosleep
  hrtimer_wakeup
  hrtimer_get_remaining
  hrtimer_get_res
  hrtimer_init_sleeper


This is because the '>' and '>>' act just like they do in bash.
To rewrite the filters, use '>'
To append to the filters, use '>>'

To clear out a filter so that all functions will be recorded
again::

 # echo > set_ftrace_filter
 # cat set_ftrace_filter
 #

Again, now we want to append.

::

  # echo sys_nanosleep > set_ftrace_filter
  # cat set_ftrace_filter
  sys_nanosleep
  # echo 'hrtimer_*' >> set_ftrace_filter
  # cat set_ftrace_filter
  hrtimer_run_queues
  hrtimer_run_pending
  hrtimer_setup
  hrtimer_cancel
  hrtimer_try_to_cancel
  hrtimer_forward
  hrtimer_start
  hrtimer_reprogram
  hrtimer_force_reprogram
  hrtimer_get_next_event
  hrtimer_interrupt
  sys_nanosleep
  hrtimer_nanosleep
  hrtimer_wakeup
  hrtimer_get_remaining
  hrtimer_get_res
  hrtimer_init_sleeper


The set_ftrace_notrace prevents those functions from being
traced.
::

  # echo '*preempt*' '*lock*' > set_ftrace_notrace

Produces::

  # tracer: function
  #
  # entries-in-buffer/entries-written: 39608/39608   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1994  [000] ....  4342.324896: file_ra_state_init <-do_dentry_open
              bash-1994  [000] ....  4342.324897: open_check_o_direct <-do_last
              bash-1994  [000] ....  4342.324897: ima_file_check <-do_last
              bash-1994  [000] ....  4342.324898: process_measurement <-ima_file_check
              bash-1994  [000] ....  4342.324898: ima_get_action <-process_measurement
              bash-1994  [000] ....  4342.324898: ima_match_policy <-ima_get_action
              bash-1994  [000] ....  4342.324899: do_truncate <-do_last
              bash-1994  [000] ....  4342.324899: setattr_should_drop_suidgid <-do_truncate
              bash-1994  [000] ....  4342.324899: notify_change <-do_truncate
              bash-1994  [000] ....  4342.324900: current_fs_time <-notify_change
              bash-1994  [000] ....  4342.324900: current_kernel_time <-current_fs_time
              bash-1994  [000] ....  4342.324900: timespec_trunc <-current_fs_time

We can see that there's no more lock or preempt tracing.

Selecting function filters via index
------------------------------------

Because processing of strings is expensive (the address of the function
needs to be looked up before comparing to the string being passed in),
an index can be used as well to enable functions. This is useful in the
case of setting thousands of specific functions at a time. By passing
in a list of numbers, no string processing will occur. Instead, the function
at the specific location in the internal array (which corresponds to the
functions in the "available_filter_functions" file), is selected.

::

  # echo 1 > set_ftrace_filter

Will select the first function listed in "available_filter_functions"

::

  # head -1 available_filter_functions
  trace_initcall_finish_cb

  # cat set_ftrace_filter
  trace_initcall_finish_cb

  # head -50 available_filter_functions | tail -1
  x86_pmu_commit_txn

  # echo 1 50 > set_ftrace_filter
  # cat set_ftrace_filter
  trace_initcall_finish_cb
  x86_pmu_commit_txn

function graph 동적 필터와 전역 스위치

3215-3299

앞의 동적 ftrace 설명은 function tracer와 function graph tracer에 모두 적용되지만, function graph tracer에서만 사용할 수 있는 기능도 있다.

특정 함수 하나와 그 함수가 호출하는 모든 자식만 추적하려면 함수 이름을 `set_graph_function`에 쓴다. 다음 설정은 `__do_fault`를 그래프의 시작 함수로 선택한다.

 echo __do_fault > set_graph_function

결과에는 `__do_fault()` 아래에서 실행된 `filemap_fault()`, `find_lock_page()` 등의 호출 계층과 각 실행 시간이 펼쳐져 표시된다.

   0)               |  __do_fault() {
   0)               |    filemap_fault() {
   0)               |      find_lock_page() {
   0)   0.804 us    |        find_get_page();
   0)               |        __might_sleep() {
   0)   1.329 us    |        }
   0)   3.904 us    |      }
   0)   4.979 us    |    }
   0)   0.653 us    |    _spin_lock();
   0)   0.578 us    |    page_add_file_rmap();
   0)   0.525 us    |    native_set_pte_at();
   0)   0.585 us    |    _spin_unlock();
   0)               |    unlock_page() {
   0)   0.541 us    |      page_waitqueue();
   0)   0.639 us    |      __wake_up_bit();
   0)   2.786 us    |    }
   0) + 14.237 us   |  }
   0)               |  __do_fault() {
   0)               |    filemap_fault() {
   0)               |      find_lock_page() {
   0)   0.698 us    |        find_get_page();
   0)               |        __might_sleep() {
   0)   1.412 us    |        }
   0)   3.950 us    |      }
   0)   5.098 us    |    }
   0)   0.631 us    |    _spin_lock();
   0)   0.571 us    |    page_add_file_rmap();
   0)   0.526 us    |    native_set_pte_at();
   0)   0.586 us    |    _spin_unlock();
   0)               |    unlock_page() {
   0)   0.533 us    |      page_waitqueue();
   0)   0.638 us    |      __wake_up_bit();
   0)   2.793 us    |    }
   0) + 14.012 us   |  }

여러 시작 함수를 동시에 확장하려면 첫 이름을 `>`로 쓰고 다음 이름을 `>>`로 추가한다.

 echo sys_open > set_graph_function
 echo sys_close >> set_graph_function

다시 모든 함수를 추적하려면 `set_graph_function`을 비운다.

 echo > set_graph_function

proc sysctl의 `ftrace_enabled`는 function tracer 전체를 제어하는 큰 전역 스위치다. 커널에 함수 추적 기능이 켜져 있으면 기본값은 활성화 상태다. 이를 끄면 ftrace의 function tracer뿐 아니라 perf, kprobes, 스택 추적, 프로파일링처럼 함수 추적을 사용하는 다른 기능도 모두 비활성화된다. `FTRACE_OPS_FL_PERMANENT`가 설정된 콜백이 등록되어 있으면 이 스위치를 끌 수 없다.

영향 범위가 넓으므로 이 설정은 주의해서 비활성화해야 한다. `sysctl` 명령으로 끄고 다시 켤 수 있다.

  sysctl kernel.ftrace_enabled=0
  sysctl kernel.ftrace_enabled=1

같은 설정은 `/proc/sys/kernel/ftrace_enabled`에 직접 0 또는 1을 써서 바꿀 수도 있다.

  echo 0 > /proc/sys/kernel/ftrace_enabled
  echo 1 > /proc/sys/kernel/ftrace_enabled
function graph 필터와 전역 스위치
인터페이스값 또는 연산효과
`set_graph_function`함수 이름을 `>` 또는 `>>`로 기록선택한 함수와 모든 자식을 그래프로 추적
`set_graph_function`빈 값 기록그래프 시작 필터를 지워 모든 함수 추적
`kernel.ftrace_enabled`0시스템 전체 함수 추적 비활성화
`kernel.ftrace_enabled`1시스템 전체 함수 추적 활성화

그래프의 시작 범위와 시스템 전체 함수 추적 활성화를 서로 다른 인터페이스로 제어한다.

function graph 대상 선택
함수 이름을 `set_graph_function`에 기록해당 함수 진입 감지
해당 함수 진입 감지함수 본문과 모든 자식 호출 기록
함수 본문과 모든 자식 호출 기록함수 반환까지 계층과 시간 출력

시작 함수 필터는 호출 계층의 뿌리를 제한하고 자식 호출은 함께 확장한다.

Dynamic ftrace with the function graph tracer
---------------------------------------------

Although what has been explained above concerns both the
function tracer and the function-graph-tracer, there are some
special features only available in the function-graph tracer.

If you want to trace only one function and all of its children,
you just have to echo its name into set_graph_function::

 echo __do_fault > set_graph_function

will produce the following "expanded" trace of the __do_fault()
function::

   0)               |  __do_fault() {
   0)               |    filemap_fault() {
   0)               |      find_lock_page() {
   0)   0.804 us    |        find_get_page();
   0)               |        __might_sleep() {
   0)   1.329 us    |        }
   0)   3.904 us    |      }
   0)   4.979 us    |    }
   0)   0.653 us    |    _spin_lock();
   0)   0.578 us    |    page_add_file_rmap();
   0)   0.525 us    |    native_set_pte_at();
   0)   0.585 us    |    _spin_unlock();
   0)               |    unlock_page() {
   0)   0.541 us    |      page_waitqueue();
   0)   0.639 us    |      __wake_up_bit();
   0)   2.786 us    |    }
   0) + 14.237 us   |  }
   0)               |  __do_fault() {
   0)               |    filemap_fault() {
   0)               |      find_lock_page() {
   0)   0.698 us    |        find_get_page();
   0)               |        __might_sleep() {
   0)   1.412 us    |        }
   0)   3.950 us    |      }
   0)   5.098 us    |    }
   0)   0.631 us    |    _spin_lock();
   0)   0.571 us    |    page_add_file_rmap();
   0)   0.526 us    |    native_set_pte_at();
   0)   0.586 us    |    _spin_unlock();
   0)               |    unlock_page() {
   0)   0.533 us    |      page_waitqueue();
   0)   0.638 us    |      __wake_up_bit();
   0)   2.793 us    |    }
   0) + 14.012 us   |  }

You can also expand several functions at once::

 echo sys_open > set_graph_function
 echo sys_close >> set_graph_function

Now if you want to go back to trace all functions you can clear
this special filter via::

 echo > set_graph_function


ftrace_enabled
--------------

Note, the proc sysctl ftrace_enable is a big on/off switch for the
function tracer. By default it is enabled (when function tracing is
enabled in the kernel). If it is disabled, all function tracing is
disabled. This includes not only the function tracers for ftrace, but
also for any other uses (perf, kprobes, stack tracing, profiling, etc). It
cannot be disabled if there is a callback with FTRACE_OPS_FL_PERMANENT set
registered.

Please disable this with care.

This can be disable (and enabled) with::

  sysctl kernel.ftrace_enabled=0
  sysctl kernel.ftrace_enabled=1

 or

  echo 0 > /proc/sys/kernel/ftrace_enabled
  echo 1 > /proc/sys/kernel/ftrace_enabled

함수 필터 명령

3300-3421

`set_ftrace_filter` 인터페이스는 함수 이름을 고르는 것 외에도 함수에 도달했을 때 실행할 여러 trace command를 지원한다. 기본 형식은 함수, 명령, 선택적 매개변수를 콜론으로 연결한다.

  <function>:<command>:<parameter>

`mod` 명령은 모듈별 함수 필터링을 활성화하며 매개변수로 모듈 이름을 받는다. 예를 들어 ext3 모듈에서 이름이 `write`로 시작하는 함수만 고르려면 다음과 같이 설정한다.

   echo 'write*:mod:ext3' > set_ftrace_filter

모듈 필터도 일반 함수 이름 필터와 같은 방식으로 누적한다. 다른 모듈의 함수를 더하려면 `>>`를 사용하고, 특정 모듈 함수를 제거하려면 패턴 앞에 `!`를 붙인다.

   echo '!writeback*:mod:ext3' >> set_ftrace_filter

`mod`는 모듈 이름에도 glob을 지원한다. 모든 함수를 제외하되 ext3 모듈만 예외로 남기려면 모듈 패턴에도 `!`를 사용한다.

   echo '!*:mod:!ext3' >> set_ftrace_filter

모든 모듈 함수의 추적을 끄면서 커널 자체 함수는 계속 추적하려면 모든 모듈에 일치하는 `*`를 제외한다.

   echo '!*:mod:*' >> set_ftrace_filter

반대로 모듈은 제외하고 커널에서 이름에 `write`가 들어간 함수만 필터에 넣을 수도 있다.

   echo '*write*:mod:!*' >> set_ftrace_filter

모듈 glob과 함수 glob을 함께 사용하면 이름에 `snd`가 들어간 모듈의 `*write*` 함수만 선택할 수 있다.

   echo '*write*:mod:*snd*' >> set_ftrace_filter

`traceon`과 `traceoff`는 지정한 함수에 도달할 때 추적을 켜거나 끈다. 매개변수는 이 동작을 실행할 횟수이며 생략하면 제한 없이 매번 실행한다. 다음 예는 `__schedule_bug`가 처음 다섯 번 호출될 때 추적을 끈다.

   echo '__schedule_bug:traceoff:5' > set_ftrace_filter

횟수를 생략하면 `__schedule_bug`에 도달할 때마다 항상 추적을 끈다.

   echo '__schedule_bug:traceoff' > set_ftrace_filter

이 명령들은 `set_ftrace_filter`에 덧붙였는지와 관계없이 누적된다. 카운터가 있는 명령을 제거할 때는 앞에 `!`를 붙이고 카운터 자리에 0을 쓴다.

   echo '!__schedule_bug:traceoff:0' > set_ftrace_filter

카운터 없이 등록한 명령은 같은 명령 앞에 `!`를 붙인 형태로 제거한다.

   echo '!__schedule_bug:traceoff' > set_ftrace_filter

`snapshot` 명령은 지정한 함수에 도달했을 때 snapshot을 만든다.

   echo 'native_flush_tlb_others:snapshot' > set_ftrace_filter

매개변수 1을 주면 snapshot을 한 번만 만든다.

   echo 'native_flush_tlb_others:snapshot:1' > set_ftrace_filter

카운터 없는 snapshot 명령과 카운터가 있는 snapshot 명령은 각각 다음처럼 제거한다.

   echo '!native_flush_tlb_others:snapshot' > set_ftrace_filter
   echo '!native_flush_tlb_others:snapshot:0' > set_ftrace_filter

`enable_event`와 `disable_event`는 함수에 도달했을 때 trace event를 켜거나 끈다. 함수 추적 콜백은 시간에 매우 민감하므로 이 명령이 등록되면 tracepoint 자체는 활성화하되 기록만 하지 않는 soft disabled 상태에 둔다. 해당 이벤트를 트리거하는 명령이 하나라도 남아 있는 동안 이 상태가 유지된다.

   echo 'try_to_wake_up:enable_event:sched:sched_switch:2' > \
   	 set_ftrace_filter

이벤트 명령은 함수 이름, 명령, 이벤트 시스템, 이벤트 이름, 선택적 횟수 순으로 쓴다.

    <function>:enable_event:<system>:<event>[:count]
    <function>:disable_event:<system>:<event>[:count]

등록한 이벤트 명령을 제거할 때도 함수와 명령 앞에 `!`를 붙이고 카운터가 있던 명령에는 0을 지정한다.

   echo '!try_to_wake_up:enable_event:sched:sched_switch:0' > \
   	 set_ftrace_filter
   echo '!schedule:disable_event:sched:sched_switch' > \
   	 set_ftrace_filter

`dump`는 지정한 함수에 도달하면 ftrace ring buffer 전체를 콘솔에 덤프한다. triple fault 직전처럼 정상적인 덤프를 얻기 어려운 지점 앞에서 호출되는 함수를 트리거로 삼을 때 유용하다.

`cpudump`는 현재 CPU의 ftrace ring buffer만 콘솔에 덤프한다. `dump`와 달리 트리거 함수를 실행한 CPU의 버퍼만 출력한다.

`stacktrace`는 지정한 함수에 도달할 때 스택 추적을 기록한다.

set_ftrace_filter 명령
명령매개변수함수 도달 시 동작
`mod`모듈 이름 또는 glob모듈 단위로 함수 필터 선택
`traceon` / `traceoff`선택적 실행 횟수추적 기록 켜기 또는 끄기
`snapshot`선택적 실행 횟수현재 trace buffer snapshot 생성
`enable_event` / `disable_event`시스템, 이벤트, 선택적 횟수trace event 기록 상태 변경
`dump`없음전체 CPU ring buffer를 콘솔에 출력
`cpudump`없음현재 CPU ring buffer만 콘솔에 출력
`stacktrace`없음현재 스택 추적 기록

함수 호출을 조건으로 추적 상태, snapshot, 이벤트 또는 진단 출력을 제어한다.

함수 기반 trace command 실행
`set_ftrace_filter`에 함수와 명령 등록지정 함수 호출
지정 함수 호출선택적 카운터 감소
선택적 카운터 감소추적·snapshot·event·dump 동작 실행
추적·snapshot·event·dump 동작 실행카운터가 남으면 다음 호출 대기

함수 필터에 명령을 결합하여 특정 커널 경로에 도달한 순간의 동작을 자동화한다.

Filter commands
---------------

A few commands are supported by the set_ftrace_filter interface.
Trace commands have the following format::

  <function>:<command>:<parameter>

The following commands are supported:

- mod:
  This command enables function filtering per module. The
  parameter defines the module. For example, if only the write*
  functions in the ext3 module are desired, run:

   echo 'write*:mod:ext3' > set_ftrace_filter

  This command interacts with the filter in the same way as
  filtering based on function names. Thus, adding more functions
  in a different module is accomplished by appending (>>) to the
  filter file. Remove specific module functions by prepending
  '!'::

   echo '!writeback*:mod:ext3' >> set_ftrace_filter

  Mod command supports module globbing. Disable tracing for all
  functions except a specific module::

   echo '!*:mod:!ext3' >> set_ftrace_filter

  Disable tracing for all modules, but still trace kernel::

   echo '!*:mod:*' >> set_ftrace_filter

  Enable filter only for kernel::

   echo '*write*:mod:!*' >> set_ftrace_filter

  Enable filter for module globbing::

   echo '*write*:mod:*snd*' >> set_ftrace_filter

- traceon/traceoff:
  These commands turn tracing on and off when the specified
  functions are hit. The parameter determines how many times the
  tracing system is turned on and off. If unspecified, there is
  no limit. For example, to disable tracing when a schedule bug
  is hit the first 5 times, run::

   echo '__schedule_bug:traceoff:5' > set_ftrace_filter

  To always disable tracing when __schedule_bug is hit::

   echo '__schedule_bug:traceoff' > set_ftrace_filter

  These commands are cumulative whether or not they are appended
  to set_ftrace_filter. To remove a command, prepend it by '!'
  and drop the parameter::

   echo '!__schedule_bug:traceoff:0' > set_ftrace_filter

  The above removes the traceoff command for __schedule_bug
  that have a counter. To remove commands without counters::

   echo '!__schedule_bug:traceoff' > set_ftrace_filter

- snapshot:
  Will cause a snapshot to be triggered when the function is hit.
  ::

   echo 'native_flush_tlb_others:snapshot' > set_ftrace_filter

  To only snapshot once:
  ::

   echo 'native_flush_tlb_others:snapshot:1' > set_ftrace_filter

  To remove the above commands::

   echo '!native_flush_tlb_others:snapshot' > set_ftrace_filter
   echo '!native_flush_tlb_others:snapshot:0' > set_ftrace_filter

- enable_event/disable_event:
  These commands can enable or disable a trace event. Note, because
  function tracing callbacks are very sensitive, when these commands
  are registered, the trace point is activated, but disabled in
  a "soft" mode. That is, the tracepoint will be called, but
  just will not be traced. The event tracepoint stays in this mode
  as long as there's a command that triggers it.
  ::

   echo 'try_to_wake_up:enable_event:sched:sched_switch:2' > \
   	 set_ftrace_filter

  The format is::

    <function>:enable_event:<system>:<event>[:count]
    <function>:disable_event:<system>:<event>[:count]

  To remove the events commands::

   echo '!try_to_wake_up:enable_event:sched:sched_switch:0' > \
   	 set_ftrace_filter
   echo '!schedule:disable_event:sched:sched_switch' > \
   	 set_ftrace_filter

- dump:
  When the function is hit, it will dump the contents of the ftrace
  ring buffer to the console. This is useful if you need to debug
  something, and want to dump the trace when a certain function
  is hit. Perhaps it's a function that is called before a triple
  fault happens and does not allow you to get a regular dump.

- cpudump:
  When the function is hit, it will dump the contents of the ftrace
  ring buffer for the current CPU to the console. Unlike the "dump"
  command, it only prints out the contents of the ring buffer for the
  CPU that executed the function that triggered the dump.

- stacktrace:
  When the function is hit, a stack trace is recorded.

trace_pipe 실시간 스트림

3422-3468

`trace_pipe`는 `trace` 파일과 같은 형식의 내용을 출력하지만 추적에 미치는 효과는 다르다. `trace_pipe`에서 읽은 항목은 소비되므로 다음 읽기에는 다시 나타나지 않으며, 새 항목이 생기는 대로 전달하는 실시간 스트림으로 동작한다.

아래 예에서는 `trace_pipe`를 백그라운드에서 `/tmp/trace.out`으로 읽은 뒤 짧게 function tracing을 수행한다. 소비가 끝난 기본 `trace` 버퍼에는 항목이 남아 있지 않지만 출력 파일에는 읽어 간 함수 기록이 저장되어 있다.

  # echo function > current_tracer
  # cat trace_pipe > /tmp/trace.out &
  [1] 4153
  # echo 1 > tracing_on
  # usleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: function
  #
  # entries-in-buffer/entries-written: 0/0   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |

  #
  # cat /tmp/trace.out
             bash-1994  [000] ....  5281.568961: mutex_unlock <-rb_simple_write
             bash-1994  [000] ....  5281.568963: __mutex_unlock_slowpath <-mutex_unlock
             bash-1994  [000] ....  5281.568963: __fsnotify_parent <-fsnotify_modify
             bash-1994  [000] ....  5281.568964: fsnotify <-fsnotify_modify
             bash-1994  [000] ....  5281.568964: __srcu_read_lock <-fsnotify
             bash-1994  [000] ....  5281.568964: add_preempt_count <-__srcu_read_lock
             bash-1994  [000] ...1  5281.568965: sub_preempt_count <-__srcu_read_lock
             bash-1994  [000] ....  5281.568965: __srcu_read_unlock <-fsnotify
             bash-1994  [000] ....  5281.568967: sys_dup2 <-system_call_fastpath

`trace_pipe` 읽기는 새 입력이 추가될 때까지 block한다. 반면 어떤 프로세스가 `trace` 파일을 읽기 위해 열면 추적이 비활성화되어 새 항목이 추가되지 않는다. `trace_pipe`에는 이 제한이 없으므로 추적을 계속하면서 결과를 소비할 수 있다.

trace와 trace_pipe 비교
파일읽기 결과새 항목 대기추적에 미치는 영향
`trace`현재 버퍼의 비소비 snapshot즉시 반환읽기용 open이 새 기록을 막음
`trace_pipe`읽을 때마다 항목 소비새 입력까지 block추적을 계속하며 실시간 전달

두 파일은 같은 출력 형식을 사용하지만 읽기 수명과 추적 동작이 다르다.

trace_pipe 소비 흐름
ftrace ring buffer에 새 항목 기록`trace_pipe` 대기 중인 reader에 전달
`trace_pipe` 대기 중인 reader에 전달읽은 항목을 buffer에서 소비
읽은 항목을 buffer에서 소비다음 새 항목까지 다시 block

ring buffer 항목을 한 번씩 읽어 장시간 실시간 수집기로 전달한다.

trace_pipe
----------

The trace_pipe outputs the same content as the trace file, but
the effect on the tracing is different. Every read from
trace_pipe is consumed. This means that subsequent reads will be
different. The trace is live.
::

  # echo function > current_tracer
  # cat trace_pipe > /tmp/trace.out &
  [1] 4153
  # echo 1 > tracing_on
  # usleep 1
  # echo 0 > tracing_on
  # cat trace
  # tracer: function
  #
  # entries-in-buffer/entries-written: 0/0   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |

  #
  # cat /tmp/trace.out
             bash-1994  [000] ....  5281.568961: mutex_unlock <-rb_simple_write
             bash-1994  [000] ....  5281.568963: __mutex_unlock_slowpath <-mutex_unlock
             bash-1994  [000] ....  5281.568963: __fsnotify_parent <-fsnotify_modify
             bash-1994  [000] ....  5281.568964: fsnotify <-fsnotify_modify
             bash-1994  [000] ....  5281.568964: __srcu_read_lock <-fsnotify
             bash-1994  [000] ....  5281.568964: add_preempt_count <-__srcu_read_lock
             bash-1994  [000] ...1  5281.568965: sub_preempt_count <-__srcu_read_lock
             bash-1994  [000] ....  5281.568965: __srcu_read_unlock <-fsnotify
             bash-1994  [000] ....  5281.568967: sys_dup2 <-system_call_fastpath


Note, reading the trace_pipe file will block until more input is
added. This is contrary to the trace file. If any process opened
the trace file for reading, it will actually disable tracing and
prevent new entries from being added. The trace_pipe file does
not have this limitation.

trace 항목과 버퍼 크기

3469-3526

커널 문제를 진단할 때 데이터가 너무 많거나 너무 적으면 모두 곤란하다. `buffer_size_kb`는 내부 trace buffer의 CPU별 크기를 조절한다. 전체 크기는 가능한 CPU 수를 CPU별 값에 곱해 계산할 수 있다.

현재 CPU별 버퍼 크기는 다음처럼 확인한다.

  # cat buffer_size_kb
  1408 (units kilobytes)

전체 버퍼 크기는 `buffer_total_size_kb`에서 바로 읽을 수도 있다.

  # cat buffer_total_size_kb 
  5632

버퍼 크기를 바꾸려면 1,024바이트 단위의 숫자를 `buffer_size_kb`에 쓴다.

  # echo 10000 > buffer_size_kb
  # cat buffer_size_kb
  10000 (units kilobytes)

커널은 요청한 크기를 가능한 만큼 할당하려고 한다. 지나치게 큰 값을 요청하면 Out-Of-Memory가 발생할 수 있으며, 할당 실패 뒤에는 실제로 확보된 값이 표시된다.

  # echo 1000000000000 > buffer_size_kb
  -bash: echo: write error: Cannot allocate memory
  # cat buffer_size_kb
  85

`per_cpu/cpuN/buffer_size_kb`를 사용하면 CPU별 버퍼 크기를 서로 다르게 설정할 수 있다.

  # echo 10000 > per_cpu/cpu0/buffer_size_kb
  # echo 100 > per_cpu/cpu1/buffer_size_kb

CPU별 크기가 같지 않으면 최상위 `buffer_size_kb`는 단일 숫자 대신 `X`를 표시한다.

  # cat buffer_size_kb
  X

이때 전체 할당량을 확인하려면 `buffer_total_size_kb`가 유용하다.

  # cat buffer_total_size_kb 
  12916

최상위 `buffer_size_kb`에 다시 값을 쓰면 모든 CPU 버퍼가 같은 크기로 재설정된다.

trace buffer 크기 인터페이스
인터페이스읽기쓰기
`buffer_size_kb`CPU별 공통 크기, 서로 다르면 `X`모든 CPU를 같은 크기로 재설정
`buffer_total_size_kb`모든 CPU 버퍼의 합계전체 크기 확인용
`per_cpu/cpuN/buffer_size_kb`특정 CPU의 크기해당 CPU만 크기 변경

CPU별 크기와 전체 크기를 구분하여 메모리 사용량을 조절한다.

trace buffer 크기 조절
수집할 trace 양 산정CPU별 `buffer_size_kb` 설정
CPU별 `buffer_size_kb` 설정`buffer_total_size_kb`로 총량 확인
`buffer_total_size_kb`로 총량 확인할당 실패 또는 OOM 여부 확인
할당 실패 또는 OOM 여부 확인필요하면 더 작은 값으로 조정

필요한 관찰 구간과 사용 가능한 메모리를 함께 고려해 CPU별 버퍼를 배분한다.

trace entries
-------------

Having too much or not enough data can be troublesome in
diagnosing an issue in the kernel. The file buffer_size_kb is
used to modify the size of the internal trace buffers. The
number listed is the number of entries that can be recorded per
CPU. To know the full size, multiply the number of possible CPUs
with the number of entries.
::

  # cat buffer_size_kb
  1408 (units kilobytes)

Or simply read buffer_total_size_kb
::

  # cat buffer_total_size_kb 
  5632

To modify the buffer, simple echo in a number (in 1024 byte segments).
::

  # echo 10000 > buffer_size_kb
  # cat buffer_size_kb
  10000 (units kilobytes)

It will try to allocate as much as possible. If you allocate too
much, it can cause Out-Of-Memory to trigger.
::

  # echo 1000000000000 > buffer_size_kb
  -bash: echo: write error: Cannot allocate memory
  # cat buffer_size_kb
  85

The per_cpu buffers can be changed individually as well:
::

  # echo 10000 > per_cpu/cpu0/buffer_size_kb
  # echo 100 > per_cpu/cpu1/buffer_size_kb

When the per_cpu buffers are not the same, the buffer_size_kb
at the top level will just show an X
::

  # cat buffer_size_kb
  X

This is where the buffer_total_size_kb is useful:
::

  # cat buffer_total_size_kb 
  12916

Writing to the top level buffer_size_kb will reset all the buffers
to be the same again.

Snapshot

3527-3613

`CONFIG_TRACER_SNAPSHOT`은 latency tracer가 아닌 모든 tracer에 공통 snapshot 기능을 제공한다. `irqsoff`나 `wakeup`처럼 최대 지연을 기록하는 latency tracer는 내부적으로 이미 snapshot 메커니즘을 사용하므로 이 기능을 함께 사용할 수 없다.

Snapshot은 추적을 멈추지 않고 특정 시점의 현재 trace buffer를 보존한다. Ftrace가 현재 버퍼와 예비 버퍼를 맞바꾸면 추적은 이전 예비 버퍼였던 새 현재 버퍼에서 계속된다.

`tracing/snapshot` 파일은 snapshot 생성과 결과 읽기에 모두 사용한다. 1을 쓰면 예비 버퍼를 할당하고 현재 버퍼와 교환하며, 파일을 읽으면 앞서 설명한 `trace`와 같은 형식으로 보존된 내용을 얻는다. snapshot 읽기와 현재 추적은 동시에 수행할 수 있다.

예비 버퍼가 할당된 상태에서 0을 쓰면 버퍼를 해제하고, 1을 쓰면 다시 교환하며, 그 밖의 양수 값을 쓰면 snapshot 내용을 지운다. 아직 할당되지 않은 상태에서는 1만 할당과 교환을 수행하고 다른 값은 아무 동작도 하지 않는다.

snapshot 상태와 입력
현재 상태입력 0입력 1그 밖의 양수
미할당동작 없음할당 후 swap동작 없음
할당됨예비 버퍼 해제현재 버퍼와 swapsnapshot 내용 지우기

원문의 ASCII 표를 예비 버퍼 할당 상태별 동작으로 구조화했다.

다음 예는 scheduler 이벤트를 켜고 snapshot을 만든 뒤 snapshot과 계속 진행 중인 현재 `trace`를 각각 읽는다. 두 버퍼의 항목 수와 시점이 다른 것은 snapshot 이후에도 추적이 새 현재 버퍼에서 계속되기 때문이다.

  # echo 1 > events/sched/enable
  # echo 1 > snapshot
  # cat snapshot
  # tracer: nop
  #
  # entries-in-buffer/entries-written: 71/71   #P:8
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            <idle>-0     [005] d...  2440.603828: sched_switch: prev_comm=swapper/5 prev_pid=0 prev_prio=120   prev_state=R ==> next_comm=snapshot-test-2 next_pid=2242 next_prio=120
             sleep-2242  [005] d...  2440.603846: sched_switch: prev_comm=snapshot-test-2 prev_pid=2242 prev_prio=120   prev_state=R ==> next_comm=kworker/5:1 next_pid=60 next_prio=120
  [...]
          <idle>-0     [002] d...  2440.707230: sched_switch: prev_comm=swapper/2 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2229 next_prio=120  

  # cat trace  
  # tracer: nop
  #
  # entries-in-buffer/entries-written: 77/77   #P:8
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            <idle>-0     [007] d...  2440.707395: sched_switch: prev_comm=swapper/7 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2243 next_prio=120
   snapshot-test-2-2229  [002] d...  2440.707438: sched_switch: prev_comm=snapshot-test-2 prev_pid=2229 prev_prio=120 prev_state=S ==> next_comm=swapper/2 next_pid=0 next_prio=120
  [...]

현재 tracer가 latency tracer라면 예비 snapshot 버퍼를 이미 내부 용도로 사용하므로 snapshot 생성과 읽기가 모두 `Device or resource busy`로 실패한다.

  # echo wakeup > current_tracer
  # echo 1 > snapshot
  bash: echo: write error: Device or resource busy
  # cat snapshot
  cat: snapshot: Device or resource busy
중단 없는 snapshot
현재 trace buffer에 이벤트 기록snapshot 트리거
snapshot 트리거현재 버퍼와 예비 버퍼 swap
현재 버퍼와 예비 버퍼 swap보존 버퍼는 `snapshot`에서 읽기
현재 버퍼와 예비 버퍼 swap새 현재 버퍼에서 추적 계속

두 버퍼의 역할을 즉시 맞바꾸어 과거 시점을 고정하면서 새 이벤트 수집을 계속한다.

Snapshot
--------
CONFIG_TRACER_SNAPSHOT makes a generic snapshot feature
available to all non latency tracers. (Latency tracers which
record max latency, such as "irqsoff" or "wakeup", can't use
this feature, since those are already using the snapshot
mechanism internally.)

Snapshot preserves a current trace buffer at a particular point
in time without stopping tracing. Ftrace swaps the current
buffer with a spare buffer, and tracing continues in the new
current (=previous spare) buffer.

The following tracefs files in "tracing" are related to this
feature:

  snapshot:

	This is used to take a snapshot and to read the output
	of the snapshot. Echo 1 into this file to allocate a
	spare buffer and to take a snapshot (swap), then read
	the snapshot from this file in the same format as
	"trace" (described above in the section "The File
	System"). Both reads snapshot and tracing are executable
	in parallel. When the spare buffer is allocated, echoing
	0 frees it, and echoing else (positive) values clear the
	snapshot contents.
	More details are shown in the table below.

	+--------------+------------+------------+------------+
	|status\\input |     0      |     1      |    else    |
	+==============+============+============+============+
	|not allocated |(do nothing)| alloc+swap |(do nothing)|
	+--------------+------------+------------+------------+
	|allocated     |    free    |    swap    |   clear    |
	+--------------+------------+------------+------------+

Here is an example of using the snapshot feature.
::

  # echo 1 > events/sched/enable
  # echo 1 > snapshot
  # cat snapshot
  # tracer: nop
  #
  # entries-in-buffer/entries-written: 71/71   #P:8
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            <idle>-0     [005] d...  2440.603828: sched_switch: prev_comm=swapper/5 prev_pid=0 prev_prio=120   prev_state=R ==> next_comm=snapshot-test-2 next_pid=2242 next_prio=120
             sleep-2242  [005] d...  2440.603846: sched_switch: prev_comm=snapshot-test-2 prev_pid=2242 prev_prio=120   prev_state=R ==> next_comm=kworker/5:1 next_pid=60 next_prio=120
  [...]
          <idle>-0     [002] d...  2440.707230: sched_switch: prev_comm=swapper/2 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2229 next_prio=120  

  # cat trace  
  # tracer: nop
  #
  # entries-in-buffer/entries-written: 77/77   #P:8
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
            <idle>-0     [007] d...  2440.707395: sched_switch: prev_comm=swapper/7 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=snapshot-test-2 next_pid=2243 next_prio=120
   snapshot-test-2-2229  [002] d...  2440.707438: sched_switch: prev_comm=snapshot-test-2 prev_pid=2229 prev_prio=120 prev_state=S ==> next_comm=swapper/2 next_pid=0 next_prio=120
  [...]


If you try to use this snapshot feature when current tracer is
one of the latency tracers, you will get the following results.
::

  # echo wakeup > current_tracer
  # echo 1 > snapshot
  bash: echo: write error: Device or resource busy
  # cat snapshot
  cat: snapshot: Device or resource busy

독립 추적 Instances

3614-3740

tracefs의 `tracing` 디렉터리에는 `instances` 디렉터리가 있다. 그 안에 `mkdir`로 새 디렉터리를 만들고 `rmdir`로 제거할 수 있으며, 새 instance 디렉터리는 만들어지는 즉시 자체 추적 파일과 하위 디렉터리를 갖는다.

  # mkdir instances/foo
  # ls instances/foo
  buffer_size_kb  buffer_total_size_kb  events  free_buffer  per_cpu
  set_event  snapshot  trace  trace_clock  trace_marker  trace_options
  trace_pipe  tracing_on

새 디렉터리는 최상위 `tracing` 디렉터리와 매우 비슷하지만 buffer와 event 구성이 주 instance 및 다른 instance와 독립적이다. 같은 이름의 파일도 기본적으로 같은 방식으로 동작하되 각 instance의 별도 새 buffer에만 영향을 준다.

예외는 `trace_options`다. 현재 이 옵션은 모든 instance와 최상위 buffer에 동일하게 적용된다. 향후 릴리스에서는 각 옵션이 자신이 놓인 instance에만 적용되도록 바뀔 수 있다.

새 instance에는 function tracer 전용 파일, `current_tracer`, `available_tracers`가 없다. 이 문서가 설명하는 현재 구현에서 instance buffer는 event만 개별적으로 활성화할 수 있기 때문이다.

다음 예는 `foo`, `bar`, `zoot` instance를 만들고 최상위 buffer와 각 instance에 서로 다른 크기와 이벤트를 설정한다. 최상위 buffer는 function tracing을, `foo`는 scheduler wakeup과 switch를, `bar`는 IRQ 이벤트를, `zoot`는 syscall 이벤트를 각각 기록한다.

  # mkdir instances/foo
  # mkdir instances/bar
  # mkdir instances/zoot
  # echo 100000 > buffer_size_kb
  # echo 1000 > instances/foo/buffer_size_kb
  # echo 5000 > instances/bar/per_cpu/cpu1/buffer_size_kb
  # echo function > current_trace
  # echo 1 > instances/foo/events/sched/sched_wakeup/enable
  # echo 1 > instances/foo/events/sched/sched_wakeup_new/enable
  # echo 1 > instances/foo/events/sched/sched_switch/enable
  # echo 1 > instances/bar/events/irq/enable
  # echo 1 > instances/zoot/events/syscalls/enable
  # cat trace_pipe
  CPU:2 [LOST 11745 EVENTS]
              bash-2044  [002] .... 10594.481032: _raw_spin_lock_irqsave <-get_page_from_freelist
              bash-2044  [002] d... 10594.481032: add_preempt_count <-_raw_spin_lock_irqsave
              bash-2044  [002] d..1 10594.481032: __rmqueue <-get_page_from_freelist
              bash-2044  [002] d..1 10594.481033: _raw_spin_unlock <-get_page_from_freelist
              bash-2044  [002] d..1 10594.481033: sub_preempt_count <-_raw_spin_unlock
              bash-2044  [002] d... 10594.481033: get_pageblock_flags_group <-get_pageblock_migratetype
              bash-2044  [002] d... 10594.481034: __mod_zone_page_state <-get_page_from_freelist
              bash-2044  [002] d... 10594.481034: zone_statistics <-get_page_from_freelist
              bash-2044  [002] d... 10594.481034: __inc_zone_state <-zone_statistics
              bash-2044  [002] d... 10594.481034: __inc_zone_state <-zone_statistics
              bash-2044  [002] .... 10594.481035: arch_dup_task_struct <-copy_process
  [...]

  # cat instances/foo/trace_pipe
              bash-1998  [000] d..4   136.676759: sched_wakeup: comm=kworker/0:1 pid=59 prio=120 success=1 target_cpu=000
              bash-1998  [000] dN.4   136.676760: sched_wakeup: comm=bash pid=1998 prio=120 success=1 target_cpu=000
            <idle>-0     [003] d.h3   136.676906: sched_wakeup: comm=rcu_preempt pid=9 prio=120 success=1 target_cpu=003
            <idle>-0     [003] d..3   136.676909: sched_switch: prev_comm=swapper/3 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=rcu_preempt next_pid=9 next_prio=120
       rcu_preempt-9     [003] d..3   136.676916: sched_switch: prev_comm=rcu_preempt prev_pid=9 prev_prio=120 prev_state=S ==> next_comm=swapper/3 next_pid=0 next_prio=120
              bash-1998  [000] d..4   136.677014: sched_wakeup: comm=kworker/0:1 pid=59 prio=120 success=1 target_cpu=000
              bash-1998  [000] dN.4   136.677016: sched_wakeup: comm=bash pid=1998 prio=120 success=1 target_cpu=000
              bash-1998  [000] d..3   136.677018: sched_switch: prev_comm=bash prev_pid=1998 prev_prio=120 prev_state=R+ ==> next_comm=kworker/0:1 next_pid=59 next_prio=120
       kworker/0:1-59    [000] d..4   136.677022: sched_wakeup: comm=sshd pid=1995 prio=120 success=1 target_cpu=001
       kworker/0:1-59    [000] d..3   136.677025: sched_switch: prev_comm=kworker/0:1 prev_pid=59 prev_prio=120 prev_state=S ==> next_comm=bash next_pid=1998 next_prio=120
  [...]

  # cat instances/bar/trace_pipe
       migration/1-14    [001] d.h3   138.732674: softirq_raise: vec=3 [action=NET_RX]
            <idle>-0     [001] dNh3   138.732725: softirq_raise: vec=3 [action=NET_RX]
              bash-1998  [000] d.h1   138.733101: softirq_raise: vec=1 [action=TIMER]
              bash-1998  [000] d.h1   138.733102: softirq_raise: vec=9 [action=RCU]
              bash-1998  [000] ..s2   138.733105: softirq_entry: vec=1 [action=TIMER]
              bash-1998  [000] ..s2   138.733106: softirq_exit: vec=1 [action=TIMER]
              bash-1998  [000] ..s2   138.733106: softirq_entry: vec=9 [action=RCU]
              bash-1998  [000] ..s2   138.733109: softirq_exit: vec=9 [action=RCU]
              sshd-1995  [001] d.h1   138.733278: irq_handler_entry: irq=21 name=uhci_hcd:usb4
              sshd-1995  [001] d.h1   138.733280: irq_handler_exit: irq=21 ret=unhandled
              sshd-1995  [001] d.h1   138.733281: irq_handler_entry: irq=21 name=eth0
              sshd-1995  [001] d.h1   138.733283: irq_handler_exit: irq=21 ret=handled
  [...]

  # cat instances/zoot/trace
  # tracer: nop
  #
  # entries-in-buffer/entries-written: 18996/18996   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1998  [000] d...   140.733501: sys_write -> 0x2
              bash-1998  [000] d...   140.733504: sys_dup2(oldfd: a, newfd: 1)
              bash-1998  [000] d...   140.733506: sys_dup2 -> 0x1
              bash-1998  [000] d...   140.733508: sys_fcntl(fd: a, cmd: 1, arg: 0)
              bash-1998  [000] d...   140.733509: sys_fcntl -> 0x1
              bash-1998  [000] d...   140.733510: sys_close(fd: a)
              bash-1998  [000] d...   140.733510: sys_close -> 0x0
              bash-1998  [000] d...   140.733514: sys_rt_sigprocmask(how: 0, nset: 0, oset: 6e2768, sigsetsize: 8)
              bash-1998  [000] d...   140.733515: sys_rt_sigprocmask -> 0x0
              bash-1998  [000] d...   140.733516: sys_rt_sigaction(sig: 2, act: 7fff718846f0, oact: 7fff71884650, sigsetsize: 8)
              bash-1998  [000] d...   140.733516: sys_rt_sigaction -> 0x0

출력을 비교하면 최상위 `trace_pipe`에는 function tracing만 나타나고 `foo`에는 wakeup과 task switch가 표시된다. `bar`에는 softirq와 IRQ handler 기록이, `zoot`에는 syscall 진입과 반환 기록이 독립적으로 남는다.

instance를 제거하려면 해당 디렉터리를 `rmdir`로 삭제한다.

  # rmdir instances/foo
  # rmdir instances/bar
  # rmdir instances/zoot

어떤 프로세스가 instance 디렉터리 안의 trace 파일을 열어 둔 상태라면 `rmdir`는 `EBUSY`로 실패한다. reader를 종료하고 열린 파일 참조가 사라진 뒤 제거해야 한다.

추적 instance 분리 예
범위선택한 추적대표 출력
최상위function tracer함수 진입과 호출자
`instances/foo`sched wakeup·switchtask wakeup과 문맥 전환
`instances/bar`IRQ eventssoftirq와 IRQ handler
`instances/zoot`syscall eventssystem call 진입과 반환

각 instance는 별도 buffer와 event 선택을 사용하여 같은 시간대의 서로 다른 관찰 관점을 유지한다.

instance 생명주기
`mkdir instances/<name>`instance 전용 buffer 생성
instance 전용 buffer 생성buffer 크기와 event 설정
buffer 크기와 event 설정instance의 `trace` 또는 `trace_pipe` 읽기
instance의 `trace` 또는 `trace_pipe` 읽기reader 종료
reader 종료`rmdir instances/<name>`

독립 buffer를 만들고 이벤트를 수집한 뒤 열린 reader가 없을 때 제거한다.

Instances
---------
In the tracefs tracing directory, there is a directory called "instances".
This directory can have new directories created inside of it using
mkdir, and removing directories with rmdir. The directory created
with mkdir in this directory will already contain files and other
directories after it is created.
::

  # mkdir instances/foo
  # ls instances/foo
  buffer_size_kb  buffer_total_size_kb  events  free_buffer  per_cpu
  set_event  snapshot  trace  trace_clock  trace_marker  trace_options
  trace_pipe  tracing_on

As you can see, the new directory looks similar to the tracing directory
itself. In fact, it is very similar, except that the buffer and
events are agnostic from the main directory, or from any other
instances that are created.

The files in the new directory work just like the files with the
same name in the tracing directory except the buffer that is used
is a separate and new buffer. The files affect that buffer but do not
affect the main buffer with the exception of trace_options. Currently,
the trace_options affect all instances and the top level buffer
the same, but this may change in future releases. That is, options
may become specific to the instance they reside in.

Notice that none of the function tracer files are there, nor is
current_tracer and available_tracers. This is because the buffers
can currently only have events enabled for them.
::

  # mkdir instances/foo
  # mkdir instances/bar
  # mkdir instances/zoot
  # echo 100000 > buffer_size_kb
  # echo 1000 > instances/foo/buffer_size_kb
  # echo 5000 > instances/bar/per_cpu/cpu1/buffer_size_kb
  # echo function > current_trace
  # echo 1 > instances/foo/events/sched/sched_wakeup/enable
  # echo 1 > instances/foo/events/sched/sched_wakeup_new/enable
  # echo 1 > instances/foo/events/sched/sched_switch/enable
  # echo 1 > instances/bar/events/irq/enable
  # echo 1 > instances/zoot/events/syscalls/enable
  # cat trace_pipe
  CPU:2 [LOST 11745 EVENTS]
              bash-2044  [002] .... 10594.481032: _raw_spin_lock_irqsave <-get_page_from_freelist
              bash-2044  [002] d... 10594.481032: add_preempt_count <-_raw_spin_lock_irqsave
              bash-2044  [002] d..1 10594.481032: __rmqueue <-get_page_from_freelist
              bash-2044  [002] d..1 10594.481033: _raw_spin_unlock <-get_page_from_freelist
              bash-2044  [002] d..1 10594.481033: sub_preempt_count <-_raw_spin_unlock
              bash-2044  [002] d... 10594.481033: get_pageblock_flags_group <-get_pageblock_migratetype
              bash-2044  [002] d... 10594.481034: __mod_zone_page_state <-get_page_from_freelist
              bash-2044  [002] d... 10594.481034: zone_statistics <-get_page_from_freelist
              bash-2044  [002] d... 10594.481034: __inc_zone_state <-zone_statistics
              bash-2044  [002] d... 10594.481034: __inc_zone_state <-zone_statistics
              bash-2044  [002] .... 10594.481035: arch_dup_task_struct <-copy_process
  [...]

  # cat instances/foo/trace_pipe
              bash-1998  [000] d..4   136.676759: sched_wakeup: comm=kworker/0:1 pid=59 prio=120 success=1 target_cpu=000
              bash-1998  [000] dN.4   136.676760: sched_wakeup: comm=bash pid=1998 prio=120 success=1 target_cpu=000
            <idle>-0     [003] d.h3   136.676906: sched_wakeup: comm=rcu_preempt pid=9 prio=120 success=1 target_cpu=003
            <idle>-0     [003] d..3   136.676909: sched_switch: prev_comm=swapper/3 prev_pid=0 prev_prio=120 prev_state=R ==> next_comm=rcu_preempt next_pid=9 next_prio=120
       rcu_preempt-9     [003] d..3   136.676916: sched_switch: prev_comm=rcu_preempt prev_pid=9 prev_prio=120 prev_state=S ==> next_comm=swapper/3 next_pid=0 next_prio=120
              bash-1998  [000] d..4   136.677014: sched_wakeup: comm=kworker/0:1 pid=59 prio=120 success=1 target_cpu=000
              bash-1998  [000] dN.4   136.677016: sched_wakeup: comm=bash pid=1998 prio=120 success=1 target_cpu=000
              bash-1998  [000] d..3   136.677018: sched_switch: prev_comm=bash prev_pid=1998 prev_prio=120 prev_state=R+ ==> next_comm=kworker/0:1 next_pid=59 next_prio=120
       kworker/0:1-59    [000] d..4   136.677022: sched_wakeup: comm=sshd pid=1995 prio=120 success=1 target_cpu=001
       kworker/0:1-59    [000] d..3   136.677025: sched_switch: prev_comm=kworker/0:1 prev_pid=59 prev_prio=120 prev_state=S ==> next_comm=bash next_pid=1998 next_prio=120
  [...]

  # cat instances/bar/trace_pipe
       migration/1-14    [001] d.h3   138.732674: softirq_raise: vec=3 [action=NET_RX]
            <idle>-0     [001] dNh3   138.732725: softirq_raise: vec=3 [action=NET_RX]
              bash-1998  [000] d.h1   138.733101: softirq_raise: vec=1 [action=TIMER]
              bash-1998  [000] d.h1   138.733102: softirq_raise: vec=9 [action=RCU]
              bash-1998  [000] ..s2   138.733105: softirq_entry: vec=1 [action=TIMER]
              bash-1998  [000] ..s2   138.733106: softirq_exit: vec=1 [action=TIMER]
              bash-1998  [000] ..s2   138.733106: softirq_entry: vec=9 [action=RCU]
              bash-1998  [000] ..s2   138.733109: softirq_exit: vec=9 [action=RCU]
              sshd-1995  [001] d.h1   138.733278: irq_handler_entry: irq=21 name=uhci_hcd:usb4
              sshd-1995  [001] d.h1   138.733280: irq_handler_exit: irq=21 ret=unhandled
              sshd-1995  [001] d.h1   138.733281: irq_handler_entry: irq=21 name=eth0
              sshd-1995  [001] d.h1   138.733283: irq_handler_exit: irq=21 ret=handled
  [...]

  # cat instances/zoot/trace
  # tracer: nop
  #
  # entries-in-buffer/entries-written: 18996/18996   #P:4
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
              bash-1998  [000] d...   140.733501: sys_write -> 0x2
              bash-1998  [000] d...   140.733504: sys_dup2(oldfd: a, newfd: 1)
              bash-1998  [000] d...   140.733506: sys_dup2 -> 0x1
              bash-1998  [000] d...   140.733508: sys_fcntl(fd: a, cmd: 1, arg: 0)
              bash-1998  [000] d...   140.733509: sys_fcntl -> 0x1
              bash-1998  [000] d...   140.733510: sys_close(fd: a)
              bash-1998  [000] d...   140.733510: sys_close -> 0x0
              bash-1998  [000] d...   140.733514: sys_rt_sigprocmask(how: 0, nset: 0, oset: 6e2768, sigsetsize: 8)
              bash-1998  [000] d...   140.733515: sys_rt_sigprocmask -> 0x0
              bash-1998  [000] d...   140.733516: sys_rt_sigaction(sig: 2, act: 7fff718846f0, oact: 7fff71884650, sigsetsize: 8)
              bash-1998  [000] d...   140.733516: sys_rt_sigaction -> 0x0

You can see that the trace of the top most trace buffer shows only
the function tracing. The foo instance displays wakeups and task
switches.

To remove the instances, simply delete their directories:
::

  # rmdir instances/foo
  # rmdir instances/bar
  # rmdir instances/zoot

Note, if a process has a trace file open in one of the instance
directories, the rmdir will fail with EBUSY.

커널 스택 사용량 추적

3741-3798

커널 스택은 크기가 고정되어 있으므로 함수에서 낭비하지 않는 것이 중요하다. 개발자가 스택에 너무 많은 데이터를 할당하면 stack overflow와 메모리 손상이 발생하고 대개 system panic으로 이어진다.

일부 도구는 interrupt를 이용해 주기적으로 스택 사용량을 점검하지만, 모든 함수 호출 지점에서 검사할 수 있다면 더 유용하다. Ftrace는 이미 function tracer를 제공하므로 stack tracer를 통해 각 함수 호출마다 스택 크기를 확인할 수 있다.

`CONFIG_STACK_TRACER`가 ftrace stack tracing 기능을 활성화한다. 실행 중에 사용하려면 `/proc/sys/kernel/stack_tracer_enabled`에 1을 쓴다.

 # echo 1 > /proc/sys/kernel/stack_tracer_enabled

부팅 과정의 커널 스택 크기를 추적하려면 kernel command line에 `stacktrace` 매개변수를 추가하여 시작부터 활성화할 수도 있다.

몇 분 동안 실행한 뒤 `stack_max_size`를 읽으면 관찰한 최대 스택 사용량을, `stack_trace`를 읽으면 그 최대값을 만든 호출 경로와 각 frame의 크기를 확인할 수 있다.

  # cat stack_max_size
  2928

  # cat stack_trace
          Depth    Size   Location    (18 entries)
          -----    ----   --------
    0)     2928     224   update_sd_lb_stats+0xbc/0x4ac
    1)     2704     160   find_busiest_group+0x31/0x1f1
    2)     2544     256   load_balance+0xd9/0x662
    3)     2288      80   idle_balance+0xbb/0x130
    4)     2208     128   __schedule+0x26e/0x5b9
    5)     2080      16   schedule+0x64/0x66
    6)     2064     128   schedule_timeout+0x34/0xe0
    7)     1936     112   wait_for_common+0x97/0xf1
    8)     1824      16   wait_for_completion+0x1d/0x1f
    9)     1808     128   flush_work+0xfe/0x119
   10)     1680      16   tty_flush_to_ldisc+0x1e/0x20
   11)     1664      48   input_available_p+0x1d/0x5c
   12)     1616      48   n_tty_poll+0x6d/0x134
   13)     1568      64   tty_poll+0x64/0x7f
   14)     1504     880   do_select+0x31e/0x511
   15)      624     400   core_sys_select+0x177/0x216
   16)      224      96   sys_select+0x91/0xb9
   17)      128     128   system_call_fastpath+0x16/0x1b

`Depth`는 해당 frame에서 남아 있는 누적 깊이, `Size`는 그 함수 frame이 차지한 크기, `Location`은 함수와 오프셋을 나타낸다. 예에서는 최대 2,928바이트의 경로가 scheduler와 select system call까지 이어진다.

GCC가 `-mfentry`를 사용하면 함수가 stack frame을 만들기 전에 추적된다. 따라서 `-mfentry` 경로에서는 leaf 함수가 stack tracer 검사 대상에 포함되지 않는다. 현재 문서 기준으로 `-mfentry`는 x86의 GCC 4.6.0 이상에서 사용된다.

stack_trace 출력 열
의미
Depth현재 위치의 누적 스택 깊이
Size해당 함수 frame이 차지한 바이트 수
Location함수 이름과 명령 오프셋

최대 스택 사용 시점의 호출 경로를 frame별 사용량과 함께 해석한다.

함수별 스택 사용량 측정
`CONFIG_STACK_TRACER`로 기능 포함stack tracer 활성화
stack tracer 활성화각 함수 호출에서 현재 깊이 검사
각 함수 호출에서 현재 깊이 검사기존 최대값과 비교
기존 최대값과 비교더 크면 `stack_max_size`와 `stack_trace` 갱신

function tracer 진입 지점에서 깊이를 검사하여 지금까지의 최대 사용 경로를 보존한다.

Stack trace
-----------
Since the kernel has a fixed sized stack, it is important not to
waste it in functions. A kernel developer must be conscious of
what they allocate on the stack. If they add too much, the system
can be in danger of a stack overflow, and corruption will occur,
usually leading to a system panic.

There are some tools that check this, usually with interrupts
periodically checking usage. But if you can perform a check
at every function call that will become very useful. As ftrace provides
a function tracer, it makes it convenient to check the stack size
at every function call. This is enabled via the stack tracer.

CONFIG_STACK_TRACER enables the ftrace stack tracing functionality.
To enable it, write a '1' into /proc/sys/kernel/stack_tracer_enabled.
::

 # echo 1 > /proc/sys/kernel/stack_tracer_enabled

You can also enable it from the kernel command line to trace
the stack size of the kernel during boot up, by adding "stacktrace"
to the kernel command line parameter.

After running it for a few minutes, the output looks like:
::

  # cat stack_max_size
  2928

  # cat stack_trace
          Depth    Size   Location    (18 entries)
          -----    ----   --------
    0)     2928     224   update_sd_lb_stats+0xbc/0x4ac
    1)     2704     160   find_busiest_group+0x31/0x1f1
    2)     2544     256   load_balance+0xd9/0x662
    3)     2288      80   idle_balance+0xbb/0x130
    4)     2208     128   __schedule+0x26e/0x5b9
    5)     2080      16   schedule+0x64/0x66
    6)     2064     128   schedule_timeout+0x34/0xe0
    7)     1936     112   wait_for_common+0x97/0xf1
    8)     1824      16   wait_for_completion+0x1d/0x1f
    9)     1808     128   flush_work+0xfe/0x119
   10)     1680      16   tty_flush_to_ldisc+0x1e/0x20
   11)     1664      48   input_available_p+0x1d/0x5c
   12)     1616      48   n_tty_poll+0x6d/0x134
   13)     1568      64   tty_poll+0x64/0x7f
   14)     1504     880   do_select+0x31e/0x511
   15)      624     400   core_sys_select+0x177/0x216
   16)      224      96   sys_select+0x91/0xb9
   17)      128     128   system_call_fastpath+0x16/0x1b

Note, if -mfentry is being used by gcc, functions get traced before
they set up the stack frame. This means that leaf level functions
are not tested by the stack tracer when -mfentry is used.

Currently, -mfentry is used by gcc 4.6.0 and above on x86 only.

추가 구현 자료

3799-3801

Ftrace의 세부 구현은 커널 소스 트리의 `kernel/trace/*.c` 파일에서 확인할 수 있다. 이 문서의 tracefs 인터페이스와 동작을 실제 코드 경로까지 추적할 때 기준이 되는 source path다.

More
----
More details can be found in the source code, in the `kernel/trace/*.c` files.