← Documents Documentation/admin-guide/reporting-issues.rst GitHub 원문 ↗

Linux 6.18.37 · Administration

Reporting issues

Linux kernel 문제의 기존 보고 검색, vanilla 재현, maintainer 탐색, taint·log·bisection 확인, 보고서 작성, stable backport와 후속 대응 절차를 설명합니다.

Source pathDocumentation/admin-guide/reporting-issues.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

reporting-issues.rst:1-1764

좋은 kernel issue report는 문제 설명만이 아니라 최신 upstream에서의 재현, 올바른 담당자와 공개 list, 환경·taint 검증, 재현 가능한 절차, 충분한 로그를 함께 제공합니다. 같은 문제를 중복 보고하지 않고 high priority 여부에 맞는 경로를 고르는 것이 첫 단계입니다.

보고 뒤에도 작업은 끝나지 않습니다. 질문과 자료 요청에 공개적으로 답하고 patch와 새 RC를 시험하며, 정해진 간격으로 정중히 상태를 갱신해야 실제 fix로 이어질 가능성이 높아집니다.

단계핵심 행동
1. 기존 보고 검색웹, LKML, subsystem list, stable list와 tracker 확인
2. 우선순위 판별regression, security issue, severe problem 여부 확인
3. 환경 정리vanilla kernel, 정상 hardware, add-on module 제거, taint 확인
4. 재현문제별 독립 절차를 만들고 최신 mainline 또는 적합한 stable에서 검증
5. 담당자 찾기`MAINTAINERS`와 `scripts/get_maintainer.pl` 사용
6. 근거 준비`.config`, `dmesg`, hardware·software 정보, 필요하면 bisection
7. 보고제목·첫 문장·요약·상세 설명 순으로 쓰고 공개 list를 올바르게 CC
8. 후속 대응공개 reply-all, 요청 자료와 patch test, RC 재시험, 상태 갱신

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: (GPL-2.0+ OR CC-BY-4.0)
2 .. See the bottom of this file for additional redistribution information.
3
4 Reporting issues
5 ++++++++++++++++
6
7
8 The short guide (aka TL;DR)
9 ===========================
10
11 Are you facing a regression with vanilla kernels from the same stable or
12 longterm series? One still supported? Then search the `LKML
13 <https://lore.kernel.org/lkml/>`_ and the `Linux stable mailing list
14 <https://lore.kernel.org/stable/>`_ archives for matching reports to join. If
15 you don't find any, install `the latest release from that series
16 <https://kernel.org/>`_. If it still shows the issue, report it to the stable
17 mailing list ([email protected]) and CC the regressions list
18 ([email protected]); ideally also CC the maintainer and the mailing
19 list for the subsystem in question.
20
21 In all other cases try your best guess which kernel part might be causing the
22 issue. Check the :ref:`MAINTAINERS <maintainers>` file for how its developers
23 expect to be told about problems, which most of the time will be by email with a
24 mailing list in CC. Check the destination's archives for matching reports;
25 search the `LKML <https://lore.kernel.org/lkml/>`_ and the web, too. If you
26 don't find any to join, install `the latest mainline kernel
27 <https://kernel.org/>`_. If the issue is present there, send a report.
28
29 The issue was fixed there, but you would like to see it resolved in a still
30 supported stable or longterm series as well? Then install its latest release.
31 If it shows the problem, search for the change that fixed it in mainline and
32 check if backporting is in the works or was discarded; if it's neither, ask
33 those who handled the change for it.
34
35 **General remarks**: When installing and testing a kernel as outlined above,
36 ensure it's vanilla (IOW: not patched and not using add-on modules). Also make
37 sure it's built and running in a healthy environment and not already tainted
38 before the issue occurs.
39
40 If you are facing multiple issues with the Linux kernel at once, report each
41 separately. While writing your report, include all information relevant to the
42 issue, like the kernel and the distro used. In case of a regression, CC the
43 regressions mailing list ([email protected]) to your report. Also try
44 to pinpoint the culprit with a bisection; if you succeed, include its
45 commit-id and CC everyone in the sign-off-by chain.
46
47 Once the report is out, answer any questions that come up and help where you
48 can. That includes keeping the ball rolling by occasionally retesting with newer
49 releases and sending a status update afterwards.
50
51 Step-by-step guide how to report issues to the kernel maintainers
52 =================================================================
53
54 The above TL;DR outlines roughly how to report issues to the Linux kernel
55 developers. It might be all that's needed for people already familiar with
56 reporting issues to Free/Libre & Open Source Software (FLOSS) projects. For
57 everyone else there is this section. It is more detailed and uses a
58 step-by-step approach. It still tries to be brief for readability and leaves
59 out a lot of details; those are described below the step-by-step guide in a
60 reference section, which explains each of the steps in more detail.
61
62 Note: this section covers a few more aspects than the TL;DR and does things in
63 a slightly different order. That's in your interest, to make sure you notice
64 early if an issue that looks like a Linux kernel problem is actually caused by
65 something else. These steps thus help to ensure the time you invest in this
66 process won't feel wasted in the end:
67
68 * Are you facing an issue with a Linux kernel a hardware or software vendor
69 provided? Then in almost all cases you are better off to stop reading this
70 document and reporting the issue to your vendor instead, unless you are
71 willing to install the latest Linux version yourself. Be aware the latter
72 will often be needed anyway to hunt down and fix issues.
73
74 * Perform a rough search for existing reports with your favorite internet
75 search engine; additionally, check the archives of the `Linux Kernel Mailing
76 List (LKML) <https://lore.kernel.org/lkml/>`_. If you find matching reports,
77 join the discussion instead of sending a new one.
78
79 * See if the issue you are dealing with qualifies as regression, security
80 issue, or a really severe problem: those are 'issues of high priority' that
81 need special handling in some steps that are about to follow.
82
83 * Make sure it's not the kernel's surroundings that are causing the issue
84 you face.
85
86 * Create a fresh backup and put system repair and restore tools at hand.
87
88 * Ensure your system does not enhance its kernels by building additional
89 kernel modules on-the-fly, which solutions like DKMS might be doing locally
90 without your knowledge.
91
92 * Check if your kernel was 'tainted' when the issue occurred, as the event
93 that made the kernel set this flag might be causing the issue you face.
94
95 * Write down coarsely how to reproduce the issue. If you deal with multiple
96 issues at once, create separate notes for each of them and make sure they
97 work independently on a freshly booted system. That's needed, as each issue
98 needs to get reported to the kernel developers separately, unless they are
99 strongly entangled.
100
101 * If you are facing a regression within a stable or longterm version line
102 (say something broke when updating from 5.10.4 to 5.10.5), scroll down to
103 'Dealing with regressions within a stable and longterm kernel line'.
104
105 * Locate the driver or kernel subsystem that seems to be causing the issue.
106 Find out how and where its developers expect reports. Note: most of the
107 time this won't be bugzilla.kernel.org, as issues typically need to be sent
108 by mail to a maintainer and a public mailing list.
109
110 * Search the archives of the bug tracker or mailing list in question
111 thoroughly for reports that might match your issue. If you find anything,
112 join the discussion instead of sending a new report.
113
114 After these preparations you'll now enter the main part:
115
116 * Unless you are already running the latest 'mainline' Linux kernel, better
117 go and install it for the reporting process. Testing and reporting with
118 the latest 'stable' Linux can be an acceptable alternative in some
119 situations; during the merge window that actually might be even the best
120 approach, but in that development phase it can be an even better idea to
121 suspend your efforts for a few days anyway. Whatever version you choose,
122 ideally use a 'vanilla' build. Ignoring these advices will dramatically
123 increase the risk your report will be rejected or ignored.
124
125 * Ensure the kernel you just installed does not 'taint' itself when
126 running.
127
128 * Reproduce the issue with the kernel you just installed. If it doesn't show
129 up there, scroll down to the instructions for issues only happening with
130 stable and longterm kernels.
131
132 * Optimize your notes: try to find and write the most straightforward way to
133 reproduce your issue. Make sure the end result has all the important
134 details, and at the same time is easy to read and understand for others
135 that hear about it for the first time. And if you learned something in this
136 process, consider searching again for existing reports about the issue.
137
138 * If your failure involves a 'panic', 'Oops', 'warning', or 'BUG', consider
139 decoding the kernel log to find the line of code that triggered the error.
140
141 * If your problem is a regression, try to narrow down when the issue was
142 introduced as much as possible.
143
144 * Start to compile the report by writing a detailed description about the
145 issue. Always mention a few things: the latest kernel version you installed
146 for reproducing, the Linux Distribution used, and your notes on how to
147 reproduce the issue. Ideally, make the kernel's build configuration
148 (.config) and the output from ``dmesg`` available somewhere on the net and
149 link to it. Include or upload all other information that might be relevant,
150 like the output/screenshot of an Oops or the output from ``lspci``. Once
151 you wrote this main part, insert a normal length paragraph on top of it
152 outlining the issue and the impact quickly. On top of this add one sentence
153 that briefly describes the problem and gets people to read on. Now give the
154 thing a descriptive title or subject that yet again is shorter. Then you're
155 ready to send or file the report like the MAINTAINERS file told you, unless
156 you are dealing with one of those 'issues of high priority': they need
157 special care which is explained in 'Special handling for high priority
158 issues' below.
159
160 * Wait for reactions and keep the thing rolling until you can accept the
161 outcome in one way or the other. Thus react publicly and in a timely manner
162 to any inquiries. Test proposed fixes. Do proactive testing: retest with at
163 least every first release candidate (RC) of a new mainline version and
164 report your results. Send friendly reminders if things stall. And try to
165 help yourself, if you don't get any help or if it's unsatisfying.
166
167
168 Reporting regressions within a stable and longterm kernel line
169 --------------------------------------------------------------
170
171 This subsection is for you, if you followed above process and got sent here at
172 the point about regression within a stable or longterm kernel version line. You
173 face one of those if something breaks when updating from 5.10.4 to 5.10.5 (a
174 switch from 5.9.15 to 5.10.5 does not qualify). The developers want to fix such
175 regressions as quickly as possible, hence there is a streamlined process to
176 report them:
177
178 * Check if the kernel developers still maintain the Linux kernel version
179 line you care about: go to the `front page of kernel.org
180 <https://kernel.org/>`_ and make sure it mentions
181 the latest release of the particular version line without an '[EOL]' tag.
182
183 * Check the archives of the `Linux stable mailing list
184 <https://lore.kernel.org/stable/>`_ for existing reports.
185
186 * Install the latest release from the particular version line as a vanilla
187 kernel. Ensure this kernel is not tainted and still shows the problem, as
188 the issue might have already been fixed there. If you first noticed the
189 problem with a vendor kernel, check a vanilla build of the last version
190 known to work performs fine as well.
191
192 * Send a short problem report to the Linux stable mailing list
193 ([email protected]) and CC the Linux regressions mailing list
194 ([email protected]); if you suspect the cause in a particular
195 subsystem, CC its maintainer and its mailing list. Roughly describe the
196 issue and ideally explain how to reproduce it. Mention the first version
197 that shows the problem and the last version that's working fine. Then
198 wait for further instructions.
199
200 The reference section below explains each of these steps in more detail.
201
202
203 Reporting issues only occurring in older kernel version lines
204 -------------------------------------------------------------
205
206 This subsection is for you, if you tried the latest mainline kernel as outlined
207 above, but failed to reproduce your issue there; at the same time you want to
208 see the issue fixed in a still supported stable or longterm series or vendor
209 kernels regularly rebased on those. If that is the case, follow these steps:
210
211 * Prepare yourself for the possibility that going through the next few steps
212 might not get the issue solved in older releases: the fix might be too big
213 or risky to get backported there.
214
215 * Perform the first three steps in the section "Dealing with regressions
216 within a stable and longterm kernel line" above.
217
218 * Search the Linux kernel version control system for the change that fixed
219 the issue in mainline, as its commit message might tell you if the fix is
220 scheduled for backporting already. If you don't find anything that way,
221 search the appropriate mailing lists for posts that discuss such an issue
222 or peer-review possible fixes; then check the discussions if the fix was
223 deemed unsuitable for backporting. If backporting was not considered at
224 all, join the newest discussion, asking if it's in the cards.
225
226 * One of the former steps should lead to a solution. If that doesn't work
227 out, ask the maintainers for the subsystem that seems to be causing the
228 issue for advice; CC the mailing list for the particular subsystem as well
229 as the stable mailing list.
230
231 The reference section below explains each of these steps in more detail.
232
233
234 Reference section: Reporting issues to the kernel maintainers
235 =============================================================
236
237 The detailed guides above outline all the major steps in brief fashion, which
238 should be enough for most people. But sometimes there are situations where even
239 experienced users might wonder how to actually do one of those steps. That's
240 what this section is for, as it will provide a lot more details on each of the
241 above steps. Consider this as reference documentation: it's possible to read it
242 from top to bottom. But it's mainly meant to skim over and a place to look up
243 details how to actually perform those steps.
244
245 A few words of general advice before digging into the details:
246
247 * The Linux kernel developers are well aware this process is complicated and
248 demands more than other FLOSS projects. We'd love to make it simpler. But
249 that would require work in various places as well as some infrastructure,
250 which would need constant maintenance; nobody has stepped up to do that
251 work, so that's just how things are for now.
252
253 * A warranty or support contract with some vendor doesn't entitle you to
254 request fixes from developers in the upstream Linux kernel community: such
255 contracts are completely outside the scope of the Linux kernel, its
256 development community, and this document. That's why you can't demand
257 anything such a contract guarantees in this context, not even if the
258 developer handling the issue works for the vendor in question. If you want
259 to claim your rights, use the vendor's support channel instead. When doing
260 so, you might want to mention you'd like to see the issue fixed in the
261 upstream Linux kernel; motivate them by saying it's the only way to ensure
262 the fix in the end will get incorporated in all Linux distributions.
263
264 * If you never reported an issue to a FLOSS project before you should consider
265 reading `How to Report Bugs Effectively
266 <https://www.chiark.greenend.org.uk/~sgtatham/bugs.html>`_, `How To Ask
267 Questions The Smart Way
268 <http://www.catb.org/esr/faqs/smart-questions.html>`_, and `How to ask good
269 questions <https://jvns.ca/blog/good-questions/>`_.
270
271 With that off the table, find below the details on how to properly report
272 issues to the Linux kernel developers.
273
274
275 Make sure you're using the upstream Linux kernel
276 ------------------------------------------------
277
278 *Are you facing an issue with a Linux kernel a hardware or software vendor
279 provided? Then in almost all cases you are better off to stop reading this
280 document and reporting the issue to your vendor instead, unless you are
281 willing to install the latest Linux version yourself. Be aware the latter
282 will often be needed anyway to hunt down and fix issues.*
283
284 Like most programmers, Linux kernel developers don't like to spend time dealing
285 with reports for issues that don't even happen with their current code. It's
286 just a waste everybody's time, especially yours. Unfortunately such situations
287 easily happen when it comes to the kernel and often leads to frustration on both
288 sides. That's because almost all Linux-based kernels pre-installed on devices
289 (Computers, Laptops, Smartphones, Routers, …) and most shipped by Linux
290 distributors are quite distant from the official Linux kernel as distributed by
291 kernel.org: these kernels from these vendors are often ancient from the point of
292 Linux development or heavily modified, often both.
293
294 Most of these vendor kernels are quite unsuitable for reporting issues to the
295 Linux kernel developers: an issue you face with one of them might have been
296 fixed by the Linux kernel developers months or years ago already; additionally,
297 the modifications and enhancements by the vendor might be causing the issue you
298 face, even if they look small or totally unrelated. That's why you should report
299 issues with these kernels to the vendor. Its developers should look into the
300 report and, in case it turns out to be an upstream issue, fix it directly
301 upstream or forward the report there. In practice that often does not work out
302 or might not what you want. You thus might want to consider circumventing the
303 vendor by installing the very latest Linux kernel core yourself. If that's an
304 option for you move ahead in this process, as a later step in this guide will
305 explain how to do that once it rules out other potential causes for your issue.
306
307 Note, the previous paragraph is starting with the word 'most', as sometimes
308 developers in fact are willing to handle reports about issues occurring with
309 vendor kernels. If they do in the end highly depends on the developers and the
310 issue in question. Your chances are quite good if the distributor applied only
311 small modifications to a kernel based on a recent Linux version; that for
312 example often holds true for the mainline kernels shipped by Debian GNU/Linux
313 Sid or Fedora Rawhide. Some developers will also accept reports about issues
314 with kernels from distributions shipping the latest stable kernel, as long as
315 it's only slightly modified; that for example is often the case for Arch Linux,
316 regular Fedora releases, and openSUSE Tumbleweed. But keep in mind, you better
317 want to use a mainline Linux and avoid using a stable kernel for this
318 process, as outlined in the section 'Install a fresh kernel for testing' in more
319 detail.
320
321 Obviously you are free to ignore all this advice and report problems with an old
322 or heavily modified vendor kernel to the upstream Linux developers. But note,
323 those often get rejected or ignored, so consider yourself warned. But it's still
324 better than not reporting the issue at all: sometimes such reports directly or
325 indirectly will help to get the issue fixed over time.
326
327
328 Search for existing reports, first run
329 --------------------------------------
330
331 *Perform a rough search for existing reports with your favorite internet
332 search engine; additionally, check the archives of the Linux Kernel Mailing
333 List (LKML). If you find matching reports, join the discussion instead of
334 sending a new one.*
335
336 Reporting an issue that someone else already brought forward is often a waste of
337 time for everyone involved, especially you as the reporter. So it's in your own
338 interest to thoroughly check if somebody reported the issue already. At this
339 step of the process it's okay to just perform a rough search: a later step will
340 tell you to perform a more detailed search once you know where your issue needs
341 to be reported to. Nevertheless, do not hurry with this step of the reporting
342 process, it can save you time and trouble.
343
344 Simply search the internet with your favorite search engine first. Afterwards,
345 search the `Linux Kernel Mailing List (LKML) archives
346 <https://lore.kernel.org/lkml/>`_.
347
348 If you get flooded with results consider telling your search engine to limit
349 search timeframe to the past month or year. And wherever you search, make sure
350 to use good search terms; vary them a few times, too. While doing so try to
351 look at the issue from the perspective of someone else: that will help you to
352 come up with other words to use as search terms. Also make sure not to use too
353 many search terms at once. Remember to search with and without information like
354 the name of the kernel driver or the name of the affected hardware component.
355 But its exact brand name (say 'ASUS Red Devil Radeon RX 5700 XT Gaming OC')
356 often is not much helpful, as it is too specific. Instead try search terms like
357 the model line (Radeon 5700 or Radeon 5000) and the code name of the main chip
358 ('Navi' or 'Navi10') with and without its manufacturer ('AMD').
359
360 In case you find an existing report about your issue, join the discussion, as
361 you might be able to provide valuable additional information. That can be
362 important even when a fix is prepared or in its final stages already, as
363 developers might look for people that can provide additional information or
364 test a proposed fix. Jump to the section 'Duties after the report went out' for
365 details on how to get properly involved.
366
367 Note, searching `bugzilla.kernel.org <https://bugzilla.kernel.org/>`_ might also
368 be a good idea, as that might provide valuable insights or turn up matching
369 reports. If you find the latter, just keep in mind: most subsystems expect
370 reports in different places, as described below in the section "Check where you
371 need to report your issue". The developers that should take care of the issue
372 thus might not even be aware of the bugzilla ticket. Hence, check the ticket if
373 the issue already got reported as outlined in this document and if not consider
374 doing so.
375
376
377 Issue of high priority?
378 -----------------------
379
380 *See if the issue you are dealing with qualifies as regression, security
381 issue, or a really severe problem: those are 'issues of high priority' that
382 need special handling in some steps that are about to follow.*
383
384 Linus Torvalds and the leading Linux kernel developers want to see some issues
385 fixed as soon as possible, hence there are 'issues of high priority' that get
386 handled slightly differently in the reporting process. Three type of cases
387 qualify: regressions, security issues, and really severe problems.
388
389 You deal with a regression if some application or practical use case running
390 fine with one Linux kernel works worse or not at all with a newer version
391 compiled using a similar configuration. The document
392 Documentation/admin-guide/reporting-regressions.rst explains this in more
393 detail. It also provides a good deal of other information about regressions you
394 might want to be aware of; it for example explains how to add your issue to the
395 list of tracked regressions, to ensure it won't fall through the cracks.
396
397 What qualifies as security issue is left to your judgment. Consider reading
398 Documentation/process/security-bugs.rst before proceeding, as it
399 provides additional details how to best handle security issues.
400
401 An issue is a 'really severe problem' when something totally unacceptably bad
402 happens. That's for example the case when a Linux kernel corrupts the data it's
403 handling or damages hardware it's running on. You're also dealing with a severe
404 issue when the kernel suddenly stops working with an error message ('kernel
405 panic') or without any farewell note at all. Note: do not confuse a 'panic' (a
406 fatal error where the kernel stop itself) with a 'Oops' (a recoverable error),
407 as the kernel remains running after the latter.
408
409
410 Ensure a healthy environment
411 ----------------------------
412
413 *Make sure it's not the kernel's surroundings that are causing the issue
414 you face.*
415
416 Problems that look a lot like a kernel issue are sometimes caused by build or
417 runtime environment. It's hard to rule out that problem completely, but you
418 should minimize it:
419
420 * Use proven tools when building your kernel, as bugs in the compiler or the
421 binutils can cause the resulting kernel to misbehave.
422
423 * Ensure your computer components run within their design specifications;
424 that's especially important for the main processor, the main memory, and the
425 motherboard. Therefore, stop undervolting or overclocking when facing a
426 potential kernel issue.
427
428 * Try to make sure it's not faulty hardware that is causing your issue. Bad
429 main memory for example can result in a multitude of issues that will
430 manifest itself in problems looking like kernel issues.
431
432 * If you're dealing with a filesystem issue, you might want to check the file
433 system in question with ``fsck``, as it might be damaged in a way that leads
434 to unexpected kernel behavior.
435
436 * When dealing with a regression, make sure it's not something else that
437 changed in parallel to updating the kernel. The problem for example might be
438 caused by other software that was updated at the same time. It can also
439 happen that a hardware component coincidentally just broke when you rebooted
440 into a new kernel for the first time. Updating the systems BIOS or changing
441 something in the BIOS Setup can also lead to problems that on look a lot
442 like a kernel regression.
443
444
445 Prepare for emergencies
446 -----------------------
447
448 *Create a fresh backup and put system repair and restore tools at hand.*
449
450 Reminder, you are dealing with computers, which sometimes do unexpected things,
451 especially if you fiddle with crucial parts like the kernel of its operating
452 system. That's what you are about to do in this process. Thus, make sure to
453 create a fresh backup; also ensure you have all tools at hand to repair or
454 reinstall the operating system as well as everything you need to restore the
455 backup.
456
457
458 Make sure your kernel doesn't get enhanced
459 ------------------------------------------
460
461 *Ensure your system does not enhance its kernels by building additional
462 kernel modules on-the-fly, which solutions like DKMS might be doing locally
463 without your knowledge.*
464
465 The risk your issue report gets ignored or rejected dramatically increases if
466 your kernel gets enhanced in any way. That's why you should remove or disable
467 mechanisms like akmods and DKMS: those build add-on kernel modules
468 automatically, for example when you install a new Linux kernel or boot it for
469 the first time. Also remove any modules they might have installed. Then reboot
470 before proceeding.
471
472 Note, you might not be aware that your system is using one of these solutions:
473 they often get set up silently when you install Nvidia's proprietary graphics
474 driver, VirtualBox, or other software that requires a some support from a
475 module not part of the Linux kernel. That why your might need to uninstall the
476 packages with such software to get rid of any 3rd party kernel module.
477
478
479 Check 'taint' flag
480 ------------------
481
482 *Check if your kernel was 'tainted' when the issue occurred, as the event
483 that made the kernel set this flag might be causing the issue you face.*
484
485 The kernel marks itself with a 'taint' flag when something happens that might
486 lead to follow-up errors that look totally unrelated. The issue you face might
487 be such an error if your kernel is tainted. That's why it's in your interest to
488 rule this out early before investing more time into this process. This is the
489 only reason why this step is here, as this process later will tell you to
490 install the latest mainline kernel; you will need to check the taint flag again
491 then, as that's when it matters because it's the kernel the report will focus
492 on.
493
494 On a running system is easy to check if the kernel tainted itself: if ``cat
495 /proc/sys/kernel/tainted`` returns '0' then the kernel is not tainted and
496 everything is fine. Checking that file is impossible in some situations; that's
497 why the kernel also mentions the taint status when it reports an internal
498 problem (a 'kernel bug'), a recoverable error (a 'kernel Oops') or a
499 non-recoverable error before halting operation (a 'kernel panic'). Look near
500 the top of the error messages printed when one of these occurs and search for a
501 line starting with 'CPU:'. It should end with 'Not tainted' if the kernel was
502 not tainted when it noticed the problem; it was tainted if you see 'Tainted:'
503 followed by a few spaces and some letters.
504
505 If your kernel is tainted, study Documentation/admin-guide/tainted-kernels.rst
506 to find out why. Try to eliminate the reason. Often it's caused by one these
507 three things:
508
509 1. A recoverable error (a 'kernel Oops') occurred and the kernel tainted
510 itself, as the kernel knows it might misbehave in strange ways after that
511 point. In that case check your kernel or system log and look for a section
512 that starts with this::
513
514 Oops: 0000 [#1] SMP
515
516 That's the first Oops since boot-up, as the '#1' between the brackets shows.
517 Every Oops and any other problem that happens after that point might be a
518 follow-up problem to that first Oops, even if both look totally unrelated.
519 Rule this out by getting rid of the cause for the first Oops and reproducing
520 the issue afterwards. Sometimes simply restarting will be enough, sometimes
521 a change to the configuration followed by a reboot can eliminate the Oops.
522 But don't invest too much time into this at this point of the process, as
523 the cause for the Oops might already be fixed in the newer Linux kernel
524 version you are going to install later in this process.
525
526 2. Your system uses a software that installs its own kernel modules, for
527 example Nvidia's proprietary graphics driver or VirtualBox. The kernel
528 taints itself when it loads such module from external sources (even if
529 they are Open Source): they sometimes cause errors in unrelated kernel
530 areas and thus might be causing the issue you face. You therefore have to
531 prevent those modules from loading when you want to report an issue to the
532 Linux kernel developers. Most of the time the easiest way to do that is:
533 temporarily uninstall such software including any modules they might have
534 installed. Afterwards reboot.
535
536 3. The kernel also taints itself when it's loading a module that resides in
537 the staging tree of the Linux kernel source. That's a special area for
538 code (mostly drivers) that does not yet fulfill the normal Linux kernel
539 quality standards. When you report an issue with such a module it's
540 obviously okay if the kernel is tainted; just make sure the module in
541 question is the only reason for the taint. If the issue happens in an
542 unrelated area reboot and temporarily block the module from being loaded
543 by specifying ``foo.blacklist=1`` as kernel parameter (replace 'foo' with
544 the name of the module in question).
545
546
547 Document how to reproduce issue
548 -------------------------------
549
550 *Write down coarsely how to reproduce the issue. If you deal with multiple
551 issues at once, create separate notes for each of them and make sure they
552 work independently on a freshly booted system. That's needed, as each issue
553 needs to get reported to the kernel developers separately, unless they are
554 strongly entangled.*
555
556 If you deal with multiple issues at once, you'll have to report each of them
557 separately, as they might be handled by different developers. Describing
558 various issues in one report also makes it quite difficult for others to tear
559 it apart. Hence, only combine issues in one report if they are very strongly
560 entangled.
561
562 Additionally, during the reporting process you will have to test if the issue
563 happens with other kernel versions. Therefore, it will make your work easier if
564 you know exactly how to reproduce an issue quickly on a freshly booted system.
565
566 Note: it's often fruitless to report issues that only happened once, as they
567 might be caused by a bit flip due to cosmic radiation. That's why you should
568 try to rule that out by reproducing the issue before going further. Feel free
569 to ignore this advice if you are experienced enough to tell a one-time error
570 due to faulty hardware apart from a kernel issue that rarely happens and thus
571 is hard to reproduce.
572
573
574 Regression in stable or longterm kernel?
575 ----------------------------------------
576
577 *If you are facing a regression within a stable or longterm version line
578 (say something broke when updating from 5.10.4 to 5.10.5), scroll down to
579 'Dealing with regressions within a stable and longterm kernel line'.*
580
581 Regression within a stable and longterm kernel version line are something the
582 Linux developers want to fix badly, as such issues are even more unwanted than
583 regression in the main development branch, as they can quickly affect a lot of
584 people. The developers thus want to learn about such issues as quickly as
585 possible, hence there is a streamlined process to report them. Note,
586 regressions with newer kernel version line (say something broke when switching
587 from 5.9.15 to 5.10.5) do not qualify.
588
589
590 Check where you need to report your issue
591 -----------------------------------------
592
593 *Locate the driver or kernel subsystem that seems to be causing the issue.
594 Find out how and where its developers expect reports. Note: most of the
595 time this won't be bugzilla.kernel.org, as issues typically need to be sent
596 by mail to a maintainer and a public mailing list.*
597
598 It's crucial to send your report to the right people, as the Linux kernel is a
599 big project and most of its developers are only familiar with a small subset of
600 it. Quite a few programmers for example only care for just one driver, for
601 example one for a WiFi chip; its developer likely will only have small or no
602 knowledge about the internals of remote or unrelated "subsystems", like the TCP
603 stack, the PCIe/PCI subsystem, memory management or file systems.
604
605 Problem is: the Linux kernel lacks a central bug tracker where you can simply
606 file your issue and make it reach the developers that need to know about it.
607 That's why you have to find the right place and way to report issues yourself.
608 You can do that with the help of a script (see below), but it mainly targets
609 kernel developers and experts. For everybody else the MAINTAINERS file is the
610 better place.
611
612 How to read the MAINTAINERS file
613 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
614 To illustrate how to use the :ref:`MAINTAINERS <maintainers>` file, let's assume
615 the WiFi in your Laptop suddenly misbehaves after updating the kernel. In that
616 case it's likely an issue in the WiFi driver. Obviously it could also be some
617 code it builds upon, but unless you suspect something like that stick to the
618 driver. If it's really something else, the driver's developers will get the
619 right people involved.
620
621 Sadly, there is no way to check which code is driving a particular hardware
622 component that is both universal and easy.
623
624 In case of a problem with the WiFi driver you for example might want to look at
625 the output of ``lspci -k``, as it lists devices on the PCI/PCIe bus and the
626 kernel module driving it::
627
628 [user@something ~]$ lspci -k
629 [...]
630 3a:00.0 Network controller: Qualcomm Atheros QCA6174 802.11ac Wireless Network Adapter (rev 32)
631 Subsystem: Bigfoot Networks, Inc. Device 1535
632 Kernel driver in use: ath10k_pci
633 Kernel modules: ath10k_pci
634 [...]
635
636 But this approach won't work if your WiFi chip is connected over USB or some
637 other internal bus. In those cases you might want to check your WiFi manager or
638 the output of ``ip link``. Look for the name of the problematic network
639 interface, which might be something like 'wlp58s0'. This name can be used like
640 this to find the module driving it::
641
642 [user@something ~]$ realpath --relative-to=/sys/module/ /sys/class/net/wlp58s0/device/driver/module
643 ath10k_pci
644
645 In case tricks like these don't bring you any further, try to search the
646 internet on how to narrow down the driver or subsystem in question. And if you
647 are unsure which it is: just try your best guess, somebody will help you if you
648 guessed poorly.
649
650 Once you know the driver or subsystem, you want to search for it in the
651 MAINTAINERS file. In the case of 'ath10k_pci' you won't find anything, as the
652 name is too specific. Sometimes you will need to search on the net for help;
653 but before doing so, try a somewhat shorted or modified name when searching the
654 MAINTAINERS file, as then you might find something like this::
655
656 QUALCOMM ATHEROS ATH10K WIRELESS DRIVER
657 Mail: A. Some Human <[email protected]>
658 Mailing list: [email protected]
659 Status: Supported
660 Web-page: https://wireless.wiki.kernel.org/en/users/Drivers/ath10k
661 SCM: git git://git.kernel.org/pub/scm/linux/kernel/git/kvalo/ath.git
662 Files: drivers/net/wireless/ath/ath10k/
663
664 Note: the line description will be abbreviations, if you read the plain
665 MAINTAINERS file found in the root of the Linux source tree. 'Mail:' for
666 example will be 'M:', 'Mailing list:' will be 'L', and 'Status:' will be 'S:'.
667 A section near the top of the file explains these and other abbreviations.
668
669 First look at the line 'Status'. Ideally it should be 'Supported' or
670 'Maintained'. If it states 'Obsolete' then you are using some outdated approach
671 that was replaced by a newer solution you need to switch to. Sometimes the code
672 only has someone who provides 'Odd Fixes' when feeling motivated. And with
673 'Orphan' you are totally out of luck, as nobody takes care of the code anymore.
674 That only leaves these options: arrange yourself to live with the issue, fix it
675 yourself, or find a programmer somewhere willing to fix it.
676
677 After checking the status, look for a line starting with 'bugs:': it will tell
678 you where to find a subsystem specific bug tracker to file your issue. The
679 example above does not have such a line. That is the case for most sections, as
680 Linux kernel development is completely driven by mail. Very few subsystems use
681 a bug tracker, and only some of those rely on bugzilla.kernel.org.
682
683 In this and many other cases you thus have to look for lines starting with
684 'Mail:' instead. Those mention the name and the email addresses for the
685 maintainers of the particular code. Also look for a line starting with 'Mailing
686 list:', which tells you the public mailing list where the code is developed.
687 Your report later needs to go by mail to those addresses. Additionally, for all
688 issue reports sent by email, make sure to add the Linux Kernel Mailing List
689 (LKML) <[email protected]> to CC. Don't omit either of the mailing
690 lists when sending your issue report by mail later! Maintainers are busy people
691 and might leave some work for other developers on the subsystem specific list;
692 and LKML is important to have one place where all issue reports can be found.
693
694
695 Finding the maintainers with the help of a script
696 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
697
698 For people that have the Linux sources at hand there is a second option to find
699 the proper place to report: the script 'scripts/get_maintainer.pl' which tries
700 to find all people to contact. It queries the MAINTAINERS file and needs to be
701 called with a path to the source code in question. For drivers compiled as
702 module if often can be found with a command like this::
703
704 $ modinfo ath10k_pci | grep filename | sed 's!/lib/modules/.*/kernel/!!; s!filename:!!; s!\.ko\(\|\.xz\)!!'
705 drivers/net/wireless/ath/ath10k/ath10k_pci.ko
706
707 Pass parts of this to the script::
708
709 $ ./scripts/get_maintainer.pl -f drivers/net/wireless/ath/ath10k*
710 Some Human <[email protected]> (supporter:QUALCOMM ATHEROS ATH10K WIRELESS DRIVER)
711 Another S. Human <[email protected]> (maintainer:NETWORKING DRIVERS)
712 [email protected] (open list:QUALCOMM ATHEROS ATH10K WIRELESS DRIVER)
713 [email protected] (open list:NETWORKING DRIVERS (WIRELESS))
714 [email protected] (open list:NETWORKING DRIVERS)
715 [email protected] (open list)
716
717 Don't sent your report to all of them. Send it to the maintainers, which the
718 script calls "supporter:"; additionally CC the most specific mailing list for
719 the code as well as the Linux Kernel Mailing List (LKML). In this case you thus
720 would need to send the report to 'Some Human <[email protected]>' with
722
723 Note: in case you cloned the Linux sources with git you might want to call
724 ``get_maintainer.pl`` a second time with ``--git``. The script then will look
725 at the commit history to find which people recently worked on the code in
726 question, as they might be able to help. But use these results with care, as it
727 can easily send you in a wrong direction. That for example happens quickly in
728 areas rarely changed (like old or unmaintained drivers): sometimes such code is
729 modified during tree-wide cleanups by developers that do not care about the
730 particular driver at all.
731
732
733 Search for existing reports, second run
734 ---------------------------------------
735
736 *Search the archives of the bug tracker or mailing list in question
737 thoroughly for reports that might match your issue. If you find anything,
738 join the discussion instead of sending a new report.*
739
740 As mentioned earlier already: reporting an issue that someone else already
741 brought forward is often a waste of time for everyone involved, especially you
742 as the reporter. That's why you should search for existing report again, now
743 that you know where they need to be reported to. If it's mailing list, you will
744 often find its archives on `lore.kernel.org <https://lore.kernel.org/>`_.
745
746 But some list are hosted in different places. That for example is the case for
747 the ath10k WiFi driver used as example in the previous step. But you'll often
748 find the archives for these lists easily on the net. Searching for 'archive
749 [email protected]' for example will lead you to the `Info page for the
750 ath10k mailing list <https://lists.infradead.org/mailman/listinfo/ath10k>`_,
751 which at the top links to its
752 `list archives <https://lists.infradead.org/pipermail/ath10k/>`_. Sadly this and
753 quite a few other lists miss a way to search the archives. In those cases use a
754 regular internet search engine and add something like
755 'site:lists.infradead.org/pipermail/ath10k/' to your search terms, which limits
756 the results to the archives at that URL.
757
758 It's also wise to check the internet, LKML and maybe bugzilla.kernel.org again
759 at this point. If your report needs to be filed in a bug tracker, you may want
760 to check the mailing list archives for the subsystem as well, as someone might
761 have reported it only there.
762
763 For details how to search and what to do if you find matching reports see
764 "Search for existing reports, first run" above.
765
766 Do not hurry with this step of the reporting process: spending 30 to 60 minutes
767 or even more time can save you and others quite a lot of time and trouble.
768
769
770 Install a fresh kernel for testing
771 ----------------------------------
772
773 *Unless you are already running the latest 'mainline' Linux kernel, better
774 go and install it for the reporting process. Testing and reporting with
775 the latest 'stable' Linux can be an acceptable alternative in some
776 situations; during the merge window that actually might be even the best
777 approach, but in that development phase it can be an even better idea to
778 suspend your efforts for a few days anyway. Whatever version you choose,
779 ideally use a 'vanilla' built. Ignoring these advices will dramatically
780 increase the risk your report will be rejected or ignored.*
781
782 As mentioned in the detailed explanation for the first step already: Like most
783 programmers, Linux kernel developers don't like to spend time dealing with
784 reports for issues that don't even happen with the current code. It's just a
785 waste everybody's time, especially yours. That's why it's in everybody's
786 interest that you confirm the issue still exists with the latest upstream code
787 before reporting it. You are free to ignore this advice, but as outlined
788 earlier: doing so dramatically increases the risk that your issue report might
789 get rejected or simply ignored.
790
791 In the scope of the kernel "latest upstream" normally means:
792
793 * Install a mainline kernel; the latest stable kernel can be an option, but
794 most of the time is better avoided. Longterm kernels (sometimes called 'LTS
795 kernels') are unsuitable at this point of the process. The next subsection
796 explains all of this in more detail.
797
798 * The over next subsection describes way to obtain and install such a kernel.
799 It also outlines that using a pre-compiled kernel are fine, but better are
800 vanilla, which means: it was built using Linux sources taken straight `from
801 kernel.org <https://kernel.org/>`_ and not modified or enhanced in any way.
802
803 Choosing the right version for testing
804 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
805
806 Head over to `kernel.org <https://kernel.org/>`_ to find out which version you
807 want to use for testing. Ignore the big yellow button that says 'Latest release'
808 and look a little lower at the table. At its top you'll see a line starting with
809 mainline, which most of the time will point to a pre-release with a version
810 number like '5.8-rc2'. If that's the case, you'll want to use this mainline
811 kernel for testing, as that where all fixes have to be applied first. Do not let
812 that 'rc' scare you, these 'development kernels' are pretty reliable — and you
813 made a backup, as you were instructed above, didn't you?
814
815 In about two out of every nine to ten weeks, mainline might point you to a
816 proper release with a version number like '5.7'. If that happens, consider
817 suspending the reporting process until the first pre-release of the next
818 version (5.8-rc1) shows up on kernel.org. That's because the Linux development
819 cycle then is in its two-week long 'merge window'. The bulk of the changes and
820 all intrusive ones get merged for the next release during this time. It's a bit
821 more risky to use mainline during this period. Kernel developers are also often
822 quite busy then and might have no spare time to deal with issue reports. It's
823 also quite possible that one of the many changes applied during the merge
824 window fixes the issue you face; that's why you soon would have to retest with
825 a newer kernel version anyway, as outlined below in the section 'Duties after
826 the report went out'.
827
828 That's why it might make sense to wait till the merge window is over. But don't
829 to that if you're dealing with something that shouldn't wait. In that case
830 consider obtaining the latest mainline kernel via git (see below) or use the
831 latest stable version offered on kernel.org. Using that is also acceptable in
832 case mainline for some reason does currently not work for you. An in general:
833 using it for reproducing the issue is also better than not reporting it issue
834 at all.
835
836 Better avoid using the latest stable kernel outside merge windows, as all fixes
837 must be applied to mainline first. That's why checking the latest mainline
838 kernel is so important: any issue you want to see fixed in older version lines
839 needs to be fixed in mainline first before it can get backported, which can
840 take a few days or weeks. Another reason: the fix you hope for might be too
841 hard or risky for backporting; reporting the issue again hence is unlikely to
842 change anything.
843
844 These aspects are also why longterm kernels (sometimes called "LTS kernels")
845 are unsuitable for this part of the reporting process: they are to distant from
846 the current code. Hence go and test mainline first and follow the process
847 further: if the issue doesn't occur with mainline it will guide you how to get
848 it fixed in older version lines, if that's in the cards for the fix in question.
849
850 How to obtain a fresh Linux kernel
851 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
852
853 **Using a pre-compiled kernel**: This is often the quickest, easiest, and safest
854 way for testing — especially is you are unfamiliar with the Linux kernel. The
855 problem: most of those shipped by distributors or add-on repositories are build
856 from modified Linux sources. They are thus not vanilla and therefore often
857 unsuitable for testing and issue reporting: the changes might cause the issue
858 you face or influence it somehow.
859
860 But you are in luck if you are using a popular Linux distribution: for quite a
861 few of them you'll find repositories on the net that contain packages with the
862 latest mainline or stable Linux built as vanilla kernel. It's totally okay to
863 use these, just make sure from the repository's description they are vanilla or
864 at least close to it. Additionally ensure the packages contain the latest
865 versions as offered on kernel.org. The packages are likely unsuitable if they
866 are older than a week, as new mainline and stable kernels typically get released
867 at least once a week.
868
869 Please note that you might need to build your own kernel manually later: that's
870 sometimes needed for debugging or testing fixes, as described later in this
871 document. Also be aware that pre-compiled kernels might lack debug symbols that
872 are needed to decode messages the kernel prints when a panic, Oops, warning, or
873 BUG occurs; if you plan to decode those, you might be better off compiling a
874 kernel yourself (see the end of this subsection and the section titled 'Decode
875 failure messages' for details).
876
877 **Using git**: Developers and experienced Linux users familiar with git are
878 often best served by obtaining the latest Linux kernel sources straight from the
879 `official development repository on kernel.org
880 <https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/>`_.
881 Those are likely a bit ahead of the latest mainline pre-release. Don't worry
882 about it: they are as reliable as a proper pre-release, unless the kernel's
883 development cycle is currently in the middle of a merge window. But even then
884 they are quite reliable.
885
886 **Conventional**: People unfamiliar with git are often best served by
887 downloading the sources as tarball from `kernel.org <https://kernel.org/>`_.
888
889 How to actually build a kernel is not described here, as many websites explain
890 the necessary steps already. If you are new to it, consider following one of
891 those how-to's that suggest to use ``make localmodconfig``, as that tries to
892 pick up the configuration of your current kernel and then tries to adjust it
893 somewhat for your system. That does not make the resulting kernel any better,
894 but quicker to compile.
895
896 Note: If you are dealing with a panic, Oops, warning, or BUG from the kernel,
897 please try to enable CONFIG_KALLSYMS when configuring your kernel.
898 Additionally, enable CONFIG_DEBUG_KERNEL and CONFIG_DEBUG_INFO, too; the
899 latter is the relevant one of those two, but can only be reached if you enable
900 the former. Be aware CONFIG_DEBUG_INFO increases the storage space required to
901 build a kernel by quite a bit. But that's worth it, as these options will allow
902 you later to pinpoint the exact line of code that triggers your issue. The
903 section 'Decode failure messages' below explains this in more detail.
904
905 But keep in mind: Always keep a record of the issue encountered in case it is
906 hard to reproduce. Sending an undecoded report is better than not reporting
907 the issue at all.
908
909
910 Check 'taint' flag
911 ------------------
912
913 *Ensure the kernel you just installed does not 'taint' itself when
914 running.*
915
916 As outlined above in more detail already: the kernel sets a 'taint' flag when
917 something happens that can lead to follow-up errors that look totally
918 unrelated. That's why you need to check if the kernel you just installed does
919 not set this flag. And if it does, you in almost all the cases needs to
920 eliminate the reason for it before you reporting issues that occur with it. See
921 the section above for details how to do that.
922
923
924 Reproduce issue with the fresh kernel
925 -------------------------------------
926
927 *Reproduce the issue with the kernel you just installed. If it doesn't show
928 up there, scroll down to the instructions for issues only happening with
929 stable and longterm kernels.*
930
931 Check if the issue occurs with the fresh Linux kernel version you just
932 installed. If it was fixed there already, consider sticking with this version
933 line and abandoning your plan to report the issue. But keep in mind that other
934 users might still be plagued by it, as long as it's not fixed in either stable
935 and longterm version from kernel.org (and thus vendor kernels derived from
936 those). If you prefer to use one of those or just want to help their users,
937 head over to the section "Details about reporting issues only occurring in
938 older kernel version lines" below.
939
940
941 Optimize description to reproduce issue
942 ---------------------------------------
943
944 *Optimize your notes: try to find and write the most straightforward way to
945 reproduce your issue. Make sure the end result has all the important
946 details, and at the same time is easy to read and understand for others
947 that hear about it for the first time. And if you learned something in this
948 process, consider searching again for existing reports about the issue.*
949
950 An unnecessarily complex report will make it hard for others to understand your
951 report. Thus try to find a reproducer that's straight forward to describe and
952 thus easy to understand in written form. Include all important details, but at
953 the same time try to keep it as short as possible.
954
955 In this in the previous steps you likely have learned a thing or two about the
956 issue you face. Use this knowledge and search again for existing reports
957 instead you can join.
958
959
960 Decode failure messages
961 -----------------------
962
963 *If your failure involves a 'panic', 'Oops', 'warning', or 'BUG', consider
964 decoding the kernel log to find the line of code that triggered the error.*
965
966 When the kernel detects an internal problem, it will log some information about
967 the executed code. This makes it possible to pinpoint the exact line in the
968 source code that triggered the issue and shows how it was called. But that only
969 works if you enabled CONFIG_DEBUG_INFO and CONFIG_KALLSYMS when configuring
970 your kernel. If you did so, consider to decode the information from the
971 kernel's log. That will make it a lot easier to understand what lead to the
972 'panic', 'Oops', 'warning', or 'BUG', which increases the chances that someone
973 can provide a fix.
974
975 Decoding can be done with a script you find in the Linux source tree. If you
976 are running a kernel you compiled yourself earlier, call it like this::
977
978 [user@something ~]$ sudo dmesg | ./linux-5.10.5/scripts/decode_stacktrace.sh ./linux-5.10.5/vmlinux
979
980 If you are running a packaged vanilla kernel, you will likely have to install
981 the corresponding packages with debug symbols. Then call the script (which you
982 might need to get from the Linux sources if your distro does not package it)
983 like this::
984
985 [user@something ~]$ sudo dmesg | ./linux-5.10.5/scripts/decode_stacktrace.sh \
986 /usr/lib/debug/lib/modules/5.10.10-4.1.x86_64/vmlinux /usr/src/kernels/5.10.10-4.1.x86_64/
987
988 The script will work on log lines like the following, which show the address of
989 the code the kernel was executing when the error occurred::
990
991 [ 68.387301] RIP: 0010:test_module_init+0x5/0xffa [test_module]
992
993 Once decoded, these lines will look like this::
994
995 [ 68.387301] RIP: 0010:test_module_init (/home/username/linux-5.10.5/test-module/test-module.c:16) test_module
996
997 In this case the executed code was built from the file
998 '~/linux-5.10.5/test-module/test-module.c' and the error occurred by the
999 instructions found in line '16'.
1001 The script will similarly decode the addresses mentioned in the section
1002 starting with 'Call trace', which show the path to the function where the
1003 problem occurred. Additionally, the script will show the assembler output for
1004 the code section the kernel was executing.
1006 Note, if you can't get this to work, simply skip this step and mention the
1007 reason for it in the report. If you're lucky, it might not be needed. And if it
1008 is, someone might help you to get things going. Also be aware this is just one
1009 of several ways to decode kernel stack traces. Sometimes different steps will
1010 be required to retrieve the relevant details. Don't worry about that, if that's
1011 needed in your case, developers will tell you what to do.
1014 Special care for regressions
1015 ----------------------------
1017 *If your problem is a regression, try to narrow down when the issue was
1018 introduced as much as possible.*
1020 Linux lead developer Linus Torvalds insists that the Linux kernel never
1021 worsens, that's why he deems regressions as unacceptable and wants to see them
1022 fixed quickly. That's why changes that introduced a regression are often
1023 promptly reverted if the issue they cause can't get solved quickly any other
1024 way. Reporting a regression is thus a bit like playing a kind of trump card to
1025 get something quickly fixed. But for that to happen the change that's causing
1026 the regression needs to be known. Normally it's up to the reporter to track
1027 down the culprit, as maintainers often won't have the time or setup at hand to
1028 reproduce it themselves.
1030 To find the change there is a process called 'bisection' which the document
1031 Documentation/admin-guide/bug-bisect.rst describes in detail. That process
1032 will often require you to build about ten to twenty kernel images, trying to
1033 reproduce the issue with each of them before building the next. Yes, that takes
1034 some time, but don't worry, it works a lot quicker than most people assume.
1035 Thanks to a 'binary search' this will lead you to the one commit in the source
1036 code management system that's causing the regression. Once you find it, search
1037 the net for the subject of the change, its commit id and the shortened commit id
1038 (the first 12 characters of the commit id). This will lead you to existing
1039 reports about it, if there are any.
1041 Note, a bisection needs a bit of know-how, which not everyone has, and quite a
1042 bit of effort, which not everyone is willing to invest. Nevertheless, it's
1043 highly recommended performing a bisection yourself. If you really can't or
1044 don't want to go down that route at least find out which mainline kernel
1045 introduced the regression. If something for example breaks when switching from
1046 5.5.15 to 5.8.4, then try at least all the mainline releases in that area (5.6,
1047 5.7 and 5.8) to check when it first showed up. Unless you're trying to find a
1048 regression in a stable or longterm kernel, avoid testing versions which number
1049 has three sections (5.6.12, 5.7.8), as that makes the outcome hard to
1050 interpret, which might render your testing useless. Once you found the major
1051 version which introduced the regression, feel free to move on in the reporting
1052 process. But keep in mind: it depends on the issue at hand if the developers
1053 will be able to help without knowing the culprit. Sometimes they might
1054 recognize from the report want went wrong and can fix it; other times they will
1055 be unable to help unless you perform a bisection.
1057 When dealing with regressions make sure the issue you face is really caused by
1058 the kernel and not by something else, as outlined above already.
1060 In the whole process keep in mind: an issue only qualifies as regression if the
1061 older and the newer kernel got built with a similar configuration. This can be
1062 achieved by using ``make olddefconfig``, as explained in more detail by
1063 Documentation/admin-guide/reporting-regressions.rst; that document also
1064 provides a good deal of other information about regressions you might want to be
1065 aware of.
1068 Write and send the report
1069 -------------------------
1071 *Start to compile the report by writing a detailed description about the
1072 issue. Always mention a few things: the latest kernel version you installed
1073 for reproducing, the Linux Distribution used, and your notes on how to
1074 reproduce the issue. Ideally, make the kernel's build configuration
1075 (.config) and the output from ``dmesg`` available somewhere on the net and
1076 link to it. Include or upload all other information that might be relevant,
1077 like the output/screenshot of an Oops or the output from ``lspci``. Once
1078 you wrote this main part, insert a normal length paragraph on top of it
1079 outlining the issue and the impact quickly. On top of this add one sentence
1080 that briefly describes the problem and gets people to read on. Now give the
1081 thing a descriptive title or subject that yet again is shorter. Then you're
1082 ready to send or file the report like the MAINTAINERS file told you, unless
1083 you are dealing with one of those 'issues of high priority': they need
1084 special care which is explained in 'Special handling for high priority
1085 issues' below.*
1087 Now that you have prepared everything it's time to write your report. How to do
1088 that is partly explained by the three documents linked to in the preface above.
1089 That's why this text will only mention a few of the essentials as well as
1090 things specific to the Linux kernel.
1092 There is one thing that fits both categories: the most crucial parts of your
1093 report are the title/subject, the first sentence, and the first paragraph.
1094 Developers often get quite a lot of mail. They thus often just take a few
1095 seconds to skim a mail before deciding to move on or look closer. Thus: the
1096 better the top section of your report, the higher are the chances that someone
1097 will look into it and help you. And that is why you should ignore them for now
1098 and write the detailed report first. ;-)
1100 Things each report should mention
1101 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1103 Describe in detail how your issue happens with the fresh vanilla kernel you
1104 installed. Try to include the step-by-step instructions you wrote and optimized
1105 earlier that outline how you and ideally others can reproduce the issue; in
1106 those rare cases where that's impossible try to describe what you did to
1107 trigger it.
1109 Also include all the relevant information others might need to understand the
1110 issue and its environment. What's actually needed depends a lot on the issue,
1111 but there are some things you should include always:
1113 * the output from ``cat /proc/version``, which contains the Linux kernel
1114 version number and the compiler it was built with.
1116 * the Linux distribution the machine is running (``hostnamectl | grep
1117 "Operating System"``)
1119 * the architecture of the CPU and the operating system (``uname -mi``)
1121 * if you are dealing with a regression and performed a bisection, mention the
1122 subject and the commit-id of the change that is causing it.
1124 In a lot of cases it's also wise to make two more things available to those
1125 that read your report:
1127 * the configuration used for building your Linux kernel (the '.config' file)
1129 * the kernel's messages that you get from ``dmesg`` written to a file. Make
1130 sure that it starts with a line like 'Linux version 5.8-1
1131 ([email protected]) (gcc (GCC) 10.2.1, GNU ld version 2.34) #1 SMP Mon Aug
1132 3 14:54:37 UTC 2020' If it's missing, then important messages from the first
1133 boot phase already got discarded. In this case instead consider using
1134 ``journalctl -b 0 -k``; alternatively you can also reboot, reproduce the
1135 issue and call ``dmesg`` right afterwards.
1137 These two files are big, that's why it's a bad idea to put them directly into
1138 your report. If you are filing the issue in a bug tracker then attach them to
1139 the ticket. If you report the issue by mail do not attach them, as that makes
1140 the mail too large; instead do one of these things:
1142 * Upload the files somewhere public (your website, a public file paste
1143 service, a ticket created just for this purpose on `bugzilla.kernel.org
1144 <https://bugzilla.kernel.org/>`_, ...) and include a link to them in your
1145 report. Ideally use something where the files stay available for years, as
1146 they could be useful to someone many years from now; this for example can
1147 happen if five or ten years from now a developer works on some code that was
1148 changed just to fix your issue.
1150 * Put the files aside and mention you will send them later in individual
1151 replies to your own mail. Just remember to actually do that once the report
1152 went out. ;-)
1154 Things that might be wise to provide
1155 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1157 Depending on the issue you might need to add more background data. Here are a
1158 few suggestions what often is good to provide:
1160 * If you are dealing with a 'warning', an 'OOPS' or a 'panic' from the kernel,
1161 include it. If you can't copy'n'paste it, try to capture a netconsole trace
1162 or at least take a picture of the screen.
1164 * If the issue might be related to your computer hardware, mention what kind
1165 of system you use. If you for example have problems with your graphics card,
1166 mention its manufacturer, the card's model, and what chip is uses. If it's a
1167 laptop mention its name, but try to make sure it's meaningful. 'Dell XPS 13'
1168 for example is not, because it might be the one from 2012; that one looks
1169 not that different from the one sold today, but apart from that the two have
1170 nothing in common. Hence, in such cases add the exact model number, which
1171 for example are '9380' or '7390' for XPS 13 models introduced during 2019.
1172 Names like 'Lenovo Thinkpad T590' are also somewhat ambiguous: there are
1173 variants of this laptop with and without a dedicated graphics chip, so try
1174 to find the exact model name or specify the main components.
1176 * Mention the relevant software in use. If you have problems with loading
1177 modules, you want to mention the versions of kmod, systemd, and udev in use.
1178 If one of the DRM drivers misbehaves, you want to state the versions of
1179 libdrm and Mesa; also specify your Wayland compositor or the X-Server and
1180 its driver. If you have a filesystem issue, mention the version of
1181 corresponding filesystem utilities (e2fsprogs, btrfs-progs, xfsprogs, ...).
1183 * Gather additional information from the kernel that might be of interest. The
1184 output from ``lspci -nn`` will for example help others to identify what
1185 hardware you use. If you have a problem with hardware you even might want to
1186 make the output from ``sudo lspci -vvv`` available, as that provides
1187 insights how the components were configured. For some issues it might be
1188 good to include the contents of files like ``/proc/cpuinfo``,
1189 ``/proc/ioports``, ``/proc/iomem``, ``/proc/modules``, or
1190 ``/proc/scsi/scsi``. Some subsystem also offer tools to collect relevant
1191 information. One such tool is ``alsa-info.sh`` `which the audio/sound
1192 subsystem developers provide <https://www.alsa-project.org/wiki/AlsaInfo>`_.
1194 Those examples should give your some ideas of what data might be wise to
1195 attach, but you have to think yourself what will be helpful for others to know.
1196 Don't worry too much about forgetting something, as developers will ask for
1197 additional details they need. But making everything important available from
1198 the start increases the chance someone will take a closer look.
1201 The important part: the head of your report
1202 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1204 Now that you have the detailed part of the report prepared let's get to the
1205 most important section: the first few sentences. Thus go to the top, add
1206 something like 'The detailed description:' before the part you just wrote and
1207 insert two newlines at the top. Now write one normal length paragraph that
1208 describes the issue roughly. Leave out all boring details and focus on the
1209 crucial parts readers need to know to understand what this is all about; if you
1210 think this bug affects a lot of users, mention this to get people interested.
1212 Once you did that insert two more lines at the top and write a one sentence
1213 summary that explains quickly what the report is about. After that you have to
1214 get even more abstract and write an even shorter subject/title for the report.
1216 Now that you have written this part take some time to optimize it, as it is the
1217 most important parts of your report: a lot of people will only read this before
1218 they decide if reading the rest is time well spent.
1220 Now send or file the report like the :ref:`MAINTAINERS <maintainers>` file told
1221 you, unless it's one of those 'issues of high priority' outlined earlier: in
1222 that case please read the next subsection first before sending the report on
1223 its way.
1225 Special handling for high priority issues
1226 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1228 Reports for high priority issues need special handling.
1230 **Severe issues**: make sure the subject or ticket title as well as the first
1231 paragraph makes the severeness obvious.
1233 **Regressions**: make the report's subject start with '[REGRESSION]'.
1235 In case you performed a successful bisection, use the title of the change that
1236 introduced the regression as the second part of your subject. Make the report
1237 also mention the commit id of the culprit. In case of an unsuccessful bisection,
1238 make your report mention the latest tested version that's working fine (say 5.7)
1239 and the oldest where the issue occurs (say 5.8-rc1).
1241 When sending the report by mail, CC the Linux regressions mailing list
1242 ([email protected]). In case the report needs to be filed to some web
1243 tracker, proceed to do so. Once filed, forward the report by mail to the
1244 regressions list; CC the maintainer and the mailing list for the subsystem in
1245 question. Make sure to inline the forwarded report, hence do not attach it.
1246 Also add a short note at the top where you mention the URL to the ticket.
1248 When mailing or forwarding the report, in case of a successful bisection add the
1249 author of the culprit to the recipients; also CC everyone in the signed-off-by
1250 chain, which you find at the end of its commit message.
1252 **Security issues**: for these issues your will have to evaluate if a
1253 short-term risk to other users would arise if details were publicly disclosed.
1254 If that's not the case simply proceed with reporting the issue as described.
1255 For issues that bear such a risk you will need to adjust the reporting process
1256 slightly:
1258 * If the MAINTAINERS file instructed you to report the issue by mail, do not
1259 CC any public mailing lists.
1261 * If you were supposed to file the issue in a bug tracker make sure to mark
1262 the ticket as 'private' or 'security issue'. If the bug tracker does not
1263 offer a way to keep reports private, forget about it and send your report as
1264 a private mail to the maintainers instead.
1266 In both cases make sure to also mail your report to the addresses the
1267 MAINTAINERS file lists in the section 'security contact'. Ideally directly CC
1268 them when sending the report by mail. If you filed it in a bug tracker, forward
1269 the report's text to these addresses; but on top of it put a small note where
1270 you mention that you filed it with a link to the ticket.
1272 See Documentation/process/security-bugs.rst for more information.
1275 Duties after the report went out
1276 --------------------------------
1278 *Wait for reactions and keep the thing rolling until you can accept the
1279 outcome in one way or the other. Thus react publicly and in a timely manner
1280 to any inquiries. Test proposed fixes. Do proactive testing: retest with at
1281 least every first release candidate (RC) of a new mainline version and
1282 report your results. Send friendly reminders if things stall. And try to
1283 help yourself, if you don't get any help or if it's unsatisfying.*
1285 If your report was good and you are really lucky then one of the developers
1286 might immediately spot what's causing the issue; they then might write a patch
1287 to fix it, test it, and send it straight for integration in mainline while
1288 tagging it for later backport to stable and longterm kernels that need it. Then
1289 all you need to do is reply with a 'Thank you very much' and switch to a version
1290 with the fix once it gets released.
1292 But this ideal scenario rarely happens. That's why the job is only starting
1293 once you got the report out. What you'll have to do depends on the situations,
1294 but often it will be the things listed below. But before digging into the
1295 details, here are a few important things you need to keep in mind for this part
1296 of the process.
1299 General advice for further interactions
1300 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1302 **Always reply in public**: When you filed the issue in a bug tracker, always
1303 reply there and do not contact any of the developers privately about it. For
1304 mailed reports always use the 'Reply-all' function when replying to any mails
1305 you receive. That includes mails with any additional data you might want to add
1306 to your report: go to your mail applications 'Sent' folder and use 'reply-all'
1307 on your mail with the report. This approach will make sure the public mailing
1308 list(s) and everyone else that gets involved over time stays in the loop; it
1309 also keeps the mail thread intact, which among others is really important for
1310 mailing lists to group all related mails together.
1312 There are just two situations where a comment in a bug tracker or a 'Reply-all'
1313 is unsuitable:
1315 * Someone tells you to send something privately.
1317 * You were told to send something, but noticed it contains sensitive
1318 information that needs to be kept private. In that case it's okay to send it
1319 in private to the developer that asked for it. But note in the ticket or a
1320 mail that you did that, so everyone else knows you honored the request.
1322 **Do research before asking for clarifications or help**: In this part of the
1323 process someone might tell you to do something that requires a skill you might
1324 not have mastered yet. For example, you might be asked to use some test tools
1325 you never have heard of yet; or you might be asked to apply a patch to the
1326 Linux kernel sources to test if it helps. In some cases it will be fine sending
1327 a reply asking for instructions how to do that. But before going that route try
1328 to find the answer own your own by searching the internet; alternatively
1329 consider asking in other places for advice. For example ask a friend or post
1330 about it to a chatroom or forum you normally hang out.
1332 **Be patient**: If you are really lucky you might get a reply to your report
1333 within a few hours. But most of the time it will take longer, as maintainers
1334 are scattered around the globe and thus might be in a different time zone – one
1335 where they already enjoy their night away from keyboard.
1337 In general, kernel developers will take one to five business days to respond to
1338 reports. Sometimes it will take longer, as they might be busy with the merge
1339 windows, other work, visiting developer conferences, or simply enjoying a long
1340 summer holiday.
1342 The 'issues of high priority' (see above for an explanation) are an exception
1343 here: maintainers should address them as soon as possible; that's why you
1344 should wait a week at maximum (or just two days if it's something urgent)
1345 before sending a friendly reminder.
1347 Sometimes the maintainer might not be responding in a timely manner; other
1348 times there might be disagreements, for example if an issue qualifies as
1349 regression or not. In such cases raise your concerns on the mailing list and
1350 ask others for public or private replies how to move on. If that fails, it
1351 might be appropriate to get a higher authority involved. In case of a WiFi
1352 driver that would be the wireless maintainers; if there are no higher level
1353 maintainers or all else fails, it might be one of those rare situations where
1354 it's okay to get Linus Torvalds involved.
1356 **Proactive testing**: Every time the first pre-release (the 'rc1') of a new
1357 mainline kernel version gets released, go and check if the issue is fixed there
1358 or if anything of importance changed. Mention the outcome in the ticket or in a
1359 mail you sent as reply to your report (make sure it has all those in the CC
1360 that up to that point participated in the discussion). This will show your
1361 commitment and that you are willing to help. It also tells developers if the
1362 issue persists and makes sure they do not forget about it. A few other
1363 occasional retests (for example with rc3, rc5 and the final) are also a good
1364 idea, but only report your results if something relevant changed or if you are
1365 writing something anyway.
1367 With all these general things off the table let's get into the details of how
1368 to help to get issues resolved once they were reported.
1370 Inquires and testing request
1371 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1373 Here are your duties in case you got replies to your report:
1375 **Check who you deal with**: Most of the time it will be the maintainer or a
1376 developer of the particular code area that will respond to your report. But as
1377 issues are normally reported in public it could be anyone that's replying —
1378 including people that want to help, but in the end might guide you totally off
1379 track with their questions or requests. That rarely happens, but it's one of
1380 many reasons why it's wise to quickly run an internet search to see who you're
1381 interacting with. By doing this you also get aware if your report was heard by
1382 the right people, as a reminder to the maintainer (see below) might be in order
1383 later if discussion fades out without leading to a satisfying solution for the
1384 issue.
1386 **Inquiries for data**: Often you will be asked to test something or provide
1387 additional details. Try to provide the requested information soon, as you have
1388 the attention of someone that might help and risk losing it the longer you
1389 wait; that outcome is even likely if you do not provide the information within
1390 a few business days.
1392 **Requests for testing**: When you are asked to test a diagnostic patch or a
1393 possible fix, try to test it in timely manner, too. But do it properly and make
1394 sure to not rush it: mixing things up can happen easily and can lead to a lot
1395 of confusion for everyone involved. A common mistake for example is thinking a
1396 proposed patch with a fix was applied, but in fact wasn't. Things like that
1397 happen even to experienced testers occasionally, but they most of the time will
1398 notice when the kernel with the fix behaves just as one without it.
1400 What to do when nothing of substance happens
1401 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1403 Some reports will not get any reaction from the responsible Linux kernel
1404 developers; or a discussion around the issue evolved, but faded out with
1405 nothing of substance coming out of it.
1407 In these cases wait two (better: three) weeks before sending a friendly
1408 reminder: maybe the maintainer was just away from keyboard for a while when
1409 your report arrived or had something more important to take care of. When
1410 writing the reminder, kindly ask if anything else from your side is needed to
1411 get the ball running somehow. If the report got out by mail, do that in the
1412 first lines of a mail that is a reply to your initial mail (see above) which
1413 includes a full quote of the original report below: that's on of those few
1414 situations where such a 'TOFU' (Text Over, Fullquote Under) is the right
1415 approach, as then all the recipients will have the details at hand immediately
1416 in the proper order.
1418 After the reminder wait three more weeks for replies. If you still don't get a
1419 proper reaction, you first should reconsider your approach. Did you maybe try
1420 to reach out to the wrong people? Was the report maybe offensive or so
1421 confusing that people decided to completely stay away from it? The best way to
1422 rule out such factors: show the report to one or two people familiar with FLOSS
1423 issue reporting and ask for their opinion. Also ask them for their advice how
1424 to move forward. That might mean: prepare a better report and make those people
1425 review it before you send it out. Such an approach is totally fine; just
1426 mention that this is the second and improved report on the issue and include a
1427 link to the first report.
1429 If the report was proper you can send a second reminder; in it ask for advice
1430 why the report did not get any replies. A good moment for this second reminder
1431 mail is shortly after the first pre-release (the 'rc1') of a new Linux kernel
1432 version got published, as you should retest and provide a status update at that
1433 point anyway (see above).
1435 If the second reminder again results in no reaction within a week, try to
1436 contact a higher-level maintainer asking for advice: even busy maintainers by
1437 then should at least have sent some kind of acknowledgment.
1439 Remember to prepare yourself for a disappointment: maintainers ideally should
1440 react somehow to every issue report, but they are only obliged to fix those
1441 'issues of high priority' outlined earlier. So don't be too devastating if you
1442 get a reply along the lines of 'thanks for the report, I have more important
1443 issues to deal with currently and won't have time to look into this for the
1444 foreseeable future'.
1446 It's also possible that after some discussion in the bug tracker or on a list
1447 nothing happens anymore and reminders don't help to motivate anyone to work out
1448 a fix. Such situations can be devastating, but is within the cards when it
1449 comes to Linux kernel development. This and several other reasons for not
1450 getting help are explained in 'Why some issues won't get any reaction or remain
1451 unfixed after being reported' near the end of this document.
1453 Don't get devastated if you don't find any help or if the issue in the end does
1454 not get solved: the Linux kernel is FLOSS and thus you can still help yourself.
1455 You for example could try to find others that are affected and team up with
1456 them to get the issue resolved. Such a team could prepare a fresh report
1457 together that mentions how many you are and why this is something that in your
1458 option should get fixed. Maybe together you can also narrow down the root cause
1459 or the change that introduced a regression, which often makes developing a fix
1460 easier. And with a bit of luck there might be someone in the team that knows a
1461 bit about programming and might be able to write a fix.
1464 Reference for "Reporting regressions within a stable and longterm kernel line"
1465 ------------------------------------------------------------------------------
1467 This subsection provides details for the steps you need to perform if you face
1468 a regression within a stable and longterm kernel line.
1470 Make sure the particular version line still gets support
1471 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1473 *Check if the kernel developers still maintain the Linux kernel version
1474 line you care about: go to the front page of kernel.org and make sure it
1475 mentions the latest release of the particular version line without an
1476 '[EOL]' tag.*
1478 Most kernel version lines only get supported for about three months, as
1479 maintaining them longer is quite a lot of work. Hence, only one per year is
1480 chosen and gets supported for at least two years (often six). That's why you
1481 need to check if the kernel developers still support the version line you care
1482 for.
1484 Note, if kernel.org lists two stable version lines on the front page, you
1485 should consider switching to the newer one and forget about the older one:
1486 support for it is likely to be abandoned soon. Then it will get a "end-of-life"
1487 (EOL) stamp. Version lines that reached that point still get mentioned on the
1488 kernel.org front page for a week or two, but are unsuitable for testing and
1489 reporting.
1491 Search stable mailing list
1492 ~~~~~~~~~~~~~~~~~~~~~~~~~~
1494 *Check the archives of the Linux stable mailing list for existing reports.*
1496 Maybe the issue you face is already known and was fixed or is about to. Hence,
1497 `search the archives of the Linux stable mailing list
1498 <https://lore.kernel.org/stable/>`_ for reports about an issue like yours. If
1499 you find any matches, consider joining the discussion, unless the fix is
1500 already finished and scheduled to get applied soon.
1502 Reproduce issue with the newest release
1503 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1505 *Install the latest release from the particular version line as a vanilla
1506 kernel. Ensure this kernel is not tainted and still shows the problem, as
1507 the issue might have already been fixed there. If you first noticed the
1508 problem with a vendor kernel, check a vanilla build of the last version
1509 known to work performs fine as well.*
1511 Before investing any more time in this process you want to check if the issue
1512 was already fixed in the latest release of version line you're interested in.
1513 This kernel needs to be vanilla and shouldn't be tainted before the issue
1514 happens, as detailed outlined already above in the section "Install a fresh
1515 kernel for testing".
1517 Did you first notice the regression with a vendor kernel? Then changes the
1518 vendor applied might be interfering. You need to rule that out by performing
1519 a recheck. Say something broke when you updated from 5.10.4-vendor.42 to
1520 5.10.5-vendor.43. Then after testing the latest 5.10 release as outlined in
1521 the previous paragraph check if a vanilla build of Linux 5.10.4 works fine as
1522 well. If things are broken there, the issue does not qualify as upstream
1523 regression and you need switch back to the main step-by-step guide to report
1524 the issue.
1526 Report the regression
1527 ~~~~~~~~~~~~~~~~~~~~~
1529 *Send a short problem report to the Linux stable mailing list
1530 ([email protected]) and CC the Linux regressions mailing list
1531 ([email protected]); if you suspect the cause in a particular
1532 subsystem, CC its maintainer and its mailing list. Roughly describe the
1533 issue and ideally explain how to reproduce it. Mention the first version
1534 that shows the problem and the last version that's working fine. Then
1535 wait for further instructions.*
1537 When reporting a regression that happens within a stable or longterm kernel
1538 line (say when updating from 5.10.4 to 5.10.5) a brief report is enough for
1539 the start to get the issue reported quickly. Hence a rough description to the
1540 stable and regressions mailing list is all it takes; but in case you suspect
1541 the cause in a particular subsystem, CC its maintainers and its mailing list
1542 as well, because that will speed things up.
1544 And note, it helps developers a great deal if you can specify the exact version
1545 that introduced the problem. Hence if possible within a reasonable time frame,
1546 try to find that version using vanilla kernels. Let's assume something broke when
1547 your distributor released a update from Linux kernel 5.10.5 to 5.10.8. Then as
1548 instructed above go and check the latest kernel from that version line, say
1549 5.10.9. If it shows the problem, try a vanilla 5.10.5 to ensure that no patches
1550 the distributor applied interfere. If the issue doesn't manifest itself there,
1551 try 5.10.7 and then (depending on the outcome) 5.10.8 or 5.10.6 to find the
1552 first version where things broke. Mention it in the report and state that 5.10.9
1553 is still broken.
1555 What the previous paragraph outlines is basically a rough manual 'bisection'.
1556 Once your report is out your might get asked to do a proper one, as it allows to
1557 pinpoint the exact change that causes the issue (which then can easily get
1558 reverted to fix the issue quickly). Hence consider to do a proper bisection
1559 right away if time permits. See the section 'Special care for regressions' and
1560 the document Documentation/admin-guide/bug-bisect.rst for details how to
1561 perform one. In case of a successful bisection add the author of the culprit to
1562 the recipients; also CC everyone in the signed-off-by chain, which you find at
1563 the end of its commit message.
1566 Reference for "Reporting issues only occurring in older kernel version lines"
1567 -----------------------------------------------------------------------------
1569 This section provides details for the steps you need to take if you could not
1570 reproduce your issue with a mainline kernel, but want to see it fixed in older
1571 version lines (aka stable and longterm kernels).
1573 Some fixes are too complex
1574 ~~~~~~~~~~~~~~~~~~~~~~~~~~
1576 *Prepare yourself for the possibility that going through the next few steps
1577 might not get the issue solved in older releases: the fix might be too big
1578 or risky to get backported there.*
1580 Even small and seemingly obvious code-changes sometimes introduce new and
1581 totally unexpected problems. The maintainers of the stable and longterm kernels
1582 are very aware of that and thus only apply changes to these kernels that are
1583 within rules outlined in Documentation/process/stable-kernel-rules.rst.
1585 Complex or risky changes for example do not qualify and thus only get applied
1586 to mainline. Other fixes are easy to get backported to the newest stable and
1587 longterm kernels, but too risky to integrate into older ones. So be aware the
1588 fix you are hoping for might be one of those that won't be backported to the
1589 version line your care about. In that case you'll have no other choice then to
1590 live with the issue or switch to a newer Linux version, unless you want to
1591 patch the fix into your kernels yourself.
1593 Common preparations
1594 ~~~~~~~~~~~~~~~~~~~
1596 *Perform the first three steps in the section "Reporting issues only
1597 occurring in older kernel version lines" above.*
1599 You need to carry out a few steps already described in another section of this
1600 guide. Those steps will let you:
1602 * Check if the kernel developers still maintain the Linux kernel version line
1603 you care about.
1605 * Search the Linux stable mailing list for exiting reports.
1607 * Check with the latest release.
1610 Check code history and search for existing discussions
1611 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1613 *Search the Linux kernel version control system for the change that fixed
1614 the issue in mainline, as its commit message might tell you if the fix is
1615 scheduled for backporting already. If you don't find anything that way,
1616 search the appropriate mailing lists for posts that discuss such an issue
1617 or peer-review possible fixes; then check the discussions if the fix was
1618 deemed unsuitable for backporting. If backporting was not considered at
1619 all, join the newest discussion, asking if it's in the cards.*
1621 In a lot of cases the issue you deal with will have happened with mainline, but
1622 got fixed there. The commit that fixed it would need to get backported as well
1623 to get the issue solved. That's why you want to search for it or any
1624 discussions abound it.
1626 * First try to find the fix in the Git repository that holds the Linux kernel
1627 sources. You can do this with the web interfaces `on kernel.org
1628 <https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/>`_
1629 or its mirror `on GitHub <https://github.com/torvalds/linux>`_; if you have
1630 a local clone you alternatively can search on the command line with ``git
1631 log --grep=<pattern>``.
1633 If you find the fix, look if the commit message near the end contains a
1634 'stable tag' that looks like this:
1636 Cc: <[email protected]> # 5.4+
1638 If that's case the developer marked the fix safe for backporting to version
1639 line 5.4 and later. Most of the time it's getting applied there within two
1640 weeks, but sometimes it takes a bit longer.
1642 * If the commit doesn't tell you anything or if you can't find the fix, look
1643 again for discussions about the issue. Search the net with your favorite
1644 internet search engine as well as the archives for the `Linux kernel
1645 developers mailing list <https://lore.kernel.org/lkml/>`_. Also read the
1646 section `Locate kernel area that causes the issue` above and follow the
1647 instructions to find the subsystem in question: its bug tracker or mailing
1648 list archive might have the answer you are looking for.
1650 * If you see a proposed fix, search for it in the version control system as
1651 outlined above, as the commit might tell you if a backport can be expected.
1653 * Check the discussions for any indicators the fix might be too risky to get
1654 backported to the version line you care about. If that's the case you have
1655 to live with the issue or switch to the kernel version line where the fix
1656 got applied.
1658 * If the fix doesn't contain a stable tag and backporting was not discussed,
1659 join the discussion: mention the version where you face the issue and that
1660 you would like to see it fixed, if suitable.
1663 Ask for advice
1664 ~~~~~~~~~~~~~~
1666 *One of the former steps should lead to a solution. If that doesn't work
1667 out, ask the maintainers for the subsystem that seems to be causing the
1668 issue for advice; CC the mailing list for the particular subsystem as well
1669 as the stable mailing list.*
1671 If the previous three steps didn't get you closer to a solution there is only
1672 one option left: ask for advice. Do that in a mail you sent to the maintainers
1673 for the subsystem where the issue seems to have its roots; CC the mailing list
1674 for the subsystem as well as the stable mailing list ([email protected]).
1677 Why some issues won't get any reaction or remain unfixed after being reported
1678 =============================================================================
1680 When reporting a problem to the Linux developers, be aware only 'issues of high
1681 priority' (regressions, security issues, severe problems) are definitely going
1682 to get resolved. The maintainers or if all else fails Linus Torvalds himself
1683 will make sure of that. They and the other kernel developers will fix a lot of
1684 other issues as well. But be aware that sometimes they can't or won't help; and
1685 sometimes there isn't even anyone to send a report to.
1687 This is best explained with kernel developers that contribute to the Linux
1688 kernel in their spare time. Quite a few of the drivers in the kernel were
1689 written by such programmers, often because they simply wanted to make their
1690 hardware usable on their favorite operating system.
1692 These programmers most of the time will happily fix problems other people
1693 report. But nobody can force them to do, as they are contributing voluntarily.
1695 Then there are situations where such developers really want to fix an issue,
1696 but can't: sometimes they lack hardware programming documentation to do so.
1697 This often happens when the publicly available docs are superficial or the
1698 driver was written with the help of reverse engineering.
1700 Sooner or later spare time developers will also stop caring for the driver.
1701 Maybe their test hardware broke, got replaced by something more fancy, or is so
1702 old that it's something you don't find much outside of computer museums
1703 anymore. Sometimes developer stops caring for their code and Linux at all, as
1704 something different in their life became way more important. In some cases
1705 nobody is willing to take over the job as maintainer – and nobody can be forced
1706 to, as contributing to the Linux kernel is done on a voluntary basis. Abandoned
1707 drivers nevertheless remain in the kernel: they are still useful for people and
1708 removing would be a regression.
1710 The situation is not that different with developers that are paid for their
1711 work on the Linux kernel. Those contribute most changes these days. But their
1712 employers sooner or later also stop caring for their code or make its
1713 programmer focus on other things. Hardware vendors for example earn their money
1714 mainly by selling new hardware; quite a few of them hence are not investing
1715 much time and energy in maintaining a Linux kernel driver for something they
1716 stopped selling years ago. Enterprise Linux distributors often care for a
1717 longer time period, but in new versions often leave support for old and rare
1718 hardware aside to limit the scope. Often spare time contributors take over once
1719 a company orphans some code, but as mentioned above: sooner or later they will
1720 leave the code behind, too.
1722 Priorities are another reason why some issues are not fixed, as maintainers
1723 quite often are forced to set those, as time to work on Linux is limited.
1724 That's true for spare time or the time employers grant their developers to
1725 spend on maintenance work on the upstream kernel. Sometimes maintainers also
1726 get overwhelmed with reports, even if a driver is working nearly perfectly. To
1727 not get completely stuck, the programmer thus might have no other choice than
1728 to prioritize issue reports and reject some of them.
1730 But don't worry too much about all of this, a lot of drivers have active
1731 maintainers who are quite interested in fixing as many issues as possible.
1734 Closing words
1735 =============
1737 Compared with other Free/Libre & Open Source Software it's hard to report
1738 issues to the Linux kernel developers: the length and complexity of this
1739 document and the implications between the lines illustrate that. But that's how
1740 it is for now. The main author of this text hopes documenting the state of the
1741 art will lay some groundwork to improve the situation over time.
1744 ..
1745 end-of-content
1746 ..
1747 This document is maintained by Thorsten Leemhuis <[email protected]>. If
1748 you spot a typo or small mistake, feel free to let him know directly and
1749 he'll fix it. You are free to do the same in a mostly informal way if you
1750 want to contribute changes to the text, but for copyright reasons please CC
1751 [email protected] and "sign-off" your contribution as
1752 Documentation/process/submitting-patches.rst outlines in the section "Sign
1753 your work - the Developer's Certificate of Origin".
1754 ..
1755 This text is available under GPL-2.0+ or CC-BY-4.0, as stated at the top
1756 of the file. If you want to distribute this text under CC-BY-4.0 only,
1757 please use "The Linux kernel developers" for author attribution and link
1758 this as source:
1759 https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/plain/Documentation/admin-guide/reporting-issues.rst
1760 ..
1761 Note: Only the content of this RST file as found in the Linux kernel sources
1762 is available under CC-BY-4.0, as versions of this text that were processed
1763 (for example by the kernel's build system) might contain content taken from
1764 files which use a more restrictive license.

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

짧은 안내서(TL;DR)

1-50

이 문서는 `(GPL-2.0+ OR CC-BY-4.0)` 조건으로 제공되며 Linux kernel 문제를 올바른 담당자에게 보고하는 절차를 설명합니다.

같은 stable 또는 longterm 계열의 vanilla kernel 사이에서 regression이 발생했고 그 계열이 아직 지원된다면 먼저 LKML(`https://lore.kernel.org/lkml/`)과 Linux stable mailing list(`https://lore.kernel.org/stable/`) archive에서 같은 보고를 찾습니다. 없다면 그 계열의 최신 release를 `https://kernel.org/`에서 설치해 재현합니다.

최신 release에서도 문제가 나타나면 `[email protected]`에 보고하고 regressions list인 `[email protected]`를 CC합니다. 가능하면 해당 subsystem의 maintainer와 mailing list도 CC합니다.

그 밖의 문제는 원인이 될 만한 kernel 영역을 최대한 추정하고 `MAINTAINERS`에서 개발자가 원하는 보고 방법을 찾습니다. 대부분 maintainer에게 email을 보내고 공개 mailing list를 CC하는 방식입니다. 목적지 archive, LKML, 웹을 검색한 뒤 같은 보고가 없다면 최신 mainline kernel에서 재현하고 보고합니다.

Mainline에서는 고쳐졌지만 아직 지원되는 stable 또는 longterm 계열에도 수정이 필요하다면 그 계열의 최신 release를 먼저 검사합니다. 문제가 남아 있으면 mainline의 수정 commit을 찾아 backport가 진행 중인지, 거부됐는지 확인하고 어느 쪽도 아니면 변경을 처리한 사람에게 backport를 요청합니다.

시험 kernel은 vanilla, 즉 patch와 add-on module이 없어야 합니다. 정상적인 build·runtime 환경에서 실행하고 문제가 생기기 전부터 tainted 상태가 아니어야 합니다.

동시에 여러 문제가 있으면 각각 따로 보고합니다. Kernel과 배포판 등 관련 정보를 모두 포함하고, regression이면 `[email protected]`를 CC합니다. 가능하면 bisection으로 원인 commit을 찾고 commit-id를 적으며 sign-off-by chain의 모든 사람을 CC합니다.

보고 뒤에는 질문에 답하고 가능한 도움을 계속 제공해야 합니다. 새 release로 가끔 다시 시험하고 상태를 갱신해 논의가 멈추지 않게 합니다.

Kernel maintainer에게 보고하는 단계별 안내

51-167

앞의 TL;DR은 FLOSS project 문제 보고에 익숙한 사람에게 충분할 수 있습니다. 이 절은 같은 절차를 더 세분화하며, 자세한 이유와 예시는 뒤의 reference section에서 설명합니다. 순서가 조금 다른 이유는 Linux kernel 문제처럼 보이는 현상이 실제로는 다른 원인인지 일찍 확인해 시간을 낭비하지 않게 하기 위해서입니다.

Hardware 또는 software vendor가 제공한 kernel에서 생긴 문제라면 최신 Linux를 직접 설치해 볼 의향이 없는 한 보통 vendor에게 먼저 보고하는 편이 낫습니다. 문제를 추적하고 고치려면 결국 최신 upstream kernel 시험이 필요할 때가 많습니다.

선호하는 검색 엔진과 LKML archive에서 기존 보고를 대략 검색합니다. 일치하는 보고가 있으면 새 보고를 만들지 말고 기존 논의에 참여합니다.

문제가 regression, security issue, 정말 심각한 문제인지 판단합니다. 이 세 종류는 뒤 단계에서 특별히 다루는 high priority issue입니다.

Kernel 주변 환경이 원인이 아닌지 확인하고, 최신 backup과 시스템 복구·복원 도구를 준비합니다. DKMS처럼 모르는 사이 add-on kernel module을 build하는 기능을 제거하고, 문제가 생겼을 때 kernel의 taint 상태를 확인합니다.

문제를 재현하는 대략적인 절차를 기록합니다. 여러 문제라면 새로 boot한 시스템에서 서로 독립적으로 동작하는 별도 기록을 만듭니다. 강하게 얽혀 있지 않은 문제는 각각 다른 개발자가 처리할 수 있으므로 따로 보고해야 합니다.

5.10.4에서 5.10.5로 갱신할 때처럼 같은 stable 또는 longterm 계열 안에서 생긴 regression이면 해당 전용 절차로 이동합니다. 5.9.15에서 5.10.5로 바뀐 경우는 이 범주가 아닙니다.

원인으로 보이는 driver 또는 subsystem을 찾아 개발자가 보고를 기대하는 위치와 방식을 확인합니다. 대부분 `bugzilla.kernel.org`가 아니라 maintainer와 공개 mailing list로 보내는 mail입니다. 해당 tracker나 list archive를 자세히 검색하고 기존 보고가 있으면 그곳에 참여합니다.

준비가 끝나면 최신 mainline Linux kernel을 설치하는 것이 좋습니다. 상황에 따라 최신 stable이 대안이 될 수 있고 merge window에는 오히려 더 나을 수도 있지만, 며칠 기다리는 편이 더 합리적일 때도 있습니다. 어떤 버전을 고르든 vanilla build를 쓰는 것이 이상적이며, 이를 무시하면 보고가 거부되거나 무시될 위험이 크게 높아집니다.

새 kernel이 실행 중 스스로 taint되지 않는지 확인하고 문제를 재현합니다. 재현되지 않으면 stable·longterm에서만 발생하는 문제 절차로 이동합니다.

재현 기록을 가장 단순하고 이해하기 쉬운 절차로 다듬되 중요한 세부 사항은 모두 남깁니다. 과정에서 새 사실을 알았다면 기존 보고를 다시 검색합니다. Panic, Oops, warning, BUG가 관련되면 kernel log를 decode해 오류를 일으킨 코드 줄을 찾는 것도 고려합니다.

Regression이면 도입 시점을 최대한 좁힙니다. 보고서에는 최신 재현 kernel, 사용한 Linux distribution, 재현 절차를 반드시 적습니다. 가능하면 `.config`와 `dmesg`를 웹에 올려 link하고 Oops 출력이나 `lspci` 같은 관련 정보도 포함합니다.

상세 설명 위에는 영향과 문제를 빠르게 설명하는 보통 길이 문단을, 그 위에는 한 문장 요약을, 마지막으로 더 짧고 구체적인 제목을 둡니다. `MAINTAINERS`가 지시한 위치로 보내되 high priority issue는 별도 규칙을 먼저 확인합니다.

보고 뒤에는 공개적으로 신속히 응답하고 제안된 수정안을 시험합니다. 새 mainline의 첫 release candidate(RC)는 적어도 매번 다시 시험해 결과를 알리고, 멈춘 논의에는 정중히 reminder를 보내며 필요한 경우 스스로 해결에 참여합니다.

일반 문제 보고 흐름
기존 보고 검색우선순위 판별환경·taint 정리최신 vanilla 재현담당자 확인근거와 보고서 작성공개 후속 대응

기존 보고 검색부터 후속 시험까지의 권장 순서입니다.

Stable·longterm 계열 내부 regression

168-202

이 절은 앞 절차에서 같은 stable 또는 longterm 계열 내부 regression으로 분류된 경우에 사용합니다. 5.10.4에서 5.10.5로 갱신해 문제가 생긴 경우가 해당하며 5.9.15에서 5.10.5로 전환한 경우는 해당하지 않습니다. 개발자는 이런 regression을 빠르게 고치려 하므로 간소화된 보고 절차가 있습니다.

`kernel.org` 첫 화면에서 해당 version line의 최신 release가 표시되고 `[EOL]` tag가 없는지 확인해 아직 유지보수 중인지 검사합니다.

Linux stable mailing list archive에서 같은 보고를 검색합니다. 그 뒤 해당 계열 최신 release를 vanilla kernel로 설치하고 taint되지 않았으며 문제가 남아 있는지 확인합니다. Vendor kernel에서 처음 발견했다면 마지막으로 정상인 버전의 vanilla build도 정상인지 검사합니다.

`[email protected]`에 짧게 보고하고 `[email protected]`를 CC합니다. 특정 subsystem이 의심되면 maintainer와 list도 CC합니다. 문제와 재현법을 대략 설명하고 처음 문제를 보이는 버전과 마지막 정상 버전을 적은 뒤 추가 지시를 기다립니다.

오래된 kernel 계열에서만 발생하는 문제

203-233

최신 mainline에서는 재현되지 않지만 아직 지원되는 stable·longterm 계열이나 이를 정기적으로 rebase하는 vendor kernel에도 수정이 필요할 때 이 절차를 사용합니다. 수정이 너무 크거나 위험해 오래된 release로 backport되지 못할 가능성을 먼저 받아들여야 합니다.

앞의 stable·longterm regression 절차에서 지원 상태 확인, stable list 검색, 최신 release 재현의 첫 세 단계를 수행합니다.

Linux kernel version control system에서 mainline 문제를 고친 변경을 찾습니다. Commit message는 이미 backport가 예정됐는지 알려줄 수 있습니다. 찾지 못하면 관련 mailing list에서 문제와 수정안의 review 논의를 찾아 backport가 부적합하다고 판단됐는지 확인합니다.

Backport가 전혀 논의되지 않았다면 가장 최근 논의에 참여해 가능한지 묻습니다. 그래도 해결되지 않으면 원인 subsystem maintainer에게 조언을 요청하고 해당 subsystem list와 stable list를 CC합니다.

Reference 소개와 upstream kernel 확인

234-327

앞의 안내는 주요 단계를 짧게 설명합니다. 이 reference section은 각 단계를 실제로 수행할 때 필요한 세부 사항을 제공합니다. 처음부터 끝까지 읽을 수도 있지만 주로 필요한 항목을 찾아보는 용도입니다.

Linux kernel 개발자도 이 절차가 다른 FLOSS project보다 복잡하고 요구 사항이 많다는 점을 알고 있습니다. 더 단순하게 만들려면 여러 영역의 작업과 지속적으로 유지할 infrastructure가 필요하지만 맡을 사람이 없어 현재 방식이 유지되고 있습니다.

Vendor와 맺은 보증 또는 지원 계약은 upstream Linux kernel community에 수정을 요구할 권리를 주지 않습니다. 그런 권리를 주장하려면 vendor 지원 채널을 사용해야 합니다. 이때 upstream에서도 고치길 원한다고 말하면 모든 Linux distribution에 수정이 들어갈 유일한 길이라는 점으로 vendor를 설득할 수 있습니다.

FLOSS project에 처음 보고한다면 `How to Report Bugs Effectively`, `How To Ask Questions The Smart Way`, `How to ask good questions`도 읽어보는 것이 좋습니다.

Device에 미리 설치되거나 distribution이 제공한 많은 kernel은 `kernel.org`의 공식 Linux와 상당히 다릅니다. 오래됐거나 크게 수정됐으며 둘 다인 경우도 많습니다. 그 문제는 upstream에서 이미 고쳤을 수 있고 vendor 변경이 원인일 수도 있으므로 보통 vendor에게 보고해야 합니다.

Vendor 개발자는 보고를 조사해 upstream 문제라면 직접 upstream에서 고치거나 보고를 전달해야 합니다. 현실적으로 잘 되지 않거나 원하는 방식이 아닐 수 있으므로 가능하다면 최신 Linux kernel core를 직접 설치해 upstream 여부를 검사할 수 있습니다.

예외적으로 최근 Linux에 작은 변경만 적용한 vendor kernel 보고를 받는 개발자도 있습니다. Debian GNU/Linux Sid나 Fedora Rawhide의 mainline kernel이 이런 경우가 많고, Arch Linux, 일반 Fedora release, openSUSE Tumbleweed처럼 최신 stable에 작은 변경만 적용한 배포판도 받아들여질 수 있습니다.

그래도 이 절차에서는 stable보다 mainline Linux를 사용하는 편이 낫습니다. 오래됐거나 크게 수정된 vendor kernel 보고는 거부되거나 무시될 수 있습니다. 다만 전혀 보고하지 않는 것보다는 나아 직접 또는 간접적으로 장기적인 수정에 도움이 될 때도 있습니다.

기존 보고 검색과 high priority 판단

328-409

중복 보고는 관련된 모든 사람, 특히 보고자의 시간을 낭비합니다. 이 단계에서는 대략 검색하고 담당 위치를 안 뒤 다시 자세히 검색합니다. 서두르지 말고 먼저 일반 웹 검색과 LKML archive를 확인합니다.

결과가 너무 많으면 최근 한 달이나 1년으로 기간을 제한합니다. 검색어를 여러 방식으로 바꾸고 다른 사람의 관점에서 문제를 표현해 봅니다. Driver나 hardware component 이름을 넣거나 빼서 검색하되 너무 많은 용어를 한꺼번에 쓰지 않습니다.

`ASUS Red Devil Radeon RX 5700 XT Gaming OC` 같은 정확한 상품명은 지나치게 구체적일 수 있습니다. `Radeon 5700`, `Radeon 5000`, chip codename인 `Navi` 또는 `Navi10`, 제조사 `AMD` 같은 더 일반적인 조합도 검색합니다.

기존 보고를 찾으면 논의에 참여합니다. 수정이 거의 끝난 경우에도 추가 정보나 제안된 fix 시험이 필요할 수 있습니다. `bugzilla.kernel.org`도 유용하지만 많은 subsystem은 다른 위치에서 보고를 받으므로 ticket이 실제 담당자에게 전달됐는지 확인해야 합니다.

High priority issue는 regression, security issue, really severe problem의 세 종류입니다. Regression은 비슷한 설정으로 build한 이전 Linux에서는 잘 되던 application이나 실제 사용 사례가 새 버전에서 나빠지거나 동작하지 않는 경우입니다. 자세한 내용과 tracking 방법은 `Documentation/admin-guide/reporting-regressions.rst`를 참조합니다.

Security issue 여부는 보고자가 판단하되 `Documentation/process/security-bugs.rst`를 먼저 읽습니다. 정말 심각한 문제는 kernel이 data를 손상하거나 hardware를 망가뜨리는 경우, `kernel panic`과 함께 멈추거나 아무 메시지 없이 멈추는 경우입니다.

`panic`은 kernel이 스스로 중지하는 fatal error이고 `Oops`는 recoverable error라 kernel이 계속 실행됩니다. 두 상태를 혼동하지 않아야 합니다.

환경, backup, add-on module, taint와 재현

410-589

Kernel 문제처럼 보이는 현상이 build 또는 runtime 환경에서 생길 수 있습니다. Compile에는 검증된 compiler와 binutils를 사용하고 CPU, main memory, motherboard를 설계 사양 안에서 운용하며 undervolting과 overclocking을 중지합니다.

고장 난 hardware, 특히 bad memory가 kernel 문제처럼 보이는 다양한 오류를 만들 수 있으므로 확인합니다. Filesystem 문제라면 `fsck`로 손상을 검사합니다. Regression이라면 kernel과 동시에 갱신된 software, 우연히 고장 난 hardware, BIOS update나 BIOS Setup 변경이 원인이 아닌지도 확인합니다.

Kernel처럼 운영체제의 핵심 부분을 바꾸기 전에 새 backup을 만들고 운영체제를 복구·재설치할 도구와 backup 복원 수단을 준비합니다.

Kernel이 어떤 방식으로든 추가 확장되면 보고가 무시되거나 거부될 위험이 크게 높아집니다. `akmods`와 `DKMS`처럼 새 kernel 설치나 첫 boot 때 add-on module을 자동 build하는 mechanism을 제거하거나 비활성화하고 설치된 module도 지운 뒤 reboot합니다.

Nvidia proprietary graphics driver, VirtualBox처럼 upstream Linux에 없는 module이 필요한 software는 이런 mechanism을 조용히 설치할 수 있습니다. Third-party kernel module을 없애려면 해당 package를 임시로 제거해야 할 수 있습니다.

Kernel은 이후의 무관해 보이는 오류를 일으킬 수 있는 사건이 생기면 `taint` flag를 설정합니다. 실행 중인 시스템에서 `cat /proc/sys/kernel/tainted`가 `0`을 반환하면 tainted되지 않은 상태입니다.

파일을 확인할 수 없는 상황에서는 kernel bug, Oops, panic log 상단의 `CPU:`로 시작하는 줄을 봅니다. 끝이 `Not tainted`면 당시 tainted되지 않았고, `Tainted:` 뒤에 문자들이 있으면 tainted 상태였습니다. 원인은 `Documentation/admin-guide/tainted-kernels.rst`에서 해석합니다.

첫째 원인은 recoverable `kernel Oops`일 수 있습니다. Log에서 다음과 같이 시작하는 첫 Oops를 찾습니다.

Oops: 0000 [#1] SMP

`[#1]`은 boot 뒤 첫 Oops라는 뜻입니다. 이후의 모든 Oops와 문제는 겉보기에 무관해도 첫 Oops의 후속 현상일 수 있습니다. 첫 원인을 제거하고 다시 재현하며, 단순 reboot나 설정 변경 뒤 reboot로 없어질 수 있습니다. 뒤에서 설치할 최신 kernel에 이미 수정됐을 수도 있으므로 이 단계에서 지나치게 오래 매달리지는 않습니다.

둘째 원인은 Nvidia proprietary driver나 VirtualBox처럼 자체 kernel module을 설치하는 software입니다. External source module은 Open Source여도 kernel을 taint시키고 무관한 영역의 오류를 만들 수 있으므로, module과 software를 임시 제거하고 reboot해 loading을 막습니다.

셋째 원인은 Linux source의 staging tree에 있는 module입니다. 이 영역은 정상 kernel 품질 기준을 아직 만족하지 못한 code를 담습니다. 그 module 자체 문제를 보고할 때는 taint가 허용되지만 유일한 taint 원인인지 확인합니다. 무관한 영역 문제라면 kernel parameter에 `foo.blacklist=1`을 지정해 해당 module loading을 임시 차단합니다.

여러 문제는 강하게 얽힌 경우가 아니면 각각 보고합니다. 다른 kernel 버전에서도 시험해야 하므로 새로 boot한 시스템에서 빠르게 재현할 정확한 절차를 만듭니다.

한 번만 생긴 문제는 cosmic radiation의 bit flip 같은 일회성 사건일 수 있어 보통 보고 효과가 낮습니다. 가능하면 먼저 재현해 배제합니다. 경험이 충분해 드문 kernel issue와 faulty hardware의 일회성 오류를 구별할 수 있다면 예외로 할 수 있습니다.

같은 stable·longterm version line 내부 regression은 많은 사용자에게 빠르게 영향을 주므로 전용 간소화 절차를 사용합니다. 더 새 version line으로 바꿀 때 생긴 regression은 이 범주가 아닙니다.

시험용 최신 kernel 선택과 확보

770-909

현재 최신 mainline Linux가 아니라면 보고를 위해 설치하는 편이 좋습니다. 최신 stable은 일부 상황에서 대안이 되고 merge window에는 더 나을 수 있지만, 며칠 작업을 미루는 편이 나을 때도 있습니다. 어떤 버전이든 vanilla build가 이상적입니다.

개발자는 현재 code에서 발생하지 않는 문제에 시간을 쓰기 어렵습니다. 보고 전에 최신 upstream에서도 문제가 존재하는지 확인하지 않으면 거부되거나 무시될 가능성이 크게 높아집니다.

여기서 최신 upstream은 보통 mainline kernel을 뜻합니다. 최신 stable도 선택할 수 있지만 대체로 피하는 편이 낫고, longterm 또는 LTS kernel은 이 단계에 부적합합니다. Vanilla는 `kernel.org`에서 직접 받은 source를 수정하거나 확장하지 않고 build했다는 뜻입니다.

`kernel.org`의 큰 `Latest release` button보다 아래 표의 `mainline` 행을 확인합니다. 대개 `5.8-rc2` 같은 pre-release를 가리키며 모든 fix가 먼저 들어가는 이 mainline을 시험해야 합니다. `rc`가 붙은 development kernel도 비교적 신뢰할 수 있고 앞에서 backup을 준비했습니다.

9~10주 주기 중 약 2주는 mainline이 `5.7` 같은 정식 release를 가리키는 merge window입니다. 이때 다음 release의 크고 침습적인 변경이 합쳐져 위험이 약간 높고 개발자도 바쁩니다. 많은 변경 중 하나가 문제를 고칠 수 있으므로 곧 다음 `rc1`에서 다시 시험해야 할 수도 있습니다.

급하지 않다면 merge window 종료를 기다릴 수 있습니다. 기다릴 수 없는 문제라면 Git의 최신 mainline을 받거나 `kernel.org`의 최신 stable을 사용합니다. Mainline이 동작하지 않을 때 stable에서 재현하는 것도 보고하지 않는 것보다 낫습니다.

Merge window 밖에서는 최신 stable만 시험하지 않는 편이 좋습니다. 모든 fix는 먼저 mainline에 들어가야 하며 오래된 계열에 backport되기까지 며칠이나 몇 주가 걸릴 수 있고, 너무 어렵거나 위험해 backport되지 않을 수도 있습니다. LTS는 현재 code와 더 멀어 이 단계에 부적합합니다.

Pre-compiled kernel은 빠르고 쉽고 안전하지만 distribution이나 add-on repository의 package는 수정된 source로 build돼 vanilla가 아닐 수 있습니다. 인기 distribution에는 최신 mainline 또는 stable을 vanilla로 package한 repository가 있으므로 설명을 확인하고 `kernel.org` release보다 일주일 이상 오래되지 않았는지 봅니다.

나중에 debug 또는 fix 시험을 위해 직접 build해야 할 수 있습니다. Pre-compiled package는 panic, Oops, warning, BUG를 decode하는 debug symbol이 없을 수도 있습니다.

Git에 익숙한 개발자와 사용자는 `https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/`의 공식 development repository에서 최신 source를 받는 것이 좋습니다. 최신 pre-release보다 조금 앞설 수 있지만 merge window 중간이 아니라면 동등하게 신뢰할 수 있고, 그때도 대체로 안정적입니다.

Git에 익숙하지 않다면 `kernel.org`의 tarball을 받습니다. Build 방법은 다른 자료를 참고하되 처음이라면 현재 설정을 읽어 시스템에 맞게 줄여 compile 시간을 단축하는 `make localmodconfig` 절차가 편리합니다. 결과 kernel의 품질 자체가 높아지는 것은 아닙니다.

Panic, Oops, warning, BUG를 다룬다면 설정에서 `CONFIG_KALLSYMS`를 켭니다. `CONFIG_DEBUG_KERNEL`과 `CONFIG_DEBUG_INFO`도 활성화합니다. `CONFIG_DEBUG_INFO`는 공간을 많이 쓰지만 나중에 문제를 일으킨 정확한 source 줄을 찾게 해줍니다.

재현이 어려운 경우를 대비해 발생한 문제 기록을 항상 보관합니다. Decode하지 못한 보고라도 전혀 보고하지 않는 것보다 낫습니다.

새 kernel 재현, log decode, regression bisection

910-1067

새로 설치한 kernel도 taint flag를 설정하지 않는지 다시 확인합니다. Taint되면 대부분 보고 전에 원인을 제거해야 하며 방법은 앞 절을 따릅니다.

새 Linux에서 문제를 재현합니다. 이미 고쳐졌다면 이 계열을 계속 사용하고 보고를 중단할 수 있지만 stable·longterm과 그 기반 vendor kernel 사용자는 여전히 영향받을 수 있습니다. 그 사용자도 돕고 싶다면 오래된 version line 전용 절차로 이동합니다.

재현법은 글로 설명하기 쉽고 처음 보는 사람도 이해할 정도로 단순하게 만듭니다. 중요한 정보는 모두 포함하되 가능한 짧게 유지합니다. 이 과정에서 얻은 새 정보로 기존 보고를 다시 검색합니다.

`CONFIG_DEBUG_INFO`와 `CONFIG_KALLSYMS`를 켰다면 kernel이 기록한 실행 주소를 decode해 panic, Oops, warning, BUG를 일으킨 정확한 source 줄과 call path를 찾을 수 있습니다. 직접 build한 kernel은 다음처럼 source tree의 script를 실행합니다.

[user@something ~]$ sudo dmesg | ./linux-5.10.5/scripts/decode_stacktrace.sh ./linux-5.10.5/vmlinux

Package된 vanilla kernel은 대응하는 debug symbol package를 설치한 뒤 다음처럼 `vmlinux`와 source path를 전달합니다.

[user@something ~]$ sudo dmesg | ./linux-5.10.5/scripts/decode_stacktrace.sh \
 /usr/lib/debug/lib/modules/5.10.10-4.1.x86_64/vmlinux /usr/src/kernels/5.10.10-4.1.x86_64/

Script는 오류 당시 실행 주소를 나타내는 다음 log 줄을 처리합니다.

[   68.387301] RIP: 0010:test_module_init+0x5/0xffa [test_module]

Decode 결과는 다음처럼 source file과 line number를 표시합니다.

[   68.387301] RIP: 0010:test_module_init (/home/username/linux-5.10.5/test-module/test-module.c:16) test_module

이 예에서는 `~/linux-5.10.5/test-module/test-module.c`의 line `16`에서 오류가 발생했습니다. Script는 `Call trace`의 주소도 decode해 문제가 난 function까지의 경로를 보여주고 해당 code section의 assembler output도 표시합니다.

Decode가 되지 않으면 이 단계를 건너뛰고 보고서에 이유를 적습니다. 다른 stack trace decode 방법이나 추가 단계가 필요할 수 있으며, 필요하면 개발자가 방법을 안내할 것입니다.

Linus Torvalds는 Linux kernel이 나빠지는 regression을 받아들일 수 없는 것으로 보고 빠른 수정을 원합니다. 빠르게 해결되지 않으면 원인 변경을 revert하기도 하지만 이를 위해 culprit를 알아야 합니다. Maintainer가 직접 재현할 시간이나 환경이 없으므로 보통 보고자가 추적합니다.

`Documentation/admin-guide/bug-bisect.rst`가 설명하는 bisection은 보통 10~20개의 kernel image를 build하고 각각 재현 시험합니다. Binary search로 regression을 일으킨 하나의 commit을 찾은 뒤 subject, 전체 commit-id, 앞 12자의 짧은 commit-id를 검색해 기존 보고를 찾습니다.

Bisection이 어렵다면 적어도 어느 mainline release에서 시작됐는지 찾습니다. 5.5.15에서 5.8.4로 전환해 문제가 생겼다면 5.6, 5.7, 5.8을 시험합니다. Stable·longterm 내부 regression을 찾는 경우가 아니라면 5.6.12나 5.7.8처럼 세 부분 version은 결과 해석을 어렵게 하므로 피합니다.

문제가 실제 kernel regression인지 다시 확인하고 이전·새 kernel을 비슷한 설정으로 build합니다. `make olddefconfig`를 쓰는 방법과 추가 기준은 `Documentation/admin-guide/reporting-regressions.rst`를 참조합니다.

보고서 작성과 high priority 특별 처리

1068-1274

먼저 상세 보고를 작성합니다. 가장 중요한 부분은 제목, 첫 문장, 첫 문단이지만 개발자가 몇 초만 훑고 더 읽을지 결정하므로 상세 내용을 완성한 뒤 마지막에 이 앞부분을 다듬는 편이 좋습니다.

새로 설치한 vanilla kernel에서 문제가 어떻게 발생하는지 자세히 설명하고, 다른 사람이 재현할 단계별 절차를 포함합니다. 재현이 불가능한 드문 경우에는 무엇을 하다가 문제가 발생했는지 설명합니다.

항상 `cat /proc/version` 출력으로 kernel version과 compiler를, `hostnamectl | grep "Operating System"`으로 distribution을, `uname -mi`로 CPU와 운영체제 architecture를 제공합니다. Regression을 bisection했다면 원인 변경의 subject와 commit-id를 적습니다.

Kernel build에 사용한 `.config`와 파일로 저장한 `dmesg`도 보통 제공하는 것이 좋습니다. `dmesg`는 `Linux version ...`으로 시작해야 하며 초반 boot 메시지가 사라졌다면 `journalctl -b 0 -k`를 쓰거나 reboot 직후 재현해 바로 `dmesg`를 수집합니다.

이 두 파일은 커서 email 본문에 직접 넣기 좋지 않습니다. Tracker에는 attachment로 올리고, mail 보고에는 장기간 공개되는 website나 paste service 또는 `bugzilla.kernel.org` ticket에 올려 link합니다. 또는 별도 reply로 나중에 보낸다고 명시하고 실제로 보내야 합니다.

Kernel `warning`, `OOPS`, `panic`이 있으면 포함합니다. Copy할 수 없으면 netconsole trace를 수집하거나 적어도 화면 사진을 찍습니다.

Hardware 관련 문제는 system 종류를 정확히 적습니다. Graphics card는 제조사, model, chip을 적고 laptop은 연식과 구성을 구분할 exact model number를 제공합니다. `Dell XPS 13`만으로는 부족하며 `9380`, `7390`처럼 정확히 구분합니다. `Lenovo Thinkpad T590`도 discrete graphics 유무 등 주요 구성을 명시합니다.

관련 software version도 적습니다. Module loading 문제면 kmod, systemd, udev를, DRM driver 문제면 libdrm, Mesa, Wayland compositor 또는 X-Server와 driver를, filesystem 문제면 e2fsprogs, btrfs-progs, xfsprogs 같은 utility version을 제공합니다.

`lspci -nn`은 hardware 식별에 도움이 되고 `sudo lspci -vvv`는 component 설정을 자세히 보여줍니다. 필요에 따라 `/proc/cpuinfo`, `/proc/ioports`, `/proc/iomem`, `/proc/modules`, `/proc/scsi/scsi`를 제공합니다. Sound 문제에는 `alsa-info.sh` 같은 subsystem 수집 도구가 있습니다.

필요한 추가 자료는 문제마다 다릅니다. 빠진 것이 있으면 개발자가 요청하겠지만 처음부터 중요한 자료를 제공하면 자세히 검토받을 가능성이 높아집니다.

보고서 영역내용
제목가장 짧고 구체적인 문제 식별자; regression은 `[REGRESSION]`으로 시작
첫 문장보고서가 무엇에 관한 것인지 한 문장으로 요약
첫 문단영향과 핵심 상황을 보통 길이의 문단으로 설명
상세 설명재현 절차, 환경, kernel 버전, 로그와 첨부 자료를 제공

상세 설명 앞에 `The detailed description:` 같은 표식을 두고 그 위에 핵심과 영향을 설명하는 보통 길이 문단을 씁니다. 그 위에는 한 문장 요약, 가장 위에는 더 짧은 subject/title을 둡니다. 많은 사람이 이 앞부분만 보고 전체를 읽을지 결정하므로 충분히 다듬습니다.

Severe issue는 subject 또는 ticket title과 첫 문단에서 심각성이 분명해야 합니다. Regression 보고의 subject는 `[REGRESSION]`으로 시작합니다.

Bisection에 성공했다면 regression을 도입한 변경 제목을 subject 두 번째 부분으로 쓰고 culprit commit id를 적습니다. 실패했다면 마지막 정상 버전과 처음 문제가 보인 가장 오래된 시험 버전을 명시합니다.

Mail 보고에는 `[email protected]`를 CC합니다. Web tracker에 filed했다면 report를 regressions list로 inline forward하고 maintainer와 subsystem list를 CC하며 ticket URL을 위에 적습니다. Attachment로 전달하지 않습니다.

Bisection에 성공했다면 culprit author와 commit message의 signed-off-by chain 전체를 recipient에 추가합니다.

Security issue는 공개가 다른 사용자에게 단기 위험을 만드는지 평가합니다. 위험이 없으면 일반 절차를 따릅니다. 위험이 있으면 mail 보고에서 공개 list를 CC하지 않고, tracker ticket은 `private` 또는 `security issue`로 지정합니다. 비공개 기능이 없으면 tracker를 쓰지 말고 maintainer에게 private mail을 보냅니다.

두 경우 모두 `MAINTAINERS`의 `security contact` section 주소에도 보냅니다. Tracker에 filed했다면 ticket link를 적은 짧은 note와 report text를 그 주소로 전달합니다. 자세한 내용은 `Documentation/process/security-bugs.rst`를 따릅니다.

보고 뒤의 의무, 응답, 시험과 reminder

1275-1463

이상적인 경우 개발자가 즉시 원인을 찾고 patch를 작성·시험해 mainline에 보내며 필요한 stable·longterm backport tag까지 붙입니다. 이때는 감사를 전하고 fix가 포함된 release로 옮기면 됩니다. 그러나 이런 경우는 드물고 보고가 나간 뒤 실제 작업이 시작됩니다.

Tracker 보고에는 항상 그곳에서 공개 reply합니다. Mail 보고에는 언제나 `Reply-all`을 사용하고 추가 자료도 Sent folder의 원 보고에 reply-all해 보냅니다. 그래야 공개 list와 새 참여자가 계속 논의에 포함되고 thread가 유지됩니다.

예외는 누군가 private 전송을 요구했거나 요청 자료에 공개하면 안 되는 민감 정보가 있을 때입니다. 후자라면 요청한 개발자에게 private로 보내되 ticket이나 공개 mail에 그렇게 보냈다고 기록합니다.

익숙하지 않은 test tool이나 patch 적용을 요청받으면 바로 설명을 요구하기 전에 웹 검색이나 익숙한 community에서 도움을 구해 스스로 방법을 조사합니다. 상황에 따라 직접 지침을 요청하는 것도 괜찮습니다.

응답에는 인내가 필요합니다. Maintainer는 세계 여러 시간대에 있고 merge window, 다른 업무, conference, 휴가 때문에 보통 1~5 business day 또는 그 이상 걸릴 수 있습니다. High priority issue는 예외로, 최대 1주 또는 긴급하면 2일 뒤 정중히 reminder를 보냅니다.

Maintainer 응답이 늦거나 regression 여부에 이견이 있으면 mailing list에서 공개적으로 우려를 제기하고 조언을 구합니다. 실패하면 상위 maintainer에게 올릴 수 있으며 WiFi driver라면 wireless maintainer가 해당합니다. 상위 담당자도 없거나 모두 실패한 드문 경우 Linus Torvalds에게 알리는 것이 적절할 수 있습니다.

새 mainline의 `rc1`이 나올 때마다 문제가 고쳐졌는지, 중요한 변화가 있는지 시험하고 기존 report thread에 결과를 적습니다. `rc3`, `rc5`, final에서도 가끔 다시 시험하되 관련 변화가 있거나 다른 reply를 보낼 때만 결과를 알리는 것이 좋습니다.

답변한 사람이 실제 maintainer나 해당 code 개발자인지 웹 검색으로 확인합니다. 공개 보고에는 선의로 돕지만 잘못된 방향을 제시할 사람도 참여할 수 있습니다. 담당자가 논의를 들었는지 파악해 나중 reminder 필요 여부도 판단합니다.

추가 data 요청은 관심이 유지되는 동안 가능한 빨리, 보통 며칠 안에 제공합니다. Diagnostic patch나 fix 시험도 신속하되 서두르지 말고 제대로 적용됐는지 확인합니다. Patch가 실제로 적용되지 않았는데 적용됐다고 생각하는 실수는 경험자에게도 생깁니다.

아무 실질적 반응이 없으면 2주, 가능하면 3주 기다린 뒤 정중한 reminder를 보냅니다. Mail은 최초 보고에 reply하고 첫 줄에서 추가로 필요한 일이 있는지 묻고 원문 전체를 아래에 인용합니다. 이런 경우에는 `TOFU`(Text Over, Fullquote Under)가 적절합니다.

Reminder 뒤 3주 더 기다려도 반응이 없으면 담당자를 잘못 골랐는지, 보고가 공격적이거나 혼란스러웠는지 다시 검토합니다. FLOSS 보고에 익숙한 사람 한두 명에게 보여주고 개선된 두 번째 보고를 보낸다면 첫 보고 link와 개선판임을 밝힙니다.

보고가 적절했다면 두 번째 reminder에서 반응이 없는 이유와 조언을 묻습니다. 새 kernel의 `rc1` 직후 재시험 결과와 함께 보내기 좋습니다. 일주일 뒤에도 반응이 없으면 상위 maintainer에게 조언을 요청합니다.

Maintainer는 모든 보고에 응답하는 것이 이상적이지만 반드시 고쳐야 하는 것은 앞의 high priority issue입니다. 우선순위 때문에 당분간 처리할 시간이 없다는 답을 받을 수 있고, 논의와 reminder 뒤에도 fix가 나오지 않을 수 있습니다.

Linux kernel은 FLOSS이므로 영향받는 사용자를 찾아 협력하고, 몇 명이 왜 문제를 중요하게 보는지 적은 새 보고를 만들며, root cause나 regression 도입 변경을 함께 좁힐 수 있습니다. 팀에 programming 가능한 사람이 있다면 직접 fix를 작성할 수도 있습니다.

Stable·longterm 내부 regression 세부 reference

1464-1565

대부분의 kernel version line은 유지보수 부담 때문에 약 3개월만 지원됩니다. 매년 하나 정도가 최소 2년, 흔히 6년 지원됩니다. `kernel.org` 첫 화면에서 관심 계열이 최신 release로 표시되고 `[EOL]`이 없는지 확인합니다.

Stable line이 두 개 표시되면 오래된 쪽은 곧 지원이 끝날 수 있으므로 새 계열로 이동하는 편을 고려합니다. EOL line은 1~2주 더 화면에 남아도 시험과 보고에는 부적합합니다.

`https://lore.kernel.org/stable/`의 Linux stable mailing list archive에서 같은 문제를 검색하고, 이미 fix가 완료돼 곧 적용될 예정이 아니라면 기존 논의에 참여합니다.

관심 계열의 최신 release를 vanilla로 설치하고 문제가 생기기 전 taint되지 않았으며 여전히 재현되는지 확인합니다. Vendor kernel에서 처음 발견했다면 마지막 정상 버전의 vanilla build도 정상인지 검사합니다.

예를 들어 `5.10.4-vendor.42`에서 `5.10.5-vendor.43`으로 바꿔 깨졌다면 최신 5.10을 시험한 뒤 vanilla Linux 5.10.4도 정상인지 확인합니다. 거기서도 깨지면 upstream regression이 아니므로 일반 단계별 안내로 돌아갑니다.

Stable·longterm 내부 regression은 빠른 보고를 위해 짧은 설명으로 시작해도 됩니다. `[email protected]`와 `[email protected]`로 보내고 특정 subsystem이 의심되면 maintainer와 list도 CC합니다.

정확한 도입 버전을 알면 큰 도움이 됩니다. Distribution이 5.10.5에서 5.10.8로 갱신한 뒤 깨졌다면 최신 5.10.9와 vanilla 5.10.5를 확인하고 5.10.7, 결과에 따라 5.10.8 또는 5.10.6을 시험해 첫 문제 버전을 찾습니다. 보고에는 첫 broken 버전과 5.10.9도 여전히 broken임을 적습니다.

이는 거친 수동 bisection입니다. 가능하면 `Documentation/admin-guide/bug-bisect.rst`에 따라 정확한 culprit를 찾습니다. 성공하면 culprit author를 recipient에 추가하고 commit message 끝의 signed-off-by chain 전체를 CC합니다.

오래된 version line 전용 세부 reference

1566-1676

Mainline에서는 재현되지 않지만 stable·longterm 같은 오래된 line에서 고치고 싶을 때 사용합니다. 작은 변경도 예상 밖 문제를 만들 수 있어 maintainer는 `Documentation/process/stable-kernel-rules.rst` 규칙에 맞는 변경만 적용합니다.

복잡하거나 위험한 변경은 mainline에만 들어갑니다. 어떤 fix는 최신 stable·longterm에는 쉽게 backport돼도 더 오래된 line에는 위험할 수 있습니다. Backport가 불가능하면 문제를 감수하거나 새 Linux로 이동하거나 직접 patch해야 합니다.

먼저 해당 line이 지원되는지 확인하고 stable list의 기존 보고를 검색하며 그 계열 최신 release에서도 문제를 검사합니다.

Mainline에서 문제를 고친 commit과 논의를 찾습니다. `kernel.org` Git web interface나 `https://github.com/torvalds/linux` mirror를 쓰거나 local clone에서 `git log --grep=<pattern>`을 실행합니다.

Fix commit message 끝에 다음과 같은 stable tag가 있는지 확인합니다.

Cc: <[email protected]> # 5.4+

이 예는 fix가 5.4 이상 계열에 안전하게 backport될 수 있다고 표시합니다. 대부분 2주 안에 적용되지만 더 오래 걸릴 수도 있습니다.

Commit에 정보가 없거나 fix를 찾지 못하면 웹, LKML archive, subsystem tracker와 mailing list에서 다시 논의를 찾습니다. 제안된 fix가 있으면 version control system에서 commit을 찾아 backport 예정 여부를 확인합니다.

논의에서 fix가 관심 line에 너무 위험하다고 판단됐으면 문제를 감수하거나 fix가 적용된 새 line으로 이동합니다. Stable tag도 없고 backport 논의도 없다면 최신 논의에 참여해 해당 버전에서 문제를 겪고 있으며 적합하다면 수정되길 원한다고 말합니다.

그래도 해결되지 않으면 원인 subsystem maintainer에게 mail로 조언을 요청하고 해당 subsystem mailing list와 `[email protected]`를 CC합니다.

일부 문제가 반응이나 fix를 얻지 못하는 이유

1677-1733

Regression, security issue, severe problem 같은 high priority issue는 해결 대상이며 maintainer나 필요하면 Linus Torvalds가 챙깁니다. 다른 많은 문제도 고쳐지지만 개발자가 도울 수 없거나 돕지 않거나 보고할 사람 자체가 없을 수 있습니다.

많은 driver는 자기 hardware를 Linux에서 쓰려고 여가 시간에 개발한 사람이 작성했습니다. 보통 다른 사용자의 문제도 기꺼이 고치지만 자발적 기여이므로 강제할 수 없습니다.

고치고 싶어도 hardware programming 문서가 부족해 불가능할 수 있습니다. 공개 문서가 피상적이거나 reverse engineering으로 작성된 driver에서 흔합니다.

시간이 지나 test hardware가 고장 나거나 교체되고, hardware가 너무 오래되거나 개발자의 삶에서 다른 일이 중요해지면서 관심이 끝날 수 있습니다. 맡을 사람이 없어도 유용한 사용자가 있고 제거가 regression이 되므로 abandoned driver가 kernel에 남기도 합니다.

급여를 받고 일하는 개발자도 employer가 오래된 제품을 더 이상 중요하게 보지 않거나 다른 업무를 맡기면 유지보수가 줄어듭니다. Hardware vendor는 새 제품 판매가 중심이라 판매 중단 제품 driver에 투자하지 않을 수 있고 enterprise distribution도 새 version 범위를 줄이기 위해 오래되고 드문 hardware 지원을 뺄 수 있습니다.

유지보수 시간과 처리 가능한 보고 수가 한정되어 우선순위도 필요합니다. 거의 완벽히 동작하는 driver라도 보고가 너무 많으면 일부를 거부해야 전체 작업이 멈추지 않습니다. 그래도 많은 driver에는 가능한 많은 문제를 고치려는 active maintainer가 있습니다.

맺음말과 재배포 정보

1734-1764

다른 FLOSS에 비해 Linux kernel 개발자에게 문제를 보고하는 일은 어렵습니다. 이 문서의 길이와 복잡성이 그 현실을 보여줍니다. 주 작성자는 현재 절차를 기록하는 일이 시간이 지나 상황을 개선할 기반이 되기를 바랍니다.

이 문서는 Thorsten Leemhuis `<[email protected]>`가 유지합니다. 오타나 작은 실수는 직접 알려도 됩니다. 비공식적인 방식으로 text 변경에 기여할 수도 있지만 copyright를 위해 `[email protected]`를 CC하고 `Documentation/process/submitting-patches.rst`의 `Sign your work - the Developer's Certificate of Origin` 절에 따라 sign-off해야 합니다.

Text는 파일 상단에 적힌 대로 `GPL-2.0+` 또는 `CC-BY-4.0`으로 제공됩니다. `CC-BY-4.0`만으로 배포하려면 author attribution을 `The Linux kernel developers`로 쓰고 `https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/plain/Documentation/admin-guide/reporting-issues.rst`를 source로 link합니다.

`CC-BY-4.0`은 Linux kernel source에 있는 이 RST 파일 내용에만 적용됩니다. Kernel build system 등으로 처리된 version은 더 제한적인 license의 다른 파일 내용을 포함할 수 있습니다.