← Documents Documentation/networking/ip-sysctl.rst GitHub 원문 ↗

Linux 6.18.37 · Networking

IP Sysctl

IPv4·IPv6와 TCP·UDP·ICMP·SCTP 등 Linux IP stack의 `/proc/sys/net` 조정 변수, 기본값, 상호작용과 운영상 주의점을 설명합니다.

Source pathDocumentation/networking/ip-sysctl.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

ip-sysctl.rst:1-3731

이 문서는 Linux network stack의 동작을 runtime에 바꾸는 sysctl의 의미와 기본값을 protocol별로 정리합니다. 같은 이름이라도 global, `conf/all`, `conf/default`, `conf/interface`의 결합 규칙이 다르고, forwarding처럼 값을 바꾸며 다른 기본값을 초기화하는 항목도 있어 설정 단위를 함께 읽어야 합니다.

설정 영역
영역대표 설정주요 목적
IPv4 routingip_forward, FIB multipath, rp_filter전달, ECMP, source 검증
TCPECN, RACK, buffer, TFO, pacing연결 복구와 성능·memory 제어
Local transportUDP/RAW, port range, early demuxsocket 분배와 namespace 격리
Control protocolICMP, IGMP, MLD응답 제한과 multicast membership
IPv6RA, ND, DAD, privacy addresshost/router 동작과 자동 구성
Link integrationARP, bridge netfilter, L3 masterneighbor와 VRF/bridge 처리
SCTPfailover, RTO, buffer, encapsulationmultihoming과 extension 관리

긴 sysctl 목록을 운영 목적에 따라 분류했습니다.

Interface 설정 적용
conf/default/* 변경이후 생성되는 interface의 초기값
conf/interface/* 변경해당 interface 동작
conf/all/* 변경문서가 정의한 AND, OR 또는 max 규칙실제 interface 동작

새 interface 기본값과 현재 interface 설정의 역할을 구분해야 합니다.

성능 관련 숫자는 무조건 크게 하는 값이 아닙니다. queue와 memory 상한 증가는 burst 흡수와 동시 연결 수를 늘리지만 memory pressure와 지연을 키울 수 있고, PMTU·ECN·reordering·SYN 보호 설정은 peer와 middlebox의 실제 동작을 전제로 합니다. 변경 전 namespace 범위, socket option override, kernel build option, 기본값과 폐기 여부를 확인해야 합니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 =========
4 IP Sysctl
5 =========
6
7 /proc/sys/net/ipv4/* Variables
8 ==============================
9
10 ip_forward - BOOLEAN
11 Forward Packets between interfaces.
12
13 This variable is special, its change resets all configuration
14 parameters to their default state (RFC1122 for hosts, RFC1812
15 for routers)
16
17 Possible values:
18
19 - 0 (disabled)
20 - 1 (enabled)
21
22 Default: 0 (disabled)
23
24 ip_default_ttl - INTEGER
25 Default value of TTL field (Time To Live) for outgoing (but not
26 forwarded) IP packets. Should be between 1 and 255 inclusive.
27 Default: 64 (as recommended by RFC1700)
28
29 ip_no_pmtu_disc - INTEGER
30 Disable Path MTU Discovery. If enabled in mode 1 and a
31 fragmentation-required ICMP is received, the PMTU to this
32 destination will be set to the smallest of the old MTU to
33 this destination and min_pmtu (see below). You will need
34 to raise min_pmtu to the smallest interface MTU on your system
35 manually if you want to avoid locally generated fragments.
36
37 In mode 2 incoming Path MTU Discovery messages will be
38 discarded. Outgoing frames are handled the same as in mode 1,
39 implicitly setting IP_PMTUDISC_DONT on every created socket.
40
41 Mode 3 is a hardened pmtu discover mode. The kernel will only
42 accept fragmentation-needed errors if the underlying protocol
43 can verify them besides a plain socket lookup. Current
44 protocols for which pmtu events will be honored are TCP and
45 SCTP as they verify e.g. the sequence number or the
46 association. This mode should not be enabled globally but is
47 only intended to secure e.g. name servers in namespaces where
48 TCP path mtu must still work but path MTU information of other
49 protocols should be discarded. If enabled globally this mode
50 could break other protocols.
51
52 Possible values: 0-3
53
54 Default: FALSE
55
56 min_pmtu - INTEGER
57 default 552 - minimum Path MTU. Unless this is changed manually,
58 each cached pmtu will never be lower than this setting.
59
60 ip_forward_use_pmtu - BOOLEAN
61 By default we don't trust protocol path MTUs while forwarding
62 because they could be easily forged and can lead to unwanted
63 fragmentation by the router.
64 You only need to enable this if you have user-space software
65 which tries to discover path mtus by itself and depends on the
66 kernel honoring this information. This is normally not the
67 case.
68
69 Possible values:
70
71 - 0 (disabled)
72 - 1 (enabled)
73
74 Default: 0 (disabled)
75
76 fwmark_reflect - BOOLEAN
77 Controls the fwmark of kernel-generated IPv4 reply packets that are not
78 associated with a socket for example, TCP RSTs or ICMP echo replies).
79 If disabled, these packets have a fwmark of zero. If enabled, they have the
80 fwmark of the packet they are replying to.
81
82 Possible values:
83
84 - 0 (disabled)
85 - 1 (enabled)
86
87 Default: 0 (disabled)
88
89 fib_multipath_use_neigh - BOOLEAN
90 Use status of existing neighbor entry when determining nexthop for
91 multipath routes. If disabled, neighbor information is not used and
92 packets could be directed to a failed nexthop. Only valid for kernels
93 built with CONFIG_IP_ROUTE_MULTIPATH enabled.
94
95 Possible values:
96
97 - 0 (disabled)
98 - 1 (enabled)
99
100 Default: 0 (disabled)
101
102 fib_multipath_hash_policy - INTEGER
103 Controls which hash policy to use for multipath routes. Only valid
104 for kernels built with CONFIG_IP_ROUTE_MULTIPATH enabled.
105
106 Default: 0 (Layer 3)
107
108 Possible values:
109
110 - 0 - Layer 3
111 - 1 - Layer 4
112 - 2 - Layer 3 or inner Layer 3 if present
113 - 3 - Custom multipath hash. Fields used for multipath hash calculation
114 are determined by fib_multipath_hash_fields sysctl
115
116 fib_multipath_hash_fields - UNSIGNED INTEGER
117 When fib_multipath_hash_policy is set to 3 (custom multipath hash), the
118 fields used for multipath hash calculation are determined by this
119 sysctl.
120
121 This value is a bitmask which enables various fields for multipath hash
122 calculation.
123
124 Possible fields are:
125
126 ====== ============================
127 0x0001 Source IP address
128 0x0002 Destination IP address
129 0x0004 IP protocol
130 0x0008 Unused (Flow Label)
131 0x0010 Source port
132 0x0020 Destination port
133 0x0040 Inner source IP address
134 0x0080 Inner destination IP address
135 0x0100 Inner IP protocol
136 0x0200 Inner Flow Label
137 0x0400 Inner source port
138 0x0800 Inner destination port
139 ====== ============================
140
141 Default: 0x0007 (source IP, destination IP and IP protocol)
142
143 fib_multipath_hash_seed - UNSIGNED INTEGER
144 The seed value used when calculating hash for multipath routes. Applies
145 to both IPv4 and IPv6 datapath. Only present for kernels built with
146 CONFIG_IP_ROUTE_MULTIPATH enabled.
147
148 When set to 0, the seed value used for multipath routing defaults to an
149 internal random-generated one.
150
151 The actual hashing algorithm is not specified -- there is no guarantee
152 that a next hop distribution effected by a given seed will keep stable
153 across kernel versions.
154
155 Default: 0 (random)
156
157 fib_sync_mem - UNSIGNED INTEGER
158 Amount of dirty memory from fib entries that can be backlogged before
159 synchronize_rcu is forced.
160
161 Default: 512kB Minimum: 64kB Maximum: 64MB
162
163 ip_forward_update_priority - INTEGER
164 Whether to update SKB priority from "TOS" field in IPv4 header after it
165 is forwarded. The new SKB priority is mapped from TOS field value
166 according to an rt_tos2priority table (see e.g. man tc-prio).
167
168 Default: 1 (Update priority.)
169
170 Possible values:
171
172 - 0 - Do not update priority.
173 - 1 - Update priority.
174
175 route/max_size - INTEGER
176 Maximum number of routes allowed in the kernel. Increase
177 this when using large numbers of interfaces and/or routes.
178
179 From linux kernel 3.6 onwards, this is deprecated for ipv4
180 as route cache is no longer used.
181
182 From linux kernel 6.3 onwards, this is deprecated for ipv6
183 as garbage collection manages cached route entries.
184
185 neigh/default/gc_thresh1 - INTEGER
186 Minimum number of entries to keep. Garbage collector will not
187 purge entries if there are fewer than this number.
188
189 Default: 128
190
191 neigh/default/gc_thresh2 - INTEGER
192 Threshold when garbage collector becomes more aggressive about
193 purging entries. Entries older than 5 seconds will be cleared
194 when over this number.
195
196 Default: 512
197
198 neigh/default/gc_thresh3 - INTEGER
199 Maximum number of non-PERMANENT neighbor entries allowed. Increase
200 this when using large numbers of interfaces and when communicating
201 with large numbers of directly-connected peers.
202
203 Default: 1024
204
205 neigh/default/unres_qlen_bytes - INTEGER
206 The maximum number of bytes which may be used by packets
207 queued for each unresolved address by other network layers.
208 (added in linux 3.3)
209
210 Setting negative value is meaningless and will return error.
211
212 Default: SK_WMEM_DEFAULT, (same as net.core.wmem_default).
213
214 Exact value depends on architecture and kernel options,
215 but should be enough to allow queuing 256 packets
216 of medium size.
217
218 neigh/default/unres_qlen - INTEGER
219 The maximum number of packets which may be queued for each
220 unresolved address by other network layers.
221
222 (deprecated in linux 3.3) : use unres_qlen_bytes instead.
223
224 Prior to linux 3.3, the default value is 3 which may cause
225 unexpected packet loss. The current default value is calculated
226 according to default value of unres_qlen_bytes and true size of
227 packet.
228
229 Default: 101
230
231 neigh/default/interval_probe_time_ms - INTEGER
232 The probe interval for neighbor entries with NTF_MANAGED flag,
233 the min value is 1.
234
235 Default: 5000
236
237 mtu_expires - INTEGER
238 Time, in seconds, that cached PMTU information is kept.
239
240 min_adv_mss - INTEGER
241 The advertised MSS depends on the first hop route MTU, but will
242 never be lower than this setting.
243
244 fib_notify_on_flag_change - INTEGER
245 Whether to emit RTM_NEWROUTE notifications whenever RTM_F_OFFLOAD/
246 RTM_F_TRAP/RTM_F_OFFLOAD_FAILED flags are changed.
247
248 After installing a route to the kernel, user space receives an
249 acknowledgment, which means the route was installed in the kernel,
250 but not necessarily in hardware.
251 It is also possible for a route already installed in hardware to change
252 its action and therefore its flags. For example, a host route that is
253 trapping packets can be "promoted" to perform decapsulation following
254 the installation of an IPinIP/VXLAN tunnel.
255 The notifications will indicate to user-space the state of the route.
256
257 Default: 0 (Do not emit notifications.)
258
259 Possible values:
260
261 - 0 - Do not emit notifications.
262 - 1 - Emit notifications.
263 - 2 - Emit notifications only for RTM_F_OFFLOAD_FAILED flag change.
264
265 IP Fragmentation:
266
267 ipfrag_high_thresh - LONG INTEGER
268 Maximum memory used to reassemble IP fragments.
269
270 ipfrag_low_thresh - LONG INTEGER
271 (Obsolete since linux-4.17)
272 Maximum memory used to reassemble IP fragments before the kernel
273 begins to remove incomplete fragment queues to free up resources.
274 The kernel still accepts new fragments for defragmentation.
275
276 ipfrag_time - INTEGER
277 Time in seconds to keep an IP fragment in memory.
278
279 ipfrag_max_dist - INTEGER
280 ipfrag_max_dist is a non-negative integer value which defines the
281 maximum "disorder" which is allowed among fragments which share a
282 common IP source address. Note that reordering of packets is
283 not unusual, but if a large number of fragments arrive from a source
284 IP address while a particular fragment queue remains incomplete, it
285 probably indicates that one or more fragments belonging to that queue
286 have been lost. When ipfrag_max_dist is positive, an additional check
287 is done on fragments before they are added to a reassembly queue - if
288 ipfrag_max_dist (or more) fragments have arrived from a particular IP
289 address between additions to any IP fragment queue using that source
290 address, it's presumed that one or more fragments in the queue are
291 lost. The existing fragment queue will be dropped, and a new one
292 started. An ipfrag_max_dist value of zero disables this check.
293
294 Using a very small value, e.g. 1 or 2, for ipfrag_max_dist can
295 result in unnecessarily dropping fragment queues when normal
296 reordering of packets occurs, which could lead to poor application
297 performance. Using a very large value, e.g. 50000, increases the
298 likelihood of incorrectly reassembling IP fragments that originate
299 from different IP datagrams, which could result in data corruption.
300 Default: 64
301
302 bc_forwarding - INTEGER
303 bc_forwarding enables the feature described in rfc1812#section-5.3.5.2
304 and rfc2644. It allows the router to forward directed broadcast.
305 To enable this feature, the 'all' entry and the input interface entry
306 should be set to 1.
307 Default: 0
308
309 INET peer storage
310 =================
311
312 inet_peer_threshold - INTEGER
313 The approximate size of the storage. Starting from this threshold
314 entries will be thrown aggressively. This threshold also determines
315 entries' time-to-live and time intervals between garbage collection
316 passes. More entries, less time-to-live, less GC interval.
317
318 inet_peer_minttl - INTEGER
319 Minimum time-to-live of entries. Should be enough to cover fragment
320 time-to-live on the reassembling side. This minimum time-to-live is
321 guaranteed if the pool size is less than inet_peer_threshold.
322 Measured in seconds.
323
324 inet_peer_maxttl - INTEGER
325 Maximum time-to-live of entries. Unused entries will expire after
326 this period of time if there is no memory pressure on the pool (i.e.
327 when the number of entries in the pool is very small).
328 Measured in seconds.
329
330 TCP variables
331 =============
332
333 somaxconn - INTEGER
334 Limit of socket listen() backlog, known in userspace as SOMAXCONN.
335 Defaults to 4096. (Was 128 before linux-5.4)
336 See also tcp_max_syn_backlog for additional tuning for TCP sockets.
337
338 tcp_abort_on_overflow - BOOLEAN
339 If listening service is too slow to accept new connections,
340 reset them. Default state is FALSE. It means that if overflow
341 occurred due to a burst, connection will recover. Enable this
342 option _only_ if you are really sure that listening daemon
343 cannot be tuned to accept connections faster. Enabling this
344 option can harm clients of your server.
345
346 tcp_adv_win_scale - INTEGER
347 Obsolete since linux-6.6
348 Count buffering overhead as bytes/2^tcp_adv_win_scale
349 (if tcp_adv_win_scale > 0) or bytes-bytes/2^(-tcp_adv_win_scale),
350 if it is <= 0.
351
352 Possible values are [-31, 31], inclusive.
353
354 Default: 1
355
356 tcp_allowed_congestion_control - STRING
357 Show/set the congestion control choices available to non-privileged
358 processes. The list is a subset of those listed in
359 tcp_available_congestion_control.
360
361 Default is "reno" and the default setting (tcp_congestion_control).
362
363 tcp_app_win - INTEGER
364 Reserve max(window/2^tcp_app_win, mss) of window for application
365 buffer. Value 0 is special, it means that nothing is reserved.
366
367 Possible values are [0, 31], inclusive.
368
369 Default: 31
370
371 tcp_autocorking - BOOLEAN
372 Enable TCP auto corking :
373 When applications do consecutive small write()/sendmsg() system calls,
374 we try to coalesce these small writes as much as possible, to lower
375 total amount of sent packets. This is done if at least one prior
376 packet for the flow is waiting in Qdisc queues or device transmit
377 queue. Applications can still use TCP_CORK for optimal behavior
378 when they know how/when to uncork their sockets.
379
380 Possible values:
381
382 - 0 (disabled)
383 - 1 (enabled)
384
385 Default: 1 (enabled)
386
387 tcp_available_congestion_control - STRING
388 Shows the available congestion control choices that are registered.
389 More congestion control algorithms may be available as modules,
390 but not loaded.
391
392 tcp_base_mss - INTEGER
393 The initial value of search_low to be used by the packetization layer
394 Path MTU discovery (MTU probing). If MTU probing is enabled,
395 this is the initial MSS used by the connection.
396
397 tcp_mtu_probe_floor - INTEGER
398 If MTU probing is enabled this caps the minimum MSS used for search_low
399 for the connection.
400
401 Default : 48
402
403 tcp_min_snd_mss - INTEGER
404 TCP SYN and SYNACK messages usually advertise an ADVMSS option,
405 as described in RFC 1122 and RFC 6691.
406
407 If this ADVMSS option is smaller than tcp_min_snd_mss,
408 it is silently capped to tcp_min_snd_mss.
409
410 Default : 48 (at least 8 bytes of payload per segment)
411
412 tcp_congestion_control - STRING
413 Set the congestion control algorithm to be used for new
414 connections. The algorithm "reno" is always available, but
415 additional choices may be available based on kernel configuration.
416 Default is set as part of kernel configuration.
417 For passive connections, the listener congestion control choice
418 is inherited.
419
420 [see setsockopt(listenfd, SOL_TCP, TCP_CONGESTION, "name" ...) ]
421
422 tcp_dsack - BOOLEAN
423 Allows TCP to send "duplicate" SACKs.
424
425 Possible values:
426
427 - 0 (disabled)
428 - 1 (enabled)
429
430 Default: 1 (enabled)
431
432 tcp_early_retrans - INTEGER
433 Tail loss probe (TLP) converts RTOs occurring due to tail
434 losses into fast recovery (RFC8985). Note that
435 TLP requires RACK to function properly (see tcp_recovery below)
436
437 Possible values:
438
439 - 0 disables TLP
440 - 3 or 4 enables TLP
441
442 Default: 3
443
444 tcp_ecn - INTEGER
445 Control use of Explicit Congestion Notification (ECN) by TCP.
446 ECN is used only when both ends of the TCP connection indicate support
447 for it. This feature is useful in avoiding losses due to congestion by
448 allowing supporting routers to signal congestion before having to drop
449 packets. A host that supports ECN both sends ECN at the IP layer and
450 feeds back ECN at the TCP layer. The highest variant of ECN feedback
451 that both peers support is chosen by the ECN negotiation (Accurate ECN,
452 ECN, or no ECN).
453
454 The highest negotiated variant for incoming connection requests
455 and the highest variant requested by outgoing connection
456 attempts:
457
458 ===== ==================== ====================
459 Value Incoming connections Outgoing connections
460 ===== ==================== ====================
461 0 No ECN No ECN
462 1 ECN ECN
463 2 ECN No ECN
464 3 AccECN AccECN
465 4 AccECN ECN
466 5 AccECN No ECN
467 ===== ==================== ====================
468
469 Default: 2
470
471 tcp_ecn_option - INTEGER
472 Control Accurate ECN (AccECN) option sending when AccECN has been
473 successfully negotiated during handshake. Send logic inhibits
474 sending AccECN options regarless of this setting when no AccECN
475 option has been seen for the reverse direction.
476
477 Possible values are:
478
479 = ============================================================
480 0 Never send AccECN option. This also disables sending AccECN
481 option in SYN/ACK during handshake.
482 1 Send AccECN option sparingly according to the minimum option
483 rules outlined in draft-ietf-tcpm-accurate-ecn.
484 2 Send AccECN option on every packet whenever it fits into TCP
485 option space.
486 = ============================================================
487
488 Default: 2
489
490 tcp_ecn_option_beacon - INTEGER
491 Control Accurate ECN (AccECN) option sending frequency per RTT and it
492 takes effect only when tcp_ecn_option is set to 2.
493
494 Default: 3 (AccECN will be send at least 3 times per RTT)
495
496 tcp_ecn_fallback - BOOLEAN
497 If the kernel detects that ECN connection misbehaves, enable fall
498 back to non-ECN. Currently, this knob implements the fallback
499 from RFC3168, section 6.1.1.1., but we reserve that in future,
500 additional detection mechanisms could be implemented under this
501 knob. The value is not used, if tcp_ecn or per route (or congestion
502 control) ECN settings are disabled.
503
504 Possible values:
505
506 - 0 (disabled)
507 - 1 (enabled)
508
509 Default: 1 (enabled)
510
511 tcp_fack - BOOLEAN
512 This is a legacy option, it has no effect anymore.
513
514 tcp_fin_timeout - INTEGER
515 The length of time an orphaned (no longer referenced by any
516 application) connection will remain in the FIN_WAIT_2 state
517 before it is aborted at the local end. While a perfectly
518 valid "receive only" state for an un-orphaned connection, an
519 orphaned connection in FIN_WAIT_2 state could otherwise wait
520 forever for the remote to close its end of the connection.
521
522 Cf. tcp_max_orphans
523
524 Default: 60 seconds
525
526 tcp_frto - INTEGER
527 Enables Forward RTO-Recovery (F-RTO) defined in RFC5682.
528 F-RTO is an enhanced recovery algorithm for TCP retransmission
529 timeouts. It is particularly beneficial in networks where the
530 RTT fluctuates (e.g., wireless). F-RTO is sender-side only
531 modification. It does not require any support from the peer.
532
533 By default it's enabled with a non-zero value. 0 disables F-RTO.
534
535 tcp_fwmark_accept - BOOLEAN
536 If enabled, incoming connections to listening sockets that do not have a
537 socket mark will set the mark of the accepting socket to the fwmark of
538 the incoming SYN packet. This will cause all packets on that connection
539 (starting from the first SYNACK) to be sent with that fwmark. The
540 listening socket's mark is unchanged. Listening sockets that already
541 have a fwmark set via setsockopt(SOL_SOCKET, SO_MARK, ...) are
542 unaffected.
543
544 Possible values:
545
546 - 0 (disabled)
547 - 1 (enabled)
548
549 Default: 0 (disabled)
550
551 tcp_invalid_ratelimit - INTEGER
552 Limit the maximal rate for sending duplicate acknowledgments
553 in response to incoming TCP packets that are for an existing
554 connection but that are invalid due to any of these reasons:
555
556 (a) out-of-window sequence number,
557 (b) out-of-window acknowledgment number, or
558 (c) PAWS (Protection Against Wrapped Sequence numbers) check failure
559
560 This can help mitigate simple "ack loop" DoS attacks, wherein
561 a buggy or malicious middlebox or man-in-the-middle can
562 rewrite TCP header fields in manner that causes each endpoint
563 to think that the other is sending invalid TCP segments, thus
564 causing each side to send an unterminating stream of duplicate
565 acknowledgments for invalid segments.
566
567 Using 0 disables rate-limiting of dupacks in response to
568 invalid segments; otherwise this value specifies the minimal
569 space between sending such dupacks, in milliseconds.
570
571 Default: 500 (milliseconds).
572
573 tcp_keepalive_time - INTEGER
574 How often TCP sends out keepalive messages when keepalive is enabled.
575 Default: 2hours.
576
577 tcp_keepalive_probes - INTEGER
578 How many keepalive probes TCP sends out, until it decides that the
579 connection is broken. Default value: 9.
580
581 tcp_keepalive_intvl - INTEGER
582 How frequently the probes are send out. Multiplied by
583 tcp_keepalive_probes it is time to kill not responding connection,
584 after probes started. Default value: 75sec i.e. connection
585 will be aborted after ~11 minutes of retries.
586
587 tcp_l3mdev_accept - BOOLEAN
588 Enables child sockets to inherit the L3 master device index.
589 Enabling this option allows a "global" listen socket to work
590 across L3 master domains (e.g., VRFs) with connected sockets
591 derived from the listen socket to be bound to the L3 domain in
592 which the packets originated. Only valid when the kernel was
593 compiled with CONFIG_NET_L3_MASTER_DEV.
594
595 Possible values:
596
597 - 0 (disabled)
598 - 1 (enabled)
599
600 Default: 0 (disabled)
601
602 tcp_low_latency - BOOLEAN
603 This is a legacy option, it has no effect anymore.
604
605 tcp_max_orphans - INTEGER
606 Maximal number of TCP sockets not attached to any user file handle,
607 held by system. If this number is exceeded orphaned connections are
608 reset immediately and warning is printed. This limit exists
609 only to prevent simple DoS attacks, you _must_ not rely on this
610 or lower the limit artificially, but rather increase it
611 (probably, after increasing installed memory),
612 if network conditions require more than default value,
613 and tune network services to linger and kill such states
614 more aggressively. Let me to remind again: each orphan eats
615 up to ~64K of unswappable memory.
616
617 tcp_max_syn_backlog - INTEGER
618 Maximal number of remembered connection requests (SYN_RECV),
619 which have not received an acknowledgment from connecting client.
620
621 This is a per-listener limit.
622
623 The minimal value is 128 for low memory machines, and it will
624 increase in proportion to the memory of machine.
625
626 If server suffers from overload, try increasing this number.
627
628 Remember to also check /proc/sys/net/core/somaxconn
629 A SYN_RECV request socket consumes about 304 bytes of memory.
630
631 tcp_max_tw_buckets - INTEGER
632 Maximal number of timewait sockets held by system simultaneously.
633 If this number is exceeded time-wait socket is immediately destroyed
634 and warning is printed. This limit exists only to prevent
635 simple DoS attacks, you _must_ not lower the limit artificially,
636 but rather increase it (probably, after increasing installed memory),
637 if network conditions require more than default value.
638
639 tcp_mem - vector of 3 INTEGERs: min, pressure, max
640 min: below this number of pages TCP is not bothered about its
641 memory appetite.
642
643 pressure: when amount of memory allocated by TCP exceeds this number
644 of pages, TCP moderates its memory consumption and enters memory
645 pressure mode, which is exited when memory consumption falls
646 under "min".
647
648 max: number of pages allowed for queueing by all TCP sockets.
649
650 Defaults are calculated at boot time from amount of available
651 memory.
652
653 tcp_min_rtt_wlen - INTEGER
654 The window length of the windowed min filter to track the minimum RTT.
655 A shorter window lets a flow more quickly pick up new (higher)
656 minimum RTT when it is moved to a longer path (e.g., due to traffic
657 engineering). A longer window makes the filter more resistant to RTT
658 inflations such as transient congestion. The unit is seconds.
659
660 Possible values: 0 - 86400 (1 day)
661
662 Default: 300
663
664 tcp_moderate_rcvbuf - BOOLEAN
665 If enabled, TCP performs receive buffer auto-tuning, attempting to
666 automatically size the buffer (no greater than tcp_rmem[2]) to
667 match the size required by the path for full throughput.
668
669 Possible values:
670
671 - 0 (disabled)
672 - 1 (enabled)
673
674 Default: 1 (enabled)
675
676 tcp_mtu_probing - INTEGER
677 Controls TCP Packetization-Layer Path MTU Discovery. Takes three
678 values:
679
680 - 0 - Disabled
681 - 1 - Disabled by default, enabled when an ICMP black hole detected
682 - 2 - Always enabled, use initial MSS of tcp_base_mss.
683
684 tcp_probe_interval - UNSIGNED INTEGER
685 Controls how often to start TCP Packetization-Layer Path MTU
686 Discovery reprobe. The default is reprobing every 10 minutes as
687 per RFC4821.
688
689 tcp_probe_threshold - INTEGER
690 Controls when TCP Packetization-Layer Path MTU Discovery probing
691 will stop in respect to the width of search range in bytes. Default
692 is 8 bytes.
693
694 tcp_no_metrics_save - BOOLEAN
695 By default, TCP saves various connection metrics in the route cache
696 when the connection closes, so that connections established in the
697 near future can use these to set initial conditions. Usually, this
698 increases overall performance, but may sometimes cause performance
699 degradation. If enabled, TCP will not cache metrics on closing
700 connections.
701
702 Possible values:
703
704 - 0 (disabled)
705 - 1 (enabled)
706
707 Default: 0 (disabled)
708
709 tcp_no_ssthresh_metrics_save - BOOLEAN
710 Controls whether TCP saves ssthresh metrics in the route cache.
711 If enabled, ssthresh metrics are disabled.
712
713 Possible values:
714
715 - 0 (disabled)
716 - 1 (enabled)
717
718 Default: 1 (enabled)
719
720 tcp_orphan_retries - INTEGER
721 This value influences the timeout of a locally closed TCP connection,
722 when RTO retransmissions remain unacknowledged.
723 See tcp_retries2 for more details.
724
725 The default value is 8.
726
727 If your machine is a loaded WEB server,
728 you should think about lowering this value, such sockets
729 may consume significant resources. Cf. tcp_max_orphans.
730
731 tcp_recovery - INTEGER
732 This value is a bitmap to enable various experimental loss recovery
733 features.
734
735 ========= =============================================================
736 RACK: 0x1 enables RACK loss detection, for fast detection of lost
737 retransmissions and tail drops, and resilience to
738 reordering. currently, setting this bit to 0 has no
739 effect, since RACK is the only supported loss detection
740 algorithm.
741
742 RACK: 0x2 makes RACK's reordering window static (min_rtt/4).
743
744 RACK: 0x4 disables RACK's DUPACK threshold heuristic
745 ========= =============================================================
746
747 Default: 0x1
748
749 tcp_reflect_tos - BOOLEAN
750 For listening sockets, reuse the DSCP value of the initial SYN message
751 for outgoing packets. This allows to have both directions of a TCP
752 stream to use the same DSCP value, assuming DSCP remains unchanged for
753 the lifetime of the connection.
754
755 This options affects both IPv4 and IPv6.
756
757 Possible values:
758
759 - 0 (disabled)
760 - 1 (enabled)
761
762 Default: 0 (disabled)
763
764 tcp_reordering - INTEGER
765 Initial reordering level of packets in a TCP stream.
766 TCP stack can then dynamically adjust flow reordering level
767 between this initial value and tcp_max_reordering
768
769 Default: 3
770
771 tcp_max_reordering - INTEGER
772 Maximal reordering level of packets in a TCP stream.
773 300 is a fairly conservative value, but you might increase it
774 if paths are using per packet load balancing (like bonding rr mode)
775
776 Default: 300
777
778 tcp_retrans_collapse - BOOLEAN
779 Bug-to-bug compatibility with some broken printers.
780 On retransmit try to send bigger packets to work around bugs in
781 certain TCP stacks.
782
783 Possible values:
784
785 - 0 (disabled)
786 - 1 (enabled)
787
788 Default: 1 (enabled)
789
790 tcp_retries1 - INTEGER
791 This value influences the time, after which TCP decides, that
792 something is wrong due to unacknowledged RTO retransmissions,
793 and reports this suspicion to the network layer.
794 See tcp_retries2 for more details.
795
796 RFC 1122 recommends at least 3 retransmissions, which is the
797 default.
798
799 tcp_retries2 - INTEGER
800 This value influences the timeout of an alive TCP connection,
801 when RTO retransmissions remain unacknowledged.
802 Given a value of N, a hypothetical TCP connection following
803 exponential backoff with an initial RTO of TCP_RTO_MIN would
804 retransmit N times before killing the connection at the (N+1)th RTO.
805
806 The default value of 15 yields a hypothetical timeout of 924.6
807 seconds and is a lower bound for the effective timeout.
808 TCP will effectively time out at the first RTO which exceeds the
809 hypothetical timeout.
810 If tcp_rto_max_ms is decreased, it is recommended to also
811 change tcp_retries2.
812
813 RFC 1122 recommends at least 100 seconds for the timeout,
814 which corresponds to a value of at least 8.
815
816 tcp_rfc1337 - BOOLEAN
817 If enabled, the TCP stack behaves conforming to RFC1337. If unset,
818 we are not conforming to RFC, but prevent TCP TIME_WAIT
819 assassination.
820
821 Possible values:
822
823 - 0 (disabled)
824 - 1 (enabled)
825
826 Default: 0 (disabled)
827
828 tcp_rmem - vector of 3 INTEGERs: min, default, max
829 min: Minimal size of receive buffer used by TCP sockets.
830 It is guaranteed to each TCP socket, even under moderate memory
831 pressure.
832
833 Default: 4K
834
835 default: initial size of receive buffer used by TCP sockets.
836 This value overrides net.core.rmem_default used by other protocols.
837 Default: 131072 bytes.
838 This value results in initial window of 65535.
839
840 max: maximal size of receive buffer allowed for automatically
841 selected receiver buffers for TCP socket.
842 Calling setsockopt() with SO_RCVBUF disables
843 automatic tuning of that socket's receive buffer size, in which
844 case this value is ignored.
845 Default: between 131072 and 32MB, depending on RAM size.
846
847 tcp_sack - BOOLEAN
848 Enable select acknowledgments (SACKS).
849
850 Possible values:
851
852 - 0 (disabled)
853 - 1 (enabled)
854
855 Default: 1 (enabled)
856
857 tcp_comp_sack_delay_ns - LONG INTEGER
858 TCP tries to reduce number of SACK sent, using a timer
859 based on 5% of SRTT, capped by this sysctl, in nano seconds.
860 The default is 1ms, based on TSO autosizing period.
861
862 Default : 1,000,000 ns (1 ms)
863
864 tcp_comp_sack_slack_ns - LONG INTEGER
865 This sysctl control the slack used when arming the
866 timer used by SACK compression. This gives extra time
867 for small RTT flows, and reduces system overhead by allowing
868 opportunistic reduction of timer interrupts.
869
870 Default : 100,000 ns (100 us)
871
872 tcp_comp_sack_nr - INTEGER
873 Max number of SACK that can be compressed.
874 Using 0 disables SACK compression.
875
876 Default : 44
877
878 tcp_backlog_ack_defer - BOOLEAN
879 If enabled, user thread processing socket backlog tries sending
880 one ACK for the whole queue. This helps to avoid potential
881 long latencies at end of a TCP socket syscall.
882
883 Possible values:
884
885 - 0 (disabled)
886 - 1 (enabled)
887
888 Default: 1 (enabled)
889
890 tcp_slow_start_after_idle - BOOLEAN
891 If enabled, provide RFC2861 behavior and time out the congestion
892 window after an idle period. An idle period is defined at
893 the current RTO. If unset, the congestion window will not
894 be timed out after an idle period.
895
896 Possible values:
897
898 - 0 (disabled)
899 - 1 (enabled)
900
901 Default: 1 (enabled)
902
903 tcp_stdurg - BOOLEAN
904 Use the Host requirements interpretation of the TCP urgent pointer field.
905 Most hosts use the older BSD interpretation, so if enabled,
906 Linux might not communicate correctly with them.
907
908 Possible values:
909
910 - 0 (disabled)
911 - 1 (enabled)
912
913 Default: 0 (disabled)
914
915 tcp_synack_retries - INTEGER
916 Number of times SYNACKs for a passive TCP connection attempt will
917 be retransmitted. Should not be higher than 255. Default value
918 is 5, which corresponds to 31seconds till the last retransmission
919 with the current initial RTO of 1second. With this the final timeout
920 for a passive TCP connection will happen after 63seconds.
921
922 tcp_syncookies - INTEGER
923 Only valid when the kernel was compiled with CONFIG_SYN_COOKIES
924 Send out syncookies when the syn backlog queue of a socket
925 overflows. This is to prevent against the common 'SYN flood attack'
926 Default: 1
927
928 Note, that syncookies is fallback facility.
929 It MUST NOT be used to help highly loaded servers to stand
930 against legal connection rate. If you see SYN flood warnings
931 in your logs, but investigation shows that they occur
932 because of overload with legal connections, you should tune
933 another parameters until this warning disappear.
934 See: tcp_max_syn_backlog, tcp_synack_retries, tcp_abort_on_overflow.
935
936 syncookies seriously violate TCP protocol, do not allow
937 to use TCP extensions, can result in serious degradation
938 of some services (f.e. SMTP relaying), visible not by you,
939 but your clients and relays, contacting you. While you see
940 SYN flood warnings in logs not being really flooded, your server
941 is seriously misconfigured.
942
943 If you want to test which effects syncookies have to your
944 network connections you can set this knob to 2 to enable
945 unconditionally generation of syncookies.
946
947 tcp_migrate_req - BOOLEAN
948 The incoming connection is tied to a specific listening socket when
949 the initial SYN packet is received during the three-way handshake.
950 When a listener is closed, in-flight request sockets during the
951 handshake and established sockets in the accept queue are aborted.
952
953 If the listener has SO_REUSEPORT enabled, other listeners on the
954 same port should have been able to accept such connections. This
955 option makes it possible to migrate such child sockets to another
956 listener after close() or shutdown().
957
958 The BPF_SK_REUSEPORT_SELECT_OR_MIGRATE type of eBPF program should
959 usually be used to define the policy to pick an alive listener.
960 Otherwise, the kernel will randomly pick an alive listener only if
961 this option is enabled.
962
963 Note that migration between listeners with different settings may
964 crash applications. Let's say migration happens from listener A to
965 B, and only B has TCP_SAVE_SYN enabled. B cannot read SYN data from
966 the requests migrated from A. To avoid such a situation, cancel
967 migration by returning SK_DROP in the type of eBPF program, or
968 disable this option.
969
970 Possible values:
971
972 - 0 (disabled)
973 - 1 (enabled)
974
975 Default: 0 (disabled)
976
977 tcp_fastopen - INTEGER
978 Enable TCP Fast Open (RFC7413) to send and accept data in the opening
979 SYN packet.
980
981 The client support is enabled by flag 0x1 (on by default). The client
982 then must use sendmsg() or sendto() with the MSG_FASTOPEN flag,
983 rather than connect() to send data in SYN.
984
985 The server support is enabled by flag 0x2 (off by default). Then
986 either enable for all listeners with another flag (0x400) or
987 enable individual listeners via TCP_FASTOPEN socket option with
988 the option value being the length of the syn-data backlog.
989
990 The values (bitmap) are
991
992 ===== ======== ======================================================
993 0x1 (client) enables sending data in the opening SYN on the client.
994 0x2 (server) enables the server support, i.e., allowing data in
995 a SYN packet to be accepted and passed to the
996 application before 3-way handshake finishes.
997 0x4 (client) send data in the opening SYN regardless of cookie
998 availability and without a cookie option.
999 0x200 (server) accept data-in-SYN w/o any cookie option present.
1000 0x400 (server) enable all listeners to support Fast Open by
1001 default without explicit TCP_FASTOPEN socket option.
1002 ===== ======== ======================================================
1004 Default: 0x1
1006 Note that additional client or server features are only
1007 effective if the basic support (0x1 and 0x2) are enabled respectively.
1009 tcp_fastopen_blackhole_timeout_sec - INTEGER
1010 Initial time period in second to disable Fastopen on active TCP sockets
1011 when a TFO firewall blackhole issue happens.
1012 This time period will grow exponentially when more blackhole issues
1013 get detected right after Fastopen is re-enabled and will reset to
1014 initial value when the blackhole issue goes away.
1015 0 to disable the blackhole detection.
1017 By default, it is set to 0 (feature is disabled).
1019 tcp_fastopen_key - list of comma separated 32-digit hexadecimal INTEGERs
1020 The list consists of a primary key and an optional backup key. The
1021 primary key is used for both creating and validating cookies, while the
1022 optional backup key is only used for validating cookies. The purpose of
1023 the backup key is to maximize TFO validation when keys are rotated.
1025 A randomly chosen primary key may be configured by the kernel if
1026 the tcp_fastopen sysctl is set to 0x400 (see above), or if the
1027 TCP_FASTOPEN setsockopt() optname is set and a key has not been
1028 previously configured via sysctl. If keys are configured via
1029 setsockopt() by using the TCP_FASTOPEN_KEY optname, then those
1030 per-socket keys will be used instead of any keys that are specified via
1031 sysctl.
1033 A key is specified as 4 8-digit hexadecimal integers which are separated
1034 by a '-' as: xxxxxxxx-xxxxxxxx-xxxxxxxx-xxxxxxxx. Leading zeros may be
1035 omitted. A primary and a backup key may be specified by separating them
1036 by a comma. If only one key is specified, it becomes the primary key and
1037 any previously configured backup keys are removed.
1039 tcp_syn_retries - INTEGER
1040 Number of times initial SYNs for an active TCP connection attempt
1041 will be retransmitted. Should not be higher than 127. Default value
1042 is 6, which corresponds to 67seconds (with tcp_syn_linear_timeouts = 4)
1043 till the last retransmission with the current initial RTO of 1second.
1044 With this the final timeout for an active TCP connection attempt
1045 will happen after 131seconds.
1047 tcp_timestamps - INTEGER
1048 Enable timestamps as defined in RFC1323.
1050 - 0: Disabled.
1051 - 1: Enable timestamps as defined in RFC1323 and use random offset for
1052 each connection rather than only using the current time.
1053 - 2: Like 1, but without random offsets.
1055 Default: 1
1057 tcp_min_tso_segs - INTEGER
1058 Minimal number of segments per TSO frame.
1060 Since linux-3.12, TCP does an automatic sizing of TSO frames,
1061 depending on flow rate, instead of filling 64Kbytes packets.
1062 For specific usages, it's possible to force TCP to build big
1063 TSO frames. Note that TCP stack might split too big TSO packets
1064 if available window is too small.
1066 Default: 2
1068 tcp_tso_rtt_log - INTEGER
1069 Adjustment of TSO packet sizes based on min_rtt
1071 Starting from linux-5.18, TCP autosizing can be tweaked
1072 for flows having small RTT.
1074 Old autosizing was splitting the pacing budget to send 1024 TSO
1075 per second.
1077 tso_packet_size = sk->sk_pacing_rate / 1024;
1079 With the new mechanism, we increase this TSO sizing using:
1081 distance = min_rtt_usec / (2^tcp_tso_rtt_log)
1082 tso_packet_size += gso_max_size >> distance;
1084 This means that flows between very close hosts can use bigger
1085 TSO packets, reducing their cpu costs.
1087 If you want to use the old autosizing, set this sysctl to 0.
1089 Default: 9 (2^9 = 512 usec)
1091 tcp_pacing_ss_ratio - INTEGER
1092 sk->sk_pacing_rate is set by TCP stack using a ratio applied
1093 to current rate. (current_rate = cwnd * mss / srtt)
1094 If TCP is in slow start, tcp_pacing_ss_ratio is applied
1095 to let TCP probe for bigger speeds, assuming cwnd can be
1096 doubled every other RTT.
1098 Default: 200
1100 tcp_pacing_ca_ratio - INTEGER
1101 sk->sk_pacing_rate is set by TCP stack using a ratio applied
1102 to current rate. (current_rate = cwnd * mss / srtt)
1103 If TCP is in congestion avoidance phase, tcp_pacing_ca_ratio
1104 is applied to conservatively probe for bigger throughput.
1106 Default: 120
1108 tcp_syn_linear_timeouts - INTEGER
1109 The number of times for an active TCP connection to retransmit SYNs with
1110 a linear backoff timeout before defaulting to an exponential backoff
1111 timeout. This has no effect on SYNACK at the passive TCP side.
1113 With an initial RTO of 1 and tcp_syn_linear_timeouts = 4 we would
1114 expect SYN RTOs to be: 1, 1, 1, 1, 1, 2, 4, ... (4 linear timeouts,
1115 and the first exponential backoff using 2^0 * initial_RTO).
1116 Default: 4
1118 tcp_tso_win_divisor - INTEGER
1119 This allows control over what percentage of the congestion window
1120 can be consumed by a single TSO frame.
1121 The setting of this parameter is a choice between burstiness and
1122 building larger TSO frames.
1124 Default: 3
1126 tcp_tw_reuse - INTEGER
1127 Enable reuse of TIME-WAIT sockets for new connections when it is
1128 safe from protocol viewpoint.
1130 - 0 - disable
1131 - 1 - global enable
1132 - 2 - enable for loopback traffic only
1134 It should not be changed without advice/request of technical
1135 experts.
1137 Default: 2
1139 tcp_tw_reuse_delay - UNSIGNED INTEGER
1140 The delay in milliseconds before a TIME-WAIT socket can be reused by a
1141 new connection, if TIME-WAIT socket reuse is enabled. The actual reuse
1142 threshold is within [N, N+1] range, where N is the requested delay in
1143 milliseconds, to ensure the delay interval is never shorter than the
1144 configured value.
1146 This setting contains an assumption about the other TCP timestamp clock
1147 tick interval. It should not be set to a value lower than the peer's
1148 clock tick for PAWS (Protection Against Wrapped Sequence numbers)
1149 mechanism work correctly for the reused connection.
1151 Default: 1000 (milliseconds)
1153 tcp_window_scaling - BOOLEAN
1154 Enable window scaling as defined in RFC1323.
1156 Possible values:
1158 - 0 (disabled)
1159 - 1 (enabled)
1161 Default: 1 (enabled)
1163 tcp_shrink_window - BOOLEAN
1164 This changes how the TCP receive window is calculated.
1166 RFC 7323, section 2.4, says there are instances when a retracted
1167 window can be offered, and that TCP implementations MUST ensure
1168 that they handle a shrinking window, as specified in RFC 1122.
1170 Possible values:
1172 - 0 (disabled) - The window is never shrunk.
1173 - 1 (enabled) - The window is shrunk when necessary to remain within
1174 the memory limit set by autotuning (sk_rcvbuf).
1175 This only occurs if a non-zero receive window
1176 scaling factor is also in effect.
1178 Default: 0 (disabled)
1180 tcp_wmem - vector of 3 INTEGERs: min, default, max
1181 min: Amount of memory reserved for send buffers for TCP sockets.
1182 Each TCP socket has rights to use it due to fact of its birth.
1184 Default: 4K
1186 default: initial size of send buffer used by TCP sockets. This
1187 value overrides net.core.wmem_default used by other protocols.
1189 It is usually lower than net.core.wmem_default.
1191 Default: 16K
1193 max: Maximal amount of memory allowed for automatically tuned
1194 send buffers for TCP sockets. This value does not override
1195 net.core.wmem_max. Calling setsockopt() with SO_SNDBUF disables
1196 automatic tuning of that socket's send buffer size, in which case
1197 this value is ignored.
1199 Default: between 64K and 4MB, depending on RAM size.
1201 tcp_notsent_lowat - UNSIGNED INTEGER
1202 A TCP socket can control the amount of unsent bytes in its write queue,
1203 thanks to TCP_NOTSENT_LOWAT socket option. poll()/select()/epoll()
1204 reports POLLOUT events if the amount of unsent bytes is below a per
1205 socket value, and if the write queue is not full. sendmsg() will
1206 also not add new buffers if the limit is hit.
1208 This global variable controls the amount of unsent data for
1209 sockets not using TCP_NOTSENT_LOWAT. For these sockets, a change
1210 to the global variable has immediate effect.
1212 Default: UINT_MAX (0xFFFFFFFF)
1214 tcp_workaround_signed_windows - BOOLEAN
1215 If enabled, assume no receipt of a window scaling option means the
1216 remote TCP is broken and treats the window as a signed quantity.
1217 If disabled, assume the remote TCP is not broken even if we do
1218 not receive a window scaling option from them.
1220 Possible values:
1222 - 0 (disabled)
1223 - 1 (enabled)
1225 Default: 0 (disabled)
1227 tcp_thin_linear_timeouts - BOOLEAN
1228 Enable dynamic triggering of linear timeouts for thin streams.
1229 If enabled, a check is performed upon retransmission by timeout to
1230 determine if the stream is thin (less than 4 packets in flight).
1231 As long as the stream is found to be thin, up to 6 linear
1232 timeouts may be performed before exponential backoff mode is
1233 initiated. This improves retransmission latency for
1234 non-aggressive thin streams, often found to be time-dependent.
1235 For more information on thin streams, see
1236 Documentation/networking/tcp-thin.rst
1238 Possible values:
1240 - 0 (disabled)
1241 - 1 (enabled)
1243 Default: 0 (disabled)
1245 tcp_limit_output_bytes - INTEGER
1246 Controls TCP Small Queue limit per tcp socket.
1247 TCP bulk sender tends to increase packets in flight until it
1248 gets losses notifications. With SNDBUF autotuning, this can
1249 result in a large amount of packets queued on the local machine
1250 (e.g.: qdiscs, CPU backlog, or device) hurting latency of other
1251 flows, for typical pfifo_fast qdiscs. tcp_limit_output_bytes
1252 limits the number of bytes on qdisc or device to reduce artificial
1253 RTT/cwnd and reduce bufferbloat.
1255 Default: 4194304 (4 MB)
1257 tcp_challenge_ack_limit - INTEGER
1258 Limits number of Challenge ACK sent per second, as recommended
1259 in RFC 5961 (Improving TCP's Robustness to Blind In-Window Attacks)
1260 Note that this per netns rate limit can allow some side channel
1261 attacks and probably should not be enabled.
1262 TCP stack implements per TCP socket limits anyway.
1263 Default: INT_MAX (unlimited)
1265 tcp_ehash_entries - INTEGER
1266 Show the number of hash buckets for TCP sockets in the current
1267 networking namespace.
1269 A negative value means the networking namespace does not own its
1270 hash buckets and shares the initial networking namespace's one.
1272 tcp_child_ehash_entries - INTEGER
1273 Control the number of hash buckets for TCP sockets in the child
1274 networking namespace, which must be set before clone() or unshare().
1276 If the value is not 0, the kernel uses a value rounded up to 2^n
1277 as the actual hash bucket size. 0 is a special value, meaning
1278 the child networking namespace will share the initial networking
1279 namespace's hash buckets.
1281 Note that the child will use the global one in case the kernel
1282 fails to allocate enough memory. In addition, the global hash
1283 buckets are spread over available NUMA nodes, but the allocation
1284 of the child hash table depends on the current process's NUMA
1285 policy, which could result in performance differences.
1287 Note also that the default value of tcp_max_tw_buckets and
1288 tcp_max_syn_backlog depend on the hash bucket size.
1290 Possible values: 0, 2^n (n: 0 - 24 (16Mi))
1292 Default: 0
1294 tcp_plb_enabled - BOOLEAN
1295 If enabled and the underlying congestion control (e.g. DCTCP) supports
1296 and enables PLB feature, TCP PLB (Protective Load Balancing) is
1297 enabled. PLB is described in the following paper:
1298 https://doi.org/10.1145/3544216.3544226. Based on PLB parameters,
1299 upon sensing sustained congestion, TCP triggers a change in
1300 flow label field for outgoing IPv6 packets. A change in flow label
1301 field potentially changes the path of outgoing packets for switches
1302 that use ECMP/WCMP for routing.
1304 PLB changes socket txhash which results in a change in IPv6 Flow Label
1305 field, and currently no-op for IPv4 headers. It is possible
1306 to apply PLB for IPv4 with other network header fields (e.g. TCP
1307 or IPv4 options) or using encapsulation where outer header is used
1308 by switches to determine next hop. In either case, further host
1309 and switch side changes will be needed.
1311 If enabled, PLB assumes that congestion signal (e.g. ECN) is made
1312 available and used by congestion control module to estimate a
1313 congestion measure (e.g. ce_ratio). PLB needs a congestion measure to
1314 make repathing decisions.
1316 Possible values:
1318 - 0 (disabled)
1319 - 1 (enabled)
1321 Default: 0 (disabled)
1323 tcp_plb_idle_rehash_rounds - INTEGER
1324 Number of consecutive congested rounds (RTT) seen after which
1325 a rehash can be performed, given there are no packets in flight.
1326 This is referred to as M in PLB paper:
1327 https://doi.org/10.1145/3544216.3544226.
1329 Possible Values: 0 - 31
1331 Default: 3
1333 tcp_plb_rehash_rounds - INTEGER
1334 Number of consecutive congested rounds (RTT) seen after which
1335 a forced rehash can be performed. Be careful when setting this
1336 parameter, as a small value increases the risk of retransmissions.
1337 This is referred to as N in PLB paper:
1338 https://doi.org/10.1145/3544216.3544226.
1340 Possible Values: 0 - 31
1342 Default: 12
1344 tcp_plb_suspend_rto_sec - INTEGER
1345 Time, in seconds, to suspend PLB in event of an RTO. In order to avoid
1346 having PLB repath onto a connectivity "black hole", after an RTO a TCP
1347 connection suspends PLB repathing for a random duration between 1x and
1348 2x of this parameter. Randomness is added to avoid concurrent rehashing
1349 of multiple TCP connections. This should be set corresponding to the
1350 amount of time it takes to repair a failed link.
1352 Possible Values: 0 - 255
1354 Default: 60
1356 tcp_plb_cong_thresh - INTEGER
1357 Fraction of packets marked with congestion over a round (RTT) to
1358 tag that round as congested. This is referred to as K in the PLB paper:
1359 https://doi.org/10.1145/3544216.3544226.
1361 The 0-1 fraction range is mapped to 0-256 range to avoid floating
1362 point operations. For example, 128 means that if at least 50% of
1363 the packets in a round were marked as congested then the round
1364 will be tagged as congested.
1366 Setting threshold to 0 means that PLB repaths every RTT regardless
1367 of congestion. This is not intended behavior for PLB and should be
1368 used only for experimentation purpose.
1370 Possible Values: 0 - 256
1372 Default: 128
1374 tcp_pingpong_thresh - INTEGER
1375 The number of estimated data replies sent for estimated incoming data
1376 requests that must happen before TCP considers that a connection is a
1377 "ping-pong" (request-response) connection for which delayed
1378 acknowledgments can provide benefits.
1380 This threshold is 1 by default, but some applications may need a higher
1381 threshold for optimal performance.
1383 Possible Values: 1 - 255
1385 Default: 1
1387 tcp_rto_min_us - INTEGER
1388 Minimal TCP retransmission timeout (in microseconds). Note that the
1389 rto_min route option has the highest precedence for configuring this
1390 setting, followed by the TCP_BPF_RTO_MIN and TCP_RTO_MIN_US socket
1391 options, followed by this tcp_rto_min_us sysctl.
1393 The recommended practice is to use a value less or equal to 200000
1394 microseconds.
1396 Possible Values: 1 - INT_MAX
1398 Default: 200000
1400 tcp_rto_max_ms - INTEGER
1401 Maximal TCP retransmission timeout (in ms).
1402 Note that TCP_RTO_MAX_MS socket option has higher precedence.
1404 When changing tcp_rto_max_ms, it is important to understand
1405 that tcp_retries2 might need a change.
1407 Possible Values: 1000 - 120,000
1409 Default: 120,000
1411 UDP variables
1412 =============
1414 udp_l3mdev_accept - BOOLEAN
1415 Enabling this option allows a "global" bound socket to work
1416 across L3 master domains (e.g., VRFs) with packets capable of
1417 being received regardless of the L3 domain in which they
1418 originated. Only valid when the kernel was compiled with
1419 CONFIG_NET_L3_MASTER_DEV.
1421 Possible values:
1423 - 0 (disabled)
1424 - 1 (enabled)
1426 Default: 0 (disabled)
1428 udp_mem - vector of 3 INTEGERs: min, pressure, max
1429 Number of pages allowed for queueing by all UDP sockets.
1431 min: Number of pages allowed for queueing by all UDP sockets.
1433 pressure: This value was introduced to follow format of tcp_mem.
1435 max: This value was introduced to follow format of tcp_mem.
1437 Default is calculated at boot time from amount of available memory.
1439 udp_rmem_min - INTEGER
1440 Minimal size of receive buffer used by UDP sockets in moderation.
1441 Each UDP socket is able to use the size for receiving data, even if
1442 total pages of UDP sockets exceed udp_mem pressure. The unit is byte.
1444 Default: 4K
1446 udp_wmem_min - INTEGER
1447 UDP does not have tx memory accounting and this tunable has no effect.
1449 udp_hash_entries - INTEGER
1450 Show the number of hash buckets for UDP sockets in the current
1451 networking namespace.
1453 A negative value means the networking namespace does not own its
1454 hash buckets and shares the initial networking namespace's one.
1456 udp_child_hash_entries - INTEGER
1457 Control the number of hash buckets for UDP sockets in the child
1458 networking namespace, which must be set before clone() or unshare().
1460 If the value is not 0, the kernel uses a value rounded up to 2^n
1461 as the actual hash bucket size. 0 is a special value, meaning
1462 the child networking namespace will share the initial networking
1463 namespace's hash buckets.
1465 Note that the child will use the global one in case the kernel
1466 fails to allocate enough memory. In addition, the global hash
1467 buckets are spread over available NUMA nodes, but the allocation
1468 of the child hash table depends on the current process's NUMA
1469 policy, which could result in performance differences.
1471 Possible values: 0, 2^n (n: 7 (128) - 16 (64K))
1473 Default: 0
1476 RAW variables
1477 =============
1479 raw_l3mdev_accept - BOOLEAN
1480 Enabling this option allows a "global" bound socket to work
1481 across L3 master domains (e.g., VRFs) with packets capable of
1482 being received regardless of the L3 domain in which they
1483 originated. Only valid when the kernel was compiled with
1484 CONFIG_NET_L3_MASTER_DEV.
1486 Possible values:
1488 - 0 (disabled)
1489 - 1 (enabled)
1491 Default: 1 (enabled)
1493 CIPSOv4 Variables
1494 =================
1496 cipso_cache_enable - BOOLEAN
1497 If enabled, enable additions to and lookups from the CIPSO label mapping
1498 cache. If disabled, additions are ignored and lookups always result in a
1499 miss. However, regardless of the setting the cache is still
1500 invalidated when required when means you can safely toggle this on and
1501 off and the cache will always be "safe".
1503 Possible values:
1505 - 0 (disabled)
1506 - 1 (enabled)
1508 Default: 1 (enabled)
1510 cipso_cache_bucket_size - INTEGER
1511 The CIPSO label cache consists of a fixed size hash table with each
1512 hash bucket containing a number of cache entries. This variable limits
1513 the number of entries in each hash bucket; the larger the value is, the
1514 more CIPSO label mappings that can be cached. When the number of
1515 entries in a given hash bucket reaches this limit adding new entries
1516 causes the oldest entry in the bucket to be removed to make room.
1518 Default: 10
1520 cipso_rbm_optfmt - BOOLEAN
1521 Enable the "Optimized Tag 1 Format" as defined in section 3.4.2.6 of
1522 the CIPSO draft specification (see Documentation/netlabel for details).
1523 This means that when set the CIPSO tag will be padded with empty
1524 categories in order to make the packet data 32-bit aligned.
1526 Possible values:
1528 - 0 (disabled)
1529 - 1 (enabled)
1531 Default: 0 (disabled)
1533 cipso_rbm_strictvalid - BOOLEAN
1534 If enabled, do a very strict check of the CIPSO option when
1535 ip_options_compile() is called. If disabled, relax the checks done during
1536 ip_options_compile(). Either way is "safe" as errors are caught else
1537 where in the CIPSO processing code but setting this to 0 (False) should
1538 result in less work (i.e. it should be faster) but could cause problems
1539 with other implementations that require strict checking.
1541 Possible values:
1543 - 0 (disabled)
1544 - 1 (enabled)
1546 Default: 0 (disabled)
1548 IP Variables
1549 ============
1551 ip_local_port_range - 2 INTEGERS
1552 Defines the local port range that is used by TCP and UDP to
1553 choose the local port. The first number is the first, the
1554 second the last local port number.
1555 If possible, it is better these numbers have different parity
1556 (one even and one odd value).
1557 Must be greater than or equal to ip_unprivileged_port_start.
1558 The default values are 32768 and 60999 respectively.
1560 ip_local_reserved_ports - list of comma separated ranges
1561 Specify the ports which are reserved for known third-party
1562 applications. These ports will not be used by automatic port
1563 assignments (e.g. when calling connect() or bind() with port
1564 number 0). Explicit port allocation behavior is unchanged.
1566 The format used for both input and output is a comma separated
1567 list of ranges (e.g. "1,2-4,10-10" for ports 1, 2, 3, 4 and
1568 10). Writing to the file will clear all previously reserved
1569 ports and update the current list with the one given in the
1570 input.
1572 Note that ip_local_port_range and ip_local_reserved_ports
1573 settings are independent and both are considered by the kernel
1574 when determining which ports are available for automatic port
1575 assignments.
1577 You can reserve ports which are not in the current
1578 ip_local_port_range, e.g.::
1580 $ cat /proc/sys/net/ipv4/ip_local_port_range
1581 32000 60999
1582 $ cat /proc/sys/net/ipv4/ip_local_reserved_ports
1583 8080,9148
1585 although this is redundant. However such a setting is useful
1586 if later the port range is changed to a value that will
1587 include the reserved ports. Also keep in mind, that overlapping
1588 of these ranges may affect probability of selecting ephemeral
1589 ports which are right after block of reserved ports.
1591 Default: Empty
1593 ip_unprivileged_port_start - INTEGER
1594 This is a per-namespace sysctl. It defines the first
1595 unprivileged port in the network namespace. Privileged ports
1596 require root or CAP_NET_BIND_SERVICE in order to bind to them.
1597 To disable all privileged ports, set this to 0. They must not
1598 overlap with the ip_local_port_range.
1600 Default: 1024
1602 ip_nonlocal_bind - BOOLEAN
1603 If enabled, allows processes to bind() to non-local IP addresses,
1604 which can be quite useful - but may break some applications.
1606 Possible values:
1608 - 0 (disabled)
1609 - 1 (enabled)
1611 Default: 0 (disabled)
1613 ip_autobind_reuse - BOOLEAN
1614 By default, bind() does not select the ports automatically even if
1615 the new socket and all sockets bound to the port have SO_REUSEADDR.
1616 ip_autobind_reuse allows bind() to reuse the port and this is useful
1617 when you use bind()+connect(), but may break some applications.
1618 The preferred solution is to use IP_BIND_ADDRESS_NO_PORT and this
1619 option should only be set by experts.
1621 Possible values:
1623 - 0 (disabled)
1624 - 1 (enabled)
1626 Default: 0 (disabled)
1628 ip_dynaddr - INTEGER
1629 If set non-zero, enables support for dynamic addresses.
1630 If set to a non-zero value larger than 1, a kernel log
1631 message will be printed when dynamic address rewriting
1632 occurs.
1634 Default: 0
1636 ip_early_demux - BOOLEAN
1637 Optimize input packet processing down to one demux for
1638 certain kinds of local sockets. Currently we only do this
1639 for established TCP and connected UDP sockets.
1641 It may add an additional cost for pure routing workloads that
1642 reduces overall throughput, in such case you should disable it.
1644 Possible values:
1646 - 0 (disabled)
1647 - 1 (enabled)
1649 Default: 1 (enabled)
1651 ping_group_range - 2 INTEGERS
1652 Restrict ICMP_PROTO datagram sockets to users in the group range.
1653 The default is "1 0", meaning, that nobody (not even root) may
1654 create ping sockets. Setting it to "100 100" would grant permissions
1655 to the single group. "0 4294967294" would enable it for the world, "100
1656 4294967294" would enable it for the users, but not daemons.
1658 tcp_early_demux - BOOLEAN
1659 Enable early demux for established TCP sockets.
1661 Possible values:
1663 - 0 (disabled)
1664 - 1 (enabled)
1666 Default: 1 (enabled)
1668 udp_early_demux - BOOLEAN
1669 Enable early demux for connected UDP sockets. Disable this if
1670 your system could experience more unconnected load.
1672 Possible values:
1674 - 0 (disabled)
1675 - 1 (enabled)
1677 Default: 1 (enabled)
1679 icmp_echo_ignore_all - BOOLEAN
1680 If enabled, then the kernel will ignore all ICMP ECHO
1681 requests sent to it.
1683 Possible values:
1685 - 0 (disabled)
1686 - 1 (enabled)
1688 Default: 0 (disabled)
1690 icmp_echo_enable_probe - BOOLEAN
1691 If enabled, then the kernel will respond to RFC 8335 PROBE
1692 requests sent to it.
1694 Possible values:
1696 - 0 (disabled)
1697 - 1 (enabled)
1699 Default: 0 (disabled)
1701 icmp_echo_ignore_broadcasts - BOOLEAN
1702 If enabled, then the kernel will ignore all ICMP ECHO and
1703 TIMESTAMP requests sent to it via broadcast/multicast.
1705 Possible values:
1707 - 0 (disabled)
1708 - 1 (enabled)
1710 Default: 1 (enabled)
1712 icmp_ratelimit - INTEGER
1713 Limit the maximal rates for sending ICMP packets whose type matches
1714 icmp_ratemask (see below) to specific targets.
1715 0 to disable any limiting,
1716 otherwise the minimal space between responses in milliseconds.
1717 Note that another sysctl, icmp_msgs_per_sec limits the number
1718 of ICMP packets sent on all targets.
1720 Default: 1000
1722 icmp_msgs_per_sec - INTEGER
1723 Limit maximal number of ICMP packets sent per second from this host.
1724 Only messages whose type matches icmp_ratemask (see below) are
1725 controlled by this limit. For security reasons, the precise count
1726 of messages per second is randomized.
1728 Default: 1000
1730 icmp_msgs_burst - INTEGER
1731 icmp_msgs_per_sec controls number of ICMP packets sent per second,
1732 while icmp_msgs_burst controls the burst size of these packets.
1733 For security reasons, the precise burst size is randomized.
1735 Default: 50
1737 icmp_ratemask - INTEGER
1738 Mask made of ICMP types for which rates are being limited.
1740 Significant bits: IHGFEDCBA9876543210
1742 Default mask: 0000001100000011000 (6168)
1744 Bit definitions (see include/linux/icmp.h):
1746 = =========================
1747 0 Echo Reply
1748 3 Destination Unreachable [1]_
1749 4 Source Quench [1]_
1750 5 Redirect
1751 8 Echo Request
1752 B Time Exceeded [1]_
1753 C Parameter Problem [1]_
1754 D Timestamp Request
1755 E Timestamp Reply
1756 F Info Request
1757 G Info Reply
1758 H Address Mask Request
1759 I Address Mask Reply
1760 = =========================
1762 .. [1] These are rate limited by default (see default mask above)
1764 icmp_ignore_bogus_error_responses - BOOLEAN
1765 Some routers violate RFC1122 by sending bogus responses to broadcast
1766 frames. Such violations are normally logged via a kernel warning.
1767 If enabled, the kernel will not give such warnings, which
1768 will avoid log file clutter.
1770 Possible values:
1772 - 0 (disabled)
1773 - 1 (enabled)
1775 Default: 1 (enabled)
1777 icmp_errors_use_inbound_ifaddr - BOOLEAN
1779 If disabled, icmp error messages are sent with the primary address of
1780 the exiting interface.
1782 If enabled, the message will be sent with the primary address of
1783 the interface that received the packet that caused the icmp error.
1784 This is the behaviour many network administrators will expect from
1785 a router. And it can make debugging complicated network layouts
1786 much easier.
1788 Note that if no primary address exists for the interface selected,
1789 then the primary address of the first non-loopback interface that
1790 has one will be used regardless of this setting.
1792 Possible values:
1794 - 0 (disabled)
1795 - 1 (enabled)
1797 Default: 0 (disabled)
1799 igmp_max_memberships - INTEGER
1800 Change the maximum number of multicast groups we can subscribe to.
1801 Default: 20
1803 Theoretical maximum value is bounded by having to send a membership
1804 report in a single datagram (i.e. the report can't span multiple
1805 datagrams, or risk confusing the switch and leaving groups you don't
1806 intend to).
1808 The number of supported groups 'M' is bounded by the number of group
1809 report entries you can fit into a single datagram of 65535 bytes.
1811 M = 65536-sizeof (ip header)/(sizeof(Group record))
1813 Group records are variable length, with a minimum of 12 bytes.
1814 So net.ipv4.igmp_max_memberships should not be set higher than:
1816 (65536-24) / 12 = 5459
1818 The value 5459 assumes no IP header options, so in practice
1819 this number may be lower.
1821 igmp_max_msf - INTEGER
1822 Maximum number of addresses allowed in the source filter list for a
1823 multicast group.
1825 Default: 10
1827 igmp_qrv - INTEGER
1828 Controls the IGMP query robustness variable (see RFC2236 8.1).
1830 Default: 2 (as specified by RFC2236 8.1)
1832 Minimum: 1 (as specified by RFC6636 4.5)
1834 force_igmp_version - INTEGER
1835 - 0 - (default) No enforcement of a IGMP version, IGMPv1/v2 fallback
1836 allowed. Will back to IGMPv3 mode again if all IGMPv1/v2 Querier
1837 Present timer expires.
1838 - 1 - Enforce to use IGMP version 1. Will also reply IGMPv1 report if
1839 receive IGMPv2/v3 query.
1840 - 2 - Enforce to use IGMP version 2. Will fallback to IGMPv1 if receive
1841 IGMPv1 query message. Will reply report if receive IGMPv3 query.
1842 - 3 - Enforce to use IGMP version 3. The same react with default 0.
1844 .. note::
1846 this is not the same with force_mld_version because IGMPv3 RFC3376
1847 Security Considerations does not have clear description that we could
1848 ignore other version messages completely as MLDv2 RFC3810. So make
1849 this value as default 0 is recommended.
1851 ``conf/interface/*``
1852 changes special settings per interface (where
1853 interface" is the name of your network interface)
1855 ``conf/all/*``
1856 is special, changes the settings for all interfaces
1858 log_martians - BOOLEAN
1859 Log packets with impossible addresses to kernel log.
1860 log_martians for the interface will be enabled if at least one of
1861 conf/{all,interface}/log_martians is set to TRUE,
1862 it will be disabled otherwise
1864 accept_redirects - BOOLEAN
1865 Accept ICMP redirect messages.
1866 accept_redirects for the interface will be enabled if:
1868 - both conf/{all,interface}/accept_redirects are TRUE in the case
1869 forwarding for the interface is enabled
1871 or
1873 - at least one of conf/{all,interface}/accept_redirects is TRUE in the
1874 case forwarding for the interface is disabled
1876 accept_redirects for the interface will be disabled otherwise
1878 default:
1880 - TRUE (host)
1881 - FALSE (router)
1883 forwarding - BOOLEAN
1884 Enable IP forwarding on this interface. This controls whether packets
1885 received _on_ this interface can be forwarded.
1887 mc_forwarding - BOOLEAN
1888 Do multicast routing. The kernel needs to be compiled with CONFIG_MROUTE
1889 and a multicast routing daemon is required.
1890 conf/all/mc_forwarding must also be set to TRUE to enable multicast
1891 routing for the interface
1893 medium_id - INTEGER
1894 Integer value used to differentiate the devices by the medium they
1895 are attached to. Two devices can have different id values when
1896 the broadcast packets are received only on one of them.
1897 The default value 0 means that the device is the only interface
1898 to its medium, value of -1 means that medium is not known.
1900 Currently, it is used to change the proxy_arp behavior:
1901 the proxy_arp feature is enabled for packets forwarded between
1902 two devices attached to different media.
1904 proxy_arp - BOOLEAN
1905 Do proxy arp.
1907 proxy_arp for the interface will be enabled if at least one of
1908 conf/{all,interface}/proxy_arp is set to TRUE,
1909 it will be disabled otherwise
1911 proxy_arp_pvlan - BOOLEAN
1912 Private VLAN proxy arp.
1914 Basically allow proxy arp replies back to the same interface
1915 (from which the ARP request/solicitation was received).
1917 This is done to support (ethernet) switch features, like RFC
1918 3069, where the individual ports are NOT allowed to
1919 communicate with each other, but they are allowed to talk to
1920 the upstream router. As described in RFC 3069, it is possible
1921 to allow these hosts to communicate through the upstream
1922 router by proxy_arp'ing. Don't need to be used together with
1923 proxy_arp.
1925 This technology is known by different names:
1927 - In RFC 3069 it is called VLAN Aggregation.
1928 - Cisco and Allied Telesyn call it Private VLAN.
1929 - Hewlett-Packard call it Source-Port filtering or port-isolation.
1930 - Ericsson call it MAC-Forced Forwarding (RFC Draft).
1932 proxy_delay - INTEGER
1933 Delay proxy response.
1935 Delay response to a neighbor solicitation when proxy_arp
1936 or proxy_ndp is enabled. A random value between [0, proxy_delay)
1937 will be chosen, setting to zero means reply with no delay.
1938 Value in jiffies. Defaults to 80.
1940 shared_media - BOOLEAN
1941 Send(router) or accept(host) RFC1620 shared media redirects.
1942 Overrides secure_redirects.
1944 shared_media for the interface will be enabled if at least one of
1945 conf/{all,interface}/shared_media is set to TRUE,
1946 it will be disabled otherwise
1948 default TRUE
1950 secure_redirects - BOOLEAN
1951 Accept ICMP redirect messages only to gateways listed in the
1952 interface's current gateway list. Even if disabled, RFC1122 redirect
1953 rules still apply.
1955 Overridden by shared_media.
1957 secure_redirects for the interface will be enabled if at least one of
1958 conf/{all,interface}/secure_redirects is set to TRUE,
1959 it will be disabled otherwise
1961 default TRUE
1963 send_redirects - BOOLEAN
1964 Send redirects, if router.
1966 send_redirects for the interface will be enabled if at least one of
1967 conf/{all,interface}/send_redirects is set to TRUE,
1968 it will be disabled otherwise
1970 Default: TRUE
1972 bootp_relay - BOOLEAN
1973 Accept packets with source address 0.b.c.d destined
1974 not to this host as local ones. It is supposed, that
1975 BOOTP relay daemon will catch and forward such packets.
1976 conf/all/bootp_relay must also be set to TRUE to enable BOOTP relay
1977 for the interface
1979 default FALSE
1981 Not Implemented Yet.
1983 accept_source_route - BOOLEAN
1984 Accept packets with SRR option.
1985 conf/all/accept_source_route must also be set to TRUE to accept packets
1986 with SRR option on the interface
1988 default
1990 - TRUE (router)
1991 - FALSE (host)
1993 accept_local - BOOLEAN
1994 Accept packets with local source addresses. In combination with
1995 suitable routing, this can be used to direct packets between two
1996 local interfaces over the wire and have them accepted properly.
1997 default FALSE
1999 route_localnet - BOOLEAN
2000 Do not consider loopback addresses as martian source or destination
2001 while routing. This enables the use of 127/8 for local routing purposes.
2003 default FALSE
2005 rp_filter - INTEGER
2006 - 0 - No source validation.
2007 - 1 - Strict mode as defined in RFC3704 Strict Reverse Path
2008 Each incoming packet is tested against the FIB and if the interface
2009 is not the best reverse path the packet check will fail.
2010 By default failed packets are discarded.
2011 - 2 - Loose mode as defined in RFC3704 Loose Reverse Path
2012 Each incoming packet's source address is also tested against the FIB
2013 and if the source address is not reachable via any interface
2014 the packet check will fail.
2016 Current recommended practice in RFC3704 is to enable strict mode
2017 to prevent IP spoofing from DDos attacks. If using asymmetric routing
2018 or other complicated routing, then loose mode is recommended.
2020 The max value from conf/{all,interface}/rp_filter is used
2021 when doing source validation on the {interface}.
2023 Default value is 0. Note that some distributions enable it
2024 in startup scripts.
2026 src_valid_mark - BOOLEAN
2027 - 0 - The fwmark of the packet is not included in reverse path
2028 route lookup. This allows for asymmetric routing configurations
2029 utilizing the fwmark in only one direction, e.g., transparent
2030 proxying.
2032 - 1 - The fwmark of the packet is included in reverse path route
2033 lookup. This permits rp_filter to function when the fwmark is
2034 used for routing traffic in both directions.
2036 This setting also affects the utilization of fmwark when
2037 performing source address selection for ICMP replies, or
2038 determining addresses stored for the IPOPT_TS_TSANDADDR and
2039 IPOPT_RR IP options.
2041 The max value from conf/{all,interface}/src_valid_mark is used.
2043 Default value is 0.
2045 arp_filter - BOOLEAN
2046 - 1 - Allows you to have multiple network interfaces on the same
2047 subnet, and have the ARPs for each interface be answered
2048 based on whether or not the kernel would route a packet from
2049 the ARP'd IP out that interface (therefore you must use source
2050 based routing for this to work). In other words it allows control
2051 of which cards (usually 1) will respond to an arp request.
2053 - 0 - (default) The kernel can respond to arp requests with addresses
2054 from other interfaces. This may seem wrong but it usually makes
2055 sense, because it increases the chance of successful communication.
2056 IP addresses are owned by the complete host on Linux, not by
2057 particular interfaces. Only for more complex setups like load-
2058 balancing, does this behaviour cause problems.
2060 arp_filter for the interface will be enabled if at least one of
2061 conf/{all,interface}/arp_filter is set to TRUE,
2062 it will be disabled otherwise
2064 arp_announce - INTEGER
2065 Define different restriction levels for announcing the local
2066 source IP address from IP packets in ARP requests sent on
2067 interface:
2069 - 0 - (default) Use any local address, configured on any interface
2070 - 1 - Try to avoid local addresses that are not in the target's
2071 subnet for this interface. This mode is useful when target
2072 hosts reachable via this interface require the source IP
2073 address in ARP requests to be part of their logical network
2074 configured on the receiving interface. When we generate the
2075 request we will check all our subnets that include the
2076 target IP and will preserve the source address if it is from
2077 such subnet. If there is no such subnet we select source
2078 address according to the rules for level 2.
2079 - 2 - Always use the best local address for this target.
2080 In this mode we ignore the source address in the IP packet
2081 and try to select local address that we prefer for talks with
2082 the target host. Such local address is selected by looking
2083 for primary IP addresses on all our subnets on the outgoing
2084 interface that include the target IP address. If no suitable
2085 local address is found we select the first local address
2086 we have on the outgoing interface or on all other interfaces,
2087 with the hope we will receive reply for our request and
2088 even sometimes no matter the source IP address we announce.
2090 The max value from conf/{all,interface}/arp_announce is used.
2092 Increasing the restriction level gives more chance for
2093 receiving answer from the resolved target while decreasing
2094 the level announces more valid sender's information.
2096 arp_ignore - INTEGER
2097 Define different modes for sending replies in response to
2098 received ARP requests that resolve local target IP addresses:
2100 - 0 - (default): reply for any local target IP address, configured
2101 on any interface
2102 - 1 - reply only if the target IP address is local address
2103 configured on the incoming interface
2104 - 2 - reply only if the target IP address is local address
2105 configured on the incoming interface and both with the
2106 sender's IP address are part from same subnet on this interface
2107 - 3 - do not reply for local addresses configured with scope host,
2108 only resolutions for global and link addresses are replied
2109 - 4-7 - reserved
2110 - 8 - do not reply for all local addresses
2112 The max value from conf/{all,interface}/arp_ignore is used
2113 when ARP request is received on the {interface}
2115 arp_notify - BOOLEAN
2116 Define mode for notification of address and device changes.
2118 == ==========================================================
2119 0 (default): do nothing
2120 1 Generate gratuitous arp requests when device is brought up
2121 or hardware address changes.
2122 == ==========================================================
2124 arp_accept - INTEGER
2125 Define behavior for accepting gratuitous ARP (garp) frames from devices
2126 that are not already present in the ARP table:
2128 - 0 - don't create new entries in the ARP table
2129 - 1 - create new entries in the ARP table
2130 - 2 - create new entries only if the source IP address is in the same
2131 subnet as an address configured on the interface that received the
2132 garp message.
2134 Both replies and requests type gratuitous arp will trigger the
2135 ARP table to be updated, if this setting is on.
2137 If the ARP table already contains the IP address of the
2138 gratuitous arp frame, the arp table will be updated regardless
2139 if this setting is on or off.
2141 arp_evict_nocarrier - BOOLEAN
2142 Clears the ARP cache on NOCARRIER events. This option is important for
2143 wireless devices where the ARP cache should not be cleared when roaming
2144 between access points on the same network. In most cases this should
2145 remain as the default (1).
2147 Possible values:
2149 - 0 (disabled) - Do not clear ARP cache on NOCARRIER events
2150 - 1 (enabled) - Clear the ARP cache on NOCARRIER events
2152 Default: 1 (enabled)
2154 mcast_solicit - INTEGER
2155 The maximum number of multicast probes in INCOMPLETE state,
2156 when the associated hardware address is unknown. Defaults
2157 to 3.
2159 ucast_solicit - INTEGER
2160 The maximum number of unicast probes in PROBE state, when
2161 the hardware address is being reconfirmed. Defaults to 3.
2163 app_solicit - INTEGER
2164 The maximum number of probes to send to the user space ARP daemon
2165 via netlink before dropping back to multicast probes (see
2166 mcast_resolicit). Defaults to 0.
2168 mcast_resolicit - INTEGER
2169 The maximum number of multicast probes after unicast and
2170 app probes in PROBE state. Defaults to 0.
2172 disable_policy - BOOLEAN
2173 Disable IPSEC policy (SPD) for this interface
2175 Possible values:
2177 - 0 (disabled)
2178 - 1 (enabled)
2180 Default: 0 (disabled)
2182 disable_xfrm - BOOLEAN
2183 Disable IPSEC encryption on this interface, whatever the policy
2185 Possible values:
2187 - 0 (disabled)
2188 - 1 (enabled)
2190 Default: 0 (disabled)
2192 igmpv2_unsolicited_report_interval - INTEGER
2193 The interval in milliseconds in which the next unsolicited
2194 IGMPv1 or IGMPv2 report retransmit will take place.
2196 Default: 10000 (10 seconds)
2198 igmpv3_unsolicited_report_interval - INTEGER
2199 The interval in milliseconds in which the next unsolicited
2200 IGMPv3 report retransmit will take place.
2202 Default: 1000 (1 seconds)
2204 ignore_routes_with_linkdown - BOOLEAN
2205 Ignore routes whose link is down when performing a FIB lookup.
2207 Possible values:
2209 - 0 (disabled)
2210 - 1 (enabled)
2212 Default: 0 (disabled)
2214 promote_secondaries - BOOLEAN
2215 When a primary IP address is removed from this interface
2216 promote a corresponding secondary IP address instead of
2217 removing all the corresponding secondary IP addresses.
2219 Possible values:
2221 - 0 (disabled)
2222 - 1 (enabled)
2224 Default: 0 (disabled)
2226 drop_unicast_in_l2_multicast - BOOLEAN
2227 Drop any unicast IP packets that are received in link-layer
2228 multicast (or broadcast) frames.
2230 This behavior (for multicast) is actually a SHOULD in RFC
2231 1122, but is disabled by default for compatibility reasons.
2233 Possible values:
2235 - 0 (disabled)
2236 - 1 (enabled)
2238 Default: 0 (disabled)
2240 drop_gratuitous_arp - BOOLEAN
2241 Drop all gratuitous ARP frames, for example if there's a known
2242 good ARP proxy on the network and such frames need not be used
2243 (or in the case of 802.11, must not be used to prevent attacks.)
2245 Possible values:
2247 - 0 (disabled)
2248 - 1 (enabled)
2250 Default: 0 (disabled)
2253 tag - INTEGER
2254 Allows you to write a number, which can be used as required.
2256 Default value is 0.
2258 xfrm4_gc_thresh - INTEGER
2259 (Obsolete since linux-4.14)
2260 The threshold at which we will start garbage collecting for IPv4
2261 destination cache entries. At twice this value the system will
2262 refuse new allocations.
2264 igmp_link_local_mcast_reports - BOOLEAN
2265 Enable IGMP reports for link local multicast groups in the
2266 224.0.0.X range.
2268 Default TRUE
2270 Alexey Kuznetsov.
2273 Updated by:
2275 - Andi Kleen
2277 - Nicolas Delon
2283 /proc/sys/net/ipv6/* Variables
2284 ==============================
2286 IPv6 has no global variables such as tcp_*. tcp_* settings under ipv4/ also
2287 apply to IPv6 [XXX?].
2289 bindv6only - BOOLEAN
2290 Default value for IPV6_V6ONLY socket option,
2291 which restricts use of the IPv6 socket to IPv6 communication
2292 only.
2294 Possible values:
2296 - 0 (disabled) - enable IPv4-mapped address feature
2297 - 1 (enabled) - disable IPv4-mapped address feature
2299 Default: 0 (disabled)
2301 flowlabel_consistency - BOOLEAN
2302 Protect the consistency (and unicity) of flow label.
2303 You have to disable it to use IPV6_FL_F_REFLECT flag on the
2304 flow label manager.
2306 Possible values:
2308 - 0 (disabled)
2309 - 1 (enabled)
2311 Default: 1 (enabled)
2313 auto_flowlabels - INTEGER
2314 Automatically generate flow labels based on a flow hash of the
2315 packet. This allows intermediate devices, such as routers, to
2316 identify packet flows for mechanisms like Equal Cost Multipath
2317 Routing (see RFC 6438).
2319 = ===========================================================
2320 0 automatic flow labels are completely disabled
2321 1 automatic flow labels are enabled by default, they can be
2322 disabled on a per socket basis using the IPV6_AUTOFLOWLABEL
2323 socket option
2324 2 automatic flow labels are allowed, they may be enabled on a
2325 per socket basis using the IPV6_AUTOFLOWLABEL socket option
2326 3 automatic flow labels are enabled and enforced, they cannot
2327 be disabled by the socket option
2328 = ===========================================================
2330 Default: 1
2332 flowlabel_state_ranges - BOOLEAN
2333 Split the flow label number space into two ranges. 0-0x7FFFF is
2334 reserved for the IPv6 flow manager facility, 0x80000-0xFFFFF
2335 is reserved for stateless flow labels as described in RFC6437.
2337 Possible values:
2339 - 0 (disabled)
2340 - 1 (enabled)
2342 Default: 1 (enabled)
2345 flowlabel_reflect - INTEGER
2346 Control flow label reflection. Needed for Path MTU
2347 Discovery to work with Equal Cost Multipath Routing in anycast
2348 environments. See RFC 7690 and:
2349 https://tools.ietf.org/html/draft-wang-6man-flow-label-reflection-01
2351 This is a bitmask.
2353 - 1: enabled for established flows
2355 Note that this prevents automatic flowlabel changes, as done
2356 in "tcp: change IPv6 flow-label upon receiving spurious retransmission"
2357 and "tcp: Change txhash on every SYN and RTO retransmit"
2359 - 2: enabled for TCP RESET packets (no active listener)
2360 If set, a RST packet sent in response to a SYN packet on a closed
2361 port will reflect the incoming flow label.
2363 - 4: enabled for ICMPv6 echo reply messages.
2365 Default: 0
2367 fib_multipath_hash_policy - INTEGER
2368 Controls which hash policy to use for multipath routes.
2370 Default: 0 (Layer 3)
2372 Possible values:
2374 - 0 - Layer 3 (source and destination addresses plus flow label)
2375 - 1 - Layer 4 (standard 5-tuple)
2376 - 2 - Layer 3 or inner Layer 3 if present
2377 - 3 - Custom multipath hash. Fields used for multipath hash calculation
2378 are determined by fib_multipath_hash_fields sysctl
2380 fib_multipath_hash_fields - UNSIGNED INTEGER
2381 When fib_multipath_hash_policy is set to 3 (custom multipath hash), the
2382 fields used for multipath hash calculation are determined by this
2383 sysctl.
2385 This value is a bitmask which enables various fields for multipath hash
2386 calculation.
2388 Possible fields are:
2390 ====== ============================
2391 0x0001 Source IP address
2392 0x0002 Destination IP address
2393 0x0004 IP protocol
2394 0x0008 Flow Label
2395 0x0010 Source port
2396 0x0020 Destination port
2397 0x0040 Inner source IP address
2398 0x0080 Inner destination IP address
2399 0x0100 Inner IP protocol
2400 0x0200 Inner Flow Label
2401 0x0400 Inner source port
2402 0x0800 Inner destination port
2403 ====== ============================
2405 Default: 0x0007 (source IP, destination IP and IP protocol)
2407 anycast_src_echo_reply - BOOLEAN
2408 Controls the use of anycast addresses as source addresses for ICMPv6
2409 echo reply
2411 Possible values:
2413 - 0 (disabled)
2414 - 1 (enabled)
2416 Default: 0 (disabled)
2419 idgen_delay - INTEGER
2420 Controls the delay in seconds after which time to retry
2421 privacy stable address generation if a DAD conflict is
2422 detected.
2424 Default: 1 (as specified in RFC7217)
2426 idgen_retries - INTEGER
2427 Controls the number of retries to generate a stable privacy
2428 address if a DAD conflict is detected.
2430 Default: 3 (as specified in RFC7217)
2432 mld_qrv - INTEGER
2433 Controls the MLD query robustness variable (see RFC3810 9.1).
2435 Default: 2 (as specified by RFC3810 9.1)
2437 Minimum: 1 (as specified by RFC6636 4.5)
2439 max_dst_opts_number - INTEGER
2440 Maximum number of non-padding TLVs allowed in a Destination
2441 options extension header. If this value is less than zero
2442 then unknown options are disallowed and the number of known
2443 TLVs allowed is the absolute value of this number.
2445 Default: 8
2447 max_hbh_opts_number - INTEGER
2448 Maximum number of non-padding TLVs allowed in a Hop-by-Hop
2449 options extension header. If this value is less than zero
2450 then unknown options are disallowed and the number of known
2451 TLVs allowed is the absolute value of this number.
2453 Default: 8
2455 max_dst_opts_length - INTEGER
2456 Maximum length allowed for a Destination options extension
2457 header.
2459 Default: INT_MAX (unlimited)
2461 max_hbh_length - INTEGER
2462 Maximum length allowed for a Hop-by-Hop options extension
2463 header.
2465 Default: INT_MAX (unlimited)
2467 skip_notify_on_dev_down - BOOLEAN
2468 Controls whether an RTM_DELROUTE message is generated for routes
2469 removed when a device is taken down or deleted. IPv4 does not
2470 generate this message; IPv6 does by default. Setting this sysctl
2471 to true skips the message, making IPv4 and IPv6 on par in relying
2472 on userspace caches to track link events and evict routes.
2474 Possible values:
2476 - 0 (disabled) - generate the message
2477 - 1 (enabled) - skip generating the message
2479 Default: 0 (disabled)
2481 nexthop_compat_mode - BOOLEAN
2482 New nexthop API provides a means for managing nexthops independent of
2483 prefixes. Backwards compatibility with old route format is enabled by
2484 default which means route dumps and notifications contain the new
2485 nexthop attribute but also the full, expanded nexthop definition.
2486 Further, updates or deletes of a nexthop configuration generate route
2487 notifications for each fib entry using the nexthop. Once a system
2488 understands the new API, this sysctl can be disabled to achieve full
2489 performance benefits of the new API by disabling the nexthop expansion
2490 and extraneous notifications.
2492 Note that as a backward-compatible mode, dumping of modern features
2493 might be incomplete or wrong. For example, resilient groups will not be
2494 shown as such, but rather as just a list of next hops. Also weights that
2495 do not fit into 8 bits will show incorrectly.
2497 Default: true (backward compat mode)
2499 fib_notify_on_flag_change - INTEGER
2500 Whether to emit RTM_NEWROUTE notifications whenever RTM_F_OFFLOAD/
2501 RTM_F_TRAP/RTM_F_OFFLOAD_FAILED flags are changed.
2503 After installing a route to the kernel, user space receives an
2504 acknowledgment, which means the route was installed in the kernel,
2505 but not necessarily in hardware.
2506 It is also possible for a route already installed in hardware to change
2507 its action and therefore its flags. For example, a host route that is
2508 trapping packets can be "promoted" to perform decapsulation following
2509 the installation of an IPinIP/VXLAN tunnel.
2510 The notifications will indicate to user-space the state of the route.
2512 Default: 0 (Do not emit notifications.)
2514 Possible values:
2516 - 0 - Do not emit notifications.
2517 - 1 - Emit notifications.
2518 - 2 - Emit notifications only for RTM_F_OFFLOAD_FAILED flag change.
2520 ioam6_id - INTEGER
2521 Define the IOAM id of this node. Uses only 24 bits out of 32 in total.
2523 Possible value range:
2525 - Min: 0
2526 - Max: 0xFFFFFF
2528 Default: 0xFFFFFF
2530 ioam6_id_wide - LONG INTEGER
2531 Define the wide IOAM id of this node. Uses only 56 bits out of 64 in
2532 total. Can be different from ioam6_id.
2534 Possible value range:
2536 - Min: 0
2537 - Max: 0xFFFFFFFFFFFFFF
2539 Default: 0xFFFFFFFFFFFFFF
2541 IPv6 Fragmentation:
2543 ip6frag_high_thresh - INTEGER
2544 Maximum memory used to reassemble IPv6 fragments. When
2545 ip6frag_high_thresh bytes of memory is allocated for this purpose,
2546 the fragment handler will toss packets until ip6frag_low_thresh
2547 is reached.
2549 ip6frag_low_thresh - INTEGER
2550 See ip6frag_high_thresh
2552 ip6frag_time - INTEGER
2553 Time in seconds to keep an IPv6 fragment in memory.
2555 ``conf/default/*``:
2556 Change the interface-specific default settings.
2558 These settings would be used during creating new interfaces.
2561 ``conf/all/*``:
2562 Change all the interface-specific settings.
2564 [XXX: Other special features than forwarding?]
2566 conf/all/disable_ipv6 - BOOLEAN
2567 Changing this value is same as changing ``conf/default/disable_ipv6``
2568 setting and also all per-interface ``disable_ipv6`` settings to the same
2569 value.
2571 Reading this value does not have any particular meaning. It does not say
2572 whether IPv6 support is enabled or disabled. Returned value can be 1
2573 also in the case when some interface has ``disable_ipv6`` set to 0 and
2574 has configured IPv6 addresses.
2576 conf/all/forwarding - BOOLEAN
2577 Enable global IPv6 forwarding between all interfaces.
2579 IPv4 and IPv6 work differently here; the ``force_forwarding`` flag must
2580 be used to control which interfaces may forward packets.
2582 This also sets all interfaces' Host/Router setting
2583 'forwarding' to the specified value. See below for details.
2585 This referred to as global forwarding.
2587 proxy_ndp - BOOLEAN
2588 Do proxy ndp.
2590 Possible values:
2592 - 0 (disabled)
2593 - 1 (enabled)
2595 Default: 0 (disabled)
2597 force_forwarding - BOOLEAN
2598 Enable forwarding on this interface only -- regardless of the setting on
2599 ``conf/all/forwarding``. When setting ``conf.all.forwarding`` to 0,
2600 the ``force_forwarding`` flag will be reset on all interfaces.
2602 fwmark_reflect - BOOLEAN
2603 Controls the fwmark of kernel-generated IPv6 reply packets that are not
2604 associated with a socket for example, TCP RSTs or ICMPv6 echo replies).
2605 If disabled, these packets have a fwmark of zero. If enabled, they have the
2606 fwmark of the packet they are replying to.
2608 Possible values:
2610 - 0 (disabled)
2611 - 1 (enabled)
2613 Default: 0 (disabled)
2615 ``conf/interface/*``:
2616 Change special settings per interface.
2618 The functional behaviour for certain settings is different
2619 depending on whether local forwarding is enabled or not.
2621 accept_ra - INTEGER
2622 Accept Router Advertisements; autoconfigure using them.
2624 It also determines whether or not to transmit Router
2625 Solicitations. If and only if the functional setting is to
2626 accept Router Advertisements, Router Solicitations will be
2627 transmitted.
2629 Possible values are:
2631 == ===========================================================
2632 0 Do not accept Router Advertisements.
2633 1 Accept Router Advertisements if forwarding is disabled.
2634 2 Overrule forwarding behaviour. Accept Router Advertisements
2635 even if forwarding is enabled.
2636 == ===========================================================
2638 Functional default:
2640 - enabled if local forwarding is disabled.
2641 - disabled if local forwarding is enabled.
2643 accept_ra_defrtr - BOOLEAN
2644 Learn default router in Router Advertisement.
2646 Functional default:
2648 - enabled if accept_ra is enabled.
2649 - disabled if accept_ra is disabled.
2651 ra_defrtr_metric - UNSIGNED INTEGER
2652 Route metric for default route learned in Router Advertisement. This value
2653 will be assigned as metric for the default route learned via IPv6 Router
2654 Advertisement. Takes affect only if accept_ra_defrtr is enabled.
2656 Possible values:
2657 1 to 0xFFFFFFFF
2659 Default: IP6_RT_PRIO_USER i.e. 1024.
2661 accept_ra_from_local - BOOLEAN
2662 Accept RA with source-address that is found on local machine
2663 if the RA is otherwise proper and able to be accepted.
2665 Default is to NOT accept these as it may be an un-intended
2666 network loop.
2668 Functional default:
2670 - enabled if accept_ra_from_local is enabled
2671 on a specific interface.
2672 - disabled if accept_ra_from_local is disabled
2673 on a specific interface.
2675 accept_ra_min_hop_limit - INTEGER
2676 Minimum hop limit Information in Router Advertisement.
2678 Hop limit Information in Router Advertisement less than this
2679 variable shall be ignored.
2681 Default: 1
2683 accept_ra_min_lft - INTEGER
2684 Minimum acceptable lifetime value in Router Advertisement.
2686 RA sections with a lifetime less than this value shall be
2687 ignored. Zero lifetimes stay unaffected.
2689 Default: 0
2691 accept_ra_pinfo - BOOLEAN
2692 Learn Prefix Information in Router Advertisement.
2694 Functional default:
2696 - enabled if accept_ra is enabled.
2697 - disabled if accept_ra is disabled.
2699 ra_honor_pio_life - BOOLEAN
2700 Whether to use RFC4862 Section 5.5.3e to determine the valid
2701 lifetime of an address matching a prefix sent in a Router
2702 Advertisement Prefix Information Option.
2704 Possible values:
2706 - 0 (disabled) - RFC4862 section 5.5.3e is used to determine
2707 the valid lifetime of the address.
2708 - 1 (enabled) - the PIO valid lifetime will always be honored.
2710 Default: 0 (disabled)
2712 ra_honor_pio_pflag - BOOLEAN
2713 The Prefix Information Option P-flag indicates the network can
2714 allocate a unique IPv6 prefix per client using DHCPv6-PD.
2715 This sysctl can be enabled when a userspace DHCPv6-PD client
2716 is running to cause the P-flag to take effect: i.e. the
2717 P-flag suppresses any effects of the A-flag within the same
2718 PIO. For a given PIO, P=1 and A=1 is treated as A=0.
2720 Possible values:
2722 - 0 (disabled) - the P-flag is ignored.
2723 - 1 (enabled) - the P-flag will disable SLAAC autoconfiguration
2724 for the given Prefix Information Option.
2726 Default: 0 (disabled)
2728 accept_ra_rt_info_min_plen - INTEGER
2729 Minimum prefix length of Route Information in RA.
2731 Route Information w/ prefix smaller than this variable shall
2732 be ignored.
2734 Functional default:
2736 * 0 if accept_ra_rtr_pref is enabled.
2737 * -1 if accept_ra_rtr_pref is disabled.
2739 accept_ra_rt_info_max_plen - INTEGER
2740 Maximum prefix length of Route Information in RA.
2742 Route Information w/ prefix larger than this variable shall
2743 be ignored.
2745 Functional default:
2747 * 0 if accept_ra_rtr_pref is enabled.
2748 * -1 if accept_ra_rtr_pref is disabled.
2750 accept_ra_rtr_pref - BOOLEAN
2751 Accept Router Preference in RA.
2753 Functional default:
2755 - enabled if accept_ra is enabled.
2756 - disabled if accept_ra is disabled.
2758 accept_ra_mtu - BOOLEAN
2759 Apply the MTU value specified in RA option 5 (RFC4861). If
2760 disabled, the MTU specified in the RA will be ignored.
2762 Functional default:
2764 - enabled if accept_ra is enabled.
2765 - disabled if accept_ra is disabled.
2767 accept_redirects - BOOLEAN
2768 Accept Redirects.
2770 Functional default:
2772 - enabled if local forwarding is disabled.
2773 - disabled if local forwarding is enabled.
2775 accept_source_route - INTEGER
2776 Accept source routing (routing extension header).
2778 - >= 0: Accept only routing header type 2.
2779 - < 0: Do not accept routing header.
2781 Default: 0
2783 autoconf - BOOLEAN
2784 Autoconfigure addresses using Prefix Information in Router
2785 Advertisements.
2787 Functional default:
2789 - enabled if accept_ra_pinfo is enabled.
2790 - disabled if accept_ra_pinfo is disabled.
2792 dad_transmits - INTEGER
2793 The amount of Duplicate Address Detection probes to send.
2795 Default: 1
2797 forwarding - INTEGER
2798 Configure interface-specific Host/Router behaviour.
2800 .. note::
2802 It is recommended to have the same setting on all
2803 interfaces; mixed router/host scenarios are rather uncommon.
2805 Possible values are:
2807 - 0 Forwarding disabled
2808 - 1 Forwarding enabled
2810 **FALSE (0)**:
2812 By default, Host behaviour is assumed. This means:
2814 1. IsRouter flag is not set in Neighbour Advertisements.
2815 2. If accept_ra is TRUE (default), transmit Router
2816 Solicitations.
2817 3. If accept_ra is TRUE (default), accept Router
2818 Advertisements (and do autoconfiguration).
2819 4. If accept_redirects is TRUE (default), accept Redirects.
2821 **TRUE (1)**:
2823 If local forwarding is enabled, Router behaviour is assumed.
2824 This means exactly the reverse from the above:
2826 1. IsRouter flag is set in Neighbour Advertisements.
2827 2. Router Solicitations are not sent unless accept_ra is 2.
2828 3. Router Advertisements are ignored unless accept_ra is 2.
2829 4. Redirects are ignored.
2831 Default: 0 (disabled) if global forwarding is disabled (default),
2832 otherwise 1 (enabled).
2834 hop_limit - INTEGER
2835 Default Hop Limit to set.
2837 Default: 64
2839 mtu - INTEGER
2840 Default Maximum Transfer Unit
2842 Default: 1280 (IPv6 required minimum)
2844 ip_nonlocal_bind - BOOLEAN
2845 If enabled, allows processes to bind() to non-local IPv6 addresses,
2846 which can be quite useful - but may break some applications.
2848 Possible values:
2850 - 0 (disabled)
2851 - 1 (enabled)
2853 Default: 0 (disabled)
2855 router_probe_interval - INTEGER
2856 Minimum interval (in seconds) between Router Probing described
2857 in RFC4191.
2859 Default: 60
2861 router_solicitation_delay - INTEGER
2862 Number of seconds to wait after interface is brought up
2863 before sending Router Solicitations.
2865 Default: 1
2867 router_solicitation_interval - INTEGER
2868 Number of seconds to wait between Router Solicitations.
2870 Default: 4
2872 router_solicitations - INTEGER
2873 Number of Router Solicitations to send until assuming no
2874 routers are present.
2876 Default: 3
2878 use_oif_addrs_only - BOOLEAN
2879 When enabled, the candidate source addresses for destinations
2880 routed via this interface are restricted to the set of addresses
2881 configured on this interface (vis. RFC 6724, section 4).
2883 Possible values:
2885 - 0 (disabled)
2886 - 1 (enabled)
2888 Default: 0 (disabled)
2890 use_tempaddr - INTEGER
2891 Preference for Privacy Extensions (RFC3041).
2893 * <= 0 : disable Privacy Extensions
2894 * == 1 : enable Privacy Extensions, but prefer public
2895 addresses over temporary addresses.
2896 * > 1 : enable Privacy Extensions and prefer temporary
2897 addresses over public addresses.
2899 Default:
2901 * 0 (for most devices)
2902 * -1 (for point-to-point devices and loopback devices)
2904 temp_valid_lft - INTEGER
2905 valid lifetime (in seconds) for temporary addresses. If less than the
2906 minimum required lifetime (typically 5-7 seconds), temporary addresses
2907 will not be created.
2909 Default: 172800 (2 days)
2911 temp_prefered_lft - INTEGER
2912 Preferred lifetime (in seconds) for temporary addresses. If
2913 temp_prefered_lft is less than the minimum required lifetime (typically
2914 5-7 seconds), the preferred lifetime is the minimum required. If
2915 temp_prefered_lft is greater than temp_valid_lft, the preferred lifetime
2916 is temp_valid_lft.
2918 Default: 86400 (1 day)
2920 keep_addr_on_down - INTEGER
2921 Keep all IPv6 addresses on an interface down event. If set static
2922 global addresses with no expiration time are not flushed.
2924 * >0 : enabled
2925 * 0 : system default
2926 * <0 : disabled
2928 Default: 0 (addresses are removed)
2930 max_desync_factor - INTEGER
2931 Maximum value for DESYNC_FACTOR, which is a random value
2932 that ensures that clients don't synchronize with each
2933 other and generate new addresses at exactly the same time.
2934 value is in seconds.
2936 Default: 600
2938 regen_min_advance - INTEGER
2939 How far in advance (in seconds), at minimum, to create a new temporary
2940 address before the current one is deprecated. This value is added to
2941 the amount of time that may be required for duplicate address detection
2942 to determine when to create a new address. Linux permits setting this
2943 value to less than the default of 2 seconds, but a value less than 2
2944 does not conform to RFC 8981.
2946 Default: 2
2948 regen_max_retry - INTEGER
2949 Number of attempts before give up attempting to generate
2950 valid temporary addresses.
2952 Default: 5
2954 max_addresses - INTEGER
2955 Maximum number of autoconfigured addresses per interface. Setting
2956 to zero disables the limitation. It is not recommended to set this
2957 value too large (or to zero) because it would be an easy way to
2958 crash the kernel by allowing too many addresses to be created.
2960 Default: 16
2962 disable_ipv6 - BOOLEAN
2963 Disable IPv6 operation. If accept_dad is set to 2, this value
2964 will be dynamically set to TRUE if DAD fails for the link-local
2965 address.
2967 Default: FALSE (enable IPv6 operation)
2969 When this value is changed from 1 to 0 (IPv6 is being enabled),
2970 it will dynamically create a link-local address on the given
2971 interface and start Duplicate Address Detection, if necessary.
2973 When this value is changed from 0 to 1 (IPv6 is being disabled),
2974 it will dynamically delete all addresses and routes on the given
2975 interface. From now on it will not possible to add addresses/routes
2976 to the selected interface.
2978 accept_dad - INTEGER
2979 Whether to accept DAD (Duplicate Address Detection).
2981 == ==============================================================
2982 0 Disable DAD
2983 1 Enable DAD (default)
2984 2 Enable DAD, and disable IPv6 operation if MAC-based duplicate
2985 link-local address has been found.
2986 == ==============================================================
2988 DAD operation and mode on a given interface will be selected according
2989 to the maximum value of conf/{all,interface}/accept_dad.
2991 force_tllao - BOOLEAN
2992 Enable sending the target link-layer address option even when
2993 responding to a unicast neighbor solicitation.
2995 Default: FALSE
2997 Quoting from RFC 2461, section 4.4, Target link-layer address:
2999 "The option MUST be included for multicast solicitations in order to
3000 avoid infinite Neighbor Solicitation "recursion" when the peer node
3001 does not have a cache entry to return a Neighbor Advertisements
3002 message. When responding to unicast solicitations, the option can be
3003 omitted since the sender of the solicitation has the correct link-
3004 layer address; otherwise it would not have be able to send the unicast
3005 solicitation in the first place. However, including the link-layer
3006 address in this case adds little overhead and eliminates a potential
3007 race condition where the sender deletes the cached link-layer address
3008 prior to receiving a response to a previous solicitation."
3010 ndisc_notify - BOOLEAN
3011 Define mode for notification of address and device changes.
3013 Possible values:
3015 - 0 (disabled) - do nothing
3016 - 1 (enabled) - Generate unsolicited neighbour advertisements when device is brought
3017 up or hardware address changes.
3019 Default: 0 (disabled)
3021 ndisc_tclass - INTEGER
3022 The IPv6 Traffic Class to use by default when sending IPv6 Neighbor
3023 Discovery (Router Solicitation, Router Advertisement, Neighbor
3024 Solicitation, Neighbor Advertisement, Redirect) messages.
3025 These 8 bits can be interpreted as 6 high order bits holding the DSCP
3026 value and 2 low order bits representing ECN (which you probably want
3027 to leave cleared).
3029 * 0 - (default)
3031 ndisc_evict_nocarrier - BOOLEAN
3032 Clears the neighbor discovery table on NOCARRIER events. This option is
3033 important for wireless devices where the neighbor discovery cache should
3034 not be cleared when roaming between access points on the same network.
3035 In most cases this should remain as the default (1).
3037 Possible values:
3039 - 0 (disabled) - Do not clear neighbor discovery cache on NOCARRIER events.
3040 - 1 (enabled) - Clear neighbor discover cache on NOCARRIER events.
3042 Default: 1 (enabled)
3044 mldv1_unsolicited_report_interval - INTEGER
3045 The interval in milliseconds in which the next unsolicited
3046 MLDv1 report retransmit will take place.
3048 Default: 10000 (10 seconds)
3050 mldv2_unsolicited_report_interval - INTEGER
3051 The interval in milliseconds in which the next unsolicited
3052 MLDv2 report retransmit will take place.
3054 Default: 1000 (1 second)
3056 force_mld_version - INTEGER
3057 * 0 - (default) No enforcement of a MLD version, MLDv1 fallback allowed
3058 * 1 - Enforce to use MLD version 1
3059 * 2 - Enforce to use MLD version 2
3061 suppress_frag_ndisc - INTEGER
3062 Control RFC 6980 (Security Implications of IPv6 Fragmentation
3063 with IPv6 Neighbor Discovery) behavior:
3065 * 1 - (default) discard fragmented neighbor discovery packets
3066 * 0 - allow fragmented neighbor discovery packets
3068 optimistic_dad - BOOLEAN
3069 Whether to perform Optimistic Duplicate Address Detection (RFC 4429).
3071 Optimistic Duplicate Address Detection for the interface will be enabled
3072 if at least one of conf/{all,interface}/optimistic_dad is set to 1,
3073 it will be disabled otherwise.
3075 Possible values:
3077 - 0 (disabled)
3078 - 1 (enabled)
3080 Default: 0 (disabled)
3083 use_optimistic - BOOLEAN
3084 If enabled, do not classify optimistic addresses as deprecated during
3085 source address selection. Preferred addresses will still be chosen
3086 before optimistic addresses, subject to other ranking in the source
3087 address selection algorithm.
3089 This will be enabled if at least one of
3090 conf/{all,interface}/use_optimistic is set to 1, disabled otherwise.
3092 Possible values:
3094 - 0 (disabled)
3095 - 1 (enabled)
3097 Default: 0 (disabled)
3099 stable_secret - IPv6 address
3100 This IPv6 address will be used as a secret to generate IPv6
3101 addresses for link-local addresses and autoconfigured
3102 ones. All addresses generated after setting this secret will
3103 be stable privacy ones by default. This can be changed via the
3104 addrgenmode ip-link. conf/default/stable_secret is used as the
3105 secret for the namespace, the interface specific ones can
3106 overwrite that. Writes to conf/all/stable_secret are refused.
3108 It is recommended to generate this secret during installation
3109 of a system and keep it stable after that.
3111 By default the stable secret is unset.
3113 addr_gen_mode - INTEGER
3114 Defines how link-local and autoconf addresses are generated.
3116 = =================================================================
3117 0 generate address based on EUI64 (default)
3118 1 do no generate a link-local address, use EUI64 for addresses
3119 generated from autoconf
3120 2 generate stable privacy addresses, using the secret from
3121 stable_secret (RFC7217)
3122 3 generate stable privacy addresses, using a random secret if unset
3123 = =================================================================
3125 drop_unicast_in_l2_multicast - BOOLEAN
3126 Drop any unicast IPv6 packets that are received in link-layer
3127 multicast (or broadcast) frames.
3129 Possible values:
3131 - 0 (disabled)
3132 - 1 (enabled)
3134 Default: 0 (disabled)
3136 drop_unsolicited_na - BOOLEAN
3137 Drop all unsolicited neighbor advertisements, for example if there's
3138 a known good NA proxy on the network and such frames need not be used
3139 (or in the case of 802.11, must not be used to prevent attacks.)
3141 Possible values:
3143 - 0 (disabled)
3144 - 1 (enabled)
3146 Default: 0 (disabled).
3148 accept_untracked_na - INTEGER
3149 Define behavior for accepting neighbor advertisements from devices that
3150 are absent in the neighbor cache:
3152 - 0 - (default) Do not accept unsolicited and untracked neighbor
3153 advertisements.
3155 - 1 - Add a new neighbor cache entry in STALE state for routers on
3156 receiving a neighbor advertisement (either solicited or unsolicited)
3157 with target link-layer address option specified if no neighbor entry
3158 is already present for the advertised IPv6 address. Without this knob,
3159 NAs received for untracked addresses (absent in neighbor cache) are
3160 silently ignored.
3162 This is as per router-side behavior documented in RFC9131.
3164 This has lower precedence than drop_unsolicited_na.
3166 This will optimize the return path for the initial off-link
3167 communication that is initiated by a directly connected host, by
3168 ensuring that the first-hop router which turns on this setting doesn't
3169 have to buffer the initial return packets to do neighbor-solicitation.
3170 The prerequisite is that the host is configured to send unsolicited
3171 neighbor advertisements on interface bringup. This setting should be
3172 used in conjunction with the ndisc_notify setting on the host to
3173 satisfy this prerequisite.
3175 - 2 - Extend option (1) to add a new neighbor cache entry only if the
3176 source IP address is in the same subnet as an address configured on
3177 the interface that received the neighbor advertisement.
3179 enhanced_dad - BOOLEAN
3180 Include a nonce option in the IPv6 neighbor solicitation messages used for
3181 duplicate address detection per RFC7527. A received DAD NS will only signal
3182 a duplicate address if the nonce is different. This avoids any false
3183 detection of duplicates due to loopback of the NS messages that we send.
3184 The nonce option will be sent on an interface unless both of
3185 conf/{all,interface}/enhanced_dad are set to FALSE.
3187 Possible values:
3189 - 0 (disabled)
3190 - 1 (enabled)
3192 Default: 1 (enabled)
3194 ``icmp/*``:
3195 ===========
3197 ratelimit - INTEGER
3198 Limit the maximal rates for sending ICMPv6 messages to a particular
3199 peer.
3201 0 to disable any limiting,
3202 otherwise the space between responses in milliseconds.
3204 Default: 100
3206 ratemask - list of comma separated ranges
3207 For ICMPv6 message types matching the ranges in the ratemask, limit
3208 the sending of the message according to ratelimit parameter.
3210 The format used for both input and output is a comma separated
3211 list of ranges (e.g. "0-127,129" for ICMPv6 message type 0 to 127 and
3212 129). Writing to the file will clear all previous ranges of ICMPv6
3213 message types and update the current list with the input.
3215 Refer to: https://www.iana.org/assignments/icmpv6-parameters/icmpv6-parameters.xhtml
3216 for numerical values of ICMPv6 message types, e.g. echo request is 128
3217 and echo reply is 129.
3219 Default: 0-1,3-127 (rate limit ICMPv6 errors except Packet Too Big)
3221 echo_ignore_all - BOOLEAN
3222 If enabled, then the kernel will ignore all ICMP ECHO
3223 requests sent to it over the IPv6 protocol.
3225 Possible values:
3227 - 0 (disabled)
3228 - 1 (enabled)
3230 Default: 0 (disabled)
3232 echo_ignore_multicast - BOOLEAN
3233 If enabled, then the kernel will ignore all ICMP ECHO
3234 requests sent to it over the IPv6 protocol via multicast.
3236 Possible values:
3238 - 0 (disabled)
3239 - 1 (enabled)
3241 Default: 0 (disabled)
3243 echo_ignore_anycast - BOOLEAN
3244 If enabled, then the kernel will ignore all ICMP ECHO
3245 requests sent to it over the IPv6 protocol destined to anycast address.
3247 Possible values:
3249 - 0 (disabled)
3250 - 1 (enabled)
3252 Default: 0 (disabled)
3254 error_anycast_as_unicast - BOOLEAN
3255 If enabled, then the kernel will respond with ICMP Errors
3256 resulting from requests sent to it over the IPv6 protocol destined
3257 to anycast address essentially treating anycast as unicast.
3259 Possible values:
3261 - 0 (disabled)
3262 - 1 (enabled)
3264 Default: 0 (disabled)
3266 xfrm6_gc_thresh - INTEGER
3267 (Obsolete since linux-4.14)
3268 The threshold at which we will start garbage collecting for IPv6
3269 destination cache entries. At twice this value the system will
3270 refuse new allocations.
3273 IPv6 Update by:
3274 Pekka Savola <[email protected]>
3275 YOSHIFUJI Hideaki / USAGI Project <[email protected]>
3278 /proc/sys/net/bridge/* Variables:
3279 =================================
3281 bridge-nf-call-arptables - BOOLEAN
3283 Possible values:
3285 - 0 (disabled) - disable this.
3286 - 1 (enabled) - pass bridged ARP traffic to arptables' FORWARD chain.
3288 Default: 1 (enabled)
3290 bridge-nf-call-iptables - BOOLEAN
3292 Possible values:
3294 - 0 (disabled) - disable this.
3295 - 1 (enabled) - pass bridged IPv4 traffic to iptables' chains.
3297 Default: 1 (enabled)
3299 bridge-nf-call-ip6tables - BOOLEAN
3301 Possible values:
3303 - 0 (disabled) - disable this.
3304 - 1 (enabled) - pass bridged IPv6 traffic to ip6tables' chains.
3306 Default: 1 (enabled)
3308 bridge-nf-filter-vlan-tagged - BOOLEAN
3310 Possible values:
3312 - 0 (disabled) - disable this.
3313 - 1 (enabled) - pass bridged vlan-tagged ARP/IP/IPv6 traffic to {arp,ip,ip6}tables
3315 Default: 0 (disabled)
3317 bridge-nf-filter-pppoe-tagged - BOOLEAN
3319 Possible values:
3321 - 0 (disabled) - disable this.
3322 - 1 (enabled) - pass bridged pppoe-tagged IP/IPv6 traffic to {ip,ip6}tables.
3324 Default: 0 (disabled)
3326 bridge-nf-pass-vlan-input-dev - BOOLEAN
3327 - 1: if bridge-nf-filter-vlan-tagged is enabled, try to find a vlan
3328 interface on the bridge and set the netfilter input device to the
3329 vlan. This allows use of e.g. "iptables -i br0.1" and makes the
3330 REDIRECT target work with vlan-on-top-of-bridge interfaces. When no
3331 matching vlan interface is found, or this switch is off, the input
3332 device is set to the bridge interface.
3334 - 0: disable bridge netfilter vlan interface lookup.
3336 Default: 0
3338 ``proc/sys/net/sctp/*`` Variables:
3339 ==================================
3341 addip_enable - BOOLEAN
3342 Enable or disable extension of Dynamic Address Reconfiguration
3343 (ADD-IP) functionality specified in RFC5061. This extension provides
3344 the ability to dynamically add and remove new addresses for the SCTP
3345 associations.
3347 Possible values:
3349 - 0 (disabled) - disable extension.
3350 - 1 (enabled) - enable extension
3352 Default: 0 (disabled)
3354 pf_enable - INTEGER
3355 Enable or disable pf (pf is short for potentially failed) state. A value
3356 of pf_retrans > path_max_retrans also disables pf state. That is, one of
3357 both pf_enable and pf_retrans > path_max_retrans can disable pf state.
3358 Since pf_retrans and path_max_retrans can be changed by userspace
3359 application, sometimes user expects to disable pf state by the value of
3360 pf_retrans > path_max_retrans, but occasionally the value of pf_retrans
3361 or path_max_retrans is changed by the user application, this pf state is
3362 enabled. As such, it is necessary to add this to dynamically enable
3363 and disable pf state. See:
3364 https://datatracker.ietf.org/doc/draft-ietf-tsvwg-sctp-failover for
3365 details.
3367 Possible values:
3369 - 1: Enable pf.
3370 - 0: Disable pf.
3372 Default: 1
3374 pf_expose - INTEGER
3375 Unset or enable/disable pf (pf is short for potentially failed) state
3376 exposure. Applications can control the exposure of the PF path state
3377 in the SCTP_PEER_ADDR_CHANGE event and access of SCTP_PF-state
3378 transport info via SCTP_GET_PEER_ADDR_INFO sockopt.
3380 Possible values:
3382 - 0: Unset pf state exposure (compatible with old applications). No
3383 event will be sent but the transport info can be queried.
3384 - 1: Disable pf state exposure. No event will be sent and trying to
3385 obtain transport info will return -EACCESS.
3386 - 2: Enable pf state exposure. The event will be sent for a transport
3387 becoming SCTP_PF state and transport info can be obtained.
3389 Default: 0
3391 addip_noauth_enable - BOOLEAN
3392 Dynamic Address Reconfiguration (ADD-IP) requires the use of
3393 authentication to protect the operations of adding or removing new
3394 addresses. This requirement is mandated so that unauthorized hosts
3395 would not be able to hijack associations. However, older
3396 implementations may not have implemented this requirement while
3397 allowing the ADD-IP extension. For reasons of interoperability,
3398 we provide this variable to control the enforcement of the
3399 authentication requirement.
3401 == ===============================================================
3402 1 Allow ADD-IP extension to be used without authentication. This
3403 should only be set in a closed environment for interoperability
3404 with older implementations.
3406 0 Enforce the authentication requirement
3407 == ===============================================================
3409 Default: 0
3411 auth_enable - BOOLEAN
3412 Enable or disable Authenticated Chunks extension. This extension
3413 provides the ability to send and receive authenticated chunks and is
3414 required for secure operation of Dynamic Address Reconfiguration
3415 (ADD-IP) extension.
3417 Possible values:
3419 - 0 (disabled) - disable extension.
3420 - 1 (enabled) - enable extension
3422 Default: 0 (disabled)
3424 prsctp_enable - BOOLEAN
3425 Enable or disable the Partial Reliability extension (RFC3758) which
3426 is used to notify peers that a given DATA should no longer be expected.
3428 Possible values:
3430 - 0 (disabled) - disable extension.
3431 - 1 (enabled) - enable extension
3433 Default: 1 (enabled)
3435 max_burst - INTEGER
3436 The limit of the number of new packets that can be initially sent. It
3437 controls how bursty the generated traffic can be.
3439 Default: 4
3441 association_max_retrans - INTEGER
3442 Set the maximum number for retransmissions that an association can
3443 attempt deciding that the remote end is unreachable. If this value
3444 is exceeded, the association is terminated.
3446 Default: 10
3448 max_init_retransmits - INTEGER
3449 The maximum number of retransmissions of INIT and COOKIE-ECHO chunks
3450 that an association will attempt before declaring the destination
3451 unreachable and terminating.
3453 Default: 8
3455 path_max_retrans - INTEGER
3456 The maximum number of retransmissions that will be attempted on a given
3457 path. Once this threshold is exceeded, the path is considered
3458 unreachable, and new traffic will use a different path when the
3459 association is multihomed.
3461 Default: 5
3463 pf_retrans - INTEGER
3464 The number of retransmissions that will be attempted on a given path
3465 before traffic is redirected to an alternate transport (should one
3466 exist). Note this is distinct from path_max_retrans, as a path that
3467 passes the pf_retrans threshold can still be used. Its only
3468 deprioritized when a transmission path is selected by the stack. This
3469 setting is primarily used to enable fast failover mechanisms without
3470 having to reduce path_max_retrans to a very low value. See:
3471 http://www.ietf.org/id/draft-nishida-tsvwg-sctp-failover-05.txt
3472 for details. Note also that a value of pf_retrans > path_max_retrans
3473 disables this feature. Since both pf_retrans and path_max_retrans can
3474 be changed by userspace application, a variable pf_enable is used to
3475 disable pf state.
3477 Default: 0
3479 ps_retrans - INTEGER
3480 Primary.Switchover.Max.Retrans (PSMR), it's a tunable parameter coming
3481 from section-5 "Primary Path Switchover" in rfc7829. The primary path
3482 will be changed to another active path when the path error counter on
3483 the old primary path exceeds PSMR, so that "the SCTP sender is allowed
3484 to continue data transmission on a new working path even when the old
3485 primary destination address becomes active again". Note this feature
3486 is disabled by initializing 'ps_retrans' per netns as 0xffff by default,
3487 and its value can't be less than 'pf_retrans' when changing by sysctl.
3489 Default: 0xffff
3491 rto_initial - INTEGER
3492 The initial round trip timeout value in milliseconds that will be used
3493 in calculating round trip times. This is the initial time interval
3494 for retransmissions.
3496 Default: 3000
3498 rto_max - INTEGER
3499 The maximum value (in milliseconds) of the round trip timeout. This
3500 is the largest time interval that can elapse between retransmissions.
3502 Default: 60000
3504 rto_min - INTEGER
3505 The minimum value (in milliseconds) of the round trip timeout. This
3506 is the smallest time interval the can elapse between retransmissions.
3508 Default: 1000
3510 hb_interval - INTEGER
3511 The interval (in milliseconds) between HEARTBEAT chunks. These chunks
3512 are sent at the specified interval on idle paths to probe the state of
3513 a given path between 2 associations.
3515 Default: 30000
3517 sack_timeout - INTEGER
3518 The amount of time (in milliseconds) that the implementation will wait
3519 to send a SACK.
3521 Default: 200
3523 valid_cookie_life - INTEGER
3524 The default lifetime of the SCTP cookie (in milliseconds). The cookie
3525 is used during association establishment.
3527 Default: 60000
3529 cookie_preserve_enable - BOOLEAN
3530 Enable or disable the ability to extend the lifetime of the SCTP cookie
3531 that is used during the establishment phase of SCTP association
3533 Possible values:
3535 - 0 (disabled) - disable.
3536 - 1 (enabled) - enable cookie lifetime extension.
3538 Default: 1 (enabled)
3540 cookie_hmac_alg - STRING
3541 Select the hmac algorithm used when generating the cookie value sent by
3542 a listening sctp socket to a connecting client in the INIT-ACK chunk.
3543 Valid values are:
3545 * sha256
3546 * none
3548 Default: sha256
3550 rcvbuf_policy - INTEGER
3551 Determines if the receive buffer is attributed to the socket or to
3552 association. SCTP supports the capability to create multiple
3553 associations on a single socket. When using this capability, it is
3554 possible that a single stalled association that's buffering a lot
3555 of data may block other associations from delivering their data by
3556 consuming all of the receive buffer space. To work around this,
3557 the rcvbuf_policy could be set to attribute the receiver buffer space
3558 to each association instead of the socket. This prevents the described
3559 blocking.
3561 - 1: rcvbuf space is per association
3562 - 0: rcvbuf space is per socket
3564 Default: 0
3566 sndbuf_policy - INTEGER
3567 Similar to rcvbuf_policy above, this applies to send buffer space.
3569 - 1: Send buffer is tracked per association
3570 - 0: Send buffer is tracked per socket.
3572 Default: 0
3574 sctp_mem - vector of 3 INTEGERs: min, pressure, max
3575 Number of pages allowed for queueing by all SCTP sockets.
3577 * min: Below this number of pages SCTP is not bothered about its
3578 memory usage. When amount of memory allocated by SCTP exceeds
3579 this number, SCTP starts to moderate memory usage.
3580 * pressure: This value was introduced to follow format of tcp_mem.
3581 * max: Maximum number of allowed pages.
3583 Default is calculated at boot time from amount of available memory.
3585 sctp_rmem - vector of 3 INTEGERs: min, default, max
3586 Only the first value ("min") is used, "default" and "max" are
3587 ignored.
3589 * min: Minimal size of receive buffer used by SCTP socket.
3590 It is guaranteed to each SCTP socket (but not association) even
3591 under moderate memory pressure.
3593 Default: 4K
3595 sctp_wmem - vector of 3 INTEGERs: min, default, max
3596 Only the first value ("min") is used, "default" and "max" are
3597 ignored.
3599 * min: Minimum size of send buffer that can be used by SCTP sockets.
3600 It is guaranteed to each SCTP socket (but not association) even
3601 under moderate memory pressure.
3603 Default: 4K
3605 addr_scope_policy - INTEGER
3606 Control IPv4 address scoping (see
3607 https://datatracker.ietf.org/doc/draft-stewart-tsvwg-sctp-ipv4/00/
3608 for details).
3610 - 0 - Disable IPv4 address scoping
3611 - 1 - Enable IPv4 address scoping
3612 - 2 - Follow draft but allow IPv4 private addresses
3613 - 3 - Follow draft but allow IPv4 link local addresses
3615 Default: 1
3617 udp_port - INTEGER
3618 The listening port for the local UDP tunneling sock. Normally it's
3619 using the IANA-assigned UDP port number 9899 (sctp-tunneling).
3621 This UDP sock is used for processing the incoming UDP-encapsulated
3622 SCTP packets (from RFC6951), and shared by all applications in the
3623 same net namespace. This UDP sock will be closed when the value is
3624 set to 0.
3626 The value will also be used to set the src port of the UDP header
3627 for the outgoing UDP-encapsulated SCTP packets. For the dest port,
3628 please refer to 'encap_port' below.
3630 Default: 0
3632 encap_port - INTEGER
3633 The default remote UDP encapsulation port.
3635 This value is used to set the dest port of the UDP header for the
3636 outgoing UDP-encapsulated SCTP packets by default. Users can also
3637 change the value for each sock/asoc/transport by using setsockopt.
3638 For further information, please refer to RFC6951.
3640 Note that when connecting to a remote server, the client should set
3641 this to the port that the UDP tunneling sock on the peer server is
3642 listening to and the local UDP tunneling sock on the client also
3643 must be started. On the server, it would get the encap_port from
3644 the incoming packet's source port.
3646 Default: 0
3648 plpmtud_probe_interval - INTEGER
3649 The time interval (in milliseconds) for the PLPMTUD probe timer,
3650 which is configured to expire after this period to receive an
3651 acknowledgment to a probe packet. This is also the time interval
3652 between the probes for the current pmtu when the probe search
3653 is done.
3655 PLPMTUD will be disabled when 0 is set, and other values for it
3656 must be >= 5000.
3658 Default: 0
3660 reconf_enable - BOOLEAN
3661 Enable or disable extension of Stream Reconfiguration functionality
3662 specified in RFC6525. This extension provides the ability to "reset"
3663 a stream, and it includes the Parameters of "Outgoing/Incoming SSN
3664 Reset", "SSN/TSN Reset" and "Add Outgoing/Incoming Streams".
3666 Possible values:
3668 - 0 (disabled) - Disable extension.
3669 - 1 (enabled) - Enable extension.
3671 Default: 0 (disabled)
3673 intl_enable - BOOLEAN
3674 Enable or disable extension of User Message Interleaving functionality
3675 specified in RFC8260. This extension allows the interleaving of user
3676 messages sent on different streams. With this feature enabled, I-DATA
3677 chunk will replace DATA chunk to carry user messages if also supported
3678 by the peer. Note that to use this feature, one needs to set this option
3679 to 1 and also needs to set socket options SCTP_FRAGMENT_INTERLEAVE to 2
3680 and SCTP_INTERLEAVING_SUPPORTED to 1.
3682 Possible values:
3684 - 0 (disabled) - Disable extension.
3685 - 1 (enabled) - Enable extension.
3687 Default: 0 (disabled)
3689 ecn_enable - BOOLEAN
3690 Control use of Explicit Congestion Notification (ECN) by SCTP.
3691 Like in TCP, ECN is used only when both ends of the SCTP connection
3692 indicate support for it. This feature is useful in avoiding losses
3693 due to congestion by allowing supporting routers to signal congestion
3694 before having to drop packets.
3696 Possible values:
3698 - 0 (disabled) - Disable ecn.
3699 - 1 (enabled) - Enable ecn.
3701 Default: 1 (enabled)
3703 l3mdev_accept - BOOLEAN
3704 Enabling this option allows a "global" bound socket to work
3705 across L3 master domains (e.g., VRFs) with packets capable of
3706 being received regardless of the L3 domain in which they
3707 originated. Only valid when the kernel was compiled with
3708 CONFIG_NET_L3_MASTER_DEV.
3710 Possible values:
3712 - 0 (disabled)
3713 - 1 (enabled)
3715 Default: 1 (enabled)
3718 ``/proc/sys/net/core/*``
3719 ========================
3721 Please see: Documentation/admin-guide/sysctl/net.rst for descriptions of these entries.
3724 ``/proc/sys/net/unix/*``
3725 ========================
3727 max_dgram_qlen - INTEGER
3728 The maximum length of dgram socket receive queue
3730 Default: 10

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

문서 범위

1-6

이 문서는 Linux IP 계층이 `/proc/sys/net` 아래에 제공하는 sysctl 변수를 설명합니다. IPv4·IPv6의 전달과 경로 선택뿐 아니라 TCP, UDP, RAW, CIPSO, ICMP, IGMP/MLD, bridge netfilter, SCTP, Unix datagram queue까지 네트워크 namespace별 조정 지점을 한데 모은 참조 문서입니다.

.. SPDX-License-Identifier: GPL-2.0

=========
IP Sysctl
=========

IPv4 전달과 경로 MTU

7-88

`ip_forward`는 interface 사이의 IPv4 packet 전달을 켭니다. 값을 바꾸면 RFC 1122 host 또는 RFC 1812 router 기본값으로 IPv4 구성 변수가 다시 초기화되므로 단순한 한 bit 변경으로 보아서는 안 됩니다. 기본은 0이고 `ip_default_ttl`의 기본 TTL은 64입니다.

`ip_no_pmtu_disc`는 Path MTU Discovery 동작을 0~3으로 고릅니다. 0은 정상 PMTUD, 1은 수신 PMTU를 `min_pmtu`로 제한, 2는 들어온 PMTU를 폐기하고 outgoing packet의 DF도 세우지 않는 모드, 3은 TCP·SCTP처럼 sequence로 검증할 수 있는 protocol만 hardened PMTU update를 허용하는 모드입니다. mode 2와 3은 namespace 전체에 적용하면 다른 traffic을 망가뜨릴 수 있어 보호가 필요한 제한된 환경에만 씁니다. `min_pmtu` 기본값은 552입니다.

`ip_forward_use_pmtu`는 forwarding 시 protocol PMTU 정보를 신뢰할지 정합니다. protocol spoofing 위험 때문에 기본은 0입니다. `fwmark_reflect`를 켜면 socket과 연결되지 않은 TCP RST, ICMP echo reply 같은 kernel 생성 응답이 원 packet의 fwmark를 이어받습니다.

/proc/sys/net/ipv4/* Variables
==============================

ip_forward - BOOLEAN
        Forward Packets between interfaces.

        This variable is special, its change resets all configuration
        parameters to their default state (RFC1122 for hosts, RFC1812
        for routers)

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

ip_default_ttl - INTEGER
        Default value of TTL field (Time To Live) for outgoing (but not
        forwarded) IP packets. Should be between 1 and 255 inclusive.
        Default: 64 (as recommended by RFC1700)

ip_no_pmtu_disc - INTEGER
        Disable Path MTU Discovery. If enabled in mode 1 and a
        fragmentation-required ICMP is received, the PMTU to this
        destination will be set to the smallest of the old MTU to
        this destination and min_pmtu (see below). You will need
        to raise min_pmtu to the smallest interface MTU on your system
        manually if you want to avoid locally generated fragments.

        In mode 2 incoming Path MTU Discovery messages will be
        discarded. Outgoing frames are handled the same as in mode 1,
        implicitly setting IP_PMTUDISC_DONT on every created socket.

        Mode 3 is a hardened pmtu discover mode. The kernel will only
        accept fragmentation-needed errors if the underlying protocol
        can verify them besides a plain socket lookup. Current
        protocols for which pmtu events will be honored are TCP and
        SCTP as they verify e.g. the sequence number or the
        association. This mode should not be enabled globally but is
        only intended to secure e.g. name servers in namespaces where
        TCP path mtu must still work but path MTU information of other
        protocols should be discarded. If enabled globally this mode
        could break other protocols.

        Possible values: 0-3

        Default: FALSE

min_pmtu - INTEGER
        default 552 - minimum Path MTU. Unless this is changed manually,
        each cached pmtu will never be lower than this setting.

ip_forward_use_pmtu - BOOLEAN
        By default we don't trust protocol path MTUs while forwarding
        because they could be easily forged and can lead to unwanted
        fragmentation by the router.
        You only need to enable this if you have user-space software
        which tries to discover path mtus by itself and depends on the
        kernel honoring this information. This is normally not the
        case.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

fwmark_reflect - BOOLEAN
        Controls the fwmark of kernel-generated IPv4 reply packets that are not
        associated with a socket for example, TCP RSTs or ICMP echo replies).
        If disabled, these packets have a fwmark of zero. If enabled, they have the
        fwmark of the packet they are replying to.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

IPv4 multipath hash와 FIB 동기화

89-174

`fib_multipath_use_neigh`는 ECMP nexthop 선택에 neighbor 상태를 반영합니다. 기본 0은 route group 안의 nexthop을 neighbor 상태와 무관하게 사용할 수 있고, 1은 실패한 neighbor를 피합니다.

`fib_multipath_hash_policy`는 multipath hash를 0=L3, 1=L4 5-tuple, 2=encapsulation이 있으면 inner L3, 3=사용자 지정으로 고릅니다. mode 3의 `fib_multipath_hash_fields` bitmask는 source/destination IP(0x1/0x2), protocol(0x4), source/destination port(0x10/0x20), inner IP·protocol·port(0x40~0x800)를 조합하며 기본 0x7은 양쪽 IP와 protocol입니다. `fib_multipath_hash_seed`는 namespace별 seed이고 0이면 random입니다. 같은 seed라도 kernel version 사이에 hash 결과가 안정적이라고 가정하면 안 됩니다.

`fib_sync_mem`은 route entry를 fib notifier에 동기화하면서 `synchronize_rcu()`를 호출하기 전 쌓을 수 있는 dirty memory를 byte 단위로 제한합니다. 기본 512 KiB, 최소 64 KiB, 최대 64 MiB입니다. `ip_forward_update_priority`가 1이면 forwarding 중 IPv4 header priority를 TOS에서 갱신합니다.

fib_multipath_use_neigh - BOOLEAN
        Use status of existing neighbor entry when determining nexthop for
        multipath routes. If disabled, neighbor information is not used and
        packets could be directed to a failed nexthop. Only valid for kernels
        built with CONFIG_IP_ROUTE_MULTIPATH enabled.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

fib_multipath_hash_policy - INTEGER
        Controls which hash policy to use for multipath routes. Only valid
        for kernels built with CONFIG_IP_ROUTE_MULTIPATH enabled.

        Default: 0 (Layer 3)

        Possible values:

        - 0 - Layer 3
        - 1 - Layer 4
        - 2 - Layer 3 or inner Layer 3 if present
        - 3 - Custom multipath hash. Fields used for multipath hash calculation
          are determined by fib_multipath_hash_fields sysctl

fib_multipath_hash_fields - UNSIGNED INTEGER
        When fib_multipath_hash_policy is set to 3 (custom multipath hash), the
        fields used for multipath hash calculation are determined by this
        sysctl.

        This value is a bitmask which enables various fields for multipath hash
        calculation.

        Possible fields are:

        ====== ============================
        0x0001 Source IP address
        0x0002 Destination IP address
        0x0004 IP protocol
        0x0008 Unused (Flow Label)
        0x0010 Source port
        0x0020 Destination port
        0x0040 Inner source IP address
        0x0080 Inner destination IP address
        0x0100 Inner IP protocol
        0x0200 Inner Flow Label
        0x0400 Inner source port
        0x0800 Inner destination port
        ====== ============================

        Default: 0x0007 (source IP, destination IP and IP protocol)

fib_multipath_hash_seed - UNSIGNED INTEGER
        The seed value used when calculating hash for multipath routes. Applies
        to both IPv4 and IPv6 datapath. Only present for kernels built with
        CONFIG_IP_ROUTE_MULTIPATH enabled.

        When set to 0, the seed value used for multipath routing defaults to an
        internal random-generated one.

        The actual hashing algorithm is not specified -- there is no guarantee
        that a next hop distribution effected by a given seed will keep stable
        across kernel versions.

        Default: 0 (random)

fib_sync_mem - UNSIGNED INTEGER
        Amount of dirty memory from fib entries that can be backlogged before
        synchronize_rcu is forced.

        Default: 512kB   Minimum: 64kB   Maximum: 64MB

ip_forward_update_priority - INTEGER
        Whether to update SKB priority from "TOS" field in IPv4 header after it
        is forwarded. The new SKB priority is mapped from TOS field value
        according to an rt_tos2priority table (see e.g. man tc-prio).

        Default: 1 (Update priority.)

        Possible values:

        - 0 - Do not update priority.
        - 1 - Update priority.

Route cache, neighbor queue와 offload 알림

175-266

IPv4의 `route/max_size`는 3.6부터 route cache가 없어져 더 이상 의미가 없고 IPv6에는 별도 의미가 있습니다. neighbor cache의 `gc_thresh1`, `gc_thresh2`, `gc_thresh3`는 각각 최소 보존 수, 5초 이상 cache가 유지될 때 허용하는 soft maximum, 강제 GC 기준입니다. IPv6에서는 `gc_thresh3`가 동작에 필요한 최소 permanent entry보다 작아지지 않습니다.

`unres_qlen_bytes`는 address resolution이 끝나지 않은 neighbor마다 보류할 packet의 총 byte를 제한하고 기본값은 `SK_WMEM_MAX`와 같습니다. `unres_qlen`은 호환용 packet 수 제한으로 기본 101이지만 packet 메모리 overhead를 계산하지 못하므로 byte 제한이 더 정확합니다. `interval_probe_time_ms`는 NTF_USE로 queue된 probe의 최소 간격이며 기본 5000 ms입니다.

`mtu_expires`는 PMTU cache 수명, `min_adv_mss`는 첫 MSS 광고의 최소값입니다. `fib_notify_on_flag_change`는 route의 `RTM_F_OFFLOAD`, `RTM_F_TRAP`, `RTM_F_OFFLOAD_FAILED` flag 변화 때 `RTM_NEWROUTE`를 보낼지 정하며 0=없음, 1=모두, 2=offload 실패만입니다. kernel 설치 ACK는 hardware 설치까지 보장하지 않으므로 이 알림이 offload 상태 추적에 쓰입니다.

route/max_size - INTEGER
        Maximum number of routes allowed in the kernel.  Increase
        this when using large numbers of interfaces and/or routes.

        From linux kernel 3.6 onwards, this is deprecated for ipv4
        as route cache is no longer used.

        From linux kernel 6.3 onwards, this is deprecated for ipv6
        as garbage collection manages cached route entries.

neigh/default/gc_thresh1 - INTEGER
        Minimum number of entries to keep.  Garbage collector will not
        purge entries if there are fewer than this number.

        Default: 128

neigh/default/gc_thresh2 - INTEGER
        Threshold when garbage collector becomes more aggressive about
        purging entries. Entries older than 5 seconds will be cleared
        when over this number.

        Default: 512

neigh/default/gc_thresh3 - INTEGER
        Maximum number of non-PERMANENT neighbor entries allowed.  Increase
        this when using large numbers of interfaces and when communicating
        with large numbers of directly-connected peers.

        Default: 1024

neigh/default/unres_qlen_bytes - INTEGER
        The maximum number of bytes which may be used by packets
        queued for each        unresolved address by other network layers.
        (added in linux 3.3)

        Setting negative value is meaningless and will return error.

        Default: SK_WMEM_DEFAULT, (same as net.core.wmem_default).

                Exact value depends on architecture and kernel options,
                but should be enough to allow queuing 256 packets
                of medium size.

neigh/default/unres_qlen - INTEGER
        The maximum number of packets which may be queued for each
        unresolved address by other network layers.

        (deprecated in linux 3.3) : use unres_qlen_bytes instead.

        Prior to linux 3.3, the default value is 3 which may cause
        unexpected packet loss. The current default value is calculated
        according to default value of unres_qlen_bytes and true size of
        packet.

        Default: 101

neigh/default/interval_probe_time_ms - INTEGER
        The probe interval for neighbor entries with NTF_MANAGED flag,
        the min value is 1.

        Default: 5000

mtu_expires - INTEGER
        Time, in seconds, that cached PMTU information is kept.

min_adv_mss - INTEGER
        The advertised MSS depends on the first hop route MTU, but will
        never be lower than this setting.

fib_notify_on_flag_change - INTEGER
        Whether to emit RTM_NEWROUTE notifications whenever RTM_F_OFFLOAD/
        RTM_F_TRAP/RTM_F_OFFLOAD_FAILED flags are changed.

        After installing a route to the kernel, user space receives an
        acknowledgment, which means the route was installed in the kernel,
        but not necessarily in hardware.
        It is also possible for a route already installed in hardware to change
        its action and therefore its flags. For example, a host route that is
        trapping packets can be "promoted" to perform decapsulation following
        the installation of an IPinIP/VXLAN tunnel.
        The notifications will indicate to user-space the state of the route.

        Default: 0 (Do not emit notifications.)

        Possible values:

        - 0 - Do not emit notifications.
        - 1 - Emit notifications.
        - 2 - Emit notifications only for RTM_F_OFFLOAD_FAILED flag change.

IP Fragmentation:

IPv4 fragment와 directed broadcast

267-308

`ipfrag_high_thresh`는 IPv4 fragment 재조립에 쓸 최대 memory이고 `ipfrag_low_thresh`는 4.17부터 폐기되어 효과가 없습니다. `ipfrag_time`은 fragment를 보관하는 초 단위 시간입니다.

`ipfrag_max_dist`는 같은 source의 fragment 사이에서 다른 IP fragment가 몇 개까지 도착할 수 있는지를 제한해 fragment ID 재사용으로 잘못 합쳐지는 것을 줄입니다. 0은 검사를 끄고 기본 64는 reorder가 큰 환경에서 조기 폐기를 일으킬 수 있습니다. source별 분리 경로가 흔하면 더 크게 조정할 수 있지만 지나치게 크면 오조립 위험이 커집니다. `bc_forwarding`은 directed broadcast forwarding이며 `conf/all`과 ingress interface가 모두 켜져야 하고 기본은 0입니다.

ipfrag_high_thresh - LONG INTEGER
        Maximum memory used to reassemble IP fragments.

ipfrag_low_thresh - LONG INTEGER
        (Obsolete since linux-4.17)
        Maximum memory used to reassemble IP fragments before the kernel
        begins to remove incomplete fragment queues to free up resources.
        The kernel still accepts new fragments for defragmentation.

ipfrag_time - INTEGER
        Time in seconds to keep an IP fragment in memory.

ipfrag_max_dist - INTEGER
        ipfrag_max_dist is a non-negative integer value which defines the
        maximum "disorder" which is allowed among fragments which share a
        common IP source address. Note that reordering of packets is
        not unusual, but if a large number of fragments arrive from a source
        IP address while a particular fragment queue remains incomplete, it
        probably indicates that one or more fragments belonging to that queue
        have been lost. When ipfrag_max_dist is positive, an additional check
        is done on fragments before they are added to a reassembly queue - if
        ipfrag_max_dist (or more) fragments have arrived from a particular IP
        address between additions to any IP fragment queue using that source
        address, it's presumed that one or more fragments in the queue are
        lost. The existing fragment queue will be dropped, and a new one
        started. An ipfrag_max_dist value of zero disables this check.

        Using a very small value, e.g. 1 or 2, for ipfrag_max_dist can
        result in unnecessarily dropping fragment queues when normal
        reordering of packets occurs, which could lead to poor application
        performance. Using a very large value, e.g. 50000, increases the
        likelihood of incorrectly reassembling IP fragments that originate
        from different IP datagrams, which could result in data corruption.
        Default: 64

bc_forwarding - INTEGER
        bc_forwarding enables the feature described in rfc1812#section-5.3.5.2
        and rfc2644. It allows the router to forward directed broadcast.
        To enable this feature, the 'all' entry and the input interface entry
        should be set to 1.
        Default: 0

INET peer 저장소

309-329

INET peer storage는 route 사이에서 오래 유지해야 하는 peer별 정보를 보관합니다. `inet_peer_threshold`를 넘으면 entry를 공격적으로 폐기하며 기본 threshold는 memory 규모에 따라 정해집니다. `inet_peer_minttl`과 `inet_peer_maxttl`은 각각 peer entry의 최소·최대 수명으로, 기본은 120초와 600초입니다. cache pressure가 커질수록 실제 TTL은 두 값 사이에서 짧아집니다.

INET peer storage
=================

inet_peer_threshold - INTEGER
        The approximate size of the storage.  Starting from this threshold
        entries will be thrown aggressively.  This threshold also determines
        entries' time-to-live and time intervals between garbage collection
        passes.  More entries, less time-to-live, less GC interval.

inet_peer_minttl - INTEGER
        Minimum time-to-live of entries.  Should be enough to cover fragment
        time-to-live on the reassembling side.  This minimum time-to-live  is
        guaranteed if the pool size is less than inet_peer_threshold.
        Measured in seconds.

inet_peer_maxttl - INTEGER
        Maximum time-to-live of entries.  Unused entries will expire after
        this period of time if there is no memory pressure on the pool (i.e.
        when the number of entries in the pool is very small).
        Measured in seconds.

TCP listen, window와 congestion control

330-443

`somaxconn`은 listen backlog의 상한으로 기본 4096입니다. `tcp_abort_on_overflow`를 켜면 application이 accept를 따라가지 못할 때 연결을 조용히 재시도하게 두는 대신 reset하지만 client에 해를 줄 수 있습니다. `tcp_adv_win_scale`는 6.6부터 폐기되었고, `tcp_app_win`은 receive window 중 application buffer로 남길 비율을 정합니다.

`tcp_allowed_congestion_control`은 unprivileged process가 고를 수 있는 congestion algorithm 목록이고 `tcp_available_congestion_control`은 등록된 전체 목록, `tcp_congestion_control`은 새 연결의 기본 algorithm입니다. `tcp_autocorking`은 같은 flow의 앞 packet이 qdisc나 transmit queue에 있을 때 작은 write를 합칩니다.

`tcp_base_mss`는 MTU probing의 search_low, `tcp_mtu_probe_floor`는 마지막 probe 하한, `tcp_min_snd_mss`는 수신 MSS에 관계없이 outgoing segment가 지켜야 할 최소값입니다. `tcp_dsack`은 RFC 2883 duplicate SACK을, `tcp_early_retrans`는 Tail Loss Probe를 포함한 early retransmit 수준을 제어하며 값 3이 기본 TLP 모드입니다.

TCP variables
=============

somaxconn - INTEGER
        Limit of socket listen() backlog, known in userspace as SOMAXCONN.
        Defaults to 4096. (Was 128 before linux-5.4)
        See also tcp_max_syn_backlog for additional tuning for TCP sockets.

tcp_abort_on_overflow - BOOLEAN
        If listening service is too slow to accept new connections,
        reset them. Default state is FALSE. It means that if overflow
        occurred due to a burst, connection will recover. Enable this
        option _only_ if you are really sure that listening daemon
        cannot be tuned to accept connections faster. Enabling this
        option can harm clients of your server.

tcp_adv_win_scale - INTEGER
        Obsolete since linux-6.6
        Count buffering overhead as bytes/2^tcp_adv_win_scale
        (if tcp_adv_win_scale > 0) or bytes-bytes/2^(-tcp_adv_win_scale),
        if it is <= 0.

        Possible values are [-31, 31], inclusive.

        Default: 1

tcp_allowed_congestion_control - STRING
        Show/set the congestion control choices available to non-privileged
        processes. The list is a subset of those listed in
        tcp_available_congestion_control.

        Default is "reno" and the default setting (tcp_congestion_control).

tcp_app_win - INTEGER
        Reserve max(window/2^tcp_app_win, mss) of window for application
        buffer. Value 0 is special, it means that nothing is reserved.

        Possible values are [0, 31], inclusive.

        Default: 31

tcp_autocorking - BOOLEAN
        Enable TCP auto corking :
        When applications do consecutive small write()/sendmsg() system calls,
        we try to coalesce these small writes as much as possible, to lower
        total amount of sent packets. This is done if at least one prior
        packet for the flow is waiting in Qdisc queues or device transmit
        queue. Applications can still use TCP_CORK for optimal behavior
        when they know how/when to uncork their sockets.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_available_congestion_control - STRING
        Shows the available congestion control choices that are registered.
        More congestion control algorithms may be available as modules,
        but not loaded.

tcp_base_mss - INTEGER
        The initial value of search_low to be used by the packetization layer
        Path MTU discovery (MTU probing).  If MTU probing is enabled,
        this is the initial MSS used by the connection.

tcp_mtu_probe_floor - INTEGER
        If MTU probing is enabled this caps the minimum MSS used for search_low
        for the connection.

        Default : 48

tcp_min_snd_mss - INTEGER
        TCP SYN and SYNACK messages usually advertise an ADVMSS option,
        as described in RFC 1122 and RFC 6691.

        If this ADVMSS option is smaller than tcp_min_snd_mss,
        it is silently capped to tcp_min_snd_mss.

        Default : 48 (at least 8 bytes of payload per segment)

tcp_congestion_control - STRING
        Set the congestion control algorithm to be used for new
        connections. The algorithm "reno" is always available, but
        additional choices may be available based on kernel configuration.
        Default is set as part of kernel configuration.
        For passive connections, the listener congestion control choice
        is inherited.

        [see setsockopt(listenfd, SOL_TCP, TCP_CONGESTION, "name" ...) ]

tcp_dsack - BOOLEAN
        Allows TCP to send "duplicate" SACKs.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_early_retrans - INTEGER
        Tail loss probe (TLP) converts RTOs occurring due to tail
        losses into fast recovery (RFC8985). Note that
        TLP requires RACK to function properly (see tcp_recovery below)

        Possible values:

                - 0 disables TLP
                - 3 or 4 enables TLP

        Default: 3

TCP ECN, 종료와 keepalive

444-603

`tcp_ecn`은 incoming/outgoing ECN 협상을 0~5로 세분합니다. 0은 ECN을 요청하지도 수락하지도 않고, 1/2는 classic ECN을 양방향 또는 incoming 위주로, 3/4/5는 AccECN을 양방향·incoming·fallback 조합으로 사용합니다. 기본 2는 outgoing SYN에서는 요청하지 않지만 peer 요청은 받아들입니다. `tcp_ecn_option`은 AccECN option을 0=없음, 1=필요할 때, 2=가능하면 매 packet으로 보내고 `tcp_ecn_option_beacon`은 RTT당 최소 beacon 수입니다. `tcp_ecn_fallback`은 RFC 3168 fallback을 켭니다.

`tcp_fack`은 효과가 없는 legacy 변수입니다. `tcp_fin_timeout`은 orphaned FIN_WAIT_2 연결의 보존 시간(기본 60초), `tcp_frto`는 F-RTO 복구, `tcp_fwmark_accept`는 passive connection의 fwmark 상속을 제어합니다. `tcp_invalid_ratelimit`은 invalid duplicate ACK에 대한 challenge ACK 사이의 최소 간격으로 기본 500 ms이며 0은 제한을 끕니다.

keepalive는 `tcp_keepalive_time` 7200초, `tcp_keepalive_probes` 9회, `tcp_keepalive_intvl` 75초가 기본입니다. socket option으로 override할 수 있습니다. `tcp_l3mdev_accept`는 global bound TCP socket이 VRF 같은 L3 master domain을 가로질러 받을 수 있게 하고 기본 0입니다. `tcp_low_latency`는 4.14부터 효과가 없는 legacy 설정입니다.

tcp_ecn - INTEGER
        Control use of Explicit Congestion Notification (ECN) by TCP.
        ECN is used only when both ends of the TCP connection indicate support
        for it. This feature is useful in avoiding losses due to congestion by
        allowing supporting routers to signal congestion before having to drop
        packets. A host that supports ECN both sends ECN at the IP layer and
        feeds back ECN at the TCP layer. The highest variant of ECN feedback
        that both peers support is chosen by the ECN negotiation (Accurate ECN,
        ECN, or no ECN).

        The highest negotiated variant for incoming connection requests
        and the highest variant requested by outgoing connection
        attempts:

        ===== ==================== ====================
        Value Incoming connections Outgoing connections
        ===== ==================== ====================
        0     No ECN               No ECN
        1     ECN                  ECN
        2     ECN                  No ECN
        3     AccECN               AccECN
        4     AccECN               ECN
        5     AccECN               No ECN
        ===== ==================== ====================

        Default: 2

tcp_ecn_option - INTEGER
        Control Accurate ECN (AccECN) option sending when AccECN has been
        successfully negotiated during handshake. Send logic inhibits
        sending AccECN options regarless of this setting when no AccECN
        option has been seen for the reverse direction.

        Possible values are:

        = ============================================================
        0 Never send AccECN option. This also disables sending AccECN
          option in SYN/ACK during handshake.
        1 Send AccECN option sparingly according to the minimum option
          rules outlined in draft-ietf-tcpm-accurate-ecn.
        2 Send AccECN option on every packet whenever it fits into TCP
          option space.
        = ============================================================

        Default: 2

tcp_ecn_option_beacon - INTEGER
        Control Accurate ECN (AccECN) option sending frequency per RTT and it
        takes effect only when tcp_ecn_option is set to 2.

        Default: 3 (AccECN will be send at least 3 times per RTT)

tcp_ecn_fallback - BOOLEAN
        If the kernel detects that ECN connection misbehaves, enable fall
        back to non-ECN. Currently, this knob implements the fallback
        from RFC3168, section 6.1.1.1., but we reserve that in future,
        additional detection mechanisms could be implemented under this
        knob. The value        is not used, if tcp_ecn or per route (or congestion
        control) ECN settings are disabled.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_fack - BOOLEAN
        This is a legacy option, it has no effect anymore.

tcp_fin_timeout - INTEGER
        The length of time an orphaned (no longer referenced by any
        application) connection will remain in the FIN_WAIT_2 state
        before it is aborted at the local end.  While a perfectly
        valid "receive only" state for an un-orphaned connection, an
        orphaned connection in FIN_WAIT_2 state could otherwise wait
        forever for the remote to close its end of the connection.

        Cf. tcp_max_orphans

        Default: 60 seconds

tcp_frto - INTEGER
        Enables Forward RTO-Recovery (F-RTO) defined in RFC5682.
        F-RTO is an enhanced recovery algorithm for TCP retransmission
        timeouts.  It is particularly beneficial in networks where the
        RTT fluctuates (e.g., wireless). F-RTO is sender-side only
        modification. It does not require any support from the peer.

        By default it's enabled with a non-zero value. 0 disables F-RTO.

tcp_fwmark_accept - BOOLEAN
        If enabled, incoming connections to listening sockets that do not have a
        socket mark will set the mark of the accepting socket to the fwmark of
        the incoming SYN packet. This will cause all packets on that connection
        (starting from the first SYNACK) to be sent with that fwmark. The
        listening socket's mark is unchanged. Listening sockets that already
        have a fwmark set via setsockopt(SOL_SOCKET, SO_MARK, ...) are
        unaffected.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_invalid_ratelimit - INTEGER
        Limit the maximal rate for sending duplicate acknowledgments
        in response to incoming TCP packets that are for an existing
        connection but that are invalid due to any of these reasons:

          (a) out-of-window sequence number,
          (b) out-of-window acknowledgment number, or
          (c) PAWS (Protection Against Wrapped Sequence numbers) check failure

        This can help mitigate simple "ack loop" DoS attacks, wherein
        a buggy or malicious middlebox or man-in-the-middle can
        rewrite TCP header fields in manner that causes each endpoint
        to think that the other is sending invalid TCP segments, thus
        causing each side to send an unterminating stream of duplicate
        acknowledgments for invalid segments.

        Using 0 disables rate-limiting of dupacks in response to
        invalid segments; otherwise this value specifies the minimal
        space between sending such dupacks, in milliseconds.

        Default: 500 (milliseconds).

tcp_keepalive_time - INTEGER
        How often TCP sends out keepalive messages when keepalive is enabled.
        Default: 2hours.

tcp_keepalive_probes - INTEGER
        How many keepalive probes TCP sends out, until it decides that the
        connection is broken. Default value: 9.

tcp_keepalive_intvl - INTEGER
        How frequently the probes are send out. Multiplied by
        tcp_keepalive_probes it is time to kill not responding connection,
        after probes started. Default value: 75sec i.e. connection
        will be aborted after ~11 minutes of retries.

tcp_l3mdev_accept - BOOLEAN
        Enables child sockets to inherit the L3 master device index.
        Enabling this option allows a "global" listen socket to work
        across L3 master domains (e.g., VRFs) with connected sockets
        derived from the listen socket to be bound to the L3 domain in
        which the packets originated. Only valid when the kernel was
        compiled with CONFIG_NET_L3_MASTER_DEV.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_low_latency - BOOLEAN
        This is a legacy option, it has no effect anymore.

TCP memory, PMTU와 metric cache

604-730

`tcp_max_orphans`는 user file handle에 붙지 않은 TCP socket 수의 상한입니다. orphan 하나가 대략 64 KiB의 unswappable memory를 쓸 수 있어 과도한 증가는 DoS 위험을 키웁니다. `tcp_max_syn_backlog`은 아직 client ACK를 받지 못한 remembered connection request 수, `tcp_max_tw_buckets`는 TIME_WAIT socket 수의 전역 상한입니다.

`tcp_mem`의 min/pressure/max는 모든 TCP socket queue가 쓸 수 있는 page 수입니다. `tcp_min_rtt_wlen`은 min RTT filter window, `tcp_moderate_rcvbuf`는 receive buffer autotuning을 켭니다. `tcp_mtu_probing`은 0=끔, 1=ICMP black hole 감지 시 시작, 2=항상 시작이고 초기 MSS는 `tcp_base_mss`입니다. `tcp_probe_interval`은 probe 재시도 간격, `tcp_probe_threshold`는 search range를 멈출 임계값입니다.

`tcp_no_metrics_save`는 connection 종료 때 metric cache 저장을 막고 `tcp_no_ssthresh_metrics_save`는 ssthresh만 저장하지 않습니다. `tcp_orphan_retries`는 local TCP connection의 orphan probe 횟수이며 0이어도 최대 8회까지 재시도할 수 있어 즉시 종료를 뜻하지 않습니다.


tcp_max_orphans - INTEGER
        Maximal number of TCP sockets not attached to any user file handle,
        held by system.        If this number is exceeded orphaned connections are
        reset immediately and warning is printed. This limit exists
        only to prevent simple DoS attacks, you _must_ not rely on this
        or lower the limit artificially, but rather increase it
        (probably, after increasing installed memory),
        if network conditions require more than default value,
        and tune network services to linger and kill such states
        more aggressively. Let me to remind again: each orphan eats
        up to ~64K of unswappable memory.

tcp_max_syn_backlog - INTEGER
        Maximal number of remembered connection requests (SYN_RECV),
        which have not received an acknowledgment from connecting client.

        This is a per-listener limit.

        The minimal value is 128 for low memory machines, and it will
        increase in proportion to the memory of machine.

        If server suffers from overload, try increasing this number.

        Remember to also check /proc/sys/net/core/somaxconn
        A SYN_RECV request socket consumes about 304 bytes of memory.

tcp_max_tw_buckets - INTEGER
        Maximal number of timewait sockets held by system simultaneously.
        If this number is exceeded time-wait socket is immediately destroyed
        and warning is printed. This limit exists only to prevent
        simple DoS attacks, you _must_ not lower the limit artificially,
        but rather increase it (probably, after increasing installed memory),
        if network conditions require more than default value.

tcp_mem - vector of 3 INTEGERs: min, pressure, max
        min: below this number of pages TCP is not bothered about its
        memory appetite.

        pressure: when amount of memory allocated by TCP exceeds this number
        of pages, TCP moderates its memory consumption and enters memory
        pressure mode, which is exited when memory consumption falls
        under "min".

        max: number of pages allowed for queueing by all TCP sockets.

        Defaults are calculated at boot time from amount of available
        memory.

tcp_min_rtt_wlen - INTEGER
        The window length of the windowed min filter to track the minimum RTT.
        A shorter window lets a flow more quickly pick up new (higher)
        minimum RTT when it is moved to a longer path (e.g., due to traffic
        engineering). A longer window makes the filter more resistant to RTT
        inflations such as transient congestion. The unit is seconds.

        Possible values: 0 - 86400 (1 day)

        Default: 300

tcp_moderate_rcvbuf - BOOLEAN
        If enabled, TCP performs receive buffer auto-tuning, attempting to
        automatically size the buffer (no greater than tcp_rmem[2]) to
        match the size required by the path for full throughput.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_mtu_probing - INTEGER
        Controls TCP Packetization-Layer Path MTU Discovery.  Takes three
        values:

        - 0 - Disabled
        - 1 - Disabled by default, enabled when an ICMP black hole detected
        - 2 - Always enabled, use initial MSS of tcp_base_mss.

tcp_probe_interval - UNSIGNED INTEGER
        Controls how often to start TCP Packetization-Layer Path MTU
        Discovery reprobe. The default is reprobing every 10 minutes as
        per RFC4821.

tcp_probe_threshold - INTEGER
        Controls when TCP Packetization-Layer Path MTU Discovery probing
        will stop in respect to the width of search range in bytes. Default
        is 8 bytes.

tcp_no_metrics_save - BOOLEAN
        By default, TCP saves various connection metrics in the route cache
        when the connection closes, so that connections established in the
        near future can use these to set initial conditions.  Usually, this
        increases overall performance, but may sometimes cause performance
        degradation.  If enabled, TCP will not cache metrics on closing
        connections.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_no_ssthresh_metrics_save - BOOLEAN
        Controls whether TCP saves ssthresh metrics in the route cache.
        If enabled, ssthresh metrics are disabled.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_orphan_retries - INTEGER
        This value influences the timeout of a locally closed TCP connection,
        when RTO retransmissions remain unacknowledged.
        See tcp_retries2 for more details.

        The default value is 8.

        If your machine is a loaded WEB server,
        you should think about lowering this value, such sockets
        may consume significant resources. Cf. tcp_max_orphans.

TCP RACK, reordering과 재전송

731-815

`tcp_recovery`는 RACK loss detection option bitmap입니다. bit 0x1은 RACK, 0x2는 RACK의 static reordering window, 0x4는 DUPACK threshold로 reorder window를 끄는 동작을 비활성화합니다. 기본 0x1입니다. `tcp_reflect_tos`는 incoming SYN의 DSCP를 child socket에 반영하며 ECN bit는 반영하지 않습니다.

`tcp_reordering`은 packet reordering의 초기 추정값(기본 3), `tcp_max_reordering`은 최대 추정값(기본 300)입니다. `tcp_retrans_collapse`는 retransmit 때 full-sized packet으로 합치려 하지만 수신 측 bug 우회가 필요할 수 있습니다. `tcp_retries1`은 IP layer에 network 상태를 다시 평가시키기 전 재시도, `tcp_retries2`는 established connection을 포기하기 전 재시도이며 RFC 1122 기준상 최소 R2는 100초 이상에 해당해야 합니다.

tcp_recovery - INTEGER
        This value is a bitmap to enable various experimental loss recovery
        features.

        =========   =============================================================
        RACK: 0x1   enables RACK loss detection, for fast detection of lost
                    retransmissions and tail drops, and resilience to
                    reordering. currently, setting this bit to 0 has no
                    effect, since RACK is the only supported loss detection
                    algorithm.

        RACK: 0x2   makes RACK's reordering window static (min_rtt/4).

        RACK: 0x4   disables RACK's DUPACK threshold heuristic
        =========   =============================================================

        Default: 0x1

tcp_reflect_tos - BOOLEAN
        For listening sockets, reuse the DSCP value of the initial SYN message
        for outgoing packets. This allows to have both directions of a TCP
        stream to use the same DSCP value, assuming DSCP remains unchanged for
        the lifetime of the connection.

        This options affects both IPv4 and IPv6.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_reordering - INTEGER
        Initial reordering level of packets in a TCP stream.
        TCP stack can then dynamically adjust flow reordering level
        between this initial value and tcp_max_reordering

        Default: 3

tcp_max_reordering - INTEGER
        Maximal reordering level of packets in a TCP stream.
        300 is a fairly conservative value, but you might increase it
        if paths are using per packet load balancing (like bonding rr mode)

        Default: 300

tcp_retrans_collapse - BOOLEAN
        Bug-to-bug compatibility with some broken printers.
        On retransmit try to send bigger packets to work around bugs in
        certain TCP stacks.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_retries1 - INTEGER
        This value influences the time, after which TCP decides, that
        something is wrong due to unacknowledged RTO retransmissions,
        and reports this suspicion to the network layer.
        See tcp_retries2 for more details.

        RFC 1122 recommends at least 3 retransmissions, which is the
        default.

tcp_retries2 - INTEGER
        This value influences the timeout of an alive TCP connection,
        when RTO retransmissions remain unacknowledged.
        Given a value of N, a hypothetical TCP connection following
        exponential backoff with an initial RTO of TCP_RTO_MIN would
        retransmit N times before killing the connection at the (N+1)th RTO.

        The default value of 15 yields a hypothetical timeout of 924.6
        seconds and is a lower bound for the effective timeout.
        TCP will effectively time out at the first RTO which exceeds the
        hypothetical timeout.
        If tcp_rto_max_ms is decreased, it is recommended to also
        change tcp_retries2.

        RFC 1122 recommends at least 100 seconds for the timeout,
        which corresponds to a value of at least 8.

TCP receive buffer, SACK와 SYN 보호

816-976

`tcp_rfc1337`은 TIME_WAIT assassination으로부터 RFC 1337 방식으로 보호합니다. `tcp_rmem`의 min/default/max는 receive buffer의 최소 보장, 기본값, autotuning 최대값이며 `SO_RCVBUF`는 이를 override합니다. `tcp_sack`은 selective ACK을 켭니다.

SACK compression은 `tcp_comp_sack_delay_ns`만큼 ACK을 늦춰 압축하고 `tcp_comp_sack_slack_ns`는 timer slack, `tcp_comp_sack_nr`는 한 packet에 담을 최대 SACK 수입니다. `tcp_backlog_ack_defer`는 process context가 packet을 처리할 때 ACK 전송을 backlog 끝까지 미루고, `tcp_slow_start_after_idle`은 idle 뒤 congestion window를 timeout 기준으로 줄입니다. `tcp_stdurg`는 RFC 1122 urgent pointer 해석을 사용하지만 기본은 BSD 호환 방식입니다.

`tcp_synack_retries`는 passive SYN/ACK 재전송 횟수입니다. `tcp_syncookies`는 0=끔, 1=SYN backlog overflow 때만, 2=항상이며 정상 부하 확장용이 아니라 마지막 방어 수단입니다. cookies는 TCP option 정보를 잃을 수 있고 listener overload를 고치지 않습니다. `tcp_migrate_req`는 `SO_REUSEPORT` listener가 죽을 때 handshake request를 다른 listener로 옮기며 BPF program이 대상 socket을 선택할 수 있습니다.

tcp_rfc1337 - BOOLEAN
        If enabled, the TCP stack behaves conforming to RFC1337. If unset,
        we are not conforming to RFC, but prevent TCP TIME_WAIT
        assassination.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_rmem - vector of 3 INTEGERs: min, default, max
        min: Minimal size of receive buffer used by TCP sockets.
        It is guaranteed to each TCP socket, even under moderate memory
        pressure.

        Default: 4K

        default: initial size of receive buffer used by TCP sockets.
        This value overrides net.core.rmem_default used by other protocols.
        Default: 131072 bytes.
        This value results in initial window of 65535.

        max: maximal size of receive buffer allowed for automatically
        selected receiver buffers for TCP socket.
        Calling setsockopt() with SO_RCVBUF disables
        automatic tuning of that socket's receive buffer size, in which
        case this value is ignored.
        Default: between 131072 and 32MB, depending on RAM size.

tcp_sack - BOOLEAN
        Enable select acknowledgments (SACKS).

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_comp_sack_delay_ns - LONG INTEGER
        TCP tries to reduce number of SACK sent, using a timer
        based on 5% of SRTT, capped by this sysctl, in nano seconds.
        The default is 1ms, based on TSO autosizing period.

        Default : 1,000,000 ns (1 ms)

tcp_comp_sack_slack_ns - LONG INTEGER
        This sysctl control the slack used when arming the
        timer used by SACK compression. This gives extra time
        for small RTT flows, and reduces system overhead by allowing
        opportunistic reduction of timer interrupts.

        Default : 100,000 ns (100 us)

tcp_comp_sack_nr - INTEGER
        Max number of SACK that can be compressed.
        Using 0 disables SACK compression.

        Default : 44

tcp_backlog_ack_defer - BOOLEAN
        If enabled, user thread processing socket backlog tries sending
        one ACK for the whole queue. This helps to avoid potential
        long latencies at end of a TCP socket syscall.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_slow_start_after_idle - BOOLEAN
        If enabled, provide RFC2861 behavior and time out the congestion
        window after an idle period.  An idle period is defined at
        the current RTO.  If unset, the congestion window will not
        be timed out after an idle period.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_stdurg - BOOLEAN
        Use the Host requirements interpretation of the TCP urgent pointer field.
        Most hosts use the older BSD interpretation, so if enabled,
        Linux might not communicate correctly with them.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_synack_retries - INTEGER
        Number of times SYNACKs for a passive TCP connection attempt will
        be retransmitted. Should not be higher than 255. Default value
        is 5, which corresponds to 31seconds till the last retransmission
        with the current initial RTO of 1second. With this the final timeout
        for a passive TCP connection will happen after 63seconds.

tcp_syncookies - INTEGER
        Only valid when the kernel was compiled with CONFIG_SYN_COOKIES
        Send out syncookies when the syn backlog queue of a socket
        overflows. This is to prevent against the common 'SYN flood attack'
        Default: 1

        Note, that syncookies is fallback facility.
        It MUST NOT be used to help highly loaded servers to stand
        against legal connection rate. If you see SYN flood warnings
        in your logs, but investigation        shows that they occur
        because of overload with legal connections, you should tune
        another parameters until this warning disappear.
        See: tcp_max_syn_backlog, tcp_synack_retries, tcp_abort_on_overflow.

        syncookies seriously violate TCP protocol, do not allow
        to use TCP extensions, can result in serious degradation
        of some services (f.e. SMTP relaying), visible not by you,
        but your clients and relays, contacting you. While you see
        SYN flood warnings in logs not being really flooded, your server
        is seriously misconfigured.

        If you want to test which effects syncookies have to your
        network connections you can set this knob to 2 to enable
        unconditionally generation of syncookies.

tcp_migrate_req - BOOLEAN
        The incoming connection is tied to a specific listening socket when
        the initial SYN packet is received during the three-way handshake.
        When a listener is closed, in-flight request sockets during the
        handshake and established sockets in the accept queue are aborted.

        If the listener has SO_REUSEPORT enabled, other listeners on the
        same port should have been able to accept such connections. This
        option makes it possible to migrate such child sockets to another
        listener after close() or shutdown().

        The BPF_SK_REUSEPORT_SELECT_OR_MIGRATE type of eBPF program should
        usually be used to define the policy to pick an alive listener.
        Otherwise, the kernel will randomly pick an alive listener only if
        this option is enabled.

        Note that migration between listeners with different settings may
        crash applications. Let's say migration happens from listener A to
        B, and only B has TCP_SAVE_SYN enabled. B cannot read SYN data from
        the requests migrated from A. To avoid such a situation, cancel
        migration by returning SK_DROP in the type of eBPF program, or
        disable this option.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

TCP Fast Open, timestamp와 pacing

977-1125

`tcp_fastopen`은 bitmap입니다. 0x1 client, 0x2 server, 0x4 cookie 없는 client, 0x200 cookie 없는 server, 0x400은 `TCP_FASTOPEN` socket option 없이 모든 listener를 켭니다. server 쪽은 listener 또는 0x400 설정이 추가로 필요합니다. `tcp_fastopen_blackhole_timeout_sec`는 TFO blackhole 감지 뒤 기능을 끄는 초기 timeout이고 연속 실패하면 지수 backoff합니다. `tcp_fastopen_key`는 최대 4개의 128-bit key를 지정하며 첫 key로 cookie를 만들고 나머지는 검증에만 씁니다.

`tcp_syn_retries`는 active SYN 재전송 횟수, `tcp_timestamps`는 0=끔, 1=random offset timestamp, 2=offset 없는 timestamp입니다. `tcp_min_tso_segs`는 TSO frame당 최소 segment 수이고 `tcp_tso_rtt_log`는 RTT에 따라 autosizing하는 packet 크기의 지수 기준을 바꿉니다.

`tcp_pacing_ss_ratio`와 `tcp_pacing_ca_ratio`는 slow start와 congestion avoidance에서 internal pacing rate를 current rate의 몇 %로 잡을지 정합니다. `tcp_syn_linear_timeouts` 횟수까지 SYN timeout을 선형으로 둔 뒤 exponential backoff하며 기본 4는 초기 RTO 1초에서 6번째 재전송까지 약 10초가 걸립니다. `tcp_tso_win_divisor`는 한 TSO frame이 congestion window의 어느 분수까지 쓸지 정합니다.

tcp_fastopen - INTEGER
        Enable TCP Fast Open (RFC7413) to send and accept data in the opening
        SYN packet.

        The client support is enabled by flag 0x1 (on by default). The client
        then must use sendmsg() or sendto() with the MSG_FASTOPEN flag,
        rather than connect() to send data in SYN.

        The server support is enabled by flag 0x2 (off by default). Then
        either enable for all listeners with another flag (0x400) or
        enable individual listeners via TCP_FASTOPEN socket option with
        the option value being the length of the syn-data backlog.

        The values (bitmap) are

        =====  ======== ======================================================
          0x1  (client) enables sending data in the opening SYN on the client.
          0x2  (server) enables the server support, i.e., allowing data in
                        a SYN packet to be accepted and passed to the
                        application before 3-way handshake finishes.
          0x4  (client) send data in the opening SYN regardless of cookie
                        availability and without a cookie option.
        0x200  (server) accept data-in-SYN w/o any cookie option present.
        0x400  (server) enable all listeners to support Fast Open by
                        default without explicit TCP_FASTOPEN socket option.
        =====  ======== ======================================================

        Default: 0x1

        Note that additional client or server features are only
        effective if the basic support (0x1 and 0x2) are enabled respectively.

tcp_fastopen_blackhole_timeout_sec - INTEGER
        Initial time period in second to disable Fastopen on active TCP sockets
        when a TFO firewall blackhole issue happens.
        This time period will grow exponentially when more blackhole issues
        get detected right after Fastopen is re-enabled and will reset to
        initial value when the blackhole issue goes away.
        0 to disable the blackhole detection.

        By default, it is set to 0 (feature is disabled).

tcp_fastopen_key - list of comma separated 32-digit hexadecimal INTEGERs
        The list consists of a primary key and an optional backup key. The
        primary key is used for both creating and validating cookies, while the
        optional backup key is only used for validating cookies. The purpose of
        the backup key is to maximize TFO validation when keys are rotated.

        A randomly chosen primary key may be configured by the kernel if
        the tcp_fastopen sysctl is set to 0x400 (see above), or if the
        TCP_FASTOPEN setsockopt() optname is set and a key has not been
        previously configured via sysctl. If keys are configured via
        setsockopt() by using the TCP_FASTOPEN_KEY optname, then those
        per-socket keys will be used instead of any keys that are specified via
        sysctl.

        A key is specified as 4 8-digit hexadecimal integers which are separated
        by a '-' as: xxxxxxxx-xxxxxxxx-xxxxxxxx-xxxxxxxx. Leading zeros may be
        omitted. A primary and a backup key may be specified by separating them
        by a comma. If only one key is specified, it becomes the primary key and
        any previously configured backup keys are removed.

tcp_syn_retries - INTEGER
        Number of times initial SYNs for an active TCP connection attempt
        will be retransmitted. Should not be higher than 127. Default value
        is 6, which corresponds to 67seconds (with tcp_syn_linear_timeouts = 4)
        till the last retransmission with the current initial RTO of 1second.
        With this the final timeout for an active TCP connection attempt
        will happen after 131seconds.

tcp_timestamps - INTEGER
        Enable timestamps as defined in RFC1323.

        - 0: Disabled.
        - 1: Enable timestamps as defined in RFC1323 and use random offset for
          each connection rather than only using the current time.
        - 2: Like 1, but without random offsets.

        Default: 1

tcp_min_tso_segs - INTEGER
        Minimal number of segments per TSO frame.

        Since linux-3.12, TCP does an automatic sizing of TSO frames,
        depending on flow rate, instead of filling 64Kbytes packets.
        For specific usages, it's possible to force TCP to build big
        TSO frames. Note that TCP stack might split too big TSO packets
        if available window is too small.

        Default: 2

tcp_tso_rtt_log - INTEGER
        Adjustment of TSO packet sizes based on min_rtt

        Starting from linux-5.18, TCP autosizing can be tweaked
        for flows having small RTT.

        Old autosizing was splitting the pacing budget to send 1024 TSO
        per second.

        tso_packet_size = sk->sk_pacing_rate / 1024;

        With the new mechanism, we increase this TSO sizing using:

        distance = min_rtt_usec / (2^tcp_tso_rtt_log)
        tso_packet_size += gso_max_size >> distance;

        This means that flows between very close hosts can use bigger
        TSO packets, reducing their cpu costs.

        If you want to use the old autosizing, set this sysctl to 0.

        Default: 9  (2^9 = 512 usec)

tcp_pacing_ss_ratio - INTEGER
        sk->sk_pacing_rate is set by TCP stack using a ratio applied
        to current rate. (current_rate = cwnd * mss / srtt)
        If TCP is in slow start, tcp_pacing_ss_ratio is applied
        to let TCP probe for bigger speeds, assuming cwnd can be
        doubled every other RTT.

        Default: 200

tcp_pacing_ca_ratio - INTEGER
        sk->sk_pacing_rate is set by TCP stack using a ratio applied
        to current rate. (current_rate = cwnd * mss / srtt)
        If TCP is in congestion avoidance phase, tcp_pacing_ca_ratio
        is applied to conservatively probe for bigger throughput.

        Default: 120

tcp_syn_linear_timeouts - INTEGER
        The number of times for an active TCP connection to retransmit SYNs with
        a linear backoff timeout before defaulting to an exponential backoff
        timeout. This has no effect on SYNACK at the passive TCP side.

        With an initial RTO of 1 and tcp_syn_linear_timeouts = 4 we would
        expect SYN RTOs to be: 1, 1, 1, 1, 1, 2, 4, ... (4 linear timeouts,
        and the first exponential backoff using 2^0 * initial_RTO).
        Default: 4

tcp_tso_win_divisor - INTEGER
        This allows control over what percentage of the congestion window
        can be consumed by a single TSO frame.
        The setting of this parameter is a choice between burstiness and
        building larger TSO frames.

        Default: 3

TCP TIME_WAIT, window와 send buffer

1126-1200

`tcp_tw_reuse`는 TIME_WAIT socket 재사용을 0=끔, 1=전역, 2=loopback만으로 정하며 기본 2입니다. `tcp_tw_reuse_delay`는 재사용 전 최소 지연으로 기본 1000 ms이며 peer timestamp clock tick보다 작게 두면 안 됩니다.

`tcp_window_scaling`은 RFC 1323 window scaling을 사용합니다. `tcp_shrink_window`는 receive buffer pressure 때 window를 이미 광고한 크기보다 줄일 수 있게 하지만 peer가 shrink를 잘못 처리하면 문제가 생길 수 있습니다. `tcp_wmem`의 min/default/max는 send buffer memory 보장과 autotuning 범위이며 `SO_SNDBUF`가 override할 수 있습니다.

tcp_tw_reuse - INTEGER
        Enable reuse of TIME-WAIT sockets for new connections when it is
        safe from protocol viewpoint.

        - 0 - disable
        - 1 - global enable
        - 2 - enable for loopback traffic only

        It should not be changed without advice/request of technical
        experts.

        Default: 2

tcp_tw_reuse_delay - UNSIGNED INTEGER
        The delay in milliseconds before a TIME-WAIT socket can be reused by a
        new connection, if TIME-WAIT socket reuse is enabled. The actual reuse
        threshold is within [N, N+1] range, where N is the requested delay in
        milliseconds, to ensure the delay interval is never shorter than the
        configured value.

        This setting contains an assumption about the other TCP timestamp clock
        tick interval. It should not be set to a value lower than the peer's
        clock tick for PAWS (Protection Against Wrapped Sequence numbers)
        mechanism work correctly for the reused connection.

        Default: 1000 (milliseconds)

tcp_window_scaling - BOOLEAN
        Enable window scaling as defined in RFC1323.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

tcp_shrink_window - BOOLEAN
        This changes how the TCP receive window is calculated.

        RFC 7323, section 2.4, says there are instances when a retracted
        window can be offered, and that TCP implementations MUST ensure
        that they handle a shrinking window, as specified in RFC 1122.

        Possible values:

        - 0 (disabled) - The window is never shrunk.
        - 1 (enabled)  - The window is shrunk when necessary to remain within
          the memory limit set by autotuning (sk_rcvbuf).
          This only occurs if a non-zero receive window
          scaling factor is also in effect.

        Default: 0 (disabled)

tcp_wmem - vector of 3 INTEGERs: min, default, max
        min: Amount of memory reserved for send buffers for TCP sockets.
        Each TCP socket has rights to use it due to fact of its birth.

        Default: 4K

        default: initial size of send buffer used by TCP sockets.  This
        value overrides net.core.wmem_default used by other protocols.

        It is usually lower than net.core.wmem_default.

        Default: 16K

        max: Maximal amount of memory allowed for automatically tuned
        send buffers for TCP sockets. This value does not override
        net.core.wmem_max.  Calling setsockopt() with SO_SNDBUF disables
        automatic tuning of that socket's send buffer size, in which case
        this value is ignored.

        Default: between 64K and 4MB, depending on RAM size.

TCP unsent queue, hash와 PLB

1201-1373

`tcp_notsent_lowat`는 `TCP_NOTSENT_LOWAT`을 따로 설정하지 않은 socket의 unsent byte 기준이며 기본 `UINT_MAX`는 제한 없음입니다. `tcp_workaround_signed_windows`는 window scaling 없는 잘못된 peer의 signed window를 우회합니다. `tcp_thin_linear_timeouts`는 얇은 stream에서 최대 6회의 RTO를 선형으로 두고, `tcp_limit_output_bytes`는 qdisc/device queue에 쌓을 socket별 byte를 제한해 bufferbloat를 줄입니다. `tcp_challenge_ack_limit`은 4.7부터 per-netns random 값으로 바뀌어 legacy global limit은 사실상 무한입니다.

`tcp_ehash_entries`는 현재 namespace TCP established hash bucket 수를 읽으며 초기 namespace에서만 boot parameter로 쓸 수 있습니다. `tcp_child_ehash_entries`는 child namespace의 hash 크기를 예약하며 0이면 initial namespace hash를 공유하고 양수는 2의 거듭제곱 bucket으로 반올림됩니다.

Protective Load Balancing은 congestion이 지속되는 ECMP path에서 flow hash를 바꿉니다. `tcp_plb_enabled`가 master switch이고 `tcp_plb_idle_rehash_rounds`는 idle 뒤 rehash, `tcp_plb_rehash_rounds`는 congestion round 수, `tcp_plb_suspend_rto_sec`는 RTO 뒤 PLB 중지 시간, `tcp_plb_cong_thresh`는 round가 congestion으로 판정될 CE 비율입니다. PLB는 ECN을 전제로 하며 기본은 꺼져 있습니다.

tcp_notsent_lowat - UNSIGNED INTEGER
        A TCP socket can control the amount of unsent bytes in its write queue,
        thanks to TCP_NOTSENT_LOWAT socket option. poll()/select()/epoll()
        reports POLLOUT events if the amount of unsent bytes is below a per
        socket value, and if the write queue is not full. sendmsg() will
        also not add new buffers if the limit is hit.

        This global variable controls the amount of unsent data for
        sockets not using TCP_NOTSENT_LOWAT. For these sockets, a change
        to the global variable has immediate effect.

        Default: UINT_MAX (0xFFFFFFFF)

tcp_workaround_signed_windows - BOOLEAN
        If enabled, assume no receipt of a window scaling option means the
        remote TCP is broken and treats the window as a signed quantity.
        If disabled, assume the remote TCP is not broken even if we do
        not receive a window scaling option from them.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_thin_linear_timeouts - BOOLEAN
        Enable dynamic triggering of linear timeouts for thin streams.
        If enabled, a check is performed upon retransmission by timeout to
        determine if the stream is thin (less than 4 packets in flight).
        As long as the stream is found to be thin, up to 6 linear
        timeouts may be performed before exponential backoff mode is
        initiated. This improves retransmission latency for
        non-aggressive thin streams, often found to be time-dependent.
        For more information on thin streams, see
        Documentation/networking/tcp-thin.rst

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_limit_output_bytes - INTEGER
        Controls TCP Small Queue limit per tcp socket.
        TCP bulk sender tends to increase packets in flight until it
        gets losses notifications. With SNDBUF autotuning, this can
        result in a large amount of packets queued on the local machine
        (e.g.: qdiscs, CPU backlog, or device) hurting latency of other
        flows, for typical pfifo_fast qdiscs.  tcp_limit_output_bytes
        limits the number of bytes on qdisc or device to reduce artificial
        RTT/cwnd and reduce bufferbloat.

        Default: 4194304 (4 MB)

tcp_challenge_ack_limit - INTEGER
        Limits number of Challenge ACK sent per second, as recommended
        in RFC 5961 (Improving TCP's Robustness to Blind In-Window Attacks)
        Note that this per netns rate limit can allow some side channel
        attacks and probably should not be enabled.
        TCP stack implements per TCP socket limits anyway.
        Default: INT_MAX (unlimited)

tcp_ehash_entries - INTEGER
        Show the number of hash buckets for TCP sockets in the current
        networking namespace.

        A negative value means the networking namespace does not own its
        hash buckets and shares the initial networking namespace's one.

tcp_child_ehash_entries - INTEGER
        Control the number of hash buckets for TCP sockets in the child
        networking namespace, which must be set before clone() or unshare().

        If the value is not 0, the kernel uses a value rounded up to 2^n
        as the actual hash bucket size.  0 is a special value, meaning
        the child networking namespace will share the initial networking
        namespace's hash buckets.

        Note that the child will use the global one in case the kernel
        fails to allocate enough memory.  In addition, the global hash
        buckets are spread over available NUMA nodes, but the allocation
        of the child hash table depends on the current process's NUMA
        policy, which could result in performance differences.

        Note also that the default value of tcp_max_tw_buckets and
        tcp_max_syn_backlog depend on the hash bucket size.

        Possible values: 0, 2^n (n: 0 - 24 (16Mi))

        Default: 0

tcp_plb_enabled - BOOLEAN
        If enabled and the underlying congestion control (e.g. DCTCP) supports
        and enables PLB feature, TCP PLB (Protective Load Balancing) is
        enabled. PLB is described in the following paper:
        https://doi.org/10.1145/3544216.3544226. Based on PLB parameters,
        upon sensing sustained congestion, TCP triggers a change in
        flow label field for outgoing IPv6 packets. A change in flow label
        field potentially changes the path of outgoing packets for switches
        that use ECMP/WCMP for routing.

        PLB changes socket txhash which results in a change in IPv6 Flow Label
        field, and currently no-op for IPv4 headers. It is possible
        to apply PLB for IPv4 with other network header fields (e.g. TCP
        or IPv4 options) or using encapsulation where outer header is used
        by switches to determine next hop. In either case, further host
        and switch side changes will be needed.

        If enabled, PLB assumes that congestion signal (e.g. ECN) is made
        available and used by congestion control module to estimate a
        congestion measure (e.g. ce_ratio). PLB needs a congestion measure to
        make repathing decisions.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

tcp_plb_idle_rehash_rounds - INTEGER
        Number of consecutive congested rounds (RTT) seen after which
        a rehash can be performed, given there are no packets in flight.
        This is referred to as M in PLB paper:
        https://doi.org/10.1145/3544216.3544226.

        Possible Values: 0 - 31

        Default: 3

tcp_plb_rehash_rounds - INTEGER
        Number of consecutive congested rounds (RTT) seen after which
        a forced rehash can be performed. Be careful when setting this
        parameter, as a small value increases the risk of retransmissions.
        This is referred to as N in PLB paper:
        https://doi.org/10.1145/3544216.3544226.

        Possible Values: 0 - 31

        Default: 12

tcp_plb_suspend_rto_sec - INTEGER
        Time, in seconds, to suspend PLB in event of an RTO. In order to avoid
        having PLB repath onto a connectivity "black hole", after an RTO a TCP
        connection suspends PLB repathing for a random duration between 1x and
        2x of this parameter. Randomness is added to avoid concurrent rehashing
        of multiple TCP connections. This should be set corresponding to the
        amount of time it takes to repair a failed link.

        Possible Values: 0 - 255

        Default: 60

tcp_plb_cong_thresh - INTEGER
        Fraction of packets marked with congestion over a round (RTT) to
        tag that round as congested. This is referred to as K in the PLB paper:
        https://doi.org/10.1145/3544216.3544226.

        The 0-1 fraction range is mapped to 0-256 range to avoid floating
        point operations. For example, 128 means that if at least 50% of
        the packets in a round were marked as congested then the round
        will be tagged as congested.

        Setting threshold to 0 means that PLB repaths every RTT regardless
        of congestion. This is not intended behavior for PLB and should be
        used only for experimentation purpose.

        Possible Values: 0 - 256

        Default: 128

TCP ping-pong 판정과 RTO 범위

1374-1410

`tcp_pingpong_thresh`는 request/response 교환 횟수가 기준 이하인 connection을 ping-pong으로 분류해 delayed ACK 동작에 반영합니다. 기본 1입니다. `tcp_rto_min_us`는 TCP RTO 하한을 microsecond로 정하며 0이면 route의 `rto_min`을 사용합니다. `tcp_rto_max_ms`는 RTO 상한으로 기본 120000 ms이고 1000 ms보다 작게 둘 수 없습니다.

tcp_pingpong_thresh - INTEGER
        The number of estimated data replies sent for estimated incoming data
        requests that must happen before TCP considers that a connection is a
        "ping-pong" (request-response) connection for which delayed
        acknowledgments can provide benefits.

        This threshold is 1 by default, but some applications may need a higher
        threshold for optimal performance.

        Possible Values: 1 - 255

        Default: 1

tcp_rto_min_us - INTEGER
        Minimal TCP retransmission timeout (in microseconds). Note that the
        rto_min route option has the highest precedence for configuring this
        setting, followed by the TCP_BPF_RTO_MIN and TCP_RTO_MIN_US socket
        options, followed by this tcp_rto_min_us sysctl.

        The recommended practice is to use a value less or equal to 200000
        microseconds.

        Possible Values: 1 - INT_MAX

        Default: 200000

tcp_rto_max_ms - INTEGER
        Maximal TCP retransmission timeout (in ms).
        Note that TCP_RTO_MAX_MS socket option has higher precedence.

        When changing tcp_rto_max_ms, it is important to understand
        that tcp_retries2 might need a change.

        Possible Values: 1000 - 120,000

        Default: 120,000

UDP 변수

1411-1475

`udp_l3mdev_accept`를 켜면 global bound UDP socket이 VRF 같은 L3 master domain을 가로질러 packet을 받을 수 있으며 기본 0입니다. `udp_mem`은 전체 UDP socket queue의 min/pressure/max page 수, `udp_rmem_min`은 memory pressure에서도 socket마다 보장할 receive buffer입니다. `udp_wmem_min`은 문서상 효과가 없습니다.

`udp_hash_entries`는 현재 namespace의 UDP hash table bucket 수입니다. `udp_child_hash_entries`는 child namespace가 별도 hash table을 만들 때 요청할 bucket 수이고 0이면 initial namespace table을 공유합니다. 별도 table은 1024보다 작게 만들 수 없으며 2의 거듭제곱으로 반올림됩니다.

UDP variables
=============

udp_l3mdev_accept - BOOLEAN
        Enabling this option allows a "global" bound socket to work
        across L3 master domains (e.g., VRFs) with packets capable of
        being received regardless of the L3 domain in which they
        originated. Only valid when the kernel was compiled with
        CONFIG_NET_L3_MASTER_DEV.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

udp_mem - vector of 3 INTEGERs: min, pressure, max
        Number of pages allowed for queueing by all UDP sockets.

        min: Number of pages allowed for queueing by all UDP sockets.

        pressure: This value was introduced to follow format of tcp_mem.

        max: This value was introduced to follow format of tcp_mem.

        Default is calculated at boot time from amount of available memory.

udp_rmem_min - INTEGER
        Minimal size of receive buffer used by UDP sockets in moderation.
        Each UDP socket is able to use the size for receiving data, even if
        total pages of UDP sockets exceed udp_mem pressure. The unit is byte.

        Default: 4K

udp_wmem_min - INTEGER
        UDP does not have tx memory accounting and this tunable has no effect.

udp_hash_entries - INTEGER
        Show the number of hash buckets for UDP sockets in the current
        networking namespace.

        A negative value means the networking namespace does not own its
        hash buckets and shares the initial networking namespace's one.

udp_child_hash_entries - INTEGER
        Control the number of hash buckets for UDP sockets in the child
        networking namespace, which must be set before clone() or unshare().

        If the value is not 0, the kernel uses a value rounded up to 2^n
        as the actual hash bucket size.  0 is a special value, meaning
        the child networking namespace will share the initial networking
        namespace's hash buckets.

        Note that the child will use the global one in case the kernel
        fails to allocate enough memory.  In addition, the global hash
        buckets are spread over available NUMA nodes, but the allocation
        of the child hash table depends on the current process's NUMA
        policy, which could result in performance differences.

        Possible values: 0, 2^n (n: 7 (128) - 16 (64K))

        Default: 0

RAW socket 변수

1476-1492

`raw_l3mdev_accept`는 global bound RAW socket이 모든 L3 master domain의 packet을 받을 수 있게 합니다. kernel이 `CONFIG_NET_L3_MASTER_DEV`로 빌드되었을 때만 유효하며 기본 1입니다.

RAW variables
=============

raw_l3mdev_accept - BOOLEAN
        Enabling this option allows a "global" bound socket to work
        across L3 master domains (e.g., VRFs) with packets capable of
        being received regardless of the L3 domain in which they
        originated. Only valid when the kernel was compiled with
        CONFIG_NET_L3_MASTER_DEV.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

CIPSOv4 변수

1493-1547

`cipso_cache_enable`은 incoming packet의 DOI/level/category mapping cache를 켜고 기본 1입니다. `cipso_cache_bucket_size`는 hash bucket마다 보관할 entry 수로 기본 10이며 0은 사실상 cache를 끕니다.

`cipso_rbm_optfmt`는 restricted bitmap category를 가능한 경우 trailing zero 없이 optimized form으로 만들고 기본 0은 fixed 32-bit multiple 형식입니다. `cipso_rbm_strictvalid`는 incoming CIPSO option을 엄격히 검사합니다. 켜면 malformed option을 더 잘 막지만 처리량이 줄어 기본은 1입니다.

CIPSOv4 Variables
=================

cipso_cache_enable - BOOLEAN
        If enabled, enable additions to and lookups from the CIPSO label mapping
        cache.  If disabled, additions are ignored and lookups always result in a
        miss.  However, regardless of the setting the cache is still
        invalidated when required when means you can safely toggle this on and
        off and the cache will always be "safe".

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

cipso_cache_bucket_size - INTEGER
        The CIPSO label cache consists of a fixed size hash table with each
        hash bucket containing a number of cache entries.  This variable limits
        the number of entries in each hash bucket; the larger the value is, the
        more CIPSO label mappings that can be cached.  When the number of
        entries in a given hash bucket reaches this limit adding new entries
        causes the oldest entry in the bucket to be removed to make room.

        Default: 10

cipso_rbm_optfmt - BOOLEAN
        Enable the "Optimized Tag 1 Format" as defined in section 3.4.2.6 of
        the CIPSO draft specification (see Documentation/netlabel for details).
        This means that when set the CIPSO tag will be padded with empty
        categories in order to make the packet data 32-bit aligned.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

cipso_rbm_strictvalid - BOOLEAN
        If enabled, do a very strict check of the CIPSO option when
        ip_options_compile() is called.  If disabled, relax the checks done during
        ip_options_compile().  Either way is "safe" as errors are caught else
        where in the CIPSO processing code but setting this to 0 (False) should
        result in less work (i.e. it should be faster) but could cause problems
        with other implementations that require strict checking.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

Local port, bind와 early demux

1548-1678

`ip_local_port_range`는 TCP/UDP가 자동 선택할 local port의 첫·끝 값이며 기본 32768~60999입니다. 두 끝의 parity를 다르게 두는 편이 좋고 시작은 `ip_unprivileged_port_start` 이상이어야 합니다. `ip_local_reserved_ports`는 자동 할당에서 빼 둘 port/range 목록입니다. explicit bind에는 영향을 주지 않고 쓸 때 기존 목록 전체를 교체합니다. 두 range는 독립적으로 평가되며 예약 block 직후 port의 선택 확률에 영향을 줄 수 있습니다.

namespace별 `ip_unprivileged_port_start`는 capability 없이 bind 가능한 첫 port이고 기본 1024, 0이면 privileged port가 없습니다. `ip_nonlocal_bind`는 local address가 아닌 주소에 bind를 허용합니다. `ip_autobind_reuse`는 `SO_REUSEADDR` socket의 bind+connect 자동 port 재사용을 허용하지만 `IP_BIND_ADDRESS_NO_PORT`가 권장되며 전문가용입니다. `ip_dynaddr`는 dynamic address rewrite를 켜고 1보다 크면 rewrite를 log합니다.

`ip_early_demux`는 established TCP와 connected UDP의 local input demux를 앞당겨 socket workload를 줄이지만 순수 router에서는 추가 비용이 될 수 있습니다. `ping_group_range`는 ICMP datagram ping socket을 만들 수 있는 group ID 범위입니다. 기본 `1 0`은 root를 포함해 아무도 허용하지 않습니다. `tcp_early_demux`와 `udp_early_demux`로 protocol별 최적화를 끌 수 있습니다.

IP Variables
============

ip_local_port_range - 2 INTEGERS
        Defines the local port range that is used by TCP and UDP to
        choose the local port. The first number is the first, the
        second the last local port number.
        If possible, it is better these numbers have different parity
        (one even and one odd value).
        Must be greater than or equal to ip_unprivileged_port_start.
        The default values are 32768 and 60999 respectively.

ip_local_reserved_ports - list of comma separated ranges
        Specify the ports which are reserved for known third-party
        applications. These ports will not be used by automatic port
        assignments (e.g. when calling connect() or bind() with port
        number 0). Explicit port allocation behavior is unchanged.

        The format used for both input and output is a comma separated
        list of ranges (e.g. "1,2-4,10-10" for ports 1, 2, 3, 4 and
        10). Writing to the file will clear all previously reserved
        ports and update the current list with the one given in the
        input.

        Note that ip_local_port_range and ip_local_reserved_ports
        settings are independent and both are considered by the kernel
        when determining which ports are available for automatic port
        assignments.

        You can reserve ports which are not in the current
        ip_local_port_range, e.g.::

            $ cat /proc/sys/net/ipv4/ip_local_port_range
            32000        60999
            $ cat /proc/sys/net/ipv4/ip_local_reserved_ports
            8080,9148

        although this is redundant. However such a setting is useful
        if later the port range is changed to a value that will
        include the reserved ports. Also keep in mind, that overlapping
        of these ranges may affect probability of selecting ephemeral
        ports which are right after block of reserved ports.

        Default: Empty

ip_unprivileged_port_start - INTEGER
        This is a per-namespace sysctl.  It defines the first
        unprivileged port in the network namespace.  Privileged ports
        require root or CAP_NET_BIND_SERVICE in order to bind to them.
        To disable all privileged ports, set this to 0.  They must not
        overlap with the ip_local_port_range.

        Default: 1024

ip_nonlocal_bind - BOOLEAN
        If enabled, allows processes to bind() to non-local IP addresses,
        which can be quite useful - but may break some applications.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

ip_autobind_reuse - BOOLEAN
        By default, bind() does not select the ports automatically even if
        the new socket and all sockets bound to the port have SO_REUSEADDR.
        ip_autobind_reuse allows bind() to reuse the port and this is useful
        when you use bind()+connect(), but may break some applications.
        The preferred solution is to use IP_BIND_ADDRESS_NO_PORT and this
        option should only be set by experts.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

ip_dynaddr - INTEGER
        If set non-zero, enables support for dynamic addresses.
        If set to a non-zero value larger than 1, a kernel log
        message will be printed when dynamic address rewriting
        occurs.

        Default: 0

ip_early_demux - BOOLEAN
        Optimize input packet processing down to one demux for
        certain kinds of local sockets.  Currently we only do this
        for established TCP and connected UDP sockets.

        It may add an additional cost for pure routing workloads that
        reduces overall throughput, in such case you should disable it.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

ping_group_range - 2 INTEGERS
        Restrict ICMP_PROTO datagram sockets to users in the group range.
        The default is "1 0", meaning, that nobody (not even root) may
        create ping sockets.  Setting it to "100 100" would grant permissions
        to the single group. "0 4294967294" would enable it for the world, "100
        4294967294" would enable it for the users, but not daemons.

tcp_early_demux - BOOLEAN
        Enable early demux for established TCP sockets.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

udp_early_demux - BOOLEAN
        Enable early demux for connected UDP sockets. Disable this if
        your system could experience more unconnected load.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

ICMPv4 rate와 응답 정책

1679-1798

`icmp_echo_ignore_all`은 모든 IPv4 echo request를 무시합니다. `icmp_echo_enable_probe`는 RFC 8335 extended echo request에 응답할지 정하고 `icmp_echo_ignore_broadcasts`는 broadcast/multicast echo를 무시하며 기본 1입니다.

`icmp_ratelimit`은 `icmp_ratemask`에 해당하는 ICMP type 응답 사이의 최소 ms이고 0이면 제한이 없습니다. `icmp_msgs_per_sec`와 `icmp_msgs_burst`는 per-CPU global rate와 burst를 대략 제한하며 randomization 때문에 정확한 한도로 보아서는 안 됩니다. 기본 mask 6168은 Destination Unreachable, Source Quench, Time Exceeded, Parameter Problem에 제한을 적용합니다.

`icmp_ignore_bogus_error_responses`는 RFC 1122를 어기고 broadcast frame에 보낸 bogus error를 무시합니다. `icmp_errors_use_inbound_ifaddr`는 ICMP error source address를 packet이 들어온 interface의 primary address로 고릅니다. 기본 0은 packet을 보내는 interface의 primary address를 쓰며 router 식별과 traceroute 결과가 달라질 수 있습니다.

icmp_echo_ignore_all - BOOLEAN
        If enabled, then the kernel will ignore all ICMP ECHO
        requests sent to it.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

icmp_echo_enable_probe - BOOLEAN
        If enabled, then the kernel will respond to RFC 8335 PROBE
        requests sent to it.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

icmp_echo_ignore_broadcasts - BOOLEAN
        If enabled, then the kernel will ignore all ICMP ECHO and
        TIMESTAMP requests sent to it via broadcast/multicast.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

icmp_ratelimit - INTEGER
        Limit the maximal rates for sending ICMP packets whose type matches
        icmp_ratemask (see below) to specific targets.
        0 to disable any limiting,
        otherwise the minimal space between responses in milliseconds.
        Note that another sysctl, icmp_msgs_per_sec limits the number
        of ICMP packets        sent on all targets.

        Default: 1000

icmp_msgs_per_sec - INTEGER
        Limit maximal number of ICMP packets sent per second from this host.
        Only messages whose type matches icmp_ratemask (see below) are
        controlled by this limit. For security reasons, the precise count
        of messages per second is randomized.

        Default: 1000

icmp_msgs_burst - INTEGER
        icmp_msgs_per_sec controls number of ICMP packets sent per second,
        while icmp_msgs_burst controls the burst size of these packets.
        For security reasons, the precise burst size is randomized.

        Default: 50

icmp_ratemask - INTEGER
        Mask made of ICMP types for which rates are being limited.

        Significant bits: IHGFEDCBA9876543210

        Default mask:     0000001100000011000 (6168)

        Bit definitions (see include/linux/icmp.h):

                = =========================
                0 Echo Reply
                3 Destination Unreachable [1]_
                4 Source Quench [1]_
                5 Redirect
                8 Echo Request
                B Time Exceeded [1]_
                C Parameter Problem [1]_
                D Timestamp Request
                E Timestamp Reply
                F Info Request
                G Info Reply
                H Address Mask Request
                I Address Mask Reply
                = =========================

        .. [1] These are rate limited by default (see default mask above)

icmp_ignore_bogus_error_responses - BOOLEAN
        Some routers violate RFC1122 by sending bogus responses to broadcast
        frames.  Such violations are normally logged via a kernel warning.
        If enabled, the kernel will not give such warnings, which
        will avoid log file clutter.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

icmp_errors_use_inbound_ifaddr - BOOLEAN

        If disabled, icmp error messages are sent with the primary address of
        the exiting interface.

        If enabled, the message will be sent with the primary address of
        the interface that received the packet that caused the icmp error.
        This is the behaviour many network administrators will expect from
        a router. And it can make debugging complicated network layouts
        much easier.

        Note that if no primary address exists for the interface selected,
        then the primary address of the first non-loopback interface that
        has one will be used regardless of this setting.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

IGMP membership과 version

1799-1850

`igmp_max_memberships`는 socket이 join할 수 있는 multicast group 수로 기본 20입니다. 보고 packet 하나가 담을 수 있는 최대치는 65535-byte IP packet 한계 때문에 대략 5459이고 너무 크게 하면 report가 여러 packet으로 나뉘어 switch가 처리하지 못할 수 있습니다. `igmp_max_msf`는 group별 source filter 수로 기본 10입니다.

`igmp_qrv`는 IGMP Query Robustness Variable이며 기본 2, 7보다 크면 IGMPv3 report의 3-bit QRV에 0으로 표시됩니다. `force_igmp_version`은 0=자동, 1=IGMPv1, 2=IGMPv2, 3=IGMPv3이고 특수 0/1 값의 legacy 차이를 문서에 따릅니다.

igmp_max_memberships - INTEGER
        Change the maximum number of multicast groups we can subscribe to.
        Default: 20

        Theoretical maximum value is bounded by having to send a membership
        report in a single datagram (i.e. the report can't span multiple
        datagrams, or risk confusing the switch and leaving groups you don't
        intend to).

        The number of supported groups 'M' is bounded by the number of group
        report entries you can fit into a single datagram of 65535 bytes.

        M = 65536-sizeof (ip header)/(sizeof(Group record))

        Group records are variable length, with a minimum of 12 bytes.
        So net.ipv4.igmp_max_memberships should not be set higher than:

        (65536-24) / 12 = 5459

        The value 5459 assumes no IP header options, so in practice
        this number may be lower.

igmp_max_msf - INTEGER
        Maximum number of addresses allowed in the source filter list for a
        multicast group.

        Default: 10

igmp_qrv - INTEGER
        Controls the IGMP query robustness variable (see RFC2236 8.1).

        Default: 2 (as specified by RFC2236 8.1)

        Minimum: 1 (as specified by RFC6636 4.5)

force_igmp_version - INTEGER
        - 0 - (default) No enforcement of a IGMP version, IGMPv1/v2 fallback
          allowed. Will back to IGMPv3 mode again if all IGMPv1/v2 Querier
          Present timer expires.
        - 1 - Enforce to use IGMP version 1. Will also reply IGMPv1 report if
          receive IGMPv2/v3 query.
        - 2 - Enforce to use IGMP version 2. Will fallback to IGMPv1 if receive
          IGMPv1 query message. Will reply report if receive IGMPv3 query.
        - 3 - Enforce to use IGMP version 3. The same react with default 0.

        .. note::

           this is not the same with force_mld_version because IGMPv3 RFC3376
           Security Considerations does not have clear description that we could
           ignore other version messages completely as MLDv2 RFC3810. So make
           this value as default 0 is recommended.

IPv4 interface별 forwarding과 redirect

1851-2044

`conf/interface/*`는 한 interface, `conf/all/*`은 모든 interface에 적용합니다. `log_martians`는 불가능한 source address를 log합니다. `accept_redirects`는 host에서는 all 또는 interface 중 하나가 true면, forwarding router에서는 둘 다 true일 때만 켜집니다. 기본은 host=true, router=false입니다. `forwarding`은 해당 ingress interface packet의 전달, `mc_forwarding`은 `CONFIG_MROUTE`와 daemon을 전제로 multicast routing을 켭니다.

`medium_id`는 interface가 공유하는 medium을 구분해 서로 다른 medium 사이 proxy ARP 판단에 쓰고 0=유일 interface, -1=미상입니다. `proxy_arp`는 proxy ARP, `proxy_arp_pvlan`은 같은 interface로 다시 응답해 RFC 3069 port isolation host가 upstream router를 통해 통신하게 합니다. `proxy_delay`는 proxy ARP/NDP 응답에 [0, 값) jiffy random delay를 둡니다.

`shared_media`는 RFC 1620 redirect를 보내거나 받아 `secure_redirects`보다 우선합니다. `secure_redirects`는 현재 gateway list의 gateway로 가는 redirect만 받고, `send_redirects`는 router가 redirect를 보냅니다. `bootp_relay`는 아직 구현되지 않았습니다. `accept_source_route`는 SRR option을 받고, `accept_local`은 local source packet을 wire에서 받아들이며, `route_localnet`은 127/8을 routing에서 martian으로 보지 않습니다.

`rp_filter`는 0=검증 없음, 1=best reverse path인 strict RFC 3704, 2=어느 interface로든 reachable하면 되는 loose mode입니다. spoofing 방지에는 strict가 좋지만 asymmetric routing에는 loose가 적합합니다. all/interface 중 큰 값을 씁니다. `src_valid_mark`를 켜면 reverse lookup과 ICMP source/일부 IP option address 선택에 packet fwmark를 포함해 양방향 policy routing과 맞춥니다.

``conf/interface/*``
        changes special settings per interface (where
        interface" is the name of your network interface)

``conf/all/*``
          is special, changes the settings for all interfaces

log_martians - BOOLEAN
        Log packets with impossible addresses to kernel log.
        log_martians for the interface will be enabled if at least one of
        conf/{all,interface}/log_martians is set to TRUE,
        it will be disabled otherwise

accept_redirects - BOOLEAN
        Accept ICMP redirect messages.
        accept_redirects for the interface will be enabled if:

        - both conf/{all,interface}/accept_redirects are TRUE in the case
          forwarding for the interface is enabled

        or

        - at least one of conf/{all,interface}/accept_redirects is TRUE in the
          case forwarding for the interface is disabled

        accept_redirects for the interface will be disabled otherwise

        default:

                - TRUE (host)
                - FALSE (router)

forwarding - BOOLEAN
        Enable IP forwarding on this interface.  This controls whether packets
        received _on_ this interface can be forwarded.

mc_forwarding - BOOLEAN
        Do multicast routing. The kernel needs to be compiled with CONFIG_MROUTE
        and a multicast routing daemon is required.
        conf/all/mc_forwarding must also be set to TRUE to enable multicast
        routing        for the interface

medium_id - INTEGER
        Integer value used to differentiate the devices by the medium they
        are attached to. Two devices can have different id values when
        the broadcast packets are received only on one of them.
        The default value 0 means that the device is the only interface
        to its medium, value of -1 means that medium is not known.

        Currently, it is used to change the proxy_arp behavior:
        the proxy_arp feature is enabled for packets forwarded between
        two devices attached to different media.

proxy_arp - BOOLEAN
        Do proxy arp.

        proxy_arp for the interface will be enabled if at least one of
        conf/{all,interface}/proxy_arp is set to TRUE,
        it will be disabled otherwise

proxy_arp_pvlan - BOOLEAN
        Private VLAN proxy arp.

        Basically allow proxy arp replies back to the same interface
        (from which the ARP request/solicitation was received).

        This is done to support (ethernet) switch features, like RFC
        3069, where the individual ports are NOT allowed to
        communicate with each other, but they are allowed to talk to
        the upstream router.  As described in RFC 3069, it is possible
        to allow these hosts to communicate through the upstream
        router by proxy_arp'ing. Don't need to be used together with
        proxy_arp.

        This technology is known by different names:

        - In RFC 3069 it is called VLAN Aggregation.
        - Cisco and Allied Telesyn call it Private VLAN.
        - Hewlett-Packard call it Source-Port filtering or port-isolation.
        - Ericsson call it MAC-Forced Forwarding (RFC Draft).

proxy_delay - INTEGER
        Delay proxy response.

        Delay response to a neighbor solicitation when proxy_arp
        or proxy_ndp is enabled. A random value between [0, proxy_delay)
        will be chosen, setting to zero means reply with no delay.
        Value in jiffies. Defaults to 80.

shared_media - BOOLEAN
        Send(router) or accept(host) RFC1620 shared media redirects.
        Overrides secure_redirects.

        shared_media for the interface will be enabled if at least one of
        conf/{all,interface}/shared_media is set to TRUE,
        it will be disabled otherwise

        default TRUE

secure_redirects - BOOLEAN
        Accept ICMP redirect messages only to gateways listed in the
        interface's current gateway list. Even if disabled, RFC1122 redirect
        rules still apply.

        Overridden by shared_media.

        secure_redirects for the interface will be enabled if at least one of
        conf/{all,interface}/secure_redirects is set to TRUE,
        it will be disabled otherwise

        default TRUE

send_redirects - BOOLEAN
        Send redirects, if router.

        send_redirects for the interface will be enabled if at least one of
        conf/{all,interface}/send_redirects is set to TRUE,
        it will be disabled otherwise

        Default: TRUE

bootp_relay - BOOLEAN
        Accept packets with source address 0.b.c.d destined
        not to this host as local ones. It is supposed, that
        BOOTP relay daemon will catch and forward such packets.
        conf/all/bootp_relay must also be set to TRUE to enable BOOTP relay
        for the interface

        default FALSE

        Not Implemented Yet.

accept_source_route - BOOLEAN
        Accept packets with SRR option.
        conf/all/accept_source_route must also be set to TRUE to accept packets
        with SRR option on the interface

        default

                - TRUE (router)
                - FALSE (host)

accept_local - BOOLEAN
        Accept packets with local source addresses. In combination with
        suitable routing, this can be used to direct packets between two
        local interfaces over the wire and have them accepted properly.
        default FALSE

route_localnet - BOOLEAN
        Do not consider loopback addresses as martian source or destination
        while routing. This enables the use of 127/8 for local routing purposes.

        default FALSE

rp_filter - INTEGER
        - 0 - No source validation.
        - 1 - Strict mode as defined in RFC3704 Strict Reverse Path
          Each incoming packet is tested against the FIB and if the interface
          is not the best reverse path the packet check will fail.
          By default failed packets are discarded.
        - 2 - Loose mode as defined in RFC3704 Loose Reverse Path
          Each incoming packet's source address is also tested against the FIB
          and if the source address is not reachable via any interface
          the packet check will fail.

        Current recommended practice in RFC3704 is to enable strict mode
        to prevent IP spoofing from DDos attacks. If using asymmetric routing
        or other complicated routing, then loose mode is recommended.

        The max value from conf/{all,interface}/rp_filter is used
        when doing source validation on the {interface}.

        Default value is 0. Note that some distributions enable it
        in startup scripts.

src_valid_mark - BOOLEAN
        - 0 - The fwmark of the packet is not included in reverse path
          route lookup.  This allows for asymmetric routing configurations
          utilizing the fwmark in only one direction, e.g., transparent
          proxying.

        - 1 - The fwmark of the packet is included in reverse path route
          lookup.  This permits rp_filter to function when the fwmark is
          used for routing traffic in both directions.

        This setting also affects the utilization of fmwark when
        performing source address selection for ICMP replies, or
        determining addresses stored for the IPOPT_TS_TSANDADDR and
        IPOPT_RR IP options.

        The max value from conf/{all,interface}/src_valid_mark is used.

        Default value is 0.

IPv4 ARP 응답과 neighbor probe

2045-2191

`arp_filter`는 같은 subnet의 여러 interface 중 route lookup상 outgoing interface인 card만 ARP에 응답하게 하며 source-based routing이 필요합니다. 기본 0에서는 Linux host 전체가 IP address를 소유한다는 모델에 따라 다른 interface 주소에도 응답할 수 있습니다. `arp_announce`는 ARP request sender address 선택을 0=아무 local address, 1=target subnet 밖 주소 회피, 2=outgoing interface에서 target에 가장 적합한 주소로 제한합니다.

`arp_ignore`는 reply 범위를 0=어느 local 주소든, 1=incoming interface 주소만, 2=같은 subnet까지, 3=host-scope 제외, 8=모두 무시로 강화합니다. 4~7은 예약입니다. `arp_notify`는 device up/MAC 변경 때 gratuitous ARP를 보내고 `arp_accept`는 cache에 없는 gratuitous ARP를 0=새 entry 없음, 1=허용, 2=같은 subnet만 허용으로 정합니다. 기존 entry는 설정과 무관하게 갱신됩니다.

`arp_evict_nocarrier`는 carrier가 사라질 때 ARP cache를 비우며 기본 1입니다. 같은 network의 AP 사이를 roaming하는 wireless 환경에서는 0이 유용할 수 있습니다. `mcast_solicit`, `ucast_solicit`, `app_solicit`, `mcast_resolicit`은 INCOMPLETE/PROBE 상태에서 multicast, unicast, userspace ARP daemon, 재multicast probe의 최대 횟수입니다. `disable_policy`와 `disable_xfrm`은 각각 interface의 IPsec SPD와 encryption을 끕니다.

arp_filter - BOOLEAN
        - 1 - Allows you to have multiple network interfaces on the same
          subnet, and have the ARPs for each interface be answered
          based on whether or not the kernel would route a packet from
          the ARP'd IP out that interface (therefore you must use source
          based routing for this to work). In other words it allows control
          of which cards (usually 1) will respond to an arp request.

        - 0 - (default) The kernel can respond to arp requests with addresses
          from other interfaces. This may seem wrong but it usually makes
          sense, because it increases the chance of successful communication.
          IP addresses are owned by the complete host on Linux, not by
          particular interfaces. Only for more complex setups like load-
          balancing, does this behaviour cause problems.

        arp_filter for the interface will be enabled if at least one of
        conf/{all,interface}/arp_filter is set to TRUE,
        it will be disabled otherwise

arp_announce - INTEGER
        Define different restriction levels for announcing the local
        source IP address from IP packets in ARP requests sent on
        interface:

        - 0 - (default) Use any local address, configured on any interface
        - 1 - Try to avoid local addresses that are not in the target's
          subnet for this interface. This mode is useful when target
          hosts reachable via this interface require the source IP
          address in ARP requests to be part of their logical network
          configured on the receiving interface. When we generate the
          request we will check all our subnets that include the
          target IP and will preserve the source address if it is from
          such subnet. If there is no such subnet we select source
          address according to the rules for level 2.
        - 2 - Always use the best local address for this target.
          In this mode we ignore the source address in the IP packet
          and try to select local address that we prefer for talks with
          the target host. Such local address is selected by looking
          for primary IP addresses on all our subnets on the outgoing
          interface that include the target IP address. If no suitable
          local address is found we select the first local address
          we have on the outgoing interface or on all other interfaces,
          with the hope we will receive reply for our request and
          even sometimes no matter the source IP address we announce.

        The max value from conf/{all,interface}/arp_announce is used.

        Increasing the restriction level gives more chance for
        receiving answer from the resolved target while decreasing
        the level announces more valid sender's information.

arp_ignore - INTEGER
        Define different modes for sending replies in response to
        received ARP requests that resolve local target IP addresses:

        - 0 - (default): reply for any local target IP address, configured
          on any interface
        - 1 - reply only if the target IP address is local address
          configured on the incoming interface
        - 2 - reply only if the target IP address is local address
          configured on the incoming interface and both with the
          sender's IP address are part from same subnet on this interface
        - 3 - do not reply for local addresses configured with scope host,
          only resolutions for global and link addresses are replied
        - 4-7 - reserved
        - 8 - do not reply for all local addresses

        The max value from conf/{all,interface}/arp_ignore is used
        when ARP request is received on the {interface}

arp_notify - BOOLEAN
        Define mode for notification of address and device changes.

         ==  ==========================================================
          0  (default): do nothing
          1  Generate gratuitous arp requests when device is brought up
             or hardware address changes.
         ==  ==========================================================

arp_accept - INTEGER
        Define behavior for accepting gratuitous ARP (garp) frames from devices
        that are not already present in the ARP table:

        - 0 - don't create new entries in the ARP table
        - 1 - create new entries in the ARP table
        - 2 - create new entries only if the source IP address is in the same
          subnet as an address configured on the interface that received the
          garp message.

        Both replies and requests type gratuitous arp will trigger the
        ARP table to be updated, if this setting is on.

        If the ARP table already contains the IP address of the
        gratuitous arp frame, the arp table will be updated regardless
        if this setting is on or off.

arp_evict_nocarrier - BOOLEAN
        Clears the ARP cache on NOCARRIER events. This option is important for
        wireless devices where the ARP cache should not be cleared when roaming
        between access points on the same network. In most cases this should
        remain as the default (1).

        Possible values:

        - 0 (disabled) - Do not clear ARP cache on NOCARRIER events
        - 1 (enabled)  - Clear the ARP cache on NOCARRIER events

        Default: 1 (enabled)

mcast_solicit - INTEGER
        The maximum number of multicast probes in INCOMPLETE state,
        when the associated hardware address is unknown.  Defaults
        to 3.

ucast_solicit - INTEGER
        The maximum number of unicast probes in PROBE state, when
        the hardware address is being reconfirmed.  Defaults to 3.

app_solicit - INTEGER
        The maximum number of probes to send to the user space ARP daemon
        via netlink before dropping back to multicast probes (see
        mcast_resolicit).  Defaults to 0.

mcast_resolicit - INTEGER
        The maximum number of multicast probes after unicast and
        app probes in PROBE state.  Defaults to 0.

disable_policy - BOOLEAN
        Disable IPSEC policy (SPD) for this interface

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

disable_xfrm - BOOLEAN
        Disable IPSEC encryption on this interface, whatever the policy

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

IPv4 interface 기타 정책

2192-2282

`igmpv2_unsolicited_report_interval`은 IGMPv1/v2 report 재전송 간격 10초, `igmpv3_unsolicited_report_interval`은 IGMPv3 간격 1초가 기본입니다. `ignore_routes_with_linkdown`은 down link route를 FIB lookup에서 제외합니다. `promote_secondaries`는 primary address 삭제 때 관련 secondary를 함께 없애는 대신 하나를 primary로 승격합니다.

`drop_unicast_in_l2_multicast`는 L2 multicast/broadcast frame에 실린 unicast IPv4를 폐기하고 `drop_gratuitous_arp`는 모든 gratuitous ARP를 버립니다. 둘 다 compatibility 때문에 기본 0입니다. `tag`는 자유롭게 쓸 수 있는 정수 tag입니다. `xfrm4_gc_thresh`는 4.14부터 폐기되었습니다. `igmp_link_local_mcast_reports`는 224.0.0.X link-local group에도 IGMP report를 보내며 기본 true입니다. 마지막에는 원 작성자와 갱신자 credit이 기록되어 있습니다.

igmpv2_unsolicited_report_interval - INTEGER
        The interval in milliseconds in which the next unsolicited
        IGMPv1 or IGMPv2 report retransmit will take place.

        Default: 10000 (10 seconds)

igmpv3_unsolicited_report_interval - INTEGER
        The interval in milliseconds in which the next unsolicited
        IGMPv3 report retransmit will take place.

        Default: 1000 (1 seconds)

ignore_routes_with_linkdown - BOOLEAN
        Ignore routes whose link is down when performing a FIB lookup.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

promote_secondaries - BOOLEAN
        When a primary IP address is removed from this interface
        promote a corresponding secondary IP address instead of
        removing all the corresponding secondary IP addresses.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

drop_unicast_in_l2_multicast - BOOLEAN
        Drop any unicast IP packets that are received in link-layer
        multicast (or broadcast) frames.

        This behavior (for multicast) is actually a SHOULD in RFC
        1122, but is disabled by default for compatibility reasons.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

drop_gratuitous_arp - BOOLEAN
        Drop all gratuitous ARP frames, for example if there's a known
        good ARP proxy on the network and such frames need not be used
        (or in the case of 802.11, must not be used to prevent attacks.)

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)


tag - INTEGER
        Allows you to write a number, which can be used as required.

        Default value is 0.

xfrm4_gc_thresh - INTEGER
        (Obsolete since linux-4.14)
        The threshold at which we will start garbage collecting for IPv4
        destination cache entries.  At twice this value the system will
        refuse new allocations.

igmp_link_local_mcast_reports - BOOLEAN
        Enable IGMP reports for link local multicast groups in the
        224.0.0.X range.

        Default TRUE

Alexey Kuznetsov.
[email protected]

Updated by:

- Andi Kleen
  [email protected]
- Nicolas Delon
  [email protected]



IPv6 socket, flow label과 ECMP

2283-2406

IPv6에는 별도 global `tcp_*`가 없고 `ipv4/` 아래 TCP 설정이 IPv6에도 적용됩니다. `bindv6only`는 `IPV6_V6ONLY` 기본값이며 0이면 IPv4-mapped address를 허용합니다. `flowlabel_consistency`는 flow label의 일관성과 유일성을 보호하고, `auto_flowlabels`는 packet flow hash로 label을 0=완전 끔, 1=기본 켬/socket에서 끌 수 있음, 2=기본 끔/socket에서 켤 수 있음, 3=강제 켬으로 정합니다.

`flowlabel_state_ranges`는 0~0x7ffff를 flow manager, 0x80000~0xfffff를 RFC 6437 stateless label에 예약합니다. `flowlabel_reflect`는 bit 1=established flow, 2=closed port TCP RST, 4=ICMPv6 echo reply에 incoming label을 반영해 anycast ECMP에서 PMTUD가 같은 path를 타게 합니다.

IPv6 `fib_multipath_hash_policy`는 0=L3 주소+flow label, 1=L4 5-tuple, 2=inner L3 우선, 3=custom입니다. custom `fib_multipath_hash_fields`는 outer/inner source·destination IP, protocol, flow label, port bit를 조합하고 기본 0x7은 source IP, destination IP, protocol입니다.

/proc/sys/net/ipv6/* Variables
==============================

IPv6 has no global variables such as tcp_*.  tcp_* settings under ipv4/ also
apply to IPv6 [XXX?].

bindv6only - BOOLEAN
        Default value for IPV6_V6ONLY socket option,
        which restricts use of the IPv6 socket to IPv6 communication
        only.

        Possible values:

        - 0 (disabled) - enable IPv4-mapped address feature
        - 1 (enabled)  - disable IPv4-mapped address feature

        Default: 0 (disabled)

flowlabel_consistency - BOOLEAN
        Protect the consistency (and unicity) of flow label.
        You have to disable it to use IPV6_FL_F_REFLECT flag on the
        flow label manager.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

auto_flowlabels - INTEGER
        Automatically generate flow labels based on a flow hash of the
        packet. This allows intermediate devices, such as routers, to
        identify packet flows for mechanisms like Equal Cost Multipath
        Routing (see RFC 6438).

        =  ===========================================================
        0  automatic flow labels are completely disabled
        1  automatic flow labels are enabled by default, they can be
           disabled on a per socket basis using the IPV6_AUTOFLOWLABEL
           socket option
        2  automatic flow labels are allowed, they may be enabled on a
           per socket basis using the IPV6_AUTOFLOWLABEL socket option
        3  automatic flow labels are enabled and enforced, they cannot
           be disabled by the socket option
        =  ===========================================================

        Default: 1

flowlabel_state_ranges - BOOLEAN
        Split the flow label number space into two ranges. 0-0x7FFFF is
        reserved for the IPv6 flow manager facility, 0x80000-0xFFFFF
        is reserved for stateless flow labels as described in RFC6437.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)


flowlabel_reflect - INTEGER
        Control flow label reflection. Needed for Path MTU
        Discovery to work with Equal Cost Multipath Routing in anycast
        environments. See RFC 7690 and:
        https://tools.ietf.org/html/draft-wang-6man-flow-label-reflection-01

        This is a bitmask.

        - 1: enabled for established flows

          Note that this prevents automatic flowlabel changes, as done
          in "tcp: change IPv6 flow-label upon receiving spurious retransmission"
          and "tcp: Change txhash on every SYN and RTO retransmit"

        - 2: enabled for TCP RESET packets (no active listener)
          If set, a RST packet sent in response to a SYN packet on a closed
          port will reflect the incoming flow label.

        - 4: enabled for ICMPv6 echo reply messages.

        Default: 0

fib_multipath_hash_policy - INTEGER
        Controls which hash policy to use for multipath routes.

        Default: 0 (Layer 3)

        Possible values:

        - 0 - Layer 3 (source and destination addresses plus flow label)
        - 1 - Layer 4 (standard 5-tuple)
        - 2 - Layer 3 or inner Layer 3 if present
        - 3 - Custom multipath hash. Fields used for multipath hash calculation
          are determined by fib_multipath_hash_fields sysctl

fib_multipath_hash_fields - UNSIGNED INTEGER
        When fib_multipath_hash_policy is set to 3 (custom multipath hash), the
        fields used for multipath hash calculation are determined by this
        sysctl.

        This value is a bitmask which enables various fields for multipath hash
        calculation.

        Possible fields are:

        ====== ============================
        0x0001 Source IP address
        0x0002 Destination IP address
        0x0004 IP protocol
        0x0008 Flow Label
        0x0010 Source port
        0x0020 Destination port
        0x0040 Inner source IP address
        0x0080 Inner destination IP address
        0x0100 Inner IP protocol
        0x0200 Inner Flow Label
        0x0400 Inner source port
        0x0800 Inner destination port
        ====== ============================

        Default: 0x0007 (source IP, destination IP and IP protocol)

IPv6 option 제한, route 알림과 fragment

2407-2586

`anycast_src_echo_reply`는 ICMPv6 echo reply source로 anycast address를 쓸지 정합니다. `idgen_delay`와 `idgen_retries`는 RFC 7217 stable privacy address가 DAD 충돌했을 때 재시도 간격과 횟수로 기본 1초, 3회입니다. `mld_qrv`의 기본은 2입니다.

`max_dst_opts_number`와 `max_hbh_opts_number`는 Destination/Hop-by-Hop extension header의 non-padding TLV 수를 제한합니다. 음수면 unknown option을 금지하고 절댓값만큼 known TLV를 허용합니다. `max_dst_opts_length`와 `max_hbh_length`는 header 길이 상한이며 기본 `INT_MAX`입니다.

`skip_notify_on_dev_down`은 device down/delete로 route가 사라질 때 IPv6 `RTM_DELROUTE`를 생략해 IPv4와 같은 동작으로 만듭니다. `nexthop_compat_mode`는 새 nexthop API를 옛 route dump 형식으로 펼치고 변경 route마다 알림을 보내는 호환 모드입니다. 끄면 성능은 좋아지지만 old client 호환이 사라집니다. resilient group과 8-bit를 넘는 weight는 호환 dump에서 부정확할 수 있습니다. `fib_notify_on_flag_change`는 IPv4와 같은 0/1/2 offload flag 알림 정책입니다.

`ioam6_id`는 24-bit node ID, `ioam6_id_wide`는 56-bit wide ID입니다. IPv6 fragmentation의 `ip6frag_high_thresh`, 폐기 기준인 `ip6frag_low_thresh`, 보존 초인 `ip6frag_time`을 제공합니다. `conf/default/*`은 새 interface 기본값, `conf/all/*`은 모든 interface 설정입니다. `conf/all/disable_ipv6` 읽기는 전체 활성 여부를 확정하지 못하고, global `forwarding`은 모든 interface의 host/router forwarding을 함께 바꿉니다.

anycast_src_echo_reply - BOOLEAN
        Controls the use of anycast addresses as source addresses for ICMPv6
        echo reply

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)


idgen_delay - INTEGER
        Controls the delay in seconds after which time to retry
        privacy stable address generation if a DAD conflict is
        detected.

        Default: 1 (as specified in RFC7217)

idgen_retries - INTEGER
        Controls the number of retries to generate a stable privacy
        address if a DAD conflict is detected.

        Default: 3 (as specified in RFC7217)

mld_qrv - INTEGER
        Controls the MLD query robustness variable (see RFC3810 9.1).

        Default: 2 (as specified by RFC3810 9.1)

        Minimum: 1 (as specified by RFC6636 4.5)

max_dst_opts_number - INTEGER
        Maximum number of non-padding TLVs allowed in a Destination
        options extension header. If this value is less than zero
        then unknown options are disallowed and the number of known
        TLVs allowed is the absolute value of this number.

        Default: 8

max_hbh_opts_number - INTEGER
        Maximum number of non-padding TLVs allowed in a Hop-by-Hop
        options extension header. If this value is less than zero
        then unknown options are disallowed and the number of known
        TLVs allowed is the absolute value of this number.

        Default: 8

max_dst_opts_length - INTEGER
        Maximum length allowed for a Destination options extension
        header.

        Default: INT_MAX (unlimited)

max_hbh_length - INTEGER
        Maximum length allowed for a Hop-by-Hop options extension
        header.

        Default: INT_MAX (unlimited)

skip_notify_on_dev_down - BOOLEAN
        Controls whether an RTM_DELROUTE message is generated for routes
        removed when a device is taken down or deleted. IPv4 does not
        generate this message; IPv6 does by default. Setting this sysctl
        to true skips the message, making IPv4 and IPv6 on par in relying
        on userspace caches to track link events and evict routes.

        Possible values:

        - 0 (disabled) - generate the message
        - 1 (enabled)  - skip generating the message

        Default: 0 (disabled)

nexthop_compat_mode - BOOLEAN
        New nexthop API provides a means for managing nexthops independent of
        prefixes. Backwards compatibility with old route format is enabled by
        default which means route dumps and notifications contain the new
        nexthop attribute but also the full, expanded nexthop definition.
        Further, updates or deletes of a nexthop configuration generate route
        notifications for each fib entry using the nexthop. Once a system
        understands the new API, this sysctl can be disabled to achieve full
        performance benefits of the new API by disabling the nexthop expansion
        and extraneous notifications.

        Note that as a backward-compatible mode, dumping of modern features
        might be incomplete or wrong. For example, resilient groups will not be
        shown as such, but rather as just a list of next hops. Also weights that
        do not fit into 8 bits will show incorrectly.

        Default: true (backward compat mode)

fib_notify_on_flag_change - INTEGER
        Whether to emit RTM_NEWROUTE notifications whenever RTM_F_OFFLOAD/
        RTM_F_TRAP/RTM_F_OFFLOAD_FAILED flags are changed.

        After installing a route to the kernel, user space receives an
        acknowledgment, which means the route was installed in the kernel,
        but not necessarily in hardware.
        It is also possible for a route already installed in hardware to change
        its action and therefore its flags. For example, a host route that is
        trapping packets can be "promoted" to perform decapsulation following
        the installation of an IPinIP/VXLAN tunnel.
        The notifications will indicate to user-space the state of the route.

        Default: 0 (Do not emit notifications.)

        Possible values:

        - 0 - Do not emit notifications.
        - 1 - Emit notifications.
        - 2 - Emit notifications only for RTM_F_OFFLOAD_FAILED flag change.

ioam6_id - INTEGER
        Define the IOAM id of this node. Uses only 24 bits out of 32 in total.

        Possible value range:

        - Min: 0
        - Max: 0xFFFFFF

        Default: 0xFFFFFF

ioam6_id_wide - LONG INTEGER
        Define the wide IOAM id of this node. Uses only 56 bits out of 64 in
        total. Can be different from ioam6_id.

        Possible value range:

        - Min: 0
        - Max: 0xFFFFFFFFFFFFFF

        Default: 0xFFFFFFFFFFFFFF

IPv6 Fragmentation:

ip6frag_high_thresh - INTEGER
        Maximum memory used to reassemble IPv6 fragments. When
        ip6frag_high_thresh bytes of memory is allocated for this purpose,
        the fragment handler will toss packets until ip6frag_low_thresh
        is reached.

ip6frag_low_thresh - INTEGER
        See ip6frag_high_thresh

ip6frag_time - INTEGER
        Time in seconds to keep an IPv6 fragment in memory.

``conf/default/*``:
        Change the interface-specific default settings.

        These settings would be used during creating new interfaces.


``conf/all/*``:
        Change all the interface-specific settings.

        [XXX:  Other special features than forwarding?]

conf/all/disable_ipv6 - BOOLEAN
        Changing this value is same as changing ``conf/default/disable_ipv6``
        setting and also all per-interface ``disable_ipv6`` settings to the same
        value.

        Reading this value does not have any particular meaning. It does not say
        whether IPv6 support is enabled or disabled. Returned value can be 1
        also in the case when some interface has ``disable_ipv6`` set to 0 and
        has configured IPv6 addresses.

conf/all/forwarding - BOOLEAN
        Enable global IPv6 forwarding between all interfaces.

        IPv4 and IPv6 work differently here; the ``force_forwarding`` flag must
        be used to control which interfaces may forward packets.

        This also sets all interfaces' Host/Router setting
        'forwarding' to the specified value.  See below for details.

        This referred to as global forwarding.

IPv6 forwarding과 Router Advertisement

2587-2796

`proxy_ndp`는 proxy Neighbor Discovery, `force_forwarding`은 global forwarding과 무관하게 해당 interface만 전달하게 합니다. global forwarding을 0으로 바꾸면 모든 force flag도 reset됩니다. IPv6 `fwmark_reflect`는 socket 없는 RST/ICMPv6 reply가 incoming fwmark를 상속하게 합니다.

interface별 `accept_ra`는 0=RA 거부, 1=forwarding이 꺼졌을 때 수락, 2=forwarding 중에도 강제 수락입니다. 이 값이 실제로 RA를 수락할 때만 Router Solicitation도 보냅니다. `accept_ra_defrtr`는 default router 학습, `ra_defrtr_metric`은 학습 route metric(기본 1024), `accept_ra_from_local`은 local source RA 허용, `accept_ra_min_hop_limit`과 `accept_ra_min_lft`는 받아들일 hop limit·lifetime 하한입니다.

`accept_ra_pinfo`는 Prefix Information 학습입니다. `ra_honor_pio_life`가 0이면 RFC 4862 5.5.3e lifetime 규칙, 1이면 PIO lifetime을 그대로 존중합니다. `ra_honor_pio_pflag`는 DHCPv6-PD client가 P=1인 PIO의 A flag 효과를 억제하게 합니다. `accept_ra_rt_info_min_plen/max_plen`은 Route Information prefix 길이 범위, `accept_ra_rtr_pref`는 router preference, `accept_ra_mtu`는 RA option 5 MTU 적용을 제어합니다.

`accept_redirects`는 host mode에서 기본 허용, forwarding에서는 기본 거부입니다. `accept_source_route`는 음수면 routing header를 모두 거부하고 0 이상이면 type 2만 받습니다. `autoconf`는 PIO 기반 address 자동 구성을, `dad_transmits`는 DAD probe 횟수를 정합니다.

proxy_ndp - BOOLEAN
        Do proxy ndp.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

force_forwarding - BOOLEAN
        Enable forwarding on this interface only -- regardless of the setting on
        ``conf/all/forwarding``. When setting ``conf.all.forwarding`` to 0,
        the ``force_forwarding`` flag will be reset on all interfaces.

fwmark_reflect - BOOLEAN
        Controls the fwmark of kernel-generated IPv6 reply packets that are not
        associated with a socket for example, TCP RSTs or ICMPv6 echo replies).
        If disabled, these packets have a fwmark of zero. If enabled, they have the
        fwmark of the packet they are replying to.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

``conf/interface/*``:
        Change special settings per interface.

        The functional behaviour for certain settings is different
        depending on whether local forwarding is enabled or not.

accept_ra - INTEGER
        Accept Router Advertisements; autoconfigure using them.

        It also determines whether or not to transmit Router
        Solicitations. If and only if the functional setting is to
        accept Router Advertisements, Router Solicitations will be
        transmitted.

        Possible values are:

                ==  ===========================================================
                 0  Do not accept Router Advertisements.
                 1  Accept Router Advertisements if forwarding is disabled.
                 2  Overrule forwarding behaviour. Accept Router Advertisements
                    even if forwarding is enabled.
                ==  ===========================================================

        Functional default:

                - enabled if local forwarding is disabled.
                - disabled if local forwarding is enabled.

accept_ra_defrtr - BOOLEAN
        Learn default router in Router Advertisement.

        Functional default:

                - enabled if accept_ra is enabled.
                - disabled if accept_ra is disabled.

ra_defrtr_metric - UNSIGNED INTEGER
        Route metric for default route learned in Router Advertisement. This value
        will be assigned as metric for the default route learned via IPv6 Router
        Advertisement. Takes affect only if accept_ra_defrtr is enabled.

        Possible values:
                1 to 0xFFFFFFFF

                Default: IP6_RT_PRIO_USER i.e. 1024.

accept_ra_from_local - BOOLEAN
        Accept RA with source-address that is found on local machine
        if the RA is otherwise proper and able to be accepted.

        Default is to NOT accept these as it may be an un-intended
        network loop.

        Functional default:

           - enabled if accept_ra_from_local is enabled
             on a specific interface.
           - disabled if accept_ra_from_local is disabled
             on a specific interface.

accept_ra_min_hop_limit - INTEGER
        Minimum hop limit Information in Router Advertisement.

        Hop limit Information in Router Advertisement less than this
        variable shall be ignored.

        Default: 1

accept_ra_min_lft - INTEGER
        Minimum acceptable lifetime value in Router Advertisement.

        RA sections with a lifetime less than this value shall be
        ignored. Zero lifetimes stay unaffected.

        Default: 0

accept_ra_pinfo - BOOLEAN
        Learn Prefix Information in Router Advertisement.

        Functional default:

                - enabled if accept_ra is enabled.
                - disabled if accept_ra is disabled.

ra_honor_pio_life - BOOLEAN
        Whether to use RFC4862 Section 5.5.3e to determine the valid
        lifetime of an address matching a prefix sent in a Router
        Advertisement Prefix Information Option.

        Possible values:

        - 0 (disabled) - RFC4862 section 5.5.3e is used to determine
          the valid lifetime of the address.
        - 1 (enabled)  - the PIO valid lifetime will always be honored.

        Default: 0 (disabled)

ra_honor_pio_pflag - BOOLEAN
        The Prefix Information Option P-flag indicates the network can
        allocate a unique IPv6 prefix per client using DHCPv6-PD.
        This sysctl can be enabled when a userspace DHCPv6-PD client
        is running to cause the P-flag to take effect: i.e. the
        P-flag suppresses any effects of the A-flag within the same
        PIO. For a given PIO, P=1 and A=1 is treated as A=0.

        Possible values:

        - 0 (disabled) - the P-flag is ignored.
        - 1 (enabled)  - the P-flag will disable SLAAC autoconfiguration
          for the given Prefix Information Option.

        Default: 0 (disabled)

accept_ra_rt_info_min_plen - INTEGER
        Minimum prefix length of Route Information in RA.

        Route Information w/ prefix smaller than this variable shall
        be ignored.

        Functional default:

                * 0 if accept_ra_rtr_pref is enabled.
                * -1 if accept_ra_rtr_pref is disabled.

accept_ra_rt_info_max_plen - INTEGER
        Maximum prefix length of Route Information in RA.

        Route Information w/ prefix larger than this variable shall
        be ignored.

        Functional default:

                * 0 if accept_ra_rtr_pref is enabled.
                * -1 if accept_ra_rtr_pref is disabled.

accept_ra_rtr_pref - BOOLEAN
        Accept Router Preference in RA.

        Functional default:

                - enabled if accept_ra is enabled.
                - disabled if accept_ra is disabled.

accept_ra_mtu - BOOLEAN
        Apply the MTU value specified in RA option 5 (RFC4861). If
        disabled, the MTU specified in the RA will be ignored.

        Functional default:

                - enabled if accept_ra is enabled.
                - disabled if accept_ra is disabled.

accept_redirects - BOOLEAN
        Accept Redirects.

        Functional default:

                - enabled if local forwarding is disabled.
                - disabled if local forwarding is enabled.

accept_source_route - INTEGER
        Accept source routing (routing extension header).

        - >= 0: Accept only routing header type 2.
        - < 0: Do not accept routing header.

        Default: 0

autoconf - BOOLEAN
        Autoconfigure addresses using Prefix Information in Router
        Advertisements.

        Functional default:

                - enabled if accept_ra_pinfo is enabled.
                - disabled if accept_ra_pinfo is disabled.

dad_transmits - INTEGER
        The amount of Duplicate Address Detection probes to send.

        Default: 1

IPv6 host/router 상태와 privacy address

2797-2990

interface `forwarding=0`은 host mode로 NA의 IsRouter flag를 지우고 RA/redirect를 수락하며 RS를 보냅니다. `forwarding=1`은 반대로 router mode가 되어 IsRouter를 세우고 `accept_ra=2`가 아니면 RS를 보내지 않고 RA를 무시하며 redirect도 무시합니다. 모든 interface에서 같은 mode를 쓰는 것이 권장됩니다. `hop_limit` 기본 64, IPv6 `mtu` 기본·최소는 1280입니다.

IPv6 `ip_nonlocal_bind`는 non-local IPv6 address bind를 허용합니다. Router probing/solicitation은 `router_probe_interval` 60초, `router_solicitation_delay` 1초, `router_solicitation_interval` 4초, `router_solicitations` 3회가 기본입니다. `use_oif_addrs_only`는 source candidate를 outgoing interface 주소로 제한합니다.

`use_tempaddr`는 0 이하=privacy extension 끔, 1=켜되 public 우선, 1보다 큼=temporary 우선입니다. `temp_valid_lft` 기본 2일, `temp_prefered_lft` 기본 1일이며 valid보다 길 수 없습니다. `keep_addr_on_down`은 interface down 때 static global 주소를 보존할지 정합니다. `max_desync_factor`는 host 간 주소 재생성 시점 분산, `regen_min_advance`는 deprecate 전에 새 temporary address를 만들 최소 선행 시간, `regen_max_retry`는 생성 시도 횟수입니다.

`max_addresses`는 interface별 autoconfigured address 상한이며 기본 16, 0은 무제한이지만 memory DoS 위험이 있습니다. `disable_ipv6`를 1로 바꾸면 해당 interface의 IPv6 주소와 route를 지우고 새 구성을 막으며 0으로 되돌리면 link-local address와 DAD를 시작합니다. `accept_dad`는 0=끔, 1=켬, 2=MAC 기반 link-local 중복이면 IPv6까지 끔이고 all/interface 중 큰 값을 씁니다.

forwarding - INTEGER
        Configure interface-specific Host/Router behaviour.

        .. note::

           It is recommended to have the same setting on all
           interfaces; mixed router/host scenarios are rather uncommon.

        Possible values are:

                - 0 Forwarding disabled
                - 1 Forwarding enabled

        **FALSE (0)**:

        By default, Host behaviour is assumed.  This means:

        1. IsRouter flag is not set in Neighbour Advertisements.
        2. If accept_ra is TRUE (default), transmit Router
           Solicitations.
        3. If accept_ra is TRUE (default), accept Router
           Advertisements (and do autoconfiguration).
        4. If accept_redirects is TRUE (default), accept Redirects.

        **TRUE (1)**:

        If local forwarding is enabled, Router behaviour is assumed.
        This means exactly the reverse from the above:

        1. IsRouter flag is set in Neighbour Advertisements.
        2. Router Solicitations are not sent unless accept_ra is 2.
        3. Router Advertisements are ignored unless accept_ra is 2.
        4. Redirects are ignored.

        Default: 0 (disabled) if global forwarding is disabled (default),
        otherwise 1 (enabled).

hop_limit - INTEGER
        Default Hop Limit to set.

        Default: 64

mtu - INTEGER
        Default Maximum Transfer Unit

        Default: 1280 (IPv6 required minimum)

ip_nonlocal_bind - BOOLEAN
        If enabled, allows processes to bind() to non-local IPv6 addresses,
        which can be quite useful - but may break some applications.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

router_probe_interval - INTEGER
        Minimum interval (in seconds) between Router Probing described
        in RFC4191.

        Default: 60

router_solicitation_delay - INTEGER
        Number of seconds to wait after interface is brought up
        before sending Router Solicitations.

        Default: 1

router_solicitation_interval - INTEGER
        Number of seconds to wait between Router Solicitations.

        Default: 4

router_solicitations - INTEGER
        Number of Router Solicitations to send until assuming no
        routers are present.

        Default: 3

use_oif_addrs_only - BOOLEAN
        When enabled, the candidate source addresses for destinations
        routed via this interface are restricted to the set of addresses
        configured on this interface (vis. RFC 6724, section 4).

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

use_tempaddr - INTEGER
        Preference for Privacy Extensions (RFC3041).

          * <= 0 : disable Privacy Extensions
          * == 1 : enable Privacy Extensions, but prefer public
            addresses over temporary addresses.
          * >  1 : enable Privacy Extensions and prefer temporary
            addresses over public addresses.

        Default:

                * 0 (for most devices)
                * -1 (for point-to-point devices and loopback devices)

temp_valid_lft - INTEGER
        valid lifetime (in seconds) for temporary addresses. If less than the
        minimum required lifetime (typically 5-7 seconds), temporary addresses
        will not be created.

        Default: 172800 (2 days)

temp_prefered_lft - INTEGER
        Preferred lifetime (in seconds) for temporary addresses. If
        temp_prefered_lft is less than the minimum required lifetime (typically
        5-7 seconds), the preferred lifetime is the minimum required. If
        temp_prefered_lft is greater than temp_valid_lft, the preferred lifetime
        is temp_valid_lft.

        Default: 86400 (1 day)

keep_addr_on_down - INTEGER
        Keep all IPv6 addresses on an interface down event. If set static
        global addresses with no expiration time are not flushed.

        *   >0 : enabled
        *    0 : system default
        *   <0 : disabled

        Default: 0 (addresses are removed)

max_desync_factor - INTEGER
        Maximum value for DESYNC_FACTOR, which is a random value
        that ensures that clients don't synchronize with each
        other and generate new addresses at exactly the same time.
        value is in seconds.

        Default: 600

regen_min_advance - INTEGER
        How far in advance (in seconds), at minimum, to create a new temporary
        address before the current one is deprecated. This value is added to
        the amount of time that may be required for duplicate address detection
        to determine when to create a new address. Linux permits setting this
        value to less than the default of 2 seconds, but a value less than 2
        does not conform to RFC 8981.

        Default: 2

regen_max_retry - INTEGER
        Number of attempts before give up attempting to generate
        valid temporary addresses.

        Default: 5

max_addresses - INTEGER
        Maximum number of autoconfigured addresses per interface.  Setting
        to zero disables the limitation.  It is not recommended to set this
        value too large (or to zero) because it would be an easy way to
        crash the kernel by allowing too many addresses to be created.

        Default: 16

disable_ipv6 - BOOLEAN
        Disable IPv6 operation.  If accept_dad is set to 2, this value
        will be dynamically set to TRUE if DAD fails for the link-local
        address.

        Default: FALSE (enable IPv6 operation)

        When this value is changed from 1 to 0 (IPv6 is being enabled),
        it will dynamically create a link-local address on the given
        interface and start Duplicate Address Detection, if necessary.

        When this value is changed from 0 to 1 (IPv6 is being disabled),
        it will dynamically delete all addresses and routes on the given
        interface. From now on it will not possible to add addresses/routes
        to the selected interface.

accept_dad - INTEGER
        Whether to accept DAD (Duplicate Address Detection).

         == ==============================================================
          0  Disable DAD
          1  Enable DAD (default)
          2  Enable DAD, and disable IPv6 operation if MAC-based duplicate
             link-local address has been found.
         == ==============================================================

        DAD operation and mode on a given interface will be selected according
        to the maximum value of conf/{all,interface}/accept_dad.

IPv6 Neighbor Discovery, DAD와 address 생성

2991-3193

`force_tllao`는 unicast Neighbor Solicitation 응답에도 target link-layer address option을 넣어 cache 삭제와 응답 사이 race를 줄입니다. `ndisc_notify`는 device up/MAC 변경 때 unsolicited NA를 보내고 `ndisc_tclass`는 ND message의 Traffic Class를 정합니다. `ndisc_evict_nocarrier`는 carrier 상실 때 ND cache를 비우며 wireless roaming에서는 0이 유용할 수 있습니다.

`mldv1_unsolicited_report_interval`과 `mldv2_unsolicited_report_interval` 기본은 10초와 1초입니다. `force_mld_version`은 0=자동, 1=MLDv1, 2=MLDv2입니다. `suppress_frag_ndisc` 기본 1은 RFC 6980에 따라 fragmented ND packet을 폐기합니다. `optimistic_dad`는 RFC 4429를 켜고 `use_optimistic`은 source 선택에서 optimistic address를 deprecated로 낮추지 않되 preferred보다는 뒤에 둡니다.

`stable_secret`은 RFC 7217 stable privacy address 생성 secret입니다. namespace 기본 secret은 `conf/default`, interface별 값은 이를 override하고 `conf/all` 쓰기는 거부됩니다. 설치 때 생성해 안정적으로 유지하는 것이 권장됩니다. `addr_gen_mode`는 0=EUI-64, 1=link-local 생성 안 함, 2=설정 secret의 stable privacy, 3=secret이 없으면 random secret의 stable privacy입니다.

`drop_unicast_in_l2_multicast`는 L2 multicast/broadcast에 담긴 unicast IPv6를 버리고 `drop_unsolicited_na`는 unsolicited NA를 버립니다. `accept_untracked_na`는 0=거부, 1=router에서 target LLA가 있는 untracked NA를 STALE entry로 추가, 2=source가 local subnet일 때만 추가합니다. drop 설정이 더 우선합니다. host의 `ndisc_notify`와 조합하면 off-link 통신의 첫 return packet을 위한 NS 지연을 줄일 수 있습니다. `enhanced_dad`는 RFC 7527 nonce를 넣어 자체 NS loopback을 중복으로 오인하지 않게 하며 기본 1입니다.

force_tllao - BOOLEAN
        Enable sending the target link-layer address option even when
        responding to a unicast neighbor solicitation.

        Default: FALSE

        Quoting from RFC 2461, section 4.4, Target link-layer address:

        "The option MUST be included for multicast solicitations in order to
        avoid infinite Neighbor Solicitation "recursion" when the peer node
        does not have a cache entry to return a Neighbor Advertisements
        message.  When responding to unicast solicitations, the option can be
        omitted since the sender of the solicitation has the correct link-
        layer address; otherwise it would not have be able to send the unicast
        solicitation in the first place. However, including the link-layer
        address in this case adds little overhead and eliminates a potential
        race condition where the sender deletes the cached link-layer address
        prior to receiving a response to a previous solicitation."

ndisc_notify - BOOLEAN
        Define mode for notification of address and device changes.

        Possible values:

        - 0 (disabled) - do nothing
        - 1 (enabled)  - Generate unsolicited neighbour advertisements when device is brought
          up or hardware address changes.

        Default: 0 (disabled)

ndisc_tclass - INTEGER
        The IPv6 Traffic Class to use by default when sending IPv6 Neighbor
        Discovery (Router Solicitation, Router Advertisement, Neighbor
        Solicitation, Neighbor Advertisement, Redirect) messages.
        These 8 bits can be interpreted as 6 high order bits holding the DSCP
        value and 2 low order bits representing ECN (which you probably want
        to leave cleared).

        * 0 - (default)

ndisc_evict_nocarrier - BOOLEAN
        Clears the neighbor discovery table on NOCARRIER events. This option is
        important for wireless devices where the neighbor discovery cache should
        not be cleared when roaming between access points on the same network.
        In most cases this should remain as the default (1).

        Possible values:

        - 0 (disabled) - Do not clear neighbor discovery cache on NOCARRIER events.
        - 1 (enabled)  - Clear neighbor discover cache on NOCARRIER events.

        Default: 1 (enabled)

mldv1_unsolicited_report_interval - INTEGER
        The interval in milliseconds in which the next unsolicited
        MLDv1 report retransmit will take place.

        Default: 10000 (10 seconds)

mldv2_unsolicited_report_interval - INTEGER
        The interval in milliseconds in which the next unsolicited
        MLDv2 report retransmit will take place.

        Default: 1000 (1 second)

force_mld_version - INTEGER
        * 0 - (default) No enforcement of a MLD version, MLDv1 fallback allowed
        * 1 - Enforce to use MLD version 1
        * 2 - Enforce to use MLD version 2

suppress_frag_ndisc - INTEGER
        Control RFC 6980 (Security Implications of IPv6 Fragmentation
        with IPv6 Neighbor Discovery) behavior:

        * 1 - (default) discard fragmented neighbor discovery packets
        * 0 - allow fragmented neighbor discovery packets

optimistic_dad - BOOLEAN
        Whether to perform Optimistic Duplicate Address Detection (RFC 4429).

        Optimistic Duplicate Address Detection for the interface will be enabled
        if at least one of conf/{all,interface}/optimistic_dad is set to 1,
        it will be disabled otherwise.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)


use_optimistic - BOOLEAN
        If enabled, do not classify optimistic addresses as deprecated during
        source address selection.  Preferred addresses will still be chosen
        before optimistic addresses, subject to other ranking in the source
        address selection algorithm.

        This will be enabled if at least one of
        conf/{all,interface}/use_optimistic is set to 1, disabled otherwise.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

stable_secret - IPv6 address
        This IPv6 address will be used as a secret to generate IPv6
        addresses for link-local addresses and autoconfigured
        ones. All addresses generated after setting this secret will
        be stable privacy ones by default. This can be changed via the
        addrgenmode ip-link. conf/default/stable_secret is used as the
        secret for the namespace, the interface specific ones can
        overwrite that. Writes to conf/all/stable_secret are refused.

        It is recommended to generate this secret during installation
        of a system and keep it stable after that.

        By default the stable secret is unset.

addr_gen_mode - INTEGER
        Defines how link-local and autoconf addresses are generated.

        =  =================================================================
        0  generate address based on EUI64 (default)
        1  do no generate a link-local address, use EUI64 for addresses
           generated from autoconf
        2  generate stable privacy addresses, using the secret from
           stable_secret (RFC7217)
        3  generate stable privacy addresses, using a random secret if unset
        =  =================================================================

drop_unicast_in_l2_multicast - BOOLEAN
        Drop any unicast IPv6 packets that are received in link-layer
        multicast (or broadcast) frames.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

drop_unsolicited_na - BOOLEAN
        Drop all unsolicited neighbor advertisements, for example if there's
        a known good NA proxy on the network and such frames need not be used
        (or in the case of 802.11, must not be used to prevent attacks.)

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled).

accept_untracked_na - INTEGER
        Define behavior for accepting neighbor advertisements from devices that
        are absent in the neighbor cache:

        - 0 - (default) Do not accept unsolicited and untracked neighbor
          advertisements.

        - 1 - Add a new neighbor cache entry in STALE state for routers on
          receiving a neighbor advertisement (either solicited or unsolicited)
          with target link-layer address option specified if no neighbor entry
          is already present for the advertised IPv6 address. Without this knob,
          NAs received for untracked addresses (absent in neighbor cache) are
          silently ignored.

          This is as per router-side behavior documented in RFC9131.

          This has lower precedence than drop_unsolicited_na.

          This will optimize the return path for the initial off-link
          communication that is initiated by a directly connected host, by
          ensuring that the first-hop router which turns on this setting doesn't
          have to buffer the initial return packets to do neighbor-solicitation.
          The prerequisite is that the host is configured to send unsolicited
          neighbor advertisements on interface bringup. This setting should be
          used in conjunction with the ndisc_notify setting on the host to
          satisfy this prerequisite.

        - 2 - Extend option (1) to add a new neighbor cache entry only if the
          source IP address is in the same subnet as an address configured on
          the interface that received the neighbor advertisement.

enhanced_dad - BOOLEAN
        Include a nonce option in the IPv6 neighbor solicitation messages used for
        duplicate address detection per RFC7527. A received DAD NS will only signal
        a duplicate address if the nonce is different. This avoids any false
        detection of duplicates due to loopback of the NS messages that we send.
        The nonce option will be sent on an interface unless both of
        conf/{all,interface}/enhanced_dad are set to FALSE.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

ICMPv6 제한과 echo 정책

3194-3277

ICMPv6 `ratelimit`은 peer별 응답 사이의 최소 ms이고 0은 제한을 끕니다. `ratemask`는 `0-127,129` 같은 range 목록으로 제한할 ICMPv6 type을 정하며 쓰면 기존 목록을 교체합니다. 기본 `0-1,3-127`은 Packet Too Big을 제외한 ICMPv6 error를 제한합니다.

`echo_ignore_all`, `echo_ignore_multicast`, `echo_ignore_anycast`는 각각 모든, multicast 대상, anycast 대상 echo request를 무시합니다. `error_anycast_as_unicast`는 anycast 대상 요청에서 발생한 ICMP error에도 unicast처럼 응답합니다. 모두 기본 0입니다. `xfrm6_gc_thresh`는 4.14부터 폐기되었습니다. 마지막에는 IPv6 문서 갱신자 credit이 있습니다.

``icmp/*``:
===========

ratelimit - INTEGER
        Limit the maximal rates for sending ICMPv6 messages to a particular
        peer.

        0 to disable any limiting,
        otherwise the space between responses in milliseconds.

        Default: 100

ratemask - list of comma separated ranges
        For ICMPv6 message types matching the ranges in the ratemask, limit
        the sending of the message according to ratelimit parameter.

        The format used for both input and output is a comma separated
        list of ranges (e.g. "0-127,129" for ICMPv6 message type 0 to 127 and
        129). Writing to the file will clear all previous ranges of ICMPv6
        message types and update the current list with the input.

        Refer to: https://www.iana.org/assignments/icmpv6-parameters/icmpv6-parameters.xhtml
        for numerical values of ICMPv6 message types, e.g. echo request is 128
        and echo reply is 129.

        Default: 0-1,3-127 (rate limit ICMPv6 errors except Packet Too Big)

echo_ignore_all - BOOLEAN
        If enabled, then the kernel will ignore all ICMP ECHO
        requests sent to it over the IPv6 protocol.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

echo_ignore_multicast - BOOLEAN
        If enabled, then the kernel will ignore all ICMP ECHO
        requests sent to it over the IPv6 protocol via multicast.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

echo_ignore_anycast - BOOLEAN
        If enabled, then the kernel will ignore all ICMP ECHO
        requests sent to it over the IPv6 protocol destined to anycast address.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

error_anycast_as_unicast - BOOLEAN
        If enabled, then the kernel will respond with ICMP Errors
        resulting from requests sent to it over the IPv6 protocol destined
        to anycast address essentially treating anycast as unicast.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 0 (disabled)

xfrm6_gc_thresh - INTEGER
        (Obsolete since linux-4.14)
        The threshold at which we will start garbage collecting for IPv6
        destination cache entries.  At twice this value the system will
        refuse new allocations.


IPv6 Update by:
Pekka Savola <[email protected]>
YOSHIFUJI Hideaki / USAGI Project <[email protected]>

Bridge netfilter 변수

3278-3337

`bridge-nf-call-arptables`, `bridge-nf-call-iptables`, `bridge-nf-call-ip6tables`는 bridged ARP, IPv4, IPv6 traffic을 각 netfilter chain에 통과시키며 기본 1입니다. `bridge-nf-filter-vlan-tagged`와 `bridge-nf-filter-pppoe-tagged`는 VLAN/PPPoE tagged bridged traffic에도 filtering을 적용하며 기본 0입니다.

`bridge-nf-pass-vlan-input-dev`를 켜면 VLAN-tagged filtering 중 bridge 위의 matching VLAN interface를 찾아 netfilter input device로 사용합니다. 그러면 `iptables -i br0.1`과 VLAN-on-bridge의 REDIRECT가 동작합니다. match가 없거나 설정이 0이면 bridge interface를 씁니다.

/proc/sys/net/bridge/* Variables:
=================================

bridge-nf-call-arptables - BOOLEAN

        Possible values:

        - 0 (disabled) - disable this.
        - 1 (enabled)  - pass bridged ARP traffic to arptables' FORWARD chain.

        Default: 1 (enabled)

bridge-nf-call-iptables - BOOLEAN

        Possible values:

        - 0 (disabled) - disable this.
        - 1 (enabled)  - pass bridged IPv4 traffic to iptables' chains.

        Default: 1 (enabled)

bridge-nf-call-ip6tables - BOOLEAN

        Possible values:

        - 0 (disabled) - disable this.
        - 1 (enabled)  - pass bridged IPv6 traffic to ip6tables' chains.

        Default: 1 (enabled)

bridge-nf-filter-vlan-tagged - BOOLEAN

        Possible values:

        - 0 (disabled) - disable this.
        - 1 (enabled)  - pass bridged vlan-tagged ARP/IP/IPv6 traffic to {arp,ip,ip6}tables

        Default: 0 (disabled)

bridge-nf-filter-pppoe-tagged - BOOLEAN

        Possible values:

        - 0 (disabled) - disable this.
        - 1 (enabled)  - pass bridged pppoe-tagged IP/IPv6 traffic to {ip,ip6}tables.

        Default: 0 (disabled)

bridge-nf-pass-vlan-input-dev - BOOLEAN
        - 1: if bridge-nf-filter-vlan-tagged is enabled, try to find a vlan
          interface on the bridge and set the netfilter input device to the
          vlan. This allows use of e.g. "iptables -i br0.1" and makes the
          REDIRECT target work with vlan-on-top-of-bridge interfaces.  When no
          matching vlan interface is found, or this switch is off, the input
          device is set to the bridge interface.

        - 0: disable bridge netfilter vlan interface lookup.

        Default: 0

SCTP extension과 failover

3338-3509

`addip_enable`은 RFC 5061 Dynamic Address Reconfiguration을, `auth_enable`은 ADD-IP 보호에 필요한 Authenticated Chunks를 켭니다. `addip_noauth_enable`은 구형 구현과의 폐쇄망 호환을 위해 인증 없는 ADD-IP를 허용하므로 일반 환경에서는 0을 유지해야 합니다. `prsctp_enable`은 RFC 3758 Partial Reliability를 켜고 기본 1입니다.

`pf_enable`은 Potentially Failed path 상태를 켭니다. `pf_retrans`가 `path_max_retrans`보다 크면 역시 PF를 끕니다. `pf_expose`는 application에 PF transition event와 `SCTP_GET_PEER_ADDR_INFO`를 0=event 없이 조회 허용, 1=둘 다 차단, 2=event와 조회 허용으로 노출합니다. `max_burst` 기본 4는 처음 보낼 새 packet burst를 제한합니다.

`association_max_retrans`는 association 전체 재전송 상한 10, `max_init_retransmits`는 INIT/COOKIE-ECHO 상한 8, `path_max_retrans`는 path unreachable 판정 상한 5입니다. `pf_retrans`는 alternate transport로 빨리 돌릴 threshold이고 path 자체는 계속 쓸 수 있습니다. `ps_retrans`는 RFC 7829 Primary Switchover threshold이며 기본 0xffff로 기능을 끄고 `pf_retrans`보다 작게 설정할 수 없습니다.

SCTP RTO의 `rto_initial`, `rto_max`, `rto_min` 기본값은 각각 3000, 60000, 1000 ms입니다.

``proc/sys/net/sctp/*`` Variables:
==================================

addip_enable - BOOLEAN
        Enable or disable extension of  Dynamic Address Reconfiguration
        (ADD-IP) functionality specified in RFC5061.  This extension provides
        the ability to dynamically add and remove new addresses for the SCTP
        associations.

        Possible values:

        - 0 (disabled) - disable extension.
        - 1 (enabled)  - enable extension

        Default: 0 (disabled)

pf_enable - INTEGER
        Enable or disable pf (pf is short for potentially failed) state. A value
        of pf_retrans > path_max_retrans also disables pf state. That is, one of
        both pf_enable and pf_retrans > path_max_retrans can disable pf state.
        Since pf_retrans and path_max_retrans can be changed by userspace
        application, sometimes user expects to disable pf state by the value of
        pf_retrans > path_max_retrans, but occasionally the value of pf_retrans
        or path_max_retrans is changed by the user application, this pf state is
        enabled. As such, it is necessary to add this to dynamically enable
        and disable pf state. See:
        https://datatracker.ietf.org/doc/draft-ietf-tsvwg-sctp-failover for
        details.

        Possible values:

        - 1: Enable pf.
        - 0: Disable pf.

        Default: 1

pf_expose - INTEGER
        Unset or enable/disable pf (pf is short for potentially failed) state
        exposure.  Applications can control the exposure of the PF path state
        in the SCTP_PEER_ADDR_CHANGE event and access of SCTP_PF-state
        transport info via SCTP_GET_PEER_ADDR_INFO sockopt.

        Possible values:

        - 0: Unset pf state exposure (compatible with old applications). No
          event will be sent but the transport info can be queried.
        - 1: Disable pf state exposure. No event will be sent and trying to
          obtain transport info will return -EACCESS.
        - 2: Enable pf state exposure. The event will be sent for a transport
          becoming SCTP_PF state and transport info can be obtained.

        Default: 0

addip_noauth_enable - BOOLEAN
        Dynamic Address Reconfiguration (ADD-IP) requires the use of
        authentication to protect the operations of adding or removing new
        addresses.  This requirement is mandated so that unauthorized hosts
        would not be able to hijack associations.  However, older
        implementations may not have implemented this requirement while
        allowing the ADD-IP extension.  For reasons of interoperability,
        we provide this variable to control the enforcement of the
        authentication requirement.

        == ===============================================================
        1  Allow ADD-IP extension to be used without authentication.  This
           should only be set in a closed environment for interoperability
           with older implementations.

        0  Enforce the authentication requirement
        == ===============================================================

        Default: 0

auth_enable - BOOLEAN
        Enable or disable Authenticated Chunks extension.  This extension
        provides the ability to send and receive authenticated chunks and is
        required for secure operation of Dynamic Address Reconfiguration
        (ADD-IP) extension.

        Possible values:

        - 0 (disabled) - disable extension.
        - 1 (enabled)  - enable extension

        Default: 0 (disabled)

prsctp_enable - BOOLEAN
        Enable or disable the Partial Reliability extension (RFC3758) which
        is used to notify peers that a given DATA should no longer be expected.

        Possible values:

        - 0 (disabled) - disable extension.
        - 1 (enabled)  - enable extension

        Default: 1 (enabled)

max_burst - INTEGER
        The limit of the number of new packets that can be initially sent.  It
        controls how bursty the generated traffic can be.

        Default: 4

association_max_retrans - INTEGER
        Set the maximum number for retransmissions that an association can
        attempt deciding that the remote end is unreachable.  If this value
        is exceeded, the association is terminated.

        Default: 10

max_init_retransmits - INTEGER
        The maximum number of retransmissions of INIT and COOKIE-ECHO chunks
        that an association will attempt before declaring the destination
        unreachable and terminating.

        Default: 8

path_max_retrans - INTEGER
        The maximum number of retransmissions that will be attempted on a given
        path.  Once this threshold is exceeded, the path is considered
        unreachable, and new traffic will use a different path when the
        association is multihomed.

        Default: 5

pf_retrans - INTEGER
        The number of retransmissions that will be attempted on a given path
        before traffic is redirected to an alternate transport (should one
        exist).  Note this is distinct from path_max_retrans, as a path that
        passes the pf_retrans threshold can still be used.  Its only
        deprioritized when a transmission path is selected by the stack.  This
        setting is primarily used to enable fast failover mechanisms without
        having to reduce path_max_retrans to a very low value.  See:
        http://www.ietf.org/id/draft-nishida-tsvwg-sctp-failover-05.txt
        for details.  Note also that a value of pf_retrans > path_max_retrans
        disables this feature. Since both pf_retrans and path_max_retrans can
        be changed by userspace application, a variable pf_enable is used to
        disable pf state.

        Default: 0

ps_retrans - INTEGER
        Primary.Switchover.Max.Retrans (PSMR), it's a tunable parameter coming
        from section-5 "Primary Path Switchover" in rfc7829.  The primary path
        will be changed to another active path when the path error counter on
        the old primary path exceeds PSMR, so that "the SCTP sender is allowed
        to continue data transmission on a new working path even when the old
        primary destination address becomes active again".   Note this feature
        is disabled by initializing 'ps_retrans' per netns as 0xffff by default,
        and its value can't be less than 'pf_retrans' when changing by sysctl.

        Default: 0xffff

rto_initial - INTEGER
        The initial round trip timeout value in milliseconds that will be used
        in calculating round trip times.  This is the initial time interval
        for retransmissions.

        Default: 3000

rto_max - INTEGER
        The maximum value (in milliseconds) of the round trip timeout.  This
        is the largest time interval that can elapse between retransmissions.

        Default: 60000

rto_min - INTEGER
        The minimum value (in milliseconds) of the round trip timeout.  This
        is the smallest time interval the can elapse between retransmissions.

        Default: 1000

SCTP timer, memory와 encapsulation

3510-3717

`hb_interval`은 idle path HEARTBEAT 간격 30000 ms, `sack_timeout`은 SACK 지연 200 ms, `valid_cookie_life`는 association cookie 수명 60000 ms입니다. `cookie_preserve_enable`은 cookie lifetime extension을 허용하고 `cookie_hmac_alg`는 기본 `sha256` 또는 `none`을 고릅니다.

`rcvbuf_policy`와 `sndbuf_policy`는 buffer accounting을 0=socket별, 1=association별로 정합니다. 한 socket의 stalled association이 다른 association을 막는 문제에는 association별 receive 정책이 유용합니다. `sctp_mem`은 전체 queue page의 min/pressure/max, `sctp_rmem`과 `sctp_wmem`은 첫 min 값만 사용해 memory pressure에서도 socket별 4 KiB를 보장하고 default/max 값은 무시합니다.

`addr_scope_policy`는 IPv4 scope를 0=끔, 1=draft 준수, 2=private address 허용, 3=link-local 허용으로 정합니다. `udp_port`는 RFC 6951 incoming SCTP-over-UDP를 받는 local shared tunnel port이자 outgoing source port이고 0이면 socket을 닫습니다. `encap_port`는 outgoing destination port의 기본값이며 socket/association/transport별 option으로 override할 수 있습니다.

`plpmtud_probe_interval`은 PLPMTUD probe ACK timer와 완료 후 current PMTU 재probe 간격입니다. 0은 기능을 끄고 다른 값은 최소 5000 ms입니다. `reconf_enable`은 RFC 6525 Stream Reconfiguration, `intl_enable`은 RFC 8260 User Message Interleaving을 켭니다. interleaving은 sysctl 외에 `SCTP_FRAGMENT_INTERLEAVE=2`와 `SCTP_INTERLEAVING_SUPPORTED=1`도 필요합니다. `ecn_enable`은 peer 양쪽이 지원할 때 SCTP ECN을 사용하고 기본 1입니다. `l3mdev_accept`는 global bound SCTP socket이 VRF를 가로질러 받게 하며 기본 1입니다.

hb_interval - INTEGER
        The interval (in milliseconds) between HEARTBEAT chunks.  These chunks
        are sent at the specified interval on idle paths to probe the state of
        a given path between 2 associations.

        Default: 30000

sack_timeout - INTEGER
        The amount of time (in milliseconds) that the implementation will wait
        to send a SACK.

        Default: 200

valid_cookie_life - INTEGER
        The default lifetime of the SCTP cookie (in milliseconds).  The cookie
        is used during association establishment.

        Default: 60000

cookie_preserve_enable - BOOLEAN
        Enable or disable the ability to extend the lifetime of the SCTP cookie
        that is used during the establishment phase of SCTP association

        Possible values:

        - 0 (disabled) - disable.
        - 1 (enabled)  - enable cookie lifetime extension.

        Default: 1 (enabled)

cookie_hmac_alg - STRING
        Select the hmac algorithm used when generating the cookie value sent by
        a listening sctp socket to a connecting client in the INIT-ACK chunk.
        Valid values are:

        * sha256
        * none

        Default: sha256

rcvbuf_policy - INTEGER
        Determines if the receive buffer is attributed to the socket or to
        association.   SCTP supports the capability to create multiple
        associations on a single socket.  When using this capability, it is
        possible that a single stalled association that's buffering a lot
        of data may block other associations from delivering their data by
        consuming all of the receive buffer space.  To work around this,
        the rcvbuf_policy could be set to attribute the receiver buffer space
        to each association instead of the socket.  This prevents the described
        blocking.

        - 1: rcvbuf space is per association
        - 0: rcvbuf space is per socket

        Default: 0

sndbuf_policy - INTEGER
        Similar to rcvbuf_policy above, this applies to send buffer space.

        - 1: Send buffer is tracked per association
        - 0: Send buffer is tracked per socket.

        Default: 0

sctp_mem - vector of 3 INTEGERs: min, pressure, max
        Number of pages allowed for queueing by all SCTP sockets.

        * min: Below this number of pages SCTP is not bothered about its
          memory usage. When amount of memory allocated by SCTP exceeds
          this number, SCTP starts to moderate memory usage.
        * pressure: This value was introduced to follow format of tcp_mem.
        * max: Maximum number of allowed pages.

        Default is calculated at boot time from amount of available memory.

sctp_rmem - vector of 3 INTEGERs: min, default, max
        Only the first value ("min") is used, "default" and "max" are
        ignored.

        * min: Minimal size of receive buffer used by SCTP socket.
          It is guaranteed to each SCTP socket (but not association) even
          under moderate memory pressure.

        Default: 4K

sctp_wmem  - vector of 3 INTEGERs: min, default, max
        Only the first value ("min") is used, "default" and "max" are
        ignored.

        * min: Minimum size of send buffer that can be used by SCTP sockets.
          It is guaranteed to each SCTP socket (but not association) even
          under moderate memory pressure.

        Default: 4K

addr_scope_policy - INTEGER
        Control IPv4 address scoping (see
        https://datatracker.ietf.org/doc/draft-stewart-tsvwg-sctp-ipv4/00/
        for details).

        - 0   - Disable IPv4 address scoping
        - 1   - Enable IPv4 address scoping
        - 2   - Follow draft but allow IPv4 private addresses
        - 3   - Follow draft but allow IPv4 link local addresses

        Default: 1

udp_port - INTEGER
        The listening port for the local UDP tunneling sock. Normally it's
        using the IANA-assigned UDP port number 9899 (sctp-tunneling).

        This UDP sock is used for processing the incoming UDP-encapsulated
        SCTP packets (from RFC6951), and shared by all applications in the
        same net namespace. This UDP sock will be closed when the value is
        set to 0.

        The value will also be used to set the src port of the UDP header
        for the outgoing UDP-encapsulated SCTP packets. For the dest port,
        please refer to 'encap_port' below.

        Default: 0

encap_port - INTEGER
        The default remote UDP encapsulation port.

        This value is used to set the dest port of the UDP header for the
        outgoing UDP-encapsulated SCTP packets by default. Users can also
        change the value for each sock/asoc/transport by using setsockopt.
        For further information, please refer to RFC6951.

        Note that when connecting to a remote server, the client should set
        this to the port that the UDP tunneling sock on the peer server is
        listening to and the local UDP tunneling sock on the client also
        must be started. On the server, it would get the encap_port from
        the incoming packet's source port.

        Default: 0

plpmtud_probe_interval - INTEGER
        The time interval (in milliseconds) for the PLPMTUD probe timer,
        which is configured to expire after this period to receive an
        acknowledgment to a probe packet. This is also the time interval
        between the probes for the current pmtu when the probe search
        is done.

        PLPMTUD will be disabled when 0 is set, and other values for it
        must be >= 5000.

        Default: 0

reconf_enable - BOOLEAN
        Enable or disable extension of Stream Reconfiguration functionality
        specified in RFC6525. This extension provides the ability to "reset"
        a stream, and it includes the Parameters of "Outgoing/Incoming SSN
        Reset", "SSN/TSN Reset" and "Add Outgoing/Incoming Streams".

        Possible values:

        - 0 (disabled) - Disable extension.
        - 1 (enabled) - Enable extension.

        Default: 0 (disabled)

intl_enable - BOOLEAN
        Enable or disable extension of User Message Interleaving functionality
        specified in RFC8260. This extension allows the interleaving of user
        messages sent on different streams. With this feature enabled, I-DATA
        chunk will replace DATA chunk to carry user messages if also supported
        by the peer. Note that to use this feature, one needs to set this option
        to 1 and also needs to set socket options SCTP_FRAGMENT_INTERLEAVE to 2
        and SCTP_INTERLEAVING_SUPPORTED to 1.

        Possible values:

        - 0 (disabled) - Disable extension.
        - 1 (enabled) - Enable extension.

        Default: 0 (disabled)

ecn_enable - BOOLEAN
        Control use of Explicit Congestion Notification (ECN) by SCTP.
        Like in TCP, ECN is used only when both ends of the SCTP connection
        indicate support for it. This feature is useful in avoiding losses
        due to congestion by allowing supporting routers to signal congestion
        before having to drop packets.

        Possible values:

        - 0 (disabled) - Disable ecn.
        - 1 (enabled) - Enable ecn.

        Default: 1 (enabled)

l3mdev_accept - BOOLEAN
        Enabling this option allows a "global" bound socket to work
        across L3 master domains (e.g., VRFs) with packets capable of
        being received regardless of the L3 domain in which they
        originated. Only valid when the kernel was compiled with
        CONFIG_NET_L3_MASTER_DEV.

        Possible values:

        - 0 (disabled)
        - 1 (enabled)

        Default: 1 (enabled)

Core 참조와 Unix datagram queue

3718-3731

`/proc/sys/net/core/*` 항목은 `Documentation/admin-guide/sysctl/net.rst`에서 설명합니다. `/proc/sys/net/unix/max_dgram_qlen`은 Unix datagram socket receive queue의 최대 길이이며 기본 10입니다.

``/proc/sys/net/core/*``
========================

        Please see: Documentation/admin-guide/sysctl/net.rst for descriptions of these entries.


``/proc/sys/net/unix/*``
========================

max_dgram_qlen - INTEGER
        The maximum length of dgram socket receive queue

        Default: 10