← Documents Documentation/bpf/prog_flow_dissector.rst GitHub 원문 ↗

Linux 6.18.37 · BPF

BPF_PROG_TYPE_FLOW_DISSECTOR

BPF flow dissector의 context, VLAN 전후 packet 시작점, optional flags, reference 구현과 현재 제한을 설명합니다.

Source pathDocumentation/bpf/prog_flow_dissector.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

prog_flow_dissector.rst:1-147

Flow dissector BPF program은 제한된 `__sk_buff` field와 `struct bpf_flow_keys`를 사용해 packet metadata를 parse합니다. 성공 시 `BPF_OK`, parsing error 시 `BPF_DROP`을 반환합니다.

Program은 VLAN header가 아직 남아 TCI에서 시작하는 경우와 VLAN 처리가 끝나 L3 header에서 시작하는 경우를 모두 처리해야 합니다. Double VLAN의 802.1AD도 고려해야 합니다.

Reference 구현은 protocol별 sub-program을 `jmp_table`에 두고 `bpf_tail_call`로 dispatch합니다. Root network namespace에 attach한 program은 child namespace가 덮어쓸 수 없습니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 ============================
4 BPF_PROG_TYPE_FLOW_DISSECTOR
5 ============================
6
7 Overview
8 ========
9
10 Flow dissector is a routine that parses metadata out of the packets. It's
11 used in the various places in the networking subsystem (RFS, flow hash, etc).
12
13 BPF flow dissector is an attempt to reimplement C-based flow dissector logic
14 in BPF to gain all the benefits of BPF verifier (namely, limits on the
15 number of instructions and tail calls).
16
17 API
18 ===
19
20 BPF flow dissector programs operate on an ``__sk_buff``. However, only the
21 limited set of fields is allowed: ``data``, ``data_end`` and ``flow_keys``.
22 ``flow_keys`` is ``struct bpf_flow_keys`` and contains flow dissector input
23 and output arguments.
24
25 The inputs are:
26 * ``nhoff`` - initial offset of the networking header
27 * ``thoff`` - initial offset of the transport header, initialized to nhoff
28 * ``n_proto`` - L3 protocol type, parsed out of L2 header
29 * ``flags`` - optional flags
30
31 Flow dissector BPF program should fill out the rest of the ``struct
32 bpf_flow_keys`` fields. Input arguments ``nhoff/thoff/n_proto`` should be
33 also adjusted accordingly.
34
35 The return code of the BPF program is either BPF_OK to indicate successful
36 dissection, or BPF_DROP to indicate parsing error.
37
38 __sk_buff->data
39 ===============
40
41 In the VLAN-less case, this is what the initial state of the BPF flow
42 dissector looks like::
43
44 +------+------+------------+-----------+
45 | DMAC | SMAC | ETHER_TYPE | L3_HEADER |
46 +------+------+------------+-----------+
47 ^
48 |
49 +-- flow dissector starts here
50
51
52 .. code:: c
53
54 skb->data + flow_keys->nhoff point to the first byte of L3_HEADER
55 flow_keys->thoff = nhoff
56 flow_keys->n_proto = ETHER_TYPE
57
58 In case of VLAN, flow dissector can be called with the two different states.
59
60 Pre-VLAN parsing::
61
62 +------+------+------+-----+-----------+-----------+
63 | DMAC | SMAC | TPID | TCI |ETHER_TYPE | L3_HEADER |
64 +------+------+------+-----+-----------+-----------+
65 ^
66 |
67 +-- flow dissector starts here
68
69 .. code:: c
70
71 skb->data + flow_keys->nhoff point the to first byte of TCI
72 flow_keys->thoff = nhoff
73 flow_keys->n_proto = TPID
74
75 Please note that TPID can be 802.1AD and, hence, BPF program would
76 have to parse VLAN information twice for double tagged packets.
77
78
79 Post-VLAN parsing::
80
81 +------+------+------+-----+-----------+-----------+
82 | DMAC | SMAC | TPID | TCI |ETHER_TYPE | L3_HEADER |
83 +------+------+------+-----+-----------+-----------+
84 ^
85 |
86 +-- flow dissector starts here
87
88 .. code:: c
89
90 skb->data + flow_keys->nhoff point the to first byte of L3_HEADER
91 flow_keys->thoff = nhoff
92 flow_keys->n_proto = ETHER_TYPE
93
94 In this case VLAN information has been processed before the flow dissector
95 and BPF flow dissector is not required to handle it.
96
97
98 The takeaway here is as follows: BPF flow dissector program can be called with
99 the optional VLAN header and should gracefully handle both cases: when single
100 or double VLAN is present and when it is not present. The same program
101 can be called for both cases and would have to be written carefully to
102 handle both cases.
103
104
105 Flags
106 =====
107
108 ``flow_keys->flags`` might contain optional input flags that work as follows:
109
110 * ``BPF_FLOW_DISSECTOR_F_PARSE_1ST_FRAG`` - tells BPF flow dissector to
111 continue parsing first fragment; the default expected behavior is that
112 flow dissector returns as soon as it finds out that the packet is fragmented;
113 used by ``eth_get_headlen`` to estimate length of all headers for GRO.
114 * ``BPF_FLOW_DISSECTOR_F_STOP_AT_FLOW_LABEL`` - tells BPF flow dissector to
115 stop parsing as soon as it reaches IPv6 flow label; used by
116 ``___skb_get_hash`` to get flow hash.
117 * ``BPF_FLOW_DISSECTOR_F_STOP_AT_ENCAP`` - tells BPF flow dissector to stop
118 parsing as soon as it reaches encapsulated headers; used by routing
119 infrastructure.
120
121
122 Reference Implementation
123 ========================
124
125 See ``tools/testing/selftests/bpf/progs/bpf_flow.c`` for the reference
126 implementation and ``tools/testing/selftests/bpf/flow_dissector_load.[hc]``
127 for the loader. bpftool can be used to load BPF flow dissector program as well.
128
129 The reference implementation is organized as follows:
130 * ``jmp_table`` map that contains sub-programs for each supported L3 protocol
131 * ``_dissect`` routine - entry point; it does input ``n_proto`` parsing and
132 does ``bpf_tail_call`` to the appropriate L3 handler
133
134 Since BPF at this point doesn't support looping (or any jumping back),
135 jmp_table is used instead to handle multiple levels of encapsulation (and
136 IPv6 options).
137
138
139 Current Limitations
140 ===================
141 BPF flow dissector doesn't support exporting all the metadata that in-kernel
142 C-based implementation can export. Notable example is single VLAN (802.1Q)
143 and double VLAN (802.1AD) tags. Please refer to the ``struct bpf_flow_keys``
144 for a set of information that's currently can be exported from the BPF context.
145
146 When BPF flow dissector is attached to the root network namespace (machine-wide
147 policy), users can't override it in their child network namespaces.
148

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Flow dissector와 BPF verifier

1-16

`BPF_PROG_TYPE_FLOW_DISSECTOR` 문서는 `GPL-2.0` 라이선스를 사용합니다.

Flow dissector는 packet에서 metadata를 parse하는 routine입니다. Networking subsystem의 RFS, flow hash 등 여러 위치에서 사용됩니다.

BPF flow dissector는 C 기반 flow dissector logic을 BPF로 다시 구현해 BPF verifier의 이점, 특히 instruction 수와 tail call 수 제한을 얻으려는 방식입니다.

__sk_buff와 struct bpf_flow_keys

17-37

BPF flow dissector program은 `__sk_buff`에서 동작하지만 사용할 수 있는 field는 `data`, `data_end`, `flow_keys`로 제한됩니다. `flow_keys`는 `struct bpf_flow_keys`이며 flow dissector의 input 및 output argument를 담습니다.

Input은 다음과 같습니다.

  • `nhoff`: networking header의 initial offset
  • `thoff`: transport header의 initial offset이며 처음에는 `nhoff`로 초기화됩니다.
  • `n_proto`: L2 header에서 parse한 L3 protocol type
  • `flags`: optional flag

Flow dissector BPF program은 `struct bpf_flow_keys`의 나머지 field를 채워야 하며 input argument `nhoff`, `thoff`, `n_proto`도 그에 맞게 조정해야 합니다.

BPF program의 return code는 dissection 성공을 나타내는 `BPF_OK` 또는 parsing error를 나타내는 `BPF_DROP`입니다.

VLAN이 없는 packet layout

38-57

VLAN이 없는 경우 BPF flow dissector의 initial state는 다음과 같습니다.

VLAN 없는 Ethernet frame
DMACSMACETHER_TYPEL3_HEADER
DMACSMACETHER_TYPEL3_HEADER <- start
Flow dissector 시작점skb->data + flow_keys->nhoff가 L3_HEADER의 첫 byte를 가리킵니다.

원문의 packet ASCII를 field 순서와 flow dissector 시작점이 드러나도록 구조화했습니다.

skb->data + flow_keys->nhoff point to the first byte of L3_HEADER
flow_keys->thoff = nhoff
flow_keys->n_proto = ETHER_TYPE

`skb->data + flow_keys->nhoff`는 `L3_HEADER`의 첫 byte를 가리키고, `flow_keys->thoff = nhoff`, `flow_keys->n_proto = ETHER_TYPE`으로 시작합니다.

Pre-VLAN parsing

58-77

VLAN이 있으면 flow dissector는 서로 다른 두 상태로 호출될 수 있습니다. Pre-VLAN parsing 상태는 다음과 같습니다.

Pre-VLAN parsing frame
DMACSMACTPIDTCIETHER_TYPEL3_HEADER
DMACSMACTPIDTCI <- startETHER_TYPEL3_HEADER
Flow dissector 시작점skb->data + flow_keys->nhoff가 TCI의 첫 byte를 가리킵니다.

TCI 첫 byte에서 dissector가 시작하며 VLAN TPID가 아직 protocol input으로 남아 있는 상태입니다.

skb->data + flow_keys->nhoff point the to first byte of TCI
flow_keys->thoff = nhoff
flow_keys->n_proto = TPID

`skb->data + flow_keys->nhoff`는 `TCI`의 첫 byte를 가리키고, `flow_keys->thoff = nhoff`, `flow_keys->n_proto = TPID`입니다.

`TPID`는 802.1AD일 수 있으므로 double tagged packet에서는 BPF program이 VLAN information을 두 번 parse해야 합니다.

Post-VLAN parsing과 공통 처리

78-103

Post-VLAN parsing 상태는 다음과 같습니다.

Post-VLAN parsing frame
DMACSMACTPIDTCIETHER_TYPEL3_HEADER
DMACSMACTPIDTCIETHER_TYPEL3_HEADER <- start
Flow dissector 시작점skb->data + flow_keys->nhoff가 L3_HEADER의 첫 byte를 가리킵니다.

VLAN information이 이미 처리되어 dissector가 L3 header에서 시작하는 상태입니다.

skb->data + flow_keys->nhoff point the to first byte of L3_HEADER
flow_keys->thoff = nhoff
flow_keys->n_proto = ETHER_TYPE

`skb->data + flow_keys->nhoff`는 `L3_HEADER`의 첫 byte를 가리키고, `flow_keys->thoff = nhoff`, `flow_keys->n_proto = ETHER_TYPE`입니다.

이 경우 VLAN information은 flow dissector 호출 전에 처리되었으므로 BPF flow dissector가 직접 처리할 필요가 없습니다.

핵심은 BPF flow dissector program이 optional VLAN header가 있는 상태와 없는 상태 모두로 호출될 수 있다는 점입니다. 동일한 program이 VLAN 없음, single VLAN, double VLAN을 모두 안전하게 처리하도록 신중하게 작성해야 합니다.

Optional input flags

104-120

`flow_keys->flags`에는 다음 optional input flag가 들어갈 수 있습니다.

  • `BPF_FLOW_DISSECTOR_F_PARSE_1ST_FRAG`: 첫 fragment parsing을 계속합니다. 기본 동작은 packet이 fragmented임을 확인하면 즉시 반환하는 것입니다. `eth_get_headlen`이 GRO를 위한 전체 header 길이를 추정할 때 사용합니다.
  • `BPF_FLOW_DISSECTOR_F_STOP_AT_FLOW_LABEL`: IPv6 flow label에 도달하면 parsing을 중지합니다. `___skb_get_hash`가 flow hash를 얻을 때 사용합니다.
  • `BPF_FLOW_DISSECTOR_F_STOP_AT_ENCAP`: encapsulated header에 도달하면 parsing을 중지합니다. Routing infrastructure가 사용합니다.

Reference implementation

121-138

Reference implementation은 `tools/testing/selftests/bpf/progs/bpf_flow.c`에 있고 loader는 `tools/testing/selftests/bpf/flow_dissector_load.[hc]`에 있습니다. `bpftool`로도 BPF flow dissector program을 load할 수 있습니다.

Reference implementation의 구성은 다음과 같습니다.

  • `jmp_table`: 지원하는 각 L3 protocol의 sub-program을 담는 map
  • `_dissect` routine: entry point입니다. Input `n_proto`를 parse하고 적절한 L3 handler로 `bpf_tail_call`을 수행합니다.

이 시점의 BPF는 loop 또는 뒤로 jump하는 동작을 지원하지 않으므로 여러 encapsulation level과 IPv6 option을 처리하는 데 `jmp_table`을 사용합니다.

Current limitations와 namespace policy

139-147

BPF flow dissector는 in-kernel C 구현이 export할 수 있는 metadata 전부를 지원하지 않습니다. 대표적인 예가 single VLAN 802.1Q와 double VLAN 802.1AD tag입니다. 현재 BPF context에서 export할 수 있는 정보 집합은 `struct bpf_flow_keys`를 참고하십시오.

BPF flow dissector가 root network namespace에 attach되어 machine-wide policy가 되면 child network namespace의 user는 이를 override할 수 없습니다.