← Documents Documentation/bpf/standardization/instruction-set.rst GitHub 원문 ↗

Linux 6.18.37 · BPF / standardization

BPF Instruction Set Architecture (ISA)

BPF 명령어의 64·128비트 인코딩, 레지스터와 피연산자 규약, 산술·점프·메모리·원자 연산 및 맵 주소 immediate 형식을 정의합니다.

Source pathDocumentation/bpf/standardization/instruction-set.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

instruction-set.rst:1-790

BPF ISA는 8비트 opcode, 레지스터 번호, offset과 immediate로 된 기본 64비트 형식과 두 번째 immediate를 붙인 128비트 wide 형식을 사용합니다. opcode의 하위 비트가 instruction class를 고릅니다.

ALU·jump 명령어는 연산 code와 immediate/register source를 조합하고, load/store 명령어는 mode와 size를 조합합니다. 모든 산술 예외, 부호 확장, jump offset, atomic fetch 및 compare-exchange 의미가 명시적으로 정의됩니다.

준수 그룹은 런타임과 컴파일러가 공통 기능 수준을 합의하게 합니다. base32는 필수이며 base64, atomic, divmul, legacy packet 그룹을 선택적으로 더할 수 있습니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. contents::
2 .. sectnum::
3
4 ======================================
5 BPF Instruction Set Architecture (ISA)
6 ======================================
7
8 eBPF, also commonly
9 referred to as BPF, is a technology with origins in the Linux kernel
10 that can run untrusted programs in a privileged context such as an
11 operating system kernel. This document specifies the BPF instruction
12 set architecture (ISA).
13
14 As a historical note, BPF originally stood for Berkeley Packet Filter,
15 but now that it can do so much more than packet filtering, the acronym
16 no longer makes sense. BPF is now considered a standalone term that
17 does not stand for anything. The original BPF is sometimes referred to
18 as cBPF (classic BPF) to distinguish it from the now widely deployed
19 eBPF (extended BPF).
20
21 Documentation conventions
22 =========================
23
24 The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
25 "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
26 "OPTIONAL" in this document are to be interpreted as described in
27 BCP 14 `<https://www.rfc-editor.org/info/rfc2119>`_
28 `<https://www.rfc-editor.org/info/rfc8174>`_
29 when, and only when, they appear in all capitals, as shown here.
30
31 For brevity and consistency, this document refers to families
32 of types using a shorthand syntax and refers to several expository,
33 mnemonic functions when describing the semantics of instructions.
34 The range of valid values for those types and the semantics of those
35 functions are defined in the following subsections.
36
37 Types
38 -----
39 This document refers to integer types with the notation `SN` to specify
40 a type's signedness (`S`) and bit width (`N`), respectively.
41
42 .. table:: Meaning of signedness notation
43
44 ==== =========
45 S Meaning
46 ==== =========
47 u unsigned
48 s signed
49 ==== =========
50
51 .. table:: Meaning of bit-width notation
52
53 ===== =========
54 N Bit width
55 ===== =========
56 8 8 bits
57 16 16 bits
58 32 32 bits
59 64 64 bits
60 128 128 bits
61 ===== =========
62
63 For example, `u32` is a type whose valid values are all the 32-bit unsigned
64 numbers and `s16` is a type whose valid values are all the 16-bit signed
65 numbers.
66
67 Functions
68 ---------
69
70 The following byteswap functions are direction-agnostic. That is,
71 the same function is used for conversion in either direction discussed
72 below.
73
74 * be16: Takes an unsigned 16-bit number and converts it between
75 host byte order and big-endian
76 (`IEN137 <https://www.rfc-editor.org/ien/ien137.txt>`_) byte order.
77 * be32: Takes an unsigned 32-bit number and converts it between
78 host byte order and big-endian byte order.
79 * be64: Takes an unsigned 64-bit number and converts it between
80 host byte order and big-endian byte order.
81 * bswap16: Takes an unsigned 16-bit number in either big- or little-endian
82 format and returns the equivalent number with the same bit width but
83 opposite endianness.
84 * bswap32: Takes an unsigned 32-bit number in either big- or little-endian
85 format and returns the equivalent number with the same bit width but
86 opposite endianness.
87 * bswap64: Takes an unsigned 64-bit number in either big- or little-endian
88 format and returns the equivalent number with the same bit width but
89 opposite endianness.
90 * le16: Takes an unsigned 16-bit number and converts it between
91 host byte order and little-endian byte order.
92 * le32: Takes an unsigned 32-bit number and converts it between
93 host byte order and little-endian byte order.
94 * le64: Takes an unsigned 64-bit number and converts it between
95 host byte order and little-endian byte order.
96
97 Definitions
98 -----------
99
100 .. glossary::
101
102 Sign Extend
103 To `sign extend an` ``X`` `-bit number, A, to a` ``Y`` `-bit number, B ,` means to
104
105 #. Copy all ``X`` bits from `A` to the lower ``X`` bits of `B`.
106 #. Set the value of the remaining ``Y`` - ``X`` bits of `B` to the value of
107 the most-significant bit of `A`.
108
109 .. admonition:: Example
110
111 Sign extend an 8-bit number ``A`` to a 16-bit number ``B`` on a big-endian platform:
112 ::
113
114 A: 10000110
115 B: 11111111 10000110
116
117 Conformance groups
118 ------------------
119
120 An implementation does not need to support all instructions specified in this
121 document (e.g., deprecated instructions). Instead, a number of conformance
122 groups are specified. An implementation MUST support the base32 conformance
123 group and MAY support additional conformance groups, where supporting a
124 conformance group means it MUST support all instructions in that conformance
125 group.
126
127 The use of named conformance groups enables interoperability between a runtime
128 that executes instructions, and tools such as compilers that generate
129 instructions for the runtime. Thus, capability discovery in terms of
130 conformance groups might be done manually by users or automatically by tools.
131
132 Each conformance group has a short ASCII label (e.g., "base32") that
133 corresponds to a set of instructions that are mandatory. That is, each
134 instruction has one or more conformance groups of which it is a member.
135
136 This document defines the following conformance groups:
137
138 * base32: includes all instructions defined in this
139 specification unless otherwise noted.
140 * base64: includes base32, plus instructions explicitly noted
141 as being in the base64 conformance group.
142 * atomic32: includes 32-bit atomic operation instructions (see `Atomic operations`_).
143 * atomic64: includes atomic32, plus 64-bit atomic operation instructions.
144 * divmul32: includes 32-bit division, multiplication, and modulo instructions.
145 * divmul64: includes divmul32, plus 64-bit division, multiplication,
146 and modulo instructions.
147 * packet: deprecated packet access instructions.
148
149 Instruction encoding
150 ====================
151
152 BPF has two instruction encodings:
153
154 * the basic instruction encoding, which uses 64 bits to encode an instruction
155 * the wide instruction encoding, which appends a second 64 bits
156 after the basic instruction for a total of 128 bits.
157
158 Basic instruction encoding
159 --------------------------
160
161 A basic instruction is encoded as follows::
162
163 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
164 | opcode | regs | offset |
165 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
166 | imm |
167 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
168
169 **opcode**
170 operation to perform, encoded as follows::
171
172 +-+-+-+-+-+-+-+-+
173 |specific |class|
174 +-+-+-+-+-+-+-+-+
175
176 **specific**
177 The format of these bits varies by instruction class
178
179 **class**
180 The instruction class (see `Instruction classes`_)
181
182 **regs**
183 The source and destination register numbers, encoded as follows
184 on a little-endian host::
185
186 +-+-+-+-+-+-+-+-+
187 |src_reg|dst_reg|
188 +-+-+-+-+-+-+-+-+
189
190 and as follows on a big-endian host::
191
192 +-+-+-+-+-+-+-+-+
193 |dst_reg|src_reg|
194 +-+-+-+-+-+-+-+-+
195
196 **src_reg**
197 the source register number (0-10), except where otherwise specified
198 (`64-bit immediate instructions`_ reuse this field for other purposes)
199
200 **dst_reg**
201 destination register number (0-10), unless otherwise specified
202 (future instructions might reuse this field for other purposes)
203
204 **offset**
205 signed integer offset used with pointer arithmetic, except where
206 otherwise specified (some arithmetic instructions reuse this field
207 for other purposes)
208
209 **imm**
210 signed integer immediate value
211
212 Note that the contents of multi-byte fields ('offset' and 'imm') are
213 stored using big-endian byte ordering on big-endian hosts and
214 little-endian byte ordering on little-endian hosts.
215
216 For example::
217
218 opcode offset imm assembly
219 src_reg dst_reg
220 07 0 1 00 00 44 33 22 11 r1 += 0x11223344 // little
221 dst_reg src_reg
222 07 1 0 00 00 11 22 33 44 r1 += 0x11223344 // big
223
224 Note that most instructions do not use all of the fields.
225 Unused fields SHALL be cleared to zero.
226
227 Wide instruction encoding
228 --------------------------
229
230 Some instructions are defined to use the wide instruction encoding,
231 which uses two 32-bit immediate values. The 64 bits following
232 the basic instruction format contain a pseudo instruction
233 with 'opcode', 'dst_reg', 'src_reg', and 'offset' all set to zero.
234
235 This is depicted in the following figure::
236
237 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
238 | opcode | regs | offset |
239 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
240 | imm |
241 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
242 | reserved |
243 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
244 | next_imm |
245 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
246
247 **opcode**
248 operation to perform, encoded as explained above
249
250 **regs**
251 The source and destination register numbers (unless otherwise
252 specified), encoded as explained above
253
254 **offset**
255 signed integer offset used with pointer arithmetic, unless
256 otherwise specified
257
258 **imm**
259 signed integer immediate value
260
261 **reserved**
262 unused, set to zero
263
264 **next_imm**
265 second signed integer immediate value
266
267 Instruction classes
268 -------------------
269
270 The three least significant bits of the 'opcode' field store the instruction class:
271
272 .. table:: Instruction class
273
274 ===== ===== =============================== ===================================
275 class value description reference
276 ===== ===== =============================== ===================================
277 LD 0x0 non-standard load operations `Load and store instructions`_
278 LDX 0x1 load into register operations `Load and store instructions`_
279 ST 0x2 store from immediate operations `Load and store instructions`_
280 STX 0x3 store from register operations `Load and store instructions`_
281 ALU 0x4 32-bit arithmetic operations `Arithmetic and jump instructions`_
282 JMP 0x5 64-bit jump operations `Arithmetic and jump instructions`_
283 JMP32 0x6 32-bit jump operations `Arithmetic and jump instructions`_
284 ALU64 0x7 64-bit arithmetic operations `Arithmetic and jump instructions`_
285 ===== ===== =============================== ===================================
286
287 Arithmetic and jump instructions
288 ================================
289
290 For arithmetic and jump instructions (``ALU``, ``ALU64``, ``JMP`` and
291 ``JMP32``), the 8-bit 'opcode' field is divided into three parts::
292
293 +-+-+-+-+-+-+-+-+
294 | code |s|class|
295 +-+-+-+-+-+-+-+-+
296
297 **code**
298 the operation code, whose meaning varies by instruction class
299
300 **s (source)**
301 the source operand location, which unless otherwise specified is one of:
302
303 .. table:: Source operand location
304
305 ====== ===== ==============================================
306 source value description
307 ====== ===== ==============================================
308 K 0 use 32-bit 'imm' value as source operand
309 X 1 use 'src_reg' register value as source operand
310 ====== ===== ==============================================
311
312 **instruction class**
313 the instruction class (see `Instruction classes`_)
314
315 Arithmetic instructions
316 -----------------------
317
318 ``ALU`` uses 32-bit wide operands while ``ALU64`` uses 64-bit wide operands for
319 otherwise identical operations. ``ALU64`` instructions belong to the
320 base64 conformance group unless noted otherwise.
321 The 'code' field encodes the operation as below, where 'src' refers to the
322 the source operand and 'dst' refers to the value of the destination
323 register.
324
325 .. table:: Arithmetic instructions
326
327 ===== ===== ======= ===================================================================================
328 name code offset description
329 ===== ===== ======= ===================================================================================
330 ADD 0x0 0 dst += src
331 SUB 0x1 0 dst -= src
332 MUL 0x2 0 dst \*= src
333 DIV 0x3 0 dst = (src != 0) ? (dst / src) : 0
334 SDIV 0x3 1 dst = (src == 0) ? 0 : ((src == -1 && dst == LLONG_MIN) ? LLONG_MIN : (dst s/ src))
335 OR 0x4 0 dst \|= src
336 AND 0x5 0 dst &= src
337 LSH 0x6 0 dst <<= (src & mask)
338 RSH 0x7 0 dst >>= (src & mask)
339 NEG 0x8 0 dst = -dst
340 MOD 0x9 0 dst = (src != 0) ? (dst % src) : dst
341 SMOD 0x9 1 dst = (src == 0) ? dst : ((src == -1 && dst == LLONG_MIN) ? 0: (dst s% src))
342 XOR 0xa 0 dst ^= src
343 MOV 0xb 0 dst = src
344 MOVSX 0xb 8/16/32 dst = (s8,s16,s32)src
345 ARSH 0xc 0 :term:`sign extending<Sign Extend>` dst >>= (src & mask)
346 END 0xd 0 byte swap operations (see `Byte swap instructions`_ below)
347 ===== ===== ======= ===================================================================================
348
349 Underflow and overflow are allowed during arithmetic operations, meaning
350 the 64-bit or 32-bit value will wrap. If BPF program execution would
351 result in division by zero, the destination register is instead set to zero.
352 Otherwise, for ``ALU64``, if execution would result in ``LLONG_MIN``
353 divided by -1, the destination register is instead set to ``LLONG_MIN``. For
354 ``ALU``, if execution would result in ``INT_MIN`` divided by -1, the
355 destination register is instead set to ``INT_MIN``.
356
357 If execution would result in modulo by zero, for ``ALU64`` the value of
358 the destination register is unchanged whereas for ``ALU`` the upper
359 32 bits of the destination register are zeroed. Otherwise, for ``ALU64``,
360 if execution would resuslt in ``LLONG_MIN`` modulo -1, the destination
361 register is instead set to 0. For ``ALU``, if execution would result in
362 ``INT_MIN`` modulo -1, the destination register is instead set to 0.
363
364 ``{ADD, X, ALU}``, where 'code' = ``ADD``, 'source' = ``X``, and 'class' = ``ALU``, means::
365
366 dst = (u32) ((u32) dst + (u32) src)
367
368 where '(u32)' indicates that the upper 32 bits are zeroed.
369
370 ``{ADD, X, ALU64}`` means::
371
372 dst = dst + src
373
374 ``{XOR, K, ALU}`` means::
375
376 dst = (u32) dst ^ (u32) imm
377
378 ``{XOR, K, ALU64}`` means::
379
380 dst = dst ^ imm
381
382 Note that most arithmetic instructions have 'offset' set to 0. Only three instructions
383 (``SDIV``, ``SMOD``, ``MOVSX``) have a non-zero 'offset'.
384
385 Division, multiplication, and modulo operations for ``ALU`` are part
386 of the "divmul32" conformance group, and division, multiplication, and
387 modulo operations for ``ALU64`` are part of the "divmul64" conformance
388 group.
389 The division and modulo operations support both unsigned and signed flavors.
390
391 For unsigned operations (``DIV`` and ``MOD``), for ``ALU``,
392 'imm' is interpreted as a 32-bit unsigned value. For ``ALU64``,
393 'imm' is first :term:`sign extended<Sign Extend>` from 32 to 64 bits, and then
394 interpreted as a 64-bit unsigned value.
395
396 For signed operations (``SDIV`` and ``SMOD``), for ``ALU``,
397 'imm' is interpreted as a 32-bit signed value. For ``ALU64``, 'imm'
398 is first :term:`sign extended<Sign Extend>` from 32 to 64 bits, and then
399 interpreted as a 64-bit signed value.
400
401 Note that there are varying definitions of the signed modulo operation
402 when the dividend or divisor are negative, where implementations often
403 vary by language such that Python, Ruby, etc. differ from C, Go, Java,
404 etc. This specification requires that signed modulo MUST use truncated division
405 (where -13 % 3 == -1) as implemented in C, Go, etc.::
406
407 a % n = a - n * trunc(a / n)
408
409 The ``MOVSX`` instruction does a move operation with sign extension.
410 ``{MOVSX, X, ALU}`` :term:`sign extends<Sign Extend>` 8-bit and 16-bit operands into
411 32-bit operands, and zeroes the remaining upper 32 bits.
412 ``{MOVSX, X, ALU64}`` :term:`sign extends<Sign Extend>` 8-bit, 16-bit, and 32-bit
413 operands into 64-bit operands. Unlike other arithmetic instructions,
414 ``MOVSX`` is only defined for register source operands (``X``).
415
416 ``{MOV, K, ALU64}`` means::
417
418 dst = (s64)imm
419
420 ``{MOV, X, ALU}`` means::
421
422 dst = (u32)src
423
424 ``{MOVSX, X, ALU}`` with 'offset' 8 means::
425
426 dst = (u32)(s32)(s8)src
427
428
429 The ``NEG`` instruction is only defined when the source bit is clear
430 (``K``).
431
432 Shift operations use a mask of 0x3F (63) for 64-bit operations and 0x1F (31)
433 for 32-bit operations.
434
435 Byte swap instructions
436 ----------------------
437
438 The byte swap instructions use instruction classes of ``ALU`` and ``ALU64``
439 and a 4-bit 'code' field of ``END``.
440
441 The byte swap instructions operate on the destination register
442 only and do not use a separate source register or immediate value.
443
444 For ``ALU``, the 1-bit source operand field in the opcode is used to
445 select what byte order the operation converts from or to. For
446 ``ALU64``, the 1-bit source operand field in the opcode is reserved
447 and MUST be set to 0.
448
449 .. table:: Byte swap instructions
450
451 ===== ======== ===== =================================================
452 class source value description
453 ===== ======== ===== =================================================
454 ALU LE 0 convert between host byte order and little endian
455 ALU BE 1 convert between host byte order and big endian
456 ALU64 Reserved 0 do byte swap unconditionally
457 ===== ======== ===== =================================================
458
459 The 'imm' field encodes the width of the swap operations. The following widths
460 are supported: 16, 32 and 64. Width 64 operations belong to the base64
461 conformance group and other swap operations belong to the base32
462 conformance group.
463
464 Examples:
465
466 ``{END, LE, ALU}`` with 'imm' = 16/32/64 means::
467
468 dst = le16(dst)
469 dst = le32(dst)
470 dst = le64(dst)
471
472 ``{END, BE, ALU}`` with 'imm' = 16/32/64 means::
473
474 dst = be16(dst)
475 dst = be32(dst)
476 dst = be64(dst)
477
478 ``{END, TO, ALU64}`` with 'imm' = 16/32/64 means::
479
480 dst = bswap16(dst)
481 dst = bswap32(dst)
482 dst = bswap64(dst)
483
484 Jump instructions
485 -----------------
486
487 ``JMP32`` uses 32-bit wide operands and indicates the base32
488 conformance group, while ``JMP`` uses 64-bit wide operands for
489 otherwise identical operations, and indicates the base64 conformance
490 group unless otherwise specified.
491 The 'code' field encodes the operation as below:
492
493 .. table:: Jump instructions
494
495 ======== ===== ======= ================================= ===================================================
496 code value src_reg description notes
497 ======== ===== ======= ================================= ===================================================
498 JA 0x0 0x0 PC += offset {JA, K, JMP} only
499 JA 0x0 0x0 PC += imm {JA, K, JMP32} only
500 JEQ 0x1 any PC += offset if dst == src
501 JGT 0x2 any PC += offset if dst > src unsigned
502 JGE 0x3 any PC += offset if dst >= src unsigned
503 JSET 0x4 any PC += offset if dst & src
504 JNE 0x5 any PC += offset if dst != src
505 JSGT 0x6 any PC += offset if dst > src signed
506 JSGE 0x7 any PC += offset if dst >= src signed
507 CALL 0x8 0x0 call helper function by static ID {CALL, K, JMP} only, see `Helper functions`_
508 CALL 0x8 0x1 call PC += imm {CALL, K, JMP} only, see `Program-local functions`_
509 CALL 0x8 0x2 call helper function by BTF ID {CALL, K, JMP} only, see `Helper functions`_
510 EXIT 0x9 0x0 return {CALL, K, JMP} only
511 JLT 0xa any PC += offset if dst < src unsigned
512 JLE 0xb any PC += offset if dst <= src unsigned
513 JSLT 0xc any PC += offset if dst < src signed
514 JSLE 0xd any PC += offset if dst <= src signed
515 ======== ===== ======= ================================= ===================================================
516
517 where 'PC' denotes the program counter, and the offset to increment by
518 is in units of 64-bit instructions relative to the instruction following
519 the jump instruction. Thus 'PC += 1' skips execution of the next
520 instruction if it's a basic instruction or results in undefined behavior
521 if the next instruction is a 128-bit wide instruction.
522
523 Example:
524
525 ``{JSGE, X, JMP32}`` means::
526
527 if (s32)dst s>= (s32)src goto +offset
528
529 where 's>=' indicates a signed '>=' comparison.
530
531 ``{JLE, K, JMP}`` means::
532
533 if dst <= (u64)(s64)imm goto +offset
534
535 ``{JA, K, JMP32}`` means::
536
537 gotol +imm
538
539 where 'imm' means the branch offset comes from the 'imm' field.
540
541 Note that there are two flavors of ``JA`` instructions. The
542 ``JMP`` class permits a 16-bit jump offset specified by the 'offset'
543 field, whereas the ``JMP32`` class permits a 32-bit jump offset
544 specified by the 'imm' field. A > 16-bit conditional jump may be
545 converted to a < 16-bit conditional jump plus a 32-bit unconditional
546 jump.
547
548 All ``CALL`` and ``JA`` instructions belong to the
549 base32 conformance group.
550
551 Helper functions
552 ~~~~~~~~~~~~~~~~
553
554 Helper functions are a concept whereby BPF programs can call into a
555 set of function calls exposed by the underlying platform.
556
557 Historically, each helper function was identified by a static ID
558 encoded in the 'imm' field. Further documentation of helper functions
559 is outside the scope of this document and standardization is left for
560 future work, but use is widely deployed and more information can be
561 found in platform-specific documentation (e.g., Linux kernel documentation).
562
563 Platforms that support the BPF Type Format (BTF) support identifying
564 a helper function by a BTF ID encoded in the 'imm' field, where the BTF ID
565 identifies the helper name and type. Further documentation of BTF
566 is outside the scope of this document and standardization is left for
567 future work, but use is widely deployed and more information can be
568 found in platform-specific documentation (e.g., Linux kernel documentation).
569
570 Program-local functions
571 ~~~~~~~~~~~~~~~~~~~~~~~
572 Program-local functions are functions exposed by the same BPF program as the
573 caller, and are referenced by offset from the instruction following the call
574 instruction, similar to ``JA``. The offset is encoded in the 'imm' field of
575 the call instruction. An ``EXIT`` within the program-local function will
576 return to the caller.
577
578 Load and store instructions
579 ===========================
580
581 For load and store instructions (``LD``, ``LDX``, ``ST``, and ``STX``), the
582 8-bit 'opcode' field is divided as follows::
583
584 +-+-+-+-+-+-+-+-+
585 |mode |sz |class|
586 +-+-+-+-+-+-+-+-+
587
588 **mode**
589 The mode modifier is one of:
590
591 .. table:: Mode modifier
592
593 ============= ===== ==================================== =============
594 mode modifier value description reference
595 ============= ===== ==================================== =============
596 IMM 0 64-bit immediate instructions `64-bit immediate instructions`_
597 ABS 1 legacy BPF packet access (absolute) `Legacy BPF Packet access instructions`_
598 IND 2 legacy BPF packet access (indirect) `Legacy BPF Packet access instructions`_
599 MEM 3 regular load and store operations `Regular load and store operations`_
600 MEMSX 4 sign-extension load operations `Sign-extension load operations`_
601 ATOMIC 6 atomic operations `Atomic operations`_
602 ============= ===== ==================================== =============
603
604 **sz (size)**
605 The size modifier is one of:
606
607 .. table:: Size modifier
608
609 ==== ===== =====================
610 size value description
611 ==== ===== =====================
612 W 0 word (4 bytes)
613 H 1 half word (2 bytes)
614 B 2 byte
615 DW 3 double word (8 bytes)
616 ==== ===== =====================
617
618 Instructions using ``DW`` belong to the base64 conformance group.
619
620 **class**
621 The instruction class (see `Instruction classes`_)
622
623 Regular load and store operations
624 ---------------------------------
625
626 The ``MEM`` mode modifier is used to encode regular load and store
627 instructions that transfer data between a register and memory.
628
629 ``{MEM, <size>, STX}`` means::
630
631 *(size *) (dst + offset) = src
632
633 ``{MEM, <size>, ST}`` means::
634
635 *(size *) (dst + offset) = imm
636
637 ``{MEM, <size>, LDX}`` means::
638
639 dst = *(unsigned size *) (src + offset)
640
641 Where '<size>' is one of: ``B``, ``H``, ``W``, or ``DW``, and
642 'unsigned size' is one of: u8, u16, u32, or u64.
643
644 Sign-extension load operations
645 ------------------------------
646
647 The ``MEMSX`` mode modifier is used to encode :term:`sign-extension<Sign Extend>` load
648 instructions that transfer data between a register and memory.
649
650 ``{MEMSX, <size>, LDX}`` means::
651
652 dst = *(signed size *) (src + offset)
653
654 Where '<size>' is one of: ``B``, ``H``, or ``W``, and
655 'signed size' is one of: s8, s16, or s32.
656
657 Atomic operations
658 -----------------
659
660 Atomic operations are operations that operate on memory and can not be
661 interrupted or corrupted by other access to the same memory region
662 by other BPF programs or means outside of this specification.
663
664 All atomic operations supported by BPF are encoded as store operations
665 that use the ``ATOMIC`` mode modifier as follows:
666
667 * ``{ATOMIC, W, STX}`` for 32-bit operations, which are
668 part of the "atomic32" conformance group.
669 * ``{ATOMIC, DW, STX}`` for 64-bit operations, which are
670 part of the "atomic64" conformance group.
671 * 8-bit and 16-bit wide atomic operations are not supported.
672
673 The 'imm' field is used to encode the actual atomic operation.
674 Simple atomic operation use a subset of the values defined to encode
675 arithmetic operations in the 'imm' field to encode the atomic operation:
676
677 .. table:: Simple atomic operations
678
679 ======== ===== ===========
680 imm value description
681 ======== ===== ===========
682 ADD 0x00 atomic add
683 OR 0x40 atomic or
684 AND 0x50 atomic and
685 XOR 0xa0 atomic xor
686 ======== ===== ===========
687
688
689 ``{ATOMIC, W, STX}`` with 'imm' = ADD means::
690
691 *(u32 *)(dst + offset) += src
692
693 ``{ATOMIC, DW, STX}`` with 'imm' = ADD means::
694
695 *(u64 *)(dst + offset) += src
696
697 In addition to the simple atomic operations, there also is a modifier and
698 two complex atomic operations:
699
700 .. table:: Complex atomic operations
701
702 =========== ================ ===========================
703 imm value description
704 =========== ================ ===========================
705 FETCH 0x01 modifier: return old value
706 XCHG 0xe0 | FETCH atomic exchange
707 CMPXCHG 0xf0 | FETCH atomic compare and exchange
708 =========== ================ ===========================
709
710 The ``FETCH`` modifier is optional for simple atomic operations, and
711 always set for the complex atomic operations. If the ``FETCH`` flag
712 is set, then the operation also overwrites ``src`` with the value that
713 was in memory before it was modified.
714
715 The ``XCHG`` operation atomically exchanges ``src`` with the value
716 addressed by ``dst + offset``.
717
718 The ``CMPXCHG`` operation atomically compares the value addressed by
719 ``dst + offset`` with ``R0``. If they match, the value addressed by
720 ``dst + offset`` is replaced with ``src``. In either case, the
721 value that was at ``dst + offset`` before the operation is zero-extended
722 and loaded back to ``R0``.
723
724 64-bit immediate instructions
725 -----------------------------
726
727 Instructions with the ``IMM`` 'mode' modifier use the wide instruction
728 encoding defined in `Instruction encoding`_, and use the 'src_reg' field of the
729 basic instruction to hold an opcode subtype.
730
731 The following table defines a set of ``{IMM, DW, LD}`` instructions
732 with opcode subtypes in the 'src_reg' field, using new terms such as "map"
733 defined further below:
734
735 .. table:: 64-bit immediate instructions
736
737 ======= ========================================= =========== ==============
738 src_reg pseudocode imm type dst type
739 ======= ========================================= =========== ==============
740 0x0 dst = (next_imm << 32) | imm integer integer
741 0x1 dst = map_by_fd(imm) map fd map
742 0x2 dst = map_val(map_by_fd(imm)) + next_imm map fd data address
743 0x3 dst = var_addr(imm) variable id data address
744 0x4 dst = code_addr(imm) integer code address
745 0x5 dst = map_by_idx(imm) map index map
746 0x6 dst = map_val(map_by_idx(imm)) + next_imm map index data address
747 ======= ========================================= =========== ==============
748
749 where
750
751 * map_by_fd(imm) means to convert a 32-bit file descriptor into an address of a map (see `Maps`_)
752 * map_by_idx(imm) means to convert a 32-bit index into an address of a map
753 * map_val(map) gets the address of the first value in a given map
754 * var_addr(imm) gets the address of a platform variable (see `Platform Variables`_) with a given id
755 * code_addr(imm) gets the address of the instruction at a specified relative offset in number of (64-bit) instructions
756 * the 'imm type' can be used by disassemblers for display
757 * the 'dst type' can be used for verification and JIT compilation purposes
758
759 Maps
760 ~~~~
761
762 Maps are shared memory regions accessible by BPF programs on some platforms.
763 A map can have various semantics as defined in a separate document, and may or
764 may not have a single contiguous memory region, but the 'map_val(map)' is
765 currently only defined for maps that do have a single contiguous memory region.
766
767 Each map can have a file descriptor (fd) if supported by the platform, where
768 'map_by_fd(imm)' means to get the map with the specified file descriptor. Each
769 BPF program can also be defined to use a set of maps associated with the
770 program at load time, and 'map_by_idx(imm)' means to get the map with the given
771 index in the set associated with the BPF program containing the instruction.
772
773 Platform Variables
774 ~~~~~~~~~~~~~~~~~~
775
776 Platform variables are memory regions, identified by integer ids, exposed by
777 the runtime and accessible by BPF programs on some platforms. The
778 'var_addr(imm)' operation means to get the address of the memory region
779 identified by the given id.
780
781 Legacy BPF Packet access instructions
782 -------------------------------------
783
784 BPF previously introduced special instructions for access to packet data that were
785 carried over from classic BPF. These instructions used an instruction
786 class of ``LD``, a size modifier of ``W``, ``H``, or ``B``, and a
787 mode modifier of ``ABS`` or ``IND``. The 'dst_reg' and 'offset' fields were
788 set to zero, and 'src_reg' was set to zero for ``ABS``. However, these
789 instructions are deprecated and SHOULD no longer be used. All legacy packet
790 access instructions belong to the "packet" conformance group.
791

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

BPF ISA와 문서 규약

1-36
.. contents::
.. sectnum::

`BPF Instruction Set Architecture (ISA)` 문서는 Linux 커널에서 시작되어 운영체제 커널 같은 권한 있는 컨텍스트에서 신뢰되지 않은 프로그램을 실행할 수 있는 eBPF, 즉 BPF의 명령어 집합 아키텍처를 규정합니다.

역사적으로 BPF는 Berkeley Packet Filter를 뜻했지만 패킷 필터링보다 훨씬 많은 일을 하게 되어 현재는 어떤 말의 약어가 아닌 독립 용어로 취급됩니다. 원래 BPF는 널리 배포된 eBPF (extended BPF)와 구분하기 위해 cBPF (classic BPF)라고도 합니다.

문서에서 모두 대문자로 쓴 `MUST`, `MUST NOT`, `REQUIRED`, `SHALL`, `SHALL NOT`, `SHOULD`, `SHOULD NOT`, `RECOMMENDED`, `NOT RECOMMENDED`, `MAY`, `OPTIONAL`은 BCP 14, RFC 2119와 RFC 8174의 의미로 해석해야 합니다: https://www.rfc-editor.org/info/rfc2119, https://www.rfc-editor.org/info/rfc8174

간결성과 일관성을 위해 이 문서는 형식 계열을 축약 표기하고 명령어 의미를 설명할 때 몇 가지 설명용 mnemonic 함수를 사용합니다. 유효 값 범위와 함수 의미는 이어지는 절에서 정의합니다.

정수 형식, 바이트 교환과 부호 확장

37-116

정수 형식은 `SN` 표기를 사용하며 `S`는 signedness, `N`은 bit width를 나타냅니다.

S의미
uunsigned
ssigned
N비트 폭
88 bits
1616 bits
3232 bits
6464 bits
128128 bits

예를 들어 `u32`의 유효 값은 모든 32비트 부호 없는 수이고, `s16`의 유효 값은 모든 16비트 부호 있는 수입니다.

다음 byteswap 함수는 방향에 무관하여 호스트 순서에서 대상 순서로 갈 때와 반대로 돌아올 때 같은 함수를 사용합니다.

  • `be16`: unsigned 16-bit 값을 host byte order와 big-endian IEN137 byte order 사이에서 변환합니다.
  • `be32`: unsigned 32-bit 값을 host byte order와 big-endian 사이에서 변환합니다.
  • `be64`: unsigned 64-bit 값을 host byte order와 big-endian 사이에서 변환합니다.
  • `bswap16`: 16-bit 값의 폭은 유지하면서 big-endian과 little-endian을 서로 뒤집습니다.
  • `bswap32`: 32-bit 값의 폭은 유지하면서 endianness를 뒤집습니다.
  • `bswap64`: 64-bit 값의 폭은 유지하면서 endianness를 뒤집습니다.
  • `le16`: unsigned 16-bit 값을 host byte order와 little-endian 사이에서 변환합니다.
  • `le32`: unsigned 32-bit 값을 host byte order와 little-endian 사이에서 변환합니다.
  • `le64`: unsigned 64-bit 값을 host byte order와 little-endian 사이에서 변환합니다.

`X`비트 수 `A`를 `Y`비트 수 `B`로 Sign Extend한다는 것은 다음 절차를 뜻합니다.

  • `A`의 `X`개 비트를 모두 `B`의 하위 `X`개 비트로 복사합니다.
  • 남은 `Y - X`개 비트를 `A`의 최상위 비트 값으로 채웁니다.

big-endian 플랫폼에서 8비트 값 `A`를 16비트 값 `B`로 부호 확장하는 예입니다.

Sign Extend 비트 예
비트
A (8-bit)10000110
B (16-bit)11111111 10000110
결과하위 8비트는 그대로이며 상위 8비트가 모두 1로 채워집니다.

최상위 부호 비트 1을 새 상위 8비트 전체로 복제합니다.

준수 그룹

117-148

구현은 폐기된 명령어 등을 포함해 이 문서의 모든 명령어를 지원할 필요가 없습니다. 대신 base32 준수 그룹은 반드시 지원해야 하며 추가 그룹은 선택적으로 지원할 수 있습니다. 어떤 그룹을 지원한다면 그 그룹의 모든 명령어를 지원해야 합니다.

이름 있는 conformance groups는 명령어를 실행하는 런타임과 명령어를 생성하는 컴파일러 같은 도구 사이의 상호 운용성을 가능하게 합니다. 사용자가 수동으로 또는 도구가 자동으로 그룹 단위 기능을 탐지할 수 있습니다. 각 그룹은 `base32` 같은 짧은 ASCII label과 필수 명령어 집합을 갖습니다.

  • `base32`: 별도 표기가 없는 한 이 규격의 모든 명령어
  • `base64`: base32와 base64 그룹으로 명시된 명령어
  • `atomic32`: 32-bit atomic operation instructions
  • `atomic64`: atomic32와 64-bit atomic operation instructions
  • `divmul32`: 32-bit division, multiplication, modulo 명령어
  • `divmul64`: divmul32와 64-bit division, multiplication, modulo 명령어
  • `packet`: 폐기된 packet access instructions

기본 64비트 명령어 인코딩

149-226

BPF는 다음 두 instruction encodings를 사용합니다.

  • basic instruction encoding: 한 명령어를 64 bits에 인코딩
  • wide instruction encoding: 기본 명령어 뒤에 64 bits를 더 붙여 총 128 bits 사용

기본 명령어의 필드 배치는 다음과 같습니다.

Basic instruction encoding
비트 범위필드
0-7opcode8 bits
8-15regs8 bits
16-31offset16 bits
32-63imm32 bits

64비트 명령어를 opcode, regs, offset, imm 필드로 나눕니다.

`opcode`는 수행할 연산이며 상위 specific 비트와 하위 class 비트로 나뉩니다.

Opcode field
비트 범위이름의미
7-3specific명령어 class별 세부 형식
2-0classinstruction class

8비트 opcode의 상위 5비트 의미는 명령어 class마다 달라집니다.

`regs`는 source와 destination register 번호를 담습니다. little-endian host와 big-endian host는 4비트 nibble 순서가 반대입니다.

Register field by host endianness
호스트상위 4비트하위 4비트
little-endiansrc_regdst_reg
big-endiandst_regsrc_reg

각 레지스터 번호는 0-10이며 wide immediate 명령어 등은 src_reg를 다른 목적으로 재사용할 수 있습니다.

`src_reg`는 별도 규정이 없으면 source register 0-10이고, `dst_reg`는 destination register 0-10입니다. `offset`은 주로 포인터 산술에 쓰는 signed integer offset이며, `imm`은 signed integer immediate value입니다.

다중 바이트 필드인 `offset`과 `imm`은 big-endian host에서는 big-endian byte ordering, little-endian host에서는 little-endian byte ordering으로 저장합니다.

다음 예는 같은 `r1 += 0x11223344`를 두 호스트 순서로 인코딩합니다.

opcode                  offset imm          assembly
       src_reg dst_reg
07     0       1        00 00  44 33 22 11  r1 += 0x11223344 // little
       dst_reg src_reg
07     1       0        00 00  11 22 33 44  r1 += 0x11223344 // big

대부분의 명령어는 모든 필드를 사용하지 않습니다. 사용하지 않는 필드는 반드시 0으로 지워야 합니다.

Wide 인코딩과 명령어 class

227-286

일부 명령어는 두 개의 32-bit immediate 값을 쓰는 wide instruction encoding을 사용합니다. 기본 형식 뒤의 64 bits는 `opcode`, `dst_reg`, `src_reg`, `offset`이 모두 0인 pseudo instruction입니다.

Wide instruction encoding
64-bit word상위 영역하위 32 bits
첫 번째opcode + regs + offsetimm
두 번째reserved (0)next_imm

첫 64비트에 기본 명령어와 imm을, 다음 64비트에 reserved와 next_imm을 둡니다.

`opcode`, `regs`, `offset`, `imm`은 기본 인코딩과 같은 의미입니다. `reserved`는 사용하지 않으므로 0이고, `next_imm`은 두 번째 signed integer immediate value입니다.

`opcode`의 최하위 3 bits가 instruction class를 저장합니다.

classvalue설명참조
LD0x0non-standard load operationsLoad and store instructions
LDX0x1load into register operationsLoad and store instructions
ST0x2store from immediate operationsLoad and store instructions
STX0x3store from register operationsLoad and store instructions
ALU0x432-bit arithmetic operationsArithmetic and jump instructions
JMP0x564-bit jump operationsArithmetic and jump instructions
JMP320x632-bit jump operationsArithmetic and jump instructions
ALU640x764-bit arithmetic operationsArithmetic and jump instructions

산술·점프 opcode 형식

287-314

`ALU`, `ALU64`, `JMP`, `JMP32`의 8-bit opcode는 code, source bit `s`, class 세 부분으로 나뉩니다.

Arithmetic and jump opcode
비트 범위필드의미
7-4codeclass별 operation code
3ssource operand location
2-0classinstruction class

상위 code는 연산, 가운데 s는 source 위치, 하위 class는 명령어 class를 선택합니다.

sourcevalue설명
K032-bit imm 값을 source operand로 사용
X1src_reg 값을 source operand로 사용

산술 명령어 의미

315-434

`ALU`는 32-bit operands, `ALU64`는 64-bit operands를 사용하며 나머지 연산 의미는 같습니다. 별도 표기가 없으면 ALU64 명령어는 base64 준수 그룹에 속합니다. 아래에서 `src`는 source operand, `dst`는 destination register 값입니다.

namecodeoffset설명
ADD0x00dst += src
SUB0x10dst -= src
MUL0x20dst *= src
DIV0x30dst = (src != 0) ? (dst / src) : 0
SDIV0x31dst = (src == 0) ? 0 : ((src == -1 && dst == LLONG_MIN) ? LLONG_MIN : (dst s/ src))
OR0x40dst |= src
AND0x50dst &= src
LSH0x60dst <<= (src & mask)
RSH0x70dst >>= (src & mask)
NEG0x80dst = -dst
MOD0x90dst = (src != 0) ? (dst % src) : dst
SMOD0x91dst = (src == 0) ? dst : ((src == -1 && dst == LLONG_MIN) ? 0 : (dst s% src))
XOR0xa0dst ^= src
MOV0xb0dst = src
MOVSX0xb8/16/32dst = (s8,s16,s32)src
ARSH0xc0sign extending dst >>= (src & mask)
END0xd0byte swap operations

산술 underflow와 overflow는 허용되어 32-bit 또는 64-bit 값이 wrap됩니다. 0으로 나누면 destination register는 0이 됩니다. `LLONG_MIN / -1`은 ALU64에서 `LLONG_MIN`, `INT_MIN / -1`은 ALU에서 `INT_MIN`이 됩니다.

0으로 modulo하면 ALU64에서는 destination이 유지되고 ALU에서는 상위 32 bits가 0이 됩니다. `LLONG_MIN % -1`과 `INT_MIN % -1`은 각각 0이 됩니다.

`{ADD, X, ALU}`는 다음과 같습니다.

dst = (u32) ((u32) dst + (u32) src)

`(u32)`는 상위 32 bits를 0으로 만든다는 뜻입니다. `{ADD, X, ALU64}`는 다음과 같습니다.

dst = dst + src

`{XOR, K, ALU}`는 다음과 같습니다.

dst = (u32) dst ^ (u32) imm

`{XOR, K, ALU64}`는 다음과 같습니다.

dst = dst ^ imm

대부분의 산술 명령어는 offset이 0이며 `SDIV`, `SMOD`, `MOVSX`만 0이 아닌 offset을 사용합니다. ALU의 division, multiplication, modulo는 divmul32, ALU64의 해당 연산은 divmul64 그룹입니다.

unsigned `DIV`와 `MOD`에서 ALU의 imm은 32-bit unsigned 값입니다. ALU64에서는 imm을 32에서 64 bits로 sign extend한 뒤 64-bit unsigned로 해석합니다. signed `SDIV`와 `SMOD`에서는 각각 32-bit signed와 sign-extended 64-bit signed로 해석합니다.

음수 피제수나 제수에 대한 signed modulo 정의는 언어마다 다릅니다. 이 규격은 C와 Go처럼 truncated division을 반드시 사용하며 `-13 % 3 == -1`입니다.

a % n = a - n * trunc(a / n)

`MOVSX`는 sign extension을 포함한 move입니다. `{MOVSX, X, ALU}`는 8/16-bit 피연산자를 32-bit로 확장하고 상위 32 bits를 0으로 만듭니다. `{MOVSX, X, ALU64}`는 8/16/32-bit 피연산자를 64-bit로 확장합니다. MOVSX는 register source operand `X`에만 정의됩니다.

`{MOV, K, ALU64}`는 다음과 같습니다.

dst = (s64)imm

`{MOV, X, ALU}`는 다음과 같습니다.

dst = (u32)src

offset이 8인 `{MOVSX, X, ALU}`는 다음과 같습니다.

dst = (u32)(s32)(s8)src

`NEG`는 source bit가 `K`로 clear된 경우에만 정의됩니다. Shift는 64-bit 연산에 `0x3F (63)`, 32-bit 연산에 `0x1F (31)` mask를 사용합니다.

바이트 순서 변환

435-483

byte swap instructions는 `ALU` 또는 `ALU64` class와 4-bit code `END`를 사용합니다. destination register에만 작용하며 별도 source register나 immediate value를 쓰지 않습니다.

ALU에서는 opcode의 1-bit source operand field가 변환할 byte order를 선택합니다. ALU64에서는 이 비트가 reserved이며 반드시 0이어야 합니다.

classsourcevalue설명
ALULE0host byte order와 little endian 사이 변환
ALUBE1host byte order와 big endian 사이 변환
ALU64Reserved0무조건 byte swap

imm은 swap 폭을 인코딩하며 16, 32, 64를 지원합니다. 64-bit 연산은 base64, 나머지는 base32 그룹입니다.

`{END, LE, ALU}`에서 imm = 16/32/64인 경우입니다.

dst = le16(dst)
dst = le32(dst)
dst = le64(dst)

`{END, BE, ALU}`에서 imm = 16/32/64인 경우입니다.

dst = be16(dst)
dst = be32(dst)
dst = be64(dst)

`{END, TO, ALU64}`에서 imm = 16/32/64인 경우입니다.

dst = bswap16(dst)
dst = bswap32(dst)
dst = bswap64(dst)

점프, helper와 프로그램 로컬 함수

484-577

`JMP32`는 32-bit 피연산자와 base32 그룹을 사용하고, `JMP`는 같은 연산에 64-bit 피연산자를 사용하며 별도 표기가 없으면 base64 그룹입니다.

codevaluesrc_reg설명비고
JA0x00x0PC += offset{JA, K, JMP} only
JA0x00x0PC += imm{JA, K, JMP32} only
JEQ0x1anyPC += offset if dst == src
JGT0x2anyPC += offset if dst > srcunsigned
JGE0x3anyPC += offset if dst >= srcunsigned
JSET0x4anyPC += offset if dst & src
JNE0x5anyPC += offset if dst != src
JSGT0x6anyPC += offset if dst > srcsigned
JSGE0x7anyPC += offset if dst >= srcsigned
CALL0x80x0static ID의 helper function 호출{CALL, K, JMP} only
CALL0x80x1call PC += immprogram-local function
CALL0x80x2BTF ID의 helper function 호출{CALL, K, JMP} only
EXIT0x90x0return{CALL, K, JMP} only
JLT0xaanyPC += offset if dst < srcunsigned
JLE0xbanyPC += offset if dst <= srcunsigned
JSLT0xcanyPC += offset if dst < srcsigned
JSLE0xdanyPC += offset if dst <= srcsigned

`PC`는 program counter입니다. offset 단위는 jump 다음 명령어를 기준으로 한 64-bit instructions입니다. `PC += 1`은 다음 basic instruction을 건너뛰지만 다음 명령어가 128-bit wide instruction이면 undefined behavior입니다.

`{JSGE, X, JMP32}`는 signed 비교를 수행합니다.

if (s32)dst s>= (s32)src goto +offset

`{JLE, K, JMP}`는 다음과 같습니다.

if dst <= (u64)(s64)imm goto +offset

`{JA, K, JMP32}`는 imm 필드의 32-bit branch offset을 사용합니다.

gotol +imm

JA에는 두 형식이 있습니다. JMP class는 offset 필드의 16-bit jump offset을, JMP32 class는 imm 필드의 32-bit jump offset을 사용합니다. 16 bits보다 큰 conditional jump는 작은 conditional jump와 32-bit unconditional jump로 바꿀 수 있습니다. 모든 CALL과 JA는 base32 그룹입니다.

Helper functions를 통해 BPF 프로그램은 기반 플랫폼이 노출한 함수 집합을 호출합니다. 전통적으로 helper는 imm에 인코딩한 static ID로 식별합니다. BPF Type Format (BTF)을 지원하는 플랫폼은 helper 이름과 형식을 나타내는 BTF ID를 imm에 넣을 수 있습니다. 구체적인 helper와 BTF 표준화는 이 문서 범위 밖이며 플랫폼 문서를 참조해야 합니다.

Program-local functions는 호출자와 같은 BPF 프로그램이 노출하며, call 다음 명령어를 기준으로 한 offset을 imm에 인코딩합니다. 로컬 함수 안의 `EXIT`는 호출자에게 돌아갑니다.

Load/store opcode 형식

578-622

`LD`, `LDX`, `ST`, `STX`의 8-bit opcode는 mode, size `sz`, class로 나뉩니다.

Load and store opcode
비트 범위필드의미
7-5modemode modifier
4-3szsize modifier
2-0classinstruction class

상위 mode는 접근 방식, 가운데 sz는 폭, 하위 class는 load/store 종류를 선택합니다.

mode modifiervalue설명참조
IMM064-bit immediate instructions64-bit immediate instructions
ABS1legacy BPF packet access (absolute)Legacy BPF Packet access instructions
IND2legacy BPF packet access (indirect)Legacy BPF Packet access instructions
MEM3regular load and store operationsRegular load and store operations
MEMSX4sign-extension load operationsSign-extension load operations
ATOMIC6atomic operationsAtomic operations
sizevalue설명
W0word (4 bytes)
H1half word (2 bytes)
B2byte
DW3double word (8 bytes)

`DW`를 사용하는 명령어는 base64 준수 그룹에 속하며 class 필드는 앞서 정의한 instruction class입니다.

일반 및 부호 확장 메모리 접근

623-656

`MEM` mode modifier는 register와 memory 사이에서 데이터를 옮기는 일반 load/store 명령어를 인코딩합니다.

`{MEM, <size>, STX}`는 register source를 저장합니다.

*(size *) (dst + offset) = src

`{MEM, <size>, ST}`는 immediate를 저장합니다.

*(size *) (dst + offset) = imm

`{MEM, <size>, LDX}`는 unsigned 값을 register로 읽습니다.

dst = *(unsigned size *) (src + offset)

`<size>`는 `B`, `H`, `W`, `DW` 중 하나이며 대응하는 unsigned size는 u8, u16, u32, u64입니다.

`MEMSX` mode modifier는 register와 memory 사이의 sign-extension load를 인코딩합니다. `{MEMSX, <size>, LDX}`의 의미는 다음과 같습니다.

dst = *(signed size *) (src + offset)

이 경우 `<size>`는 `B`, `H`, `W` 중 하나이고 signed size는 s8, s16, s32입니다.

원자적 메모리 연산

657-723

Atomic operations는 memory에 작용하며 다른 BPF 프로그램이나 규격 외부 수단이 같은 memory region에 접근해도 중단되거나 손상되지 않습니다.

BPF의 모든 atomic operation은 `ATOMIC` mode modifier를 쓰는 store operation으로 인코딩합니다.

  • `{ATOMIC, W, STX}`: atomic32 그룹의 32-bit operation
  • `{ATOMIC, DW, STX}`: atomic64 그룹의 64-bit operation
  • 8-bit와 16-bit atomic operations는 지원하지 않음

imm 필드가 실제 atomic operation을 인코딩합니다.

immvalue설명
ADD0x00atomic add
OR0x40atomic or
AND0x50atomic and
XOR0xa0atomic xor

`{ATOMIC, W, STX}`에서 imm = ADD인 경우입니다.

*(u32 *)(dst + offset) += src

`{ATOMIC, DW, STX}`에서 imm = ADD인 경우입니다.

*(u64 *)(dst + offset) += src

단순 연산 외에 modifier 하나와 복합 연산 두 개가 있습니다.

immvalue설명
FETCH0x01modifier: 이전 값 반환
XCHG0xe0 | FETCHatomic exchange
CMPXCHG0xf0 | FETCHatomic compare and exchange

`FETCH`는 단순 atomic operation에서는 선택 사항이고 복합 연산에서는 항상 설정됩니다. 설정되면 연산 전 memory 값을 `src`에 덮어씁니다. `XCHG`는 `src`와 `dst + offset`의 값을 원자적으로 교환합니다.

`CMPXCHG`는 `dst + offset`의 값을 `R0`과 원자적으로 비교합니다. 같으면 해당 위치를 `src`로 교체합니다. 일치 여부와 무관하게 연산 전 값은 zero-extend되어 `R0`에 적재됩니다.

64비트 immediate, 맵과 플랫폼 변수

724-780

`IMM` mode modifier 명령어는 wide instruction encoding을 쓰며 basic instruction의 src_reg에 opcode subtype을 저장합니다.

src_regpseudocodeimm typedst type
0x0dst = (next_imm << 32) | immintegerinteger
0x1dst = map_by_fd(imm)map fdmap
0x2dst = map_val(map_by_fd(imm)) + next_immmap fddata address
0x3dst = var_addr(imm)variable iddata address
0x4dst = code_addr(imm)integercode address
0x5dst = map_by_idx(imm)map indexmap
0x6dst = map_val(map_by_idx(imm)) + next_immmap indexdata address

표에서 사용하는 함수와 형식은 다음 의미입니다.

  • `map_by_fd(imm)`: 32-bit file descriptor를 map address로 변환
  • `map_by_idx(imm)`: 32-bit index를 map address로 변환
  • `map_val(map)`: 주어진 map의 첫 value address를 반환
  • `var_addr(imm)`: 주어진 ID의 platform variable address를 반환
  • `code_addr(imm)`: 64-bit instruction 수로 나타낸 상대 offset의 instruction address를 반환
  • `imm type`: disassembler 표시용
  • `dst type`: verification과 JIT compilation용

Maps는 일부 플랫폼에서 BPF 프로그램이 접근하는 shared memory regions입니다. 맵의 구체적 의미는 별도 문서가 정하며 단일 연속 memory region이 없을 수도 있습니다. 현재 `map_val(map)`은 연속 영역이 하나인 맵에만 정의됩니다.

플랫폼이 지원하면 각 map은 file descriptor를 가질 수 있고 `map_by_fd(imm)`이 해당 fd의 맵을 얻습니다. 로드 시 프로그램에 맵 집합을 연결할 수도 있으며 `map_by_idx(imm)`은 그 집합에서 주어진 index의 맵을 얻습니다.

Platform Variables는 runtime이 노출하고 일부 플랫폼의 BPF 프로그램이 접근할 수 있는 integer ID 식별 memory region입니다. `var_addr(imm)`은 해당 ID의 영역 주소를 얻습니다.

폐기된 classic BPF 패킷 접근

781-790

BPF는 과거 classic BPF에서 이어받은 특수 packet data access instructions를 도입했습니다. 이 명령어는 `LD` class, `W`/`H`/`B` size modifier, `ABS` 또는 `IND` mode modifier를 사용했습니다. dst_reg와 offset은 0이고 ABS에서는 src_reg도 0이었습니다.

이 legacy packet access instructions는 deprecated되었으며 더 이상 사용해서는 안 됩니다. 모두 `packet` conformance group에 속합니다.