← Documents Documentation/filesystems/f2fs.rst GitHub 원문 ↗

Linux 6.18.37 · Filesystems

Flash-Friendly File System (F2FS)

F2FS의 LFS 설계, mount option, on-disk layout, node·directory, GC, compression, ZNS와 device aliasing 전문 번역입니다.

Source pathDocumentation/filesystems/f2fs.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

f2fs.rst:1-1028

F2FS는 NAND flash의 out-of-place write와 FTL geometry에 맞춘 LFS입니다. NAT로 wandering tree를 끊고, SIT·SSA 기반 GC와 six-log temperature 분리로 cleaning 비용을 제어합니다.

문서는 mount option 전체, 6개 on-disk area, checkpoint shadow copy, 3.94TiB node index, multi-level directory hash, write hint, pinned fallocate, cluster compression, ZNS와 device aliasing을 하나의 설계 흐름으로 설명합니다.

F2FS 전체 구조
SB·CP·SIT·NAT·SSA·Main area 배치NAT로 node address 변환Hot/Warm/Cold six-log allocationSIT·SSA로 victim 검증과 GCcheckpoint shadow copy로 consistency 유지compression·ZNS·device aliasing 확장 기능 적용

format부터 runtime data 관리까지의 핵심 관계입니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 =================================
4 Flash-Friendly File System (F2FS)
5 =================================
6
7 Overview
8 ========
9
10 NAND flash memory-based storage devices, such as SSD, eMMC, and SD cards, have
11 been equipped on a variety systems ranging from mobile to server systems. Since
12 they are known to have different characteristics from the conventional rotating
13 disks, a file system, an upper layer to the storage device, should adapt to the
14 changes from the sketch in the design level.
15
16 F2FS is a file system exploiting NAND flash memory-based storage devices, which
17 is based on Log-structured File System (LFS). The design has been focused on
18 addressing the fundamental issues in LFS, which are snowball effect of wandering
19 tree and high cleaning overhead.
20
21 Since a NAND flash memory-based storage device shows different characteristic
22 according to its internal geometry or flash memory management scheme, namely FTL,
23 F2FS and its tools support various parameters not only for configuring on-disk
24 layout, but also for selecting allocation and cleaning algorithms.
25
26 The following git tree provides the file system formatting tool (mkfs.f2fs),
27 a consistency checking tool (fsck.f2fs), and a debugging tool (dump.f2fs).
28
29 - git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs-tools.git
30
31 For sending patches, please use the following mailing list:
32
34
35 For reporting bugs, please use the following f2fs bug tracker link:
36
37 - https://bugzilla.kernel.org/enter_bug.cgi?product=File%20System&component=f2fs
38
39 Background and Design issues
40 ============================
41
42 Log-structured File System (LFS)
43 --------------------------------
44 "A log-structured file system writes all modifications to disk sequentially in
45 a log-like structure, thereby speeding up both file writing and crash recovery.
46 The log is the only structure on disk; it contains indexing information so that
47 files can be read back from the log efficiently. In order to maintain large free
48 areas on disk for fast writing, we divide the log into segments and use a
49 segment cleaner to compress the live information from heavily fragmented
50 segments." from Rosenblum, M. and Ousterhout, J. K., 1992, "The design and
51 implementation of a log-structured file system", ACM Trans. Computer Systems
52 10, 1, 26–52.
53
54 Wandering Tree Problem
55 ----------------------
56 In LFS, when a file data is updated and written to the end of log, its direct
57 pointer block is updated due to the changed location. Then the indirect pointer
58 block is also updated due to the direct pointer block update. In this manner,
59 the upper index structures such as inode, inode map, and checkpoint block are
60 also updated recursively. This problem is called as wandering tree problem [1],
61 and in order to enhance the performance, it should eliminate or relax the update
62 propagation as much as possible.
63
64 [1] Bityutskiy, A. 2005. JFFS3 design issues. http://www.linux-mtd.infradead.org/
65
66 Cleaning Overhead
67 -----------------
68 Since LFS is based on out-of-place writes, it produces so many obsolete blocks
69 scattered across the whole storage. In order to serve new empty log space, it
70 needs to reclaim these obsolete blocks seamlessly to users. This job is called
71 as a cleaning process.
72
73 The process consists of three operations as follows.
74
75 1. A victim segment is selected through referencing segment usage table.
76 2. It loads parent index structures of all the data in the victim identified by
77 segment summary blocks.
78 3. It checks the cross-reference between the data and its parent index structure.
79 4. It moves valid data selectively.
80
81 This cleaning job may cause unexpected long delays, so the most important goal
82 is to hide the latencies to users. And also definitely, it should reduce the
83 amount of valid data to be moved, and move them quickly as well.
84
85 Key Features
86 ============
87
88 Flash Awareness
89 ---------------
90 - Enlarge the random write area for better performance, but provide the high
91 spatial locality
92 - Align FS data structures to the operational units in FTL as best efforts
93
94 Wandering Tree Problem
95 ----------------------
96 - Use a term, “node”, that represents inodes as well as various pointer blocks
97 - Introduce Node Address Table (NAT) containing the locations of all the “node”
98 blocks; this will cut off the update propagation.
99
100 Cleaning Overhead
101 -----------------
102 - Support a background cleaning process
103 - Support greedy and cost-benefit algorithms for victim selection policies
104 - Support multi-head logs for static/dynamic hot and cold data separation
105 - Introduce adaptive logging for efficient block allocation
106
107 Mount Options
108 =============
109
110
111 ======================== ============================================================
112 background_gc=%s Turn on/off cleaning operations, namely garbage
113 collection, triggered in background when I/O subsystem is
114 idle. If background_gc=on, it will turn on the garbage
115 collection and if background_gc=off, garbage collection
116 will be turned off. If background_gc=sync, it will turn
117 on synchronous garbage collection running in background.
118 Default value for this option is on. So garbage
119 collection is on by default.
120 gc_merge When background_gc is on, this option can be enabled to
121 let background GC thread to handle foreground GC requests,
122 it can eliminate the sluggish issue caused by slow foreground
123 GC operation when GC is triggered from a process with limited
124 I/O and CPU resources.
125 nogc_merge Disable GC merge feature.
126 disable_roll_forward Disable the roll-forward recovery routine
127 norecovery Disable the roll-forward recovery routine, mounted read-
128 only (i.e., -o ro,disable_roll_forward)
129 discard/nodiscard Enable/disable real-time discard in f2fs, if discard is
130 enabled, f2fs will issue discard/TRIM commands when a
131 segment is cleaned.
132 heap/no_heap Deprecated.
133 nouser_xattr Disable Extended User Attributes. Note: xattr is enabled
134 by default if CONFIG_F2FS_FS_XATTR is selected.
135 noacl Disable POSIX Access Control List. Note: acl is enabled
136 by default if CONFIG_F2FS_FS_POSIX_ACL is selected.
137 active_logs=%u Support configuring the number of active logs. In the
138 current design, f2fs supports only 2, 4, and 6 logs.
139 Default number is 6.
140 disable_ext_identify Disable the extension list configured by mkfs, so f2fs
141 is not aware of cold files such as media files.
142 inline_xattr Enable the inline xattrs feature.
143 noinline_xattr Disable the inline xattrs feature.
144 inline_xattr_size=%u Support configuring inline xattr size, it depends on
145 flexible inline xattr feature.
146 inline_data Enable the inline data feature: Newly created small (<~3.4k)
147 files can be written into inode block.
148 inline_dentry Enable the inline dir feature: data in newly created
149 directory entries can be written into inode block. The
150 space of inode block which is used to store inline
151 dentries is limited to ~3.4k.
152 noinline_dentry Disable the inline dentry feature.
153 flush_merge Merge concurrent cache_flush commands as much as possible
154 to eliminate redundant command issues. If the underlying
155 device handles the cache_flush command relatively slowly,
156 recommend to enable this option.
157 nobarrier This option can be used if underlying storage guarantees
158 its cached data should be written to the novolatile area.
159 If this option is set, no cache_flush commands are issued
160 but f2fs still guarantees the write ordering of all the
161 data writes.
162 barrier If this option is set, cache_flush commands are allowed to be
163 issued.
164 fastboot This option is used when a system wants to reduce mount
165 time as much as possible, even though normal performance
166 can be sacrificed.
167 extent_cache Enable an extent cache based on rb-tree, it can cache
168 as many as extent which map between contiguous logical
169 address and physical address per inode, resulting in
170 increasing the cache hit ratio. Set by default.
171 noextent_cache Disable an extent cache based on rb-tree explicitly, see
172 the above extent_cache mount option.
173 noinline_data Disable the inline data feature, inline data feature is
174 enabled by default.
175 data_flush Enable data flushing before checkpoint in order to
176 persist data of regular and symlink.
177 reserve_root=%d Support configuring reserved space which is used for
178 allocation from a privileged user with specified uid or
179 gid, unit: 4KB, the default limit is 12.5% of user blocks.
180 reserve_node=%d Support configuring reserved nodes which are used for
181 allocation from a privileged user with specified uid or
182 gid, the default limit is 12.5% of all nodes.
183 resuid=%d The user ID which may use the reserved blocks and nodes.
184 resgid=%d The group ID which may use the reserved blocks and nodes.
185 fault_injection=%d Enable fault injection in all supported types with
186 specified injection rate.
187 fault_type=%d Support configuring fault injection type, should be
188 enabled with fault_injection option, fault type value
189 is shown below, it supports single or combined type.
190
191 =========================== ==========
192 Type_Name Type_Value
193 =========================== ==========
194 FAULT_KMALLOC 0x00000001
195 FAULT_KVMALLOC 0x00000002
196 FAULT_PAGE_ALLOC 0x00000004
197 FAULT_PAGE_GET 0x00000008
198 FAULT_ALLOC_BIO 0x00000010 (obsolete)
199 FAULT_ALLOC_NID 0x00000020
200 FAULT_ORPHAN 0x00000040
201 FAULT_BLOCK 0x00000080
202 FAULT_DIR_DEPTH 0x00000100
203 FAULT_EVICT_INODE 0x00000200
204 FAULT_TRUNCATE 0x00000400
205 FAULT_READ_IO 0x00000800
206 FAULT_CHECKPOINT 0x00001000
207 FAULT_DISCARD 0x00002000
208 FAULT_WRITE_IO 0x00004000
209 FAULT_SLAB_ALLOC 0x00008000
210 FAULT_DQUOT_INIT 0x00010000
211 FAULT_LOCK_OP 0x00020000
212 FAULT_BLKADDR_VALIDITY 0x00040000
213 FAULT_BLKADDR_CONSISTENCE 0x00080000
214 FAULT_NO_SEGMENT 0x00100000
215 FAULT_INCONSISTENT_FOOTER 0x00200000
216 FAULT_TIMEOUT 0x00400000 (1000ms)
217 FAULT_VMALLOC 0x00800000
218 =========================== ==========
219 mode=%s Control block allocation mode which supports "adaptive"
220 and "lfs". In "lfs" mode, there should be no random
221 writes towards main area.
222 "fragment:segment" and "fragment:block" are newly added here.
223 These are developer options for experiments to simulate filesystem
224 fragmentation/after-GC situation itself. The developers use these
225 modes to understand filesystem fragmentation/after-GC condition well,
226 and eventually get some insights to handle them better.
227 In "fragment:segment", f2fs allocates a new segment in random
228 position. With this, we can simulate the after-GC condition.
229 In "fragment:block", we can scatter block allocation with
230 "max_fragment_chunk" and "max_fragment_hole" sysfs nodes.
231 We added some randomness to both chunk and hole size to make
232 it close to realistic IO pattern. So, in this mode, f2fs will allocate
233 1..<max_fragment_chunk> blocks in a chunk and make a hole in the
234 length of 1..<max_fragment_hole> by turns. With this, the newly
235 allocated blocks will be scattered throughout the whole partition.
236 Note that "fragment:block" implicitly enables "fragment:segment"
237 option for more randomness.
238 Please, use these options for your experiments and we strongly
239 recommend to re-format the filesystem after using these options.
240 usrquota Enable plain user disk quota accounting.
241 grpquota Enable plain group disk quota accounting.
242 prjquota Enable plain project quota accounting.
243 usrjquota=<file> Appoint specified file and type during mount, so that quota
244 grpjquota=<file> information can be properly updated during recovery flow,
245 prjjquota=<file> <quota file>: must be in root directory;
246 jqfmt=<quota type> <quota type>: [vfsold,vfsv0,vfsv1].
247 usrjquota= Turn off user journalled quota.
248 grpjquota= Turn off group journalled quota.
249 prjjquota= Turn off project journalled quota.
250 quota Enable plain user disk quota accounting.
251 noquota Disable all plain disk quota option.
252 alloc_mode=%s Adjust block allocation policy, which supports "reuse"
253 and "default".
254 fsync_mode=%s Control the policy of fsync. Currently supports "posix",
255 "strict", and "nobarrier". In "posix" mode, which is
256 default, fsync will follow POSIX semantics and does a
257 light operation to improve the filesystem performance.
258 In "strict" mode, fsync will be heavy and behaves in line
259 with xfs, ext4 and btrfs, where xfstest generic/342 will
260 pass, but the performance will regress. "nobarrier" is
261 based on "posix", but doesn't issue flush command for
262 non-atomic files likewise "nobarrier" mount option.
263 test_dummy_encryption
264 test_dummy_encryption=%s
265 Enable dummy encryption, which provides a fake fscrypt
266 context. The fake fscrypt context is used by xfstests.
267 The argument may be either "v1" or "v2", in order to
268 select the corresponding fscrypt policy version.
269 checkpoint=%s[:%u[%]] Set to "disable" to turn off checkpointing. Set to "enable"
270 to re-enable checkpointing. Is enabled by default. While
271 disabled, any unmounting or unexpected shutdowns will cause
272 the filesystem contents to appear as they did when the
273 filesystem was mounted with that option.
274 While mounting with checkpoint=disable, the filesystem must
275 run garbage collection to ensure that all available space can
276 be used. If this takes too much time, the mount may return
277 EAGAIN. You may optionally add a value to indicate how much
278 of the disk you would be willing to temporarily give up to
279 avoid additional garbage collection. This can be given as a
280 number of blocks, or as a percent. For instance, mounting
281 with checkpoint=disable:100% would always succeed, but it may
282 hide up to all remaining free space. The actual space that
283 would be unusable can be viewed at /sys/fs/f2fs/<disk>/unusable
284 This space is reclaimed once checkpoint=enable.
285 checkpoint_merge When checkpoint is enabled, this can be used to create a kernel
286 daemon and make it to merge concurrent checkpoint requests as
287 much as possible to eliminate redundant checkpoint issues. Plus,
288 we can eliminate the sluggish issue caused by slow checkpoint
289 operation when the checkpoint is done in a process context in
290 a cgroup having low i/o budget and cpu shares. To make this
291 do better, we set the default i/o priority of the kernel daemon
292 to "3", to give one higher priority than other kernel threads.
293 This is the same way to give a I/O priority to the jbd2
294 journaling thread of ext4 filesystem.
295 nocheckpoint_merge Disable checkpoint merge feature.
296 compress_algorithm=%s Control compress algorithm, currently f2fs supports "lzo",
297 "lz4", "zstd" and "lzo-rle" algorithm.
298 compress_algorithm=%s:%d Control compress algorithm and its compress level, now, only
299 "lz4" and "zstd" support compress level config.
300
301 ========= ===========
302 algorithm level range
303 ========= ===========
304 lz4 3 - 16
305 zstd 1 - 22
306 ========= ===========
307 compress_log_size=%u Support configuring compress cluster size. The size will
308 be 4KB * (1 << %u). The default and minimum sizes are 16KB.
309 compress_extension=%s Support adding specified extension, so that f2fs can enable
310 compression on those corresponding files, e.g. if all files
311 with '.ext' has high compression rate, we can set the '.ext'
312 on compression extension list and enable compression on
313 these file by default rather than to enable it via ioctl.
314 For other files, we can still enable compression via ioctl.
315 Note that, there is one reserved special extension '*', it
316 can be set to enable compression for all files.
317 nocompress_extension=%s Support adding specified extension, so that f2fs can disable
318 compression on those corresponding files, just contrary to compression extension.
319 If you know exactly which files cannot be compressed, you can use this.
320 The same extension name can't appear in both compress and nocompress
321 extension at the same time.
322 If the compress extension specifies all files, the types specified by the
323 nocompress extension will be treated as special cases and will not be compressed.
324 Don't allow use '*' to specifie all file in nocompress extension.
325 After add nocompress_extension, the priority should be:
326 dir_flag < comp_extention,nocompress_extension < comp_file_flag,no_comp_file_flag.
327 See more in compression sections.
328
329 compress_chksum Support verifying chksum of raw data in compressed cluster.
330 compress_mode=%s Control file compression mode. This supports "fs" and "user"
331 modes. In "fs" mode (default), f2fs does automatic compression
332 on the compression enabled files. In "user" mode, f2fs disables
333 the automaic compression and gives the user discretion of
334 choosing the target file and the timing. The user can do manual
335 compression/decompression on the compression enabled files using
336 ioctls.
337 compress_cache Support to use address space of a filesystem managed inode to
338 cache compressed block, in order to improve cache hit ratio of
339 random read.
340 inlinecrypt When possible, encrypt/decrypt the contents of encrypted
341 files using the blk-crypto framework rather than
342 filesystem-layer encryption. This allows the use of
343 inline encryption hardware. The on-disk format is
344 unaffected. For more details, see
345 Documentation/block/inline-encryption.rst.
346 atgc Enable age-threshold garbage collection, it provides high
347 effectiveness and efficiency on background GC.
348 discard_unit=%s Control discard unit, the argument can be "block", "segment"
349 and "section", issued discard command's offset/size will be
350 aligned to the unit, by default, "discard_unit=block" is set,
351 so that small discard functionality is enabled.
352 For blkzoned device, "discard_unit=section" will be set by
353 default, it is helpful for large sized SMR or ZNS devices to
354 reduce memory cost by getting rid of fs metadata supports small
355 discard.
356 memory=%s Control memory mode. This supports "normal" and "low" modes.
357 "low" mode is introduced to support low memory devices.
358 Because of the nature of low memory devices, in this mode, f2fs
359 will try to save memory sometimes by sacrificing performance.
360 "normal" mode is the default mode and same as before.
361 age_extent_cache Enable an age extent cache based on rb-tree. It records
362 data block update frequency of the extent per inode, in
363 order to provide better temperature hints for data block
364 allocation.
365 errors=%s Specify f2fs behavior on critical errors. This supports modes:
366 "panic", "continue" and "remount-ro", respectively, trigger
367 panic immediately, continue without doing anything, and remount
368 the partition in read-only mode. By default it uses "continue"
369 mode.
370
371 ====================== =============== =============== ========
372 mode continue remount-ro panic
373 ====================== =============== =============== ========
374 access ops normal normal N/A
375 syscall errors -EIO -EROFS N/A
376 mount option rw ro N/A
377 pending dir write keep keep N/A
378 pending non-dir write drop keep N/A
379 pending node write drop keep N/A
380 pending meta write keep keep N/A
381 ====================== =============== =============== ========
382 nat_bits Enable nat_bits feature to enhance full/empty nat blocks access,
383 by default it's disabled.
384 lookup_mode=%s Control the directory lookup behavior for casefolded
385 directories. This option has no effect on directories
386 that do not have the casefold feature enabled.
387
388 ================== ========================================
389 Value Description
390 ================== ========================================
391 perf (Default) Enforces a hash-only lookup.
392 The linear search fallback is always
393 disabled, ignoring the on-disk flag.
394 compat Enables the linear search fallback for
395 compatibility with directory entries
396 created by older kernel that used a
397 different case-folding algorithm.
398 This mode ignores the on-disk flag.
399 auto F2FS determines the mode based on the
400 on-disk `SB_ENC_NO_COMPAT_FALLBACK_FL`
401 flag.
402 ================== ========================================
403 ======================== ============================================================
404
405 Debugfs Entries
406 ===============
407
408 /sys/kernel/debug/f2fs/ contains information about all the partitions mounted as
409 f2fs. Each file shows the whole f2fs information.
410
411 /sys/kernel/debug/f2fs/status includes:
412
413 - major file system information managed by f2fs currently
414 - average SIT information about whole segments
415 - current memory footprint consumed by f2fs.
416
417 Sysfs Entries
418 =============
419
420 Information about mounted f2fs file systems can be found in
421 /sys/fs/f2fs. Each mounted filesystem will have a directory in
422 /sys/fs/f2fs based on its device name (i.e., /sys/fs/f2fs/sda).
423 The files in each per-device directory are shown in table below.
424
425 Files in /sys/fs/f2fs/<devname>
426 (see also Documentation/ABI/testing/sysfs-fs-f2fs)
427
428 Usage
429 =====
430
431 1. Download userland tools and compile them.
432
433 2. Skip, if f2fs was compiled statically inside kernel.
434 Otherwise, insert the f2fs.ko module::
435
436 # insmod f2fs.ko
437
438 3. Create a directory to use when mounting::
439
440 # mkdir /mnt/f2fs
441
442 4. Format the block device, and then mount as f2fs::
443
444 # mkfs.f2fs -l label /dev/block_device
445 # mount -t f2fs /dev/block_device /mnt/f2fs
446
447 mkfs.f2fs
448 ---------
449 The mkfs.f2fs is for the use of formatting a partition as the f2fs filesystem,
450 which builds a basic on-disk layout.
451
452 The quick options consist of:
453
454 =============== ===========================================================
455 ``-l [label]`` Give a volume label, up to 512 unicode name.
456 ``-a [0 or 1]`` Split start location of each area for heap-based allocation.
457
458 1 is set by default, which performs this.
459 ``-o [int]`` Set overprovision ratio in percent over volume size.
460
461 5 is set by default.
462 ``-s [int]`` Set the number of segments per section.
463
464 1 is set by default.
465 ``-z [int]`` Set the number of sections per zone.
466
467 1 is set by default.
468 ``-e [str]`` Set basic extension list. e.g. "mp3,gif,mov"
469 ``-t [0 or 1]`` Disable discard command or not.
470
471 1 is set by default, which conducts discard.
472 =============== ===========================================================
473
474 Note: please refer to the manpage of mkfs.f2fs(8) to get full option list.
475
476 fsck.f2fs
477 ---------
478 The fsck.f2fs is a tool to check the consistency of an f2fs-formatted
479 partition, which examines whether the filesystem metadata and user-made data
480 are cross-referenced correctly or not.
481 Note that, initial version of the tool does not fix any inconsistency.
482
483 The quick options consist of::
484
485 -d debug level [default:0]
486
487 Note: please refer to the manpage of fsck.f2fs(8) to get full option list.
488
489 dump.f2fs
490 ---------
491 The dump.f2fs shows the information of specific inode and dumps SSA and SIT to
492 file. Each file is dump_ssa and dump_sit.
493
494 The dump.f2fs is used to debug on-disk data structures of the f2fs filesystem.
495 It shows on-disk inode information recognized by a given inode number, and is
496 able to dump all the SSA and SIT entries into predefined files, ./dump_ssa and
497 ./dump_sit respectively.
498
499 The options consist of::
500
501 -d debug level [default:0]
502 -i inode no (hex)
503 -s [SIT dump segno from #1~#2 (decimal), for all 0~-1]
504 -a [SSA dump segno from #1~#2 (decimal), for all 0~-1]
505
506 Examples::
507
508 # dump.f2fs -i [ino] /dev/sdx
509 # dump.f2fs -s 0~-1 /dev/sdx (SIT dump)
510 # dump.f2fs -a 0~-1 /dev/sdx (SSA dump)
511
512 Note: please refer to the manpage of dump.f2fs(8) to get full option list.
513
514 sload.f2fs
515 ----------
516 The sload.f2fs gives a way to insert files and directories in the existing disk
517 image. This tool is useful when building f2fs images given compiled files.
518
519 Note: please refer to the manpage of sload.f2fs(8) to get full option list.
520
521 resize.f2fs
522 -----------
523 The resize.f2fs lets a user resize the f2fs-formatted disk image, while preserving
524 all the files and directories stored in the image.
525
526 Note: please refer to the manpage of resize.f2fs(8) to get full option list.
527
528 defrag.f2fs
529 -----------
530 The defrag.f2fs can be used to defragment scattered written data as well as
531 filesystem metadata across the disk. This can improve the write speed by giving
532 more free consecutive space.
533
534 Note: please refer to the manpage of defrag.f2fs(8) to get full option list.
535
536 f2fs_io
537 -------
538 The f2fs_io is a simple tool to issue various filesystem APIs as well as
539 f2fs-specific ones, which is very useful for QA tests.
540
541 Note: please refer to the manpage of f2fs_io(8) to get full option list.
542
543 Design
544 ======
545
546 On-disk Layout
547 --------------
548
549 F2FS divides the whole volume into a number of segments, each of which is fixed
550 to 2MB in size. A section is composed of consecutive segments, and a zone
551 consists of a set of sections. By default, section and zone sizes are set to one
552 segment size identically, but users can easily modify the sizes by mkfs.
553
554 F2FS splits the entire volume into six areas, and all the areas except superblock
555 consist of multiple segments as described below::
556
557 align with the zone size <-|
558 |-> align with the segment size
559 _________________________________________________________________________
560 | | | Segment | Node | Segment | |
561 | Superblock | Checkpoint | Info. | Address | Summary | Main |
562 | (SB) | (CP) | Table (SIT) | Table (NAT) | Area (SSA) | |
563 |____________|_____2______|______N______|______N______|______N_____|__N___|
564 . .
565 . .
566 . .
567 ._________________________________________.
568 |_Segment_|_..._|_Segment_|_..._|_Segment_|
569 . .
570 ._________._________
571 |_section_|__...__|_
572 . .
573 .________.
574 |__zone__|
575
576 - Superblock (SB)
577 It is located at the beginning of the partition, and there exist two copies
578 to avoid file system crash. It contains basic partition information and some
579 default parameters of f2fs.
580
581 - Checkpoint (CP)
582 It contains file system information, bitmaps for valid NAT/SIT sets, orphan
583 inode lists, and summary entries of current active segments.
584
585 - Segment Information Table (SIT)
586 It contains segment information such as valid block count and bitmap for the
587 validity of all the blocks.
588
589 - Node Address Table (NAT)
590 It is composed of a block address table for all the node blocks stored in
591 Main area.
592
593 - Segment Summary Area (SSA)
594 It contains summary entries which contains the owner information of all the
595 data and node blocks stored in Main area.
596
597 - Main Area
598 It contains file and directory data including their indices.
599
600 In order to avoid misalignment between file system and flash-based storage, F2FS
601 aligns the start block address of CP with the segment size. Also, it aligns the
602 start block address of Main area with the zone size by reserving some segments
603 in SSA area.
604
605 Reference the following survey for additional technical details.
606 https://wiki.linaro.org/WorkingGroups/Kernel/Projects/FlashCardSurvey
607
608 File System Metadata Structure
609 ------------------------------
610
611 F2FS adopts the checkpointing scheme to maintain file system consistency. At
612 mount time, F2FS first tries to find the last valid checkpoint data by scanning
613 CP area. In order to reduce the scanning time, F2FS uses only two copies of CP.
614 One of them always indicates the last valid data, which is called as shadow copy
615 mechanism. In addition to CP, NAT and SIT also adopt the shadow copy mechanism.
616
617 For file system consistency, each CP points to which NAT and SIT copies are
618 valid, as shown as below::
619
620 +--------+----------+---------+
621 | CP | SIT | NAT |
622 +--------+----------+---------+
623 . . . .
624 . . . .
625 . . . .
626 +-------+-------+--------+--------+--------+--------+
627 | CP #0 | CP #1 | SIT #0 | SIT #1 | NAT #0 | NAT #1 |
628 +-------+-------+--------+--------+--------+--------+
629 | ^ ^
630 | | |
631 `----------------------------------------'
632
633 Index Structure
634 ---------------
635
636 The key data structure to manage the data locations is a "node". Similar to
637 traditional file structures, F2FS has three types of node: inode, direct node,
638 indirect node. F2FS assigns 4KB to an inode block which contains 923 data block
639 indices, two direct node pointers, two indirect node pointers, and one double
640 indirect node pointer as described below. One direct node block contains 1018
641 data blocks, and one indirect node block contains also 1018 node blocks. Thus,
642 one inode block (i.e., a file) covers::
643
644 4KB * (923 + 2 * 1018 + 2 * 1018 * 1018 + 1018 * 1018 * 1018) := 3.94TB.
645
646 Inode block (4KB)
647 |- data (923)
648 |- direct node (2)
649 | `- data (1018)
650 |- indirect node (2)
651 | `- direct node (1018)
652 | `- data (1018)
653 `- double indirect node (1)
654 `- indirect node (1018)
655 `- direct node (1018)
656 `- data (1018)
657
658 Note that all the node blocks are mapped by NAT which means the location of
659 each node is translated by the NAT table. In the consideration of the wandering
660 tree problem, F2FS is able to cut off the propagation of node updates caused by
661 leaf data writes.
662
663 Directory Structure
664 -------------------
665
666 A directory entry occupies 11 bytes, which consists of the following attributes.
667
668 - hash hash value of the file name
669 - ino inode number
670 - len the length of file name
671 - type file type such as directory, symlink, etc
672
673 A dentry block consists of 214 dentry slots and file names. Therein a bitmap is
674 used to represent whether each dentry is valid or not. A dentry block occupies
675 4KB with the following composition.
676
677 ::
678
679 Dentry Block(4 K) = bitmap (27 bytes) + reserved (3 bytes) +
680 dentries(11 * 214 bytes) + file name (8 * 214 bytes)
681
682 [Bucket]
683 +--------------------------------+
684 |dentry block 1 | dentry block 2 |
685 +--------------------------------+
686 . .
687 . .
688 . [Dentry Block Structure: 4KB] .
689 +--------+----------+----------+------------+
690 | bitmap | reserved | dentries | file names |
691 +--------+----------+----------+------------+
692 [Dentry Block: 4KB] . .
693 . .
694 . .
695 +------+------+-----+------+
696 | hash | ino | len | type |
697 +------+------+-----+------+
698 [Dentry Structure: 11 bytes]
699
700 F2FS implements multi-level hash tables for directory structure. Each level has
701 a hash table with dedicated number of hash buckets as shown below. Note that
702 "A(2B)" means a bucket includes 2 data blocks.
703
704 ::
705
706 ----------------------
707 A : bucket
708 B : block
709 N : MAX_DIR_HASH_DEPTH
710 ----------------------
711
712 level #0 | A(2B)
713 |
714 level #1 | A(2B) - A(2B)
715 |
716 level #2 | A(2B) - A(2B) - A(2B) - A(2B)
717 . | . . . .
718 level #N/2 | A(2B) - A(2B) - A(2B) - A(2B) - A(2B) - ... - A(2B)
719 . | . . . .
720 level #N | A(4B) - A(4B) - A(4B) - A(4B) - A(4B) - ... - A(4B)
721
722 The number of blocks and buckets are determined by::
723
724 ,- 2, if n < MAX_DIR_HASH_DEPTH / 2,
725 # of blocks in level #n = |
726 `- 4, Otherwise
727
728 ,- 2^(n + dir_level),
729 | if n + dir_level < MAX_DIR_HASH_DEPTH / 2,
730 # of buckets in level #n = |
731 `- 2^((MAX_DIR_HASH_DEPTH / 2) - 1),
732 Otherwise
733
734 When F2FS finds a file name in a directory, at first a hash value of the file
735 name is calculated. Then, F2FS scans the hash table in level #0 to find the
736 dentry consisting of the file name and its inode number. If not found, F2FS
737 scans the next hash table in level #1. In this way, F2FS scans hash tables in
738 each levels incrementally from 1 to N. In each level F2FS needs to scan only
739 one bucket determined by the following equation, which shows O(log(# of files))
740 complexity::
741
742 bucket number to scan in level #n = (hash value) % (# of buckets in level #n)
743
744 In the case of file creation, F2FS finds empty consecutive slots that cover the
745 file name. F2FS searches the empty slots in the hash tables of whole levels from
746 1 to N in the same way as the lookup operation.
747
748 The following figure shows an example of two cases holding children::
749
750 --------------> Dir <--------------
751 | |
752 child child
753
754 child - child [hole] - child
755
756 child - child - child [hole] - [hole] - child
757
758 Case 1: Case 2:
759 Number of children = 6, Number of children = 3,
760 File size = 7 File size = 7
761
762 Default Block Allocation
763 ------------------------
764
765 At runtime, F2FS manages six active logs inside "Main" area: Hot/Warm/Cold node
766 and Hot/Warm/Cold data.
767
768 - Hot node contains direct node blocks of directories.
769 - Warm node contains direct node blocks except hot node blocks.
770 - Cold node contains indirect node blocks
771 - Hot data contains dentry blocks
772 - Warm data contains data blocks except hot and cold data blocks
773 - Cold data contains multimedia data or migrated data blocks
774
775 LFS has two schemes for free space management: threaded log and copy-and-compac-
776 tion. The copy-and-compaction scheme which is known as cleaning, is well-suited
777 for devices showing very good sequential write performance, since free segments
778 are served all the time for writing new data. However, it suffers from cleaning
779 overhead under high utilization. Contrarily, the threaded log scheme suffers
780 from random writes, but no cleaning process is needed. F2FS adopts a hybrid
781 scheme where the copy-and-compaction scheme is adopted by default, but the
782 policy is dynamically changed to the threaded log scheme according to the file
783 system status.
784
785 In order to align F2FS with underlying flash-based storage, F2FS allocates a
786 segment in a unit of section. F2FS expects that the section size would be the
787 same as the unit size of garbage collection in FTL. Furthermore, with respect
788 to the mapping granularity in FTL, F2FS allocates each section of the active
789 logs from different zones as much as possible, since FTL can write the data in
790 the active logs into one allocation unit according to its mapping granularity.
791
792 Cleaning process
793 ----------------
794
795 F2FS does cleaning both on demand and in the background. On-demand cleaning is
796 triggered when there are not enough free segments to serve VFS calls. Background
797 cleaner is operated by a kernel thread, and triggers the cleaning job when the
798 system is idle.
799
800 F2FS supports two victim selection policies: greedy and cost-benefit algorithms.
801 In the greedy algorithm, F2FS selects a victim segment having the smallest number
802 of valid blocks. In the cost-benefit algorithm, F2FS selects a victim segment
803 according to the segment age and the number of valid blocks in order to address
804 log block thrashing problem in the greedy algorithm. F2FS adopts the greedy
805 algorithm for on-demand cleaner, while background cleaner adopts cost-benefit
806 algorithm.
807
808 In order to identify whether the data in the victim segment are valid or not,
809 F2FS manages a bitmap. Each bit represents the validity of a block, and the
810 bitmap is composed of a bit stream covering whole blocks in main area.
811
812 Write-hint Policy
813 -----------------
814
815 F2FS sets the whint all the time with the below policy.
816
817 ===================== ======================== ===================
818 User F2FS Block
819 ===================== ======================== ===================
820 N/A META WRITE_LIFE_NONE|REQ_META
821 N/A HOT_NODE WRITE_LIFE_NONE
822 N/A WARM_NODE WRITE_LIFE_MEDIUM
823 N/A COLD_NODE WRITE_LIFE_LONG
824 ioctl(COLD) COLD_DATA WRITE_LIFE_EXTREME
825 extension list " "
826
827 -- buffered io
828 ------------------------------------------------------------------
829 N/A COLD_DATA WRITE_LIFE_EXTREME
830 N/A HOT_DATA WRITE_LIFE_SHORT
831 N/A WARM_DATA WRITE_LIFE_NOT_SET
832
833 -- direct io
834 ------------------------------------------------------------------
835 WRITE_LIFE_EXTREME COLD_DATA WRITE_LIFE_EXTREME
836 WRITE_LIFE_SHORT HOT_DATA WRITE_LIFE_SHORT
837 WRITE_LIFE_NOT_SET WARM_DATA WRITE_LIFE_NOT_SET
838 WRITE_LIFE_NONE " WRITE_LIFE_NONE
839 WRITE_LIFE_MEDIUM " WRITE_LIFE_MEDIUM
840 WRITE_LIFE_LONG " WRITE_LIFE_LONG
841 ===================== ======================== ===================
842
843 Fallocate(2) Policy
844 -------------------
845
846 The default policy follows the below POSIX rule.
847
848 Allocating disk space
849 The default operation (i.e., mode is zero) of fallocate() allocates
850 the disk space within the range specified by offset and len. The
851 file size (as reported by stat(2)) will be changed if offset+len is
852 greater than the file size. Any subregion within the range specified
853 by offset and len that did not contain data before the call will be
854 initialized to zero. This default behavior closely resembles the
855 behavior of the posix_fallocate(3) library function, and is intended
856 as a method of optimally implementing that function.
857
858 However, once F2FS receives ioctl(fd, F2FS_IOC_SET_PIN_FILE) in prior to
859 fallocate(fd, DEFAULT_MODE), it allocates on-disk block addresses having
860 zero or random data, which is useful to the below scenario where:
861
862 1. create(fd)
863 2. ioctl(fd, F2FS_IOC_SET_PIN_FILE)
864 3. fallocate(fd, 0, 0, size)
865 4. address = fibmap(fd, offset)
866 5. open(blkdev)
867 6. write(blkdev, address)
868
869 Compression implementation
870 --------------------------
871
872 - New term named cluster is defined as basic unit of compression, file can
873 be divided into multiple clusters logically. One cluster includes 4 << n
874 (n >= 0) logical pages, compression size is also cluster size, each of
875 cluster can be compressed or not.
876
877 - In cluster metadata layout, one special block address is used to indicate
878 a cluster is a compressed one or normal one; for compressed cluster, following
879 metadata maps cluster to [1, 4 << n - 1] physical blocks, in where f2fs
880 stores data including compress header and compressed data.
881
882 - In order to eliminate write amplification during overwrite, F2FS only
883 support compression on write-once file, data can be compressed only when
884 all logical blocks in cluster contain valid data and compress ratio of
885 cluster data is lower than specified threshold.
886
887 - To enable compression on regular inode, there are four ways:
888
889 * chattr +c file
890 * chattr +c dir; touch dir/file
891 * mount w/ -o compress_extension=ext; touch file.ext
892 * mount w/ -o compress_extension=*; touch any_file
893
894 - To disable compression on regular inode, there are two ways:
895
896 * chattr -c file
897 * mount w/ -o nocompress_extension=ext; touch file.ext
898
899 - Priority in between FS_COMPR_FL, FS_NOCOMP_FS, extensions:
900
901 * compress_extension=so; nocompress_extension=zip; chattr +c dir; touch
902 dir/foo.so; touch dir/bar.zip; touch dir/baz.txt; then foo.so and baz.txt
903 should be compresse, bar.zip should be non-compressed. chattr +c dir/bar.zip
904 can enable compress on bar.zip.
905 * compress_extension=so; nocompress_extension=zip; chattr -c dir; touch
906 dir/foo.so; touch dir/bar.zip; touch dir/baz.txt; then foo.so should be
907 compresse, bar.zip and baz.txt should be non-compressed.
908 chattr+c dir/bar.zip; chattr+c dir/baz.txt; can enable compress on bar.zip
909 and baz.txt.
910
911 - At this point, compression feature doesn't expose compressed space to user
912 directly in order to guarantee potential data updates later to the space.
913 Instead, the main goal is to reduce data writes to flash disk as much as
914 possible, resulting in extending disk life time as well as relaxing IO
915 congestion. Alternatively, we've added ioctl(F2FS_IOC_RELEASE_COMPRESS_BLOCKS)
916 interface to reclaim compressed space and show it to user after setting a
917 special flag to the inode. Once the compressed space is released, the flag
918 will block writing data to the file until either the compressed space is
919 reserved via ioctl(F2FS_IOC_RESERVE_COMPRESS_BLOCKS) or the file size is
920 truncated to zero.
921
922 Compress metadata layout::
923
924 [Dnode Structure]
925 +-----------------------------------------------+
926 | cluster 1 | cluster 2 | ......... | cluster N |
927 +-----------------------------------------------+
928 . . . .
929 . . . .
930 . Compressed Cluster . . Normal Cluster .
931 +----------+---------+---------+---------+ +---------+---------+---------+---------+
932 |compr flag| block 1 | block 2 | block 3 | | block 1 | block 2 | block 3 | block 4 |
933 +----------+---------+---------+---------+ +---------+---------+---------+---------+
934 . .
935 . .
936 . .
937 +-------------+-------------+----------+----------------------------+
938 | data length | data chksum | reserved | compressed data |
939 +-------------+-------------+----------+----------------------------+
940
941 Compression mode
942 --------------------------
943
944 f2fs supports "fs" and "user" compression modes with "compression_mode" mount option.
945 With this option, f2fs provides a choice to select the way how to compress the
946 compression enabled files (refer to "Compression implementation" section for how to
947 enable compression on a regular inode).
948
949 1) compress_mode=fs
950
951 This is the default option. f2fs does automatic compression in the writeback of the
952 compression enabled files.
953
954 2) compress_mode=user
955
956 This disables the automatic compression and gives the user discretion of choosing the
957 target file and the timing. The user can do manual compression/decompression on the
958 compression enabled files using F2FS_IOC_DECOMPRESS_FILE and F2FS_IOC_COMPRESS_FILE
959 ioctls like the below.
960
961 To decompress a file::
962
963 fd = open(filename, O_WRONLY, 0);
964 ret = ioctl(fd, F2FS_IOC_DECOMPRESS_FILE);
965
966 To compress a file::
967
968 fd = open(filename, O_WRONLY, 0);
969 ret = ioctl(fd, F2FS_IOC_COMPRESS_FILE);
970
971 NVMe Zoned Namespace devices
972 ----------------------------
973
974 - ZNS defines a per-zone capacity which can be equal or less than the
975 zone-size. Zone-capacity is the number of usable blocks in the zone.
976 F2FS checks if zone-capacity is less than zone-size, if it is, then any
977 segment which starts after the zone-capacity is marked as not-free in
978 the free segment bitmap at initial mount time. These segments are marked
979 as permanently used so they are not allocated for writes and
980 consequently are not needed to be garbage collected. In case the
981 zone-capacity is not aligned to default segment size(2MB), then a segment
982 can start before the zone-capacity and span across zone-capacity boundary.
983 Such spanning segments are also considered as usable segments. All blocks
984 past the zone-capacity are considered unusable in these segments.
985
986 Device aliasing feature
987 -----------------------
988
989 f2fs can utilize a special file called a "device aliasing file." This file allows
990 the entire storage device to be mapped with a single, large extent, not using
991 the usual f2fs node structures. This mapped area is pinned and primarily intended
992 for holding the space.
993
994 Essentially, this mechanism allows a portion of the f2fs area to be temporarily
995 reserved and used by another filesystem or for different purposes. Once that
996 external usage is complete, the device aliasing file can be deleted, releasing
997 the reserved space back to F2FS for its own use.
998
999 .. code-block::
1001 # ls /dev/vd*
1002 /dev/vdb (32GB) /dev/vdc (32GB)
1003 # mkfs.ext4 /dev/vdc
1004 # mkfs.f2fs -c /dev/[email protected] /dev/vdb
1005 # mount /dev/vdb /mnt/f2fs
1006 # ls -l /mnt/f2fs
1007 vdc.file
1008 # df -h
1009 /dev/vdb 64G 33G 32G 52% /mnt/f2fs
1011 # mount -o loop /dev/vdc /mnt/ext4
1012 # df -h
1013 /dev/vdb 64G 33G 32G 52% /mnt/f2fs
1014 /dev/loop7 32G 24K 30G 1% /mnt/ext4
1015 # umount /mnt/ext4
1017 # f2fs_io getflags /mnt/f2fs/vdc.file
1018 get a flag on /mnt/f2fs/vdc.file ret=0, flags=nocow(pinned),immutable
1019 # f2fs_io setflags noimmutable /mnt/f2fs/vdc.file
1020 get a flag on noimmutable ret=0, flags=800010
1021 set a flag on /mnt/f2fs/vdc.file ret=0, flags=noimmutable
1022 # rm /mnt/f2fs/vdc.file
1023 # df -h
1024 /dev/vdb 64G 753M 64G 2% /mnt/f2fs
1026 So, the key idea is, user can do any file operations on /dev/vdc, and
1027 reclaim the space after the use, while the space is counted as /data.
1028 That doesn't require modifying partition size and filesystem format.

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

개요·LFS 문제와 F2FS 핵심 기능

1-106

SSD, eMMC, SD card 같은 NAND flash storage는 rotating disk와 특성이 다르므로 상위 layer인 filesystem 설계도 이에 맞춰야 합니다. F2FS는 Log-structured File System(LFS)을 기반으로 하되 wandering tree의 연쇄 갱신과 높은 cleaning overhead를 완화하도록 설계됐습니다.

device 내부 geometry와 Flash Translation Layer(FTL) 관리 방식이 서로 다르기 때문에 F2FS와 userland tool은 on-disk layout, allocation algorithm, cleaning algorithm을 조정하는 다양한 parameter를 제공합니다.

`f2fs-tools` git tree는 `mkfs.f2fs`, `fsck.f2fs`, `dump.f2fs`를 제공합니다. patch는 `[email protected]`, bug는 kernel Bugzilla의 f2fs component로 보냅니다.

LFS는 모든 변경을 log 형태로 순차 기록해 file write와 crash recovery를 빠르게 합니다. 빠른 새 write를 위한 큰 free area를 유지하려고 log를 segment로 나누고 cleaner가 심하게 fragmented된 segment의 live data를 압축 이동합니다.

file data를 log 끝으로 옮기면 direct pointer, indirect pointer, inode, inode map, checkpoint가 재귀적으로 갱신되는 현상을 wandering tree problem이라 합니다. F2FS는 inode와 pointer block을 모두 `node`로 다루고 모든 node 위치를 NAT에 기록해 leaf data write의 update propagation을 끊습니다.

out-of-place write는 obsolete block을 넓게 흩뿌립니다. cleaner는 SIT로 victim segment를 고르고 SSA의 summary로 parent index를 읽은 뒤 data-parent cross-reference를 검사해 valid data만 이동합니다.

F2FS cleaning 기본 절차
SIT를 참조해 victim segment 선택SSA summary로 victim data의 parent index 구조 loaddata와 parent index의 cross-reference 검증valid data만 새 위치로 이동

LFS가 새 log 공간을 회수하는 네 단계입니다.

F2FS는 random write area를 넓히면서 spatial locality를 유지하고 FTL operation unit에 data structure를 정렬합니다. background cleaning, greedy·cost-benefit victim policy, hot/warm/cold multi-head log, adaptive logging을 지원합니다.

.. SPDX-License-Identifier: GPL-2.0

=================================
Flash-Friendly File System (F2FS)
=================================

Overview
========

NAND flash memory-based storage devices, such as SSD, eMMC, and SD cards, have
been equipped on a variety systems ranging from mobile to server systems. Since
they are known to have different characteristics from the conventional rotating
disks, a file system, an upper layer to the storage device, should adapt to the
changes from the sketch in the design level.

F2FS is a file system exploiting NAND flash memory-based storage devices, which
is based on Log-structured File System (LFS). The design has been focused on
addressing the fundamental issues in LFS, which are snowball effect of wandering
tree and high cleaning overhead.

Since a NAND flash memory-based storage device shows different characteristic
according to its internal geometry or flash memory management scheme, namely FTL,
F2FS and its tools support various parameters not only for configuring on-disk
layout, but also for selecting allocation and cleaning algorithms.

The following git tree provides the file system formatting tool (mkfs.f2fs),
a consistency checking tool (fsck.f2fs), and a debugging tool (dump.f2fs).

- git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs-tools.git

For sending patches, please use the following mailing list:

- [email protected]

For reporting bugs, please use the following f2fs bug tracker link:

- https://bugzilla.kernel.org/enter_bug.cgi?product=File%20System&component=f2fs

Background and Design issues
============================

Log-structured File System (LFS)
--------------------------------
"A log-structured file system writes all modifications to disk sequentially in
a log-like structure, thereby speeding up  both file writing and crash recovery.
The log is the only structure on disk; it contains indexing information so that
files can be read back from the log efficiently. In order to maintain large free
areas on disk for fast writing, we divide  the log into segments and use a
segment cleaner to compress the live information from heavily fragmented
segments." from Rosenblum, M. and Ousterhout, J. K., 1992, "The design and
implementation of a log-structured file system", ACM Trans. Computer Systems
10, 1, 26–52.

Wandering Tree Problem
----------------------
In LFS, when a file data is updated and written to the end of log, its direct
pointer block is updated due to the changed location. Then the indirect pointer
block is also updated due to the direct pointer block update. In this manner,
the upper index structures such as inode, inode map, and checkpoint block are
also updated recursively. This problem is called as wandering tree problem [1],
and in order to enhance the performance, it should eliminate or relax the update
propagation as much as possible.

[1] Bityutskiy, A. 2005. JFFS3 design issues. http://www.linux-mtd.infradead.org/

Cleaning Overhead
-----------------
Since LFS is based on out-of-place writes, it produces so many obsolete blocks
scattered across the whole storage. In order to serve new empty log space, it
needs to reclaim these obsolete blocks seamlessly to users. This job is called
as a cleaning process.

The process consists of three operations as follows.

1. A victim segment is selected through referencing segment usage table.
2. It loads parent index structures of all the data in the victim identified by
   segment summary blocks.
3. It checks the cross-reference between the data and its parent index structure.
4. It moves valid data selectively.

This cleaning job may cause unexpected long delays, so the most important goal
is to hide the latencies to users. And also definitely, it should reduce the
amount of valid data to be moved, and move them quickly as well.

Key Features
============

Flash Awareness
---------------
- Enlarge the random write area for better performance, but provide the high
  spatial locality
- Align FS data structures to the operational units in FTL as best efforts

Wandering Tree Problem
----------------------
- Use a term, “node”, that represents inodes as well as various pointer blocks
- Introduce Node Address Table (NAT) containing the locations of all the “node”
  blocks; this will cut off the update propagation.

Cleaning Overhead
-----------------
- Support a background cleaning process
- Support greedy and cost-benefit algorithms for victim selection policies
- Support multi-head logs for static/dynamic hot and cold data separation
- Introduce adaptive logging for efficient block allocation

Mount option: GC·inline·allocation·fault

107-239

첫 mount option 집합은 background/foreground GC 협력, roll-forward recovery, discard, xattr·ACL, active log, inline storage, flush ordering, extent cache, reserved 공간, fault injection과 allocation mode를 제어합니다.

F2FS mount option 전반부
Mount option동작
`background_gc=on|off|sync`idle I/O 중 background garbage collection을 켜거나 끕니다. `sync`는 background에서 synchronous GC를 수행합니다. 기본값은 `on`입니다.
`gc_merge``background_gc=on`일 때 background GC thread가 foreground GC request도 처리해, I/O·CPU resource가 제한된 process의 느린 foreground GC를 줄입니다.
`nogc_merge`GC merge 기능을 끕니다.
`disable_roll_forward`roll-forward recovery routine을 끕니다.
`norecovery`roll-forward recovery를 끄고 read-only로 mount합니다. 즉 `-o ro,disable_roll_forward`입니다.
`discard` / `nodiscard`real-time discard를 켜거나 끕니다. `discard`이면 segment 정리 시 discard/TRIM command를 발행합니다.
`heap` / `no_heap`deprecated option입니다.
`nouser_xattr`extended user attribute를 끕니다. `CONFIG_F2FS_FS_XATTR` 선택 시 xattr은 기본 활성화됩니다.
`noacl`POSIX ACL을 끕니다. `CONFIG_F2FS_FS_POSIX_ACL` 선택 시 ACL은 기본 활성화됩니다.
`active_logs=%u`active log 수를 지정합니다. 현재 2, 4, 6만 지원하며 기본값은 6입니다.
`disable_ext_identify``mkfs`가 구성한 extension list를 무시해 media file 같은 cold file을 식별하지 않습니다.
`inline_xattr` / `noinline_xattr`inline xattr 기능을 켜거나 끕니다.
`inline_xattr_size=%u`flexible inline xattr feature에 따라 inline xattr 크기를 지정합니다.
`inline_data`새로 만든 작은 file(약 3.4KiB 미만)의 data를 inode block에 씁니다.
`inline_dentry`새 directory의 dentry data를 inode block 안 약 3.4KiB 공간에 저장합니다.
`noinline_dentry`inline dentry 기능을 끕니다.
`flush_merge`동시 cache-flush command를 최대한 합쳐 중복 command를 줄입니다. device의 cache flush가 느릴 때 권장됩니다.
`nobarrier`storage가 cache data를 nonvolatile 영역에 기록함을 보장할 때 cache-flush command를 생략합니다. F2FS는 data write 순서는 계속 보장합니다.
`barrier`cache-flush command 발행을 허용합니다.
`fastboot`일반 성능을 일부 희생해 mount 시간을 최대한 줄입니다.
`extent_cache`inode별 logical-to-physical contiguous extent를 rb-tree에 많이 cache해 hit ratio를 높입니다. 기본 활성화입니다.
`noextent_cache`rb-tree 기반 extent cache를 명시적으로 끕니다.
`noinline_data`기본 활성화되는 inline data 기능을 끕니다.
`data_flush`checkpoint 전에 regular file과 symlink data를 flush해 영속화합니다.
`reserve_root=%d`지정 uid/gid의 privileged user가 사용할 reserved 공간을 4KiB 단위로 설정합니다. 기본 한계는 user block의 12.5%입니다.
`reserve_node=%d`지정 uid/gid의 privileged user가 사용할 reserved node 수를 설정합니다. 기본 한계는 전체 node의 12.5%입니다.
`resuid=%d` / `resgid=%d`reserved block과 node를 사용할 수 있는 user ID와 group ID를 지정합니다.
`fault_injection=%d`지정한 injection rate로 지원되는 모든 fault type의 fault injection을 켭니다.
`fault_type=%d``fault_injection`과 함께 단일 또는 조합 fault type을 지정합니다.
`mode=adaptive|lfs`block allocation mode를 정합니다. `lfs`에서는 Main area에 random write가 없어야 합니다.
`mode=fragment:segment`실험용으로 새 segment를 무작위 위치에 할당해 after-GC 상태를 모사합니다.
`mode=fragment:block``max_fragment_chunk`와 `max_fragment_hole` 범위에 randomness를 넣어 partition 전체에 block을 흩뿌립니다. 더 큰 randomness를 위해 `fragment:segment`도 암묵적으로 켜며, 실험 뒤 재format을 강하게 권장합니다.

원문의 option name, argument와 기본 동작을 보존합니다.

Fault injection type
Type nameType value
`FAULT_KMALLOC``0x00000001`
`FAULT_KVMALLOC``0x00000002`
`FAULT_PAGE_ALLOC``0x00000004`
`FAULT_PAGE_GET``0x00000008`
`FAULT_ALLOC_BIO``0x00000010` (obsolete)
`FAULT_ALLOC_NID``0x00000020`
`FAULT_ORPHAN``0x00000040`
`FAULT_BLOCK``0x00000080`
`FAULT_DIR_DEPTH``0x00000100`
`FAULT_EVICT_INODE``0x00000200`
`FAULT_TRUNCATE``0x00000400`
`FAULT_READ_IO``0x00000800`
`FAULT_CHECKPOINT``0x00001000`
`FAULT_DISCARD``0x00002000`
`FAULT_WRITE_IO``0x00004000`
`FAULT_SLAB_ALLOC``0x00008000`
`FAULT_DQUOT_INIT``0x00010000`
`FAULT_LOCK_OP``0x00020000`
`FAULT_BLKADDR_VALIDITY``0x00040000`
`FAULT_BLKADDR_CONSISTENCE``0x00080000`
`FAULT_NO_SEGMENT``0x00100000`
`FAULT_INCONSISTENT_FOOTER``0x00200000`
`FAULT_TIMEOUT``0x00400000` (1000ms)
`FAULT_VMALLOC``0x00800000`

`fault_type`은 다음 bit를 단독 또는 조합해 사용합니다.

Mount Options
=============


======================== ============================================================
background_gc=%s         Turn on/off cleaning operations, namely garbage
                         collection, triggered in background when I/O subsystem is
                         idle. If background_gc=on, it will turn on the garbage
                         collection and if background_gc=off, garbage collection
                         will be turned off. If background_gc=sync, it will turn
                         on synchronous garbage collection running in background.
                         Default value for this option is on. So garbage
                         collection is on by default.
gc_merge                 When background_gc is on, this option can be enabled to
                         let background GC thread to handle foreground GC requests,
                         it can eliminate the sluggish issue caused by slow foreground
                         GC operation when GC is triggered from a process with limited
                         I/O and CPU resources.
nogc_merge                 Disable GC merge feature.
disable_roll_forward         Disable the roll-forward recovery routine
norecovery                 Disable the roll-forward recovery routine, mounted read-
                         only (i.e., -o ro,disable_roll_forward)
discard/nodiscard         Enable/disable real-time discard in f2fs, if discard is
                         enabled, f2fs will issue discard/TRIM commands when a
                         segment is cleaned.
heap/no_heap                 Deprecated.
nouser_xattr                 Disable Extended User Attributes. Note: xattr is enabled
                         by default if CONFIG_F2FS_FS_XATTR is selected.
noacl                         Disable POSIX Access Control List. Note: acl is enabled
                         by default if CONFIG_F2FS_FS_POSIX_ACL is selected.
active_logs=%u                 Support configuring the number of active logs. In the
                         current design, f2fs supports only 2, 4, and 6 logs.
                         Default number is 6.
disable_ext_identify         Disable the extension list configured by mkfs, so f2fs
                         is not aware of cold files such as media files.
inline_xattr                 Enable the inline xattrs feature.
noinline_xattr                 Disable the inline xattrs feature.
inline_xattr_size=%u         Support configuring inline xattr size, it depends on
                         flexible inline xattr feature.
inline_data                 Enable the inline data feature: Newly created small (<~3.4k)
                         files can be written into inode block.
inline_dentry                 Enable the inline dir feature: data in newly created
                         directory entries can be written into inode block. The
                         space of inode block which is used to store inline
                         dentries is limited to ~3.4k.
noinline_dentry                 Disable the inline dentry feature.
flush_merge                 Merge concurrent cache_flush commands as much as possible
                         to eliminate redundant command issues. If the underlying
                         device handles the cache_flush command relatively slowly,
                         recommend to enable this option.
nobarrier                 This option can be used if underlying storage guarantees
                         its cached data should be written to the novolatile area.
                         If this option is set, no cache_flush commands are issued
                         but f2fs still guarantees the write ordering of all the
                         data writes.
barrier                         If this option is set, cache_flush commands are allowed to be
                         issued.
fastboot                 This option is used when a system wants to reduce mount
                         time as much as possible, even though normal performance
                         can be sacrificed.
extent_cache                 Enable an extent cache based on rb-tree, it can cache
                         as many as extent which map between contiguous logical
                         address and physical address per inode, resulting in
                         increasing the cache hit ratio. Set by default.
noextent_cache                 Disable an extent cache based on rb-tree explicitly, see
                         the above extent_cache mount option.
noinline_data                 Disable the inline data feature, inline data feature is
                         enabled by default.
data_flush                 Enable data flushing before checkpoint in order to
                         persist data of regular and symlink.
reserve_root=%d                 Support configuring reserved space which is used for
                         allocation from a privileged user with specified uid or
                         gid, unit: 4KB, the default limit is 12.5% of user blocks.
reserve_node=%d                 Support configuring reserved nodes which are used for
                         allocation from a privileged user with specified uid or
                         gid, the default limit is 12.5% of all nodes.
resuid=%d                 The user ID which may use the reserved blocks and nodes.
resgid=%d                 The group ID which may use the reserved blocks and nodes.
fault_injection=%d         Enable fault injection in all supported types with
                         specified injection rate.
fault_type=%d                 Support configuring fault injection type, should be
                         enabled with fault_injection option, fault type value
                         is shown below, it supports single or combined type.

                         ===========================      ==========
                         Type_Name                        Type_Value
                         ===========================      ==========
                         FAULT_KMALLOC                    0x00000001
                         FAULT_KVMALLOC                   0x00000002
                         FAULT_PAGE_ALLOC                 0x00000004
                         FAULT_PAGE_GET                   0x00000008
                         FAULT_ALLOC_BIO                  0x00000010 (obsolete)
                         FAULT_ALLOC_NID                  0x00000020
                         FAULT_ORPHAN                     0x00000040
                         FAULT_BLOCK                      0x00000080
                         FAULT_DIR_DEPTH                  0x00000100
                         FAULT_EVICT_INODE                0x00000200
                         FAULT_TRUNCATE                   0x00000400
                         FAULT_READ_IO                    0x00000800
                         FAULT_CHECKPOINT                 0x00001000
                         FAULT_DISCARD                    0x00002000
                         FAULT_WRITE_IO                   0x00004000
                         FAULT_SLAB_ALLOC                 0x00008000
                         FAULT_DQUOT_INIT                 0x00010000
                         FAULT_LOCK_OP                    0x00020000
                         FAULT_BLKADDR_VALIDITY           0x00040000
                         FAULT_BLKADDR_CONSISTENCE        0x00080000
                         FAULT_NO_SEGMENT                 0x00100000
                         FAULT_INCONSISTENT_FOOTER        0x00200000
                         FAULT_TIMEOUT                    0x00400000 (1000ms)
                         FAULT_VMALLOC                    0x00800000
                         ===========================      ==========
mode=%s                         Control block allocation mode which supports "adaptive"
                         and "lfs". In "lfs" mode, there should be no random
                         writes towards main area.
                         "fragment:segment" and "fragment:block" are newly added here.
                         These are developer options for experiments to simulate filesystem
                         fragmentation/after-GC situation itself. The developers use these
                         modes to understand filesystem fragmentation/after-GC condition well,
                         and eventually get some insights to handle them better.
                         In "fragment:segment", f2fs allocates a new segment in random
                         position. With this, we can simulate the after-GC condition.
                         In "fragment:block", we can scatter block allocation with
                         "max_fragment_chunk" and "max_fragment_hole" sysfs nodes.
                         We added some randomness to both chunk and hole size to make
                         it close to realistic IO pattern. So, in this mode, f2fs will allocate
                         1..<max_fragment_chunk> blocks in a chunk and make a hole in the
                         length of 1..<max_fragment_hole> by turns. With this, the newly
                         allocated blocks will be scattered throughout the whole partition.
                         Note that "fragment:block" implicitly enables "fragment:segment"
                         option for more randomness.
                         Please, use these options for your experiments and we strongly
                         recommend to re-format the filesystem after using these options.

Mount option: quota·checkpoint·compression·error

240-404

후반 mount option은 quota, fsync, dummy encryption, checkpoint disable/merge, compression, inline crypto, zoned discard, memory 절약, critical error policy, NAT와 casefold lookup을 제어합니다.

compression enable 우선순위는 directory flag보다 `compress_extension`·`nocompress_extension`이 높고, 개별 file의 compress/no-compress flag가 가장 높습니다. `nocompress_extension`은 `compress_extension=*`의 예외로도 동작합니다.

F2FS mount option 후반부
Mount option동작
`usrquota` / `grpquota` / `prjquota`일반 user, group, project disk quota accounting을 각각 켭니다.
`usrjquota=<file>` / `grpjquota=<file>` / `prjjquota=<file>`recovery 중 quota 정보가 갱신되도록 mount 시 quota file을 지정합니다. file은 root directory에 있어야 합니다.
`jqfmt=vfsold|vfsv0|vfsv1`journalled quota format을 지정합니다.
`usrjquota=` / `grpjquota=` / `prjjquota=`각 journalled quota를 끕니다.
`quota` / `noquota`일반 user disk quota를 켜거나 모든 일반 disk quota option을 끕니다.
`alloc_mode=reuse|default`block allocation policy를 조정합니다.
`fsync_mode=posix`기본값. POSIX semantics를 따르며 성능을 위한 가벼운 fsync를 수행합니다.
`fsync_mode=strict`xfs, ext4, btrfs와 같은 무거운 fsync를 수행해 xfstest `generic/342`를 통과하지만 성능이 낮아집니다.
`fsync_mode=nobarrier``posix` 기반이지만 non-atomic file에 flush command를 발행하지 않습니다.
`test_dummy_encryption[=v1|v2]`xfstests용 가짜 fscrypt context를 켭니다. argument로 fscrypt policy version을 고릅니다.
`checkpoint=disable[:blocks|percent]`checkpointing을 끕니다. unmount나 비정상 종료 뒤 내용은 이 option으로 mount했던 시점처럼 보입니다. 모든 공간 사용을 위해 GC가 필요하며 너무 오래 걸리면 `EAGAIN`을 반환할 수 있습니다.
`checkpoint=enable`기본 활성화된 checkpointing을 다시 켜고 disable 동안 숨겨진 공간을 회수합니다. unusable 공간은 `/sys/fs/f2fs/<disk>/unusable`에서 확인합니다.
`checkpoint_merge`kernel daemon이 동시 checkpoint request를 합쳐 중복과 낮은 cgroup I/O·CPU budget에 따른 지연을 줄입니다. daemon 기본 I/O priority는 ext4 jbd2 thread와 같은 방식으로 3입니다.
`nocheckpoint_merge`checkpoint merge를 끕니다.
`compress_algorithm=lzo|lz4|zstd|lzo-rle`compression algorithm을 고릅니다.
`compress_algorithm=lz4:<3-16>`LZ4 compression level을 지정합니다.
`compress_algorithm=zstd:<1-22>`Zstd compression level을 지정합니다.
`compress_log_size=%u`compression cluster 크기를 `4KiB * (1 << value)`로 설정합니다. 기본·최소 크기는 16KiB입니다.
`compress_extension=%s`지정 extension의 file을 기본 압축 대상으로 추가합니다. `*`는 모든 file을 뜻하며 다른 file은 ioctl로 켤 수 있습니다.
`nocompress_extension=%s`지정 extension을 압축 제외 대상으로 추가합니다. 같은 extension은 두 list에 동시에 넣을 수 없고 `*`는 허용하지 않습니다.
`compress_chksum`compressed cluster의 raw data checksum 검증을 지원합니다.
`compress_mode=fs|user``fs`는 자동 압축하는 기본 mode, `user`는 ioctl을 통한 수동 압축·해제를 허용하는 mode입니다.
`compress_cache`filesystem 관리 inode의 address space에 compressed block을 cache해 random read hit ratio를 높입니다.
`inlinecrypt`가능하면 filesystem layer 대신 `blk-crypto`와 inline encryption hardware로 암복호화합니다. on-disk format은 바뀌지 않습니다.
`atgc`background GC의 효율을 높이는 age-threshold garbage collection을 켭니다.
`discard_unit=block|segment|section`discard offset·size 정렬 단위를 고릅니다. 기본은 `block`, zoned device 기본은 memory 비용을 줄이는 `section`입니다.
`memory=normal|low``normal`이 기본입니다. `low`는 저메모리 device에서 때때로 성능을 희생해 memory를 절약합니다.
`age_extent_cache`inode별 extent의 data block update 빈도를 rb-tree에 기록해 allocation temperature hint를 개선합니다.
`errors=continue|remount-ro|panic`critical error 처리 방식입니다. 기본 `continue`는 계속하며 syscall에 `-EIO`, `remount-ro`는 read-only로 전환하며 `-EROFS`, `panic`은 즉시 panic합니다.
`nat_bits`full/empty NAT block 접근을 개선하는 `nat_bits`를 켭니다. 기본은 꺼짐입니다.
`lookup_mode=perf`기본값. hash-only lookup을 강제하고 on-disk flag와 무관하게 linear fallback을 끕니다.
`lookup_mode=compat`예전 case-folding algorithm이 만든 dentry와의 호환을 위해 linear fallback을 켜며 on-disk flag를 무시합니다.
`lookup_mode=auto`on-disk `SB_ENC_NO_COMPAT_FALLBACK_FL` flag에 따라 F2FS가 mode를 결정합니다.

quota부터 casefold lookup mode까지의 전문 번역입니다.

Critical error mode 비교
항목`continue``remount-ro``panic`
access operationnormalnormalN/A
syscall error`-EIO``-EROFS`N/A
mount optionread-writeread-onlyN/A
pending directory writekeepkeepN/A
pending non-directory writedropkeepN/A
pending node writedropkeepN/A
pending metadata writekeepkeepN/A

`errors=` 선택에 따른 관찰 가능한 동작입니다.

usrquota                 Enable plain user disk quota accounting.
grpquota                 Enable plain group disk quota accounting.
prjquota                 Enable plain project quota accounting.
usrjquota=<file>         Appoint specified file and type during mount, so that quota
grpjquota=<file>         information can be properly updated during recovery flow,
prjjquota=<file>         <quota file>: must be in root directory;
jqfmt=<quota type>         <quota type>: [vfsold,vfsv0,vfsv1].
usrjquota=                 Turn off user journalled quota.
grpjquota=                 Turn off group journalled quota.
prjjquota=                 Turn off project journalled quota.
quota                         Enable plain user disk quota accounting.
noquota                         Disable all plain disk quota option.
alloc_mode=%s                 Adjust block allocation policy, which supports "reuse"
                         and "default".
fsync_mode=%s                 Control the policy of fsync. Currently supports "posix",
                         "strict", and "nobarrier". In "posix" mode, which is
                         default, fsync will follow POSIX semantics and does a
                         light operation to improve the filesystem performance.
                         In "strict" mode, fsync will be heavy and behaves in line
                         with xfs, ext4 and btrfs, where xfstest generic/342 will
                         pass, but the performance will regress. "nobarrier" is
                         based on "posix", but doesn't issue flush command for
                         non-atomic files likewise "nobarrier" mount option.
test_dummy_encryption
test_dummy_encryption=%s
                         Enable dummy encryption, which provides a fake fscrypt
                         context. The fake fscrypt context is used by xfstests.
                         The argument may be either "v1" or "v2", in order to
                         select the corresponding fscrypt policy version.
checkpoint=%s[:%u[%]]         Set to "disable" to turn off checkpointing. Set to "enable"
                         to re-enable checkpointing. Is enabled by default. While
                         disabled, any unmounting or unexpected shutdowns will cause
                         the filesystem contents to appear as they did when the
                         filesystem was mounted with that option.
                         While mounting with checkpoint=disable, the filesystem must
                         run garbage collection to ensure that all available space can
                         be used. If this takes too much time, the mount may return
                         EAGAIN. You may optionally add a value to indicate how much
                         of the disk you would be willing to temporarily give up to
                         avoid additional garbage collection. This can be given as a
                         number of blocks, or as a percent. For instance, mounting
                         with checkpoint=disable:100% would always succeed, but it may
                         hide up to all remaining free space. The actual space that
                         would be unusable can be viewed at /sys/fs/f2fs/<disk>/unusable
                         This space is reclaimed once checkpoint=enable.
checkpoint_merge         When checkpoint is enabled, this can be used to create a kernel
                         daemon and make it to merge concurrent checkpoint requests as
                         much as possible to eliminate redundant checkpoint issues. Plus,
                         we can eliminate the sluggish issue caused by slow checkpoint
                         operation when the checkpoint is done in a process context in
                         a cgroup having low i/o budget and cpu shares. To make this
                         do better, we set the default i/o priority of the kernel daemon
                         to "3", to give one higher priority than other kernel threads.
                         This is the same way to give a I/O priority to the jbd2
                         journaling thread of ext4 filesystem.
nocheckpoint_merge         Disable checkpoint merge feature.
compress_algorithm=%s         Control compress algorithm, currently f2fs supports "lzo",
                         "lz4", "zstd" and "lzo-rle" algorithm.
compress_algorithm=%s:%d Control compress algorithm and its compress level, now, only
                         "lz4" and "zstd" support compress level config.

                         =========      ===========
                         algorithm        level range
                         =========      ===========
                         lz4                3 - 16
                         zstd                1 - 22
                         =========      ===========
compress_log_size=%u         Support configuring compress cluster size. The size will
                         be 4KB * (1 << %u). The default and minimum sizes are 16KB.
compress_extension=%s         Support adding specified extension, so that f2fs can enable
                         compression on those corresponding files, e.g. if all files
                         with '.ext' has high compression rate, we can set the '.ext'
                         on compression extension list and enable compression on
                         these file by default rather than to enable it via ioctl.
                         For other files, we can still enable compression via ioctl.
                         Note that, there is one reserved special extension '*', it
                         can be set to enable compression for all files.
nocompress_extension=%s         Support adding specified extension, so that f2fs can disable
                         compression on those corresponding files, just contrary to compression extension.
                         If you know exactly which files cannot be compressed, you can use this.
                         The same extension name can't appear in both compress and nocompress
                         extension at the same time.
                         If the compress extension specifies all files, the types specified by the
                         nocompress extension will be treated as special cases and will not be compressed.
                         Don't allow use '*' to specifie all file in nocompress extension.
                         After add nocompress_extension, the priority should be:
                         dir_flag < comp_extention,nocompress_extension < comp_file_flag,no_comp_file_flag.
                         See more in compression sections.

compress_chksum                 Support verifying chksum of raw data in compressed cluster.
compress_mode=%s         Control file compression mode. This supports "fs" and "user"
                         modes. In "fs" mode (default), f2fs does automatic compression
                         on the compression enabled files. In "user" mode, f2fs disables
                         the automaic compression and gives the user discretion of
                         choosing the target file and the timing. The user can do manual
                         compression/decompression on the compression enabled files using
                         ioctls.
compress_cache                 Support to use address space of a filesystem managed inode to
                         cache compressed block, in order to improve cache hit ratio of
                         random read.
inlinecrypt                 When possible, encrypt/decrypt the contents of encrypted
                         files using the blk-crypto framework rather than
                         filesystem-layer encryption. This allows the use of
                         inline encryption hardware. The on-disk format is
                         unaffected. For more details, see
                         Documentation/block/inline-encryption.rst.
atgc                         Enable age-threshold garbage collection, it provides high
                         effectiveness and efficiency on background GC.
discard_unit=%s                 Control discard unit, the argument can be "block", "segment"
                         and "section", issued discard command's offset/size will be
                         aligned to the unit, by default, "discard_unit=block" is set,
                         so that small discard functionality is enabled.
                         For blkzoned device, "discard_unit=section" will be set by
                         default, it is helpful for large sized SMR or ZNS devices to
                         reduce memory cost by getting rid of fs metadata supports small
                         discard.
memory=%s                 Control memory mode. This supports "normal" and "low" modes.
                         "low" mode is introduced to support low memory devices.
                         Because of the nature of low memory devices, in this mode, f2fs
                         will try to save memory sometimes by sacrificing performance.
                         "normal" mode is the default mode and same as before.
age_extent_cache         Enable an age extent cache based on rb-tree. It records
                         data block update frequency of the extent per inode, in
                         order to provide better temperature hints for data block
                         allocation.
errors=%s                 Specify f2fs behavior on critical errors. This supports modes:
                         "panic", "continue" and "remount-ro", respectively, trigger
                         panic immediately, continue without doing anything, and remount
                         the partition in read-only mode. By default it uses "continue"
                         mode.

                         ====================== =============== =============== ========
                         mode                        continue        remount-ro        panic
                         ====================== =============== =============== ========
                         access ops                normal                normal                N/A
                         syscall errors                -EIO                -EROFS                N/A
                         mount option                rw                ro                N/A
                         pending dir write        keep                keep                N/A
                         pending non-dir write        drop                keep                N/A
                         pending node write        drop                keep                N/A
                         pending meta write        keep                keep                N/A
                         ====================== =============== =============== ========
nat_bits                 Enable nat_bits feature to enhance full/empty nat blocks access,
                         by default it's disabled.
lookup_mode=%s                 Control the directory lookup behavior for casefolded
                         directories. This option has no effect on directories
                         that do not have the casefold feature enabled.

                         ================== ========================================
                         Value                    Description
                         ================== ========================================
                         perf                    (Default) Enforces a hash-only lookup.
                                            The linear search fallback is always
                                            disabled, ignoring the on-disk flag.
                         compat                    Enables the linear search fallback for
                                            compatibility with directory entries
                                            created by older kernel that used a
                                            different case-folding algorithm.
                                            This mode ignores the on-disk flag.
                         auto                    F2FS determines the mode based on the
                                            on-disk `SB_ENC_NO_COMPAT_FALLBACK_FL`
                                            flag.
                         ================== ========================================
======================== ============================================================

debugfs·sysfs·사용법과 userland tool

405-542

`/sys/kernel/debug/f2fs/`는 mount된 모든 F2FS partition 정보를 보여 줍니다. `status`에는 현재 주요 filesystem 정보, 전체 segment의 평균 SIT 정보, F2FS memory footprint가 포함됩니다.

mount된 filesystem의 sysfs 정보는 device name별 `/sys/fs/f2fs/<devname>`에 있으며 자세한 file 목록은 `Documentation/ABI/testing/sysfs-fs-f2fs`를 참조합니다.

사용 순서는 userland tool build, 필요 시 `f2fs.ko` module 삽입, mount directory 생성, `mkfs.f2fs` format, `mount -t f2fs`입니다.

기본 F2FS 준비
`f2fs-tools` download·build필요 시 `insmod f2fs.ko``mkdir /mnt/f2fs``mkfs.f2fs -l label /dev/block_device``mount -t f2fs /dev/block_device /mnt/f2fs`

module 방식 설치에서 mount까지의 명령 흐름입니다.

F2FS userland tool
Tool역할
`mkfs.f2fs`partition을 F2FS로 format하고 기본 on-disk layout을 만듭니다.
`fsck.f2fs`filesystem metadata와 user data의 cross-reference 일관성을 검사합니다. 초기 version은 inconsistency를 고치지 않았습니다.
`dump.f2fs`지정 inode의 on-disk 정보를 보여 주고 SSA와 SIT를 `dump_ssa`, `dump_sit`에 dump합니다.
`sload.f2fs`기존 disk image에 file과 directory를 넣어 compiled file로 image를 만들 때 사용합니다.
`resize.f2fs`저장된 file과 directory를 보존하며 F2FS disk image 크기를 바꿉니다.
`defrag.f2fs`흩어진 data와 metadata를 defragment해 연속 free space와 write speed를 늘립니다.
`f2fs_io`일반 filesystem API와 F2FS 전용 API를 발행하는 QA test 도구입니다.

각 tool의 주된 역할입니다.

`mkfs.f2fs` quick option
Option설명
`-l [label]`최대 512 unicode name의 volume label.
`-a [0|1]`heap allocation을 위해 각 area 시작 위치를 분리합니다. 기본 1.
`-o [int]`volume size 대비 overprovision ratio(%). 기본 5.
`-s [int]`section당 segment 수. 기본 1.
`-z [int]`zone당 section 수. 기본 1.
`-e [str]`기본 extension list. 예: `mp3,gif,mov`.
`-t [0|1]`discard command 사용 여부. 기본 1은 discard 수행.

전체 option은 `mkfs.f2fs(8)` manpage를 참조합니다.

`fsck.f2fs` quick option은 `-d debug_level`(기본 0)입니다. `dump.f2fs`는 `-d`, hex inode용 `-i`, SIT segment 범위용 `-s`, SSA segment 범위용 `-a`를 지원하며 `0~-1`은 전체 범위입니다.

Debugfs Entries
===============

/sys/kernel/debug/f2fs/ contains information about all the partitions mounted as
f2fs. Each file shows the whole f2fs information.

/sys/kernel/debug/f2fs/status includes:

 - major file system information managed by f2fs currently
 - average SIT information about whole segments
 - current memory footprint consumed by f2fs.

Sysfs Entries
=============

Information about mounted f2fs file systems can be found in
/sys/fs/f2fs.  Each mounted filesystem will have a directory in
/sys/fs/f2fs based on its device name (i.e., /sys/fs/f2fs/sda).
The files in each per-device directory are shown in table below.

Files in /sys/fs/f2fs/<devname>
(see also Documentation/ABI/testing/sysfs-fs-f2fs)

Usage
=====

1. Download userland tools and compile them.

2. Skip, if f2fs was compiled statically inside kernel.
   Otherwise, insert the f2fs.ko module::

        # insmod f2fs.ko

3. Create a directory to use when mounting::

        # mkdir /mnt/f2fs

4. Format the block device, and then mount as f2fs::

        # mkfs.f2fs -l label /dev/block_device
        # mount -t f2fs /dev/block_device /mnt/f2fs

mkfs.f2fs
---------
The mkfs.f2fs is for the use of formatting a partition as the f2fs filesystem,
which builds a basic on-disk layout.

The quick options consist of:

===============    ===========================================================
``-l [label]``     Give a volume label, up to 512 unicode name.
``-a [0 or 1]``    Split start location of each area for heap-based allocation.

                   1 is set by default, which performs this.
``-o [int]``       Set overprovision ratio in percent over volume size.

                   5 is set by default.
``-s [int]``       Set the number of segments per section.

                   1 is set by default.
``-z [int]``       Set the number of sections per zone.

                   1 is set by default.
``-e [str]``       Set basic extension list. e.g. "mp3,gif,mov"
``-t [0 or 1]``    Disable discard command or not.

                   1 is set by default, which conducts discard.
===============    ===========================================================

Note: please refer to the manpage of mkfs.f2fs(8) to get full option list.

fsck.f2fs
---------
The fsck.f2fs is a tool to check the consistency of an f2fs-formatted
partition, which examines whether the filesystem metadata and user-made data
are cross-referenced correctly or not.
Note that, initial version of the tool does not fix any inconsistency.

The quick options consist of::

  -d debug level [default:0]

Note: please refer to the manpage of fsck.f2fs(8) to get full option list.

dump.f2fs
---------
The dump.f2fs shows the information of specific inode and dumps SSA and SIT to
file. Each file is dump_ssa and dump_sit.

The dump.f2fs is used to debug on-disk data structures of the f2fs filesystem.
It shows on-disk inode information recognized by a given inode number, and is
able to dump all the SSA and SIT entries into predefined files, ./dump_ssa and
./dump_sit respectively.

The options consist of::

  -d debug level [default:0]
  -i inode no (hex)
  -s [SIT dump segno from #1~#2 (decimal), for all 0~-1]
  -a [SSA dump segno from #1~#2 (decimal), for all 0~-1]

Examples::

    # dump.f2fs -i [ino] /dev/sdx
    # dump.f2fs -s 0~-1 /dev/sdx (SIT dump)
    # dump.f2fs -a 0~-1 /dev/sdx (SSA dump)

Note: please refer to the manpage of dump.f2fs(8) to get full option list.

sload.f2fs
----------
The sload.f2fs gives a way to insert files and directories in the existing disk
image. This tool is useful when building f2fs images given compiled files.

Note: please refer to the manpage of sload.f2fs(8) to get full option list.

resize.f2fs
-----------
The resize.f2fs lets a user resize the f2fs-formatted disk image, while preserving
all the files and directories stored in the image.

Note: please refer to the manpage of resize.f2fs(8) to get full option list.

defrag.f2fs
-----------
The defrag.f2fs can be used to defragment scattered written data as well as
filesystem metadata across the disk. This can improve the write speed by giving
more free consecutive space.

Note: please refer to the manpage of defrag.f2fs(8) to get full option list.

f2fs_io
-------
The f2fs_io is a simple tool to issue various filesystem APIs as well as
f2fs-specific ones, which is very useful for QA tests.

Note: please refer to the manpage of f2fs_io(8) to get full option list.

On-disk layout과 metadata shadow copy

543-632

F2FS volume은 고정 2MiB segment로 나뉘고, 연속 segment가 section을, section 집합이 zone을 이룹니다. 기본 section·zone 크기는 각각 segment 하나지만 `mkfs`로 바꿀 수 있습니다.

volume은 SB, CP, SIT, NAT, SSA, Main의 여섯 area로 나뉩니다. SB를 제외한 area는 여러 segment로 구성됩니다.

F2FS on-disk area
Area내용
Superblock (`SB`)partition 시작에 crash 대비 두 사본을 두며 기본 partition 정보와 F2FS 기본 parameter를 담습니다.
Checkpoint (`CP`)filesystem 정보, 유효 NAT/SIT set bitmap, orphan inode list, 현재 active segment summary를 담습니다.
Segment Information Table (`SIT`)segment별 valid block 수와 모든 block의 validity bitmap을 담습니다.
Node Address Table (`NAT`)Main area에 저장된 모든 node block의 block address table입니다.
Segment Summary Area (`SSA`)Main area의 모든 data·node block owner 정보를 담는 summary entry를 저장합니다.
Main areaindex를 포함한 file과 directory data를 저장합니다.

원문의 ASCII volume layout을 같은 순서의 구조화 표로 재구성했습니다.

Volume 정렬
partition 시작: 두 Superblock 사본CP 시작 block을 segment size에 정렬SIT와 NAT 배치SSA에서 필요한 segment를 예약Main area 시작 block을 zone size에 정렬

filesystem과 flash storage의 operation unit이 어긋나지 않게 하는 배치입니다.

mount 시 CP area를 scan해 마지막 valid checkpoint를 찾습니다. scan 시간을 줄이려고 CP 사본은 두 개만 두고 하나가 항상 마지막 valid data를 가리키는 shadow copy mechanism을 사용합니다. NAT와 SIT도 각각 두 사본으로 같은 방식을 씁니다.

CP·SIT·NAT shadow copy
Checkpoint선택하는 SIT선택하는 NAT
`CP #0` 또는 `CP #1` 중 최신 valid copy`SIT #0` 또는 `SIT #1``NAT #0` 또는 `NAT #1`

각 checkpoint가 현재 유효한 SIT와 NAT 사본을 선택합니다.

Design
======

On-disk Layout
--------------

F2FS divides the whole volume into a number of segments, each of which is fixed
to 2MB in size. A section is composed of consecutive segments, and a zone
consists of a set of sections. By default, section and zone sizes are set to one
segment size identically, but users can easily modify the sizes by mkfs.

F2FS splits the entire volume into six areas, and all the areas except superblock
consist of multiple segments as described below::

                                            align with the zone size <-|
                 |-> align with the segment size
     _________________________________________________________________________
    |            |            |   Segment   |    Node     |   Segment  |      |
    | Superblock | Checkpoint |    Info.    |   Address   |   Summary  | Main |
    |    (SB)    |   (CP)     | Table (SIT) | Table (NAT) | Area (SSA) |      |
    |____________|_____2______|______N______|______N______|______N_____|__N___|
                                                                       .      .
                                                             .                .
                                                 .                            .
                                    ._________________________________________.
                                    |_Segment_|_..._|_Segment_|_..._|_Segment_|
                                    .           .
                                    ._________._________
                                    |_section_|__...__|_
                                    .            .
                                    .________.
                                    |__zone__|

- Superblock (SB)
   It is located at the beginning of the partition, and there exist two copies
   to avoid file system crash. It contains basic partition information and some
   default parameters of f2fs.

- Checkpoint (CP)
   It contains file system information, bitmaps for valid NAT/SIT sets, orphan
   inode lists, and summary entries of current active segments.

- Segment Information Table (SIT)
   It contains segment information such as valid block count and bitmap for the
   validity of all the blocks.

- Node Address Table (NAT)
   It is composed of a block address table for all the node blocks stored in
   Main area.

- Segment Summary Area (SSA)
   It contains summary entries which contains the owner information of all the
   data and node blocks stored in Main area.

- Main Area
   It contains file and directory data including their indices.

In order to avoid misalignment between file system and flash-based storage, F2FS
aligns the start block address of CP with the segment size. Also, it aligns the
start block address of Main area with the zone size by reserving some segments
in SSA area.

Reference the following survey for additional technical details.
https://wiki.linaro.org/WorkingGroups/Kernel/Projects/FlashCardSurvey

File System Metadata Structure
------------------------------

F2FS adopts the checkpointing scheme to maintain file system consistency. At
mount time, F2FS first tries to find the last valid checkpoint data by scanning
CP area. In order to reduce the scanning time, F2FS uses only two copies of CP.
One of them always indicates the last valid data, which is called as shadow copy
mechanism. In addition to CP, NAT and SIT also adopt the shadow copy mechanism.

For file system consistency, each CP points to which NAT and SIT copies are
valid, as shown as below::

  +--------+----------+---------+
  |   CP   |    SIT   |   NAT   |
  +--------+----------+---------+
  .         .          .          .
  .            .              .              .
  .               .                 .                 .
  +-------+-------+--------+--------+--------+--------+
  | CP #0 | CP #1 | SIT #0 | SIT #1 | NAT #0 | NAT #1 |
  +-------+-------+--------+--------+--------+--------+
     |             ^                          ^
     |             |                          |
     `----------------------------------------'

Node index와 multi-level directory hash

633-761

data 위치를 관리하는 핵심 구조는 `node`입니다. inode, direct node, indirect node의 세 종류가 있습니다. 4KiB inode block은 data index 923개, direct node pointer 2개, indirect node pointer 2개, double-indirect pointer 1개를 담습니다.

F2FS node index tree
계층하위 항목
Inode block1data index 923개
Direct node2각 data index 1,018개
Indirect node2각 direct node 1,018개, direct node마다 data 1,018개
Double-indirect node1indirect 1,018개 × direct 1,018개 × data 1,018개

원문의 index ASCII tree를 pointer 수와 fan-out 표로 재구성했습니다.

file 하나가 다루는 최대 범위는 `4KiB * (923 + 2*1018 + 2*1018*1018 + 1018*1018*1018)`, 약 3.94TiB입니다. 모든 node block 위치는 NAT가 변환하므로 leaf data write가 상위 node 주소 갱신으로 전파되지 않습니다.

directory entry 하나는 11바이트이며 file name hash, inode number, name length, directory·symlink 등의 file type으로 구성됩니다. 4KiB dentry block에는 valid bitmap 27바이트, reserved 3바이트, 11바이트 dentry 214개, 8바이트 name slot 214개가 들어갑니다.

4KiB dentry block
구성크기·개수의미
valid bitmap27 bytes각 dentry slot의 유효성
reserved3 bytes정렬·예약 공간
dentry`11 * 214` bytes`hash`, `ino`, `len`, `type`
file name slot`8 * 214` bytesfile name 저장

원문의 bucket·dentry ASCII layout을 field 표로 재구성했습니다.

directory는 multi-level hash table을 사용합니다. level `n < MAX_DIR_HASH_DEPTH/2`이면 bucket당 2 block, 이후에는 4 block입니다. bucket 수는 전반부에서 `2^(n + dir_level)`, 후반부에서 `2^((MAX_DIR_HASH_DEPTH/2)-1)`입니다.

Directory lookup
file name hash 계산level 0의 `hash % bucket_count` bucket scan없으면 level을 1씩 증가각 level에서도 bucket 하나만 scandentry의 name과 inode number 확인

각 level에서 계산된 bucket 하나만 scan해 O(log(file 수)) 복잡도를 냅니다.

file 생성은 name을 담을 연속 empty slot을 level 1부터 N까지 lookup과 같은 방식으로 찾습니다. directory의 logical file size는 hole이 있어도 마지막 배치 위치까지 반영되므로 child 수와 같지 않을 수 있습니다.

Index Structure
---------------

The key data structure to manage the data locations is a "node". Similar to
traditional file structures, F2FS has three types of node: inode, direct node,
indirect node. F2FS assigns 4KB to an inode block which contains 923 data block
indices, two direct node pointers, two indirect node pointers, and one double
indirect node pointer as described below. One direct node block contains 1018
data blocks, and one indirect node block contains also 1018 node blocks. Thus,
one inode block (i.e., a file) covers::

  4KB * (923 + 2 * 1018 + 2 * 1018 * 1018 + 1018 * 1018 * 1018) := 3.94TB.

   Inode block (4KB)
     |- data (923)
     |- direct node (2)
     |          `- data (1018)
     |- indirect node (2)
     |            `- direct node (1018)
     |                       `- data (1018)
     `- double indirect node (1)
                         `- indirect node (1018)
                                      `- direct node (1018)
                                                 `- data (1018)

Note that all the node blocks are mapped by NAT which means the location of
each node is translated by the NAT table. In the consideration of the wandering
tree problem, F2FS is able to cut off the propagation of node updates caused by
leaf data writes.

Directory Structure
-------------------

A directory entry occupies 11 bytes, which consists of the following attributes.

- hash                hash value of the file name
- ino                inode number
- len                the length of file name
- type                file type such as directory, symlink, etc

A dentry block consists of 214 dentry slots and file names. Therein a bitmap is
used to represent whether each dentry is valid or not. A dentry block occupies
4KB with the following composition.

::

  Dentry Block(4 K) = bitmap (27 bytes) + reserved (3 bytes) +
                      dentries(11 * 214 bytes) + file name (8 * 214 bytes)

                         [Bucket]
             +--------------------------------+
             |dentry block 1 | dentry block 2 |
             +--------------------------------+
             .               .
       .                             .
  .       [Dentry Block Structure: 4KB]       .
  +--------+----------+----------+------------+
  | bitmap | reserved | dentries | file names |
  +--------+----------+----------+------------+
  [Dentry Block: 4KB] .   .
                 .               .
            .                          .
            +------+------+-----+------+
            | hash | ino  | len | type |
            +------+------+-----+------+
            [Dentry Structure: 11 bytes]

F2FS implements multi-level hash tables for directory structure. Each level has
a hash table with dedicated number of hash buckets as shown below. Note that
"A(2B)" means a bucket includes 2 data blocks.

::

    ----------------------
    A : bucket
    B : block
    N : MAX_DIR_HASH_DEPTH
    ----------------------

    level #0   | A(2B)
            |
    level #1   | A(2B) - A(2B)
            |
    level #2   | A(2B) - A(2B) - A(2B) - A(2B)
        .     |   .       .       .       .
    level #N/2 | A(2B) - A(2B) - A(2B) - A(2B) - A(2B) - ... - A(2B)
        .     |   .       .       .       .
    level #N   | A(4B) - A(4B) - A(4B) - A(4B) - A(4B) - ... - A(4B)

The number of blocks and buckets are determined by::

                            ,- 2, if n < MAX_DIR_HASH_DEPTH / 2,
  # of blocks in level #n = |
                            `- 4, Otherwise

                             ,- 2^(n + dir_level),
                             |        if n + dir_level < MAX_DIR_HASH_DEPTH / 2,
  # of buckets in level #n = |
                             `- 2^((MAX_DIR_HASH_DEPTH / 2) - 1),
                                      Otherwise

When F2FS finds a file name in a directory, at first a hash value of the file
name is calculated. Then, F2FS scans the hash table in level #0 to find the
dentry consisting of the file name and its inode number. If not found, F2FS
scans the next hash table in level #1. In this way, F2FS scans hash tables in
each levels incrementally from 1 to N. In each level F2FS needs to scan only
one bucket determined by the following equation, which shows O(log(# of files))
complexity::

  bucket number to scan in level #n = (hash value) % (# of buckets in level #n)

In the case of file creation, F2FS finds empty consecutive slots that cover the
file name. F2FS searches the empty slots in the hash tables of whole levels from
1 to N in the same way as the lookup operation.

The following figure shows an example of two cases holding children::

       --------------> Dir <--------------
       |                                 |
    child                             child

    child - child                     [hole] - child

    child - child - child             [hole] - [hole] - child

   Case 1:                           Case 2:
   Number of children = 6,           Number of children = 3,
   File size = 7                     File size = 7

Block allocation·cleaning·write hint·fallocate

762-868

runtime에는 Main area 안에 Hot/Warm/Cold node와 Hot/Warm/Cold data의 active log 여섯 개가 있습니다.

여섯 active log
Log기본 내용
Hot nodedirectory의 direct node block
Warm nodehot node를 제외한 direct node block
Cold nodeindirect node block
Hot datadentry block
Warm datahot·cold data를 제외한 data block
Cold datamultimedia data 또는 migrated data block

node와 data의 temperature별 기본 분류입니다.

free-space 관리에서 copy-and-compaction(cleaning)은 sequential write가 빠른 device에 적합하지만 사용률이 높으면 cleaning overhead가 큽니다. threaded log는 random write가 생기지만 cleaning이 필요 없습니다. F2FS는 기본적으로 cleaning을 쓰다가 filesystem 상태에 따라 threaded log로 전환하는 hybrid policy를 사용합니다.

F2FS는 FTL GC unit과 맞추려고 section 단위로 segment를 할당하고, FTL mapping granularity를 고려해 active log section을 가능한 한 서로 다른 zone에서 할당합니다.

on-demand cleaner는 VFS call에 제공할 free segment가 부족할 때 실행되고 greedy algorithm으로 valid block이 가장 적은 victim을 고릅니다. idle 때 kernel thread가 실행하는 background cleaner는 age와 valid block 수를 함께 보는 cost-benefit algorithm으로 log block thrashing을 줄입니다. Main area 전체 block validity는 bitmap으로 추적합니다.

Write-hint policy
UserF2FSBlock
N/A`META``WRITE_LIFE_NONE | REQ_META`
N/A`HOT_NODE``WRITE_LIFE_NONE`
N/A`WARM_NODE``WRITE_LIFE_MEDIUM`
N/A`COLD_NODE``WRITE_LIFE_LONG`
`ioctl(COLD)` 또는 extension list`COLD_DATA``WRITE_LIFE_EXTREME`
buffered I/O: N/A`HOT_DATA``WRITE_LIFE_SHORT`
buffered I/O: N/A`WARM_DATA``WRITE_LIFE_NOT_SET`
direct I/O: `WRITE_LIFE_EXTREME``COLD_DATA``WRITE_LIFE_EXTREME`
direct I/O: `WRITE_LIFE_SHORT``HOT_DATA``WRITE_LIFE_SHORT`
direct I/O: `WRITE_LIFE_NOT_SET|NONE|MEDIUM|LONG``WARM_DATA`입력된 write life를 그대로 사용

User hint, F2FS temperature와 block write-life mapping입니다.

기본 `fallocate()` mode 0은 `offset..offset+len`에 disk space를 할당하고 필요하면 file size를 늘리며 기존 data가 없던 부분을 0으로 초기화합니다. 이는 `posix_fallocate(3)`를 효율적으로 구현하는 정책입니다.

먼저 `ioctl(fd, F2FS_IOC_SET_PIN_FILE)`을 호출하면 이후 `fallocate(fd, 0, 0, size)`가 0 또는 random data를 가진 on-disk block address를 고정 할당합니다. `fibmap()`으로 주소를 얻어 block device를 직접 여는 사용 시나리오에 유용합니다.

Pinned fallocate 사용 순서
`create(fd)``ioctl(fd, F2FS_IOC_SET_PIN_FILE)``fallocate(fd, 0, 0, size)``address = fibmap(fd, offset)``open(blkdev)``write(blkdev, address)`

원문의 syscall sequence를 보존한 흐름입니다.

Default Block Allocation
------------------------

At runtime, F2FS manages six active logs inside "Main" area: Hot/Warm/Cold node
and Hot/Warm/Cold data.

- Hot node        contains direct node blocks of directories.
- Warm node        contains direct node blocks except hot node blocks.
- Cold node        contains indirect node blocks
- Hot data        contains dentry blocks
- Warm data        contains data blocks except hot and cold data blocks
- Cold data        contains multimedia data or migrated data blocks

LFS has two schemes for free space management: threaded log and copy-and-compac-
tion. The copy-and-compaction scheme which is known as cleaning, is well-suited
for devices showing very good sequential write performance, since free segments
are served all the time for writing new data. However, it suffers from cleaning
overhead under high utilization. Contrarily, the threaded log scheme suffers
from random writes, but no cleaning process is needed. F2FS adopts a hybrid
scheme where the copy-and-compaction scheme is adopted by default, but the
policy is dynamically changed to the threaded log scheme according to the file
system status.

In order to align F2FS with underlying flash-based storage, F2FS allocates a
segment in a unit of section. F2FS expects that the section size would be the
same as the unit size of garbage collection in FTL. Furthermore, with respect
to the mapping granularity in FTL, F2FS allocates each section of the active
logs from different zones as much as possible, since FTL can write the data in
the active logs into one allocation unit according to its mapping granularity.

Cleaning process
----------------

F2FS does cleaning both on demand and in the background. On-demand cleaning is
triggered when there are not enough free segments to serve VFS calls. Background
cleaner is operated by a kernel thread, and triggers the cleaning job when the
system is idle.

F2FS supports two victim selection policies: greedy and cost-benefit algorithms.
In the greedy algorithm, F2FS selects a victim segment having the smallest number
of valid blocks. In the cost-benefit algorithm, F2FS selects a victim segment
according to the segment age and the number of valid blocks in order to address
log block thrashing problem in the greedy algorithm. F2FS adopts the greedy
algorithm for on-demand cleaner, while background cleaner adopts cost-benefit
algorithm.

In order to identify whether the data in the victim segment are valid or not,
F2FS manages a bitmap. Each bit represents the validity of a block, and the
bitmap is composed of a bit stream covering whole blocks in main area.

Write-hint Policy
-----------------

F2FS sets the whint all the time with the below policy.

===================== ======================== ===================
User                  F2FS                     Block
===================== ======================== ===================
N/A                   META                     WRITE_LIFE_NONE|REQ_META
N/A                   HOT_NODE                 WRITE_LIFE_NONE
N/A                   WARM_NODE                WRITE_LIFE_MEDIUM
N/A                   COLD_NODE                WRITE_LIFE_LONG
ioctl(COLD)           COLD_DATA                WRITE_LIFE_EXTREME
extension list        "                        "

-- buffered io
------------------------------------------------------------------
N/A                   COLD_DATA                WRITE_LIFE_EXTREME
N/A                   HOT_DATA                 WRITE_LIFE_SHORT
N/A                   WARM_DATA                WRITE_LIFE_NOT_SET

-- direct io
------------------------------------------------------------------
WRITE_LIFE_EXTREME    COLD_DATA                WRITE_LIFE_EXTREME
WRITE_LIFE_SHORT      HOT_DATA                 WRITE_LIFE_SHORT
WRITE_LIFE_NOT_SET    WARM_DATA                WRITE_LIFE_NOT_SET
WRITE_LIFE_NONE       "                        WRITE_LIFE_NONE
WRITE_LIFE_MEDIUM     "                        WRITE_LIFE_MEDIUM
WRITE_LIFE_LONG       "                        WRITE_LIFE_LONG
===================== ======================== ===================

Fallocate(2) Policy
-------------------

The default policy follows the below POSIX rule.

Allocating disk space
    The default operation (i.e., mode is zero) of fallocate() allocates
    the disk space within the range specified by offset and len.  The
    file size (as reported by stat(2)) will be changed if offset+len is
    greater than the file size.  Any subregion within the range specified
    by offset and len that did not contain data before the call will be
    initialized to zero.  This default behavior closely resembles the
    behavior of the posix_fallocate(3) library function, and is intended
    as a method of optimally implementing that function.

However, once F2FS receives ioctl(fd, F2FS_IOC_SET_PIN_FILE) in prior to
fallocate(fd, DEFAULT_MODE), it allocates on-disk block addresses having
zero or random data, which is useful to the below scenario where:

 1. create(fd)
 2. ioctl(fd, F2FS_IOC_SET_PIN_FILE)
 3. fallocate(fd, 0, 0, size)
 4. address = fibmap(fd, offset)
 5. open(blkdev)
 6. write(blkdev, address)

Compression 구현과 fs/user mode

869-970

compression 기본 단위인 cluster는 `4 << n`(`n >= 0`) logical page로 구성되며 file을 여러 cluster로 논리 분할합니다. cluster마다 독립적으로 압축 여부를 결정합니다.

특수 block address가 compressed cluster인지 normal cluster인지 표시합니다. compressed cluster의 metadata는 cluster를 `1 .. (4 << n)-1` physical block에 mapping하고 compress header와 compressed data를 저장합니다.

Compression metadata layout
계층Compressed clusterNormal cluster
Dnode entrycompress flag + physical block address들logical page마다 physical block address
Physical block 수`1 .. (4 << n)-1``4 << n`
Payload headerdata length, data checksum, reserved없음
Payloadcompressed datanormal page data

원문의 dnode·cluster ASCII layout을 구조화했습니다.

overwrite write amplification을 피하려고 write-once file만 압축하며, cluster의 모든 logical block에 valid data가 있고 압축률이 threshold보다 낮을 때만 압축합니다.

regular inode compression은 `chattr +c file`, 상위 directory의 `+c` 상속, `compress_extension=ext`, `compress_extension=*`로 켭니다. `chattr -c file` 또는 `nocompress_extension=ext`로 끕니다. 개별 file flag가 extension과 directory flag보다 우선합니다.

압축 공간은 향후 update 가능성을 보장하려고 기본적으로 user-visible free space로 돌려주지 않습니다. `F2FS_IOC_RELEASE_COMPRESS_BLOCKS`로 회수하면 inode flag가 write를 막으며, `F2FS_IOC_RESERVE_COMPRESS_BLOCKS`로 다시 예약하거나 file size를 0으로 truncate해야 write가 다시 가능합니다.

`compress_mode=fs`는 기본 mode로 writeback 때 자동 압축합니다. `compress_mode=user`는 자동 압축을 끄고 user가 `F2FS_IOC_DECOMPRESS_FILE`과 `F2FS_IOC_COMPRESS_FILE` ioctl로 대상과 시점을 정합니다.

User compression mode
`fd = open(filename, O_WRONLY, 0)`해제: `ioctl(fd, F2FS_IOC_DECOMPRESS_FILE)`압축: `ioctl(fd, F2FS_IOC_COMPRESS_FILE)`

manual decompress/compress 호출 흐름입니다.

Compression implementation
--------------------------

- New term named cluster is defined as basic unit of compression, file can
  be divided into multiple clusters logically. One cluster includes 4 << n
  (n >= 0) logical pages, compression size is also cluster size, each of
  cluster can be compressed or not.

- In cluster metadata layout, one special block address is used to indicate
  a cluster is a compressed one or normal one; for compressed cluster, following
  metadata maps cluster to [1, 4 << n - 1] physical blocks, in where f2fs
  stores data including compress header and compressed data.

- In order to eliminate write amplification during overwrite, F2FS only
  support compression on write-once file, data can be compressed only when
  all logical blocks in cluster contain valid data and compress ratio of
  cluster data is lower than specified threshold.

- To enable compression on regular inode, there are four ways:

  * chattr +c file
  * chattr +c dir; touch dir/file
  * mount w/ -o compress_extension=ext; touch file.ext
  * mount w/ -o compress_extension=*; touch any_file

- To disable compression on regular inode, there are two ways:

  * chattr -c file
  * mount w/ -o nocompress_extension=ext; touch file.ext

- Priority in between FS_COMPR_FL, FS_NOCOMP_FS, extensions:

  * compress_extension=so; nocompress_extension=zip; chattr +c dir; touch
    dir/foo.so; touch dir/bar.zip; touch dir/baz.txt; then foo.so and baz.txt
    should be compresse, bar.zip should be non-compressed. chattr +c dir/bar.zip
    can enable compress on bar.zip.
  * compress_extension=so; nocompress_extension=zip; chattr -c dir; touch
    dir/foo.so; touch dir/bar.zip; touch dir/baz.txt; then foo.so should be
    compresse, bar.zip and baz.txt should be non-compressed.
    chattr+c dir/bar.zip; chattr+c dir/baz.txt; can enable compress on bar.zip
    and baz.txt.

- At this point, compression feature doesn't expose compressed space to user
  directly in order to guarantee potential data updates later to the space.
  Instead, the main goal is to reduce data writes to flash disk as much as
  possible, resulting in extending disk life time as well as relaxing IO
  congestion. Alternatively, we've added ioctl(F2FS_IOC_RELEASE_COMPRESS_BLOCKS)
  interface to reclaim compressed space and show it to user after setting a
  special flag to the inode. Once the compressed space is released, the flag
  will block writing data to the file until either the compressed space is
  reserved via ioctl(F2FS_IOC_RESERVE_COMPRESS_BLOCKS) or the file size is
  truncated to zero.

Compress metadata layout::

                                [Dnode Structure]
                +-----------------------------------------------+
                | cluster 1 | cluster 2 | ......... | cluster N |
                +-----------------------------------------------+
                .           .                       .           .
          .                      .                .                      .
    .         Compressed Cluster       .        .        Normal Cluster            .
    +----------+---------+---------+---------+  +---------+---------+---------+---------+
    |compr flag| block 1 | block 2 | block 3 |  | block 1 | block 2 | block 3 | block 4 |
    +----------+---------+---------+---------+  +---------+---------+---------+---------+
               .                             .
            .                                           .
        .                                                           .
        +-------------+-------------+----------+----------------------------+
        | data length | data chksum | reserved |      compressed data       |
        +-------------+-------------+----------+----------------------------+

Compression mode
--------------------------

f2fs supports "fs" and "user" compression modes with "compression_mode" mount option.
With this option, f2fs provides a choice to select the way how to compress the
compression enabled files (refer to "Compression implementation" section for how to
enable compression on a regular inode).

1) compress_mode=fs

   This is the default option. f2fs does automatic compression in the writeback of the
   compression enabled files.

2) compress_mode=user

   This disables the automatic compression and gives the user discretion of choosing the
   target file and the timing. The user can do manual compression/decompression on the
   compression enabled files using F2FS_IOC_DECOMPRESS_FILE and F2FS_IOC_COMPRESS_FILE
   ioctls like the below.

To decompress a file::

  fd = open(filename, O_WRONLY, 0);
  ret = ioctl(fd, F2FS_IOC_DECOMPRESS_FILE);

To compress a file::

  fd = open(filename, O_WRONLY, 0);
  ret = ioctl(fd, F2FS_IOC_COMPRESS_FILE);

NVMe ZNS와 device aliasing

971-1028

ZNS의 zone capacity는 zone size와 같거나 작으며 zone에서 사용할 수 있는 block 수를 뜻합니다. mount 시 capacity보다 뒤에서 시작하는 segment는 free bitmap에서 permanently used로 표시해 write allocation과 GC 대상에서 제외합니다.

zone capacity가 기본 2MiB segment size에 정렬되지 않으면 capacity 앞에서 시작해 경계를 가로지르는 segment도 usable합니다. 다만 그 segment 안에서도 capacity 뒤 block은 unusable입니다.

device aliasing file은 일반 F2FS node structure 없이 storage device 전체를 하나의 큰 extent로 mapping합니다. 이 영역은 pinned 상태로 공간을 보유하여 F2FS 영역 일부를 다른 filesystem이나 용도로 임시 예약할 수 있습니다.

Device aliasing 수명
`mkfs.ext4 /dev/vdc`로 외부 filesystem 준비`mkfs.f2fs -c /dev/[email protected] /dev/vdb`로 alias 생성F2FS mount 후 `vdc.file`이 32GiB pinned extent를 보유`/dev/vdc`를 loop mount해 독립적으로 사용`f2fs_io setflags noimmutable .../vdc.file``rm .../vdc.file`로 예약 공간을 F2FS에 반환

외부 device를 F2FS file로 보유했다가 공간을 되돌리는 과정입니다.

핵심은 partition size나 filesystem format을 바꾸지 않고 `/dev/vdc`에서 임의 file operation을 수행하면서 그 공간은 F2FS의 `/data` 사용량으로 계산하고, 사용 후 alias file 삭제로 회수하는 것입니다.

NVMe Zoned Namespace devices
----------------------------

- ZNS defines a per-zone capacity which can be equal or less than the
  zone-size. Zone-capacity is the number of usable blocks in the zone.
  F2FS checks if zone-capacity is less than zone-size, if it is, then any
  segment which starts after the zone-capacity is marked as not-free in
  the free segment bitmap at initial mount time. These segments are marked
  as permanently used so they are not allocated for writes and
  consequently are not needed to be garbage collected. In case the
  zone-capacity is not aligned to default segment size(2MB), then a segment
  can start before the zone-capacity and span across zone-capacity boundary.
  Such spanning segments are also considered as usable segments. All blocks
  past the zone-capacity are considered unusable in these segments.

Device aliasing feature
-----------------------

f2fs can utilize a special file called a "device aliasing file." This file allows
the entire storage device to be mapped with a single, large extent, not using
the usual f2fs node structures. This mapped area is pinned and primarily intended
for holding the space.

Essentially, this mechanism allows a portion of the f2fs area to be temporarily
reserved and used by another filesystem or for different purposes. Once that
external usage is complete, the device aliasing file can be deleted, releasing
the reserved space back to F2FS for its own use.

.. code-block::

   # ls /dev/vd*
   /dev/vdb (32GB) /dev/vdc (32GB)
   # mkfs.ext4 /dev/vdc
   # mkfs.f2fs -c /dev/[email protected] /dev/vdb
   # mount /dev/vdb /mnt/f2fs
   # ls -l /mnt/f2fs
   vdc.file
   # df -h
   /dev/vdb                            64G   33G   32G  52% /mnt/f2fs

   # mount -o loop /dev/vdc /mnt/ext4
   # df -h
   /dev/vdb                            64G   33G   32G  52% /mnt/f2fs
   /dev/loop7                          32G   24K   30G   1% /mnt/ext4
   # umount /mnt/ext4

   # f2fs_io getflags /mnt/f2fs/vdc.file
   get a flag on /mnt/f2fs/vdc.file ret=0, flags=nocow(pinned),immutable
   # f2fs_io setflags noimmutable /mnt/f2fs/vdc.file
   get a flag on noimmutable ret=0, flags=800010
   set a flag on /mnt/f2fs/vdc.file ret=0, flags=noimmutable
   # rm /mnt/f2fs/vdc.file
   # df -h
   /dev/vdb                            64G  753M   64G   2% /mnt/f2fs

So, the key idea is, user can do any file operations on /dev/vdc, and
reclaim the space after the use, while the space is counted as /data.
That doesn't require modifying partition size and filesystem format.