요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
.. SPDX-License-Identifier: GPL-2.0
========================
VCPU Dispatch Statistics
========================
For Shared Processor LPARs, the POWER Hypervisor maintains a relatively
static mapping of the LPAR processors (vcpus) to physical processor
chips (representing the "home" node) and tries to always dispatch vcpus
on their associated physical processor chip. However, under certain
scenarios, vcpus may be dispatched on a different processor chip (away
from its home node).
/proc/powerpc/vcpudispatch_stats can be used to obtain statistics
related to the vcpu dispatch behavior. Writing '1' to this file enables
collecting the statistics, while writing '0' disables the statistics.
By default, the DTLB log for each vcpu is processed 50 times a second so
as not to miss any entries. This processing frequency can be changed
through /proc/powerpc/vcpudispatch_stats_freq.
The statistics themselves are available by reading the procfs file
/proc/powerpc/vcpudispatch_stats. Each line in the output corresponds to
a vcpu as represented by the first field, followed by 8 numbers.
The first number corresponds to:
1. total vcpu dispatches since the beginning of statistics collection
The next 4 numbers represent vcpu dispatch dispersions:
2. number of times this vcpu was dispatched on the same processor as last
time
3. number of times this vcpu was dispatched on a different processor core
as last time, but within the same chip
4. number of times this vcpu was dispatched on a different chip
5. number of times this vcpu was dispatches on a different socket/drawer
(next numa boundary)
The final 3 numbers represent statistics in relation to the home node of
the vcpu:
6. number of times this vcpu was dispatched in its home node (chip)
7. number of times this vcpu was dispatched in a different node
8. number of times this vcpu was dispatched in a node further away (numa
distance)
An example output::
$ sudo cat /proc/powerpc/vcpudispatch_stats
cpu0 6839 4126 2683 30 0 6821 18 0
cpu1 2515 1274 1229 12 0 2509 6 0
cpu2 2317 1198 1109 10 0 2312 5 0
cpu3 2259 1165 1088 6 0 2256 3 0
cpu4 2205 1143 1056 6 0 2202 3 0
cpu5 2165 1121 1038 6 0 2162 3 0
cpu6 2183 1127 1050 6 0 2180 3 0
cpu7 2193 1133 1052 8 0 2187 6 0
cpu8 2165 1115 1032 18 0 2156 9 0
cpu9 2301 1252 1033 16 0 2293 8 0
cpu10 2197 1138 1041 18 0 2187 10 0
cpu11 2273 1185 1062 26 0 2260 13 0
cpu12 2186 1125 1043 18 0 2177 9 0
cpu13 2161 1115 1030 16 0 2153 8 0
cpu14 2206 1153 1033 20 0 2196 10 0
cpu15 2163 1115 1032 16 0 2155 8 0
In the output above, for vcpu0, there have been 6839 dispatches since
statistics were enabled. 4126 of those dispatches were on the same
physical cpu as the last time. 2683 were on a different core, but within
the same chip, while 30 dispatches were on a different chip compared to
its last dispatch.
Also, out of the total of 6839 dispatches, we see that there have been
6821 dispatches on the vcpu's home node, while 18 dispatches were
outside its home node, on a neighbouring chip.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
VCPU Dispatch Statistics
1-6이 문서는 POWER Shared Processor LPAR에서 vCPU가 home chip과 다른 위치에 dispatch되는 양상을 procfs statistics로 해석하는 방법을 설명합니다.
LPAR vCPU home-node mapping
7-13POWER Hypervisor는 Shared Processor LPAR의 logical processor(vCPU)를 home node인 physical processor chip에 비교적 static하게 mapping하고 가능한 한 해당 chip에 dispatch합니다. 특정 상황에서는 home node가 아닌 chip으로 dispatch될 수 있습니다.
Dispatch는 home processor를 우선하지만 core, chip, socket 경계를 넘어갈 수 있습니다.
Procfs collection interface
14-24`/proc/powerpc/vcpudispatch_stats`에 `1`을 쓰면 collection을 enable하고 `0`을 쓰면 disable합니다.
기본적으로 각 vCPU의 DTLB log를 초당 50회 처리하여 entry를 놓치지 않게 합니다. Frequency는 `/proc/powerpc/vcpudispatch_stats_freq`로 바꿀 수 있습니다.
Statistics file을 읽으면 각 line의 첫 field가 vCPU이고 그 뒤에 8개 number가 옵니다.
출력 8개 counter
25-46| Field | 의미 |
|---|---|
| 1 | Statistics collection 시작 이후 total vCPU dispatch 수 |
| 2 | 직전과 같은 physical processor에 dispatch된 횟수 |
| 3 | 같은 chip 안에서 직전과 다른 processor core에 dispatch된 횟수 |
| 4 | 직전과 다른 chip에 dispatch된 횟수 |
| 5 | 다른 socket/drawer, 즉 다음 NUMA boundary에 dispatch된 횟수 |
| 6 | vCPU home node(chip)에 dispatch된 횟수 |
| 7 | 다른 node에 dispatch된 횟수 |
| 8 | NUMA distance가 더 먼 node에 dispatch된 횟수 |
Total, 직전 dispatch 대비 dispersion, home node 대비 locality의 세 묶음입니다.
Example output
47-66예제는 CPU별로 8개 dispatch counter를 한 줄에 출력하며, 다음 `vcpu0` 값은 직전 실행 위치와 home node 기준 locality를 함께 보여줍니다.
$ sudo cat /proc/powerpc/vcpudispatch_stats
cpu0 6839 4126 2683 30 0 6821 18 0
cpu1 2515 1274 1229 12 0 2509 6 0
cpu2 2317 1198 1109 10 0 2312 5 0
cpu3 2259 1165 1088 6 0 2256 3 0
cpu4 2205 1143 1056 6 0 2202 3 0
cpu5 2165 1121 1038 6 0 2162 3 0
cpu6 2183 1127 1050 6 0 2180 3 0
cpu7 2193 1133 1052 8 0 2187 6 0
cpu8 2165 1115 1032 18 0 2156 9 0
cpu9 2301 1252 1033 16 0 2293 8 0
cpu10 2197 1138 1041 18 0 2187 10 0
cpu11 2273 1185 1062 26 0 2260 13 0
cpu12 2186 1125 1043 18 0 2177 9 0
cpu13 2161 1115 1030 16 0 2153 8 0
cpu14 2206 1153 1033 20 0 2196 10 0
cpu15 2163 1115 1032 16 0 2155 8 0
vcpu0 결과 해석
67-75`vcpu0`는 collection 시작 뒤 6,839회 dispatch되었습니다. 그중 4,126회는 직전과 같은 physical CPU, 2,683회는 같은 chip의 다른 core, 30회는 다른 chip이었습니다.
Home-node 기준으로는 6,821회가 home chip에서 실행됐고 18회가 인접 chip의 home 밖에서 실행됐습니다.
요약과 해설
vcpudispatch_stats.rst:1-75`vcpudispatch_stats`는 total dispatch, 직전 execution location 대비 dispersion, 고정 home node 대비 locality를 분리합니다. 두 관점을 함께 봐야 scheduler 이동과 NUMA 원격 실행을 구분할 수 있습니다.