← Documents Documentation/virt/kvm/devices/vcpu.rst GitHub 원문 ↗

Linux 6.18.37 · 가상화 / KVM / Device

Generic vCPU device interface

ARM64 PMUv3·timer·stolen-time attribute와 x86 TSC offset migration algorithm을 정의하는 vCPU device ABI입니다.

Source pathDocumentation/virt/kvm/devices/vcpu.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

vcpu.rst:1-293

ARM64 PMUv3·timer·stolen-time attribute와 x86 TSC offset migration algorithm을 정의하는 vCPU device ABI입니다.

구조체 field, ioctl, attribute, error code와 migration 수식은 원문 표기를 유지했습니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 ======================
4 Generic vcpu interface
5 ======================
6
7 The virtual cpu "device" also accepts the ioctls KVM_SET_DEVICE_ATTR,
8 KVM_GET_DEVICE_ATTR, and KVM_HAS_DEVICE_ATTR. The interface uses the same struct
9 kvm_device_attr as other devices, but targets VCPU-wide settings and controls.
10
11 The groups and attributes per virtual cpu, if any, are architecture specific.
12
13 1. GROUP: KVM_ARM_VCPU_PMU_V3_CTRL
14 ==================================
15
16 :Architectures: ARM64
17
18 1.1. ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_IRQ
19 ---------------------------------------
20
21 :Parameters: in kvm_device_attr.addr the address for PMU overflow interrupt is a
22 pointer to an int
23
24 Returns:
25
26 ======= ========================================================
27 -EBUSY The PMU overflow interrupt is already set
28 -EFAULT Error reading interrupt number
29 -ENXIO PMUv3 not supported or the overflow interrupt not set
30 when attempting to get it
31 -ENODEV KVM_ARM_VCPU_PMU_V3 feature missing from VCPU
32 -EINVAL Invalid PMU overflow interrupt number supplied or
33 trying to set the IRQ number without using an in-kernel
34 irqchip.
35 ======= ========================================================
36
37 A value describing the PMUv3 (Performance Monitor Unit v3) overflow interrupt
38 number for this vcpu. This interrupt could be a PPI or SPI, but the interrupt
39 type must be same for each vcpu. As a PPI, the interrupt number is the same for
40 all vcpus, while as an SPI it must be a separate number per vcpu.
41
42 1.2 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_INIT
43 ---------------------------------------
44
45 :Parameters: no additional parameter in kvm_device_attr.addr
46
47 Returns:
48
49 ======= ======================================================
50 -EEXIST Interrupt number already used
51 -ENODEV PMUv3 not supported or GIC not initialized
52 -ENXIO PMUv3 not supported, missing VCPU feature or interrupt
53 number not set
54 -EBUSY PMUv3 already initialized
55 ======= ======================================================
56
57 Request the initialization of the PMUv3. If using the PMUv3 with an in-kernel
58 virtual GIC implementation, this must be done after initializing the in-kernel
59 irqchip.
60
61 1.3 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_FILTER
62 -----------------------------------------
63
64 :Parameters: in kvm_device_attr.addr the address for a PMU event filter is a
65 pointer to a struct kvm_pmu_event_filter
66
67 :Returns:
68
69 ======= ======================================================
70 -ENODEV PMUv3 not supported or GIC not initialized
71 -ENXIO PMUv3 not properly configured or in-kernel irqchip not
72 configured as required prior to calling this attribute
73 -EBUSY PMUv3 already initialized or a VCPU has already run
74 -EINVAL Invalid filter range
75 ======= ======================================================
76
77 Request the installation of a PMU event filter described as follows::
78
79 struct kvm_pmu_event_filter {
80 __u16 base_event;
81 __u16 nevents;
82
83 #define KVM_PMU_EVENT_ALLOW 0
84 #define KVM_PMU_EVENT_DENY 1
85
86 __u8 action;
87 __u8 pad[3];
88 };
89
90 A filter range is defined as the range [@base_event, @base_event + @nevents),
91 together with an @action (KVM_PMU_EVENT_ALLOW or KVM_PMU_EVENT_DENY). The
92 first registered range defines the global policy (global ALLOW if the first
93 @action is DENY, global DENY if the first @action is ALLOW). Multiple ranges
94 can be programmed, and must fit within the event space defined by the PMU
95 architecture (10 bits on ARMv8.0, 16 bits from ARMv8.1 onwards).
96
97 Note: "Cancelling" a filter by registering the opposite action for the same
98 range doesn't change the default action. For example, installing an ALLOW
99 filter for event range [0:10) as the first filter and then applying a DENY
100 action for the same range will leave the whole range as disabled.
101
102 Restrictions: Event 0 (SW_INCR) is never filtered, as it doesn't count a
103 hardware event. Filtering event 0x1E (CHAIN) has no effect either, as it
104 isn't strictly speaking an event. Filtering the cycle counter is possible
105 using event 0x11 (CPU_CYCLES).
106
107 1.4 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_SET_PMU
108 ------------------------------------------
109
110 :Parameters: in kvm_device_attr.addr the address to an int representing the PMU
111 identifier.
112
113 :Returns:
114
115 ======= ====================================================
116 -EBUSY PMUv3 already initialized, a VCPU has already run or
117 an event filter has already been set
118 -EFAULT Error accessing the PMU identifier
119 -ENXIO PMU not found
120 -ENODEV PMUv3 not supported or GIC not initialized
121 -ENOMEM Could not allocate memory
122 ======= ====================================================
123
124 Request that the VCPU uses the specified hardware PMU when creating guest events
125 for the purpose of PMU emulation. The PMU identifier can be read from the "type"
126 file for the desired PMU instance under /sys/devices (or, equivalent,
127 /sys/bus/even_source). This attribute is particularly useful on heterogeneous
128 systems where there are at least two CPU PMUs on the system. The PMU that is set
129 for one VCPU will be used by all the other VCPUs. It isn't possible to set a PMU
130 if a PMU event filter is already present.
131
132 Note that KVM will not make any attempts to run the VCPU on the physical CPUs
133 associated with the PMU specified by this attribute. This is entirely left to
134 userspace. However, attempting to run the VCPU on a physical CPU not supported
135 by the PMU will fail and KVM_RUN will return with
136 exit_reason = KVM_EXIT_FAIL_ENTRY and populate the fail_entry struct by setting
137 hardare_entry_failure_reason field to KVM_EXIT_FAIL_ENTRY_CPU_UNSUPPORTED and
138 the cpu field to the processor id.
139
140 1.5 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_SET_NR_COUNTERS
141 --------------------------------------------------
142
143 :Parameters: in kvm_device_attr.addr the address to an unsigned int
144 representing the maximum value taken by PMCR_EL0.N
145
146 :Returns:
147
148 ======= ====================================================
149 -EBUSY PMUv3 already initialized, a VCPU has already run or
150 an event filter has already been set
151 -EFAULT Error accessing the value pointed to by addr
152 -ENODEV PMUv3 not supported or GIC not initialized
153 -EINVAL No PMUv3 explicitly selected, or value of N out of
154 range
155 ======= ====================================================
156
157 Set the number of implemented event counters in the virtual PMU. This
158 mandates that a PMU has explicitly been selected via
159 KVM_ARM_VCPU_PMU_V3_SET_PMU, and will fail when no PMU has been
160 explicitly selected, or the number of counters is out of range for the
161 selected PMU. Selecting a new PMU cancels the effect of setting this
162 attribute.
163
164 2. GROUP: KVM_ARM_VCPU_TIMER_CTRL
165 =================================
166
167 :Architectures: ARM64
168
169 2.1. ATTRIBUTES: KVM_ARM_VCPU_TIMER_IRQ_{VTIMER,PTIMER,HVTIMER,HPTIMER}
170 -----------------------------------------------------------------------
171
172 :Parameters: in kvm_device_attr.addr the address for the timer interrupt is a
173 pointer to an int
174
175 Returns:
176
177 ======= =================================
178 -EINVAL Invalid timer interrupt number
179 -EBUSY One or more VCPUs has already run
180 ======= =================================
181
182 A value describing the architected timer interrupt number when connected to an
183 in-kernel virtual GIC. These must be a PPI (16 <= intid < 32). Setting the
184 attribute overrides the default values (see below).
185
186 ============================== ==========================================
187 KVM_ARM_VCPU_TIMER_IRQ_VTIMER The EL1 virtual timer intid (default: 27)
188 KVM_ARM_VCPU_TIMER_IRQ_PTIMER The EL1 physical timer intid (default: 30)
189 KVM_ARM_VCPU_TIMER_IRQ_HVTIMER The EL2 virtual timer intid (default: 28)
190 KVM_ARM_VCPU_TIMER_IRQ_HPTIMER The EL2 physical timer intid (default: 26)
191 ============================== ==========================================
192
193 Setting the same PPI for different timers will prevent the VCPUs from running.
194 Setting the interrupt number on a VCPU configures all VCPUs created at that
195 time to use the number provided for a given timer, overwriting any previously
196 configured values on other VCPUs. Userspace should configure the interrupt
197 numbers on at least one VCPU after creating all VCPUs and before running any
198 VCPUs.
199
200 .. _kvm_arm_vcpu_pvtime_ctrl:
201
202 3. GROUP: KVM_ARM_VCPU_PVTIME_CTRL
203 ==================================
204
205 :Architectures: ARM64
206
207 3.1 ATTRIBUTE: KVM_ARM_VCPU_PVTIME_IPA
208 --------------------------------------
209
210 :Parameters: 64-bit base address
211
212 Returns:
213
214 ======= ======================================
215 -ENXIO Stolen time not implemented
216 -EEXIST Base address already set for this VCPU
217 -EINVAL Base address not 64 byte aligned
218 ======= ======================================
219
220 Specifies the base address of the stolen time structure for this VCPU. The
221 base address must be 64 byte aligned and exist within a valid guest memory
222 region. See Documentation/virt/kvm/arm/pvtime.rst for more information
223 including the layout of the stolen time structure.
224
225 4. GROUP: KVM_VCPU_TSC_CTRL
226 ===========================
227
228 :Architectures: x86
229
230 4.1 ATTRIBUTE: KVM_VCPU_TSC_OFFSET
231
232 :Parameters: 64-bit unsigned TSC offset
233
234 Returns:
235
236 ======= ======================================
237 -EFAULT Error reading/writing the provided
238 parameter address.
239 -ENXIO Attribute not supported
240 ======= ======================================
241
242 Specifies the guest's TSC offset relative to the host's TSC. The guest's
243 TSC is then derived by the following equation:
244
245 guest_tsc = host_tsc + KVM_VCPU_TSC_OFFSET
246
247 This attribute is useful to adjust the guest's TSC on live migration,
248 so that the TSC counts the time during which the VM was paused. The
249 following describes a possible algorithm to use for this purpose.
250
251 From the source VMM process:
252
253 1. Invoke the KVM_GET_CLOCK ioctl to record the host TSC (tsc_src),
254 kvmclock nanoseconds (guest_src), and host CLOCK_REALTIME nanoseconds
255 (host_src).
256
257 2. Read the KVM_VCPU_TSC_OFFSET attribute for every vCPU to record the
258 guest TSC offset (ofs_src[i]).
259
260 3. Invoke the KVM_GET_TSC_KHZ ioctl to record the frequency of the
261 guest's TSC (freq).
262
263 From the destination VMM process:
264
265 4. Invoke the KVM_SET_CLOCK ioctl, providing the source nanoseconds from
266 kvmclock (guest_src) and CLOCK_REALTIME (host_src) in their respective
267 fields. Ensure that the KVM_CLOCK_REALTIME flag is set in the provided
268 structure.
269
270 KVM will advance the VM's kvmclock to account for elapsed time since
271 recording the clock values. Note that this will cause problems in
272 the guest (e.g., timeouts) unless CLOCK_REALTIME is synchronized
273 between the source and destination, and a reasonably short time passes
274 between the source pausing the VMs and the destination executing
275 steps 4-7.
276
277 5. Invoke the KVM_GET_CLOCK ioctl to record the host TSC (tsc_dest) and
278 kvmclock nanoseconds (guest_dest).
279
280 6. Adjust the guest TSC offsets for every vCPU to account for (1) time
281 elapsed since recording state and (2) difference in TSCs between the
282 source and destination machine:
283
284 ofs_dst[i] = ofs_src[i] -
285 (guest_src - guest_dest) * freq +
286 (tsc_src - tsc_dest)
287
288 ("ofs[i] + tsc - guest * freq" is the guest TSC value corresponding to
289 a time of 0 in kvmclock. The above formula ensures that it is the
290 same on the destination as it was on the source).
291
292 7. Write the KVM_VCPU_TSC_OFFSET attribute for every vCPU with the
293 respective value derived in the previous step.
294

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

ARM64 PMUv3 control

1-163

Virtual CPU device도 다른 KVM device와 같은 `struct kvm_device_attr`을 사용해 `KVM_SET_DEVICE_ATTR`, `KVM_GET_DEVICE_ATTR`, `KVM_HAS_DEVICE_ATTR`을 받습니다. Group과 attribute는 architecture-specific이며 vCPU 전체의 설정과 control을 대상으로 합니다.

`KVM_ARM_VCPU_PMU_V3_IRQ`는 이 vCPU의 PMUv3 overflow interrupt number를 int pointer로 get/set합니다. PPI 또는 SPI일 수 있지만 모든 vCPU에서 interrupt type은 같아야 합니다. PPI면 number도 모두 같고 SPI면 vCPU마다 별도 number가 필요합니다.

PMUv3 IRQ error
Error조건
`-EBUSY`Overflow interrupt가 이미 설정됨
`-EFAULT`Interrupt number read 실패
`-ENXIO`PMUv3 미지원 또는 get 시 overflow interrupt 미설정
`-ENODEV`vCPU에 `KVM_ARM_VCPU_PMU_V3` feature 없음
`-EINVAL`잘못된 interrupt number 또는 in-kernel irqchip 없이 IRQ 설정

Overflow interrupt 구성 오류입니다.

`KVM_ARM_VCPU_PMU_V3_INIT`은 추가 parameter 없이 PMUv3를 initialize합니다. In-kernel virtual GIC를 쓸 때는 irqchip initialization 뒤에 호출해야 합니다. 이미 사용 중인 IRQ는 `-EEXIST`, GIC·PMU 미준비는 `-ENODEV`·`-ENXIO`, 중복 init은 `-EBUSY`입니다.

`KVM_ARM_VCPU_PMU_V3_FILTER`는 `kvm_pmu_event_filter`의 `base_event`, `nevents`, `action`으로 half-open event range `[base_event, base_event + nevents)`를 ALLOW 또는 DENY합니다.

PMU event-filter policy
첫 actionGlobal policy
`KVM_PMU_EVENT_DENY`Global ALLOW; 지정 range만 deny
`KVM_PMU_EVENT_ALLOW`Global DENY; 지정 range만 allow

처음 등록한 range action의 반대가 global default가 됩니다.

여러 range를 등록할 수 있으며 ARMv8.0 event space는 10-bit, ARMv8.1 이후는 16-bit입니다. 같은 range에 반대 action을 등록해도 최초 global default는 바뀌지 않습니다. 예를 들어 첫 ALLOW [0:10) 뒤 같은 range를 DENY하면 전체 range가 disabled로 남습니다.

`SW_INCR` event 0은 hardware event를 count하지 않으므로 filter되지 않습니다. `CHAIN` 0x1E filter도 효과가 없고 cycle counter는 `CPU_CYCLES` 0x11로 filter할 수 있습니다.

Filter 설치는 PMUv3·GIC 준비가 필요하고 init 또는 vCPU 실행 뒤에는 `-EBUSY`, 잘못된 range는 `-EINVAL`입니다.

`KVM_ARM_VCPU_PMU_V3_SET_PMU`는 `/sys/devices` 또는 `/sys/bus/event_source`의 PMU `type` 값을 지정해 guest event emulation에 사용할 physical PMU를 선택합니다. Heterogeneous system에서 특히 유용하며 한 vCPU의 선택이 모든 vCPU에 적용됩니다.

KVM은 선택 PMU와 연결된 physical CPU로 vCPU를 자동 pin하지 않습니다. Userspace가 scheduling affinity를 관리해야 하며 지원하지 않는 CPU에서 실행하면 `KVM_RUN`이 `KVM_EXIT_FAIL_ENTRY`와 `KVM_EXIT_FAIL_ENTRY_CPU_UNSUPPORTED` reason 및 processor ID를 반환합니다.

PMU event filter가 이미 있으면 PMU를 선택할 수 없습니다. Init·vCPU run·filter 뒤 SET_PMU는 `-EBUSY`; PMU ID access 실패는 `-EFAULT`; PMU 미발견은 `-ENXIO`; allocation 실패는 `-ENOMEM`입니다.

`KVM_ARM_VCPU_PMU_V3_SET_NR_COUNTERS`는 virtual PMU의 `PMCR_EL0.N` 최대값으로 구현 event-counter 수를 제한합니다. SET_PMU로 physical PMU를 명시적으로 선택한 뒤에만 가능하고 선택 PMU 범위를 벗어나면 `-EINVAL`입니다. 새 PMU를 선택하면 이 설정 효과는 취소됩니다.

.. SPDX-License-Identifier: GPL-2.0

======================
Generic vcpu interface
======================

The virtual cpu "device" also accepts the ioctls KVM_SET_DEVICE_ATTR,
KVM_GET_DEVICE_ATTR, and KVM_HAS_DEVICE_ATTR. The interface uses the same struct
kvm_device_attr as other devices, but targets VCPU-wide settings and controls.

The groups and attributes per virtual cpu, if any, are architecture specific.

1. GROUP: KVM_ARM_VCPU_PMU_V3_CTRL
==================================

:Architectures: ARM64

1.1. ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_IRQ
---------------------------------------

:Parameters: in kvm_device_attr.addr the address for PMU overflow interrupt is a
	     pointer to an int

Returns:

	 =======  ========================================================
	 -EBUSY   The PMU overflow interrupt is already set
	 -EFAULT  Error reading interrupt number
	 -ENXIO   PMUv3 not supported or the overflow interrupt not set
		  when attempting to get it
	 -ENODEV  KVM_ARM_VCPU_PMU_V3 feature missing from VCPU
	 -EINVAL  Invalid PMU overflow interrupt number supplied or
		  trying to set the IRQ number without using an in-kernel
		  irqchip.
	 =======  ========================================================

A value describing the PMUv3 (Performance Monitor Unit v3) overflow interrupt
number for this vcpu. This interrupt could be a PPI or SPI, but the interrupt
type must be same for each vcpu. As a PPI, the interrupt number is the same for
all vcpus, while as an SPI it must be a separate number per vcpu.

1.2 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_INIT
---------------------------------------

:Parameters: no additional parameter in kvm_device_attr.addr

Returns:

	 =======  ======================================================
	 -EEXIST  Interrupt number already used
	 -ENODEV  PMUv3 not supported or GIC not initialized
	 -ENXIO   PMUv3 not supported, missing VCPU feature or interrupt
		  number not set
	 -EBUSY   PMUv3 already initialized
	 =======  ======================================================

Request the initialization of the PMUv3.  If using the PMUv3 with an in-kernel
virtual GIC implementation, this must be done after initializing the in-kernel
irqchip.

1.3 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_FILTER
-----------------------------------------

:Parameters: in kvm_device_attr.addr the address for a PMU event filter is a
             pointer to a struct kvm_pmu_event_filter

:Returns:

	 =======  ======================================================
	 -ENODEV  PMUv3 not supported or GIC not initialized
	 -ENXIO   PMUv3 not properly configured or in-kernel irqchip not
	 	  configured as required prior to calling this attribute
	 -EBUSY   PMUv3 already initialized or a VCPU has already run
	 -EINVAL  Invalid filter range
	 =======  ======================================================

Request the installation of a PMU event filter described as follows::

    struct kvm_pmu_event_filter {
	    __u16	base_event;
	    __u16	nevents;

    #define KVM_PMU_EVENT_ALLOW	0
    #define KVM_PMU_EVENT_DENY	1

	    __u8	action;
	    __u8	pad[3];
    };

A filter range is defined as the range [@base_event, @base_event + @nevents),
together with an @action (KVM_PMU_EVENT_ALLOW or KVM_PMU_EVENT_DENY). The
first registered range defines the global policy (global ALLOW if the first
@action is DENY, global DENY if the first @action is ALLOW). Multiple ranges
can be programmed, and must fit within the event space defined by the PMU
architecture (10 bits on ARMv8.0, 16 bits from ARMv8.1 onwards).

Note: "Cancelling" a filter by registering the opposite action for the same
range doesn't change the default action. For example, installing an ALLOW
filter for event range [0:10) as the first filter and then applying a DENY
action for the same range will leave the whole range as disabled.

Restrictions: Event 0 (SW_INCR) is never filtered, as it doesn't count a
hardware event. Filtering event 0x1E (CHAIN) has no effect either, as it
isn't strictly speaking an event. Filtering the cycle counter is possible
using event 0x11 (CPU_CYCLES).

1.4 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_SET_PMU
------------------------------------------

:Parameters: in kvm_device_attr.addr the address to an int representing the PMU
             identifier.

:Returns:

	 =======  ====================================================
	 -EBUSY   PMUv3 already initialized, a VCPU has already run or
                  an event filter has already been set
	 -EFAULT  Error accessing the PMU identifier
	 -ENXIO   PMU not found
	 -ENODEV  PMUv3 not supported or GIC not initialized
	 -ENOMEM  Could not allocate memory
	 =======  ====================================================

Request that the VCPU uses the specified hardware PMU when creating guest events
for the purpose of PMU emulation. The PMU identifier can be read from the "type"
file for the desired PMU instance under /sys/devices (or, equivalent,
/sys/bus/even_source). This attribute is particularly useful on heterogeneous
systems where there are at least two CPU PMUs on the system. The PMU that is set
for one VCPU will be used by all the other VCPUs. It isn't possible to set a PMU
if a PMU event filter is already present.

Note that KVM will not make any attempts to run the VCPU on the physical CPUs
associated with the PMU specified by this attribute. This is entirely left to
userspace. However, attempting to run the VCPU on a physical CPU not supported
by the PMU will fail and KVM_RUN will return with
exit_reason = KVM_EXIT_FAIL_ENTRY and populate the fail_entry struct by setting
hardare_entry_failure_reason field to KVM_EXIT_FAIL_ENTRY_CPU_UNSUPPORTED and
the cpu field to the processor id.

1.5 ATTRIBUTE: KVM_ARM_VCPU_PMU_V3_SET_NR_COUNTERS
--------------------------------------------------

:Parameters: in kvm_device_attr.addr the address to an unsigned int
	     representing the maximum value taken by PMCR_EL0.N

:Returns:

	 =======  ====================================================
	 -EBUSY   PMUv3 already initialized, a VCPU has already run or
                  an event filter has already been set
	 -EFAULT  Error accessing the value pointed to by addr
	 -ENODEV  PMUv3 not supported or GIC not initialized
	 -EINVAL  No PMUv3 explicitly selected, or value of N out of
	 	  range
	 =======  ====================================================

Set the number of implemented event counters in the virtual PMU. This
mandates that a PMU has explicitly been selected via
KVM_ARM_VCPU_PMU_V3_SET_PMU, and will fail when no PMU has been
explicitly selected, or the number of counters is out of range for the
selected PMU. Selecting a new PMU cancels the effect of setting this
attribute.

ARM64 architected timer IRQ

164-201

`KVM_ARM_VCPU_TIMER_CTRL`은 in-kernel virtual GIC에 연결되는 architected timer interrupt를 설정합니다. 모든 timer IRQ는 PPI 범위 `16 <= intid < 32`여야 하며 잘못된 number는 `-EINVAL`, 어느 vCPU든 이미 실행됐으면 `-EBUSY`입니다.

ARM vCPU timer 기본 PPI
AttributeTimerDefault
`KVM_ARM_VCPU_TIMER_IRQ_VTIMER`EL1 virtual timer27
`KVM_ARM_VCPU_TIMER_IRQ_PTIMER`EL1 physical timer30
`KVM_ARM_VCPU_TIMER_IRQ_HVTIMER`EL2 virtual timer28
`KVM_ARM_VCPU_TIMER_IRQ_HPTIMER`EL2 physical timer26

Attribute 설정은 아래 default를 override합니다.

서로 다른 timer에 같은 PPI를 설정하면 vCPU가 실행되지 않습니다. 한 vCPU에서 설정하면 그 시점에 생성된 모든 vCPU의 해당 timer 값이 바뀌고 다른 vCPU의 이전 설정도 덮어씁니다.

Userspace는 모든 vCPU를 만든 뒤 어느 vCPU도 실행하기 전에 최소 한 vCPU에서 timer interrupt number를 구성해야 합니다.

2. GROUP: KVM_ARM_VCPU_TIMER_CTRL
=================================

:Architectures: ARM64

2.1. ATTRIBUTES: KVM_ARM_VCPU_TIMER_IRQ_{VTIMER,PTIMER,HVTIMER,HPTIMER}
-----------------------------------------------------------------------

:Parameters: in kvm_device_attr.addr the address for the timer interrupt is a
	     pointer to an int

Returns:

	 =======  =================================
	 -EINVAL  Invalid timer interrupt number
	 -EBUSY   One or more VCPUs has already run
	 =======  =================================

A value describing the architected timer interrupt number when connected to an
in-kernel virtual GIC.  These must be a PPI (16 <= intid < 32).  Setting the
attribute overrides the default values (see below).

==============================  ==========================================
KVM_ARM_VCPU_TIMER_IRQ_VTIMER   The EL1 virtual timer intid (default: 27)
KVM_ARM_VCPU_TIMER_IRQ_PTIMER   The EL1 physical timer intid (default: 30)
KVM_ARM_VCPU_TIMER_IRQ_HVTIMER  The EL2 virtual timer intid (default: 28)
KVM_ARM_VCPU_TIMER_IRQ_HPTIMER  The EL2 physical timer intid (default: 26)
==============================  ==========================================

Setting the same PPI for different timers will prevent the VCPUs from running.
Setting the interrupt number on a VCPU configures all VCPUs created at that
time to use the number provided for a given timer, overwriting any previously
configured values on other VCPUs.  Userspace should configure the interrupt
numbers on at least one VCPU after creating all VCPUs and before running any
VCPUs.

.. _kvm_arm_vcpu_pvtime_ctrl:

ARM64 stolen-time IPA

202-224

`KVM_ARM_VCPU_PVTIME_IPA`는 해당 vCPU의 stolen-time structure base address를 64-bit 값으로 설정합니다. Address는 64-byte aligned이고 유효한 guest memory region 안에 있어야 합니다.

PVTIME IPA error
Error조건
`-ENXIO`Stolen time 미구현
`-EEXIST`해당 vCPU base가 이미 설정됨
`-EINVAL`Base address가 64-byte aligned가 아님

vCPU별 stolen-time region 설정 실패입니다.

Structure layout과 guest mapping 규칙은 `Documentation/virt/kvm/arm/pvtime.rst`를 따릅니다.

3. GROUP: KVM_ARM_VCPU_PVTIME_CTRL
==================================

:Architectures: ARM64

3.1 ATTRIBUTE: KVM_ARM_VCPU_PVTIME_IPA
--------------------------------------

:Parameters: 64-bit base address

Returns:

	 =======  ======================================
	 -ENXIO   Stolen time not implemented
	 -EEXIST  Base address already set for this VCPU
	 -EINVAL  Base address not 64 byte aligned
	 =======  ======================================

Specifies the base address of the stolen time structure for this VCPU. The
base address must be 64 byte aligned and exist within a valid guest memory
region. See Documentation/virt/kvm/arm/pvtime.rst for more information
including the layout of the stolen time structure.

x86 TSC offset과 live migration

225-293

`KVM_VCPU_TSC_OFFSET`은 host TSC에 대한 guest의 unsigned 64-bit offset입니다. `guest_tsc = host_tsc + KVM_VCPU_TSC_OFFSET` 관계를 사용하며 parameter access 실패는 `-EFAULT`, 미지원은 `-ENXIO`입니다.

이 attribute를 사용하면 live migration 동안 VM이 pause된 시간까지 guest TSC가 이어지도록 destination의 vCPU별 offset을 조정할 수 있습니다.

Source VMM clock capture
KVM_GET_CLOCK으로 `tsc_src`, kvmclock `guest_src`, CLOCK_REALTIME `host_src` 기록각 vCPU의 KVM_VCPU_TSC_OFFSET을 `ofs_src[i]`로 읽음KVM_GET_TSC_KHZ로 guest TSC frequency `freq` 기록

Pause 시점의 공통 clock state와 vCPU별 offset을 기록합니다.

Destination VMM TSC 복원
KVM_SET_CLOCK에 `guest_src`, `host_src`와 KVM_CLOCK_REALTIME flag 제공KVM_GET_CLOCK으로 `tsc_dest`, `guest_dest` 기록각 vCPU의 `ofs_dst[i]` 계산KVM_VCPU_TSC_OFFSET에 계산값 기록

Elapsed time과 source·destination TSC 차이를 offset에 반영합니다.

TSC offset 계산
의미
`ofs_dst[i] = ofs_src[i] - (guest_src - guest_dest) * freq + (tsc_src - tsc_dest)`Elapsed kvmclock와 두 host TSC origin의 차이를 보정
`ofs[i] + tsc - guest * freq`Kvmclock time 0에 대응하는 guest TSC

Kvmclock time 0에 대응하는 guest TSC가 source와 destination에서 같도록 합니다.

`KVM_SET_CLOCK`은 기록 뒤 지난 시간만큼 kvmclock을 진행시킵니다. Source와 destination의 CLOCK_REALTIME이 동기화되지 않았거나 source pause와 destination step 4~7 사이가 길면 guest timeout 같은 문제가 생길 수 있습니다.

4. GROUP: KVM_VCPU_TSC_CTRL
===========================

:Architectures: x86

4.1 ATTRIBUTE: KVM_VCPU_TSC_OFFSET

:Parameters: 64-bit unsigned TSC offset

Returns:

	 ======= ======================================
	 -EFAULT Error reading/writing the provided
		 parameter address.
	 -ENXIO  Attribute not supported
	 ======= ======================================

Specifies the guest's TSC offset relative to the host's TSC. The guest's
TSC is then derived by the following equation:

  guest_tsc = host_tsc + KVM_VCPU_TSC_OFFSET

This attribute is useful to adjust the guest's TSC on live migration,
so that the TSC counts the time during which the VM was paused. The
following describes a possible algorithm to use for this purpose.

From the source VMM process:

1. Invoke the KVM_GET_CLOCK ioctl to record the host TSC (tsc_src),
   kvmclock nanoseconds (guest_src), and host CLOCK_REALTIME nanoseconds
   (host_src).

2. Read the KVM_VCPU_TSC_OFFSET attribute for every vCPU to record the
   guest TSC offset (ofs_src[i]).

3. Invoke the KVM_GET_TSC_KHZ ioctl to record the frequency of the
   guest's TSC (freq).

From the destination VMM process:

4. Invoke the KVM_SET_CLOCK ioctl, providing the source nanoseconds from
   kvmclock (guest_src) and CLOCK_REALTIME (host_src) in their respective
   fields.  Ensure that the KVM_CLOCK_REALTIME flag is set in the provided
   structure.

   KVM will advance the VM's kvmclock to account for elapsed time since
   recording the clock values.  Note that this will cause problems in
   the guest (e.g., timeouts) unless CLOCK_REALTIME is synchronized
   between the source and destination, and a reasonably short time passes
   between the source pausing the VMs and the destination executing
   steps 4-7.

5. Invoke the KVM_GET_CLOCK ioctl to record the host TSC (tsc_dest) and
   kvmclock nanoseconds (guest_dest).

6. Adjust the guest TSC offsets for every vCPU to account for (1) time
   elapsed since recording state and (2) difference in TSCs between the
   source and destination machine:

   ofs_dst[i] = ofs_src[i] -
     (guest_src - guest_dest) * freq +
     (tsc_src - tsc_dest)

   ("ofs[i] + tsc - guest * freq" is the guest TSC value corresponding to
   a time of 0 in kvmclock.  The above formula ensures that it is the
   same on the destination as it was on the source).

7. Write the KVM_VCPU_TSC_OFFSET attribute for every vCPU with the
   respective value derived in the previous step.