요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
==================================
VDUSE - "vDPA Device in Userspace"
==================================
vDPA (virtio data path acceleration) device is a device that uses a
datapath which complies with the virtio specifications with vendor
specific control path. vDPA devices can be both physically located on
the hardware or emulated by software. VDUSE is a framework that makes it
possible to implement software-emulated vDPA devices in userspace. And
to make the device emulation more secure, the emulated vDPA device's
control path is handled in the kernel and only the data path is
implemented in the userspace.
Note that only virtio block device is supported by VDUSE framework now,
which can reduce security risks when the userspace process that implements
the data path is run by an unprivileged user. The support for other device
types can be added after the security issue of corresponding device driver
is clarified or fixed in the future.
Create/Destroy VDUSE devices
----------------------------
VDUSE devices are created as follows:
1. Create a new VDUSE instance with ioctl(VDUSE_CREATE_DEV) on
/dev/vduse/control.
2. Setup each virtqueue with ioctl(VDUSE_VQ_SETUP) on /dev/vduse/$NAME.
3. Begin processing VDUSE messages from /dev/vduse/$NAME. The first
messages will arrive while attaching the VDUSE instance to vDPA bus.
4. Send the VDPA_CMD_DEV_NEW netlink message to attach the VDUSE
instance to vDPA bus.
VDUSE devices are destroyed as follows:
1. Send the VDPA_CMD_DEV_DEL netlink message to detach the VDUSE
instance from vDPA bus.
2. Close the file descriptor referring to /dev/vduse/$NAME.
3. Destroy the VDUSE instance with ioctl(VDUSE_DESTROY_DEV) on
/dev/vduse/control.
The netlink messages can be sent via vdpa tool in iproute2 or use the
below sample codes:
.. code-block:: c
static int netlink_add_vduse(const char *name, enum vdpa_command cmd)
{
struct nl_sock *nlsock;
struct nl_msg *msg;
int famid;
nlsock = nl_socket_alloc();
if (!nlsock)
return -ENOMEM;
if (genl_connect(nlsock))
goto free_sock;
famid = genl_ctrl_resolve(nlsock, VDPA_GENL_NAME);
if (famid < 0)
goto close_sock;
msg = nlmsg_alloc();
if (!msg)
goto close_sock;
if (!genlmsg_put(msg, NL_AUTO_PORT, NL_AUTO_SEQ, famid, 0, 0, cmd, 0))
goto nla_put_failure;
NLA_PUT_STRING(msg, VDPA_ATTR_DEV_NAME, name);
if (cmd == VDPA_CMD_DEV_NEW)
NLA_PUT_STRING(msg, VDPA_ATTR_MGMTDEV_DEV_NAME, "vduse");
if (nl_send_sync(nlsock, msg))
goto close_sock;
nl_close(nlsock);
nl_socket_free(nlsock);
return 0;
nla_put_failure:
nlmsg_free(msg);
close_sock:
nl_close(nlsock);
free_sock:
nl_socket_free(nlsock);
return -1;
}
How VDUSE works
---------------
As mentioned above, a VDUSE device is created by ioctl(VDUSE_CREATE_DEV) on
/dev/vduse/control. With this ioctl, userspace can specify some basic configuration
such as device name (uniquely identify a VDUSE device), virtio features, virtio
configuration space, the number of virtqueues and so on for this emulated device.
Then a char device interface (/dev/vduse/$NAME) is exported to userspace for device
emulation. Userspace can use the VDUSE_VQ_SETUP ioctl on /dev/vduse/$NAME to
add per-virtqueue configuration such as the max size of virtqueue to the device.
After the initialization, the VDUSE device can be attached to vDPA bus via
the VDPA_CMD_DEV_NEW netlink message. Userspace needs to read()/write() on
/dev/vduse/$NAME to receive/reply some control messages from/to VDUSE kernel
module as follows:
.. code-block:: c
static int vduse_message_handler(int dev_fd)
{
int len;
struct vduse_dev_request req;
struct vduse_dev_response resp;
len = read(dev_fd, &req, sizeof(req));
if (len != sizeof(req))
return -1;
resp.request_id = req.request_id;
switch (req.type) {
/* handle different types of messages */
}
len = write(dev_fd, &resp, sizeof(resp));
if (len != sizeof(resp))
return -1;
return 0;
}
There are now three types of messages introduced by VDUSE framework:
- VDUSE_GET_VQ_STATE: Get the state for virtqueue, userspace should return
avail index for split virtqueue or the device/driver ring wrap counters and
the avail and used index for packed virtqueue.
- VDUSE_SET_STATUS: Set the device status, userspace should follow
the virtio spec: https://docs.oasis-open.org/virtio/virtio/v1.1/virtio-v1.1.html
to process this message. For example, fail to set the FEATURES_OK device
status bit if the device can not accept the negotiated virtio features
get from the VDUSE_DEV_GET_FEATURES ioctl.
- VDUSE_UPDATE_IOTLB: Notify userspace to update the memory mapping for specified
IOVA range, userspace should firstly remove the old mapping, then setup the new
mapping via the VDUSE_IOTLB_GET_FD ioctl.
After DRIVER_OK status bit is set via the VDUSE_SET_STATUS message, userspace is
able to start the dataplane processing as follows:
1. Get the specified virtqueue's information with the VDUSE_VQ_GET_INFO ioctl,
including the size, the IOVAs of descriptor table, available ring and used ring,
the state and the ready status.
2. Pass the above IOVAs to the VDUSE_IOTLB_GET_FD ioctl so that those IOVA regions
can be mapped into userspace. Some sample codes is shown below:
.. code-block:: c
static int perm_to_prot(uint8_t perm)
{
int prot = 0;
switch (perm) {
case VDUSE_ACCESS_WO:
prot |= PROT_WRITE;
break;
case VDUSE_ACCESS_RO:
prot |= PROT_READ;
break;
case VDUSE_ACCESS_RW:
prot |= PROT_READ | PROT_WRITE;
break;
}
return prot;
}
static void *iova_to_va(int dev_fd, uint64_t iova, uint64_t *len)
{
int fd;
void *addr;
size_t size;
struct vduse_iotlb_entry entry;
entry.start = iova;
entry.last = iova;
/*
* Find the first IOVA region that overlaps with the specified
* range [start, last] and return the corresponding file descriptor.
*/
fd = ioctl(dev_fd, VDUSE_IOTLB_GET_FD, &entry);
if (fd < 0)
return NULL;
size = entry.last - entry.start + 1;
*len = entry.last - iova + 1;
addr = mmap(0, size, perm_to_prot(entry.perm), MAP_SHARED,
fd, entry.offset);
close(fd);
if (addr == MAP_FAILED)
return NULL;
/*
* Using some data structures such as linked list to store
* the iotlb mapping. The munmap(2) should be called for the
* cached mapping when the corresponding VDUSE_UPDATE_IOTLB
* message is received or the device is reset.
*/
return addr + iova - entry.start;
}
3. Setup the kick eventfd for the specified virtqueues with the VDUSE_VQ_SETUP_KICKFD
ioctl. The kick eventfd is used by VDUSE kernel module to notify userspace to
consume the available ring. This is optional since userspace can choose to poll the
available ring instead.
4. Listen to the kick eventfd (optional) and consume the available ring. The buffer
described by the descriptors in the descriptor table should be also mapped into
userspace via the VDUSE_IOTLB_GET_FD ioctl before accessing.
5. Inject an interrupt for specific virtqueue with the VDUSE_INJECT_VQ_IRQ ioctl
after the used ring is filled.
For more details on the uAPI, please see include/uapi/linux/vduse.h.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
VDUSE의 목적과 지원 범위
1-19vDPA, 즉 virtio data path acceleration 장치는 vendor 전용 control path와 virtio 명세를 따르는 datapath를 사용합니다. 물리 하드웨어 장치일 수도 있고 software emulation일 수도 있습니다.
VDUSE는 software-emulated vDPA 장치를 사용자 공간에서 구현하게 하는 framework입니다. emulation 보안을 높이기 위해 control path는 커널이 처리하고 사용자 공간은 data path만 구현합니다.
현재 VDUSE framework는 virtio block device만 지원합니다. unprivileged 사용자가 datapath 프로세스를 실행할 때의 보안 위험을 줄이기 위한 제한이며, 다른 장치는 해당 driver 보안 문제가 규명되거나 해결된 뒤 추가할 수 있습니다.
==================================
VDUSE - "vDPA Device in Userspace"
==================================
vDPA (virtio data path acceleration) device is a device that uses a
datapath which complies with the virtio specifications with vendor
specific control path. vDPA devices can be both physically located on
the hardware or emulated by software. VDUSE is a framework that makes it
possible to implement software-emulated vDPA devices in userspace. And
to make the device emulation more secure, the emulated vDPA device's
control path is handled in the kernel and only the data path is
implemented in the userspace.
Note that only virtio block device is supported by VDUSE framework now,
which can reduce security risks when the userspace process that implements
the data path is run by an unprivileged user. The support for other device
types can be added after the security issue of corresponding device driver
is clarified or fixed in the future.
VDUSE 장치 생성과 삭제
20-48control device에서 instance를 만든 뒤 virtqueue와 vDPA bus를 연결합니다.
bus 연결과 열린 FD를 먼저 정리한 뒤 instance를 제거합니다.
netlink message는 iproute2의 `vdpa` 도구로 보내거나 아래 예제와 같은 코드로 보낼 수 있습니다.
Create/Destroy VDUSE devices
----------------------------
VDUSE devices are created as follows:
1. Create a new VDUSE instance with ioctl(VDUSE_CREATE_DEV) on
/dev/vduse/control.
2. Setup each virtqueue with ioctl(VDUSE_VQ_SETUP) on /dev/vduse/$NAME.
3. Begin processing VDUSE messages from /dev/vduse/$NAME. The first
messages will arrive while attaching the VDUSE instance to vDPA bus.
4. Send the VDPA_CMD_DEV_NEW netlink message to attach the VDUSE
instance to vDPA bus.
VDUSE devices are destroyed as follows:
1. Send the VDPA_CMD_DEV_DEL netlink message to detach the VDUSE
instance from vDPA bus.
2. Close the file descriptor referring to /dev/vduse/$NAME.
3. Destroy the VDUSE instance with ioctl(VDUSE_DESTROY_DEV) on
/dev/vduse/control.
The netlink messages can be sent via vdpa tool in iproute2 or use the
below sample codes:
vDPA bus 연결 Netlink 예제
49-94예제 `netlink_add_vduse()`는 Generic Netlink socket을 만들고 `VDPA_GENL_NAME` family ID를 조회한 뒤, 주어진 `VDPA_CMD_DEV_NEW` 또는 `VDPA_CMD_DEV_DEL` 명령과 `VDPA_ATTR_DEV_NAME`을 보냅니다.
새 장치를 추가할 때는 `VDPA_ATTR_MGMTDEV_DEV_NAME`을 `vduse`로 지정합니다. 실패 경로는 message, socket connection과 socket 객체를 순서대로 정리합니다.
static int netlink_add_vduse(const char *name, enum vdpa_command cmd)
{
struct nl_sock *nlsock;
struct nl_msg *msg;
int famid;
nlsock = nl_socket_alloc();
if (!nlsock)
return -ENOMEM;
if (genl_connect(nlsock))
goto free_sock;
famid = genl_ctrl_resolve(nlsock, VDPA_GENL_NAME);
if (famid < 0)
goto close_sock;
msg = nlmsg_alloc();
if (!msg)
goto close_sock;
if (!genlmsg_put(msg, NL_AUTO_PORT, NL_AUTO_SEQ, famid, 0, 0, cmd, 0))
goto nla_put_failure;
NLA_PUT_STRING(msg, VDPA_ATTR_DEV_NAME, name);
if (cmd == VDPA_CMD_DEV_NEW)
NLA_PUT_STRING(msg, VDPA_ATTR_MGMTDEV_DEV_NAME, "vduse");
if (nl_send_sync(nlsock, msg))
goto close_sock;
nl_close(nlsock);
nl_socket_free(nlsock);
return 0;
nla_put_failure:
nlmsg_free(msg);
close_sock:
nl_close(nlsock);
free_sock:
nl_socket_free(nlsock);
return -1;
}
.. code-block:: c
static int netlink_add_vduse(const char *name, enum vdpa_command cmd)
{
struct nl_sock *nlsock;
struct nl_msg *msg;
int famid;
nlsock = nl_socket_alloc();
if (!nlsock)
return -ENOMEM;
if (genl_connect(nlsock))
goto free_sock;
famid = genl_ctrl_resolve(nlsock, VDPA_GENL_NAME);
if (famid < 0)
goto close_sock;
msg = nlmsg_alloc();
if (!msg)
goto close_sock;
if (!genlmsg_put(msg, NL_AUTO_PORT, NL_AUTO_SEQ, famid, 0, 0, cmd, 0))
goto nla_put_failure;
NLA_PUT_STRING(msg, VDPA_ATTR_DEV_NAME, name);
if (cmd == VDPA_CMD_DEV_NEW)
NLA_PUT_STRING(msg, VDPA_ATTR_MGMTDEV_DEV_NAME, "vduse");
if (nl_send_sync(nlsock, msg))
goto close_sock;
nl_close(nlsock);
nl_socket_free(nlsock);
return 0;
nla_put_failure:
nlmsg_free(msg);
close_sock:
nl_close(nlsock);
free_sock:
nl_socket_free(nlsock);
return -1;
}
VDUSE 초기화와 control path
95-110`/dev/vduse/control`의 `ioctl(VDUSE_CREATE_DEV)`로 장치를 만들 때 사용자 공간은 고유 device name, virtio feature, virtio configuration space, virtqueue 수 같은 기본 구성을 지정합니다.
그 뒤 emulation용 문자 장치 `/dev/vduse/$NAME`이 노출됩니다. 이 FD의 `VDUSE_VQ_SETUP` ioctl로 최대 virtqueue 크기 같은 queue별 설정을 장치에 추가합니다.
초기화가 끝나면 `VDPA_CMD_DEV_NEW` netlink message로 vDPA bus에 연결합니다. 사용자 공간은 `/dev/vduse/$NAME`을 `read()`/`write()`해 VDUSE kernel module의 control message를 받고 응답해야 합니다.
How VDUSE works
---------------
As mentioned above, a VDUSE device is created by ioctl(VDUSE_CREATE_DEV) on
/dev/vduse/control. With this ioctl, userspace can specify some basic configuration
such as device name (uniquely identify a VDUSE device), virtio features, virtio
configuration space, the number of virtqueues and so on for this emulated device.
Then a char device interface (/dev/vduse/$NAME) is exported to userspace for device
emulation. Userspace can use the VDUSE_VQ_SETUP ioctl on /dev/vduse/$NAME to
add per-virtqueue configuration such as the max size of virtqueue to the device.
After the initialization, the VDUSE device can be attached to vDPA bus via
the VDPA_CMD_DEV_NEW netlink message. Userspace needs to read()/write() on
/dev/vduse/$NAME to receive/reply some control messages from/to VDUSE kernel
module as follows:
Control message 처리 예제
111-137예제 handler는 `struct vduse_dev_request` 전체를 읽고 같은 `request_id`를 `struct vduse_dev_response`에 복사한 뒤 요청 type별로 처리하고 응답 구조체 전체를 씁니다. 짧은 read 또는 write는 오류로 취급합니다.
static int vduse_message_handler(int dev_fd)
{
int len;
struct vduse_dev_request req;
struct vduse_dev_response resp;
len = read(dev_fd, &req, sizeof(req));
if (len != sizeof(req))
return -1;
resp.request_id = req.request_id;
switch (req.type) {
/* handle different types of messages */
}
len = write(dev_fd, &resp, sizeof(resp));
if (len != sizeof(resp))
return -1;
return 0;
}
request_id가 kernel request와 userspace response를 연결합니다.
.. code-block:: c
static int vduse_message_handler(int dev_fd)
{
int len;
struct vduse_dev_request req;
struct vduse_dev_response resp;
len = read(dev_fd, &req, sizeof(req));
if (len != sizeof(req))
return -1;
resp.request_id = req.request_id;
switch (req.type) {
/* handle different types of messages */
}
len = write(dev_fd, &resp, sizeof(resp));
if (len != sizeof(resp))
return -1;
return 0;
}
VDUSE control message 유형
138-153사용자 공간이 반환하거나 갱신해야 하는 상태가 다릅니다.
There are now three types of messages introduced by VDUSE framework:
- VDUSE_GET_VQ_STATE: Get the state for virtqueue, userspace should return
avail index for split virtqueue or the device/driver ring wrap counters and
the avail and used index for packed virtqueue.
- VDUSE_SET_STATUS: Set the device status, userspace should follow
the virtio spec: https://docs.oasis-open.org/virtio/virtio/v1.1/virtio-v1.1.html
to process this message. For example, fail to set the FEATURES_OK device
status bit if the device can not accept the negotiated virtio features
get from the VDUSE_DEV_GET_FEATURES ioctl.
- VDUSE_UPDATE_IOTLB: Notify userspace to update the memory mapping for specified
IOVA range, userspace should firstly remove the old mapping, then setup the new
mapping via the VDUSE_IOTLB_GET_FD ioctl.
Dataplane 시작: queue 정보와 IOVA
154-163`VDUSE_SET_STATUS` message로 `DRIVER_OK` bit가 설정되면 사용자 공간이 dataplane 처리를 시작할 수 있습니다.
먼저 `VDUSE_VQ_GET_INFO` ioctl로 지정 virtqueue의 크기, descriptor table·available ring·used ring IOVA, 상태와 ready 상태를 얻습니다.
그 IOVA들을 `VDUSE_IOTLB_GET_FD`에 전달해 해당 영역을 사용자 공간에 mapping합니다.
After DRIVER_OK status bit is set via the VDUSE_SET_STATUS message, userspace is
able to start the dataplane processing as follows:
1. Get the specified virtqueue's information with the VDUSE_VQ_GET_INFO ioctl,
including the size, the IOVAs of descriptor table, available ring and used ring,
the state and the ready status.
2. Pass the above IOVAs to the VDUSE_IOTLB_GET_FD ioctl so that those IOVA regions
can be mapped into userspace. Some sample codes is shown below:
IOVA를 사용자 주소로 mapping
164-220`perm_to_prot()`는 VDUSE access permission을 `PROT_READ`와 `PROT_WRITE` 조합으로 바꿉니다.
`iova_to_va()`는 요청 IOVA와 겹치는 첫 IOVA 영역을 `VDUSE_IOTLB_GET_FD`로 찾고, 반환된 `vduse_iotlb_entry`의 범위·permission·offset으로 FD를 mmap합니다. FD는 mapping 뒤 닫고, 최종적으로 요청 IOVA에 해당하는 사용자 가상 주소를 반환합니다.
static int perm_to_prot(uint8_t perm)
{
int prot = 0;
switch (perm) {
case VDUSE_ACCESS_WO:
prot |= PROT_WRITE;
break;
case VDUSE_ACCESS_RO:
prot |= PROT_READ;
break;
case VDUSE_ACCESS_RW:
prot |= PROT_READ | PROT_WRITE;
break;
}
return prot;
}
static void *iova_to_va(int dev_fd, uint64_t iova, uint64_t *len)
{
int fd;
void *addr;
size_t size;
struct vduse_iotlb_entry entry;
entry.start = iova;
entry.last = iova;
/*
* Find the first IOVA region that overlaps with the specified
* range [start, last] and return the corresponding file descriptor.
*/
fd = ioctl(dev_fd, VDUSE_IOTLB_GET_FD, &entry);
if (fd < 0)
return NULL;
size = entry.last - entry.start + 1;
*len = entry.last - iova + 1;
addr = mmap(0, size, perm_to_prot(entry.perm), MAP_SHARED,
fd, entry.offset);
close(fd);
if (addr == MAP_FAILED)
return NULL;
/*
* Using some data structures such as linked list to store
* the iotlb mapping. The munmap(2) should be called for the
* cached mapping when the corresponding VDUSE_UPDATE_IOTLB
* message is received or the device is reset.
*/
return addr + iova - entry.start;
}
실제 구현은 linked list 같은 구조로 IOTLB mapping을 cache할 수 있습니다. 대응 `VDUSE_UPDATE_IOTLB` message를 받거나 장치를 reset하면 cache된 mapping에 `munmap(2)`을 호출해야 합니다.
.. code-block:: c
static int perm_to_prot(uint8_t perm)
{
int prot = 0;
switch (perm) {
case VDUSE_ACCESS_WO:
prot |= PROT_WRITE;
break;
case VDUSE_ACCESS_RO:
prot |= PROT_READ;
break;
case VDUSE_ACCESS_RW:
prot |= PROT_READ | PROT_WRITE;
break;
}
return prot;
}
static void *iova_to_va(int dev_fd, uint64_t iova, uint64_t *len)
{
int fd;
void *addr;
size_t size;
struct vduse_iotlb_entry entry;
entry.start = iova;
entry.last = iova;
/*
* Find the first IOVA region that overlaps with the specified
* range [start, last] and return the corresponding file descriptor.
*/
fd = ioctl(dev_fd, VDUSE_IOTLB_GET_FD, &entry);
if (fd < 0)
return NULL;
size = entry.last - entry.start + 1;
*len = entry.last - iova + 1;
addr = mmap(0, size, perm_to_prot(entry.perm), MAP_SHARED,
fd, entry.offset);
close(fd);
if (addr == MAP_FAILED)
return NULL;
/*
* Using some data structures such as linked list to store
* the iotlb mapping. The munmap(2) should be called for the
* cached mapping when the corresponding VDUSE_UPDATE_IOTLB
* message is received or the device is reset.
*/
return addr + iova - entry.start;
}
Kick, descriptor 소비와 interrupt
221-233`VDUSE_VQ_SETUP_KICKFD` ioctl로 virtqueue의 kick eventfd를 선택적으로 설정합니다. kernel module은 이 eventfd로 사용자 공간에 available ring 소비를 알립니다. 사용자 공간이 available ring을 polling한다면 생략할 수 있습니다.
kick eventfd를 감시하거나 polling해 available ring을 소비합니다. descriptor table이 가리키는 buffer도 접근 전에 `VDUSE_IOTLB_GET_FD`로 사용자 공간에 mapping해야 합니다.
used ring을 채운 뒤 `VDUSE_INJECT_VQ_IRQ` ioctl로 해당 virtqueue의 interrupt를 주입합니다. 자세한 uAPI는 `include/uapi/linux/vduse.h`를 참조하십시오.
queue 정보 획득부터 guest 통지까지의 정상 경로입니다.
3. Setup the kick eventfd for the specified virtqueues with the VDUSE_VQ_SETUP_KICKFD
ioctl. The kick eventfd is used by VDUSE kernel module to notify userspace to
consume the available ring. This is optional since userspace can choose to poll the
available ring instead.
4. Listen to the kick eventfd (optional) and consume the available ring. The buffer
described by the descriptors in the descriptor table should be also mapped into
userspace via the VDUSE_IOTLB_GET_FD ioctl before accessing.
5. Inject an interrupt for specific virtqueue with the VDUSE_INJECT_VQ_IRQ ioctl
after the used ring is filled.
For more details on the uAPI, please see include/uapi/linux/vduse.h.
요약·해설
vduse.rst:1-233VDUSE는 control path를 커널에 남기고 사용자 공간에는 datapath만 맡겨 emulation 경계를 좁힙니다. 안전한 구현은 bus detach와 FD close 순서, request_id 대응, UPDATE_IOTLB 시 오래된 mapping의 munmap, used ring 기록 뒤 interrupt 주입 순서를 지켜야 합니다.