요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
Excerpt from UltraSPARC Virtual Machine Specification
Compiled from version 3.0.20+15
Publication date 2017-09-25 08:21
Copyright © 2008, 2015 Oracle and/or its affiliates. All rights reserved.
Extracted via "pdftotext -f 547 -l 572 -layout sun4v_20170925.pdf"
Authors:
Charles Kunzman
Sam Glidden
Mark Cianchetti
Chapter 36. Coprocessor services
The following APIs provide access via the Hypervisor to hardware assisted data processing functionality.
These APIs may only be provided by certain platforms, and may not be available to all virtual machines
even on supported platforms. Restrictions on the use of these APIs may be imposed in order to support
live-migration and other system management activities.
36.1. Data Analytics Accelerator
The Data Analytics Accelerator (DAX) functionality is a collection of hardware coprocessors that provide
high speed processoring of database-centric operations. The coprocessors may support one or more of
the following data query operations: search, extraction, compression, decompression, and translation. The
functionality offered may vary by virtual machine implementation.
The DAX is a virtual device to sun4v guests, with supported data operations indicated by the virtual device
compatibility property. Functionality is accessed through the submission of Command Control Blocks
(CCBs) via the ccb_submit API function. The operations are processed asynchronously, with the status
of the submitted operations reported through a Completion Area linked to each CCB. Each CCB has a
separate Completion Area and, unless execution order is specifically restricted through the use of serial-
conditional flags, the execution order of submitted CCBs is arbitrary. Likewise, the time to completion
for a given CCB is never guaranteed.
Guest software may implement a software timeout on CCB operations, and if the timeout is exceeded, the
operation may be cancelled or killed via the ccb_kill API function. It is recommended for guest software
to implement a software timeout to account for certain RAS errors which may result in lost CCBs. It is
recommended such implementation use the ccb_info API function to check the status of a CCB prior to
killing it in order to determine if the CCB is still in queue, or may have been lost due to a RAS error.
There is no fixed limit on the number of outstanding CCBs guest software may have queued in the virtual
machine, however, internal resource limitations within the virtual machine can cause CCB submissions
to be temporarily rejected with EWOULDBLOCK. In such cases, guests should continue to attempt
submissions until they succeed; waiting for an outstanding CCB to complete is not necessary, and would
not be a guarantee that a future submission would succeed.
The availability of DAX coprocessor command service is indicated by the presence of the DAX virtual
device node in the guest MD (Section 8.24.17, “Database Analytics Accelerators (DAX) virtual-device
node”).
36.1.1. DAX Compatibility Property
The query functionality may vary based on the compatibility property of the virtual device:
36.1.1.1. "ORCL,sun4v-dax" Device Compatibility
Available CCB commands:
• No-op/Sync
• Extract
• Scan Value
• Inverted Scan Value
• Scan Range
509
Coprocessor services
• Inverted Scan Range
• Translate
• Inverted Translate
• Select
See Section 36.2.1, “Query CCB Command Formats” for the corresponding CCB input and output formats.
Only version 0 CCBs are available.
36.1.1.2. "ORCL,sun4v-dax-fc" Device Compatibility
"ORCL,sun4v-dax-fc" is compatible with the "ORCL,sun4v-dax" interface, and includes additional CCB
bit fields and controls.
36.1.1.3. "ORCL,sun4v-dax2" Device Compatibility
Available CCB commands:
• No-op/Sync
• Extract
• Scan Value
• Inverted Scan Value
• Scan Range
• Inverted Scan Range
• Translate
• Inverted Translate
• Select
See Section 36.2.1, “Query CCB Command Formats” for the corresponding CCB input and output formats.
Version 0 and 1 CCBs are available. Only version 0 CCBs may use Huffman encoded data, whereas only
version 1 CCBs may use OZIP.
36.1.2. DAX Virtual Device Interrupts
The DAX virtual device has multiple interrupts associated with it which may be used by the guest if
desired. The number of device interrupts available to the guest is indicated in the virtual device node of the
guest MD (Section 8.24.17, “Database Analytics Accelerators (DAX) virtual-device node”). If the device
node indicates N interrupts available, the guest may use any value from 0 to N - 1 (inclusive) in a CCB
interrupt number field. Using values outside this range will result in the CCB being rejected for an invalid
field value.
The interrupts may be bound and managed using the standard sun4v device interrupts API (Chapter 16,
Device interrupt services). Sysino interrupts are not available for DAX devices.
36.2. Coprocessor Control Block (CCB)
CCBs are either 64 or 128 bytes long, depending on the operation type. The exact contents of the CCB
are command specific, but all CCBs contain at least one memory buffer address. All memory locations
510
Coprocessor services
referenced by a CCB must be pinned in memory until the CCB either completes execution or is killed
via the ccb_kill API call. Changes in virtual address mappings occurring after CCB submission are not
guaranteed to be visible, and as such all virtual address updates need to be synchronized with CCB
execution.
All CCBs begin with a common 32-bit header.
Table 36.1. CCB Header Format
Bits Field Description
[31:28] CCB version. For API version 2.0: set to 1 if CCB uses OZIP encoding; set to 0 if the CCB
uses Huffman encoding; otherwise either 0 or 1. For API version 1.0: always set to 0.
[27] When API version 2.0 is negotiated, this is the Pipeline Flag [512]. It is reserved in
API version 1.0
[26] Long CCB flag [512]
[25] Conditional synchronization flag [512]
[24] Serial synchronization flag
[23:16] CCB operation code:
0x00 No Operation (No-op) or Sync
0x01 Extract
0x02 Scan Value
0x12 Inverted Scan Value
0x03 Scan Range
0x13 Inverted Scan Range
0x04 Translate
0x14 Inverted Translate
0x05 Select
[15:13] Reserved
[12:11] Table address type
0b'00 No address
0b'01 Alternate context virtual address
0b'10 Real address
0b'11 Primary context virtual address
[10:8] Output/Destination address type
0b'000 No address
0b'001 Alternate context virtual address
0b'010 Real address
0b'011 Primary context virtual address
0b'100 Reserved
0b'101 Reserved
0b'110 Reserved
0b'111 Reserved
[7:5] Secondary source address type
511
Coprocessor services
Bits Field Description
0b'000 No address
0b'001 Alternate context virtual address
0b'010 Real address
0b'011 Primary context virtual address
0b'100 Reserved
0b'101 Reserved
0b'110 Reserved
0b'111 Reserved
[4:2] Primary source address type
0b'000 No address
0b'001 Alternate context virtual address
0b'010 Real address
0b'011 Primary context virtual address
0b'100 Reserved
0b'101 Reserved
0b'110 Reserved
0b'111 Reserved
[1:0] Completion area address type
0b'00 No address
0b'01 Alternate context virtual address
0b'10 Real address
0b'11 Primary context virtual address
The Long CCB flag indicates whether the submitted CCB is 64 or 128 bytes long; value is 0 for 64 bytes
and 1 for 128 bytes.
The Serial and Conditional flags allow simple relative ordering between CCBs. Any CCB with the Serial
flag set will execute sequentially relative to any previous CCB that is also marked as Serial in the same
CCB submission. CCBs without the Serial flag set execute independently, even if they are between CCBs
with the Serial flag set. CCBs marked solely with the Serial flag will execute upon the completion of the
previous Serial CCB, regardless of the completion status of that CCB. The Conditional flag allows CCBs
to conditionally execute based on the successful execution of the closest CCB marked with the Serial flag.
A CCB may only be conditional on exactly one CCB, however, a CCB may be marked both Conditional
and Serial to allow execution chaining. The flags do NOT allow fan-out chaining, where multiple CCBs
execute in parallel based on the completion of another CCB.
The Pipeline flag is an optimization that directs the output of one CCB (the "source" CCB) directly to
the input of the next CCB (the "target" CCB). The target CCB thus does not need to read the input from
memory. The Pipeline flag is advisory and may be dropped.
Both the Pipeline and Serial bits must be set in the source CCB. The Conditional bit must be set in the
target CCB. Exactly one CCB must be made conditional on the source CCB; either 0 or 2 target CCBs
is invalid. However, Pipelines can be extended beyond two CCBs: the sequence would start with a CCB
with both the Pipeline and Serial bits set, proceed through CCBs with the Pipeline, Serial, and Conditional
bits set, and terminate at a CCB that has the Conditional bit set, but not the Pipeline bit.
512
Coprocessor services
The input of the target CCB must start within 64 bytes of the output of the source CCB or the pipeline flag
will be ignored. All CCBs in a pipeline must be submitted in the same call to ccb_submit.
The various address type fields indicate how the various address values used in the CCB should be
interpreted by the virtual machine. Not all of the types specified are used by every CCB format. Types
which are not applicable to the given CCB command should be indicated as type 0 (No address). Virtual
addresses used in the CCB must have translation entries present in either the TLB or a configured TSB
for the submitting virtual processor. Virtual addresses which cannot be translated by the virtual machine
will result in the CCB submission being rejected, with the causal virtual address indicated. The CCB
may be resubmitted after inserting the translation, or the address may be translated by guest software and
resubmitted using the real address translation.
36.2.1. Query CCB Command Formats
36.2.1.1. Supported Data Formats, Elements Sizes and Offsets
Data for query commands may be encoded in multiple possible formats. The data query commands use a
common set of values to indicate the encoding formats of the data being processed. Some encoding formats
require multiple data streams for processing, requiring the specification of both primary data formats (the
encoded data) and secondary data streams (meta-data for the encoded data).
36.2.1.1.1. Primary Input Format
The primary input format code is a 4-bit field when it is used. There are 10 primary input formats available.
The packed formats are not endian neutral. Code values not listed below are reserved.
Code Format Description
0x0 Fixed width byte packed Up to 16 bytes
0x1 Fixed width bit packed Up to 15 bits (CCB version 0) or 23 bits (CCB version
1); bits are read most significant bit to least significant bit
within a byte
0x2 Variable width byte packed Data stream of lengths must be provided as a secondary
input
0x4 Fixed width byte packed with run Up to 16 bytes; data stream of run lengths must be
length encoding provided as a secondary input
0x5 Fixed width bit packed with run Up to 15 bits (CCB version 0) or 23 bits (CCB version
length encoding 1); bits are read most significant bit to least significant bit
within a byte; data stream of run lengths must be provided
as a secondary input
0x8 Fixed width byte packed with Up to 16 bytes before the encoding; compressed stream
Huffman (CCB version 0) or bits are read most significant bit to least significant bit
OZIP (CCB version 1) encoding within a byte; pointer to the encoding table must be
provided
0x9 Fixed width bit packed with Up to 15 bits (CCB version 0) or 23 bits (CCB version
Huffman (CCB version 0) or 1); compressed stream bits are read most significant bit to
OZIP (CCB version 1) encoding least significant bit within a byte; pointer to the encoding
table must be provided
0xA Variable width byte packed with Up to 16 bytes before the encoding; compressed stream
Huffman (CCB version 0) or bits are read most significant bit to least significant bit
OZIP (CCB version 1) encoding within a byte; data stream of lengths must be provided as
a secondary input; pointer to the encoding table must be
provided
513
Coprocessor services
Code Format Description
0xC Fixed width byte packed with Up to 16 bytes before the encoding; compressed stream
run length encoding, followed by bits are read most significant bit to least significant bit
Huffman (CCB version 0) or within a byte; data stream of run lengths must be provided
OZIP (CCB version 1) encoding as a secondary input; pointer to the encoding table must
be provided
0xD Fixed width bit packed with Up to 15 bits (CCB version 0) or 23 bits(CCB version 1)
run length encoding, followed by before the encoding; compressed stream bits are read most
Huffman (CCB version 0) or significant bit to least significant bit within a byte; data
OZIP (CCB version 1) encoding stream of run lengths must be provided as a secondary
input; pointer to the encoding table must be provided
If OZIP encoding is used, there must be no reserved bytes in the table.
36.2.1.1.2. Primary Input Element Size
For primary input data streams with fixed size elements, the element size must be indicated in the CCB
command. The size is encoded as the number of bits or bytes, minus one. The valid value range for this
field depends on the input format selected, as listed in the table above.
36.2.1.1.3. Secondary Input Format
For primary input data streams which require a secondary input stream, the secondary input stream is
always encoded in a fixed width, bit-packed format. The bits are read from most significant bit to least
significant bit within a byte. There are two encoding options for the secondary input stream data elements,
depending on whether the value of 0 is needed:
Secondary Input Description
Format Code
0 Element is stored as value minus 1 (0 evaluates to 1, 1 evaluates
to 2, etc)
1 Element is stored as value
36.2.1.1.4. Secondary Input Element Size
Secondary input element size is encoded as a two bit field:
Secondary Input Size Description
Code
0x0 1 bit
0x1 2 bits
0x2 4 bits
0x3 8 bits
36.2.1.1.5. Input Element Offsets
Bit-wise input data streams may have any alignment within the base addressed byte. The offset, specified
from most significant bit to least significant bit, is provided as a fixed 3 bit field for each input type. A
value of 0 indicates that the first input element begins at the most significant bit in the first byte, and a
value of 7 indicates it begins with the least significant bit.
This field should be zero for any byte-wise primary input data streams.
514
Coprocessor services
36.2.1.1.6. Output Format
Query commands support multiple sizes and encodings for output data streams. There are four possible
output encodings, and up to four supported element sizes per encoding. Not all output encodings are
supported for every command. The format is indicated by a 4-bit field in the CCB:
Output Format Code Description
0x0 Byte aligned, 1 byte elements
0x1 Byte aligned, 2 byte elements
0x2 Byte aligned, 4 byte elements
0x3 Byte aligned, 8 byte elements
0x4 16 byte aligned, 16 byte elements
0x5 Reserved
0x6 Reserved
0x7 Reserved
0x8 Packed vector of single bit elements
0x9 Reserved
0xA Reserved
0xB Reserved
0xC Reserved
0xD 2 byte elements where each element is the index value of a bit,
from an bit vector, which was 1.
0xE 4 byte elements where each element is the index value of a bit,
from an bit vector, which was 1.
0xF Reserved
36.2.1.1.7. Application Data Integrity (ADI)
On platforms which support ADI, the ADI version number may be specified for each separate memory
access type used in the CCB command. ADI checking only occurs when reading data. When writing data,
the specified ADI version number overwrites any existing ADI value in memory.
An ADI version value of 0 or 0xF indicates the ADI checking is disabled for that data access, even if it is
enabled in memory. By setting the appropriate flag in CCB_SUBMIT (Section 36.3.1, “ccb_submit”) it is
also an option to disable ADI checking for all inputs accessed via virtual address for all CCBs submitted
during that hypercall invocation.
The ADI value is only guaranteed to be checked on the first 64 bytes of each data access. Mismatches on
subsequent data chunks may not be detected, so guest software should be careful to use page size checking
to protect against buffer overruns.
36.2.1.1.8. Page size checking
All data accesses used in CCB commands must be bounded within a single memory page. When addresses
are provided using a virtual address, the page size for checking is extracted from the TTE for that virtual
address. When using real addresses, the guest must supply the page size in the same field as the address
value. The page size must be one of the sizes supported by the underlying virtual machine. Using a value
that is not supported may result in the CCB submission being rejected or the generation of a CCB parsing
error in the completion area.
515
Coprocessor services
36.2.1.2. Extract command
Converts an input vector in one format to an output vector in another format. All input format types are
supported.
The only supported output format is a padded, byte-aligned output stream, using output codes 0x0 - 0x4.
When the specified output element size is larger than the extracted input element size, zeros are padded to
the extracted input element. First, if the decompressed input size is not a whole number of bytes, 0 bits are
padded to the most significant bit side till the next byte boundary. Next, if the output element size is larger
than the byte padded input element, bytes of value 0 are added based on the Padding Direction bit in the
CCB. If the output element size is smaller than the byte-padded input element size, the input element is
truncated by dropped from the least significant byte side until the selected output size is reached.
The return value of the CCB completion area is invalid. The “number of elements processed” field in the
CCB completion area will be valid.
The extract CCB is a 64-byte “short format” CCB.
The extract CCB command format can be specified by the following packed C structure for a big-endian
machine:
struct extract_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t reserved;
uint64_t output;
uint64_t table;
};
The exact field offsets, sizes, and composition are as follows:
Offset Size Field Description
0 4 CCB header (Table 36.1, “CCB Header Format”)
4 4 Command control
Bits Field Description
[31:28] Primary Input Format (see Section 36.2.1.1.1, “Primary Input
Format”)
[27:23] Primary Input Element Size (see Section 36.2.1.1.2, “Primary
Input Element Size”)
[22:20] Primary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[19] Secondary Input Format (see Section 36.2.1.1.3, “Secondary
Input Format”)
[18:16] Secondary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
516
Coprocessor services
Offset Size Field Description
Bits Field Description
[15:14] Secondary Input Element Size (see Section 36.2.1.1.4,
“Secondary Input Element Size”
[13:10] Output Format (see Section 36.2.1.1.6, “Output Format”)
[9] Padding Direction selector: A value of 1 causes padding bytes
to be added to the left side of output elements. A value of 0
causes padding bytes to be added to the right side of output
elements.
[8:0] Reserved
8 8 Completion
Bits Field Description
[63:60] ADI version (see Section 36.2.1.1.7, “Application Data
Integrity (ADI)”)
[59] If set to 1, a virtual device interrupt will be generated using
the device interrupt number specified in the lower bits of this
completion word. If 0, the lower bits of this completion word
are ignored.
[58:6] Completion area address bits [58:6]. Address type is
determined by CCB header.
[5:0] Virtual device interrupt number for completion interrupt, if
enabled.
16 8 Primary Input
Bits Field Description
[63:60] ADI version (see Section 36.2.1.1.7, “Application Data
Integrity (ADI)”)
[59:56] If using real address, these bits should be filled in with the
page size code for the page boundary checking the guest wants
the virtual machine to use when accessing this data stream
(checking is only guaranteed to be performed when using API
version 1.1 and later). If using a virtual address, this field will
be used as as primary input address bits [59:56].
[55:0] Primary input address bits [55:0]. Address type is determined
by CCB header.
24 8 Data Access Control
Bits Field Description
[63:62] Flow Control
Value Description
0b'00 Disable flow control
0b'01 Enable flow control (only valid with "ORCL,sun4v-
dax-fc" compatible virtual device variants)
0b'10 Reserved
0b'11 Reserved
[61:60] Reserved (API 1.0)
517
Coprocessor services
Offset Size Field Description
Bits Field Description
Pipeline target (API 2.0)
Value Description
0b'00 Connect to primary input
0b'01 Connect to secondary input
0b'10 Reserved
0b'11 Reserved
[59:40] Output buffer size given in units of 64 bytes, minus 1. Value of
0 means 64 bytes, value of 1 means 128 bytes, etc. Buffer size is
only enforced if flow control is enabled in Flow Control field.
[39:32] Reserved
[31:30] Output Data Cache Allocation
Value Description
0b'00 Do not allocate cache lines for output data stream.
0b'01 Force cache lines for output data stream to be
allocated in the cache that is local to the submitting
virtual cpu.
0b'10 Allocate cache lines for output data stream, but allow
existing cache lines associated with the data to remain
in their current cache instance. Any memory not
already in cache will be allocated in the cache local
to the submitting virtual cpu.
0b'11 Reserved
[29:26] Reserved
[25:24] Primary Input Length Format
Value Description
0b'00 Number of primary symbols
0b'01 Number of primary bytes
0b'10 Number of primary bits
0b'11 Reserved
[23:0] Primary Input Length
Format Field Value
# of primary symbols Number of input elements to process,
minus 1. Command execution stops
once count is reached.
# of primary bytes Number of input bytes to process,
minus 1. Command execution stops
once count is reached. The count is
done before any decompression or
decoding.
# of primary bits Number of input bits to process,
minus 1. Command execution stops
518
Coprocessor services
Offset Size Field Description
Bits Field Description
Format Field Value
once count is reached. The count is
done before any decompression or
decoding, and does not include any
bits skipped by the Primary Input
Offset field value of the command
control word.
32 8 Secondary Input, if used by Primary Input Format. Same fields as Primary
Input.
40 8 Reserved
48 8 Output (same fields as Primary Input)
56 8 Symbol Table (if used by Primary Input)
Bits Field Description
[63:60] ADI version (see Section 36.2.1.1.7, “Application Data
Integrity (ADI)”)
[59:56] If using real address, these bits should be filled in with the
page size code for the page boundary checking the guest wants
the virtual machine to use when accessing this data stream
(checking is only guaranteed to be performed when using API
version 1.1 and later). If using a virtual address, this field will
be used as as symbol table address bits [59:56].
[55:4] Symbol table address bits [55:4]. Address type is determined
by CCB header.
[3:0] Symbol table version
Value Description
0 Huffman encoding. Must use 64 byte aligned table
address. (Only available when using version 0 CCBs)
1 OZIP encoding. Must use 16 byte aligned table
address. (Only available when using version 1 CCBs)
36.2.1.3. Scan commands
The scan commands search a stream of input data elements for values which match the selection criteria.
All the input format types are supported. There are multiple formats for the scan commands, allowing the
scan to search for exact matches to one value, exact matches to either of two values, or any value within
a specified range. The specific type of scan is indicated by the command code in the CCB header. For the
scan range commands, the boundary conditions can be specified as greater-than-or-equal-to a value, less-
than-or-equal-to a value, or both by using two boundary values.
There are two supported formats for the output stream: the bit vector and index array formats (codes 0x8,
0xD, and 0xE). For the standard scan command using the bit vector output, for each input element there
exists one bit in the vector that is set if the input element matched the scan criteria, or clear if not. The
inverted scan command inverts the polarity of the bits in the output. The most significant bit of the first
byte of the output stream corresponds to the first element in the input stream. The standard index array
output format contains one array entry for each input element that matched the scan criteria. Each array
519
Coprocessor services
entry is the index of an input element that matched the scan criteria. An inverted scan command produces
a similar array, but of all the input elements which did NOT match the scan criteria.
The return value of the CCB completion area contains the number of input elements found which match
the scan criteria (or number that did not match for the inverted scans). The “number of elements processed”
field in the CCB completion area will be valid, indicating the number of input elements processed.
These commands are 128-byte “long format” CCBs.
The scan CCB command format can be specified by the following packed C structure for a big-endian
machine:
struct scan_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t match_criteria0;
uint64_t output;
uint64_t table;
uint64_t match_criteria1;
uint64_t match_criteria2;
uint64_t match_criteria3;
uint64_t reserved[5];
};
The exact field offsets, sizes, and composition are as follows:
Offset Size Field Description
0 4 CCB header (Table 36.1, “CCB Header Format”)
4 4 Command control
Bits Field Description
[31:28] Primary Input Format (see Section 36.2.1.1.1, “Primary Input
Format”)
[27:23] Primary Input Element Size (see Section 36.2.1.1.2, “Primary
Input Element Size”)
[22:20] Primary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[19] Secondary Input Format (see Section 36.2.1.1.3, “Secondary
Input Format”)
[18:16] Secondary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[15:14] Secondary Input Element Size (see Section 36.2.1.1.4,
“Secondary Input Element Size”
[13:10] Output Format (see Section 36.2.1.1.6, “Output Format”)
[9:5] Operand size for first scan criteria value. In a scan value
operation, this is one of two potential exact match values.
In a scan range operation, this is the size of the upper range
520
Coprocessor services
Offset Size Field Description
Bits Field Description
boundary. The value of this field is the number of bytes in the
operand, minus 1. Values 0xF-0x1E are reserved. A value of
0x1F indicates this operand is not in use for this scan operation.
[4:0] Operand size for second scan criteria value. In a scan value
operation, this is one of two potential exact match values.
In a scan range operation, this is the size of the lower range
boundary. The value of this field is the number of bytes in the
operand, minus 1. Values 0xF-0x1E are reserved. A value of
0x1F indicates this operand is not in use for this scan operation.
8 8 Completion (same fields as Section 36.2.1.2, “Extract command”)
16 8 Primary Input (same fields as Section 36.2.1.2, “Extract command”)
24 8 Data Access Control (same fields as Section 36.2.1.2, “Extract command”)
32 8 Secondary Input, if used by Primary Input Format. Same fields as Primary
Input.
40 4 Most significant 4 bytes of first scan criteria operand. If first operand is less
than 4 bytes, the value is left-aligned to the lowest address bytes.
44 4 Most significant 4 bytes of second scan criteria operand. If second operand
is less than 4 bytes, the value is left-aligned to the lowest address bytes.
48 8 Output (same fields as Primary Input)
56 8 Symbol Table (if used by Primary Input). Same fields as Section 36.2.1.2,
“Extract command”
64 4 Next 4 most significant bytes of first scan criteria operand occurring after the
bytes specified at offset 40, if needed by the operand size. If first operand
is less than 8 bytes, the valid bytes are left-aligned to the lowest address.
68 4 Next 4 most significant bytes of second scan criteria operand occurring after
the bytes specified at offset 44, if needed by the operand size. If second
operand is less than 8 bytes, the valid bytes are left-aligned to the lowest
address.
72 4 Next 4 most significant bytes of first scan criteria operand occurring after the
bytes specified at offset 64, if needed by the operand size. If first operand
is less than 12 bytes, the valid bytes are left-aligned to the lowest address.
76 4 Next 4 most significant bytes of second scan criteria operand occurring after
the bytes specified at offset 68, if needed by the operand size. If second
operand is less than 12 bytes, the valid bytes are left-aligned to the lowest
address.
80 4 Next 4 most significant bytes of first scan criteria operand occurring after the
bytes specified at offset 72, if needed by the operand size. If first operand
is less than 16 bytes, the valid bytes are left-aligned to the lowest address.
84 4 Next 4 most significant bytes of second scan criteria operand occurring after
the bytes specified at offset 76, if needed by the operand size. If second
operand is less than 16 bytes, the valid bytes are left-aligned to the lowest
address.
521
Coprocessor services
36.2.1.4. Translate commands
The translate commands takes an input array of indices, and a table of single bit values indexed by those
indices, and outputs a bit vector or index array created by reading the tables bit value at each index in
the input array. The output should therefore contain exactly one bit per index in the input data stream,
when outputting as a bit vector. When outputting as an index array, the number of elements depends on the
values read in the bit table, but will always be less than, or equal to, the number of input elements. Only
a restricted subset of the possible input format types are supported. No variable width or Huffman/OZIP
encoded input streams are allowed. The primary input data element size must be 3 bytes or less.
The maximum table index size allowed is 15 bits, however, larger input elements may be used to provide
additional processing of the output values. If 2 or 3 byte values are used, the least significant 15 bits are
used as an index into the bit table. The most significant 9 bits (when using 3-byte input elements) or single
bit (when using 2-byte input elements) are compared against a fixed 9-bit test value provided in the CCB.
If the values match, the value from the bit table is used as the output element value. If the values do not
match, the output data element value is forced to 0.
In the inverted translate operation, the bit value read from bit table is inverted prior to its use. The additional
additional processing based on any additional non-index bits remains unchanged, and still forces the output
element value to 0 on a mismatch. The specific type of translate command is indicated by the command
code in the CCB header.
There are two supported formats for the output stream: the bit vector and index array formats (codes 0x8,
0xD, and 0xE). The index array format is an array of indices of bits which would have been set if the
output format was a bit array.
The return value of the CCB completion area contains the number of bits set in the output bit vector,
or number of elements in the output index array. The “number of elements processed” field in the CCB
completion area will be valid, indicating the number of input elements processed.
These commands are 64-byte “short format” CCBs.
The translate CCB command format can be specified by the following packed C structure for a big-endian
machine:
struct translate_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t reserved;
uint64_t output;
uint64_t table;
};
The exact field offsets, sizes, and composition are as follows:
Offset Size Field Description
0 4 CCB header (Table 36.1, “CCB Header Format”)
522
Coprocessor services
Offset Size Field Description
4 4 Command control
Bits Field Description
[31:28] Primary Input Format (see Section 36.2.1.1.1, “Primary Input
Format”)
[27:23] Primary Input Element Size (see Section 36.2.1.1.2, “Primary
Input Element Size”)
[22:20] Primary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[19] Secondary Input Format (see Section 36.2.1.1.3, “Secondary
Input Format”)
[18:16] Secondary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[15:14] Secondary Input Element Size (see Section 36.2.1.1.4,
“Secondary Input Element Size”
[13:10] Output Format (see Section 36.2.1.1.6, “Output Format”)
[9] Reserved
[8:0] Test value used for comparison against the most significant bits
in the input values, when using 2 or 3 byte input elements.
8 8 Completion (same fields as Section 36.2.1.2, “Extract command”
16 8 Primary Input (same fields as Section 36.2.1.2, “Extract command”
24 8 Data Access Control (same fields as Section 36.2.1.2, “Extract command”,
except Primary Input Length Format may not use the 0x0 value)
32 8 Secondary Input, if used by Primary Input Format. Same fields as Primary
Input.
40 8 Reserved
48 8 Output (same fields as Primary Input)
56 8 Bit Table
Bits Field Description
[63:60] ADI version (see Section 36.2.1.1.7, “Application Data
Integrity (ADI)”)
[59:56] If using real address, these bits should be filled in with the
page size code for the page boundary checking the guest wants
the virtual machine to use when accessing this data stream
(checking is only guaranteed to be performed when using API
version 1.1 and later). If using a virtual address, this field will
be used as as bit table address bits [59:56]
[55:4] Bit table address bits [55:4]. Address type is determined by
CCB header. Address must be 64-byte aligned (CCB version
0) or 16-byte aligned (CCB version 1).
[3:0] Bit table version
Value Description
0 4KB table size
1 8KB table size
523
Coprocessor services
36.2.1.5. Select command
The select command filters the primary input data stream by using a secondary input bit vector to determine
which input elements to include in the output. For each bit set at a given index N within the bit vector,
the Nth input element is included in the output. If the bit is not set, the element is not included. Only a
restricted subset of the possible input format types are supported. No variable width or run length encoded
input streams are allowed, since the secondary input stream is used for the filtering bit vector.
The only supported output format is a padded, byte-aligned output stream. The stream follows the same
rules and restrictions as padded output stream described in Section 36.2.1.2, “Extract command”.
The return value of the CCB completion area contains the number of bits set in the input bit vector. The
"number of elements processed" field in the CCB completion area will be valid, indicating the number
of input elements processed.
The select CCB is a 64-byte “short format” CCB.
The select CCB command format can be specified by the following packed C structure for a big-endian
machine:
struct select_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t reserved;
uint64_t output;
uint64_t table;
};
The exact field offsets, sizes, and composition are as follows:
Offset Size Field Description
0 4 CCB header (Table 36.1, “CCB Header Format”)
4 4 Command control
Bits Field Description
[31:28] Primary Input Format (see Section 36.2.1.1.1, “Primary Input
Format”)
[27:23] Primary Input Element Size (see Section 36.2.1.1.2, “Primary
Input Element Size”)
[22:20] Primary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[19] Secondary Input Format (see Section 36.2.1.1.3, “Secondary
Input Format”)
[18:16] Secondary Input Starting Offset (see Section 36.2.1.1.5, “Input
Element Offsets”)
[15:14] Secondary Input Element Size (see Section 36.2.1.1.4,
“Secondary Input Element Size”
524
Coprocessor services
Offset Size Field Description
Bits Field Description
[13:10] Output Format (see Section 36.2.1.1.6, “Output Format”)
[9] Padding Direction selector: A value of 1 causes padding bytes
to be added to the left side of output elements. A value of 0
causes padding bytes to be added to the right side of output
elements.
[8:0] Reserved
8 8 Completion (same fields as Section 36.2.1.2, “Extract command”
16 8 Primary Input (same fields as Section 36.2.1.2, “Extract command”
24 8 Data Access Control (same fields as Section 36.2.1.2, “Extract command”)
32 8 Secondary Bit Vector Input. Same fields as Primary Input.
40 8 Reserved
48 8 Output (same fields as Primary Input)
56 8 Symbol Table (if used by Primary Input). Same fields as Section 36.2.1.2,
“Extract command”
36.2.1.6. No-op and Sync commands
The no-op (no operation) command is a CCB which has no processing effect. The CCB, when processed
by the virtual machine, simply updates the completion area with its execution status. The CCB may have
the serial-conditional flags set in order to restrict when it executes.
The sync command is a variant of the no-op command which with restricted execution timing. A sync
command CCB will only execute when all previous commands submitted in the same request have
completed. This is stronger than the conditional flag sequencing, which is only dependent on a single
previous serial CCB. While the relative ordering is guaranteed, virtual machine implementations with
shared hardware resources may cause the sync command to wait for longer than the minimum required
time.
The return value of the CCB completion area is invalid for these CCBs. The “number of elements
processed” field is also invalid for these CCBs.
These commands are 64-byte “short format” CCBs.
The no-op CCB command format can be specified by the following packed C structure for a big-endian
machine:
struct nop_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t reserved[6];
};
The exact field offsets, sizes, and composition are as follows:
Offset Size Field Description
0 4 CCB header (Table 36.1, “CCB Header Format”)
525
Coprocessor services
Offset Size Field Description
4 4 Command control
Bits Field Description
[31] If set, this CCB functions as a Sync command. If clear, this
CCB functions as a No-op command.
[30:0] Reserved
8 8 Completion (same fields as Section 36.2.1.2, “Extract command”
16 46 Reserved
36.2.2. CCB Completion Area
All CCB commands use a common 128-byte Completion Area format, which can be specified by the
following packed C structure for a big-endian machine:
struct completion_area {
uint8_t status_flag;
uint8_t error_note;
uint8_t rsvd0[2];
uint32_t error_values;
uint32_t output_size;
uint32_t rsvd1;
uint64_t run_time;
uint64_t run_stats;
uint32_t elements;
uint8_t rsvd2[20];
uint64_t return_value;
uint64_t extra_return_value[8];
};
The Completion Area must be a 128-byte aligned memory location. The exact layout can be described
using byte offsets and sizes relative to the memory base:
Offset Size Field Description
0 1 CCB execution status
0x0 Command not yet completed
0x1 Command ran and succeeded
0x2 Command ran and failed (partial results may be been
produced)
0x3 Command ran and was killed (partial execution may
have occurred)
0x4 Command was not run
0x5-0xF Reserved
1 1 Error reason code
0x0 Reserved
0x1 Buffer overflow
526
Coprocessor services
Offset Size Field Description
0x2 CCB decoding error
0x3 Page overflow
0x4-0x6 Reserved
0x7 Command was killed
0x8 Command execution timeout
0x9 ADI miscompare error
0xA Data format error
0xB-0xD Reserved
0xE Unexpected hardware error (Do not retry)
0xF Unexpected hardware error (Retry is ok)
0x10-0x7F Reserved
0x80 Partial Symbol Warning
0x81-0xFF Reserved
2 2 Reserved
4 4 If a partial symbol warning was generated, this field contains the number
of remaining bits which were not decoded.
8 4 Number of bytes of output produced
12 4 Reserved
16 8 Runtime of command (unspecified time units)
24 8 Reserved
32 4 Number of elements processed
36 20 Reserved
56 8 Return value
64 64 Extended return value
The CCB completion area should be treated as read-only by guest software. The CCB execution status
byte will be cleared by the Hypervisor to reflect the pending execution status when the CCB is submitted
successfully. All other fields are considered invalid upon CCB submission until the CCB execution status
byte becomes non-zero.
CCBs which complete with status 0x2 or 0x3 may produce partial results and/or side effects due to partial
execution of the CCB command. Some valid data may be accessible depending on the fault type, however,
it is recommended that guest software treat the destination buffer as being in an unknown state. If a CCB
completes with a status byte of 0x2, the error reason code byte can be read to determine what corrective
action should be taken.
A buffer overflow indicates that the results of the operation exceeded the size of the output buffer indicated
in the CCB. The operation can be retried by resubmitting the CCB with a larger output buffer.
A CCB decoding error indicates that the CCB contained some invalid field values. It may be also be
triggered if the CCB output is directed at a non-existent secondary input and the pipelining hint is followed.
A page overflow error indicates that the operation required accessing a memory location beyond the page
size associated with a given address. No data will have been read or written past the page boundary, but
partial results may have been written to the destination buffer. The CCB can be resubmitted with a larger
page size memory allocation to complete the operation.
527
Coprocessor services
In the case of pipelined CCBs, a page overflow error will be triggered if the output from the pipeline source
CCB ends before the input of the pipeline target CCB. Page boundaries are ignored when the pipeline
hint is followed.
Command kill indicates that the CCB execution was halted or prevented by use of the ccb_kill API call.
Command timeout indicates that the CCB execution began, but did not complete within a pre-determined
limit set by the virtual machine. The command may have produced some or no output. The CCB may be
resubmitted with no alterations.
ADI miscompare indicates that the memory buffer version specified in the CCB did not match the value
in memory when accessed by the virtual machine. Guest software should not attempt to resubmit the CCB
without determining the cause of the version mismatch.
A data format error indicates that the input data stream did not follow the specified data input formatting
selected in the CCB.
Some CCBs which encounter hardware errors may be resubmitted without change. Persistent hardware
errors may result in multiple failures until RAS software can identify and isolate the faulty component.
The output size field indicates the number of bytes of valid output in the destination buffer. This field is
not valid for all possible CCB commands.
The runtime field indicates the execution time of the CCB command once it leaves the internal virtual
machine queue. The time units are fixed, but unspecified, allowing only relative timing comparisons
by guest software. The time units may also vary by hardware platform, and should not be construed to
represent any absolute time value.
Some data query commands process data in units of elements. If applicable to the command, the number of
elements processed is indicated in the listed field. This field is not valid for all possible CCB commands.
The return value and extended return value fields are output locations for commands which do not use
a destination output buffer, or have secondary return results. The field is not valid for all possible CCB
commands.
36.3. Hypervisor API Functions
36.3.1. ccb_submit
trap# FAST_TRAP
function# CCB_SUBMIT
arg0 address
arg1 length
arg2 flags
arg3 reserved
ret0 status
ret1 length
ret2 status data
ret3 reserved
Submit one or more coprocessor control blocks (CCBs) for evaluation and processing by the virtual
machine. The CCBs are passed in a linear array indicated by address. length indicates the size of
the array in bytes.
528
Coprocessor services
The address should be aligned to the size indicated by length, rounded up to the nearest power of
two. Virtual machines implementations may reject submissions which do not adhere to that alignment.
length must be a multiple of 64 bytes. If length is zero, the maximum supported array length will be
returned as length in ret1. In all other cases, the length value in ret1 will reflect the number of bytes
successfully consumed from the input CCB array.
Implementation note
Virtual machines should never reject submissions based on the alignment of address if the
entire array is contained within a single memory page of the smallest page size supported by the
virtual machine.
A guest may choose to submit addresses used in this API function, including the CCB array address,
as either a real or virtual addresses, with the type of each address indicated in flags. Virtual addresses
must be present in either the TLB or an active TSB to be processed. The translation context for virtual
addresses is determined by a combination of CCB contents and the flags argument.
The flags argument is divided into multiple fields defined as follows:
Bits Field Description
[63:16] Reserved
[15] Disable ADI for VA reads (in API 2.0)
Reserved (in API 1.0)
[14] Virtual addresses within CCBs are translated in privileged context
[13:12] Alternate translation context for virtual addresses within CCBs:
0b'00 CCBs requesting alternate context are rejected
0b'01 Reserved
0b'10 CCBs requesting alternate context use secondary context
0b'11 CCBs requesting alternate context use nucleus context
[11:9] Reserved
[8] Queue info flag
[7] All-or-nothing flag
[6] If address is a virtual address, treat its translation context as privileged
[5:4] Address type of address:
0b'00 Real address
0b'01 Virtual address in primary context
0b'10 Virtual address in secondary context
0b'11 Virtual address in nucleus context
[3:2] Reserved
[1:0] CCB command type:
0b'00 Reserved
0b'01 Reserved
0b'10 Query command
0b'11 Reserved
529
Coprocessor services
The CCB submission type and address type for the CCB array must be provided in the flags argument.
All other fields are optional values which change the default behavior of the CCB processing.
When set to one, the "Disable ADI for VA reads" bit will turn off ADI checking when using a virtual
address to load data. ADI checking will still be done when loading real-addressed memory. This bit is only
available when using major version 2 of the coprocessor API group; at major version 1 it is reserved. For
more information about using ADI and DAX, see Section 36.2.1.1.7, “Application Data Integrity (ADI)”.
By default, all virtual addresses are treated as user addresses. If the virtual address translations are
privileged, they must be marked as such in the appropriate flags field. The virtual addresses used within
the submitted CCBs must all be translated with the same privilege level.
By default, all virtual addresses used within the submitted CCBs are translated using the primary context
active at the time of the submission. The address type field within a CCB allows each address to request
translation in an alternate address context. The address context used when the alternate address context is
requested is selected in the flags argument.
The all-or-nothing flag specifies whether the virtual machine should allow partial submissions of the
input CCB array. When using CCBs with serial-conditional flags, it is strongly recommended to use
the all-or-nothing flag to avoid broken conditional chains. Using long CCB chains on a machine under
high coprocessor load may make this impractical, however, and require submitting without the flag.
When submitting serial-conditional CCBs without the all-or-nothing flag, guest software must manually
implement the serial-conditional behavior at any point where the chain was not submitted in a single API
call, and resubmission of the remaining CCBs should clear any conditional flag that might be set in the
first remaining CCB. Failure to do so will produce indeterminate CCB execution status and ordering.
When the all-or-nothing flag is not specified, callers should check the value of length in ret1 to determine
how many CCBs from the array were successfully submitted. Any remaining CCBs can be resubmitted
without modifications.
The value of length in ret1 is also valid when the API call returns an error, and callers should always
check its value to determine which CCBs in the array were already processed. This will additionally
identify which CCB encountered the processing error, and was not submitted successfully.
If the queue info flag is used during submission, and at least one CCB was successfully submitted, the
length value in ret1 will be a multi-field value defined as follows:
Bits Field Description
[63:48] DAX unit instance identifier
[47:32] DAX queue instance identifier
[31:16] Reserved
[15:0] Number of CCB bytes successfully submitted
The value of status data depends on the status value. See error status code descriptions for details.
The value is undefined for status values that do not specifically list a value for the status data.
The API has a reserved input and output register which will be used in subsequent minor versions of this
API function. Guest software implementations should treat that register as voltile across the function call
in order to maintain forward compatibility.
36.3.1.1. Errors
EOK One or more CCBs have been accepted and enqueued in the virtual machine
and no errors were been encountered during submission. Some submitted
CCBs may not have been enqueued due to internal virtual machine limitations,
and may be resubmitted without changes.
530
Coprocessor services
EWOULDBLOCK An internal resource conflict within the virtual machine has prevented it from
being able to complete the CCB submissions sufficiently quickly, requiring
it to abandon processing before it was complete. Some CCBs may have been
successfully enqueued prior to the block, and all remaining CCBs may be
resubmitted without changes.
EBADALIGN CCB array is not on a 64-byte boundary, or the array length is not a multiple
of 64 bytes.
ENORADDR A real address used either for the CCB array, or within one of the submitted
CCBs, is not valid for the guest. Some CCBs may have been enqueued prior
to the error being detected.
ENOMAP A virtual address used either for the CCB array, or within one of the submitted
CCBs, could not be translated by the virtual machine using either the TLB
or TSB contents. The submission may be retried after adding the required
mapping, or by converting the virtual address into a real address. Due to the
shared nature of address translation resources, there is no theoretical limit on
the number of times the translation may fail, and it is recommended all guests
implement some real address based backup. The virtual address which failed
translation is returned as status data in ret2. Some CCBs may have been
enqueued prior to the error being detected.
EINVAL The virtual machine detected an invalid CCB during submission, or invalid
input arguments, such as bad flag values. Note that not all invalid CCB values
will be detected during submission, and some may be reported as errors in the
completion area instead. Some CCBs may have been enqueued prior to the
error being detected. This error may be returned if the CCB version is invalid.
ETOOMANY The request was submitted with the all-or-nothing flag set, and the array size is
greater than the virtual machine can support in a single request. The maximum
supported size for the current virtual machine can be queried by submitting a
request with a zero length array, as described above.
ENOACCESS The guest does not have permission to submit CCBs, or an address used in a
CCBs lacks sufficient permissions to perform the required operation (no write
permission on the destination buffer address, for example). A virtual address
which fails permission checking is returned as status data in ret2. Some
CCBs may have been enqueued prior to the error being detected.
EUNAVAILABLE The requested CCB operation could not be performed at this time. The
restricted operation availability may apply only to the first unsuccessfully
submitted CCB, or may apply to a larger scope. The status should not be
interpreted as permanent, and the guest should attempt to submit CCBs in
the future which had previously been unable to be performed. The status
data provides additional information about scope of the restricted availability
as follows:
Value Description
0 Processing for the exact CCB instance submitted was unavailable,
and it is recommended the guest emulate the operation. The
guest should continue to submit all other CCBs, and assume no
restrictions beyond this exact CCB instance.
1 Processing is unavailable for all CCBs using the requested opcode,
and it is recommended the guest emulate the operation. The
guest should continue to submit all other CCBs that use different
opcodes, but can expect continued rejections of CCBs using the
same opcode in the near future.
531
Coprocessor services
Value Description
2 Processing is unavailable for all CCBs using the requested CCB
version, and it is recommended the guest emulate the operation.
The guest should continue to submit all other CCBs that use
different CCB versions, but can expect continued rejections of
CCBs using the same CCB version in the near future.
3 Processing is unavailable for all CCBs on the submitting vcpu,
and it is recommended the guest emulate the operation or resubmit
the CCB on a different vcpu. The guest should continue to submit
CCBs on all other vcpus but can expect continued rejections of all
CCBs on this vcpu in the near future.
4 Processing is unavailable for all CCBs, and it is recommended
the guest emulate the operation. The guest should expect all CCB
submissions to be similarly rejected in the near future.
36.3.2. ccb_info
trap# FAST_TRAP
function# CCB_INFO
arg0 address
ret0 status
ret1 CCB state
ret2 position
ret3 dax
ret4 queue
Requests status information on a previously submitted CCB. The previously submitted CCB is identified
by the 64-byte aligned real address of the CCBs completion area.
A CCB can be in one of 4 states:
State Value Description
COMPLETED 0 The CCB has been fetched and executed, and is no longer active in
the virtual machine.
ENQUEUED 1 The requested CCB is current in a queue awaiting execution.
INPROGRESS 2 The CCB has been fetched and is currently being executed. It may still
be possible to stop the execution using the ccb_kill hypercall.
NOTFOUND 3 The CCB could not be located in the virtual machine, and does not
appear to have been executed. This may occur if the CCB was lost
due to a hardware error, or the CCB may not have been successfully
submitted to the virtual machine in the first place.
Implementation note
Some platforms may not be able to report CCBs that are currently being processed, and therefore
guest software should invoke the ccb_kill hypercall prior to assuming the request CCB will never
be executed because it was in the NOTFOUND state.
532
Coprocessor services
The position return value is only valid when the state is ENQUEUED. The value returned is the number
of other CCBs ahead of the requested CCB, to provide a relative estimate of when the CCB may execute.
The dax return value is only valid when the state is ENQUEUED. The value returned is the DAX unit
instance identifier for the DAX unit processing the queue where the requested CCB is located. The value
matches the value that would have been, or was, returned by ccb_submit using the queue info flag.
The queue return value is only valid when the state is ENQUEUED. The value returned is the DAX
queue instance identifier for the DAX unit processing the queue where the requested CCB is located. The
value matches the value that would have been, or was, returned by ccb_submit using the queue info flag.
36.3.2.1. Errors
EOK The request was processed and the CCB state is valid.
EBADALIGN address is not on a 64-byte aligned.
ENORADDR The real address provided for address is not valid.
EINVAL The CCB completion area contents are not valid.
EWOULDBLOCK Internal resource constraints prevented the CCB state from being queried at this
time. The guest should retry the request.
ENOACCESS The guest does not have permission to access the coprocessor virtual device
functionality.
36.3.3. ccb_kill
trap# FAST_TRAP
function# CCB_KILL
arg0 address
ret0 status
ret1 result
Request to stop execution of a previously submitted CCB. The previously submitted CCB is identified by
the 64-byte aligned real address of the CCBs completion area.
The kill attempt can produce one of several values in the result return value, reflecting the CCB state
and actions taken by the Hypervisor:
Result Value Description
COMPLETED 0 The CCB has been fetched and executed, and is no longer active in
the virtual machine. It could not be killed and no action was taken.
DEQUEUED 1 The requested CCB was still enqueued when the kill request was
submitted, and has been removed from the queue. Since the CCB
never began execution, no memory modifications were produced by
it, and the completion area will never be updated. The same CCB may
be submitted again, if desired, with no modifications required.
KILLED 2 The CCB had been fetched and was being executed when the kill
request was submitted. The CCB execution was stopped, and the CCB
is no longer active in the virtual machine. The CCB completion area
will reflect the killed status, with the subsequent implications that
partial results may have been produced. Partial results may include full
533
Coprocessor services
Result Value Description
command execution if the command was stopped just prior to writing
to the completion area.
NOTFOUND 3 The CCB could not be located in the virtual machine, and does not
appear to have been executed. This may occur if the CCB was lost
due to a hardware error, or the CCB may not have been successfully
submitted to the virtual machine in the first place. CCBs in the state
are guaranteed to never execute in the future unless resubmitted.
36.3.3.1. Interactions with Pipelined CCBs
If the pipeline target CCB is killed but the pipeline source CCB was skipped, the completion area of the
target CCB may contain status (4,0) "Command was skipped" instead of (3,7) "Command was killed".
If the pipeline source CCB is killed, the pipeline target CCB's completion status may read (1,0) "Success".
This does not mean the target CCB was processed; since the source CCB was killed, there was no
meaningful output on which the target CCB could operate.
36.3.3.2. Errors
EOK The request was processed and the result is valid.
EBADALIGN address is not on a 64-byte aligned.
ENORADDR The real address provided for address is not valid.
EINVAL The CCB completion area contents are not valid.
EWOULDBLOCK Internal resource constraints prevented the CCB from being killed at this time.
The guest should retry the request.
ENOACCESS The guest does not have permission to access the coprocessor virtual device
functionality.
36.3.4. dax_info
trap# FAST_TRAP
function# DAX_INFO
ret0 status
ret1 Number of enabled DAX units
ret2 Number of disabled DAX units
Returns the number of DAX units that are enabled for the calling guest to submit CCBs. The number of
DAX units that are disabled for the calling guest are also returned. A disabled DAX unit would have been
available for CCB submission to the calling guest had it not been offlined.
36.3.4.1. Errors
EOK The request was processed and the number of enabled/disabled DAX units
are valid.
534
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
문서 출처와 coprocessor service
1-17이 문서는 UltraSPARC Virtual Machine Specification version `3.0.20+15`의 발췌본이며 2017-09-25에 발행되었습니다. Oracle 문서 `sun4v_20170925.pdf`의 547-572쪽을 `pdftotext -layout`으로 추출했고 Charles Kunzman, Sam Glidden, Mark Cianchetti가 저자입니다.
Chapter 36의 API는 hypervisor를 통해 hardware-assisted data processing 기능에 접근하게 합니다. 일부 platform에서만 제공될 수 있고 지원 platform에서도 모든 virtual machine에 열리지 않을 수 있습니다. live migration과 기타 system management를 지원하기 위해 사용 제한이 적용될 수 있습니다.
Data Analytics Accelerator 개요
18-47Data Analytics Accelerator(DAX)는 database 중심 연산을 고속 처리하는 hardware coprocessor 집합입니다. 구현에 따라 search, extraction, compression, decompression, translation 중 하나 이상을 지원합니다.
sun4v guest에는 DAX가 virtual device로 보이며 compatibility property가 지원 연산을 나타냅니다. guest는 `ccb_submit` API로 Command Control Block(CCB)을 제출합니다. 연산은 비동기로 처리되고 각 CCB에 연결된 별도 Completion Area에 상태가 기록됩니다. serial-conditional flag로 순서를 제한하지 않으면 CCB 실행 순서는 임의이며 완료 시간도 보장되지 않습니다.
guest software는 CCB timeout을 구현하고 초과 시 `ccb_kill`로 취소할 수 있습니다. RAS error로 CCB가 유실될 수 있으므로 timeout을 두는 것이 권장되며, kill 전에 `ccb_info`로 CCB가 queue에 남아 있는지 또는 RAS error로 사라졌는지 확인하는 것이 좋습니다.
outstanding CCB 수에는 고정 상한이 없지만 virtual machine 내부 resource가 부족하면 제출이 `EWOULDBLOCK`으로 일시 거부될 수 있습니다. 이때 성공할 때까지 다시 제출해야 하며 기존 CCB 완료를 기다릴 필요도, 기다린 뒤 성공한다는 보장도 없습니다. guest MD에 DAX virtual-device node가 있으면 command service를 사용할 수 있습니다.
sun4v-dax와 sun4v-dax-fc compatibility
48-84virtual device compatibility property에 따라 query 기능이 달라집니다. `ORCL,sun4v-dax`는 다음 CCB command를 제공합니다.
- No-op/Sync
- Extract
- Scan Value
- Inverted Scan Value
- Scan Range
- Inverted Scan Range
- Translate
- Inverted Translate
- Select
각 command의 input/output format은 Section 36.2.1의 Query CCB Command Formats를 따릅니다. `ORCL,sun4v-dax`에서는 version 0 CCB만 사용할 수 있습니다.
`ORCL,sun4v-dax-fc`는 `ORCL,sun4v-dax` interface와 호환되며 추가 CCB bit field와 control을 포함합니다.
sun4v-dax2와 virtual device interrupt
85-121`ORCL,sun4v-dax2`도 No-op/Sync, Extract, 두 Scan Value, 두 Scan Range, 두 Translate, Select command를 제공합니다. version 0과 1 CCB를 사용할 수 있으며 Huffman encoded data는 version 0만, OZIP은 version 1만 사용할 수 있습니다.
DAX virtual device에는 guest가 선택적으로 사용할 수 있는 여러 interrupt가 있습니다. guest MD의 virtual device node가 interrupt `N`개를 알리면 CCB interrupt number에는 `0`부터 `N - 1`까지 넣을 수 있습니다. 범위를 벗어나면 invalid field value로 CCB가 거부됩니다.
interrupt는 표준 sun4v device interrupt API로 bind하고 관리합니다. DAX device에는 Sysino interrupt를 사용할 수 없습니다.
CCB 공통 header와 실행 순서
122-242CCB는 operation type에 따라 64 또는 128 byte입니다. 내용은 command별로 다르지만 최소 하나의 memory buffer address를 포함합니다. 참조하는 모든 memory location은 CCB 실행이 끝나거나 `ccb_kill`로 종료될 때까지 pin해야 합니다. 제출 뒤 virtual mapping 변경은 보이지 않을 수 있으므로 address update를 CCB 실행과 동기화해야 합니다.
모든 CCB는 다음 공통 32-bit header로 시작합니다.
| bit | field와 값 |
|---|---|
| `[31:28]` | CCB version. API 2.0에서는 OZIP이면 1, Huffman이면 0, 그 밖에는 0 또는 1입니다. API 1.0에서는 항상 0입니다. |
| `[27]` | API 2.0의 Pipeline flag이며 API 1.0에서는 reserved입니다. |
| `[26]` | Long CCB flag입니다. |
| `[25]` | Conditional synchronization flag입니다. |
| `[24]` | Serial synchronization flag입니다. |
| `[23:16]` | opcode: `0x00` No-op/Sync, `0x01` Extract, `0x02` Scan Value, `0x12` Inverted Scan Value, `0x03` Scan Range, `0x13` Inverted Scan Range, `0x04` Translate, `0x14` Inverted Translate, `0x05` Select. |
| `[15:13]` | reserved입니다. |
| `[12:11]` | Table address type: `0b'00` no address, `0b'01` alternate-context VA, `0b'10` real address, `0b'11` primary-context VA. |
| `[10:8]` | Output/Destination address type: `000` no address, `001` alternate-context VA, `010` real address, `011` primary-context VA, `100-111` reserved. |
| `[7:5]` | Secondary source address type: `000` no address, `001` alternate-context VA, `010` real address, `011` primary-context VA, `100-111` reserved. |
| `[4:2]` | Primary source address type: `000` no address, `001` alternate-context VA, `010` real address, `011` primary-context VA, `100-111` reserved. |
| `[1:0]` | Completion Area address type: `00` no address, `01` alternate-context VA, `10` real address, `11` primary-context VA. |
Long flag가 0이면 CCB는 64 byte, 1이면 128 byte입니다. Serial CCB는 같은 submission에서 앞선 Serial CCB와 순차 실행되며 중간의 non-Serial CCB는 독립 실행됩니다. Conditional CCB는 가장 가까운 앞선 Serial CCB의 성공 여부에 의존합니다. 한 CCB는 정확히 하나의 CCB에만 conditional일 수 있고 Serial과 Conditional을 함께 써 chain을 만들 수 있지만 fan-out chain은 지원하지 않습니다.
Pipeline flag는 source CCB의 output을 다음 target CCB input으로 직접 전달해 memory read를 줄이는 advisory optimization입니다. source에는 Pipeline과 Serial을 모두, target에는 Conditional을 설정해야 하며 source에 conditional인 target은 정확히 하나여야 합니다. 긴 pipeline은 Pipeline+Serial source로 시작하고, Pipeline+Serial+Conditional 중간 CCB를 거쳐 Pipeline 없이 Conditional만 설정한 CCB로 끝납니다.
target input은 source output 시작점에서 64 byte 이내여야 하며 모든 pipeline CCB를 같은 `ccb_submit` 호출로 제출해야 합니다. 적용되지 않는 address type은 0(no address)으로 표시합니다. CCB의 VA는 제출 virtual processor의 TLB 또는 configured TSB에 translation entry가 있어야 하며, 실패한 VA를 mapping한 뒤 재제출하거나 guest가 real address로 변환해 제출할 수 있습니다.
Primary input format
243-300query command는 여러 encoding format을 공통 code로 나타냅니다. 일부 format은 encoded primary stream과 metadata secondary stream을 함께 요구합니다. primary input format code는 사용 시 4-bit이며 10개 format이 정의됩니다. packed format은 endian-neutral하지 않고 아래에 없는 code는 reserved입니다.
| code | format | 설명 |
|---|---|---|
| `0x0` | Fixed-width byte packed | 최대 16 byte입니다. |
| `0x1` | Fixed-width bit packed | CCB v0은 최대 15 bit, v1은 최대 23 bit입니다. byte 안에서 MSB부터 LSB 순으로 읽습니다. |
| `0x2` | Variable-width byte packed | length stream을 secondary input으로 제공해야 합니다. |
| `0x4` | Fixed-width byte packed + run-length | 최대 16 byte이며 run length stream이 secondary input으로 필요합니다. |
| `0x5` | Fixed-width bit packed + run-length | v0은 최대 15 bit, v1은 최대 23 bit이며 MSB부터 읽습니다. run length stream이 필요합니다. |
| `0x8` | Fixed-width byte packed + Huffman(v0) 또는 OZIP(v1) | encoding 전 최대 16 byte, compressed bit는 MSB부터 읽고 encoding table pointer가 필요합니다. |
| `0x9` | Fixed-width bit packed + Huffman(v0) 또는 OZIP(v1) | v0은 최대 15 bit, v1은 최대 23 bit이며 compressed bit는 MSB부터 읽고 table pointer가 필요합니다. |
| `0xA` | Variable-width byte packed + Huffman(v0) 또는 OZIP(v1) | encoding 전 최대 16 byte입니다. length secondary stream과 encoding table pointer가 필요합니다. |
| `0xC` | Fixed-width byte packed + run-length + Huffman(v0) 또는 OZIP(v1) | encoding 전 최대 16 byte입니다. run length secondary stream과 table pointer가 필요합니다. |
| `0xD` | Fixed-width bit packed + run-length + Huffman(v0) 또는 OZIP(v1) | encoding 전 v0은 최대 15 bit, v1은 최대 23 bit입니다. run length secondary stream과 table pointer가 필요합니다. |
OZIP encoding table에는 reserved byte가 없어야 합니다.
element size, secondary input과 offset
301-344fixed-size primary element의 크기는 CCB command에 실제 bit 또는 byte 수에서 1을 뺀 값으로 encode합니다. 허용 범위는 선택한 input format에 따라 달라집니다.
secondary stream이 필요한 primary format에서 secondary element는 항상 fixed-width bit-packed이며 byte 안에서 MSB부터 LSB 순으로 읽습니다.
| secondary format code | 설명 |
|---|---|
| `0` | element를 `value - 1`로 저장합니다. 저장값 0은 1, 1은 2로 평가됩니다. |
| `1` | element를 실제 value 그대로 저장합니다. |
| secondary size code | 크기 |
|---|---|
| `0x0` | 1 bit |
| `0x1` | 2 bit |
| `0x2` | 4 bit |
| `0x3` | 8 bit |
bit-wise input stream은 base byte 안에서 어떤 alignment도 가질 수 있습니다. 각 input type의 시작 offset은 MSB에서 LSB 방향으로 센 3-bit field입니다. 0은 첫 byte의 MSB, 7은 LSB에서 첫 element가 시작함을 뜻합니다. byte-wise primary stream에서는 0이어야 합니다.
output format, ADI와 page 검사
345-399query command의 output format은 4-bit field로 나타내며 command별로 일부 encoding만 지원할 수 있습니다.
| code | output format |
|---|---|
| `0x0` | byte-aligned 1-byte element |
| `0x1` | byte-aligned 2-byte element |
| `0x2` | byte-aligned 4-byte element |
| `0x3` | byte-aligned 8-byte element |
| `0x4` | 16-byte aligned 16-byte element |
| `0x5-0x7` | reserved |
| `0x8` | single-bit element의 packed vector |
| `0x9-0xC` | reserved |
| `0xD` | bit vector에서 값이 1인 bit의 index를 담는 2-byte element |
| `0xE` | bit vector에서 값이 1인 bit의 index를 담는 4-byte element |
| `0xF` | reserved |
ADI 지원 platform에서는 CCB의 memory access type마다 ADI version을 지정할 수 있습니다. read에서만 ADI를 검사하며 write에서는 지정 version이 memory의 기존 값을 덮어씁니다. `0` 또는 `0xF`는 해당 access의 ADI 검사를 끕니다. `CCB_SUBMIT` flag로 한 hypercall에 제출한 모든 CCB의 VA input 검사를 끌 수도 있습니다.
ADI 검사는 각 data access의 첫 64 byte에서만 보장됩니다. 뒤쪽 mismatch는 놓칠 수 있으므로 buffer overrun 방지에는 Page size checking을 사용해야 합니다. 모든 CCB data access는 한 page 안에 있어야 합니다. VA는 TTE에서 page size를 얻고 real address는 guest가 address field에 page size를 제공합니다. VM이 지원하지 않는 size는 submission reject 또는 Completion Area parsing error를 일으킬 수 있습니다.
Extract command 형식
400-589Extract는 한 format의 input vector를 다른 format의 output vector로 바꾸며 모든 input format을 지원합니다. output은 code `0x0-0x4`의 padded byte-aligned stream만 지원합니다. decompressed input이 byte 경계가 아니면 MSB 쪽에 0 bit를 채우고, output element가 더 크면 Padding Direction에 따라 왼쪽 또는 오른쪽에 0 byte를 넣습니다. output이 더 작으면 원하는 크기까지 LSB 쪽 byte를 버립니다.
Completion Area의 return value는 invalid이고 number of elements processed는 valid입니다. Extract는 64-byte short-format CCB입니다.
struct extract_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t reserved;
uint64_t output;
uint64_t table;
};
| offset/size | field와 세부 구성 |
|---|---|
| `0/4` | 공통 CCB header입니다. |
| `4/4` | Command control: `[31:28]` primary format, `[27:23]` primary size, `[22:20]` primary offset, `[19]` secondary format, `[18:16]` secondary offset, `[15:14]` secondary size, `[13:10]` output format, `[9]` padding direction(1=left, 0=right), `[8:0]` reserved. |
| `8/8` | Completion: `[63:60]` ADI, `[59]` interrupt enable, `[58:6]` Completion Area address, `[5:0]` virtual device interrupt number. |
| `16/8` | Primary Input: `[63:60]` ADI, `[59:56]` real-address page-size code 또는 VA address bits, `[55:0]` address. |
| `24/8` | Data Access Control: `[63:62]` flow control(00 off, 01 on for dax-fc, 10/11 reserved), `[61:60]` API 1.0 reserved/API 2.0 pipeline target(00 primary, 01 secondary), `[59:40]` 64-byte unit output size minus 1, `[39:32]` reserved, `[31:30]` output cache allocation, `[29:26]` reserved, `[25:24]` input length format, `[23:0]` input length. |
| `[31:30]` cache allocation | Output Data Cache Allocation 값은 `00` no allocation, `01` submitting vCPU local cache에 강제 allocation, `10` 기존 cache location을 유지하되 uncached memory는 local cache에 allocation, `11` reserved입니다. |
| `[25:24]` length format | `00` primary symbol 수, `01` encoded/decoded 전 primary byte 수, `10` primary bit 수, `11` reserved. field value는 각 count에서 1을 뺀 값이며 bit count는 starting offset으로 건너뛴 bit를 포함하지 않습니다. |
| `32/8` | Secondary Input. 사용 시 Primary Input과 같은 field입니다. |
| `40/8` | reserved입니다. |
| `48/8` | Output. Primary Input과 같은 address field입니다. |
| `56/8` | Symbol Table: `[63:60]` ADI, `[59:56]` real page-size 또는 VA bits, `[55:4]` table address, `[3:0]` version. version 0은 64-byte aligned Huffman(v0 CCB), version 1은 16-byte aligned OZIP(v1 CCB)입니다. |
Scan command 형식
590-721Scan command는 input element stream에서 selection criteria와 일치하는 값을 찾습니다. 모든 input format을 지원하며 한 값 또는 두 값과의 exact match, 지정 range 검색을 제공합니다. Range boundary는 lower/upper 중 하나 또는 둘 다 지정할 수 있습니다.
output은 bit vector 또는 index array(code `0x8`, `0xD`, `0xE`)입니다. 일반 bit-vector scan은 일치하면 bit를 1로 두고 inverted scan은 polarity를 뒤집습니다. 첫 output byte의 MSB가 첫 input element에 대응합니다. index array는 일치한 input index를, inverted scan은 일치하지 않은 index를 담습니다.
Completion return value는 일치 element 수이며 inverted scan에서는 불일치 수입니다. number of elements processed도 valid합니다. Scan은 128-byte long-format CCB입니다.
struct scan_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t match_criteria0;
uint64_t output;
uint64_t table;
uint64_t match_criteria1;
uint64_t match_criteria2;
uint64_t match_criteria3;
uint64_t reserved[5];
};
| offset/size | field |
|---|---|
| `0/4` | 공통 CCB header입니다. |
| `4/4` | Command control의 `[31:10]`은 Extract와 같은 primary/secondary/output format field입니다. `[9:5]`는 첫 criteria operand size minus 1, `[4:0]`는 둘째 operand size minus 1입니다. Scan Value에서는 두 exact value, Scan Range에서는 각각 upper/lower boundary입니다. `0xF-0x1E`는 reserved, `0x1F`는 operand 미사용입니다. |
| `8/8` | Completion. Extract와 같습니다. |
| `16/8` | Primary Input. Extract와 같습니다. |
| `24/8` | Data Access Control. Extract와 같습니다. |
| `32/8` | Secondary Input. 사용 시 Primary Input과 같습니다. |
| `40/4`, `44/4` | 첫째와 둘째 scan operand의 최상위 4 byte입니다. operand가 더 짧으면 low-address 쪽에 left-align합니다. |
| `48/8` | Output. Primary Input과 같습니다. |
| `56/8` | Symbol Table. Extract와 같습니다. |
| `64/4`, `68/4` | 각 operand의 다음 4 byte입니다. 8 byte 미만이면 valid byte를 low-address 쪽에 left-align합니다. |
| `72/4`, `76/4` | 각 operand의 다음 4 byte입니다. 12 byte 미만이면 valid byte를 low-address 쪽에 left-align합니다. |
| `80/4`, `84/4` | 각 operand의 마지막 4 byte입니다. 16 byte 미만이면 valid byte를 low-address 쪽에 left-align합니다. |
Translate command 형식
722-832Translate는 index array와 그 index로 조회하는 single-bit table을 받아, 각 input index의 table bit를 읽은 bit vector 또는 index array를 만듭니다. bit vector는 input index마다 정확히 1 bit가 있고 index array의 element 수는 table 값에 따라 input 수 이하입니다.
variable-width와 Huffman/OZIP input은 허용하지 않으며 primary element는 최대 3 byte입니다. table index는 최대 15 bit입니다. 2/3-byte input에서는 하위 15 bit를 index로 쓰고, 상위 1 bit 또는 9 bit를 CCB의 고정 9-bit test value와 비교합니다. 일치하면 table bit, 불일치하면 0을 output으로 씁니다. Inverted Translate는 table bit만 뒤집고 상위 bit mismatch의 0 강제는 그대로입니다.
output은 bit vector 또는 index array(code `0x8`, `0xD`, `0xE`)입니다. return value는 output의 set bit 수 또는 index-array element 수이며 number of elements processed도 valid합니다. 64-byte short-format CCB입니다.
struct translate_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t reserved;
uint64_t output;
uint64_t table;
};
| offset/size | field |
|---|---|
| `0/4` | 공통 CCB header입니다. |
| `4/4` | Command control의 `[31:10]`은 공통 format field, `[9]` reserved, `[8:0]`은 2/3-byte input의 상위 bit와 비교할 9-bit test value입니다. |
| `8/8` | Completion. Extract와 같습니다. |
| `16/8` | Primary Input. Extract와 같습니다. |
| `24/8` | Data Access Control. Extract와 같지만 Primary Input Length Format에 `0x0`을 사용할 수 없습니다. |
| `32/8` | Secondary Input. 사용 시 Primary Input과 같습니다. |
| `40/8` | reserved입니다. |
| `48/8` | Output. Primary Input과 같습니다. |
| `56/8` | Bit Table: `[63:60]` ADI, `[59:56]` real page-size 또는 VA bits, `[55:4]` table address(그리고 v0은 64-byte, v1은 16-byte aligned), `[3:0]` table version. version 0은 4KB, 1은 8KB table입니다. |
Select command 형식
833-906Select는 secondary input bit vector로 primary input stream을 filter합니다. bit vector의 index `N`이 1이면 N번째 input element를 output에 포함하고 0이면 제외합니다. secondary stream을 filter vector로 쓰므로 variable-width와 run-length input은 허용하지 않습니다.
output은 Extract와 같은 padded byte-aligned stream만 지원합니다. Completion return value는 input bit vector의 set bit 수이며 number of elements processed도 valid합니다. Select는 64-byte short-format CCB입니다.
struct select_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t primary_input;
uint64_t data_access_control;
uint64_t secondary_input;
uint64_t reserved;
uint64_t output;
uint64_t table;
};
| offset/size | field |
|---|---|
| `0/4` | 공통 CCB header입니다. |
| `4/4` | Extract와 같은 primary/secondary/output format, element size, starting offset, padding direction field를 사용하며 `[8:0]`은 reserved입니다. |
| `8/8` | Completion. Extract와 같습니다. |
| `16/8` | Primary Input. Extract와 같습니다. |
| `24/8` | Data Access Control. Extract와 같습니다. |
| `32/8` | Secondary Bit Vector Input. Primary Input과 같은 address field입니다. |
| `40/8` | reserved입니다. |
| `48/8` | Output. Primary Input과 같습니다. |
| `56/8` | Primary format이 사용할 경우 Symbol Table. Extract와 같습니다. |
No-op과 Sync command
907-954No-op CCB는 processing effect 없이 virtual machine이 Completion Area에 execution status만 기록합니다. 실행 시점을 제한하기 위해 serial-conditional flag를 설정할 수 있습니다.
Sync는 같은 request로 앞서 제출된 모든 command가 완료된 뒤에만 실행되는 No-op 변형입니다. 하나의 앞선 Serial CCB에만 의존하는 conditional sequencing보다 강합니다. shared hardware resource가 있는 구현에서는 최소 필요 시간보다 더 오래 기다릴 수 있습니다.
No-op과 Sync에서는 Completion return value와 number of elements processed가 모두 invalid이며 64-byte short-format CCB입니다.
struct nop_ccb {
uint32_t header;
uint32_t control;
uint64_t completion;
uint64_t reserved[6];
};
| offset/size | field |
|---|---|
| `0/4` | 공통 CCB header입니다. |
| `4/4` | Command control: `[31]`이 1이면 Sync, 0이면 No-op이며 `[30:0]`은 reserved입니다. |
| `8/8` | Completion. Extract와 같습니다. |
| `16/46` | reserved입니다. |
CCB Completion Area layout
955-1023모든 CCB는 공통 128-byte Completion Area를 사용하며 big-endian packed C structure는 다음과 같습니다.
struct completion_area {
uint8_t status_flag;
uint8_t error_note;
uint8_t rsvd0[2];
uint32_t error_values;
uint32_t output_size;
uint32_t rsvd1;
uint64_t run_time;
uint64_t run_stats;
uint32_t elements;
uint8_t rsvd2[20];
uint64_t return_value;
uint64_t extra_return_value[8];
};
Completion Area는 128-byte aligned memory location이어야 합니다.
| offset/size | field와 값 |
|---|---|
| `0/1` | execution status: `0x0` 미완료, `0x1` 성공, `0x2` 실패(부분 결과 가능), `0x3` killed(부분 실행 가능), `0x4` 미실행, `0x5-0xF` reserved. |
| `1/1` | error reason: `0x0` reserved, `0x1` buffer overflow, `0x2` CCB decode error, `0x3` page overflow, `0x4-0x6` reserved, `0x7` killed, `0x8` timeout, `0x9` ADI miscompare, `0xA` data format, `0xB-0xD` reserved, `0xE` unexpected hardware error(no retry), `0xF` unexpected hardware error(retry okay), `0x10-0x7F` reserved, `0x80` Partial Symbol Warning, `0x81-0xFF` reserved. |
| `2/2` | reserved입니다. |
| `4/4` | Partial Symbol Warning이면 decode되지 않은 남은 bit 수입니다. |
| `8/4` | 생성한 output byte 수입니다. |
| `12/4` | reserved입니다. |
| `16/8` | command runtime이며 time unit은 unspecified입니다. |
| `24/8` | reserved입니다. |
| `32/4` | 처리한 element 수입니다. |
| `36/20` | reserved입니다. |
| `56/8` | return value입니다. |
| `64/64` | extended return value입니다. |
Completion Area 상태와 오류 처리
1024-1085guest software는 Completion Area를 read-only로 취급해야 합니다. CCB를 성공적으로 제출하면 hypervisor가 status byte를 0으로 지우며, status가 non-zero가 될 때까지 나머지 field는 invalid입니다.
status `0x2` 또는 `0x3`은 부분 결과나 side effect를 만들 수 있으므로 destination buffer를 unknown state로 취급하는 것이 권장됩니다. `0x2`이면 error reason code로 조치를 결정합니다.
- Buffer overflow: output buffer보다 결과가 큽니다. 더 큰 buffer로 재제출할 수 있습니다.
- CCB decoding error: invalid field 또는 pipeline output을 존재하지 않는 secondary input으로 보낸 경우입니다.
- Page overflow: address에 연결된 page size를 넘어 접근하려 했습니다. boundary 밖은 읽거나 쓰지 않지만 부분 output은 있을 수 있으며 더 큰 page allocation으로 재시도할 수 있습니다. pipeline에서는 source output이 target input보다 먼저 끝나도 발생합니다.
- Command killed: `ccb_kill`로 실행이 중지되거나 막혔습니다.
- Command timeout: 실행은 시작했지만 VM limit 안에 끝나지 않았습니다. 변경 없이 재제출할 수 있습니다.
- ADI miscompare: CCB의 buffer version과 memory tag가 달랐습니다. 원인을 확인하기 전 재제출하면 안 됩니다.
- Data format error: input stream이 CCB에서 선택한 format을 따르지 않습니다.
- Hardware error: 일부는 변경 없이 재시도할 수 있지만 persistent error는 RAS software가 faulty component를 격리할 때까지 반복될 수 있습니다.
output size는 destination buffer의 valid byte 수, runtime은 internal queue를 나온 뒤의 상대 실행 시간, elements는 해당 command가 처리한 element 수입니다. runtime unit은 고정이지만 unspecified이고 platform마다 달라 절대 시간으로 해석하면 안 됩니다. return value와 extended return value는 destination buffer를 쓰지 않는 command 또는 secondary result에 사용되며 모든 command에서 valid한 것은 아닙니다.
Hypervisor API ccb_submit
1086-1207| 항목 | 값 |
|---|---|
| trap/function | `FAST_TRAP` / `CCB_SUBMIT` |
| input | `arg0=address`, `arg1=length`, `arg2=flags`, `arg3=reserved` |
| output | `ret0=status`, `ret1=length`, `ret2=status data`, `ret3=reserved` |
`ccb_submit`은 하나 이상의 CCB를 VM에 제출합니다. `address`가 linear array를 가리키고 `length`는 byte 크기입니다. address는 length를 다음 power of two로 올림한 크기에 align하는 것이 원칙이며 length는 64 byte 배수여야 합니다. length가 0이면 `ret1`에 최대 array size를 반환하고, 그 밖에는 소비한 CCB byte 수를 반환합니다.
가장 작은 지원 page 하나 안에 array 전체가 들어 있으면 VM은 address alignment만으로 제출을 거부하지 않는 것이 좋습니다.
CCB array와 내부 address는 real 또는 virtual일 수 있고 type은 flags로 지정합니다. VA는 TLB 또는 active TSB에 있어야 하며 translation context는 CCB와 flags 조합으로 결정됩니다.
| flags bit | 의미 |
|---|---|
| `[63:16]` | reserved |
| `[15]` | Disable ADI for VA reads. API 2.0에서 VA read의 ADI 검사를 끄며 API 1.0에서는 reserved입니다. |
| `[14]` | CCB 내부 VA를 privileged context에서 translation |
| `[13:12]` | alternate context: `00` request reject, `01` reserved, `10` secondary context, `11` nucleus context |
| `[11:9]` | reserved |
| `[8]` | queue info flag |
| `[7]` | all-or-nothing flag |
| `[6]` | CCB array address가 VA이면 privileged context로 translation |
| `[5:4]` | array address type: `00` real, `01` primary-context VA, `10` secondary-context VA, `11` nucleus-context VA |
| `[3:2]` | reserved |
| `[1:0]` | CCB command type: `00` reserved, `01` reserved, `10` query command, `11` reserved |
submission type과 array address type은 필수이고 나머지 flag는 기본 동작을 바꿉니다. bit 15는 API 2.0의 VA load에만 ADI 검사를 끄며 real-address load는 계속 검사합니다. 기본 VA는 user/primary context이고 privileged 또는 alternate context가 필요하면 flags로 지정합니다. 한 submission의 모든 CCB VA는 같은 privilege level을 사용해야 합니다.
all-or-nothing은 CCB array의 partial submission 허용 여부를 정합니다. serial-conditional chain에는 사용을 강하게 권장합니다. 사용하지 않으면 `ret1.length`로 제출된 byte를 확인하고 나머지를 재제출해야 하며 끊긴 chain 경계에서는 guest가 순서를 직접 구현하고 첫 remaining CCB의 Conditional flag를 지워야 합니다. error 반환에서도 `ret1`을 확인하면 처리된 CCB와 오류를 일으킨 CCB를 찾을 수 있습니다.
queue info flag를 쓰고 하나 이상 제출되면 `ret1.length`는 `[63:48]` DAX unit id, `[47:32]` DAX queue id, `[31:16]` reserved, `[15:0]` 제출 byte 수로 구성됩니다. `ret2.status data`는 status별 정의가 있을 때만 valid합니다. reserved input/output register는 향후 minor version을 위한 것이므로 guest는 call 중 volatile로 취급해야 합니다.
ccb_submit status와 오류
1208-1290| status | 의미와 조치 |
|---|---|
| `EOK` | 하나 이상의 CCB를 accept/enqueue했고 submission error가 없습니다. 내부 limit 때문에 남은 CCB는 변경 없이 재제출할 수 있습니다. |
| `EWOULDBLOCK` | VM 내부 resource conflict로 처리를 끝내지 못했습니다. 앞부분은 enqueue됐을 수 있고 나머지는 변경 없이 재제출합니다. |
| `EBADALIGN` | CCB array가 64-byte boundary가 아니거나 length가 64 byte 배수가 아닙니다. |
| `ENORADDR` | array 또는 CCB 내부 real address가 guest에 유효하지 않습니다. 오류 전 CCB 일부는 enqueue됐을 수 있습니다. |
| `ENOMAP` | TLB/TSB로 VA를 translation하지 못했습니다. mapping 후 또는 real address로 바꿔 재시도합니다. 실패 VA는 `ret2` status data로 반환되며 shared translation resource 특성상 무한 재실패 가능성이 있어 real-address fallback이 권장됩니다. |
| `EINVAL` | invalid CCB나 flags 등 invalid argument입니다. 일부 invalid 값은 submission이 아니라 Completion Area에서 보고될 수 있고 invalid CCB version도 원인입니다. |
| `ETOOMANY` | all-or-nothing array가 한 request의 VM 지원 크기보다 큽니다. zero-length query로 최대 크기를 확인합니다. |
| `ENOACCESS` | CCB 제출 권한이 없거나 destination write permission처럼 address permission이 부족합니다. 실패 VA는 `ret2`에 반환됩니다. |
| `EUNAVAILABLE` | 현재 requested operation을 수행할 수 없습니다. 영구 상태로 해석하지 말고 이후 다시 시도하며 status data로 범위를 판단합니다. |
| EUNAVAILABLE status data | 제한 범위 |
|---|---|
| `0` | 정확히 이 CCB instance만 불가합니다. guest가 emulate하고 다른 CCB는 계속 제출합니다. |
| `1` | 같은 opcode 전체가 불가합니다. 해당 opcode를 emulate하고 다른 opcode는 계속 제출합니다. |
| `2` | 같은 CCB version 전체가 불가합니다. 해당 version을 emulate하고 다른 version은 계속 제출합니다. |
| `3` | 현재 vCPU의 모든 CCB가 불가합니다. emulate하거나 다른 vCPU로 재제출하고 다른 vCPU는 계속 사용합니다. |
| `4` | 모든 CCB가 불가합니다. guest가 operation을 emulate해야 합니다. |
ccb_info 상태 조회
1291-1339| 항목 | 값 |
|---|---|
| trap/function | `FAST_TRAP` / `CCB_INFO` |
| input | `arg0=address` |
| output | `ret0=status`, `ret1=CCB state`, `ret2=position`, `ret3=dax`, `ret4=queue` |
`ccb_info`는 이전에 제출한 CCB의 상태를 조회합니다. CCB는 Completion Area의 64-byte aligned real address로 식별합니다.
| state/value | 설명 |
|---|---|
| `COMPLETED / 0` | fetch와 실행이 끝나 VM에서 더 이상 active하지 않습니다. |
| `ENQUEUED / 1` | queue에서 실행을 기다리고 있습니다. |
| `INPROGRESS / 2` | fetch 후 실행 중이며 `ccb_kill`로 멈출 수 있을 수도 있습니다. |
| `NOTFOUND / 3` | VM에서 찾지 못했고 실행 흔적도 없습니다. hardware error로 유실됐거나 처음부터 제출되지 않았을 수 있습니다. |
일부 platform은 processing 중 CCB를 보고하지 못할 수 있으므로 NOTFOUND만 보고 절대 실행되지 않는다고 판단하기 전에 `ccb_kill`을 호출해야 합니다.
`position`, `dax`, `queue` return은 ENQUEUED일 때만 valid합니다. position은 앞선 CCB 수, dax는 queue를 처리하는 DAX unit instance id, queue는 해당 queue instance id입니다. 두 id는 queue info flag를 쓴 `ccb_submit` 결과와 일치합니다.
ccb_info 오류
1340-1350| status | 설명 |
|---|---|
| `EOK` | request를 처리했고 CCB state가 valid합니다. |
| `EBADALIGN` | address가 64-byte aligned가 아닙니다. |
| `ENORADDR` | 제공한 real address가 invalid입니다. |
| `EINVAL` | Completion Area 내용이 invalid입니다. |
| `EWOULDBLOCK` | resource constraint로 지금 조회할 수 없습니다. retry합니다. |
| `ENOACCESS` | coprocessor virtual-device 기능 접근 권한이 없습니다. |
ccb_kill 실행 중지
1351-1392| 항목 | 값 |
|---|---|
| trap/function | `FAST_TRAP` / `CCB_KILL` |
| input | `arg0=address` |
| output | `ret0=status`, `ret1=result` |
`ccb_kill`은 이전에 제출한 CCB 실행을 중지합니다. Completion Area의 64-byte aligned real address로 CCB를 식별합니다.
| result/value | 설명 |
|---|---|
| `COMPLETED / 0` | 이미 실행이 끝나 active하지 않으므로 kill하지 못했고 아무 조치도 하지 않았습니다. |
| `DEQUEUED / 1` | 아직 queue에 있어 제거했습니다. 실행 전이므로 memory modification이 없고 Completion Area는 갱신되지 않습니다. 변경 없이 재제출할 수 있습니다. |
| `KILLED / 2` | 실행 중이던 CCB를 멈췄습니다. Completion Area에 killed status가 기록되고 command 전체 실행을 포함한 partial result가 생겼을 수 있습니다. |
| `NOTFOUND / 3` | VM에서 찾지 못하고 실행 흔적도 없습니다. hardware error로 유실됐거나 제출되지 않았을 수 있으며 재제출하지 않는 한 이후 실행되지 않음이 보장됩니다. |
pipeline kill 상호작용과 오류
1393-1412pipeline target을 kill했지만 source가 skip된 경우 target Completion Area는 `(3,7) Command was killed` 대신 `(4,0) Command was skipped`가 될 수 있습니다. source를 kill하면 target이 `(1,0) Success`로 보일 수 있지만 source output이 없으므로 target이 실제 처리됐다는 뜻은 아닙니다.
| status | 설명 |
|---|---|
| `EOK` | request를 처리했고 result가 valid합니다. |
| `EBADALIGN` | address가 64-byte aligned가 아닙니다. |
| `ENORADDR` | 제공한 real address가 invalid입니다. |
| `EINVAL` | Completion Area 내용이 invalid입니다. |
| `EWOULDBLOCK` | resource constraint로 지금 kill할 수 없습니다. retry합니다. |
| `ENOACCESS` | coprocessor virtual-device 기능 접근 권한이 없습니다. |
dax_info unit 수 조회
1413-1433| 항목 | 값 |
|---|---|
| trap/function | `FAST_TRAP` / `DAX_INFO` |
| output | `ret0=status`, `ret1=enabled DAX units`, `ret2=disabled DAX units` |
`dax_info`는 calling guest가 CCB를 제출할 수 있도록 활성화된 DAX unit 수와 비활성화된 수를 반환합니다. disabled unit은 offline 상태가 아니었다면 해당 guest가 CCB 제출에 사용할 수 있었던 unit입니다.
정상 처리 시 `EOK`를 반환하고 enabled/disabled unit 수가 valid합니다.
요약과 해설
dax-hv-api.txt:1-1433DAX는 database-oriented data stream을 Extract, Scan, Translate, Select하는 비동기 coprocessor입니다. guest는 64/128-byte CCB와 128-byte Completion Area를 구성하고 `ccb_submit`으로 queue에 넣으며, address type·ADI·page boundary·serial/conditional/pipeline 관계를 bit field로 지정합니다.
ABI 사용 시 memory pinning, 64-byte alignment, TLB/TSB mapping과 partial submission을 함께 관리해야 합니다. Completion status가 실패나 kill이면 destination이 partial 또는 unknown state일 수 있으며, timeout 뒤에는 `ccb_info`로 state를 확인한 후 `ccb_kill`을 사용합니다. `dax_info`는 guest에 할당된 enabled/disabled unit 수를 제공합니다.