요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
===================
Setting up NFS/RDMA
===================
:Author:
NetApp and Open Grid Computing (May 29, 2008)
.. warning::
This document is probably obsolete.
Overview
========
This document describes how to install and setup the Linux NFS/RDMA client
and server software.
The NFS/RDMA client was first included in Linux 2.6.24. The NFS/RDMA server
was first included in the following release, Linux 2.6.25.
In our testing, we have obtained excellent performance results (full 10Gbit
wire bandwidth at minimal client CPU) under many workloads. The code passes
the full Connectathon test suite and operates over both Infiniband and iWARP
RDMA adapters.
Getting Help
============
If you get stuck, you can ask questions on the
[email protected] mailing list.
Installation
============
These instructions are a step by step guide to building a machine for
use with NFS/RDMA.
- Install an RDMA device
Any device supported by the drivers in drivers/infiniband/hw is acceptable.
Testing has been performed using several Mellanox-based IB cards, the
Ammasso AMS1100 iWARP adapter, and the Chelsio cxgb3 iWARP adapter.
- Install a Linux distribution and tools
The first kernel release to contain both the NFS/RDMA client and server was
Linux 2.6.25 Therefore, a distribution compatible with this and subsequent
Linux kernel release should be installed.
The procedures described in this document have been tested with
distributions from Red Hat's Fedora Project (http://fedora.redhat.com/).
- Install nfs-utils-1.1.2 or greater on the client
An NFS/RDMA mount point can be obtained by using the mount.nfs command in
nfs-utils-1.1.2 or greater (nfs-utils-1.1.1 was the first nfs-utils
version with support for NFS/RDMA mounts, but for various reasons we
recommend using nfs-utils-1.1.2 or greater). To see which version of
mount.nfs you are using, type:
.. code-block:: sh
$ /sbin/mount.nfs -V
If the version is less than 1.1.2 or the command does not exist,
you should install the latest version of nfs-utils.
Download the latest package from: https://www.kernel.org/pub/linux/utils/nfs
Uncompress the package and follow the installation instructions.
If you will not need the idmapper and gssd executables (you do not need
these to create an NFS/RDMA enabled mount command), the installation
process can be simplified by disabling these features when running
configure:
.. code-block:: sh
$ ./configure --disable-gss --disable-nfsv4
To build nfs-utils you will need the tcp_wrappers package installed. For
more information on this see the package's README and INSTALL files.
After building the nfs-utils package, there will be a mount.nfs binary in
the utils/mount directory. This binary can be used to initiate NFS v2, v3,
or v4 mounts. To initiate a v4 mount, the binary must be called
mount.nfs4. The standard technique is to create a symlink called
mount.nfs4 to mount.nfs.
This mount.nfs binary should be installed at /sbin/mount.nfs as follows:
.. code-block:: sh
$ sudo cp utils/mount/mount.nfs /sbin/mount.nfs
In this location, mount.nfs will be invoked automatically for NFS mounts
by the system mount command.
.. note::
mount.nfs and therefore nfs-utils-1.1.2 or greater is only needed
on the NFS client machine. You do not need this specific version of
nfs-utils on the server. Furthermore, only the mount.nfs command from
nfs-utils-1.1.2 is needed on the client.
- Install a Linux kernel with NFS/RDMA
The NFS/RDMA client and server are both included in the mainline Linux
kernel version 2.6.25 and later. This and other versions of the Linux
kernel can be found at: https://www.kernel.org/pub/linux/kernel/
Download the sources and place them in an appropriate location.
- Configure the RDMA stack
Make sure your kernel configuration has RDMA support enabled. Under
Device Drivers -> InfiniBand support, update the kernel configuration
to enable InfiniBand support [NOTE: the option name is misleading. Enabling
InfiniBand support is required for all RDMA devices (IB, iWARP, etc.)].
Enable the appropriate IB HCA support (mlx4, mthca, ehca, ipath, etc.) or
iWARP adapter support (amso, cxgb3, etc.).
If you are using InfiniBand, be sure to enable IP-over-InfiniBand support.
- Configure the NFS client and server
Your kernel configuration must also have NFS file system support and/or
NFS server support enabled. These and other NFS related configuration
options can be found under File Systems -> Network File Systems.
- Build, install, reboot
The NFS/RDMA code will be enabled automatically if NFS and RDMA
are turned on. The NFS/RDMA client and server are configured via the hidden
SUNRPC_XPRT_RDMA config option that depends on SUNRPC and INFINIBAND. The
value of SUNRPC_XPRT_RDMA will be:
#. N if either SUNRPC or INFINIBAND are N, in this case the NFS/RDMA client
and server will not be built
#. M if both SUNRPC and INFINIBAND are on (M or Y) and at least one is M,
in this case the NFS/RDMA client and server will be built as modules
#. Y if both SUNRPC and INFINIBAND are Y, in this case the NFS/RDMA client
and server will be built into the kernel
Therefore, if you have followed the steps above and turned no NFS and RDMA,
the NFS/RDMA client and server will be built.
Build a new kernel, install it, boot it.
Check RDMA and NFS Setup
========================
Before configuring the NFS/RDMA software, it is a good idea to test
your new kernel to ensure that the kernel is working correctly.
In particular, it is a good idea to verify that the RDMA stack
is functioning as expected and standard NFS over TCP/IP and/or UDP/IP
is working properly.
- Check RDMA Setup
If you built the RDMA components as modules, load them at
this time. For example, if you are using a Mellanox Tavor/Sinai/Arbel
card:
.. code-block:: sh
$ modprobe ib_mthca
$ modprobe ib_ipoib
If you are using InfiniBand, make sure there is a Subnet Manager (SM)
running on the network. If your IB switch has an embedded SM, you can
use it. Otherwise, you will need to run an SM, such as OpenSM, on one
of your end nodes.
If an SM is running on your network, you should see the following:
.. code-block:: sh
$ cat /sys/class/infiniband/driverX/ports/1/state
4: ACTIVE
where driverX is mthca0, ipath5, ehca3, etc.
To further test the InfiniBand software stack, use IPoIB (this
assumes you have two IB hosts named host1 and host2):
.. code-block:: sh
host1$ ip link set dev ib0 up
host1$ ip address add dev ib0 a.b.c.x
host2$ ip link set dev ib0 up
host2$ ip address add dev ib0 a.b.c.y
host1$ ping a.b.c.y
host2$ ping a.b.c.x
For other device types, follow the appropriate procedures.
- Check NFS Setup
For the NFS components enabled above (client and/or server),
test their functionality over standard Ethernet using TCP/IP or UDP/IP.
NFS/RDMA Setup
==============
We recommend that you use two machines, one to act as the client and
one to act as the server.
One time configuration:
-----------------------
- On the server system, configure the /etc/exports file and start the NFS/RDMA server.
Exports entries with the following formats have been tested::
/vol0 192.168.0.47(fsid=0,rw,async,insecure,no_root_squash)
/vol0 192.168.0.0/255.255.255.0(fsid=0,rw,async,insecure,no_root_squash)
The IP address(es) is(are) the client's IPoIB address for an InfiniBand
HCA or the client's iWARP address(es) for an RNIC.
.. note::
The "insecure" option must be used because the NFS/RDMA client does
not use a reserved port.
Each time a machine boots:
--------------------------
- Load and configure the RDMA drivers
For InfiniBand using a Mellanox adapter:
.. code-block:: sh
$ modprobe ib_mthca
$ modprobe ib_ipoib
$ ip li set dev ib0 up
$ ip addr add dev ib0 a.b.c.d
.. note::
Please use unique addresses for the client and server!
- Start the NFS server
If the NFS/RDMA server was built as a module (CONFIG_SUNRPC_XPRT_RDMA=m in
kernel config), load the RDMA transport module:
.. code-block:: sh
$ modprobe svcrdma
Regardless of how the server was built (module or built-in), start the
server:
.. code-block:: sh
$ /etc/init.d/nfs start
or
.. code-block:: sh
$ service nfs start
Instruct the server to listen on the RDMA transport:
.. code-block:: sh
$ echo rdma 20049 > /proc/fs/nfsd/portlist
- On the client system
If the NFS/RDMA client was built as a module (CONFIG_SUNRPC_XPRT_RDMA=m in
kernel config), load the RDMA client module:
.. code-block:: sh
$ modprobe xprtrdma.ko
Regardless of how the client was built (module or built-in), use this
command to mount the NFS/RDMA server:
.. code-block:: sh
$ mount -o rdma,port=20049 <IPoIB-server-name-or-address>:/<export> /mnt
To verify that the mount is using RDMA, run "cat /proc/mounts" and check
the "proto" field for the given mount.
Congratulations! You're using NFS/RDMA!
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
문서 정보와 경고
1-10제목은 `Setting up NFS/RDMA`이며 NetApp과 Open Grid Computing이 2008년 5월 29일 작성했습니다.
원문은 이 문서가 probably obsolete라고 명시합니다. 현재 system에 적용하기 전에 배포판과 최신 kernel·nfs-utils 문서를 반드시 대조하십시오.
개요와 도움 요청
11-30이 문서는 Linux NFS/RDMA client와 server software의 설치·설정 절차를 설명합니다. Client는 Linux 2.6.24에 처음 포함됐고 server는 다음 release인 Linux 2.6.25에 포함됐습니다.
원문 시험에서는 여러 workload에서 최소 client CPU로 full 10Gbit wire bandwidth를 얻었고, code가 전체 Connectathon test suite를 통과하며 Infiniband와 iWARP RDMA adapter에서 모두 동작했습니다.
문제가 생기면 `[email protected]` mailing list에 질문할 수 있습니다.
장치·배포판·nfs-utils 설치
31-104NFS/RDMA machine 준비 조건은 다음과 같습니다.
| 구성 | 요구 사항 | 원문 확인 범위 |
|---|---|---|
| RDMA device | `drivers/infiniband/hw`의 driver가 지원하는 장치 | Mellanox IB, Ammasso AMS1100 iWARP, Chelsio cxgb3 iWARP에서 시험 |
| Linux distribution | Linux 2.6.25 이상과 호환 | Client와 server가 모두 mainline에 들어간 최초 release 기준 |
| Client utility | `nfs-utils-1.1.2` 이상 | `mount.nfs`로 NFS/RDMA mount를 생성하며 server에는 이 특정 version이 필요하지 않음 |
사용 중인 `mount.nfs` version을 다음 명령으로 확인합니다.
$ /sbin/mount.nfs -V
`nfs-utils-1.1.1`은 NFS/RDMA mount를 처음 지원한 nfs-utils version이지만, 원문은 여러 이유로 `nfs-utils-1.1.2` 이상을 권장합니다. 사용 중인 version이 1.1.2보다 낮거나 명령이 없으면 최신 nfs-utils를 설치합니다. `idmapper`와 `gssd` executable이 필요 없다면 NFS/RDMA mount command 자체에는 필요하지 않으므로 configure 때 기능을 끌 수 있습니다.
$ ./configure --disable-gss --disable-nfsv4
nfs-utils build에는 `tcp_wrappers` package가 필요하며 자세한 내용은 package의 `README`와 `INSTALL`을 봅니다. Build 뒤 `utils/mount`의 `mount.nfs`는 NFS v2·v3·v4 mount를 시작할 수 있습니다. V4는 binary 이름이 `mount.nfs4`여야 하므로 보통 `mount.nfs4`에서 `mount.nfs`로 symlink를 만듭니다.
완성한 binary를 다음처럼 `/sbin/mount.nfs`에 설치합니다.
$ sudo cp utils/mount/mount.nfs /sbin/mount.nfs
이 위치에 있으면 system `mount` command가 NFS mount 때 `mount.nfs`를 자동 호출합니다.
`mount.nfs`, 즉 nfs-utils-1.1.2 이상은 NFS client machine에만 필요합니다. Server에는 이 특정 version이 필요 없고 client도 nfs-utils 전체가 아니라 해당 `mount.nfs` command만 필요합니다.
Kernel과 RDMA·NFS 구성
105-151NFS/RDMA client와 server는 mainline Linux 2.6.25 이상에 함께 포함됩니다. Kernel source는 다음 위치에서 받을 수 있습니다.
Kernel configuration에서 `Device Drivers -> InfiniBand support`의 RDMA 지원을 켭니다. 이름과 달리 이 option은 IB와 iWARP를 포함한 모든 RDMA device에 필요합니다. 사용하는 IB HCA의 `mlx4`, `mthca`, `ehca`, `ipath` 또는 iWARP adapter의 `amso`, `cxgb3` 지원을 켭니다. InfiniBand를 쓴다면 IP-over-InfiniBand도 활성화합니다.
`File Systems -> Network File Systems`에서 NFS filesystem support와 필요하면 NFS server support를 켭니다. NFS와 RDMA가 활성화되면 NFS/RDMA code도 자동 활성화됩니다.
Client와 server build 방식은 `SUNRPC`와 `INFINIBAND`에 의존하는 hidden `SUNRPC_XPRT_RDMA` config option으로 결정됩니다.
| 결과 | 조건 | Build 동작 |
|---|---|---|
| N | `SUNRPC=N` 또는 `INFINIBAND=N` | NFS/RDMA client와 server를 build하지 않음 |
| M | 두 symbol이 M 또는 Y이고 하나 이상이 M | Client와 server를 module로 build |
| Y | `SUNRPC=Y` 및 `INFINIBAND=Y` | Client와 server를 kernel built-in으로 build |
따라서 위 절차대로 NFS와 RDMA를 모두 켜면 client와 server가 build됩니다. 새 kernel을 build·install하고 reboot합니다.
RDMA와 일반 NFS 사전 검사
152-204NFS/RDMA software를 구성하기 전에 새 kernel, RDMA stack, 표준 NFS over TCP/IP 또는 UDP/IP가 정상 동작하는지 확인하는 것이 좋습니다.
| 대상 | 검사 |
|---|---|
| RDMA | Driver load, Subnet Manager, port state `4: ACTIVE`, IPoIB ping |
| NFS | 일반 Ethernet에서 TCP/IP 또는 UDP/IP로 client/server 기능 확인 |
RDMA component를 module로 build했다면 load합니다. Mellanox Tavor/Sinai/Arbel card 예시는 다음과 같습니다.
$ modprobe ib_mthca
$ modprobe ib_ipoib
InfiniBand network에는 Subnet Manager(SM)가 필요합니다. IB switch의 embedded SM을 쓰거나 end node 하나에서 OpenSM 같은 SM을 실행합니다. SM이 동작하면 port state는 다음과 같이 보여야 합니다.
$ cat /sys/class/infiniband/driverX/ports/1/state
4: ACTIVE
`driverX`는 `mthca0`, `ipath5`, `ehca3` 등입니다. Host 이름이 `host1`, `host2`인 두 IB system에서 IPoIB stack을 더 시험하는 절차는 다음과 같습니다.
host1$ ip link set dev ib0 up
host1$ ip address add dev ib0 a.b.c.x
host2$ ip link set dev ib0 up
host2$ ip address add dev ib0 a.b.c.y
host1$ ping a.b.c.y
host2$ ping a.b.c.x
다른 device type은 해당 절차를 따릅니다. 활성화한 NFS client/server는 일반 Ethernet에서 TCP/IP 또는 UDP/IP로 기능을 시험합니다.
Server 일회성 구성
205-227Client와 server 역할을 각각 맡는 두 machine 사용을 권장합니다. Server의 `/etc/exports`를 구성하고 NFS/RDMA server를 시작합니다. 원문에서 시험한 export 형식은 다음과 같습니다.
/vol0 192.168.0.47(fsid=0,rw,async,insecure,no_root_squash)
/vol0 192.168.0.0/255.255.255.0(fsid=0,rw,async,insecure,no_root_squash)
IP address는 InfiniBand HCA를 쓸 때 client의 IPoIB address이고, RNIC를 쓸 때 client의 iWARP address입니다.
NFS/RDMA client는 reserved port를 사용하지 않으므로 export에 `insecure` option을 반드시 사용해야 합니다.
매 boot의 server·client 절차
228-292Mellanox adapter 기반 InfiniBand에서는 RDMA driver를 load하고 `ib0`를 올린 뒤 client와 server에 서로 다른 address를 설정합니다.
$ modprobe ib_mthca
$ modprobe ib_ipoib
$ ip li set dev ib0 up
$ ip addr add dev ib0 a.b.c.d
Client와 server에는 반드시 unique address를 사용하십시오.
NFS/RDMA server가 `CONFIG_SUNRPC_XPRT_RDMA=m` module로 build됐다면 RDMA transport module을 load합니다.
$ modprobe svcrdma
Module 또는 built-in 여부와 관계없이 NFS server를 시작합니다.
$ /etc/init.d/nfs start
또는 service command를 사용할 수 있습니다.
$ service nfs start
Server가 RDMA transport의 port `20049`에서 listen하도록 지시합니다.
$ echo rdma 20049 > /proc/fs/nfsd/portlist
Client가 `CONFIG_SUNRPC_XPRT_RDMA=m` module로 build됐다면 RDMA client module을 load합니다.
$ modprobe xprtrdma.ko
Client의 build 방식과 관계없이 다음 명령으로 NFS/RDMA server를 mount합니다.
$ mount -o rdma,port=20049 <IPoIB-server-name-or-address>:/<export> /mnt
RDMA mount인지 검증하려면 `cat /proc/mounts`를 실행하고 해당 mount의 `proto` field를 확인합니다.
운영 핵심
nfs-rdma.rst:1-292이 2008년 문서는 NFS/RDMA client와 server를 kernel 2.6.25 계열 방식으로 구성합니다. 원문 경고대로 최신 환경에서는 현재 배포판 절차와 대조해야 합니다.