System Requirements¶
Operating Systems Supported¶
x86-64 (all applications: Pre, Solve, Post)
Windows 10, 11
Windows Server 2016 or newer
Red Hat Enterprise Linux / Oracle Linux 8, 9, 10
Ubuntu 20.04, 22.04, 24.04
arm64 / aarch64 (Solver only)
Ubuntu 22.04, 24.04
Note
RHEL/Oracle Linux 7 and Debian are no longer supported as of M-Star CFD 4.0.
See M-Star CFD Version 4.0.x Notes for the full list of platform changes.
Python¶
Python 3.9 through 3.13 is supported on all platforms, including Windows. Python is required for the M-Star Pre API, System UDFs, and Global Scripts. See Python Install.
CUDA Runtime and GPU Driver Support¶
M-Star CFD 4.x is built against CUDA 12.9. The CUDA runtime is bundled with the distribution; only the NVIDIA driver needs to be installed on the host.
Platform |
CUDA Build |
Minimum Driver |
Minimum Compute Capability |
|---|---|---|---|
Windows 10/11, Windows Server |
12.9 |
525 |
6.0 |
RHEL/Oracle 8, 9; Ubuntu 20/22/24 |
12.9 |
525 |
6.0 |
RHEL/Oracle 10 |
13.0 |
580 |
7.5 |
M-Star is always compatible with the latest NVIDIA driver. We recommend installing the newest production driver for your platform.
Tip
Linux only: if you cannot update the host driver to 525 or later, the
cuda-compat-12-9 forward-compatibility package may allow M-Star to run
on an older driver. See the NVIDIA forward compatibility documentation.
OpenGL Support¶
OpenGL is a 3D rendering technology required by both the M-Star Pre-Processor and Post.
OpenGL version 3.3+ is required
In practice this means integrated or discrete graphics hardware with a vendor-supplied driver. Software rendering is unlikely to work and is not recommended. For remote Linux sessions, see Remote Visualization (NICE DCV).
If you have trouble starting M-Star Pre or Post, see Troubleshooting M-Star Pre/Post Startup.
NVidia Hardware¶
NVidia GPUs¶
The M-Star Solver requires GPUs with Compute Capability 6.0 or newer . (Pascal architecture, 2016, and later). Kepler and Maxwell GPUs (compute capability 3.5–5.2) supported by M-Star 3.x will not run 4.x. On RHEL/Oracle 10 the minimum is 7.5 (Turing and later). For a complete list see the NVIDIA CUDA GPU Compute Capability table.
M-Star 4.x has native (pre-compiled) support for the Pascal, Volta, Turing, Ampere, Ada Lovelace, Hopper, and Blackwell architectures.
Data Center GPUs (H100, H200, B200, and similar)
Enterprise-grade hardware intended for sustained computational load. These offer the largest memory capacity, the highest memory bandwidth, ECC memory, and — in SXM form factor — full NVLink/NVSwitch connectivity across many GPUs. Recommended for servers and for the largest models.
Workstation GPUs (RTX PRO 6000 Blackwell, RTX 6000 Ada, RTX A6000, and similar)
PCIe cards with ECC memory and up to 96 GB per GPU. On Windows these support TCC mode, which removes the display-driver overhead. Most workstation users land here; a single RTX PRO 6000 covers a very wide range of models.
GeForce GPUs (RTX 40-series, RTX 50-series)
Consumer-grade cards intended primarily for gaming and general desktop use. They are compatible with the solver and offer excellent price-to-performance for single-GPU work that fits in 16–32 GB. They lack ECC memory and NVLink.
Note
Multi-GPU jobs on GeForce hardware are supported on Linux only. On Windows, GeForce cards run in WDDM mode, which does not support the peer-to-peer access required for multi-GPU solves.
For a detailed comparison of current GPUs, see the Hardware Guide.
NVIDIA NVLink¶
NVLink is a high-speed GPU-to-GPU interconnect. When a simulation is split across multiple GPUs, inter-GPU bandwidth is frequently the limiting factor, so NVLink should be used wherever the hardware supports it.
SXM data-center GPUs (H100 SXM, H200 SXM, B200) connect through NVLink/NVSwitch to every other GPU in the node. This is the recommended configuration for multi-GPU servers.
PCIe cards support at most a two-GPU NVLink bridge, and only on specific models (RTX A6000, A100 PCIe, RTX 3090). Ada- and Blackwell-generation workstation and consumer cards do not support NVLink; multi-GPU communication on those cards goes over PCIe.
Setup¶
Minimum Requirement¶
Suitable for training, evaluation, and small domain sizes (roughly up to 20–30 million lattice points):
CPU: Quad core or better
Memory: 32 GB
Disk: 500 GB SSD
GPU: NVIDIA GeForce RTX 40-series or 50-series with 12 GB or more
Recommended Workstation¶
Tip
For more guidance on selecting GPUs, see the Hardware Guide.
CPU: 8 cores or better
Memory: 128–256 GB (plan for 1.5–2x the total GPU memory in the machine)
Disk 1: 1 TB NVMe SSD operating-system drive
Disk 2: 2 TB NVMe SSD for simulation scratch space
Display GPU: Any current GeForce or RTX workstation card with 4 GB or more (recommended on Windows so the accelerator GPU can run in TCC mode)
Accelerator GPU: NVIDIA RTX PRO 6000 Blackwell (96 GB) or RTX 6000 Ada (48 GB)
A second accelerator GPU can be added for larger models. Note that Ada/Blackwell workstation cards communicate over PCIe rather than NVLink, so users who routinely need more than one GPU should consider server-class hardware with NVLink/NVSwitch.
Recommended Servers¶
M-Star is a GPU-compute-intensive application. For server deployments the guiding principles are: maximize per-GPU memory, maximize the number of GPUs per node, connect them with NVLink/NVSwitch, and use the highest-bandwidth node interconnect available. Multi-node jobs require GPUDirect RDMA for acceptable performance. We suggest working with an HPC vendor to confirm capability against these requirements.
Starting-point guidelines:
OS: Linux (RHEL/Oracle 8, 9, 10 or Ubuntu 20.04, 22.04, 24.04)
CPU: Core count equal to or greater than the number of GPUs in the node; otherwise not a bottleneck
Memory: 1.5–2x total GPU memory in the node (for example, 1–1.5 TB for an 8x H100 80 GB node; 2 TB or more for an 8x B200 node)
Disk: 200 GB–1 TB per M-Star user on fast network storage visible to all compute nodes; requirements vary widely with output configuration
GPU: NVIDIA H100 SXM, H200 SXM, or B200 SXM with the maximum available per-GPU memory
Compute Nodes: 4x or 8x GPUs per node; larger NVLink domains where available
Intra-node Bandwidth: NVLink/NVSwitch between all GPUs in the node (see NVIDIA NVLink)
Node Interconnect: InfiniBand NDR (400 Gb/s) or better, with GPUDirect RDMA. Required for multi-node jobs.
M-Star 4.x bundles NCCL in all solver builds for multi-GPU and multi-node
communication. For multi-node jobs, the system must support GPUDirect RDMA so
that GPUs on different nodes exchange data directly over the interconnect
without staging through host memory. On NVIDIA/Mellanox fabrics this requires
the nvidia-peermem kernel module (the successor to the older
nv_peer_mem); consult your vendor documentation. See also
CUDA-aware MPI.
Running on CPUs¶
CPU execution is only partially supported and is not recommended. Many M-Star features, including UDFs, are unavailable on CPUs, and GPU execution is orders of magnitude faster.
Tip
A dedicated NVIDIA GPU is strongly recommended.
For additional information related to system requirements, see: