▸ Agent Skills
26 min read

Nutanix Enterprise AI Manual: Requirements, Sizing, and Hardware Guidelines

Prerequisites, supported Kubernetes versions, supported NVIDIA GPUs (L40S, H100, H100-NVL, A100, B300, H200, RTX PRO 6000), RWX/RWO storage classes, MetalLB networking, compute and memory sizing matrices, profile-based deployments (c1k_k200, c5k_k1k), component compute and storage requirements, Agent Gateway sizing, and platform limitations.


Figure 2: Nutanix Enterprise AI Workflow

Nutanix Enterprise AI Deployment Types

This section describes the types of Nutanix Enterprise AI deployments. The deployment types that you subscribe to determines which Nutanix Enterprise AI features are available to you.

Nutanix Enterprise AI supports distinct deployment methods to ensure the solution seamlessly integrates with your infrastructure and operational goals:

Nutanix Enterprise AI for GPT-in-a-Box 2.0

This deployment involves deploying Nutanix Enterprise AI Pro on a Nutanix validated solution stack, including Nutanix Cloud Infrastructure (NCI), Nutanix Kubernetes Platform (NKP), Nutanix Database Service (NDB), and Nutanix Unified Storage with 50TiB of Files, on your choice of compatible servers and hardware.

Nutanix Enterprise AI for Bare Metal

This deployment involves deploying Nutanix Enterprise AI Pro on bare metal servers, including hardware and GPU AI accelerators running on Nutanix Kubernetes Platform.

Nutanix Enterprise AI Standalone

This deployment involves deploying Nutanix Enterprise AI Pro on clusters running CNCF (Cloud Native Computing Foundation) compliant Kubernetes on bare metal servers, on-prem VMs, and public clouds such as Amazon EKS, Azure AKS, or Google GKE.

Nutanix Enterprise AI - Private Inference and Agent Gateway Requirements

This section describes the requirements for deploying Nutanix Enterprise AI- private inference and agent gateway.

General Requirements

Nutanix

For minimum software requirements to deploy Nutanix Enterprise AI in your environment, see

Enterprise AI - Private Inference and Agent Gateway Software Requirements

in Nutanix Enterprise AI

Release Notes .

Ensure that you deploy a Kubernetes cluster with version 1.35.

You can deploy Nutanix Enterprise AI only on the following Cloud Native Computing Foundation (CNCF) - certified Kubernetes distributions:

Nutanix Kubernetes Platform (NKP)

You can install NAI 2.8 only on NKP 2.18.

Nutanix Enterprise AI supports air-gapped installation only on NKP.

Amazon Elastic Kubernetes Service (EKS)

Azure Kubernetes Service (AKS)

Google Kubernetes Engine (GKE)

Ensure that you configure a worker node pool with the necessary resources required to host an inference endpoint based on the inference engine and the required number of GPUs. When you create an endpoint, the system displays an error message if the worker node pool does not have the required number of GPUs available.

For information on the Kubernetes platforms supported by the NVIDIA GPUs, see Supported Operating Systems and Kubernetes Platforms in NVIDIA GPU Operator documentation .

For information on how to configure a worker node pool on an NKP cluster, see

Configuring Node Pools

in the

Nutanix Kubernetes Platform Guide . For information on the minimum resources required for a worker node pool to host an inference endpoint, see

Sizing Requirements

on page 13.

After you configure a worker node pool with NVIDIA GPUs, ensure that you taint the worker nodes with the key nvidia.com/gpu and the effect NoSchedule, to enable inference pods with matching toleration to be scheduled on the tainted worker node. Any pods that do not require GPUs are not scheduled on the tainted worker node.

Ensure that you deploy the necessary GPUs within a node pool in the Kubernetes cluster that hosts Nutanix Enterprise AI.

Adding GPU Node Pool to a Nutanix Cluster

For information on how to deploy GPUs on an NKP cluster, see

in the Nutanix Kubernetes Platform Guide.

Ensure that you install the NVIDIA GPU Operator to deploy LLMs using NVIDIA GPUs.

Configuring GPU for

For information on how to install the NVIDIA GPU Operator on an NKP cluster, see

Kommander Clusters

in the Nutanix Kubernetes Platform Guide . For information on how to install the NVIDIA

GPU Operator on a supported public cloud Kubernetes platform, see the CSP Configurations section in the

NVIDIA GPU Operator

documentation.

Nutanix recommends that you have one or more AI compatible GPUs deployed in your cluster. Nutanix Enterprise AI supports the following GPUs:

NVIDIA Ada Lovelace L40S

NVIDIA Hopper H100

NVIDIA Hopper H100-NVL

NVIDIA Ampere A100

NVIDIA Blackwell Ultra B300

NVIDIA Hopper H200

NVIDIA RTX PRO 6000 Blackwell Server Edition

[!NOTE] Note:

Nutanix Enterprise AI does not block the deployment of GPU models that are not listed above. However, automatic resource calculation for validated model and GPU combinations is available only for the supported GPU models listed above. Other GPU models are not officially supported and have not been validated by Nutanix. However, if you use an unlisted GPU model, the following limitations apply:

Nutanix does not provide official support for the GPU model.

Nutanix Enterprise AI does not automatically calculate the resource requirements for the model and GPU combination. You must size and validate the deployment manually.

For information on the Kubernetes platforms supported by the NVIDIA GPUs, see Supported Operating Systems and Kubernetes Platforms in NVIDIA GPU Operator documentation .

Ensure that you create a ReadWriteMany (RWX) storage class that enables static or dynamic provisioning of NFS shares for persistent volumes. Configure the RWX storage class with

set to

.

VolumeBindingMode

Immediate

Creating a Storage Class

For more information on how to create a storage class using Nutanix CSI driver, see

for Dynamic NFS Shares

in the CSI Volume Driver Guide .

Ensure that you configure the LoadBalancer service based on your Kubernetes cluster.

For example, if you have an NKP cluster, ensure that you configure the MetalLB service in the cluster. For more information, see

MetalLB

in the Nutanix Kubernetes Platform Guide .

The LoadBalancer service assigns an external IP address to the Envoy Ingress Gateway service in Nutanix Enterprise AI that external clients can use to connect to Nutanix Enterprise AI.

Ensure that you configure a fully qualified domain name (FQDN) on the DNS domain that is accessible to your Kubernetes cluster using the external IP address of the Envoy Ingress Gateway service in Nutanix Enterprise AI.

The FQDN is necessary to connect to Nutanix Enterprise AI.

Ensure that you have a certificate authority (CA) signed TLS certificate for the configured FQDN.

The TLS certificate is necessary to update the Envoy Ingress gateway after you install Nutanix Enterprise AI in the cluster.

[!NOTE] Note: Nutanix Enterprise AI only supports HTTPS based connections with a valid TLS certificate.

Sizing Requirements

To deploy Nutanix Enterprise AI on a Kubernetes cluster in your environment, ensure that the cluster meets the minimum requirements. Additional worker nodes are required for your environment and vary based on additional cluster workloads and high-availability requirements. In addition to scaling worker node capacity, you must create a dedicated GPU node pool.

The following table outlines the minimum compute and local storage recommendation validated for a baseline deployment of Nutanix Enterprise AI - private inference and agent gateway.

[!NOTE] Note: To determine the cloud VM instance type based on the requirements mentioned in the following table, see the respective cloud service provider documentation.

Table 2: Minimum Compute and Storage Requirements for Deploying Nutanix Enterprise AI-private inference and agent gateway

Node Type Description NumberProperty 2of Nodes (VM)vCPU per NodeMemory per NodeStorage per NodeTotal vCPUTotal Memory
Control PlaneThese control plane nodes are used for running the Kubernetes Control plane.3416 GB150 GB1248 GB
WorkerThese worker nodes are for running the NAI Control Plane.31020 GB150 GB3060 GB

The node sizing assumes the following environmental conditions:

The NKP workload cluster is dedicated exclusively to NAI and its dependencies.

Worker nodes host only the NAI management plane.

Inference-related resources are excluded from these requirements.

The above setup is validated for 100 API Keys and 15 concurrent users. For power usage, scale up the worker nodes to the appropriate sizes in the table below, then redeploy NAI using the profile name. To deploy on:

NKP:

Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform

on page 46

Deploying Nutanix Enterprise AI on Amazon Elastic Kubernetes Service

EKS:

on page 72

Deploying Nutanix Enterprise AI on Azure Kubernetes Service

AKS :

on page 82

Deploying Nutanix Enterprise AI on Google Kubernetes Engine

GKE:

on page 90

Table 3: Compute and Storage Requirements for Profile-based deployment of Nutanix Enterprise AI-private inference and agent gateway

API Keys ConcurrentRequestsProfile nameNode TypevCPU per NodeMemory per NodeStorage per NodeTotal vCPUTotal Memory
2001000c1k_k200Worker6090 GB150 GB180270 GB
10005000c5k_k1kWorker90150 GB150 GB270450 GB

The following table lists the minimum compute requirements for each platform component.

The following table specifies the minimum vCPU resources required for the Model Controller component in millicores. 500m means 500 millicores, which is equivalent to 0.5 vCPU cores.

Table 4: Minimum Compute Requirements By Component

ComponentApp NameMinimum vCPU ResourcesMinimum Memory ResourcesPurpose
ClickHouse Serverchi-nai-clickhouse- server48 GiBObservability data store
ClickHouse Serverchk-nai-clickhouse- keeper100m1 GiBObservability data store
IAMiam-database- bootstrap200m0.12 GiBUser Management
IAMiam-proxy100m0.12 GiBUser Management
IAMiam-proxy-control- plane100m0.0625 GiBUser Management
IAMiam-themis200m0.0625GiBUser Management
IAMiam-themis- bootstrap100m0.031 GiBUser Management
IAMiam-ui150m0.0625 GiBUser Management
IAMiam-user-authn100m0.015 GiBUser Management
APInai-api5.17.128 GiBAPI Endpoint
NAI DB Migratornai-api-db-migrate11.00 GiBUpgrade NAI DB
DBnai-db-iep22.00 GiBApplication Data per instance
Model Controllernai-iep-model- controller500m0.49 GiBLLM Model Downloader Operator
NAI Labsnai-labs500m2.00 GiBChat & Talk to My Data apps
NAI Labsnai-agent12.00 GiBSample Agent App
Oauth2-Proxynai-oauth2-proxy100m0.3 GiBUser Management
OIDC Client Registrationnai-oidc-client- registration11.00 GiBUser Management
ComponentApp NameMinimum vCPU ResourcesMinimum Memory ResourcesPurpose
ClickHouse Operatornai-operators-nai- clickhouse-operator500m0.25 GiBObservability data store
OTEL Collectornai-otel-collector- collector42.00 GiBMetrics exporter from Nodes to ClickHouse
Frontend UInai-ui22.00 GiBFrontend
IAM & NAI Gateway Cachenai-valkey100 m0.0625 GiBCache store per instance
IAM & NAI Gateway Cachenai-valkey-sentinel25 m0.03 GiBCache store per instance
NAI Gatewayai-gateway- controller200m0.25 GiBControl Plane for AI Gateway Resources
NAI Gatewayenvoy-gateway100m0.25 GiBControl Plane for Envoy-Proxy
NAI Gatewayenvoy-nai-system- nai-ingress- gateway5.110.32 GiBData Plane Proxy
NAI Gatewayenvoy-ratelimit100m0.5 GiBRatelimit Service
Securitynai-securityscan- manager200m128 MiBModel and Endpoint scanning
NAInai-go-processor210m266 MiBBatch Inference processor per job
NAInai-audit-logs- rsyslog-otel- collector31 GiBAudit Log Exporter

The following table lists the minimum storage requirements for each platform component.

Table 5: Minimum Storage Requirements By Component

ComponentApp NameMinimum Persistent StorageAccess ModePurpose
Databasenai-db-iep-140 GiBRWOApplication data stored in PostgreSQL database
Modelnai-api20 GiBRWOModel custom resource used to initialize model storage
ClickHousechi-nai-clickhouse- server50 GiBRWOObservability Metrics store
ComponentApp NameMinimum Persistent StorageAccess ModePurpose
Valkeynai-valkey8 GiBRWOIAM and NAI Gateway cache storage per instance
NAInai-audit-logs- rsyslog-otel- collector1 GiBRWOAudit Logs exporter
ClickHouse Keeper chk-nai-clickhouse-10 GiBRWOObservability Metrics store
keeper-chkeeper
NAI Labsnai-labs20 GiBRWONAI Labs data

Nutanix Enterprise AI uses various types of storage for the following purposes:

To save the application and metrics data required to support the Nutanix Enterprise AI platform components.

To download and store pre-validated LLMs required to optimize bootstrapping of LLM inferencing pods during deployment, scaling, and upgrades.

To share LLM files, which enables high availability across multiple instances or replicas of LLM inferencing endpoints.

To manually import downloaded pre-validated or custom LLM Models from existing NFS v4 Share or S3 Compatible Storage.

The following table lists the minimum storage sizing recommendation.

Table 6: Minimum Storage Sizing Recommendations

ComponentMinimum CapacityStorage TechnologyPurpose
RWO (Block - RWO)150 GiBBlockBackend (ClickHouse, PostgreSQL, Valkey, Observability)
RWX (NFS - RWX)2 TiBNFSPre-validated or custom LLM models
S3 Compatible API1 TiBObjectPre-validated or custom LLM models

[!NOTE] Note: If you are upgrading NAI, make sure to expand the Block storage to at least 125#GiB. NAI Agent Gateway only deployments do not require RWS and S3 based storage.

GPU Requirements

The following table lists the pre-validated LLMs and the minimum number of supported GPUs required to deploy these LLMs.

[!NOTE] Note: The requirements mentioned in this table are applicable for entry-level deployments. Ensure that you plan your deployment based on your scale and traffic requirements.

Table 7: Minimum GPU Requirements

LLM ProviderLLM NameGPU Models NVIDIA L40S-48GNVIDIA A100-80GNVIDIA H100-80GNVIDIA H100 NVL-94GNVIDIA H 200-141GNVIDIA RTX PRO 6000-96GNVIDIA B300-268G
Ai2allenai/ Olmo-3-32B- Think2111111
allenai/ Olmo-3-7B- Instruct1111111
allenai/ Olmo-3-7B- Think1111111
AI21 Labsai21labs/AI21- Jamba-1.5-Mini4222121
Cross- Encodercross-encoder/ ms-marco- MiniLM-L6-v21111111
Facebook facebook/deit-1111111
base-distilled- patch16-224
Googlegoogle/ gemma-2-2b-it1111111
google/ gemma-2-9b-it1111111
google/ gemma-3-270m- it1111111
google/ gemma-4-E2B- it1111111
google/ gemma-4-26B- A4B-it2111111
google/ gemma-4-31B-it2111111
google/vit-base- patch16-2241111111
IBMibm-granite/ granite- embedding-107m- multilingual1111111
LLM ProviderLLM NameGPU Models
NVIDIA L40S-48GNVIDIA A100-80GNVIDIA H100-80GNVIDIA H100 NVL-94GNVIDIA H 200-141GNVIDIA RTX PRO 6000-96GNVIDIA B300-268G
Metameta-llama/ CodeLlama-7b- Instruct-hf1111111
meta-llama/ CodeLlama-13b- Instruct-hf1111111
meta-llama/ CodeLlama-34b- Instruct-hf2111111
meta-llama/ CodeLlama-70b- Instruct-hf4222221
meta-llama/ Llama-2-13b- chat-hf1111111
meta-llama/ Llama-3.2-11B- Vision-Instruct111111Not supported
meta-llama/ Llama-3.2-1B- Instruct1111111
meta-llama/ Llama-3.2-3b- Instruct1111111
meta-llama/ Llama-3.2-90B- Vision-Instruct444422Not supported
meta-llama/ Llama-3.3-70B- Instruct4222221
meta-llama/ Llama-4- Scout-17B-16ENot supportedNot supported44241
meta-llama/ Llama- Guard-3-8B1111111
meta- llama/Meta- Llama-3.1-70B- Instruct4222221
LLM ProviderLLM NameGPU Models
NVIDIA L40S-48GNVIDIA A100-80GNVIDIA H100-80GNVIDIA H100 NVL-94GNVIDIA H 200-141GNVIDIA RTX PRO 6000-96GNVIDIA B300-268G
meta- llama/Meta- Llama-3.1-8B- Instruct1111111
Mistral AImistralai/ Devstral- Small-2507Not supported111111
mistralai/ Magistral- Small-2506Not supported111111
mistralai/ Ministral-3-14B- Instruct-25121111111
mistralai/ Ministral-3-14B- Reasoning-25121111111
mistralai/ Ministral-3-3B- Instruct-25121111111
mistralai/ Ministral-3-3B- Reasoning-25121111111
mistralai/ Ministral-3-8B- Instruct-25121111111
mistralai/ Ministral-3-8B- Reasoning-25121111111
mistralai/ Mistral-7B- Instruct-v0.31111111
mistralai/ Mistral-Nemo- Instruct-24071111111
mistralai/ Mixtral-8x7B- Instruct-v0.14222121
mistralai/ Mixtral-8x22B- Instruct-v0.1Not supported444441
mistralai/ Mistral- Small-4-119B-2603Not supportedNot supportedNot supportedNot supported121
LLM ProviderLLM NameGPU Models
NVIDIA L40S-48GNVIDIA A100-80GNVIDIA H100-80GNVIDIA H100 NVL-94GNVIDIA H 200-141GNVIDIA RTX PRO 6000-96GNVIDIA B300-268G
mistralai/ Mistral- Large-3-675B- Instruct-2512Not supportedNot supportedNot supportedNot supported683
NVIDIAnvidia/NVIDIA- Nemotron-3- Nano-30B-A3B- FP81111111
nvidia/NVIDIA- Nemotron-3- Nano-30B-A3B- BF162111111
NVIDIA- Nemotron-3- Super-120B- A12B-BF16Not supportedNot supportedNot supportedNot supported231
NVIDIA- Nemotron-3- Ultra-550B- A55B-BF16Not supportedNot supportedNot supportedNot supported9135
NVIDIAblack-forest- labs/flux.1-devNot supportedNot supportedNot supported11Not supportedNot supported
gpt-oss-120bNot supportedNot supported21Not supportedNot supportedNot supported
gpt-oss-20bNot supportedNot supported11Not supportedNot supportedNot supported
llama- nemotron- embed-vl-1b-v211111Not supported
llama-3.1-70b- instruct4Not supportedNot supported21Not supportedNot supported
llama-3.1-8b- instruct1Not supportedNot supported111Not supported
llama-3.1- nemoguard-8b- content-safety11111Not supportedNot supported
llama-3.1- nemoguard-8b- topic-control1Not supportedNot supported11Not supportedNot supported
Llama-3.1- nemotron-70b- instructNot supportedNot supportedNot supported2Not supportedNot supportedNot supported
LLM ProviderLLM NameGPU Models
NVIDIA L40S-48GNVIDIA A100-80GNVIDIA H100-80GNVIDIA H100 NVL-94GNVIDIA H 200-141GNVIDIA RTX PRO 6000-96GNVIDIA B300-268G
llama-3.1- swallow-8b- instruct-v0.11Not supportedNot supported11Not supportedNot supported
Llama-3.1-8b- instruct-pb24h21Not supportedNot supported11Not supportedNot supported
Llama-3.1-70b- instruct-pb24h2Not supportedNot supportedNot supported2Not supportedNot supportedNot supported
llama-3.2-nv- embedqa-1b-v211111Not supportedNot supported
llama-3.2-nv- rerankqa-1b-v211111Not supportedNot supported
Llama-3.2-90b- vision-instructNot supportedNot supportedNot supported21Not supportedNot supported
llama-3.3- nemotron- super-49b-v1Not supportedNot supported221Not supportedNot supported
llama-3.3-70b- instruct4Not supported441Not supportedNot supported
Mistral- nemo-12b- instruct2Not supportedNot supported1Not supportedNot supportedNot supported
mistral-7b- instruct-v0.31Not supportedNot supported11Not supportedNot supported
mixtral-8x7b- instruct-v0.14Not supportedNot supported21Not supportedNot supported
nemoretriever- graphic- elements-v111Not supported11Not supportedNot supported
nemoretriever- ocr-v11Not supported111Not supportedNot supported
nemoretriever- page-elements- v21111Not supportedNot supported
nemoretriever- parse1Not supported111Not supportedNot supported
nemoretriever- table-structure- v11Not supported111Not supportedNot supported
openai/whisper- large-v31Not supported111Not supportedNot supported
LLM ProviderLLM NameGPU Models
NVIDIA L40S-48GNVIDIA A100-80GNVIDIA H100-80GNVIDIA H100 NVL-94GNVIDIA H 200-141GNVIDIA RTX PRO 6000-96GNVIDIA B300-268G
phi-3-mini-4k- instruct1Not supportedNot supported11Not supportedNot supported
OpenAIopenai/gpt- oss-120bNot supported111111
openai/gpt- oss-20b1111111
openai/gpt-oss- safeguard-120b2111111
openai/gpt-oss- safeguard-20b1111111
Stability AIstable-diffusion- v1-5/stable- diffusion-v1-51111111
Unslothunsloth/ Llama-3.3-70B- Instruct-bnb-4bitNot supported1111Not supported1

The following table lists the minimum resources required to host an inference endpoint based on the inference engine and the number of GPUs.

The minimum worker node pool size must be defined based on the sum of the CPU and memory resources required for all running inference workloads.

The requirements mentioned in this table are suggestive. Deploying larger LLMs requires more resources.

Ensure that you consider the additional resources required by DaemonSets, the Container Storage Interface (CSI) driver, the Container Network Interface (CNI) plugin, and so on.

Table 8: Minimum GPU Endpoint Requirements

Type of Inference EngineResource RequirementNumber of GPUs 1248
NVIDIA NIMCPU (number of cores)12121212
Memory (GiB)32323232
vLLMCPU (number of cores)8888
Memory (GiB)163264128

The following table lists the virtual machine types supported by the GPU model necessary to deploy Nutanix Enterprise AI in public cloud Kubernetes platforms.

[!NOTE] Note: For the latest virtual machine types supported by a GPU model, see the respective cloud service provider documentation.

Table 9: GPU Supported Virtual Machine Type

Property 1Cloud Service Provider Platform NameProperty 3GPU Supported Virtual Machine TypeGPU Model
AWSAmazon Elastic Kubernetes Service (EKS)EC2 G6eNVIDIA Ada Lovelace L40S
EC2 P5NVIDIA Hopper H100
EC2 P4NVIDIA Ampere A100
EC2 P5eNVIDIA Hopper H 200
EC2 p6-b300NVIDIA Blackwell B300
GCPGoogle Kubernetes Engine (GKE)A3 VMNVIDIA Hopper H100
A2 VMNVIDIA Ampere A100
AzureAzure Kubernetes Service (AKS)NCads_H100_v5-seriesNVIDIA Hopper H100
NCCads_H100_v5- series
NC_A100_v4-seriesNVIDIA Ampere A100
Nutanix Enterprise AI - Agent Gateway Requirements

This section describes the requirements for deploying Nutanix Enterprise AI- agent gateway.

General Requirements

For minimum software requirements to deploy Nutanix Enterprise AI in your environment, see

Nutanix

Enterprise AI - Agent Gateway Software Requirements

in Nutanix Enterprise AI Release Notes .

Ensure that you deploy a Kubernetes cluster with version 1.35.

You can deploy Nutanix Enterprise AI only on the following Cloud Native Computing Foundation (CNCF) - certified Kubernetes distributions:

Nutanix Kubernetes Platform (NKP)

You can install NAI 2.8 only on NKP 2.18.

Nutanix Enterprise AI supports air-gapped installation only on NKP.

Amazon Elastic Kubernetes Service (EKS)

Azure Kubernetes Service (AKS)

Google Kubernetes Engine (GKE)

For information on how to configure a worker node pool on an NKP cluster, see

Configuring Node Pools

in the

Nutanix Kubernetes Platform Guide .

Ensure that you configure the LoadBalancer service based on your Kubernetes cluster.

For example, if you have an NKP cluster, ensure that you configure the MetalLB service in the cluster. For more information, see

MetalLB

in the Nutanix Kubernetes Platform Guide .

The LoadBalancer service assigns an external IP address to the Envoy Ingress Gateway service in Nutanix Enterprise AI that external clients can use to connect to Nutanix Enterprise AI.

Ensure that you configure a fully qualified domain name (FQDN) on the DNS domain that is accessible to your Kubernetes cluster using the external IP address of the Envoy Ingress Gateway service in Nutanix Enterprise AI.

The FQDN is necessary to connect to Nutanix Enterprise AI.

Ensure that you have a certificate authority (CA) signed TLS certificate for the configured FQDN.

The TLS certificate is necessary to update the Envoy Ingress gateway after you install Nutanix Enterprise AI in the cluster.

[!NOTE] Note: Nutanix Enterprise AI only supports HTTPS based connections with a valid TLS certificate.

Sizing Requirements

To deploy Nutanix Enterprise AI on a Kubernetes cluster in your environment, ensure that the cluster meets the minimum requirements. Additional worker nodes are required for your environment and vary based on additional cluster workloads and high-availability requirements.

The following table outlines the minimum compute and local storage recommendation validated for a baseline deployment of Nutanix Enterprise AI - agent gateway.

[!NOTE] Note: To determine the cloud VM instance type based on the requirements mentioned in the following table, see the respective cloud service provider documentation.

Table 10: Minimum Compute and Storage Requirements for Deploying Nutanix Enterprise AI - agent gateway

Node Type Description NumberProperty 2of Nodes (VM)vCPU per NodeMemory per NodeStorage per NodeTotal vCPUTotal Memory
Control PlaneThese control plane nodes are used for running the Kubernetes Control plane.3416 GB150 GB1248 GB
WorkerThese worker nodes are for running the NAI Control Plane.31020 GB150 GB3060 GB

The node sizing assumes the following environmental conditions:

The NKP workload cluster is dedicated exclusively to NAI and its dependencies.

Worker nodes host only the NAI management plane.

Inference-related resources are excluded from these requirements.

The above setup is validated for 100 API Keys and 15 concurrent users. For power usage, scale up the worker nodes to the appropriate sizes in the table below, then redeploy NAI using the profile name. To deploy on:

NKP:

Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform

on page 46

EKS:

Deploying Nutanix Enterprise AI on Amazon Elastic Kubernetes Service

on page 72

Deploying Nutanix Enterprise AI on Azure Kubernetes Service

AKS :

on page 82

Deploying Nutanix Enterprise AI on Google Kubernetes Engine

GKE:

on page 90

Table 11: Compute and Storage Requirements for Profile-based deployment of Nutanix Enterprise AI - agent gateway

API Keys ConcurrentRequestsProfile nameNode TypevCPU per NodeMemory per NodeStorage per NodeTotal vCPUTotal Memory
2001000c1k_k200Worker6090 GB150 GB180270 GB
10005000c5k_k1kWorker90150 GB150 GB270450 GB

The following table lists the minimum compute requirements for each platform component.

The following table specifies the minimum vCPU resources required for the Model Controller component in millicores. 500m means 500 millicores, which is equivalent to 0.5 vCPU cores.

Table 12: Minimum Compute Requirements By Component

ComponentApp NameMinimum vCPU ResourcesMinimum Memory ResourcesPurpose
ClickHouse Serverchi-nai-clickhouse- server48 GiBObservability data store
ClickHouse Serverchk-nai-clickhouse- keeper100m1 GiBObservability data store
IAMiam-database- bootstrap200m0.12 GiBUser Management
IAMiam-proxy100m0.12 GiBUser Management
IAMiam-proxy-control- plane100m0.0625 GiBUser Management
IAMiam-themis200m0.0625GiBUser Management
IAMiam-themis- bootstrap100m0.031 GiBUser Management
IAMiam-ui150m0.0625 GiBUser Management
IAMiam-user-authn100m0.015 GiBUser Management
ComponentApp NameMinimum vCPU ResourcesMinimum Memory ResourcesPurpose
APInai-api5.17.128 GiBAPI Endpoint
NAI DB Migratornai-api-db-migrate11.00 GiBUpgrade NAI DB
DBnai-db-iep22.00 GiBApplication Data per instance
Model Controllernai-iep-model- controller500m0.49 GiBLLM Model Downloader Operator
NAI Labsnai-labs500m2.00 GiBChat & Talk to My Data apps
NAI Labsnai-agent12.00 GiBSample Agent App
Oauth2-Proxynai-oauth2-proxy100m0.3 GiBUser Management
OIDC Client Registrationnai-oidc-client- registration11.00 GiBUser Management
ClickHouse Operatornai-operators-nai- clickhouse-operator500m0.25 GiBObservability data store
OTEL Collectornai-otel-collector- collector42.00 GiBMetrics exporter from Nodes to ClickHouse
Frontend UInai-ui22.00 GiBFrontend
IAM & NAI Gateway Cachenai-valkey100 m0.0625 GiBCache store per instance
IAM & NAI Gateway Cachenai-valkey-sentinel25 m0.03 GiBCache store per instance
NAI Gatewayai-gateway- controller200m0.25 GiBControl Plane for AI Gateway Resources
NAI Gatewayenvoy-gateway100m0.25 GiBControl Plane for Envoy-Proxy
NAI Gatewayenvoy-nai-system- nai-ingress- gateway5.110.32 GiBData Plane Proxy
NAI Gatewayenvoy-ratelimit100m0.5 GiBRatelimit Service
Securitynai-securityscan- manager200m128 MiBModel and Endpoint scanning
NAInai-go-processor210m266 MiBBatch Inference processor per job
NAInai-audit-logs- rsyslog-otel- collector31 GiBAudit Log Exporter

The following table lists the minimum storage requirements for each platform component.

Table 13: Minimum Storage Requirements By Component

ComponentApp NameMinimum Persistent StorageAccess ModePurpose
Databasenai-db-iep-140 GiBRWOApplication data stored in PostgreSQL database
ClickHousechi-nai-clickhouse- server50 GiBRWOObservability Metrics store
Valkeynai-valkey8 GiBRWOIAM and NAI Gateway cache storage per instance
NAInai-audit-logs- rsyslog-otel- collector1 GiBRWOAudit Logs exporter
ClickHouse Keeper chk-nai-clickhouse-10 GiBRWOObservability Metrics store
keeper-chkeeper
NAI Labsnai-labs20 GiBRWONAI Labs data

Nutanix Enterprise AI - agent gateway uses RWO storage to save the application and metrics data.

The following table lists the minimum storage sizing recommendation.

Table 14: Minimum Storage Sizing Recommendations

ComponentMinimum CapacityStorage TechnologyPurpose
RWO (Block - RWO)130 GiBBlockBackend (ClickHouse, PostgreSQL, Valkey, Observability)

[!NOTE] Note: If you are upgrading NAI, make sure to expand the Block storage to at least 105#GiB. NAI Agent Gateway only deployments do not require RWS and S3 based storage.

Nutanix Enterprise AI Limitations

The following limitations apply when you deploy Nutanix Enterprise AI at your site.

Nutanix Enterprise AI supports Hugging Face and NVIDIA NIM formatted LLMs only.

Nutanix Enterprise AI supports x86 architecture only.

Deploy Nutanix Enterprise AI

Deploy Nutanix Enterprise AI (NAI) on supported Kubernetes platforms and environments.

If you are installing or upgrading Nutanix Enterprise AI, do one of the following:

Deploy NAI on Nutanix Kubernetes Platform (NKP)

Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform

For more information, see

on page 41.


Last updated Oct 08, 2026