▸ Agent Skills
68 min read

Nutanix Enterprise AI Manual: Deployment and Installation Guide

Nutanix Docker Hub access tokens, Helm chart configuration parameters for nai-operators and nai-core, connected deployment on Nutanix Kubernetes Platform (NKP), air-gapped NKP deployment with private registries, deployment on Amazon EKS, Azure AKS, Google GKE, self-managed PostgreSQL configurations, dashboard IP access, and Docker registry credentials rotation.


Deploy NAI on NKP in an air#gapped environment

Deploying Nutanix Enterprise AI on a Nutanix Kubernetes Platform Cluster in

For more information, see

Air-Gapped Environments

on page 50.

Deploy NAI on Amazon Elastic Kubernetes Service (EKS). For more information, see

Deploy Nutanix

Enterprise AI on Amazon Elastic Kubernetes Service

on page 67.

Deploy NAI on Azure Kubernetes Service (AKS)

For more information, see

Deploying Nutanix Enterprise AI on Azure Kubernetes Service

on page 76.

Deploy NAI on Google Kubernetes Engine (GKE)

For more information, see

Deploy Nutanix Enterprise AI on Google Kubernetes Engine

on page 85.

Generating Nutanix Docker Hub Access Tokens

Generate tokens that you can use to access the Nutanix Docker Hub private repository.

Before you begin

If your environment is air-gapped, you cannot use this procedure. For more information on the procedure for an air-gapped environment, see

Downloading Product Files

.

About this task

To generate Nutanix Docker Hub Access Tokens, follow these steps:

Procedure

  1. Log in to the Nutanix Support portal.

2. In the upper-left corner of the Nutanix Support Portal, click the

menu

Downloads

icon, then click

.

3. Click

Nutanix Enterprise AI

.

4. Click

Generate Access Token

.

You can generate a maximum of two tokens.

Manage Docker Hub Access Token

The

dialog box appears and generates a token. Once generated, this

token remains valid for all Nutanix Enterprise AI (NAI) releases. You only need to generate a replacement if the existing token is manually deleted or revoked.

What to do next

For information on how to use your tokens to download and install Nutanix Enterprise AI, see

Deploy

Nutanix Enterprise AI

on page 27.

Nutanix Enterprise AI Configuration Parameters for the

Helm Chart

nai-operators

The following tables list the configurable parameters of the

helm chart and their default
nai-operators

values.

Table 15: Global Parameters

KeyDescriptionDefault Value
imagePullSecretsList of Secrets for Container Registry hosting NAI Images

Table 16: Global Storage Parameters

KeyDescriptionProperty 3Default Value
global.storage.storageClassName Default RWO storage class usednutanix-volume
by persistent components. For
nai-operators, this is used by
CNPG cluster and Valkey PVCs when a cluster-specific storage class is empty.

Table 17: AI Gateway Parameters

KeyDescriptionDefault Value
ai-gateway- helm.extProc.image.repositoryRepository for extProc imagedocker.io/nutanix/nai-ai-gateway- extproc
ai-gateway- helm.extProc.image.tagTag for extProc imagebc729717
ai-gateway- helm.controller.image.repositoryRepository for controller imagedocker.io/nutanix/nai-ai-gateway- controller
ai-gateway- helm.controller.image.tagTag for controller imagebc729717

Table 18: Valkey Parameters

KeyDescriptionDefault Value
naiValkey.image.nameValkey image namedocker.io/nutanix/nai-valkey
naiValkey.image.tagValkey image tag9.1.0
naiValkey.image.pullPolicyImage pull policy for ValkeyIfNotPresent
naiValkey.replicaCountNumber of Valkey data nodes1
naiValkey.sentinel.replicaCountNumber of Valkey Sentinel instances1
naiValkey.sentinel.resources.limits.cpuSentinel CPU limit50m
naiValkey.sentinel.resources.limits.memorySentinel memory limit64Mi
naiValkey.sentinel.resources.requests.cpuSentinel CPU request25m
naiValkey.sentinel.resources.requests.memorySentinel memory request32Mi
naiValkey.sentinel.topologySpread.whenUnsatisfiableSentinel topology spread scheduling policyDoNotSchedule
naiValkey.sentinel.pdb.enabledEnable Pod Disruption Budget for Sentinel podsfalse
naiValkey.sentinel.pdb.minAvailableMinimum number of Sentinel pods
that must remain available
naiValkey.persistence.sizePersistent storage size for Valkey data4Gi
KeyDescriptionDefault Value
naiValkey.persistence.storageClassStorage class for Valkey
persistent volumes. Empty falls back to global.storage.storageClassName, then the cluster default StorageClass.
naiValkey.topologySpread.whenUnsatisfiableValkey data pod topology spread scheduling policyDoNotSchedule
naiValkey.pdb.enabledEnable Pod Disruption Budget for Valkey data podsfalse
naiValkey.pdb.minAvailableMinimum number of Valkey data pods that must remain available
naiValkey.resources.limits.cpuValkey CPU limit200m
naiValkey.resources.limits.memory Valkey memory limit256Mi
naiValkey.resources.requests.cpu Valkey CPU request100m
naiValkey.resources.requests.memoryValkey memory request64Mi

Table 19: Clickhouse Operators Parameters

KeyDescriptionDefault Value
nai-clickhouse- operator.operator.image.registryContainer image name or tagdocker.io
nai-clickhouse- operator.operator.image.repositoryContainer image name or tagnutanix/nai-clickhouse-operator
nai-clickhouse- operator.operator.image.tagContainer image name or tag0.24.2
nai-clickhouse- operator.metrics.image.registryContainer image name or tagdocker.io
nai-clickhouse- operator.metrics.image.repositoryContainer image name or tagnutanix/nai-clickhouse-metrics- exporter
nai-clickhouse- operator.metrics.image.tagContainer image name or tag0.24.2

Table 20: NAI Database Parameters

KeyDescriptionProperty 3Property 4Property 5Property 6Default Value
naiDatabase.externalUse an externally managed PostgreSQL database instead of provisioning CNPG clusters. The same value should be supplied to bothfalse
nai-operatorsandnai-
core.
KeyDescriptionDefault Value
naiDatabase.imagePostgreSQL operand image used by CloudNativePGdocker.io/nutanix/nai- postgresql:17.10-standard-trixie
naiDatabase.postgresql.parameters.shared_buffersPostgreSQL shared buffer size1GB
naiDatabase.postgresql.parameters.work_memPostgreSQL per-operation work memory8MB
naiDatabase.postgresql.parameters.idle_in_transaction_session_timeoutMaximum time an idle transaction can remain open5min
naiDatabase.postgresql.parameters.idle_session_timeoutMaximum time an idle database session can remain inactive5min
naiDatabase.affinity.enablePodAntiAffinityEnable PostgreSQL pod anti- affinitytrue
naiDatabase.affinity.podAntiAffinityTypePostgreSQL pod anti-affinity policyrequired
naiDatabase.affinity.topologyKeyTopology key used for PostgreSQL pod anti-affinitykubernetes.io/hostname
naiDatabase.enablePDBEnable the CNPG-managed Pod Disruption Budgettrue
naiDatabase.clusters.iep.clusterNameName of the CNPG cluster for the NAI application databasenai-db-iep
naiDatabase.clusters.iep.database Application database namenai_iep
naiDatabase.clusters.iep.username Application database usernamenai-api-user
naiDatabase.clusters.iep.password Application database passwordnai-api-password
naiDatabase.clusters.iep.sslMode SSL mode for NAI clientdisable
connections to PostgreSQL
naiDatabase.clusters.iep.sslRootCertNameClient CA certificate file name
naiDatabase.clusters.iep.sslClientCertNameClient certificate file name
naiDatabase.clusters.iep.sslClientKeyNameClient private key file name
naiDatabase.clusters.iep.hostPostgreSQL read-write host advertised to NAI clientsnai-db-iep-rw.nai-system
naiDatabase.clusters.iep.portPostgreSQL port5432
naiDatabase.clusters.iep.instances Number of PostgreSQL instances1
in the CNPG cluster
naiDatabase.clusters.iep.maxConnectionsMaximum PostgreSQL connections1000
naiDatabase.clusters.iep.synchronous.enabledEnable synchronous PostgreSQL replicationfalse
naiDatabase.clusters.iep.synchronous.methodSynchronous replication selection methodany
naiDatabase.clusters.iep.synchronous.numberNumber of synchronous standbys required1
naiDatabase.clusters.iep.synchronous.dataDurabilitySynchronous replication durability policypreferred
KeyDescriptionDefault Value
naiDatabase.clusters.iep.resources.requests.cpuPostgreSQL CPU request2
naiDatabase.clusters.iep.resources.requests.memoryPostgreSQL memory request2Gi
naiDatabase.clusters.iep.resources.limits.cpuPostgreSQL CPU limit4
naiDatabase.clusters.iep.resources.limits.memoryPostgreSQL memory limit4Gi
naiDatabase.clusters.iep.storage.sizePersistent storage size for PostgreSQL20Gi
naiDatabase.clusters.iep.storage.storageClassStorage class for PostgreSQL persistent volumes. Empty falls back to global.storage.storageClassName.
naiDatabase.clusters.iep.import.typeType of legacy PostgreSQLmonolith
import
naiDatabase.clusters.iep.import.sourceHostHostname of the legacy standalone PostgreSQL database used for importnai-db
naiDatabase.clusters.iep.import.sourcePortPort of the legacy PostgreSQL database5432
naiDatabase.clusters.iep.import.databasesDatabases imported from the legacy PostgreSQL instancenai_iep, nai_iam
naiDatabase.clusters.iep.import.rolesPostgreSQL roles imported from the legacy instancenai-api-user
naiDatabase.clusters.iep.import.sourceDatabaseDatabase used by the import session on the source PostgreSQL instancenai_iep
naiDatabase.clusters.iam.inCluster CNPG cluster hosting the IAMnai-db-iep
database
naiDatabase.clusters.iam.database IAM database namenai_iam
naiDatabase.clusters.iam.usernameIAM database usernamenai-api-user
naiDatabase.clusters.iam.passwordIAM database passwordnai-api-password
naiDatabase.clusters.iam.sslMode SSL mode for IAM clientdisable
connections to PostgreSQL
naiDatabase.clusters.iam.sslRootCertNameIAM client CA certificate file name
naiDatabase.clusters.iam.sslClientCertNameIAM client certificate file name
naiDatabase.clusters.iam.sslClientKeyNameIAM client private key file name
naiDatabase.clusters.iam.hostPostgreSQL read-write host used by IAM servicesnai-db-iep-rw.nai-system
naiDatabase.clusters.iam.portPostgreSQL port used by IAM services5432

Table 21: NAI Jobs Parameters

Property 1KeyDescriptionProperty 4Default ValueProperty 6
naiJobs.naiJobsImage.imageJob service imagedocker.io/nutanix/nai-jobs
naiJobs.naiJobsImage.tagJob service image tagv2.8.0
naiJobs.resources.limits.cpuCPU limit for nai-jobs container1
naiJobs.resources.limits.memoryMemory limit for nai-jobs container1Gi
naiJobs.resources.requests.cpuCPU request for nai-jobs container100m
naiJobs.resources.requests.memoryMemory request for nai-jobs50Mi
container
Nutanix Enterprise AI Configuration Parameters for thenai-coreHelm chart

The following tables lists the configurable parameters of the

Helm chart and their default values.

nai-core

Table 22: Global Parameters

KeyDescriptionDefault Value
imagePullSecretsList of Secrets for Container Registry hosting NAI Images

Table 23: Clickhouse keeper

KeyDescriptionDefault Value
nai-clickhouse- keeper.clickhouseKeeper.image.registryClickhouse keeper registrydocker.io
nai-clickhouse- keeper.clickhouseKeeper.image.repositoryClickhouse keeper repositorynutanix/nai-clickhouse-keeper
nai-clickhouse- keeper.clickhouseKeeper.image.tagClickhouse keeper image tag25.8.17.37
nai-clickhouse- keeper.clickhouseKeeper.storage.storageClassClickhouse keeper storage classnutanix-volume

Table 24: Clickhouse Server

KeyDescriptionDefault Value
nai-clickhouse- server.clickhouse.image.registryContainer image name or tagdocker.io
nai-clickhouse- server.clickhouse.image.repositoryContainer image name or tagnutanix/nai-clickhouse-server
KeyDescriptionDefault Value
nai-clickhouse- server.clickhouse.image.tagContainer image name or tag25.8.17.37
nai-clickhouse- server.clickhouse.initContainers.addUdf.image.registryContainer image name or tagdocker.io
nai-clickhouse- server.clickhouse.initContainers.addUdf.image.repositoryContainer image name or tagnutanix/nai-clickhouse-udf
nai-clickhouse- server.clickhouse.initContainers.addUdf.image.tagContainer image name or tagv2.8.0
nai-clickhouse- server.clickhouse.initContainers.waitForKeeper.image.registryContainer image name or tagdocker.io
nai-clickhouse- server.clickhouse.initContainers.waitForKeeper.image.repositoryContainer image name or tagnutanix/nai-jobs
nai-clickhouse- server.clickhouse.initContainers.waitForKeeper.image.tagContainer image name or tagv2.8.0
nai-clickhouse- server.clickhouse.resources.limits.cpuCPU resource specification
nai-clickhouse- server.clickhouse.resources.limits.memoryMemory resource specification8Gi
nai-clickhouse- server.clickhouse.resources.requests.cpuCPU resource specification
nai-clickhouse- server.clickhouse.resources.requests.memoryMemory resource specification8Gi
nai-clickhouse- server.clickhouse.storage.pvcStoragePersistent volume claim size50Gi
nai-clickhouse- server.clickhouse.storage.storageClassStorage class namenutanix-volume
nai-clickhouse- server.clickhouse.users.admin.passwordPassword valuenai-clickhouse-password
nai-clickhouse- server.clickhouse.users.admin.usernameUsername valuenai-clickhouse-user

Table 25: Storage Parameters

KeyDescriptionDefault Value
defaultStorageClassNameStorage class name to be used by nai-db and ClickHouse server.

Table 26: NAI IEP Operator Parameters

KeyDescriptionDefault Value
naiIepOperator.iepOperatorImage.imageIEP operator image namedocker.io/nutanix/nai-iep-operator
KeyDescriptionDefault Value
naiIepOperator.iepOperatorImage.tagIEP operator image tagv2.8.0
naiIepOperator.iepOperatorResources.limits.cpuIEP operator CPU limits500m
naiIepOperator.iepOperatorResources.limits.memoryIEP operator memory limits500Mi
naiIepOperator.iepOperatorResources.requests.cpuIEP operator CPU requests100m
naiIepOperator.iepOperatorResources.requests.memoryIEP operator memory requests100Mi
naiIepOperator.modelProcessorImage.imageModel Processor image namedocker.io/nutanix/nai-python- processor
naiIepOperator.modelProcessorImage.tagModel Processor image tagv2.8.0
naiIepOperator.modelProcessorResources.limits.cpuModel Processor CPU limits
naiIepOperator.modelProcessorResources.limits.memoryModel Processor memory limits
naiIepOperator.modelProcessorResources.requests.cpuModel Processor CPU requests
naiIepOperator.modelProcessorResources.requests.memoryModel Processor memory requests
naiIepOperator.retainModelProcessorJobDetermines whether the model processor job must be retained after completion.false
naiIepOperator.dataSourceProcessorImage.imageData Source Processor Image namedocker.io/nutanix/nai-python- processor
naiIepOperator.dataSourceProcessorImage.tagData source Processor image tagv2.8.0
naiIepOperator.dataSourceProcessorResources.limits.cpu / Data source ProcessorCPU limits
naiIepOperator.dataSourceProcessorResources.limits.memoryData source Processor memory limits
naiIepOperator.dataSourceProcessorResources.requests.cpuData source Processor CPU requests
naiIepOperator.dataSourceProcessorResources.requests.memoryData source Processor memory requests
naiIepOperator.finetuneProcessorImage.imageFinetune Processor image namedocker.io/nutanix/nai-finetuning
naiIepOperator.finetuneProcessorImage.tagFinetune Processor image tagv2.8.0
naiIepOperator.batchInferenceProcessorImage.containers.processor.imageBatch Inference Processor image namedocker.io/nutanix/nai-go- processor
naiIepOperator.batchInferenceProcessorImage.containers.processor.tagBatch Inference Processor image tagv2.8.0
naiIepOperator.batchInferenceProcessorImage.containers.processor.resources.limits.cpuBatch Inference Processor CPU limits200m
naiIepOperator.batchInferenceProcessorImage.containers.processor.resources.limits.memoryBatch Inference Processor memory limits256Mi
naiIepOperator.batchInferenceProcessorImage.containers.processor.resources.requests.cpuBatch Inference Processor CPU limits200m
naiIepOperator.batchInferenceProcessorImage.containers.processor.resources.requests.memoryBatch Inference Processor CPU memory256Mi
KeyDescriptionDefault Value
naiIepOperator.batchInferenceProcessorImage.containers.statusProvider.imageBatch Inference Status Provider image namedocker.io/nutanix/nai-go- processor
naiIepOperator.batchInferenceProcessorImage.containers.statusProvider.tagBatch Inference Status Provider image tagv2.8.0
naiIepOperator.batchInferenceProcessorImage.containers.statusProvider.resources.limits.cpuBatch Inference Status Provider CPU limits200m
naiIepOperator.batchInferenceProcessorImage.containers.statusProvider.resources.limits.memoryBatch Inference Status Provider CPU memory256Mi
naiIepOperator.batchInferenceProcessorImage.containers.statusProvider.resources.requests.cpuBatch Inference Status Provider CPU limits200m
naiIepOperator.batchInferenceProcessorImage.containers.statusProvider.resources.requests.memoryBatch Inference Status Provider CPU memory256Mi

Table 27: NAI Inference UI Parameters

KeyDescriptionDefault Value
naiInferenceUi.naiUiImage.imageThe name of the NAI UI Docker imagedocker.io/nutanix/nai-inference-ui
naiInferenceUi.naiUiImage.tagThe tag of the Docker image to be usedv2.8.0
naiInferenceUi.resources.requests.cpuNAI UI container CPU requests1
naiInferenceUi.resources.requests.memoryNAI UI container memory requests1Gi
naiInferenceUi.resources.limits.cpu NAI UI container CPU limits2
naiInferenceUi.resources.limits.memoryNAI UI container memory limits2Gi

Table 28: NAI API Parameters

KeyDescriptionDefault Value
naiApi.naiApiImage.imageThe name of the NAI API Docker imagedocker.io/nutanix/nai-api
naiApi.naiApiImage.tagThe tag of the Docker image to be usedv2.8.0
naiApi.naiApiResources.requests.cpuNAI API container CPU requests2
naiApi.naiApiResources.requests.memoryNAI API container memory requests2Gi
naiApi.naiApiResources.limits.cpu NAI API container CPU limits8
naiApi.naiApiResources.limits.memoryNAI API container memory limits4Gi
naiApi.naiMigrateInitContainerResources.requests.cpuNAI API migrate init container CPU requests100m
KeyDescriptionDefault Value
naiApi.naiMigrateInitContainerResources.requests.memoryNAI API migrate init container memory requests50Mi
naiApi.naiMigrateInitContainerResources.limits.cpuNAI API migrate init container CPU limits1
naiApi.naiMigrateInitContainerResources.limits.memoryNAI API migrate init container memory limits1Gi
naiApi.naiMigrateJobResources.requests.cpuNAI API migrate job container CPU requests1
naiApi.naiMigrateJobResources.requests.memoryNAI API migrate job container memory requests1Gi
naiApi.naiMigrateJobResources.limits.cpuNAI API migrate job container CPU limits1
naiApi.naiMigrateJobResources.limits.memoryNAI API migrate job container memory limits1Gi
naiApi.storageClassNameStorage class name to be used nai-iep for storing modelsnai-nfs-storage
naiApi.supportedTGIImageSupported TGI Runtime Imagedocker.io/nutanix/nai-tgi
naiApi.supportedTGIImageTagSupported TGI Runtime Image tag3.3.4-b2485c9
naiApi.replicaCountNumber of instances of nai-api1

Table 29: NAI Database Parameters

KeyDescriptionProperty 3Property 4Property 5Property 6Property 7Property 8Default Value
naiDatabase.externalUse an externally managed PostgreSQL database. When true,false
nai-operatorsdoes not
provision CNPG clusters. Pass the same value to both charts.
naiDatabase.clientImagePostgreSQL client image containingdocker.io/nutanix/nai- postgresql:17.10-standard-trixie
psqlandpg_isready,
used by the IAM database bootstrap job
naiDatabase.storageSpec.resources.requests.storagePersistent storage size for the legacy standalone PostgreSQL data PVC (4Gi
nai-db). The value
must match the existing PVC size during migration.
naiDatabase.clusters.iep.database Application database name usednai_iep
by NAI core services
naiDatabase.clusters.iep.portPostgreSQL port for the application database5432
naiDatabase.clusters.iep.hostRead-write PostgreSQL host for the application databasenai-db-iep-rw.nai-system
KeyDescriptionDefault Value
naiDatabase.clusters.iep.sslMode SSL mode for applicationdisable
database client connections
naiDatabase.clusters.iep.sslSecretNameKubernetes Secret containing client TLS certificate material when requirednai-db-certs
naiDatabase.clusters.iep.sslRootCertNameClient CA certificate file name
naiDatabase.clusters.iep.sslClientCertNameClient certificate file name
naiDatabase.clusters.iep.sslClientKeyNameClient private key file name
naiDatabase.clusters.iam.database IAM database name used by IAMnai_iam
services
naiDatabase.clusters.iam.portPostgreSQL port for the IAM database5432
naiDatabase.clusters.iam.hostRead-write PostgreSQL host for the IAM databasenai-db-iep-rw.nai-system
naiDatabase.clusters.iam.sslMode SSL mode for IAM database clientdisable
connections
naiDatabase.clusters.iam.sslSecretNameKubernetes Secret containing client TLS certificate material when requirednai-db-certs
naiDatabase.clusters.iam.sslRootCertNameIAM client CA certificate file name
naiDatabase.clusters.iam.sslClientCertNameIAM client certificate file name
naiDatabase.clusters.iam.sslClientKeyNameIAM client private key file name

[!NOTE] Note:

The

chart owns the database credentials Secrets.

contains the non-secret

nai-operators
nai-core

connection topology used to render service configuration.

For an internal CNPG deployment, the

and

databases are co-located in the

iep
iam
nai-db-iep

CNPG cluster by default.

For an external PostgreSQL deployment, configure

consistently in both

naiDatabase.external

charts. Configure the corresponding database connection details in both charts as described earlier.

The

chart automatically detects legacy PostgreSQL imports using Helm lookup. Use

nai-operators

the

setting only for advanced or break-glass scenarios.

importOverride

Component-specific

values are intentionally empty by default so that they inherit

storageClass

.

global.storage.storageClassName

Table 30: NAI Monitoring Parameters

KeyDescriptionDefault Value
naiMonitoring.opentelemetry.collectorImageThe name and tag of the Target Collector Docker imagedocker.io/nutanix/nai- opentelemetry-collector- contrib:0.141.0
naiMonitoring.opentelemetry.targetAllocator.image.repositoryThe name of the Target Allocator Docker image.docker.io/nutanix/nai-target- allocator
naiMonitoring.opentelemetry.targetAllocator.image.tagThe tag of the Target Allocator Docker image0.141.0
naiMonitoring.opentelemetry.targetAllocator.resources.requests.cpuTarget Allocator container CPU requests0.5
naiMonitoring.opentelemetry.targetAllocator.resources.requests.memoryTarget Allocator container memory request100Mi
naiMonitoring.opentelemetry.targetAllocator.resources.limits.cpuTarget Allocator container CPU limits1
naiMonitoring.opentelemetry.targetAllocator.resources.limits.memoryTarget Allocator container memory limits500Mi
naiMonitoring.opentelemetry.storageClassNameStorage class name to be used by Opentelemetry componentsnai-nfs-storage
naiMonitoring.opentelemetry.common.resources.requests.cpuOpentelemetry container CPU requests0.1
naiMonitoring.opentelemetry.common.resources.requests.memoryOpentelemetry container memory requests500Mi
naiMonitoring.opentelemetry.common.resources.limits.cpuOpentelemetry container CPU limits4
naiMonitoring.opentelemetry.common.resources.limits.memoryOpentelemetry container memory limits2Gi
naiMonitoring.opentelemetry.common.storageSpec.resources.requests.storageStorage spec for Opentelemetry data1Gi
naiMonitoring.opentelemetry.rsyslogAuditLogsExport.resources.requests.cpuRsyslog container CPU requests2
naiMonitoring.opentelemetry.rsyslogAuditLogsExport.resources.requests.memoryRsyslog container memory requests450Mi
naiMonitoring.opentelemetry.rsyslogAuditLogsExport.resources.limits.cpuRsyslog container CPU limits3
naiMonitoring.opentelemetry.rsyslogAuditLogsExport.resources.limits.memoryRsyslog container memory limits1Gi
naiMonitoring.opentelemetry.rsyslogAuditLogsExport.storageSpec.resources.requests.storageStorage spec for RsysLog data1Gi
naiMonitoring.nodeExporter.serviceMonitor.enabledEnable node exporter service monitorfalse
naiMonitoring.nodeExporter.serviceMonitor.namespaceSelector.matchNames[0]Namespace in which Kube- Prometheus-Stack is installedprometheus
naiMonitoring.dcgmExporter.podLevelMetricsEnable DCGM exporter pod-level metricsfalse
naiMonitoring.dcgmExporter.serviceMonitor.enabledEnable DCGM exporter service monitorfalse
naiMonitoring.dcgmExporter.serviceMonitor.namespaceSelector.matchNames[0]Namespace in which NVIDIA GPU Operator is installedgpu-operator

Table 31: NAI Labs

KeyDescriptionValue
naiLabs.labsImage.imageThe name of the NAI Labs Docker image.docker.io/nutanix/nai-rag-app
naiLabs.labsImage.tagThe tag of the Docker image to be usedv2.8.0
naiLabs.resources.requests.memoryNAI Labs memory requests2Gi
naiLabs.resources.requests.cpuNAI Labs container CPU requests 500m
naiLabs.resources.limits.memoryNAI Labs memory limits6Gi
naiLabs.resources.limits.cpuNAI Labs container CPU limits1500m
naiLabs.storageSpec.resources.requests.storageStorage spec for NAI Labs data20Gi
naiLabs.vectorDb.externalSet external: true, to use an external Milvus database instead of embedded ChromaDBfalse
naiLabs.vectorDb.milvus.hostMilvus server hostname or IP address
naiLabs.vectorDb.milvus.portMilvus server port (self hosted Milvus uses 19530, Zilliz Cloud uses 443)19530
naiLabs.vectorDb.milvus.tokenAuthentication token (leave empty if not required)
naiLabs.vectorDb.milvus.milvusTLSEnabledTLS Configuration (one-way TLS only).false
Set to true if your Milvus server requires TLS/SSL.

[!NOTE] Note: Only one-way TLS is supported (client verifies server certificate)

naiLabs.vectorDb.milvus.milvusTLSSecretName

K8s secret containing the CA certificate for SSL connection to be created in the nai-system namespac

nai-milvus-tls-certs

naiLabs.vectorDb.milvus.milvusTLSCACertName

CA certificate key name within the secret . For example, “ca.pem”.

Leave empty if your Milvus uses a publicly trusted CA (e.g., Let’s Encrypt, Zilliz Cloud)

naiLabs.vectorDb.milvus.milvusTLSServerName

Hostname for TLS certificate verification (optional).

If left blank, defaults to the host value above for hostname verification

Set this if the certificate’s CN/SAN doesn’t match the host value.

For example, host is an IP address (10.111.48.98) but certificate is issued for a hostname (milvus.example.com)

naiLabs.enabled

Flag to deploy Chat and Talk to Data app

false

Table 32: AI Gateway

KeyDescriptionDefault Value
gateway.replicaCountNumber of instances of nai- ingress-gateway1

Table 33: NAI Agent

Property 1naiAgent.enabledEnable the sample NAI Agent app trueProperty 4
naiAgent.agentImage.imageNAI Agent app image namedocker.io/nutanix/nai-agent-app
naiAgent.agentImage.tagNAI Agent app image tagv2.8.0
naiAgent.resources.requests.memoryNAI Agent app memory requests2Gi
naiAgent.resources.requests.cpuNAI Agent app CPU requests1000m
naiAgent.resources.limits.memory NAI Agent app memory limits4Gi
naiAgent.resources.limits.cpuNAI Agent app CPU limits1500m
Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform

Install or upgrade Nutanix Enterprise AI (NAI) on Nutanix Kubernetes Platform (NKP).

To install or upgrade NAI on NKP, follow these high-level steps:

1. Set up an NKP cluster. For more information, see

Setting up a Nutanix Kubernetes Platform Cluster

on

page 41.

2. Perform preflight checks. For more information, see

Performing Preflight Checks Before Deploying Nutanix

Enterprise AI on Nutanix Kubernetes Platform

on page 42.

Installing Prerequisite Components on a Nutanix

3. Install prerequisite components. For more information, see

Kubernetes Platform Cluster

on page 43.

4. Deploy NAI on NKP. For more information, see

Deploying Nutanix Enterprise AI on Nutanix Kubernetes

Platform

on page 46.

Setting up a Nutanix Kubernetes Platform Cluster Set up a Nutanix Kubernetes Platform (NKP) cluster before you deploy Nutanix Enterprise AI (NAI).

About this task

To set up your NKP cluster, follow these steps:

Procedure

  1. Create an NKP cluster.

Ensure that the Kubernetes version is 1.35and the nodes are Ubuntu based images. For more information, see the

Nutanix Kubernetes® Platform Guide

.

  1. Add the NKP Pro or Ultimate license to your NKP Cluster.

Add an NKP License

For more information, see

in the Nutanix Kubernetes® Platform Guide .

3. Ensure that the NKP cluster meets all the requirements listed in

Nutanix Enterprise AI - Private Inference and

Agent Gateway Requirements

on page 10.

  1. Create a Nutanix Files server with sufficient storage capacity to save the models.

Creating a File Server

For more information, see

in the Nutanix Files User Guide .

For information on the size of the models, see

Table 49: Pre-validated Models

on page 184.

5. Create a Network File System (NFS v4) export in the Nutanix Files server to import a custom model or to import

a model manually.

Creating an NFS Export

For more information, see

in the Nutanix Files User Guide .

6. (Optional) To override the default configuration with a custom CNI plugin, ensure that NetworkPolicy

enforcement is enabled.

By default, NKP)installs the Cilium add-on with the default configuration while creating a cluster.

What to do next

Perform preflight checks. For more information, see

Performing Preflight Checks Before Deploying Nutanix

Enterprise AI on Nutanix Kubernetes Platform

on page 42.

Performing Preflight Checks Before Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform Perform preflight checks before installing or upgrading Nutanix Enterprise AI (NAI) on Nutanix Kubernetes Platform (NKP). Preflight checks ensure the Kubernetes environment meets all the technical requirements for a successful NAI installation.

About this task

To perform preflight checks, follow these steps:

Procedure

1. Verify that the Kubernetes version is 1.33 or 1.34:

kubectl get nodes \
--selector='!node-role.kubernetes.io/control-plane,!node-role.kubernetes.io/master'
\
-o custom-
columns=NODE:.metadata.name,KUBELET_VERSION:.status.nodeInfo.kubeletVersion

The expected output is that the Kubernetes version must be 1.33 or 1.34.

2. Verify if the number of CSI pods matches the number of worker nodes:

[ $(kubectl get nodes --no-headers | wc -l) -eq $(kubectl get pods -n ntnx-system --
no-headers | grep csi-node | wc -l) ] && echo "# CSI pods = node count" || echo "#
Mismatch: CSI pods != node count"

The expected output is CSI pods = node count.

3. Verify that the

class binds volumes immediately:

nai-nfs-storage
kubectl get storageclass nai-nfs-storage -o jsonpath='{.volumeBindingMode}{"\n"}'

The expected output is

.

immediate

4. Verify if the

storage class has ReadWriteMany access:

nai-nfs-storage
kubectl apply -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-rwx-pvc
namespace: default
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Gi
storageClassName: nai-nfs-storage
EOF
kubectl get pvc test-rwx-pvc -n default

The expected output is that the storage class

has

set to

.

nai-nfs-storage
Access Modes
RWX
NAME           STATUS   VOLUME                                     CAPACITY   ACCESS
MODES   STORAGECLASS      VOLUMEATTRIBUTESCLASS   AGE
test-rwx-pvc            Bound    pvc-0df5145a-590c-4f28-9f9d-cefefc55266c   1Gi
RWX            nai-nfs-storage   <unset>                 5s
kubectl delete pvc test-rwx-pvc

Installing Prerequisite Components on a Nutanix Kubernetes Platform Cluster Install prerequisite components on a Nutanix Kubernetes Platform (NKP) cluster.

Before you begin

Complete the following:

  1. Set up an NKP cluster.

Setting up a Nutanix Kubernetes Platform Cluster

For more information, see

on page 41. cert-manager is

deployed when you set up an NKP cluster.

  1. Perform preflight checks.

For more information, see

Performing Preflight Checks Before Deploying Nutanix Enterprise AI on Nutanix

Kubernetes Platform

on page 42.

About this task

To install components on an NKP 2.18 cluster, follow these steps:

Procedure

1. Install or upgrade Envoy Gateway:

a. Install Envoy Gateway and Gateway API CRDs:

helm template eg oci://docker.io/envoyproxy/gateway-crds-helm --version v1.8.1 \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --server-side --force-conflicts -f -

b. Create

with the following configuration:

envoy-gateway-config.yaml
config:
envoyGateway:
gateway:
controllerName: "gateway.envoyproxy.io/gatewayclass-controller"
logging:
level:
default: "info"
provider:
kubernetes:
rateLimitDeployment:
container:
image: "docker.io/envoyproxy/ratelimit:1e50889b"
patch:
type: "StrategicMerge"
value:
spec:
template:
spec:
containers:
- imagePullPolicy: "IfNotPresent"
name: "envoy-ratelimit"
image: "docker.io/envoyproxy/ratelimit:1e50889b"
env:
- name: REDIS_TYPE
value: "sentinel"
- name: REDIS_PIPELINE_WINDOW
value: "150us"
type: "Kubernetes"
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
maxMessageSize: 11Mi
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- "Translation"
- "Cluster"
- "Route"
service:
fqdn:
hostname: "ai-gateway-controller.nai-system.svc.cluster.local"
port: 1063
rateLimit:
backend:
type: "Redis"
redis:
url: "mymaster,nai-valkey-sentinel.nai-system.svc.cluster.local:26379"

[!NOTE] Note: The rate-limit backend uses Valkey Sentinel with master name

mymaster

and the

nai-valkey-

service in

.

sentinel
nai-system

c. Install or Upgrade Envoy Gateway:

helm upgrade --install eg oci://docker.io/envoyproxy/gateway-helm --version v1.8.1
\
-n envoy-gateway-system --create-namespace --skip-crds \
-f "./envoy-gateway-config.yaml"

2. Install or Upgrade KServe:

The required version is KSERVE_VERSION=v0.19.0.

a. Install or upgrade the KServe CRDs:

helm upgrade --install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

b. Install or upgrade the KServe resources:

helm upgrade --install kserve oci://ghcr.io/kserve/charts/kserve-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.controller.deploymentMode=RawDeployment \
--set kserve.controller.gateway.disableIngressCreation=true

c. Install or upgrade the KServe LLMInferenceService CRD:

helm upgrade --install kserve-llmisvc-crd oci://ghcr.io/kserve/charts/kserve-
llmisvc-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

d. Install or upgrade the KServe LLMInferenceService resources:

helm upgrade --install kserve-llmisvc-resources oci://ghcr.io/kserve/charts/
kserve-llmisvc-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.createSharedResources=false \
--set kserve.llmisvc.createGIECRDs=false
  1. Install CloudNativePG Operator from NKP platform applications.

4. Install LeaderWorkerSet:

helm install lws oci://registry.k8s.io/lws/charts/lws \
--version 0.8.0 -n lws-system --create-namespace --wait

5. Install or upgrade OpenTelemetry Operator

helm upgrade --install opentelemetry-operator opentelemetry-operator \
--repo https://open-telemetry.github.io/opentelemetry-helm-charts \
--version=0.114.1 -n opentelemetry --create-namespace --wait

6. Install Prometheus Monitoring from NKP platform applications:

To optimize resource utilization on the workload cluster, configure Prometheus Monitoring with the following minimum installation settings when you enable the application:

alertmanager:
enabled: false
grafana:
enabled: false
prometheus:
enabled: false
kubeStateMetrics:
enabled: false
kubernetesServiceMonitors:
enabled: false
prometheus-node-exporter.kubeRBACProxy:
kubeRBACProxy:
enabled: true

Pro: Enabling an Application Using the UI

For more information, see

.

  1. Install NVIDIA GPU Operator from NKP platform applications.

If the GPU nodes do not have precompiled NVIDIA drivers installed, enable driver installation in the NVIDIA GPU Operator configuration. This setting ensures that the NVIDIA drivers are installed on the GPU nodes. When enabling the NVIDIA GPU Operator, add the following cluster override:

driver:
enabled: true

Pro: Enabling an Application Using the UI

For more information, see

.

What to do next

  1. Verify that the Envoy Gateway CRDs and controller are installed and ready.
  2. Verify that the KServe CRDs and controller resources are ready.
  3. Verify that the CloudNativePG operator is ready.
  4. Verify that the LeaderWorkerSet controller is ready.
  5. Verify that the OpenTelemetry Operator is ready.
  6. Verify that Prometheus monitoring is ready.
  7. Verify that the NVIDIA GPU Operator is ready.

Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform

on page 46

Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform Install or upgrade Nutanix Enterprise AI on Nutanix Kubernetes Platform (NKP).

Before you begin

Ensure to complete the following:

Installing Prerequisite Components

Install components on the Kubernetes cluster. For more information, see

on a Nutanix Kubernetes Platform Cluster

on page 43.

Ensure that the required RWO/RWX storage classes exist.

Ensure that the registry secret is available for NAI images.

Upgrades from version 2.7.0 to 2.8.0 must be executed during planned downtime. User logins will be unavailable during the upgrade, and full functionality will resume automatically once the upgrade is complete.

About this task

To install or upgrade Nutanix Enterprise AI on NKP, follow these steps:

Procedure

1. Choose a profile:

Profile-based deployment allows you to select a predefined configuration based on your environment and availability requirements. The default profile uses single replica for components and is intended for baseline deployment. The other profiles are for higher capacity usage.

Table 34: Profile and Capacity

ProfileCapacity
Default300 concurrent requests and 100 API Keys
c1k_k2001000 concurrent requests and 200 API Keys
c5k_k1k5000 concurrent requests and 1000 API Keys

2. Pull and untar both the 2.8.0 charts

helm pull ntnx-charts/nai-operators --version 2.8.0 --untar=true
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

The extracted charts contain profile files such as:

./nai-operators/profiles/c1k_k200.yaml
./nai-operators/profiles/c5k_k1k.yaml
./nai-core/profiles/c1k_k200.yaml
./nai-core/profiles/c5k_k1k.yaml

3. Deploy for c1k_k200 profile

a. Deploy NAI Operators for c1k_k200 profile

export NAI_DEFAULT_RWO_STORAGECLASS=<RWO storageclass>
helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
-f ./nai-operators/profiles/c1k_k200.yaml

b. Deploy NAI Core for c1k_k200 profile

export NAI_API_RWX_STORAGECLASS=<NFS Storageclass i.e nai-nfs-storage>
export NAI_DEFAULT_RWO_STORAGECLASS=<RWO default storageclass>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/profiles/c1k_k200.yaml
  1. Set up the NAI Helm repository. a. Add and update the Nutanix Helm repository, which contains the

and

helm chart:
nai-core
nai-operators
helm repo add ntnx-charts https://nutanix.github.io/helm-releases && helm repo
update ntnx-charts

b. Search for the version of the

and

helm chart available for installation in the
nai-core
nai-operators

Nutanix helm repository:

helm search repo ntnx-charts/nai-operators --versions
helm search repo ntnx-charts/nai-core --versions

5. Create Docker Registry Secrets:

Create the

namespace and the

secret in both

and

nai-system
docker-registry
nai-system
envoy-

namespaces.

gateway-system

The

is already present on the cluster.

envoy-gateway-system namespace
export REGISTRY_SECRET_NAME=nai-regcred
export DOCKER_SERVER=https://index.docker.io/v1/
export DOCKER_USERNAME=<docker-username>
export DOCKER_PASSWORD=<docker-password>
export DOCKER_EMAIL=<docker-email>
kubectl create namespace nai-system --dry-run=client -o yaml | kubectl apply -f -
kubectl -n nai-system create secret docker-registry ${REGISTRY_SECRET_NAME} \
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n envoy-gateway-system create secret docker-registry ${REGISTRY_SECRET_NAME}
\
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -

6. Deploy NAI Operators:

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0  -n
nai-system --create-namespace --take-ownership --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}"

[!NOTE] Note: Ensure the REGISTRY_SECRET_NAME environment variable is set before running this command.

7. Deploy NAI Core on NKP:

# Set the environment variable
export NAI_API_RWX_STORAGECLASS=<NFS Storageclass i.e nai-nfs-storage>
export NAI_DEFAULT_RWO_STORAGECLASS=<default storageclass i.e nutanix-volume>
export NKP_WORKSPACE_NAMESPACE=kommander-default-workspace # update this env as per
your workspace
export REGISTRY_SECRET_NAME=<secret created>
helm upgrade --install nai-core ntnx-charts/nai-core --version 2.8.0 -n nai-system --
create-namespace --force-conflicts --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
--set
"naiMonitoring.nodeExporter.serviceMonitor.namespaceSelector.matchNames[0]=
${NKP_WORKSPACE_NAMESPACE}" \
--set
"naiMonitoring.dcgmExporter.serviceMonitor.namespaceSelector.matchNames[0]=
${NKP_WORKSPACE_NAMESPACE}"

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

Enable the Chat and Talk to My Data application. These applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

Enable HTTPS with a self-signed certificate. By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

TLS Encryption on Nutanix

using your own certificate or a cert-manager ClusterIssuer, see

Enterprise AI

.

Scale out the ingress gateway and NAI API. To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Configure PostgreSQL database connections. To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default value is 1000. Increase this value if you expect a higher number of concurrent clients.

  1. Configure the TLS certificate.

For more information, see

TLS Encryption on Nutanix Enterprise AI

on page 119.

What to do next

Verify that the

and

Helm releases have

.

nai-operators
nai-core
STATUS=deployed

Verify that all expected pods are Ready.

Verify that persistent volumes are Bound and use the intended storage classes.

Verify that the selected profile produced the expected replica counts for critical components.

Verify that NAI services are reachable.

Pending

Failed

If you had endpoints in

status before the upgrade displaying the status as

with the message

Unable to pull runtime image with provided credentials, hibernate and resume the endpoints.

Access NAI Dashboard IP:

kubectl get svc -n envoy-gateway-system -l "gateway.envoyproxy.io/owning-gateway-
name=nai-ingress-gateway,gateway.envoyproxy.io/owning-gateway-namespace=nai-system" -
o jsonpath='{.items[0].status.loadBalancer.ingress[0].ip}'

For information on viewing the

Dashboard

, see

Log in to Nutanix Enterprise AI

on page 167.

Deploying Nutanix Enterprise AI on a Nutanix Kubernetes Platform Cluster in Air- Gapped Environments

Install Nutanix Enterprise AI (NAI) on a Nutanix Kubernetes Platform (NKP) cluster in an air-gapped environment.

Before you begin

NAI supports air-gapped installation only on NKP.

Ensure that you meet requirements listed in

Prerequisites for Deploying Nutanix Enterprise AI 2.8.0 on a

Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments

on page 52.

About this task

To install NAI on an NKP cluster in an air-gapped environment, follow these high-level steps:

Procedure

Download the NAI 2.8.0 release bundles from the Nutanix portal.

Downloading Nutanix Enterprise AI Air-Gap Release Bundles

For more information, see

on page 52.

Push Container Images to a private registry.

From this step onwards, execute the all commands from within your air-gapped environment typically a jumpbox or bastion host that has access to both your container registry and your Kubernetes cluster. Ensure that you have Docker and kubectl available on this machine before proceeding.

For more information, see

Pushing Container Images to Private Registry in Air-Gapped Environments

on

page 53

Extract the NAI Helm charts bundle to your working directory:

tar -xvf nai-helm-charts-2.8.0.tar

The extraction produces the following Helm chart archives:

gateway-crds-helm-v1.8.1.tgz
gateway-helm-v1.8.1.tgz
kserve-crd-v0.19.0.tgz
kserve-llmisvc-crd-v0.19.0.tgz
kserve-llmisvc-resources-v0.19.0.tgz
kserve-resources-v0.19.0.tgz
lws-0.8.0.tgz
nai-core-2.8.0.tgz
nai-operators-2.8.0.tgz
opentelemetry-operator-0.114.1.tgz

(Optional) Publish Helm Charts to Private Registry

To install Helm charts directly from your OCI-compatible registry instead of local files, push the charts using the following commands:

# Authenticate to your registry
helm registry login -u <username> -p <password> https://<registry>
# Push each chart to the registry
helm push gateway-crds-helm-v1.8.1.tgz oci://<registry>
helm push gateway-helm-v1.8.1.tgz oci://<registry>
helm push kserve-crd-v0.19.0.tgz oci://<registry>
helm push kserve-resources-v0.19.0.tgz oci://<registry>
helm push kserve-llmisvc-crd-v0.19.0.tgz oci://<registry>
helm push kserve-llmisvc-resources-v0.19.0.tgz oci://<registry>
helm push opentelemetry-operator-0.114.1.tgz oci://<registry>
helm push lws-0.8.0.tgz oci://<registry>
helm push nai-core-2.8.0.tgz oci://<registry>
helm push nai-operators-2.8.0.tgz oci://<registry>

The remaining steps assume installation from local Helm chart files.

Configure Registry Credentials

For more information, see

Configuring Docker Registry Credentials for Dependencies

on page 58.

Install Prometheus Monitoring from the Nutanix Kubernetes Platform (NKP) platform applications catalog. Prometheus Monitoring is not included in the Nutanix Enterprise AI air-gap image tar or the helm-charts tar. On an NKP cluster in an air-gapped environment, install Prometheus Monitoring from the NKP platform applications catalog before you install Nutanix Enterprise AI components. Enable the Prometheus Monitoring platform application on your NKP workload cluster.

To optimize resource utilization on the workload cluster, configure Prometheus Monitoring with the following minimum installation settings when you enable the application:

alertmanager:
enabled: false
grafana:
enabled: false
prometheus:
enabled: false
kubeStateMetrics:
enabled: false
kubernetesServiceMonitors:
enabled: false
prometheus-node-exporter.kubeRBACProxy:
kubeRBACProxy:
enabled: true

Pro: Enabling an Application Using the UI

For more information, see

.

Install NVIDIA GPU Operator from NKP platform applications.

If the GPU nodes do not have precompiled NVIDIA drivers installed, enable driver installation in the NVIDIA GPU Operator configuration. This setting ensures that the NVIDIA drivers are installed on the GPU nodes. When enabling the NVIDIA GPU Operator, add the following cluster override:

driver:
enabled: true

Pro: Enabling an Application Using the UI

For more information, see

.

Install CloudNativePG Operator:

helm install cnpg cloudnative-pg \
--repo https://cloudnative-pg.github.io/charts \
--version 0.28.0 -n cnpg-system --create-namespace --wait

Install LeaderWorkerSet:

helm install lws oci://registry.k8s.io/lws/charts/lws \
--version 0.8.0 -n lws-system --create-namespace --wait
  1. Install Envoy Gateway.

For more information, see

Installing Envoy Gateway in an Air-gapped Environment

on page 59.

  1. Install KServe.

Installing KServe

For more information, see

on page 61.

  1. Deploy the OpenTelemetry Operator.

For more information, see

Deploying the OpenTelemetry Operator

on page 61.

  1. Install NAI components.

Deploying Nutanix Enterprise AI Components in Air-Gapped Environments

For more information, see

on

page 62.

What to do next

1. Verify NAI operators installation. Check the status of NAI components:

# Check all pods in nai-system namespace
kubectl get pods -n nai-system
# Check NAI Operators and Core Helm release
helm list -n nai-system
# Check persistent volume claims
kubectl get pvc -n nai-system

2. Verify if NAI services are accessible:

# List all services in nai-system
kubectl get svc -n nai-system
# Check NAI API service
kubectl get svc -n nai-system nai-api
# Check NAI UI service
kubectl get svc -n nai-system nai-inference-ui
  1. Troubleshoot common issues that can occur when you install NAI 2.8 in air-gapped environments.

For more information, see

Troubleshooting Deployment of Nutanix Enterprise AI 2.8.0 in Air-Gapped

Environments

on page 66.

Prerequisites for Deploying Nutanix Enterprise AI 2.8.0 on a Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments Before proceeding with the installation, ensure the following requirements are met:

Nutanix Kubernetes Platform (NKP) cluster running Kubernetes 1.35.

kubectl CLI v1.33+ configured with cluster access

Helm CLI v4.0.5

Docker or compatible container runtime (for loading images)

Access to a private container registry

Registry credentials with appropriate permissions

Sufficient disk space for loading images (~100GB)

Downloading Nutanix Enterprise AI Air-Gap Release Bundles Download the required Nutanix Enterprise AI 2.8.0 release bundles from the Nutanix Portal.

About this task

To download the required NAI 2.8.0 release bundles from the Nutanix Portal, follow these steps:

Procedure

1. Navigate to the Nutanix Enterprise AI page on the Nutanix Support Portal

2. Select NAI Version 2.8.0 from the available releases

3. Download the following two bundles:

NAI Air-Gap Bundle (

)

nai-v2.8.0.tar

NAI Helm Charts Bundle (

)

nai-helm-charts-2.8.0.tar

Table 35: Air-Gap Download Bundles

BundleDescriptionSize
NAI Air-Gap Bundle (nai- v2.8.0.tar)~70-80GB

Contains all NAI container images and dependencies

Required for air-gapped deployments

NAI Helm Charts Bundle (nai- helm-charts-2.8.0.tar)

~5 MB

Contains NAI Helm charts and all dependency charts

Includes Envoy Gateway, KServe, LeaderWorkerSet and OpenTelemetry charts

4. Transfer both bundles to your air-gapped environment using approved methods such as USB drive, secure file

transfer, and so on.

Pushing Container Images to Private Registry in Air-Gapped Environments Before installing NAI, you must load all container images and push them to your private registry.

Before you begin

Execute all the commands from within your air-gapped environment typically a jumpbox or bastion host that has access to both your container registry and your Kubernetes cluster. Ensure you have Docker and kubectl available on this machine before proceeding.

About this task

To push container images to a private registry, follow these steps:

Procedure

1. Login to your private container registry:

docker login <registry-url>

For example,

docker login registry.example.com

The system displays a prompt to enter your registry credentials.

  1. Enter your registry credentials when prompted.
  2. Create the Image Push Script.

The Image Push Script pushes container images to your private registry.

a. Create a project/repository with the name

in your container registry, where all NAI images are

nutanix

stored. For example, in Harbor this would be a project named

, resulting in image paths like

nutanix

.

registry.example.com/nutanix/<image-name>:<tag>

b. Create a script file named

with the following content:

push-images-to-registry.sh
#!/bin/bash
#
# NAI Images - Load, Retag, and Push to Private Registry
#
# This script loads NAI container images from a tar bundle, retags them for your
# private registry, and pushes them to the registry.
#
# Prerequisites:
#   - Docker installed and running
#   - Docker logged into the target registry (docker login)
#   - NAI images tar bundle file
#
# Usage:
#   ./push-images-to-registry.sh <registry-url> <project> <tar-file>
#
# Example:
#   ./push-images-to-registry.sh registry.example.com nutanix nai-images-2.8.0.tar
#
set -uo pipefail
# ============================================================================
# Helper Functions
# ============================================================================
print_header() {
echo ""
echo "========================================"
echo "$1"
echo "========================================"
}
print_success() {
echo "# $1"
}
print_error() {
echo "# ERROR: $1" >&2
}
print_info() {
echo "# $1"
}
# ============================================================================
# Validate Arguments
# ============================================================================
if [ $# -ne 3 ]; then
echo "Usage: $0 <registry-url> <project> <tar-file>"
echo ""
echo "Arguments:"
echo "  registry-url    Your private registry URL (e.g.,
registry.example.com)"
echo "  project         Project/repository name in the registry (e.g.,
nutanix)"
echo "  tar-file        Path to the NAI images tar bundle"
echo ""
echo "Example:"
echo "  $0 registry.example.com nutanix nai-images-2.8.0.tar"
echo ""
exit 1
fi
REGISTRY="$1"
PROJECT="$2"
TAR_FILE="$3"
# Validate tar file exists
if [ ! -f "$TAR_FILE" ]; then
print_error "Tar file not found: $TAR_FILE"
exit 1
fi
# ============================================================================
# Configuration
# ============================================================================
print_header "NAI Images - Load, Retag & Push"
echo "Registry:  $REGISTRY"
echo "Project:   $PROJECT"
echo "Tar File:  $TAR_FILE"
echo "Date:      $(date)"
# Arrays to track images
LOADED_IMAGES=()
FAILED_IMAGES=()
# ============================================================================
# Step 1: Load Images from Tar Bundle
# ============================================================================
print_header "Step 1: Loading Images from Tar Bundle"
print_info "Loading images from $TAR_FILE..."
LOAD_OUTPUT=$(docker load -i "$TAR_FILE" 2>&1)
# Extract loaded image names
while IFS= read -r line; do
if [[ "$line" =~ Loaded\ image:\ (.+)$ ]]; then
LOADED_IMAGES+=("${BASH_REMATCH[1]}")
fi
done <<< "$LOAD_OUTPUT"
if [ ${#LOADED_IMAGES[@]} -eq 0 ]; then
print_error "No images were loaded from the tar file"
exit 1
fi
print_success "Loaded ${#LOADED_IMAGES[@]} images"
# ============================================================================
# Step 2: Retag and Push Images
# ============================================================================
print_header "Step 2: Retagging and Pushing Images"
PUSHED_COUNT=0
TOTAL_IMAGES=${#LOADED_IMAGES[@]}
for source_image in "${LOADED_IMAGES[@]}"; do
echo ""
print_info "[$((PUSHED_COUNT + 1))/$TOTAL_IMAGES] Processing: $source_image"
# Retag image for target registry
# Format: nutanix/nai-api:v2.8.0 # registry.example.com/<project>/nai-
api:v2.8.0
if [[ "$source_image" =~ ^nutanix/(.+)$ ]]; then
image_path="${BASH_REMATCH[1]}"
target_image="${REGISTRY}/${PROJECT}/${image_path}"
print_info "Tagging as: $target_image"
if ! docker tag "$source_image" "$target_image"; then
print_error "Failed to tag image"
FAILED_IMAGES+=("$source_image")
continue
fi
print_info "Pushing to registry..."
if docker push "$target_image"; then
print_success "Pushed successfully"
((PUSHED_COUNT++))
else
print_error "Failed to push image"
FAILED_IMAGES+=("$target_image")
fi
else
print_info "Skipping (not in nutanix/* format)"
fi
done
# ============================================================================
# Summary
# ============================================================================
print_header "Summary"
echo "Total images loaded:    $TOTAL_IMAGES"
echo "Successfully pushed:    $PUSHED_COUNT"
echo "Failed:                 ${#FAILED_IMAGES[@]}"
if [ ${#FAILED_IMAGES[@]} -gt 0 ]; then
echo ""
print_error "The following images failed:"
for img in "${FAILED_IMAGES[@]}"; do
echo "  - $img"
done
echo ""
exit 1
fi
echo ""
print_success "All images successfully pushed to $REGISTRY/$PROJECT"
echo ""
exit 0

c. Make the script executable:

chmod +x push-images-to-registry.sh

d. Execute the script to load, retag, and push all NAI images.

./push-images-to-registry.sh <registry-url> <project> nai-v2.8.0.tar

Example:

./push-images-to-registry.sh registry.example.com nutanix nai-v2.8.0.tar

Expected Output:

========================================
NAI Images - Load, Retag & Push
========================================
Registry:  registry.example.com
Project:   nutanix
Tar File:  ./nai-v2.8.0.tar
Date:      Tue Aug 18 05:16:33 PM UTC 2026
========================================
Step 1: Loading Images from Tar Bundle
========================================
# Loading images from ./nai-v2.8.0.tar...
# Loaded 41 images
========================================
Step 2: Retagging and Pushing Images
========================================
# [1/41] Processing: nutanix/nai-iam-proxy-control-plane:v2.8.0
# Tagging as: registry.example.com/nutanix/nai-iam-proxy-control-plane:v2.8.0
# Pushing to registry...
The push refers to repository [registry.example.com/nutanix/nai-iam-proxy-control-
plane]
054a97ddb80b: Pushed
5228eaa6af5b: Pushed
256f393e029f: Mounted from nutanix/nai-inference-ui
v2.8.0: digest:
sha256:587189a6559b7af653769a21f4755c66af14fe5135eaf45728c84a2eb2f3c808 size: 951
# Pushed successfully
# [2/41] Processing: nutanix/nai-iam-ui:v2.8.0
# Tagging as: registry.example.com/nutanix/nai-iam-ui:v2.8.0
# Pushing to registry...
The push refers to repository [registry.example.com/nutanix/nai-iam-ui]
7673a750ed47: Pushed
187de06a3fb0: Pushed
[... continues for all images ...]
========================================
Summary
========================================
Total images loaded:    41
Successfully pushed:    41
Failed:                 0
# All images successfully pushed to registry.example.com/nutanix

The image push process typically takes 30-60 minutes depending on your network speed and registry performance.

All images are retagged with your registry URL while preserving the original image path and tag

Original format: nutanix/nai-api:v2.8.0

Retagged format: /nutanix/nai-api:v2.8.0

The script reports failures and continues processing the remaining images.

You can safely re-run the script if it fails partway through.

Configuring Docker Registry Credentials for Dependencies Configure registry credentials.

About this task

To configure registry credentials, follow these steps:

Procedure

  1. Set environment Variables.

Export the following environment variables with your private registry credentials:

export REGISTRY=<registry-url-without-https>
export REGISTRY_USERNAME='<registry-username>'
export REGISTRY_PASSWORD='<registry-password>'
export REGISTRY_EMAIL='<registry-email>'
export IMAGE_PULL_SECRET=nai-docker-regcred
export PROJECT=nutanix # set the registry project name
  1. Replace the placeholder values with your actual registry information.

The

must not include the

protocol prefix.

REGISTRY
https://
  1. Create Image Pull Secrets.

Create Kubernetes namespaces and docker-registry secrets for Envoy Gateway System :

kubectl create namespace envoy-gateway-system --dry-run=client -o yaml | kubectl
apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n envoy-gateway-system \
--dry-run=client -o yaml | kubectl apply -f -

4. Create Kubernetes namespaces and

secrets for KServe:

docker-registry
kubectl create namespace kserve --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n kserve \
--dry-run=client -o yaml | kubectl apply -f -

5. Create Kubernetes namespaces and docker-registry secrets for OpenTelemetry:

kubectl create namespace opentelemetry --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n opentelemetry \
--dry-run=client -o yaml | kubectl apply -f -

6. Create Kubernetes namespaces and docker-registry secrets for LeaderWorkerSet:

kubectl create namespace lws-system --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n lws-system \
--dry-run=client -o yaml | kubectl apply -f -

7. Create Kubernetes namespaces and docker-registry secrets for nai-system:

kubectl create namespace nai-system --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n nai-system \
--dry-run=client -o yaml | kubectl apply -f -

Installing Envoy Gateway in an Air-gapped Environment Install Envoy Gateway.

Before you begin

Ensure that you meet requirements listed in

Prerequisites for Deploying Nutanix Enterprise AI 2.8.0 on a

Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments

on page 52.

About this task

To install Envoy Gateway, follow these steps:

Procedure

1. Install the Envoy Gateway and Gateway API Custom Resource Definitions (CRDs):

helm template eg ./gateway-crds-helm-v1.8.1.tgz \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --server-side --force-conflicts -f -

This command also installs the necessary Gateway API CRDs required for Envoy Gateway.

2. Create the configuration template file

:

eg-config-for-gateway-mode.yaml.template
# This file configures Envoy Gateway for AI Gateway mode with rate limiting
config:
envoyGateway:
gateway:
controllerName: "gateway.envoyproxy.io/gatewayclass-controller"
logging:
level:
default: "info"
provider:
kubernetes:
rateLimitDeployment:
patch:
type: "StrategicMerge"
value:
spec:
template:
spec:
containers:
- imagePullPolicy: "IfNotPresent"
name: "envoy-ratelimit"
env:
- name: REDIS_TYPE
value: "sentinel"
- name: REDIS_PIPELINE_WINDOW
value: "150us"
type: "Kubernetes"
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
maxMessageSize: 11Mi
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- "Translation"
- "Cluster"
- "Route"
service:
fqdn:
hostname: "ai-gateway-controller.nai-system.svc.cluster.local"
port: 1063
rateLimit:
backend:
type: "Redis"
redis:
url: "mymaster,nai-valkey-sentinel.nai-system.svc.cluster.local:26379"
  1. Ensure the REGISTRY environment variable is configured.

4. Deploy Envoy Gateway:

helm upgrade --install eg ./gateway-helm-v1.8.1.tgz \
-n envoy-gateway-system --create-namespace --wait \
--set global.images.envoyGateway.image=${REGISTRY}/${PROJECT}/nai-gateway:v1.8.1 \
--set global.images.ratelimit.image=${REGISTRY}/${PROJECT}/nai-ratelimit:1e50889b \
--set "global.imagePullSecrets[0].name=${IMAGE_PULL_SECRET}" \
-f ./eg-config-for-gateway-mode.yaml

The configuration file now uses your private registry for the

images through the

ratelimit
${REGISTRY}

variable substitution.

Installing KServe Install KServe.

About this task

To install KServe, follow these steps:

Procedure

1. Install the KServe Custom Resource Definitions (CRDs:)

helm upgrade --install kserve-crd ./kserve-crd-v0.19.0.tgz -n kserve --create-
namespace --wait

2. Deploy the KServe controller with RawDeployment mode:

helm upgrade --install kserve ./kserve-resources-v0.19.0.tgz \
-n kserve --wait \
--set kserve.controller.deploymentMode=RawDeployment \
--set kserve.controller.gateway.disableIngressCreation=true \
--set kserve.controller.image=${REGISTRY}/${PROJECT}/nai-kserve-controller \
--set kserve.controller.rbacProxyImage=${REGISTRY}/${PROJECT}/nai-kube-rbac-
proxy:v0.18.0 \
--set "kserve.controller.imagePullSecrets[0].name=${IMAGE_PULL_SECRET}"

3. Deploy the KServe llmisvc crds:

helm upgrade --install kserve-llmisvc-crd ./kserve-llmisvc-crd-v0.19.0.tgz -n kserve
--create-namespace --wait

4. Deploy the KServe llmisvc controller:

helm upgrade --install kserve-llmisvc-resources ./kserve-llmisvc-resources-
v0.19.0.tgz \
-n kserve --create-namespace --wait --set kserve.createSharedResources=false --set
kserve.llmisvc.createGIECRDs=false \
--set kserve.llmisvc.controller.image=${REGISTRY}/${PROJECT}/nai-llmisvc-controller
\
--set "kserve.llmisvc.controller.imagePullSecrets[0]=${IMAGE_PULL_SECRET}"

Deploying the OpenTelemetry Operator Deploy the OpenTelemetry Operator for observability and telemetry collection.

About this task

To deploy the OpenTelemetry Operator, run the following command:

Procedure

Deploy the OpenTelemetry Operator for observability and telemetry collection:

helm upgrade --install opentelemetry-operator ./opentelemetry-operator-0.114.1.tgz \
-n opentelemetry --create-namespace --wait \
--set manager.image.repository=${REGISTRY}/${PROJECT}/nai-opentelemetry-operator \
--set manager.collectorImage.repository=${REGISTRY}/${PROJECT}/nai-opentelemetry-
collector-contrib \
--set "imagePullSecrets[0].name=${IMAGE_PULL_SECRET}"

Deploying Nutanix Enterprise AI Components in Air-Gapped Environments Deploy Nutanix Enterprise AI components in air-gapped environments.

Before you begin

Upgrades from version 2.7.0 to 2.8.0 must be executed during planned downtime. User logins will be unavailable during the upgrade, and full functionality will resume automatically once the upgrade is complete.

About this task

To deploy Nutanix Enterprise AI components in air-gapped environments, follow these steps:

Procedure

using the

1. Create a values override file for NAI Operators named

darksite-nai-operators.yaml.template

following template:

This file configures all operator images to use your private registry.

global:
imagePullSecrets:
- name: ${IMAGE_PULL_SECRET}
storage:
storageClassName: ${NAI_DEFAULT_RWO_STORAGECLASS}
naiValkey:
image:
name: ${REGISTRY}/${PROJECT}/nai-valkey
naiJobs:
naiJobsImage:
image: ${REGISTRY}/${PROJECT}/nai-jobs
nai-clickhouse-operator:
operator:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-operator
metrics:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-metrics-exporter
ai-gateway-helm:
extProc:
image:
repository: ${REGISTRY}/${PROJECT}/nai-ai-gateway-extproc
controller:
image:
repository: ${REGISTRY}/${PROJECT}/nai-ai-gateway-controller
naiDatabase:
image: ${REGISTRY}/${PROJECT}/nai-postgresql:17.10-standard-trixie

2. Generate the actual values file:

Use

to replace environment variables and create the final values file:

envsubst
envsubst < darksite-nai-operators.yaml.template > darksite-nai-operators.yaml

The

command substitutes

,

and

with the actual

envsubst
${REGISTRY}
${PROJECT}
${IMAGE_PULL_SECRET}

values from your environment variables.

3. Install NAI Operators:

helm upgrade --install nai-operators ./nai-operators-2.8.0.tgz \
-n nai-system --create-namespace --wait --timeout 15m -f ./darksite-nai-
operators.yaml

4. Prepare NAI Core Values Override file:

Create a values override file for NAI Core using the provided template. This configures all NAI core component images.

a. Create the template file named

with the following content:

darksite-nai-core.yaml.template
global:
imagePullSecrets:
- name: ${IMAGE_PULL_SECRET}
storage:
storageClassName: ${NAI_DEFAULT_RWO_STORAGECLASS}
storageClassNameRWX: ${NAI_API_RWX_STORAGECLASS}
gateway:
envoyDeployment:
container:
image: ${REGISTRY}/${PROJECT}/nai-envoy:distroless-v1.38.0
naiIepOperator:
iepOperatorImage:
image: ${REGISTRY}/${PROJECT}/nai-iep-operator
modelProcessorImage:
image: ${REGISTRY}/${PROJECT}/nai-python-processor
dataSourceProcessorImage:
image: ${REGISTRY}/${PROJECT}/nai-python-processor
batchInferenceProcessor:
containers:
processor:
image: ${REGISTRY}/${PROJECT}/nai-go-processor
statusProvider:
image: ${REGISTRY}/${PROJECT}/nai-go-processor
finetuneProcessor:
containers:
processor:
image: ${REGISTRY}/${PROJECT}/nai-finetuning
statusProvider:
image: ${REGISTRY}/${PROJECT}/nai-go-processor
naiInferenceUi:
naiUiImage:
image: ${REGISTRY}/${PROJECT}/nai-inference-ui
naiJobs:
naiJobsImage:
image: ${REGISTRY}/${PROJECT}/nai-jobs
naiApi:
naiApiImage:
image: ${REGISTRY}/${PROJECT}/nai-api
supportedTGIImage: ${REGISTRY}/${PROJECT}/nai-tgi
supportedKserveRuntimeImage: ${REGISTRY}/${PROJECT}/nai-kserve-huggingfaceserver
eppImage: ${REGISTRY}/${PROJECT}/nai-epp-inference-scheduler
supportedVLLMImage: ${REGISTRY}/${PROJECT}/nai-vllm
supportedKserveCustomModelServerRuntimeImage: ${REGISTRY}/${PROJECT}/nai-kserve-
custom-model-server
naiDatabase:
clientImage: ${REGISTRY}/${PROJECT}/nai-postgresql:17.10-standard-trixie
naiIam:
iamProxy:
image: ${REGISTRY}/${PROJECT}/nai-iam-proxy
iamProxyControlPlane:
image: ${REGISTRY}/${PROJECT}/nai-iam-proxy-control-plane
iamUi:
image: ${REGISTRY}/${PROJECT}/nai-iam-ui
iamUserAuthn:
image: ${REGISTRY}/${PROJECT}/nai-iam-user-authn
iamThemis:
image: ${REGISTRY}/${PROJECT}/nai-iam-themis
iamThemisBootstrap:
image: ${REGISTRY}/${PROJECT}/nai-iam-bootstrap
naiAgent:
agentImage:
image: ${REGISTRY}/${PROJECT}/nai-agent-app
naiLabs:
labsImage:
image: ${REGISTRY}/${PROJECT}/nai-rag-app
nai-clickhouse-keeper:
clickhouseKeeper:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-keeper
oauth2-proxy:
image:
repository: ${REGISTRY}/${PROJECT}/nai-oauth2-proxy
nai-clickhouse-server:
clickhouse:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-server
initContainers:
addUdf:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-udf
waitForKeeper:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-jobs
nai-clickhouse-schemas:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-schemas
naiMonitoring:
opentelemetry:
collectorImage: ${REGISTRY}/${PROJECT}/nai-opentelemetry-collector-
contrib:0.152.0
targetAllocator:
image:
repository: ${REGISTRY}/${PROJECT}/nai-target-allocator
nodeExporter:
serviceMonitor:
namespaceSelector:
matchNames:
- prometheus
- kommander
- kommander-default-workspace
- ${NKP_WORKSPACE_NAMESPACE}
dcgmExporter:
serviceMonitor:
namespaceSelector:
matchNames:
- prometheus
- kommander
- kommander-default-workspace
- ${NKP_WORKSPACE_NAMESPACE}

b. Ensure that

,

,

,

REGISTRY
IMAGE_PULL_SECRET
NAI_API_RWX_STORAGECLASS

, and

environment variables are set before

NAI_DEFAULT_RWO_STORAGECLASS
NKP_WORKSPACE_NAMESPACE

running the next command.

c. Generate the actual values file:

Use

to replace environment variables and create the final values file:

envsubst
envsubst < darksite-nai-core.yaml.template > darksite-nai-core.yaml

5. Install NAI Core:

helm upgrade --install nai-core ./nai-core-2.8.0.tgz -n nai-system --create-namespace
--wait --timeout 15m \
-f ./darksite-nai-core.yaml

All configuration including registry URLs, storage classes, and monitoring namespaces are in the darksite-nai- core.yaml values file, making the install command much simpler.

The NAI Core installation may take 10 to15 minutes depending on your cluster resources and network speed.

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

The Chat and Talk to My Data applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

using your own certificate or a cert-manager ClusterIssuer, see

TLS Encryption on Nutanix

Enterprise AI

.

Scale out: To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Database Connections: To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default is 1000. Increase if you expect a higher number of concurrent clients.

  1. Configure the TLS certificate.

TLS Encryption on Nutanix Enterprise AI

For more information, see

on page 119.

Troubleshooting Deployment of Nutanix Enterprise AI 2.8.0 in Air-Gapped Environments The following troubleshooting tips can help you resolve a few common issues that can occur when you are installing

Image Pull Errors

Verify that:

Registry credentials are correct

Image pull secrets exist in the correct namespaces

Registry URL is accessible from the cluster

Image paths match your registry structure

All required images are present in your private registry

CRD Installation Failures

Ensure you have sufficient permissions to create cluster-scoped resources.

Helm Installation Timeouts

Increase the timeout value using –timeout flag (e.g., –timeout 20m).

Storage Class Issues

Verify the storage class exists: kubectl get storageclass

Ensure the storage class supports the required access mode (RWX for NAI API, RWO for others)

Check PVC status: kubectl get pvc -n nai-system

Pod Startup Failures

Check pod logs: kubectl logs -n nai-system

Describe pod for events: kubectl describe pod -n nai-system

Verify resource limits if running on resource-constrained clusters

Dependency Issues

Ensure all dependencies (Envoy Gateway, KServe, OpenTelemetry) are installed and running before installing NAI Core.

Deploy Nutanix Enterprise AI on Amazon Elastic Kubernetes Service

Install or upgrade Nutanix Enterprise AI (NAI) on Amazon Elastic Kubernetes Service (EKS).

To install or upgrade NAI on EKS, follow these high-level steps:

1. Set up an EKS cluster. For more information, see

Setting up an Amazon Elastic Kubernetes Service Cluster

on page 67.

2. Perform preflight checks. For more information, see

Performing Preflight Checks Before Installing NAI on

Amazon EKS

on page 68.

Installing Prerequisite Components on an EKS

3. Install prerequisite components. For more information, see

Cluster

on page 69.

4. Deploy NAI on EKS. For more information, see

Deploying Nutanix Enterprise AI on Amazon Elastic

Kubernetes Service

on page 72.

Setting up an Amazon Elastic Kubernetes Service Cluster Set up your Amazon Elastic Kubernetes Service (EKS) cluster for Nutanix Enterprise AI. Do not follow this procedure for Nutanix Kubernetes Platform (NKP) managed or attached EKS clusters.

Before you begin

Ensure that your cluster meets all the requirements mentioned in

Nutanix Enterprise AI - Private Inference

and Agent Gateway Requirements

on page 10.

About this task

To set up your EKS cluster before you deploy Nutanix Enterprise AI, follow these steps:

Procedure

Configure the AWS CLI.

AWS documentation

For more information, see

.

Configure an Amazon virtual private cloud (VPC) with public and private subnets for EKS using CloudFormation.

For more information, see

AWS documentation

.

Create an Amazon EKS cluster with Kubernetes version 1.33 or later.

For more information, see

AWS documentation

.

Create a kubeconfig file to connect kubectl to the EKS cluster.

AWS documentation

For more information, see

.

Create an EKS role for the EKS cluster.

For more information, see

AWS documentation

.

Create an Elastic Compute Cloud (EC2) role for the EKS cluster.

AWS documentation

For more information, see

.

Create a default node group. Make sure the nodes have accelerator AVX2 or newer.

The recommended instance type is from instance family C5.

Create a GPU node group on the EKS Cluster.

For more information, see

AWS documentation

.

Configure the Amazon EBS CSI driver and the RWO storageclass.

AWS documentation

For more information, see

.

  1. Configure the Amazon EFS CSI driver and the RWX storageclass.

For more information, see

AWS documentation

.

  1. Configure the Amazon VPC CNI plugin to enable network policy enforcement.

AWS documentation

For more information, see

.

Performing Preflight Checks Before Installing NAI on Amazon EKS Perform preflight checks before installing NAI on Amazon Elastic Kubernetes Service (EKS). Preflight checks ensure the Kubernetes environment meets all the technical requirements for a successful NAI installation.

About this task

To perform preflight checks for installing NAI on EKS, follow these steps:

Procedure

1. Make sure the Kubernetes version on your worker nodes is compatible:

kubectl get nodes --selector='!node-role.kubernetes.io/control-plane' \
-o custom-
columns=NAME:.metadata.name,KUBELET_VERSION:.status.nodeInfo.kubeletVersion \
--no-headers

The expected output is that the Kubernetes version must be 1.33 or later.

2. Verify if the number of EFS CSI pods match the number of worker nodes:

[ $(kubectl get nodes --no-headers | wc -l) -eq $(kubectl get pods -n kube-system --
no-headers | grep efs-csi-node | wc -l) ] && echo "# efs-csi-node pods = node count"
|| echo "# Mismatch: efs-csi-node pods != node count"

The expected output is

.

efs-csi-node pods= node count

3. Confirm that the

class binds volumes immediately:

nai-nfs-storage
kubectl get storageclass nai-nfs-storage -o jsonpath='{.volumeBindingMode}{"\n"}'

The expected output is

.

immediate

4. Check if the

storage class has

access:

nai-nfs-storage
ReadWriteMany
kubectl apply -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-rwx-pvc
namespace: default
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Gi
storageClassName: nai-nfs-storage
EOF
kubectl get pvc test-rwx-pvc -n default

The expected output is that the storage class

has

set to

.

nai-nfs-storage
Access Modes
RWX
NAME           STATUS   VOLUME                                     CAPACITY   ACCESS
MODES   STORAGECLASS      VOLUMEATTRIBUTESCLASS   AGE
test-rwx-pvc            Bound    pvc-0df5145a-590c-4f28-9f9d-cefefc55266c   1Gi
RWX            nai-nfs-storage   <unset>                 5s
kubectl delete pvc test-rwx-pvc

Installing Prerequisite Components on an EKS Cluster Install components on an EKS cluster.

About this task

To install components on an EKS cluster, follow these steps:

Procedure

1. Install cert-manager:

helm upgrade --install cert-manager cert-manager --repo https://charts.jetstack.io --
version v1.19.3 --set installCRDs=true -n cert-manager --create-namespace --wait

2. Install or upgrade Envoy Gateway:

a. Install Envoy Gateway and Gateway API CRDs:

helm template eg oci://docker.io/envoyproxy/gateway-crds-helm --version v1.8.1 \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --server-side --force-conflicts -f -

b. Create

with the following configuration:

envoy-gateway-config.yaml
config:
envoyGateway:
gateway:
controllerName: "gateway.envoyproxy.io/gatewayclass-controller"
logging:
level:
default: "info"
provider:
kubernetes:
rateLimitDeployment:
container:
image: "docker.io/envoyproxy/ratelimit:1e50889b"
patch:
type: "StrategicMerge"
value:
spec:
template:
spec:
containers:
- imagePullPolicy: "IfNotPresent"
name: "envoy-ratelimit"
image: "docker.io/envoyproxy/ratelimit:1e50889b"
env:
- name: REDIS_TYPE
value: "sentinel"
- name: REDIS_PIPELINE_WINDOW
value: "150us"
type: "Kubernetes"
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
maxMessageSize: 11Mi
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- "Translation"
- "Cluster"
- "Route"
service:
fqdn:
hostname: "ai-gateway-controller.nai-system.svc.cluster.local"
port: 1063
rateLimit:
backend:
type: "Redis"
redis:
url: "mymaster,nai-valkey-sentinel.nai-system.svc.cluster.local:26379"

[!NOTE] Note: The rate-limit backend uses Valkey Sentinel with master name

mymaster

and the

nai-valkey-

service in

.

sentinel
nai-system

c. Install or Upgrade Envoy Gateway:

helm upgrade --install eg oci://docker.io/envoyproxy/gateway-helm --version v1.8.1
\
-n envoy-gateway-system --create-namespace --skip-crds \
-f "./envoy-gateway-config.yaml"

3. Install or Upgrade KServe:

The required version is KSERVE_VERSION=v0.19.0.

a. Install or upgrade the KServe CRDs:

helm upgrade --install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

b. Install or upgrade the KServe resources:

helm upgrade --install kserve oci://ghcr.io/kserve/charts/kserve-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.controller.deploymentMode=RawDeployment \
--set kserve.controller.gateway.disableIngressCreation=true

c. Install or upgrade the KServe LLMInferenceService CRD:

helm upgrade --install kserve-llmisvc-crd oci://ghcr.io/kserve/charts/kserve-
llmisvc-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

d. Install or upgrade the KServe LLMInferenceService resources:

helm upgrade --install kserve-llmisvc-resources oci://ghcr.io/kserve/charts/
kserve-llmisvc-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.createSharedResources=false \
--set kserve.llmisvc.createGIECRDs=false

4. Install CloudNativePG Operator:

helm install cnpg cloudnative-pg \
--repo https://cloudnative-pg.github.io/charts \
--version 0.28.0 -n cnpg-system --create-namespace --wait

5. Install LeaderWorkerSet:

helm install lws oci://registry.k8s.io/lws/charts/lws \
--version 0.8.0 -n lws-system --create-namespace --wait

6. Install or upgrade OpenTelemetry Operator

helm upgrade --install opentelemetry-operator opentelemetry-operator \
--repo https://open-telemetry.github.io/opentelemetry-helm-charts \
--version=0.114.1 -n opentelemetry --create-namespace --wait

7. Install or upgrade Prometheus Monitoring:

helm upgrade --install prometheus kube-prometheus-stack \
--repo https://prometheus-community.github.io/helm-charts \
--version=82.13.6 -n prometheus --create-namespace --wait \
--set grafana.enabled=false \
--set prometheus.enabled=false \
--set kubeStateMetrics.enabled=false \
--set alertManager.enabled=false \
--set kubernetesServiceMonitors.enabled=false \
--set prometheus-node-exporter.kubeRBACProxy.enabled=true

8. Install or upgrade NVIDIA GPU Operator :

helm upgrade --install --wait gpu-operator gpu-operator \
--repo https://helm.ngc.nvidia.com/nvidia \
-n gpu-operator --create-namespace --version=v26.3.0

What to do next

  1. Verify that the Envoy Gateway CRDs and controller are installed and ready.
  2. Verify that the KServe CRDs and controller resources are ready.
  3. Verify that the CloudNativePG operator is ready.
  4. Verify that the LeaderWorkerSet controller is ready.
  5. Verify that the OpenTelemetry Operator is ready.
  6. Verify that Prometheus monitoring is ready.
  7. Verify that the NVIDIA GPU Operator is ready.
Deploying Nutanix Enterprise AI on Amazon Elastic Kubernetes Service Install or upgrade Nutanix Enterprise AI on Amazon Elastic Kubernetes Service (EKS).

Before you begin

Ensure to complete the following:

Installing Prerequisite Components

Install components on the Kubernetes cluster. For more information, see

on an EKS Cluster

on page 69.

Upgrades from version 2.7.0 to 2.8.0 must be executed during planned downtime. User logins will be unavailable during the upgrade, and full functionality will resume automatically once the upgrade is complete.

About this task

To install or upgrade Nutanix Enterprise AI on EKS, follow these steps:

Procedure

1. Choose a profile:

Profile-based deployment allows you to select a predefined configuration based on your environment and availability requirements. The default profile uses single replica for components and is intended for baseline deployment. The other profiles are for higher capacity usage.

Table 36: Profile and Capacity

ProfileCapacity
Default300 concurrent requests and 100 API Keys
c1k_k2001000 concurrent requests and 200 API Keys
c5k_k1k5000 concurrent requests and 1000 API Keys

2. Pull and untar both the 2.8.0 charts

helm pull ntnx-charts/nai-operators --version 2.8.0 --untar=true
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

The extracted charts contain profile files such as:

./nai-operators/profiles/c1k_k200.yaml
./nai-operators/profiles/c5k_k1k.yaml
./nai-core/profiles/c1k_k200.yaml
./nai-core/profiles/c5k_k1k.yaml

3. Deploy for c1k_k200 profile

a. Deploy NAI Operators for c1k_k200 profile

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
-f ./nai-operators/profiles/c1k_k200.yaml \
-f ./nai-operators/eks-values.yaml

b. Deploy NAI Core for c1k_k200 profile

export NAI_API_RWX_STORAGECLASS=<RWX/NFS storageclass>
export NAI_DEFAULT_RWO_STORAGECLASS=<RWO default storageclass>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/profiles/c1k_k200.yaml \
-f ./nai-core/eks-values.yaml
  1. Set up the NAI Helm repository. a. Add and update the Nutanix Helm repository, which contains the

and

helm chart:
nai-core
nai-operators
helm repo add ntnx-charts https://nutanix.github.io/helm-releases && helm repo
update ntnx-charts

b. Search for the version of the

and

helm chart available for installation in the
nai-core
nai-operators

Nutanix helm repository:

helm search repo ntnx-charts/nai-operators --versions
helm search repo ntnx-charts/nai-core --versions

5. Create Docker Registry Secrets:

Create the

namespace and the

secret in both

and

nai-system
docker-registry
nai-system
envoy-

namespaces.

gateway-system

The

is already present on the cluster.

envoy-gateway-system namespace
export REGISTRY_SECRET_NAME=nai-regcred
export DOCKER_SERVER=https://index.docker.io/v1/
export DOCKER_USERNAME=<docker-username>
export DOCKER_PASSWORD=<docker-password>
export DOCKER_EMAIL=<docker-email>
kubectl create namespace nai-system --dry-run=client -o yaml | kubectl apply -f -
kubectl -n nai-system create secret docker-registry ${REGISTRY_SECRET_NAME} \
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n envoy-gateway-system create secret docker-registry ${REGISTRY_SECRET_NAME}
\
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -

6. Deploy NAI Operators:

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0  -n
nai-system --create-namespace --take-ownership --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}"

[!NOTE] Note: Ensure the REGISTRY_SECRET_NAME environment variable is set before running this command.

helm chart and extract it to your local repository:

7. Pull the

nai-core
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

8. Install or upgrade NAI Core on EKS:

# Set the environment variable
NAI_API_RWX_STORAGECLASS=<NFS Storageclass i.e nai-nfs-storage>
NAI_DEFAULT_RWO_STORAGECLASS=<default storageclass>
REGISTRY_SECRET_NAME=<secret created>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 -n nai-system --
create-namespace --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/eks-values.yaml

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

Enable the Chat and Talk to My Data application. These applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

Enable HTTPS with a self-signed certificate. By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

TLS Encryption on Nutanix

using your own certificate or a cert-manager ClusterIssuer, see

Enterprise AI

.

Scale out the ingress gateway and NAI API. To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Configure PostgreSQL database connections. To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default value is 1000. Increase this value if you expect a higher number of concurrent clients.

  1. Configure the TLS certificate.

For more information, see

TLS Encryption on Nutanix Enterprise AI

on page 119.

What to do next

Ensure that all the pods are successfully deployed in the

namespace and are in the ready state by:

nai-system
kubectl get pods -n nai-system

The following is a sample output of the command. Match the pod names by ignoring the auto-generated suffix added by Kubernetes and its status.

[!NOTE] Note:

The number of

should be equal to the number of worker

nai-otel-collector-collector pods

nodes in the NAI Kubernetes cluster.

NAME                                                            READY   STATUS
RESTARTS   AGE
chi-nai-clickhouse-server-chcluster1-0-0-0                      1/1     Running     0
16h
chk-nai-clickhouse-keeper-chkeeper-0-0-0                        1/1     Running     0
16h
iam-database-bootstrap-b8etj-hk9rz                              0/1     Completed   0
16h
iam-proxy-68f9459885-zcwgs                                      1/1     Running     0
16h
iam-proxy-control-plane-6897669d64-rvglt                        1/1     Running     0
16h
iam-themis-749b7b56f8-pmclb                                     1/1     Running     0
16h
iam-themis-bootstrap-qgczx-p7bl6                                0/1     Completed   0
16h
iam-ui-6697d94478-fftl5                                         1/1     Running     0
16h
iam-user-authn-5b4dcfdfb7-jhllq                                 1/1     Running     0
16h
nai-api-784fb7b99-8tch9                                         1/1     Running     0
16h
nai-api-db-migrate-mgnba-cl7gw                                  0/1     Completed   0
16h
nai-clickhouse-schema-job-1771350985-bd4sq                      0/1     Completed   0
16h
nai-db-0                                                        1/1     Running     0
16h
nai-iep-model-controller-5cd8bcd5f-9h9cl                        1/1     Running     0
16h
nai-labs-86589cc95d-gk87q                                       1/1     Running     0
16h
nai-oauth2-proxy-5746ccc8b7-65jmf                               1/1     Running     0
16h
nai-oidc-client-registration-lvrrv-pbdsn                        0/1     Completed   0
16h
nai-otel-collector-collector-4pxtl                              1/1     Running     0
16h
nai-otel-collector-collector-8p4w8                              1/1     Running     0
16h
nai-otel-collector-collector-bwstg                              1/1     Running     0
16h
nai-otel-collector-collector-drddv                              1/1     Running     0
16h
nai-otel-collector-collector-k457n                              1/1     Running     0
16h
nai-otel-collector-collector-nmwjw                              1/1     Running     0
16h
nai-otel-collector-targetallocator-c9dcc6544-5wbrx              1/1     Running     0
16h
nai-pulse-job-29522885-6w7pn                                    0/1     Completed   0
10h
nai-ui-6d9cc89b87-b4npf                                         1/1     Running     0
16h
nutanix-ai-operators-nai-clickhouse-operator-7f9965dbdd-trwp5   2/2     Running     0
16h
redis-standalone-6df56bc96d-hp86q                               2/2     Running     0
16h

Access NAI Dashboard IP:

kubectl get svc -n envoy-gateway-system -l "gateway.envoyproxy.io/owning-gateway-
name=nai-ingress-gateway,gateway.envoyproxy.io/owning-gateway-namespace=nai-system" -
o jsonpath='{.items[0].status.loadBalancer.ingress[0].hostname}'

Log in to Nutanix Enterprise AI

For information on viewing the Dashboard, see

on page 167.

If you had endpoints in

Pending

status before the upgrade displaying the status as

Failed

with the message

Unable to pull runtime image with provided credentials, hibernate and resume the endpoints.

Deploying Nutanix Enterprise AI on Azure Kubernetes Service

Install or upgrade Nutanix Enterprise AI (NAI) on Azure Kubernetes Service (AKS).

To install or upgrade NAI on AKS, follow these step:

1. Set up an AKS cluster. For more information, see

Setting up an Azure Kubernetes Service Cluster

on

page 76.

2. Perform preflight checks. For more information, see

Performing Preflight Checks Before Installing Nutanix

Enterprise AI on Azure AKS

on page 77.

3. Install prerequisite components. For more information, see

Deploying Nutanix Enterprise AI on Azure

Kubernetes Service

on page 82.

4. Deploy NAI on AKS. For more information, see

Deploying Nutanix Enterprise AI on Azure Kubernetes

Service

on page 82.

Setting up an Azure Kubernetes Service Cluster Set up your Azure Kubernetes Service (AKS) cluster for Nutanix Enterprise AI (NAI). This procedure is not applicable for Nutanix Kubernetes Platform (NKP) managed or attached AKS clusters.

Before you begin

Ensure that your cluster meets all the requirements mentioned in

Nutanix Enterprise AI - Private Inference

and Agent Gateway Requirements

on page 10.

About this task

To set up your AKS cluster before you deploy NAI, follow these high-level steps:

Procedure

  1. Configure Azure CLI.

For more information, see

Azure documentation

.

  1. Create an Azure resource group.

Azure documentation

For more information, see

.

  1. Create an AKS cluster with Kubernetes version 1.33 or 1.34.

Azure documentation

For more information, see

.

  1. Configure the AKS cluster with a supported network policy engine enabled for enforcement.

For more information, see

Azure documentation

.

  1. Configure the AKS preview extension.

Azure documentation

For more information, see

.

  1. Create a default node group. Make sure the nodes have accelerator AVX2 or newer.
  2. Add a GPU node pool to the AKS cluster using Azure CLI.

Ensure that you skip installing the GPU driver when adding the GPU node pool to the cluster. For more information, see

Azure documentation

Azure documentation

. For more information, see

.

Performing Preflight Checks Before Installing Nutanix Enterprise AI on Azure AKS Perform preflight checks before installing Nutanix Enterprise AI on Azure AKS. Preflight checks ensure the Kubernetes environment meets all the technical requirements for a successful NAI installation.

About this task

To perform preflight checks for installing Nutanix Enterprise AI on Azure AKS, follow these steps:

Procedure

1. Choose a profile:

Profile-based deployment allows you to select a predefined configuration based on your environment and availability requirements. The default profile uses single replica for components and is intended for baseline deployment. The other profiles are for higher capacity usage.

Table 37: Profile and Capacity

ProfileCapacity
Default300 concurrent requests and 100 API Keys
c1k_k2001000 concurrent requests and 200 API Keys
c5k_k1k5000 concurrent requests and 1000 API Keys

2. Pull and untar both the 2.8.0 charts

helm pull ntnx-charts/nai-operators --version 2.8.0 --untar=true
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

The extracted charts contain profile files such as:

./nai-operators/profiles/c1k_k200.yaml
./nai-operators/profiles/c5k_k1k.yaml
./nai-core/profiles/c1k_k200.yaml
./nai-core/profiles/c5k_k1k.yaml

3. Deploy for small profile:

a. Deploy NAI Operators for small profile

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
-f ./nai-operators/profiles/small.yaml \
-f ./nai-operators/aks-values.yaml

b. Deploy NAI Core for small profile

export NAI_API_RWX_STORAGECLASS=<RWX/NFS storageclass>
export NAI_DEFAULT_RWO_STORAGECLASS=<RWO default storageclass>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/profiles/small.yaml \
-f ./nai-core/aks-values.yaml

4. Ensure your Kubernetes version is 1.33 or 1.34:

kubectl get nodes --selector='!node-role.kubernetes.io/control-plane' \
-o custom-
columns=NAME:.metadata.name,KUBELET_VERSION:.status.nodeInfo.kubeletVersion \
--no-headers

The expected output is that the Kubernetes version must be 1.33 or 1.34.

5. Verify if the number of Azure Disk CSI pods (

) match the number of worker nodes:

csi-azuredisk-node
[ $(kubectl get nodes --no-headers | wc -l) -eq $(kubectl get pods -n kube-system --
no-headers | grep csi-azuredisk-node | wc -l) ] && echo "# csi-azuredisk-node pods =
node count" || echo "# Mismatch: csi-azuredisk-node pods != node count"

The expected output is

.

csi-azuredisk-node pods= node count

6. Confirm that the

storage class binds volumes immediately:

azurefile-csi
kubectl get sc azurefile-csi -o jsonpath='{.metadata.name}: {.volumeBindingMode}
{"\n"}'

The expected output is

.

immediate

7. Check if the storage class has

access:

ReadWriteMany
kubectl apply -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-rwx-pvc
namespace: default
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Gi
storageClassName: azurefile-csi
EOF
kubectl get pvc test-rwx-pvc -n default

The expected output is that the storage class

has

set to

.

azurefile-csi
Access Modes
RWX
NAME           STATUS   VOLUME                                     CAPACITY   ACCESS
MODES   STORAGECLASS    VOLUMEATTRIBUTESCLASS   AGE
test-rwx-pvc   Bound    pvc-4811d039-5cf4-4925-9181-2b2a3689ea47   1Gi        RWX
azurefile-csi   <unset>                 4s
kubectl delete pvc test-rwx-pvc

8. Verify if a CNI is enabled on the cluster:

az aks show --resource-group <resource-group-name> --name <aks-cluster-name> --query
"networkProfile.networkPolicy"

The expected output is the network policy engine selected during cluster setup. If the output is

, network policy is not enabled for the cluster.

"networkPolicy": "none"

Installing Prerequisite Components on an AKS Cluster Install components on an AKS cluster.

About this task

To install components on an AKS cluster, follow these steps:

Procedure

1. Install cert-manager:

helm upgrade --install cert-manager cert-manager --repo https://charts.jetstack.io --
version v1.19.3 --set installCRDs=true -n cert-manager --create-namespace --wait

2. Install or upgrade Envoy Gateway:

a. Install Envoy Gateway and Gateway API CRDs:

helm template eg oci://docker.io/envoyproxy/gateway-crds-helm --version v1.8.1 \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --server-side --force-conflicts -f -

b. Create

with the following configuration:

envoy-gateway-config.yaml
config:
envoyGateway:
gateway:
controllerName: "gateway.envoyproxy.io/gatewayclass-controller"
logging:
level:
default: "info"
provider:
kubernetes:
rateLimitDeployment:
container:
image: "docker.io/envoyproxy/ratelimit:1e50889b"
patch:
type: "StrategicMerge"
value:
spec:
template:
spec:
containers:
- imagePullPolicy: "IfNotPresent"
name: "envoy-ratelimit"
image: "docker.io/envoyproxy/ratelimit:1e50889b"
env:
- name: REDIS_TYPE
value: "sentinel"
- name: REDIS_PIPELINE_WINDOW
value: "150us"
type: "Kubernetes"
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
maxMessageSize: 11Mi
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- "Translation"
- "Cluster"
- "Route"
service:
fqdn:
hostname: "ai-gateway-controller.nai-system.svc.cluster.local"
port: 1063
rateLimit:
backend:
type: "Redis"
redis:
url: "mymaster,nai-valkey-sentinel.nai-system.svc.cluster.local:26379"

[!NOTE] Note: The rate-limit backend uses Valkey Sentinel with master name

and the

mymaster
nai-valkey-

service in

.

sentinel
nai-system

c. Install or Upgrade Envoy Gateway:

helm upgrade --install eg oci://docker.io/envoyproxy/gateway-helm --version v1.8.1
\
-n envoy-gateway-system --create-namespace --skip-crds \
-f "./envoy-gateway-config.yaml"

3. Install or Upgrade KServe:

The required version is KSERVE_VERSION=v0.19.0.

a. Install or upgrade the KServe CRDs:

helm upgrade --install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

b. Install or upgrade the KServe resources:

helm upgrade --install kserve oci://ghcr.io/kserve/charts/kserve-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.controller.deploymentMode=RawDeployment \
--set kserve.controller.gateway.disableIngressCreation=true

c. Install or upgrade the KServe LLMInferenceService CRD:

helm upgrade --install kserve-llmisvc-crd oci://ghcr.io/kserve/charts/kserve-
llmisvc-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

d. Install or upgrade the KServe LLMInferenceService resources:

helm upgrade --install kserve-llmisvc-resources oci://ghcr.io/kserve/charts/
kserve-llmisvc-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.createSharedResources=false \
--set kserve.llmisvc.createGIECRDs=false

4. Install CloudNativePG Operator:

helm install cnpg cloudnative-pg \
--repo https://cloudnative-pg.github.io/charts \
--version 0.28.0 -n cnpg-system --create-namespace --wait

5. Install LeaderWorkerSet:

helm install lws oci://registry.k8s.io/lws/charts/lws \
--version 0.8.0 -n lws-system --create-namespace --wait

6. Install or upgrade OpenTelemetry Operator

helm upgrade --install opentelemetry-operator opentelemetry-operator \
--repo https://open-telemetry.github.io/opentelemetry-helm-charts \
--version=0.114.1 -n opentelemetry --create-namespace --wait

7. Install or upgrade Prometheus Monitoring:

helm upgrade --install prometheus kube-prometheus-stack \
--repo https://prometheus-community.github.io/helm-charts \
--version=82.13.6 -n prometheus --create-namespace --wait \
--set grafana.enabled=false \
--set prometheus.enabled=false \
--set kubeStateMetrics.enabled=false \
--set alertManager.enabled=false \
--set kubernetesServiceMonitors.enabled=false \
--set prometheus-node-exporter.kubeRBACProxy.enabled=true

8. Install or upgrade NVIDIA GPU Operator :

helm upgrade --install --wait gpu-operator gpu-operator \
--repo https://helm.ngc.nvidia.com/nvidia \
-n gpu-operator --create-namespace --version=v26.3.0

What to do next

  1. Verify that the Envoy Gateway CRDs and controller are installed and ready.
  2. Verify that the KServe CRDs and controller resources are ready.
  3. Verify that the CloudNativePG operator is ready.
  4. Verify that the LeaderWorkerSet controller is ready.
  5. Verify that the OpenTelemetry Operator is ready.
  6. Verify that Prometheus monitoring is ready.
  7. Verify that the NVIDIA GPU Operator is ready.
Deploying Nutanix Enterprise AI on Azure Kubernetes Service Install or upgrade Nutanix Enterprise AI on Azure Kubernetes Service (AKS).

Before you begin

Ensure to complete the following:

Perform preflight checks and obtain the expected output. For more information on preflight checks and the expected output, see

Performing Preflight Checks Before Installing Nutanix Enterprise AI on Azure AKS

on

page 77.

Installing Prerequisite Components on an

Install components on the AKS cluster. For more information, see

AKS Cluster

on page 79.

Upgrades from version 2.7.0 to 2.8.0 must be executed during planned downtime. User logins will be unavailable during the upgrade, and full functionality will resume automatically once the upgrade is complete.

About this task

To install or upgrade Nutanix Enterprise AI on AKS, follow these steps:

Procedure

1. Setup NAI Helm repository

a. Add and update the Nutanix Helm repository, which contains the

Helm chart:

nai-core
helm repo add ntnx-charts https://nutanix.github.io/helm-releases && helm repo
update ntnx-charts

b. Search for the version of the

and

helm chart available for installation in the
nai-operators
nai-core

Nutanix Helm repository:

helm search repo ntnx-charts/nai-operators --versions
helm search repo ntnx-charts/nai-core --versions

2. Create Docker Registry Secrets:

Create the

namespace and the

secret in both

and

nai-system
docker-registry
nai-system
envoy-

namespaces.

gateway-system

The

is already present on the cluster.

envoy-gateway-system namespace
export REGISTRY_SECRET_NAME=nai-regcred
export DOCKER_SERVER=https://index.docker.io/v1/
export DOCKER_USERNAME=<docker-username>
export DOCKER_PASSWORD=<docker-password>
export DOCKER_EMAIL=<docker-email>
kubectl create namespace nai-system --dry-run=client -o yaml | kubectl apply -f -
kubectl -n nai-system create secret docker-registry ${REGISTRY_SECRET_NAME} \
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n envoy-gateway-system create secret docker-registry ${REGISTRY_SECRET_NAME}
\
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -

3. Deploy NAI Operators:

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0  -n
nai-system --create-namespace --take-ownership --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}"

[!NOTE] Note: Ensure the REGISTRY_SECRET_NAME environment variable is set before running this command.

4. Pull the

helm chart and extract it to your local repository:
nai-core
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

5. Install or upgrade NAI Core on AKS:

# Set the environment variable
NAI_API_RWX_STORAGECLASS=<NFS Storageclass i.e nai-nfs-storage>
NAI_DEFAULT_RWO_STORAGECLASS=<default storageclass>
REGISTRY_SECRET_NAME=<secret created>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 -n nai-system --
create-namespace --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/aks-values.yaml

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

Enable the Chat and Talk to My Data application. These applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

Enable HTTPS with a self-signed certificate. By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

using your own certificate or a cert-manager ClusterIssuer, see

TLS Encryption on Nutanix

Enterprise AI

.

Scale out the ingress gateway and NAI API. To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Configure PostgreSQL database connections. To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default value is 1000. Increase this value if you expect a higher number of concurrent clients.

  1. Configure the TLS certificate.

TLS Encryption on Nutanix Enterprise AI

For more information, see

on page 119.

What to do next

Ensure that all the pods are successfully deployed in the

namespace and are in the

state by:

nai-system
Ready
kubectl get pods -n nai-system

The following is a sample output of the command. Match the pod names by ignoring the auto-generated suffix added by Kubernetes and its status.

[!NOTE] Note:

The number of

must be equal to the number of worker

nai-otel-collector-collector pods

nodes in the NAI Kubernetes cluster.

NAME                                                     READY   STATUS
ai-gateway-controller-6644f54c64-z2cwm                   1/1     Running
chi-nai-clickhouse-server-chcluster1-0-0-0               1/1     Running
chk-nai-clickhouse-keeper-chkeeper-0-0-0                 1/1     Running
iam-database-bootstrap-mlgpu-hk5zz                       0/1     Completed
iam-proxy-68d975978d-r7llr                               1/1     Running
iam-proxy-control-plane-d876b77dd-7swx2                  1/1     Running
iam-themis-679b9bbff8-nfknh                              1/1     Running
iam-themis-bootstrap-e1fhx-kglkm                         0/1     Completed
iam-ui-6cb6d49fcc-fhnhm                                  1/1     Running
iam-user-authn-5dbbdcbfcc-9j6jr                          1/1     Running
nai-agent-648c7b8c8d-rzq28                               1/1     Running
nai-api-85b8694cfc-5l5j8                                 1/1     Running
nai-api-db-migrate-lweo1-w8zfx                           0/1     Completed
nai-clickhouse-schema-job-1778237791-fs7p8               0/1     Completed
nai-db-0                                                 1/1     Running
nai-iep-model-controller-b5b8bf9-jxfrk                   1/1     Running
nai-labs-689dcb4644-l6f4p                                1/1     Running
nai-oauth2-proxy-579c5b4d9f-mn72l                        1/1     Running
nai-operators-nai-clickhouse-operator-858fdb9b94-zs8gx   2/2     Running
nai-otel-collector-collector-49bsx                       1/1     Running
nai-otel-collector-collector-6tgh5                       1/1     Running
nai-otel-collector-collector-8qp9x                       1/1     Running
nai-otel-collector-targetallocator-748856644d-5lnff      1/1     Running
nai-securityscan-manager-678ffd75ff-9t95c                1/1     Running
nai-ui-76d74f55bb-c28rw                                  1/1     Running
redis-standalone-67ccd5cc8f-6hp55                        2/2     Running

Verify that the Envoy Gateway, ingress gateway, and rate limit pods in the envoy-gateway-system namespace are in the

state:

Running
kubectl get pods -n envoy-gateway-system
NAME                                                            READY    STATUS
envoy-gateway-6b987d469d-5l2w9                                   1/1     Running
envoy-nai-system-nai-ingress-gateway-ff52ba1f-7fd4897bd4-7smgg   2/2     Running
envoy-ratelimit-85b55c877c-2n2bf                                 1/1     Running

Pending

Failed

If you had endpoints in

status before the upgrade displaying the status as

with the message

Unable to pull runtime image with provided credentials, hibernate and resume the endpoints.

Access NAI Dashboard IP:

kubectl get svc -n envoy-gateway-system -l "gateway.envoyproxy.io/owning-gateway-
name=nai-ingress-gateway,gateway.envoyproxy.io/owning-gateway-namespace=nai-system" -
o jsonpath='{.items[0].status.loadBalancer.ingress[0].ip}'

Dashboard

Log in to Nutanix Enterprise AI

For information on viewing the

, see

on page 167.

Deploy Nutanix Enterprise AI on Google Kubernetes Engine

Install or upgrade Nutanix Enterprise AI (NAI) on Google Kubernetes Engine (GKE).

To install or upgrade NAI, follow these steps:

1. Set up a GKE cluster. For more information, see

Setting up a Google Kubernetes Engine Cluster

on

page 85.

2. Perform preflight checks. For more information, see

Performing Preflight Checks Before Installing Nutanix

Enterprise AI on Google Kubernetes Engine

on page 85.

3. Install prerequisite components. For more information, see

Installing Prerequisite Components on a Google

Kubernetes Engine Cluster

on page 86.

4. Deploy NAI on GKE. For more information, see

Deploying Nutanix Enterprise AI on Google Kubernetes

Engine

on page 90.

Setting up a Google Kubernetes Engine Cluster Set up your Google Kubernetes Engine (GKE) cluster for Nutanix Enterprise AI. This procedure is not applicable for Nutanix Kubernetes Platform (NKP) managed or attached GKE clusters.

Before you begin

Ensure that your cluster meets all the requirements mentioned in

Nutanix Enterprise AI - Private Inference

and Agent Gateway Requirements

on page 10.

About this task

To set up your GKE cluster before you deploy Nutanix Enterprise AI, follow these steps:

Procedure

  1. Configure and authorize the Google Cloud CLI.

For more information, see

Google Cloud documentation

.

2. Enable

.

Required Services

This is required to enable Google files CSI driver and default RWX storage classes. For more information, see

Google Cloud documentation

.

3. Create a GKE cluster with Kubernetes version 1.33 or 1.34 based on the

GKE GPU Operator NVIDIA Driver Manager

workflow.

For more information, see

Google Cloud documentation

.

4. Configure the GKE cluster to enable network policy enforcement. Network policy enforcement is built into GKE

Dataplane V2 but needs to be enabled when Dataplane V2 is disabled.

Google Cloud documentation

For more information, see

.

  1. Create a default node group. Make sure the nodes have accelerator AVX2 or newer.
  2. Create a GKE node pool and add it to the cluster.

Ensure that you skip installing the Google GPU driver when adding the GKE node pool to the cluster. For more information, see

Google Cloud documentation

.

  1. Configure the node pool to create the GPU operator namespace and deploy the GPU operator resource quota.

For more information, see

Google Cloud documentation

.

Performing Preflight Checks Before Installing Nutanix Enterprise AI on Google Kubernetes Engine Perform preflight checks before installing NAI on Google Kubernetes Engine (GKE). Preflight checks ensure the Kubernetes environment meets all the technical requirements for a successful NAI installation.

About this task

To perform preflight checks for installing NAI on GKE, follow these steps:

Procedure

1. Make sure your Kubernetes version is 1.33 or 1.34:

kubectl get nodes --selector='!node-role.kubernetes.io/control-plane' \
-o custom-
columns=NAME:.metadata.name,KUBELET_VERSION:.status.nodeInfo.kubeletVersion \
--no-headers

The expected output is that the Kubernetes version must be 1.33 or 1.34.

2. Verify if the number of

CSI pods match the number of worker nodes:

filestore-node
[ $(kubectl get nodes --no-headers | wc -l) -eq $(kubectl get pods -n kube-system
--no-headers | grep filestore-node | wc -l) ] && echo "# filestore-node pods = node
count" || echo "# Mismatch: filestore-node pods != node count"

The expected output is

.

filestore-node pods= node count

3. Confirm that the

class binds volumes immediately:

nai-nfs-storage
kubectl get storageclass nai-nfs-storage -o jsonpath='{.volumeBindingMode}{"\n"}'

The expected output is

.

immediate

4. Check if the

storage class has

access:

nai-nfs-storage
ReadWriteMany
kubectl apply -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-rwx-pvc
namespace: default
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Gi
storageClassName: nai-nfs-storage
EOF
kubectl get pvc test-rwx-pvc -n default

The expected output is that the storage class

has

set to

.

nai-nfs-storage
Access Modes
RWX
NAME           STATUS   VOLUME                                     CAPACITY   ACCESS
MODES   STORAGECLASS      VOLUMEATTRIBUTESCLASS   AGE
test-rwx-pvc            Bound    pvc-0df5145a-590c-4f28-9f9d-cefefc55266c   1Gi
RWX            nai-nfs-storage   <unset>                 5s
kubectl delete pvc test-rwx-pvc

Installing Prerequisite Components on a Google Kubernetes Engine Cluster Install components on a Google Kubernetes Engine (GKE) cluster.

About this task

To install components on a GKE cluster, follow these steps:

Procedure

1. Install cert-manager:

helm upgrade --install cert-manager cert-manager --repo https://charts.jetstack.io --
version v1.19.3 --set installCRDs=true -n cert-manager --create-namespace --wait

2. Install or upgrade Envoy Gateway:

a. Install Envoy Gateway and Gateway API CRDs:

helm template eg oci://docker.io/envoyproxy/gateway-crds-helm --version v1.8.1 \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --server-side --force-conflicts -f -

b. Create

with the following configuration:

envoy-gateway-config.yaml
config:
envoyGateway:
gateway:
controllerName: "gateway.envoyproxy.io/gatewayclass-controller"
logging:
level:
default: "info"
provider:
kubernetes:
rateLimitDeployment:
container:
image: "docker.io/envoyproxy/ratelimit:1e50889b"
patch:
type: "StrategicMerge"
value:
spec:
template:
spec:
containers:
- imagePullPolicy: "IfNotPresent"
name: "envoy-ratelimit"
image: "docker.io/envoyproxy/ratelimit:1e50889b"
env:
- name: REDIS_TYPE
value: "sentinel"
- name: REDIS_PIPELINE_WINDOW
value: "150us"
type: "Kubernetes"
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
maxMessageSize: 11Mi
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- "Translation"
- "Cluster"
- "Route"
service:
fqdn:
hostname: "ai-gateway-controller.nai-system.svc.cluster.local"
port: 1063
rateLimit:
backend:
type: "Redis"
redis:
url: "mymaster,nai-valkey-sentinel.nai-system.svc.cluster.local:26379"

[!NOTE] Note: The rate-limit backend uses Valkey Sentinel with master name

mymaster

and the

nai-valkey-

service in

.

sentinel
nai-system

c. Install or Upgrade Envoy Gateway:

helm upgrade --install eg oci://docker.io/envoyproxy/gateway-helm --version v1.8.1
\
-n envoy-gateway-system --create-namespace --skip-crds \
-f "./envoy-gateway-config.yaml"

3. Install or Upgrade KServe:

The required version is KSERVE_VERSION=v0.19.0.

a. Install or upgrade the KServe CRDs:

helm upgrade --install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

b. Install or upgrade the KServe resources:

helm upgrade --install kserve oci://ghcr.io/kserve/charts/kserve-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.controller.deploymentMode=RawDeployment \
--set kserve.controller.gateway.disableIngressCreation=true

c. Install or upgrade the KServe LLMInferenceService CRD:

helm upgrade --install kserve-llmisvc-crd oci://ghcr.io/kserve/charts/kserve-
llmisvc-crd \
--version $KSERVE_VERSION -n kserve --create-namespace --wait

d. Install or upgrade the KServe LLMInferenceService resources:

helm upgrade --install kserve-llmisvc-resources oci://ghcr.io/kserve/charts/
kserve-llmisvc-resources \
--version $KSERVE_VERSION -n kserve --create-namespace --wait \
--set kserve.createSharedResources=false \
--set kserve.llmisvc.createGIECRDs=false

4. Install CloudNativePG Operator:

helm install cnpg cloudnative-pg \
--repo https://cloudnative-pg.github.io/charts \
--version 0.28.0 -n cnpg-system --create-namespace --wait

5. Install LeaderWorkerSet:

helm install lws oci://registry.k8s.io/lws/charts/lws \
--version 0.8.0 -n lws-system --create-namespace --wait

6. Install or upgrade OpenTelemetry Operator

helm upgrade --install opentelemetry-operator opentelemetry-operator \
--repo https://open-telemetry.github.io/opentelemetry-helm-charts \
--version=0.114.1 -n opentelemetry --create-namespace --wait

7. Install or upgrade Prometheus Monitoring:

helm upgrade --install prometheus kube-prometheus-stack \
--repo https://prometheus-community.github.io/helm-charts \
--version=82.13.6 -n prometheus --create-namespace --wait \
--set grafana.enabled=false \
--set prometheus.enabled=false \
--set kubeStateMetrics.enabled=false \
--set alertManager.enabled=false \
--set kubernetesServiceMonitors.enabled=false \
--set prometheus-node-exporter.kubeRBACProxy.enabled=true

8. Install or upgrade NVIDIA GPU Operator :

helm upgrade --install --wait gpu-operator gpu-operator \
--repo https://helm.ngc.nvidia.com/nvidia \
-n gpu-operator --create-namespace --version=v26.3.0

9. Apply a resource quota:

kubectl create namespace gpu-operator --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -n gpu-operator -f - << EOF
apiVersion: v1
kind: ResourceQuota
metadata:
name: gpu-operator-quota
spec:
hard:
pods: 100
scopeSelector:
matchExpressions:
- operator: In
scopeName: PriorityClass
values:
- system-node-critical
- system-cluster-critical
EOF

What to do next

  1. Verify that the Envoy Gateway CRDs and controller are installed and ready.
  2. Verify that the KServe CRDs and controller resources are ready.
  3. Verify that the CloudNativePG operator is ready.
  4. Verify that the LeaderWorkerSet controller is ready.
  5. Verify that the OpenTelemetry Operator is ready. 6. Verify that Prometheus monitoring is ready.
  6. Verify that the NVIDIA GPU Operator is ready.

Deploying Nutanix Enterprise AI on Google Kubernetes Engine Install Nutanix Enterprise AI on Google Kubernetes Engine (GKE).

Before you begin

Ensure to complete the following:

Installing Prerequisite Components

Install components on the Kubernetes cluster. For more information, see

on a Google Kubernetes Engine Cluster

on page 86.

Upgrades from version 2.7.0 to 2.8.0 must be executed during planned downtime. User logins will be unavailable during the upgrade, and full functionality will resume automatically once the upgrade is complete.

About this task

To install Nutanix Enterprise AI on GKE, follow these steps:

Procedure

1. Choose a profile:

Profile-based deployment allows you to select a predefined configuration based on your environment and availability requirements. The default profile uses single replica for components and is intended for baseline deployment. The other profiles are for higher capacity usage.

Table 38: Profile and Capacity

ProfileCapacity
Default300 concurrent requests and 100 API Keys
c1k_k2001000 concurrent requests and 200 API Keys
c5k_k1k5000 concurrent requests and 1000 API Keys

2. Pull and untar both the 2.8.0 charts

helm pull ntnx-charts/nai-operators --version 2.8.0 --untar=true
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

The extracted charts contain profile files such as:

./nai-operators/profiles/c1k_k200.yaml
./nai-operators/profiles/c5k_k1k.yaml
./nai-core/profiles/c1k_k200.yaml
./nai-core/profiles/c5k_k1k.yaml

3. Deploy for small profile

a. Deploy NAI Operators for c1k_k200 profile

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
-f ./nai-operators/profiles/c1k_k200.yaml \
-f ./nai-operators/gke-values.yaml

b. Deploy NAI Core for c1k_k200 profile

export NAI_API_RWX_STORAGECLASS=<RWX/NFS storageclass>
export NAI_DEFAULT_RWO_STORAGECLASS=<RWO default storageclass>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/profiles/c1k_k200.yaml \
-f ./nai-core/gke-values.yaml

4. Setup NAI Helm repository

a. Add and update the Nutanix Helm repository, which contains the

Helm chart:

nai-core
helm repo add ntnx-charts https://nutanix.github.io/helm-releases && helm repo
update ntnx-charts

b. Search for the version of the

and

Helm chart available for installation in the

nai-operators
nai-core

Nutanix helm repository:

helm search repo ntnx-charts/nai-operators --versions
helm search repo ntnx-charts/nai-core --versions

5. Create Docker Registry Secrets:

Create the

namespace and the

secret in both

and

nai-system
docker-registry
nai-system
envoy-

namespaces.

gateway-system

The

is already present on the cluster.

envoy-gateway-system namespace
export REGISTRY_SECRET_NAME=nai-regcred
export DOCKER_SERVER=https://index.docker.io/v1/
export DOCKER_USERNAME=<docker-username>
export DOCKER_PASSWORD=<docker-password>
export DOCKER_EMAIL=<docker-email>
kubectl create namespace nai-system --dry-run=client -o yaml | kubectl apply -f -
kubectl -n nai-system create secret docker-registry ${REGISTRY_SECRET_NAME} \
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n envoy-gateway-system create secret docker-registry ${REGISTRY_SECRET_NAME}
\
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -

6. Deploy NAI Operators:

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0  -n
nai-system --create-namespace --take-ownership --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}"

[!NOTE] Note: Ensure the REGISTRY_SECRET_NAME environment variable is set before running this command.

7. Pull the

helm chart and extract it to your local repository:
nai-core
helm pull ntnx-charts/nai-core --version 2.8.0 --untar=true

8. Install or upgrade NAI Core on GKE:

# Set the environment variable
NAI_API_RWX_STORAGECLASS=<NFS Storageclass i.e nai-nfs-storage>
NAI_DEFAULT_RWO_STORAGECLASS=<default storageclass>
REGISTRY_SECRET_NAME=<secret created>
helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 -n nai-system --
create-namespace --wait \
--set "global.imagePullSecrets[0].name=${REGISTRY_SECRET_NAME}" \
--set "global.storage.storageClassNameRWX=${NAI_API_RWX_STORAGECLASS}" \
--set "global.storage.storageClassName=${NAI_DEFAULT_RWO_STORAGECLASS}" \
-f ./nai-core/gke-values.yaml

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

Enable the Chat and Talk to My Data application. These applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

Enable HTTPS with a self-signed certificate. By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

using your own certificate or a cert-manager ClusterIssuer, see

TLS Encryption on Nutanix

Enterprise AI

.

Scale out the ingress gateway and NAI API. To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Configure PostgreSQL database connections. To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default value is 1000. Increase this value if you expect a higher number of concurrent clients.

  1. Configure the TLS certificate.

TLS Encryption on Nutanix Enterprise AI

For more information, see

on page 119.

What to do next

Ensure that all the pods are successfully deployed in the

namespace and are in the

state:

nai-system
Ready
kubectl get pods -n nai-system

The following is a sample output of the command:

NAME                                                     READY   STATUS
ai-gateway-controller-6644f54c64-z2cwm                   1/1     Running
chi-nai-clickhouse-server-chcluster1-0-0-0               1/1     Running
chk-nai-clickhouse-keeper-chkeeper-0-0-0                 1/1     Running
iam-database-bootstrap-mlgpu-hk5zz                       0/1     Completed
iam-proxy-68d975978d-r7llr                               1/1     Running
iam-proxy-control-plane-d876b77dd-7swx2                  1/1     Running
iam-themis-679b9bbff8-nfknh                              1/1     Running
iam-themis-bootstrap-e1fhx-kglkm                         0/1     Completed
iam-ui-6cb6d49fcc-fhnhm                                  1/1     Running
iam-user-authn-5dbbdcbfcc-9j6jr                          1/1     Running
nai-agent-648c7b8c8d-rzq28                               1/1     Running
nai-api-85b8694cfc-5l5j8                                 1/1     Running
nai-api-db-migrate-lweo1-w8zfx                           0/1     Completed
nai-clickhouse-schema-job-1778237791-fs7p8               0/1     Completed
nai-db-0                                                 1/1     Running
nai-iep-model-controller-b5b8bf9-jxfrk                   1/1     Running
nai-labs-689dcb4644-l6f4p                                1/1     Running
nai-oauth2-proxy-579c5b4d9f-mn72l                        1/1     Running
nai-operators-nai-clickhouse-operator-858fdb9b94-zs8gx   2/2     Running
nai-otel-collector-collector-49bsx                       1/1     Running
nai-otel-collector-collector-6tgh5                       1/1     Running
nai-otel-collector-collector-8qp9x                       1/1     Running
nai-otel-collector-targetallocator-748856644d-5lnff      1/1     Running
nai-securityscan-manager-678ffd75ff-9t95c                1/1     Running
nai-ui-76d74f55bb-c28rw                                  1/1     Running
redis-standalone-67ccd5cc8f-6hp55                        2/2     Running

Verify that the Envoy Gateway, ingress gateway, and rate limit pods in the envoy-gateway-system namespace are in the

state:

Running
kubectl get pods -n envoy-gateway-system
NAME                                                            READY    STATUS
envoy-gateway-6b987d469d-5l2w9                                   1/1     Running
envoy-nai-system-nai-ingress-gateway-ff52ba1f-7fd4897bd4-7smgg   2/2     Running
envoy-ratelimit-85b55c877c-2n2bf                                 1/1     Running

Access NAI Dashboard IP:

kubectl get svc -n envoy-gateway-system -l "gateway.envoyproxy.io/owning-gateway-
name=nai-ingress-gateway,gateway.envoyproxy.io/owning-gateway-namespace=nai-system" -
o jsonpath='{.items[0].status.loadBalancer.ingress[0].ip}'

For information on viewing the

Dashboard

, see

Log in to Nutanix Enterprise AI

on page 167.

Pending

Failed

If you had endpoints in

status before the upgrade displaying the status as

with the message

Unable to pull runtime image with provided credentials, hibernate and resume the endpoints.

Deploy Nutanix Enterprise AI with Self-managed PostgreSQL

Install or upgrade Nutanix Enterprise AI on a Kubernetes cluster with a self-hosted or managed PostgreSQL database server.

Performing Preflight

Before install or upgrade, you must first perform prelight checks. For more information, see

Checks Before Deploying Nutanix Enterprise AI on Nutanix Kubernetes Platform

on page 42. After

performing preflight checks, you can do any of the following:

Deploy in air-gapped environment. For more information, see

Deploying Nutanix Enterprise AI Components

on a Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments with PostgreSQL

on

page 95.

Deploy on connected clusters. For more information, see

Deploying Nutanix Enterprise AI for a Connected

NKP Cluster

on page 113.

Performing Preflight Checks Before Deploying Nutanix Enterprise AI with Self-managed PostgreSQL Perform preflight checks before installing or upgrading Nutanix Enterprise AI (NAI) on Nutanix Kubernetes Platform (NKP). Preflight checks ensure the Kubernetes environment meets all the technical requirements for a successful NAI installation.

About this task

To perform preflight checks for installing or upgrading NAI on NKP, follow these steps:

Procedure

1. Verify that the Kubernetes version is 1.33 or 1.34:

kubectl get nodes \
--selector='!node-role.kubernetes.io/control-plane,!node-role.kubernetes.io/master'
\
-o custom-
columns=NODE:.metadata.name,KUBELET_VERSION:.status.nodeInfo.kubeletVersion

The expected output is that the Kubernetes version must be 1.33 or 1.34.

2. Verify if the number of CSI pods matches the number of worker nodes:

[ $(kubectl get nodes --no-headers | wc -l) -eq $(kubectl get pods -n ntnx-system --
no-headers | grep csi-node | wc -l) ] && echo "# CSI pods = node count" || echo "#
Mismatch: CSI pods != node count"

The expected output is CSI pods = node count.

3. Verify that the

class binds volumes immediately:

nai-nfs-storage
kubectl get storageclass nai-nfs-storage -o jsonpath='{.volumeBindingMode}{"\n"}'

The expected output is

.

immediate

4. Verify if the

storage class has ReadWriteMany access:

nai-nfs-storage
kubectl apply -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: test-rwx-pvc
namespace: default
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 1Gi
storageClassName: nai-nfs-storage
EOF
kubectl get pvc test-rwx-pvc -n default

The expected output is that the storage class

has

set to

.

nai-nfs-storage
Access Modes
RWX
NAME           STATUS   VOLUME                                     CAPACITY   ACCESS
MODES   STORAGECLASS      VOLUMEATTRIBUTESCLASS   AGE
test-rwx-pvc            Bound    pvc-0df5145a-590c-4f28-9f9d-cefefc55266c   1Gi
RWX            nai-nfs-storage   <unset>                 5s
kubectl delete pvc test-rwx-pvc

Deploying Nutanix Enterprise AI Components on a Nutanix Kubernetes® Platform Cluster in Air- Gapped Environments with PostgreSQL Deploy Nutanix Enterprise AI(NAI) components on a Nutanix Kubernetes® Platform (NKP) cluster in an air- gapped environment with a PostgreSQL database.

Before you begin

Nutanix Enterprise AI supports air-gapped installation only on NKP.

Ensure that you meet requirements listed in

Prerequisites for Deploying Nutanix Enterprise AI 2.8.0 on a

Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments with PostgreSQL

on page 97.

About this task

To deploy Nutanix Enterprise AI components on an NKP cluster in an air-gapped environment with PostgreSQL, follow these steps:

Procedure

Download the required NAI 2.8.0 release bundles from the Nutanix portal.

Downloading Nutanix Enterprise AI Air-Gap Release Bundles in Air-Gapped

For more information, see

Environments with PostgreSQL

on page 97.

Push Container Images to a private registry.

From this step onwards, execute the all commands from within your air-gapped environment typically a jumpbox or bastion host that has access to both your container registry and your Kubernetes cluster. Ensure you have Docker and kubectl available on this machine before proceeding.

Pushing Container Images to Private Registry in Air-Gapped Environments

For more information, see

with PostgreSQL

on page 98.

Extract the NAI Helm charts bundle to your working directory:

tar -xvf nai-helm-charts-2.8.0.tar

The extraction produces the following Helm chart archives:

gateway-crds-helm-v1.6.3.tgz
gateway-helm-v1.6.3.tgz
kserve-crd-v0.15.0.tgz
kserve-v0.15.0.tgz
nai-core-2.8.0.tgz
nai-operators-2.8.0.tgz
opentelemetry-operator-0.102.0.tgz

(Optional) Publish Helm Charts to Private Registry

To install Helm charts directly from your OCI-compatible registry instead of local files, push the charts using the following commands:

# Authenticate to your registry
helm registry login -u <username> -p <password> https://<registry>
# Push each chart to the registry
helm push gateway-crds-helm-v1.7.0.tgz oci://<registry>
helm push gateway-helm-v1.7.0.tgz oci://<registry>
helm push kserve-crd-v0.15.0.tgz oci://<registry>
helm push kserve-v0.15.0.tgz oci://<registry>
helm push opentelemetry-operator-0.102.0.tgz oci://<registry>
helm push nai-core-2.8.0.tgz oci://<registry>
helm push nai-operators-2.8.0.tgz oci://<registry>

The remaining steps assume installation from local Helm chart files.

Configure Registry Credentials

Configuring Docker Registry Credentials for Dependencies in an Air-gapped

For more information, see

Environment with PostgreSQL

on page 103.

Install Prometheus Monitoring from the Nutanix Kubernetes Platform (NKP) platform applications catalog. Prometheus Monitoring is not included in the Nutanix Enterprise AI air-gap image tar or the helm-charts tar. On an NKP cluster in an air-gapped environment, install Prometheus Monitoring from the NKP platform applications catalog before you install Nutanix Enterprise AI components. Enable the Prometheus Monitoring platform application on your NKP workload cluster.

To optimize resource utilization on the workload cluster, configure Prometheus Monitoring with the following minimum installation settings when you enable the application:

alertmanager:
enabled: false
grafana:
enabled: false
prometheus:
enabled: false
kubeStateMetrics:
enabled: false
kubernetesServiceMonitors:
enabled: false
prometheus-node-exporter.kubeRBACProxy:
kubeRBACProxy:
enabled: true

For more information, see

Pro: Enabling an Application Using the UI

.

Install NVIDIA GPU Operator from NKP platform applications.

If the GPU nodes do not have precompiled NVIDIA drivers installed, enable driver installation in the NVIDIA GPU Operator configuration. This setting ensures that the NVIDIA drivers are installed on the GPU nodes. When enabling the NVIDIA GPU Operator, add the following cluster override:

driver:
enabled: true

For more information, see

Pro: Enabling an Application Using the UI

.

Install CloudNativePG Operator from NKP platform applications.

Deploy the LeaderWorkerSet controller for multi-host inference workloads:

helm upgrade --install lws ./lws-0.8.0.tgz -n lws-system --create-namespace --wait
\
--set "imagePullSecrets[0].name=${IMAGE_PULL_SECRET}" \
--set image.manager.repository=${REGISTRY}/${PROJECT}/nai-lws
  1. Install Envoy Gateway.

For more information, see

Installing Envoy Gateway in an Air-gapped Environment with PostgreSQL

on

page 104.

  1. Install KServe.

Installing KServe in Air-Gapped Environments with PostgreSQL

For more information, see

on

page 106.

  1. Deploy the OpenTelemetry Operator.

For more information, see

Deploying the OpenTelemetry Operator in Air-Gapped Environments with

PostgreSQL

on page 106.

  1. Install NAI components.

For more information, see

Deploying Nutanix Enterprise AI Components in Air-Gapped Environments

with PostgreSQL

on page 106.

What to do next

1. Verify NAI operators installation. Check the status of NAI components:

# Check all pods in nai-system namespace
kubectl get pods -n nai-system
# Check NAI Operators and Core Helm release
helm list -n nai-system
# Check persistent volume claims
kubectl get pvc -n nai-system

2. Verify if NAI services are accessible:

# List all services in nai-system
kubectl get svc -n nai-system
# Check NAI API service
kubectl get svc -n nai-system nai-api
# Check NAI UI service
kubectl get svc -n nai-system nai-inference-ui
  1. Troubleshoot common issues that can occur when you are installing NAI 2.8 in air-gapped environments.

For more information, see

Troubleshooting Deployment of Nutanix Enterprise AI 2.8.0 in Air-Gapped

Environments with PostgreSQL

on page 113.

Prerequisites for Deploying Nutanix Enterprise AI 2.8.0 on a Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments with PostgreSQL

Before proceeding with the installation, ensure the following requirements are met:

Nutanix Kubernetes Platform (NKP) cluster running Kubernetes 1.35.

kubectl CLI v1.33+ configured with cluster access

Helm CLI v4.0.5

Docker or compatible container runtime (for loading images)

Access to a private container registry

Registry credentials with appropriate permissions

Sufficient disk space for loading images (~100GB)

Downloading Nutanix Enterprise AI Air-Gap Release Bundles in Air-Gapped Environments with PostgreSQL

Download the required Nutanix Enterprise AI 2.8.0 release bundles from the Nutanix Portal.

About this task

To download the required NAI 2.8.0 release bundles from the Nutanix Portal, follow these steps:

Procedure

1. Navigate to the Nutanix Enterprise AI page on the Nutanix Support Portal

2. Select NAI Version 2.8.0 from the available releases

3. Download the following two bundles:

NAI Air-Gap Bundle (

)

nai-v2.8.0.tar

NAI Helm Charts Bundle (

)

nai-helm-charts-2.8.0.tar

Table 39: Air-Gap Download Bundles

BundleDescriptionSize
NAI Air-Gap Bundle (nai- v2.8.0.tar)~70-80GB

Contains all NAI container images and dependencies

Required for air-gapped deployments

NAI Helm Charts Bundle (nai- helm-charts-2.8.0.tar)

~5 MB

Contains NAI Helm charts and all dependency charts

Includes Envoy Gateway, KServe, LeaderWorkerSet and OpenTelemetry charts

4. Transfer both bundles to your air-gapped environment using approved methods such as USB drive, secure file

transfer, and so on.

Pushing Container Images to Private Registry in Air-Gapped Environments with PostgreSQL

Before installing NAI with PostgreSQL, you must load all container images and push them to your private registry.

Before you begin

Execute all the commands from within your air-gapped environment typically a jumpbox or bastion host that has access to both your container registry and your Kubernetes cluster. Ensure you have Docker and kubectl available on this machine before proceeding.

About this task

To push container images to a private registry, follow these steps:

Procedure

1. Login to your private container registry:

docker login <registry-url>

For example,

docker login registry.example.com

The system displays a prompt to enter your registry credentials.

  1. Enter your registry credentials when prompted.
  2. Create the Image Push Script.

The Image Push Script pushes container images to your private registry.

a. Create a project/repository with the name

in your container registry, where all NAI images are

nutanix

stored. For example, in Harbor this would be a project named

, resulting in image paths like

nutanix

.

registry.example.com/nutanix/<image-name>:<tag>

b. Create a script file named

with the following content:

push-images-to-registry.sh
#!/bin/bash
#
# NAI Images - Load, Retag, and Push to Private Registry
#
# This script loads NAI container images from a tar bundle, retags them for your
# private registry, and pushes them to the registry.
#
# Prerequisites:
#   - Docker installed and running
#   - Docker logged into the target registry (docker login)
#   - NAI images tar bundle file
#
# Usage:
#   ./push-images-to-registry.sh <registry-url> <project> <tar-file>
#
# Example:
#   ./push-images-to-registry.sh registry.example.com nutanix nai-images-2.8.0.tar
#
set -uo pipefail
# ============================================================================
# Helper Functions
# ============================================================================
print_header() {
echo ""
echo "========================================"
echo "$1"
echo "========================================"
}
print_success() {
echo "# $1"
}
print_error() {
echo "# ERROR: $1" >&2
}
print_info() {
echo "# $1"
}
# ============================================================================
# Validate Arguments
# ============================================================================
if [ $# -ne 3 ]; then
echo "Usage: $0 <registry-url> <project> <tar-file>"
echo ""
echo "Arguments:"
echo "  registry-url    Your private registry URL (e.g.,
registry.example.com)"
echo "  project         Project/repository name in the registry (e.g.,
nutanix)"
echo "  tar-file        Path to the NAI images tar bundle"
echo ""
echo "Example:"
echo "  $0 registry.example.com nutanix nai-images-2.8.0.tar"
echo ""
exit 1
fi
REGISTRY="$1"
PROJECT="$2"
TAR_FILE="$3"
# Validate tar file exists
if [ ! -f "$TAR_FILE" ]; then
print_error "Tar file not found: $TAR_FILE"
exit 1
fi
# ============================================================================
# Configuration
# ============================================================================
print_header "NAI Images - Load, Retag & Push"
echo "Registry:  $REGISTRY"
echo "Project:   $PROJECT"
echo "Tar File:  $TAR_FILE"
echo "Date:      $(date)"
# Arrays to track images
LOADED_IMAGES=()
FAILED_IMAGES=()
# ============================================================================
# Step 1: Load Images from Tar Bundle
# ============================================================================
print_header "Step 1: Loading Images from Tar Bundle"
print_info "Loading images from $TAR_FILE..."
LOAD_OUTPUT=$(docker load -i "$TAR_FILE" 2>&1)
# Extract loaded image names
while IFS= read -r line; do
if [[ "$line" =~ Loaded\ image:\ (.+)$ ]]; then
LOADED_IMAGES+=("${BASH_REMATCH[1]}")
fi
done <<< "$LOAD_OUTPUT"
if [ ${#LOADED_IMAGES[@]} -eq 0 ]; then
print_error "No images were loaded from the tar file"
exit 1
fi
print_success "Loaded ${#LOADED_IMAGES[@]} images"
# ============================================================================
# Step 2: Retag and Push Images
# ============================================================================
print_header "Step 2: Retagging and Pushing Images"
PUSHED_COUNT=0
TOTAL_IMAGES=${#LOADED_IMAGES[@]}
for source_image in "${LOADED_IMAGES[@]}"; do
echo ""
print_info "[$((PUSHED_COUNT + 1))/$TOTAL_IMAGES] Processing: $source_image"
# Retag image for target registry
# Format: nutanix/nai-api:v2.8.0 # registry.example.com/<project>/nai-
api:v2.8.0
if [[ "$source_image" =~ ^nutanix/(.+)$ ]]; then
image_path="${BASH_REMATCH[1]}"
target_image="${REGISTRY}/${PROJECT}/${image_path}"
print_info "Tagging as: $target_image"
if ! docker tag "$source_image" "$target_image"; then
print_error "Failed to tag image"
FAILED_IMAGES+=("$source_image")
continue
fi
print_info "Pushing to registry..."
if docker push "$target_image"; then
print_success "Pushed successfully"
((PUSHED_COUNT++))
else
print_error "Failed to push image"
FAILED_IMAGES+=("$target_image")
fi
else
print_info "Skipping (not in nutanix/* format)"
fi
done
# ============================================================================
# Summary
# ============================================================================
print_header "Summary"
echo "Total images loaded:    $TOTAL_IMAGES"
echo "Successfully pushed:    $PUSHED_COUNT"
echo "Failed:                 ${#FAILED_IMAGES[@]}"
if [ ${#FAILED_IMAGES[@]} -gt 0 ]; then
echo ""
print_error "The following images failed:"
for img in "${FAILED_IMAGES[@]}"; do
echo "  - $img"
done
echo ""
exit 1
fi
echo ""
print_success "All images successfully pushed to $REGISTRY/$PROJECT"
echo ""
exit 0

c. Make the script executable:

chmod +x push-images-to-registry.sh

d. Execute the script to load, retag, and push all NAI images.

./push-images-to-registry.sh <registry-url> <project> nai-v2.8.0.tar

Example:

./push-images-to-registry.sh registry.example.com nutanix nai-v2.8.0.tar

Expected Output:

========================================
NAI Images - Load, Retag & Push
========================================
Registry:  registry.example.com
Project:   nutanix
Tar File:  ./nai-v2.8.0.tar
Date:      Tue Aug 18 05:16:33 PM UTC 2026
========================================
Step 1: Loading Images from Tar Bundle
========================================
# Loading images from ./nai-v2.8.0.tar...
# Loaded 41 images
========================================
Step 2: Retagging and Pushing Images
========================================
# [1/41] Processing: nutanix/nai-iam-proxy-control-plane:v2.8.0
# Tagging as: registry.example.com/nutanix/nai-iam-proxy-control-plane:v2.8.0
# Pushing to registry...
The push refers to repository [registry.example.com/nutanix/nai-iam-proxy-control-
plane]
054a97ddb80b: Pushed
5228eaa6af5b: Pushed
256f393e029f: Mounted from nutanix/nai-inference-ui
v2.8.0: digest:
sha256:587189a6559b7af653769a21f4755c66af14fe5135eaf45728c84a2eb2f3c808 size: 951
# Pushed successfully
# [2/41] Processing: nutanix/nai-iam-ui:v2.8.0
# Tagging as: registry.example.com/nutanix/nai-iam-ui:v2.8.0
# Pushing to registry...
The push refers to repository [registry.example.com/nutanix/nai-iam-ui]
7673a750ed47: Pushed
187de06a3fb0: Pushed
[... continues for all images ...]
========================================
Summary
========================================
Total images loaded:    41
Successfully pushed:    41
Failed:                 0
# All images successfully pushed to registry.example.com/nutanix

The image push process typically takes 30-60 minutes depending on your network speed and registry performance.

All images are retagged with your registry URL while preserving the original image path and tag

Original format: nutanix/nai-api:v2.8.0

Retagged format: /nutanix/nai-api:v2.8.0

The script reports failures and continues processing the remaining images.

You can safely re-run the script if it fails partway through.

Configuring Docker Registry Credentials for Dependencies in an Air-gapped Environment with PostgreSQL

Configure registry credentials.

About this task

To configure registry credentials, follow these steps:

Procedure

  1. Set environment Variables.

Export the following environment variables with your private registry credentials:

export REGISTRY=<registry-url-without-https>
export REGISTRY_USERNAME='<registry-username>'
export REGISTRY_PASSWORD='<registry-password>'
export REGISTRY_EMAIL='<registry-email>'
export IMAGE_PULL_SECRET=nai-docker-regcred
export PROJECT=nutanix # set the registry project name
  1. Replace the placeholder values with your actual registry information.

The

must not include the

protocol prefix.

REGISTRY
https://
  1. Create Image Pull Secrets.

Create Kubernetes namespaces and docker-registry secrets for Envoy Gateway System :

kubectl create namespace envoy-gateway-system --dry-run=client -o yaml | kubectl
apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n envoy-gateway-system \
--dry-run=client -o yaml | kubectl apply -f -

4. Create Kubernetes namespaces and

secrets for KServe:

docker-registry
kubectl create namespace kserve --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n kserve \
--dry-run=client -o yaml | kubectl apply -f -

5. Create Kubernetes namespaces and docker-registry secrets for OpenTelemetry:

kubectl create namespace opentelemetry --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry ${IMAGE_PULL_SECRET} \
--docker-server=${REGISTRY} \
--docker-username=${REGISTRY_USERNAME} \
--docker-password=${REGISTRY_PASSWORD} \
--docker-email=${REGISTRY_EMAIL} \
-n opentelemetry \
--dry-run=client -o yaml | kubectl apply -f -

Installing Envoy Gateway in an Air-gapped Environment with PostgreSQL

Install Envoy Gateway.

Before you begin

Ensure that you meet requirements listed in

Prerequisites for Deploying Nutanix Enterprise AI 2.8.0 on a

Nutanix Kubernetes® Platform Cluster in Air-Gapped Environments

on page 52.

About this task

To install Envoy Gateway, follow these steps:

Procedure

1. Install the Envoy Gateway and Gateway API Custom Resource Definitions (CRDs):

helm template eg ./gateway-crds-helm-v1.8.1.tgz \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --server-side --force-conflicts -f -

This command also installs the necessary Gateway API CRDs required for Envoy Gateway.

2. Create the configuration template file

:

eg-config-for-gateway-mode.yaml.template
# This file configures Envoy Gateway for AI Gateway mode with rate limiting
config:
envoyGateway:
gateway:
controllerName: "gateway.envoyproxy.io/gatewayclass-controller"
logging:
level:
default: "info"
provider:
kubernetes:
rateLimitDeployment:
patch:
type: "StrategicMerge"
value:
spec:
template:
spec:
containers:
- imagePullPolicy: "IfNotPresent"
name: "envoy-ratelimit"
env:
- name: REDIS_TYPE
value: "sentinel"
- name: REDIS_PIPELINE_WINDOW
value: "150us"
type: "Kubernetes"
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
maxMessageSize: 11Mi
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- "Translation"
- "Cluster"
- "Route"
service:
fqdn:
hostname: "ai-gateway-controller.nai-system.svc.cluster.local"
port: 1063
rateLimit:
backend:
type: "Redis"
redis:
url: "mymaster,nai-valkey-sentinel.nai-system.svc.cluster.local:26379"
  1. Ensure the REGISTRY environment variable is configured.

4. Generate the actual configuration file using

:

envsubst
envsubst < eg-config-for-gateway-mode.yaml.template > eg-config-for-gateway-mode.yaml

5. Deploy Envoy Gateway:

helm upgrade --install eg ./gateway-helm-v1.8.1.tgz \
-n envoy-gateway-system --create-namespace --wait \
--set global.images.envoyGateway.image=${REGISTRY}/${PROJECT}/nai-gateway:v1.8.1 \
--set global.images.ratelimit.image=${REGISTRY}/${PROJECT}/nai-ratelimit:1e50889b \
--set "global.imagePullSecrets[0].name=${IMAGE_PULL_SECRET}" \
-f ./eg-config-for-gateway-mode.yaml

The configuration file now uses your private registry for the

images through the

ratelimit
${REGISTRY}

variable substitution.

Installing KServe in Air-Gapped Environments with PostgreSQL

Install KServe.

About this task

To install KServe, follow these steps:

Procedure

1. Install the KServe Custom Resource Definitions (CRDs:)

helm upgrade --install kserve-crd ./kserve-crd-v0.19.0.tgz -n kserve --create-
namespace --wait

2. Deploy the KServe controller with RawDeployment mode:

helm upgrade --install kserve ./kserve-resources-v0.19.0.tgz \
-n kserve --wait \
--set kserve.controller.deploymentMode=RawDeployment \
--set kserve.controller.gateway.disableIngressCreation=true \
--set kserve.controller.image=${REGISTRY}/${PROJECT}/nai-kserve-controller \
--set kserve.controller.rbacProxyImage=${REGISTRY}/${PROJECT}/nai-kube-rbac-
proxy:v0.18.0 \
--set "kserve.controller.imagePullSecrets[0].name=${IMAGE_PULL_SECRET}"

Deploying the OpenTelemetry Operator in Air-Gapped Environments with PostgreSQL

Deploy the OpenTelemetry Operator for observability and telemetry collection.

About this task

To deploy the OpenTelemetry Operator, run the following command:

Procedure

Deploy the OpenTelemetry Operator for observability and telemetry collection:

helm upgrade --install opentelemetry-operator ./opentelemetry-operator-0.114.1.tgz \
-n opentelemetry --create-namespace --wait \
--set manager.image.repository=${REGISTRY}/${PROJECT}/nai-opentelemetry-operator \
--set manager.collectorImage.repository=${REGISTRY}/${PROJECT}/nai-opentelemetry-
collector-contrib \
--set "imagePullSecrets[0].name=${IMAGE_PULL_SECRET}"

Deploying Nutanix Enterprise AI Components in Air-Gapped Environments with PostgreSQL

Deploy Nutanix Enterprise AI components in air-gapped environments.

Before you begin

If you are upgrading from NAI 2.7 to 2.8, use the following mapping to migrate your environment variables from the NAI 2.7 deployment

NAI 2.8 Environment Variable

NAI 2.7Environment Variable

Notes

IAM_DB_HOST_URL

POSTGRES_HOST_URL

IAM_DB_NAME

nai_iam

Hardcoded default value

IAM_DB_USERNAME

POSTGRES_USERNAME

NAI 2.8 Environment Variable

NAI 2.7Environment Variable

Notes

IAM_DB_PASSWORD

POSTGRES_PASSWORD

IAM_DB_PORT

Not Applicable

Check with your DB provider, defaults to

if

5432

left empty

IAM_DB_SSLMODE

SSL_MODE

Check with your Postgres DB provider, defaults to

if left empty

disable

IEP_DB_HOST_URL

POSTGRES_HOST_URL

IEP_DB_NAME

CONFIGURED_DB_NAME

IEP_DB_USERNAME

POSTGRES_USERNAME

IEP_DB_PASSWORD

POSTGRES_PASSWORD

IEP_DB_PORT

Not Applicable

Check with your Postgres DB provider, defaults to

if left empty

5432

IEP_DB_SSLMODE

SSL_MODE

Check with your DB provider, defaults to

disable

if left empty

About this task

To deploy Nutanix Enterprise AI components in air-gapped environments, follow these steps:

Procedure

1. Create the

namespace and configure the image pull secret:

nai-system
# Set environment variables (if not already set from Step 3)
export REGISTRY=<registry-url-without-https>
export REGISTRY_USERNAME='<registry-username>'
export REGISTRY_PASSWORD='<registry-password>'
export REGISTRY_EMAIL='<registry-email>'
export IMAGE_PULL_SECRET=nai-docker-regcred
export PROJECT=nutanix # set the registry project name
# Managed Postgres DB Details
# There are 2 DBs used by NAI, you can choose to use single instance or different as
per usecase
# IAM DB details
export IAM_DB_HOST_URL=<hostname>
export IAM_DB_NAME=<DB name>
export IAM_DB_USERNAME=<DB username>
export IAM_DB_PASSWORD=<DB password>
export IAM_DB_PORT=5432 # Update as per your DB port
export IAM_DB_SSLMODE=disable # Supported values: "disable", "require", "verify-ca",
"verify-full"
# Below fields to be used for "verify-full" and "verify-ca" mode
export IAM_SSL_SECRET_NAME=nai-db-certs # This secret to be precreated
export IAM_SSL_ROOT_CERT_NAME=<root cert name>
export IAM_SSL_CLIENT_CERT_NAME="" # client cert name
export IAM_SSL_CLIENT_KEY_NAME="" # client key name
# IEP DB details
export IEP_DB_HOST_URL=<hostname>
export IEP_DB_NAME=<DB name>
export IEP_DB_USERNAME=<DB username>
export IEP_DB_PASSWORD=<DB password>
export IEP_DB_PORT=5432 # Update as per your DB port
export IEP_DB_SSLMODE=disable # Supported values: "disable", "require", "verify-ca",
"verify-full"
# Below fields to be used for "verify-full" and "verify-ca" mode
export IEP_SSL_SECRET_NAME=nai-db-certs # This secret to be precreated
export IEP_SSL_ROOT_CERT_NAME=<root cert name>
export IEP_SSL_CLIENT_CERT_NAME="" # client cert name
export IEP_SSL_CLIENT_KEY_NAME="" # client key name
# Storage class for ReadWriteMany (RWX) volumes - used by NAI API
export NAI_API_RWX_STORAGECLASS=<your-rwx-storage-class>
# Storage class for ReadWriteOnce (RWO) volumes - default storage class
export NAI_DEFAULT_RWO_STORAGECLASS=<your-rwo-storage-class>
# NKP workspace namespace for monitoring (if using Nutanix Kubernetes Platform)
export NKP_WORKSPACE_NAMESPACE=<workspace-namespace>

Storage Class Examples:

For NFS-based storage: Use your NFS storage class name (e.g., nai-nfs-storage)

For Nutanix Volumes: Use nutanix-volumes or your configured storage class

For RWO: Common options include local-path, nutanix-volumes, or your default storage class

Verification: Check available storage classes:

kubectl get storageclass

2. Create a values override file for NAI Operators using the provided template. This file configures all operator

images to use your private registry.

a. Create a file named

with the following content:

darksite-nai-operators.yaml.template
global:
imagePullSecrets:
- name: ${IMAGE_PULL_SECRET}
storage:
storageClassName: ${NAI_DEFAULT_RWO_STORAGECLASS}
naiValkey:
image:
name: ${REGISTRY}/${PROJECT}/nai-valkey
naiJobs:
naiJobsImage:
image: ${REGISTRY}/${PROJECT}/nai-jobs
nai-clickhouse-operator:
operator:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-operator
metrics:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-metrics-exporter
ai-gateway-helm:
extProc:
image:
repository: ${REGISTRY}/${PROJECT}/nai-ai-gateway-extproc
controller:
image:
repository: ${REGISTRY}/${PROJECT}/nai-ai-gateway-controller
naiDatabase:
external: true
image: ${REGISTRY}/${PROJECT}/nai-postgresql:17.10-standard-trixie
clusters:
iep:
host: ${IEP_DB_HOST_URL}
database: ${IEP_DB_NAME}
username: ${IEP_DB_USERNAME}
password: ${IEP_DB_PASSWORD}
port: ${IEP_DB_PORT}
sslMode: ${IEP_DB_SSLMODE}
iam:
host: ${IAM_DB_HOST_URL}
database: ${IAM_DB_NAME}
username: ${IAM_DB_USERNAME}
password: ${IAM_DB_PASSWORD}
port: ${IAM_DB_PORT}
sslMode: ${IAM_DB_SSLMODE}

3. Generate the actual values file

Use

to replace environment variables and create the final values file:

envsubst
envsubst < darksite-nai-operators.yaml.template > darksite-nai-operators.yaml

The envsubst command substitutes

and

with the actual values from

${REGISTRY}
${IMAGE_PULL_SECRET}

your environment variables.

4. Install NAI Operators

helm upgrade --install nai-operators ./nai-operators-2.8.0.tgz \
-n nai-system --create-namespace --wait --timeout 15m -f ./darksite-nai-
operators.yaml

5. Prepare NAI Core Values Override File

Create a values override file for NAI Core using the provided template. This configures all NAI core component images.

a. Create the template file named

with the following content:

darksite-nai-core.yaml.template
global:
imagePullSecrets:
- name: ${IMAGE_PULL_SECRET}
storage:
storageClassName: ${NAI_DEFAULT_RWO_STORAGECLASS}
storageClassNameRWX: ${NAI_API_RWX_STORAGECLASS}
gateway:
envoyDeployment:
container:
image: ${REGISTRY}/${PROJECT}/nai-envoy:distroless-v1.38.0
naiIepOperator:
iepOperatorImage:
image: ${REGISTRY}/${PROJECT}/nai-iep-operator
modelProcessorImage:
image: ${REGISTRY}/${PROJECT}/nai-python-processor
dataSourceProcessorImage:
image: ${REGISTRY}/${PROJECT}/nai-python-processor
batchInferenceProcessor:
containers:
processor:
image: ${REGISTRY}/${PROJECT}/nai-go-processor
statusProvider:
image: ${REGISTRY}/${PROJECT}/nai-go-processor
finetuneProcessor:
containers:
processor:
image: ${REGISTRY}/${PROJECT}/nai-finetuning
statusProvider:
image: ${REGISTRY}/${PROJECT}/nai-go-processor
naiInferenceUi:
naiUiImage:
image: ${REGISTRY}/${PROJECT}/nai-inference-ui
naiJobs:
naiJobsImage:
image: ${REGISTRY}/${PROJECT}/nai-jobs
naiApi:
naiApiImage:
image: ${REGISTRY}/${PROJECT}/nai-api
supportedTGIImage: ${REGISTRY}/${PROJECT}/nai-tgi
supportedKserveRuntimeImage: ${REGISTRY}/${PROJECT}/nai-kserve-huggingfaceserver
eppImage: ${REGISTRY}/${PROJECT}/nai-epp-inference-scheduler
supportedVLLMImage: ${REGISTRY}/${PROJECT}/nai-vllm
supportedKserveCustomModelServerRuntimeImage: ${REGISTRY}/${PROJECT}/nai-kserve-
custom-model-server
naiDatabase:
external: true
clientImage: ${REGISTRY}/${PROJECT}/nai-postgresql:17.10-standard-trixie
clusters:
iep:
host: ${IEP_DB_HOST_URL}
database: ${IEP_DB_NAME}
port: ${IEP_DB_PORT}
sslMode: ${IEP_DB_SSLMODE}
sslSecretName: ${IEP_SSL_SECRET_NAME}
sslRootCertName: {IEP_SSL_ROOT_CERT_NAME}
sslClientCertName: {IEP_SSL_CLIENT_CERT_NAME}
sslClientKeyName: {IEP_SSL_CLIENT_KEY_NAME}
iam:
host: ${IAM_DB_HOST_URL}
database: ${IAM_DB_NAME}
port: ${IAM_DB_PORT}
sslMode: ${IAM_DB_SSLMODE}
sslSecretName: ${IAM_SSL_SECRET_NAME}
sslRootCertName: {IAM_SSL_ROOT_CERT_NAME}
sslClientCertName: {IAM_SSL_CLIENT_CERT_NAME}
sslClientKeyName: {IAM_SSL_CLIENT_KEY_NAME}
naiIam:
iamProxy:
image: ${REGISTRY}/${PROJECT}/nai-iam-proxy
iamProxyControlPlane:
image: ${REGISTRY}/${PROJECT}/nai-iam-proxy-control-plane
iamUi:
image: ${REGISTRY}/${PROJECT}/nai-iam-ui
iamUserAuthn:
image: ${REGISTRY}/${PROJECT}/nai-iam-user-authn
iamThemis:
image: ${REGISTRY}/${PROJECT}/nai-iam-themis
iamThemisBootstrap:
image: ${REGISTRY}/${PROJECT}/nai-iam-bootstrap
naiAgent:
agentImage:
image: ${REGISTRY}/${PROJECT}/nai-agent-app
naiLabs:
labsImage:
image: ${REGISTRY}/${PROJECT}/nai-rag-app
nai-clickhouse-keeper:
clickhouseKeeper:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-keeper
oauth2-proxy:
image:
repository: ${REGISTRY}/${PROJECT}/nai-oauth2-proxy
nai-clickhouse-server:
clickhouse:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-server
initContainers:
addUdf:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-udf
waitForKeeper:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-jobs
nai-clickhouse-schemas:
image:
registry: ${REGISTRY}
repository: ${PROJECT}/nai-clickhouse-schemas
naiMonitoring:
opentelemetry:
collectorImage: ${REGISTRY}/${PROJECT}/nai-opentelemetry-collector-
contrib:0.152.0
targetAllocator:
image:
repository: ${REGISTRY}/${PROJECT}/nai-target-allocator
nodeExporter:
serviceMonitor:
namespaceSelector:
matchNames:
- prometheus
- kommander
- kommander-default-workspace
- ${NKP_WORKSPACE_NAMESPACE}
dcgmExporter:
serviceMonitor:
namespaceSelector:
matchNames:
- prometheus
- kommander
- kommander-default-workspace
- ${NKP_WORKSPACE_NAMESPACE}

b. Ensure that REGISTRY, IMAGE_PULL_SECRET, NAI_API_RWX_STORAGECLASS,

NAI_DEFAULT_RWO_STORAGECLASS, and NKP_WORKSPACE_NAMESPACE environment variables are set before running the next command.

c. Generate the actual values file

Use

to replace environment variables and create the final values file:

envsubst
envsubst < darksite-nai-core.yaml.template > darksite-nai-core.yaml

6. Install NAI Core

helm upgrade --install nai-core ./nai-core-2.8.0.tgz -n nai-system --create-namespace
--wait --timeout 15m \
-f ./darksite-nai-core.yaml

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

Enable the Chat and Talk to My Data application. These applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

Enable HTTPS with a self-signed certificate. By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

using your own certificate or a cert-manager ClusterIssuer, see

TLS Encryption on Nutanix

Enterprise AI

.

Scale out the ingress gateway and NAI API. To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Configure PostgreSQL database connections. To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default value is 1000. Increase this value if you expect a higher number of concurrent clients.

  1. Configure the TLS certificate.

For more information, see

TLS Encryption on Nutanix Enterprise AI

on page 119.

Troubleshooting Deployment of Nutanix Enterprise AI 2.8.0 in Air-Gapped Environments with PostgreSQL

The following troubleshooting tips can help you resolve a few common issues that can occur when you are installing

Image Pull Errors

Verify that:

Registry credentials are correct

Image pull secrets exist in the correct namespaces

Registry URL is accessible from the cluster

Image paths match your registry structure

All required images are present in your private registry

CRD Installation Failures

Ensure you have sufficient permissions to create cluster-scoped resources.

Helm Installation Timeouts

Increase the timeout value using –timeout flag (e.g., –timeout 20m).

Storage Class Issues

Verify the storage class exists: kubectl get storageclass

Ensure the storage class supports the required access mode (RWX for NAI API, RWO for others)

Check PVC status: kubectl get pvc -n nai-system

Pod Startup Failures

Check pod logs: kubectl logs -n nai-system

Describe pod for events: kubectl describe pod -n nai-system

Verify resource limits if running on resource-constrained clusters

Dependency Issues

Ensure all dependencies (Envoy Gateway, KServe, OpenTelemetry) are installed and running before installing NAI Core.

Deploying Nutanix Enterprise AI for a Connected NKP Cluster Install or upgrade Nutanix Enterprise AI for a connected NKP cluster.

Before you begin

Installing Prerequisite Components

Install components on the Kubernetes cluster. For more information, see

on a Nutanix Kubernetes Platform Cluster

on page 43.

If you are upgrading from NAI 2.7 to 2.8, use the following mapping to migrate your environment variables from the NAI 2.7 deployment

NAI 2.8 Environment Variable

NAI 2.7Environment Variable

Notes

IAM_DB_HOST_URL

POSTGRES_HOST_URL

IAM_DB_NAME

nai_iam

Hardcoded default value

IAM_DB_USERNAME

POSTGRES_USERNAME

IAM_DB_PASSWORD

POSTGRES_PASSWORD

IAM_DB_PORT

Not Applicable

Check with your DB provider, defaults to

if

5432

left empty

IAM_DB_SSLMODE

SSL_MODE

Check with your Postgres DB provider, defaults to

if left empty

disable

IEP_DB_HOST_URL

POSTGRES_HOST_URL

IEP_DB_NAME

CONFIGURED_DB_NAME

IEP_DB_USERNAME

POSTGRES_USERNAME

IEP_DB_PASSWORD

POSTGRES_PASSWORD

IEP_DB_PORT

Not Applicable

Check with your Postgres DB provider, defaults to

if left empty

5432

IEP_DB_SSLMODE

SSL_MODE

Check with your DB provider, defaults to

disable

if left empty

About this task

To install or upgrade Nutanix Enterprise AI for a connected NKP cluster, follow these steps:

Procedure

  1. Set up the Nutanix Enterprise AI Helm repository. a. Add and update the Nutanix Helm repository, which contains the

Helm chart:

nai-core
helm repo add ntnx-charts https://nutanix.github.io/helm-releases && helm repo
update ntnx-charts

b. Search for the version of the

and

Helm chart available for installation in the

nai-operators
nai-core

Nutanix Helm repository:

helm search repo ntnx-charts/nai-operators --versions
helm search repo ntnx-charts/nai-core --versions

2. Install or upgrade NAI Core on NKP:

export REGISTRY_SECRET_NAME=nai-regcred
export DOCKER_SERVER=https://index.docker.io/v1/
export DOCKER_USERNAME=<docker-username>
export DOCKER_PASSWORD=<docker-password>
export DOCKER_EMAIL=<docker-email>
# Managed Postgres DB Details
# There are 2 DBs used by NAI, you can choose to use single instance or different as
per usecase
# IAM DB details
export IAM_DB_HOST_URL=<hostname>
export IAM_DB_NAME=<DB name>
export IAM_DB_USERNAME=<DB username>
export IAM_DB_PASSWORD=<DB password>
export IAM_DB_PORT=5432 # Update as per your DB port
export IAM_DB_SSLMODE=disable # Supported values: "disable", "require", "verify-ca",
"verify-full"
# Below fields to be used for "verify-full" and "verify-ca" mode
export IAM_SSL_SECRET_NAME=nai-db-certs # This secret to be precreated
export IAM_SSL_ROOT_CERT_NAME=<root cert name>
export IAM_SSL_CLIENT_CERT_NAME="" # client cert name
export IAM_SSL_CLIENT_KEY_NAME="" # client key name
# IEP DB details
export IEP_DB_HOST_URL=<hostname>
export IEP_DB_NAME=<DB name>
export IEP_DB_USERNAME=<DB username>
export IEP_DB_PASSWORD=<DB password>
export IEP_DB_PORT=5432 # Update as per your DB port
export IEP_DB_SSLMODE=disable # Supported values: "disable", "require", "verify-ca",
"verify-full"
# Below fields to be used for "verify-full" and "verify-ca" mode
export IEP_SSL_SECRET_NAME=nai-db-certs # This secret to be precreated
export IEP_SSL_ROOT_CERT_NAME=<root cert name>
export IEP_SSL_CLIENT_CERT_NAME="" # client cert name
export IEP_SSL_CLIENT_KEY_NAME="" # client key name
# Storage class for ReadWriteMany (RWX) volumes - used by NAI API
export NAI_API_RWX_STORAGECLASS=<your-rwx-storage-class>
# Storage class for ReadWriteOnce (RWO) volumes - default storage class
export NAI_DEFAULT_RWO_STORAGECLASS=<your-rwo-storage-class>
# NKP workspace namespace for monitoring (if using Nutanix Kubernetes Platform)
export NKP_WORKSPACE_NAMESPACE=<workspace-namespace>

[!NOTE] Note: Ensure that the configured PostgreSQL user credentials possess database creation privileges on the target PostgreSQL host.

[!NOTE] Note: You can append optional Helm overrides to the nai-core installation command to customize the deployment

Enable the Chat and Talk to My Data application. These applications are disabled by default. To enable it, add the following flag to the

Helm install command:

nai-core
--set "naiLabs.enabled=true"

Enable HTTPS with a self-signed certificate. By default, NAI does not provision a TLS certificate for the ingress gateway. To quickly enable HTTPS with a self-signed certificate (recommended for dev/test environments), add the following flag to the

Helm install command:

nai-core
--set "gateway.certManager.selfSigned=true"

This requires

to be installed on the cluster. For production TLS options, including

cert-manager

using your own certificate or a cert-manager ClusterIssuer, see

TLS Encryption on Nutanix

Enterprise AI

.

Scale out the ingress gateway and NAI API. To increase replicas for the ingress gateway and the NAI API, add the following flags to the

Helm install command:

nai-core
--set "gateway.replicaCount=<Number_of_replicas>"
--set "naiApi.replicaCount=<Number_of_replicas>"

The default is 1 replica each. Increase based on your scale requirements.

Configure PostgreSQL database connections. To adjust the maximum number of concurrent PostgreSQL connections, add the following flag to the

Helm install command:

nai-core
--set "naiDatabase.postgresConfig.maxConnections=<Number_of_Connections>"

The default value is 1000. Increase this value if you expect a higher number of concurrent clients.

3. Create Docker Registry Secrets

Create the

and the docker-registry secret in both

and

nai-system namespace
nai-system
envoy-gateway-

namespaces. The

namespace is already present on the cluster.

system
envoy-gateway-system
kubectl create namespace nai-system --dry-run=client -o yaml | kubectl apply -f -
kubectl -n nai-system create secret docker-registry ${REGISTRY_SECRET_NAME} \
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n envoy-gateway-system create secret docker-registry ${REGISTRY_SECRET_NAME}
\
--docker-server=${DOCKER_SERVER} \
--docker-username=${DOCKER_USERNAME} \
--docker-password=${DOCKER_PASSWORD} \
--docker-email=${DOCKER_EMAIL} \
--dry-run=client -o yaml | kubectl apply -f -

4. Create a values override file

for NAI Operators using the

nai-operators-external-db.yaml.template

provided template.

global:
imagePullSecrets:
- name: ${IMAGE_PULL_SECRET}
storage:
storageClassName: ${NAI_DEFAULT_RWO_STORAGECLASS}
naiDatabase:
external: true
clusters:
iep:
host: ${IEP_DB_HOST_URL}
database: ${IEP_DB_NAME}
username: ${IEP_DB_USERNAME}
password: ${IEP_DB_PASSWORD}
port: ${IEP_DB_PORT}
sslMode: ${IEP_DB_SSLMODE}
iam:
host: ${IAM_DB_HOST_URL}
database: ${IAM_DB_NAME}
username: ${IAM_DB_USERNAME}
password: ${IAM_DB_PASSWORD}
port: ${IAM_DB_PORT}
sslMode: ${IAM_DB_SSLMODE}

Generate the actual values file

envsubst < nai-operators-external-db.yaml.template > nai-operators-external-db.yaml

5. Install NAI Operators:

helm upgrade --install nai-operators ntnx-charts/nai-operators --version 2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
-f ./nai-operators-external-db.yaml

6. Create a values override file(nai-core-external-db.yaml.template) for NAI Operators using the provided

template.

global:
imagePullSecrets:
- name: ${IMAGE_PULL_SECRET}
storage:
storageClassName: ${NAI_DEFAULT_RWO_STORAGECLASS}
storageClassNameRWX: ${NAI_API_RWX_STORAGECLASS}
naiDatabase:
external: true
clusters:
iep:
host: ${IEP_DB_HOST_URL}
database: ${IEP_DB_NAME}
port: ${IEP_DB_PORT}
sslMode: ${IEP_DB_SSLMODE}
sslSecretName: ${IEP_SSL_SECRET_NAME}
sslRootCertName: {IEP_SSL_ROOT_CERT_NAME}
sslClientCertName: {IEP_SSL_CLIENT_CERT_NAME}
sslClientKeyName: {IEP_SSL_CLIENT_KEY_NAME}
iam:
host: ${IAM_DB_HOST_URL}
database: ${IAM_DB_NAME}
port: ${IAM_DB_PORT}
sslMode: ${IAM_DB_SSLMODE}
sslSecretName: ${IAM_SSL_SECRET_NAME}
sslRootCertName: {IAM_SSL_ROOT_CERT_NAME}
sslClientCertName: {IAM_SSL_CLIENT_CERT_NAME}
sslClientKeyName: {IAM_SSL_CLIENT_KEY_NAME}
naiMonitoring:
nodeExporter:
serviceMonitor:
namespaceSelector:
matchNames:
- ${NKP_WORKSPACE_NAMESPACE}
dcgmExporter:
serviceMonitor:
namespaceSelector:
matchNames:
- ${NKP_WORKSPACE_NAMESPACE}

Generate the actual values file

envsubst < nai-core-external-db.yaml.template > nai-core-external-db.yaml

7. Install NAI Core

helm upgrade --install nai-core ntnx-charts/nai-core --version=2.8.0 \
-n nai-system --create-namespace --wait --timeout 15m \
-f ./nai-core-external-db.yaml
  1. Configure the TLS certificate.

TLS Encryption on Nutanix Enterprise AI

For more information, see

on page 119.

What to do next

Ensure that all the pods are successfully deployed in the

namespace and are in the

state by

nai-system
Ready

running:

kubectl get pods -n nai-system

The following is a sample output of the command. Match the pod names by ignoring the auto-generated suffix added by Kubernetes and their status.

[!NOTE] Note:

The number of

should be equal to the number of worker

nai-otel-collector-collector pods

nodes in the NAI Kubernetes cluster.

NAME                                                            READY   STATUS
RESTARTS   AGE
chi-nai-clickhouse-server-chcluster1-0-0-0                      1/1     Running     0
16h
chk-nai-clickhouse-keeper-chkeeper-0-0-0                        1/1     Running     0
16h
iam-database-bootstrap-b8etj-hk9rz                              0/1     Completed   0
16h
iam-proxy-68f9459885-zcwgs                                      1/1     Running     0
16h
iam-proxy-control-plane-6897669d64-rvglt                        1/1     Running     0
16h
iam-themis-749b7b56f8-pmclb                                     1/1     Running     0
16h
iam-themis-bootstrap-qgczx-p7bl6                                0/1     Completed   0
16h
iam-ui-6697d94478-fftl5                                         1/1     Running     0
16h
iam-user-authn-5b4dcfdfb7-jhllq                                 1/1     Running     0
16h
nai-api-784fb7b99-8tch9                                         1/1     Running     0
16h
nai-api-db-migrate-mgnba-cl7gw                                  0/1     Completed   0
16h
nai-clickhouse-schema-job-1771350985-bd4sq                      0/1     Completed   0
16h
nai-db-0                                                        1/1     Running     0
16h
nai-iep-model-controller-5cd8bcd5f-9h9cl                        1/1     Running     0
16h
nai-labs-86589cc95d-gk87q                                       1/1     Running     0
16h
nai-oauth2-proxy-5746ccc8b7-65jmf                               1/1     Running     0
16h
nai-oidc-client-registration-lvrrv-pbdsn                        0/1     Completed   0
16h
nai-otel-collector-collector-4pxtl                              1/1     Running     0
16h
nai-otel-collector-collector-8p4w8                              1/1     Running     0
16h
nai-otel-collector-collector-bwstg                              1/1     Running     0
16h
nai-otel-collector-collector-drddv                              1/1     Running     0
16h
nai-otel-collector-collector-k457n                              1/1     Running     0
16h

Table 43: Gateway TLS Helm values

Property 1ValueTypeDefaultDescriptionProperty 6
gateway.tlsSecretNamestring“ingress-Name of the Kubernetes TLS secret referenced by the Gateway HTTPS listener.
certificate“
gateway.certManager.selfSignedboolfalseWhen true, creates a self-signed Issuer and Certificate for dev/test use.
gateway.certManager.issuerRefobjectnullReference to an existing cert-manager Issuer or ClusterIssuer. Must include name and kind; optionally group.
gateway.certManager.issuerRef.namestring-Name of the Issuer or ClusterIssuer.
gateway.certManager.issuerRef.kindstring-Issuer or ClusterIssuer.
gateway.certManager.issuerRef.groupstringcert-manager.ioAPI group of the issuer (usually left as default).
gateway.certManager.dnsNameslist[]DNS Subject Alternative Names for the certificate. Required when
issuerRefis set. The
first entry is used as
commonName.
Accessing the NAI Dashboard IP Address
Retrieve the IP address for the NAI dashboard.

Procedure

1. Get the NAI dashboard IP address:

kubectl get svc -n envoy-gateway-system -l "gateway.envoyproxy.io/owning-gateway-
name=nai-ingress-gateway,gateway.envoyproxy.io/owning-gateway-namespace=nai-system" -
o jsonpath='{.items[0].status.loadBalancer.ingress[0].ip}'

The expected output is the IP address.

  1. Copy the IP address and paste IP address in your browser.

The NAI login page is displayed.

  1. Log in to NAI.

Log in to Nutanix Enterprise AI

For information on viewing the dashboard, see

on page 167.

Rotating Docker Registry Credentials

Rotate the Docker registry credentials for Nutanix Enterprise AI to maintain secure access to the registry. Rotate the credentials when the existing credentials expire or when you suspect they have been compromised. Regular credential rotation replaces old authentication secrets with new ones, reducing the risk of unauthorized access.

About this task

When registry credentials expire or are compromised, you must update the secret in all namespaces where NAI components run.

To rotate Docker registry credentials, follow these steps:

Procedure

1. Rotate credentials:

export EXISTING_SECRET_NAME=<docker-registry-secret-name> # Use the same registry
secret name specified during the NAI installation or upgrade.
export REGISTRY=<container-registry> # For Docker: https://index.docker.io/v1/
kubectl -n nai-system create secret docker-registry ${EXISTING_SECRET_NAME} \
--docker-server=${REGISTRY} \
--docker-username=<username> \
--docker-password=<password> \
--docker-email=<email> \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n nai-admin create secret docker-registry ${EXISTING_SECRET_NAME} \
--docker-server=${REGISTRY} \
--docker-username=<username> \
--docker-password=<password> \
--docker-email=<email> \
--dry-run=client -o yaml | kubectl apply -f -
kubectl -n envoy-system-system create secret docker-registry ${EXISTING_SECRET_NAME}
\
--docker-server=${REGISTRY} \
--docker-username=<username> \
--docker-password=<password> \
--docker-email=<email> \
--dry-run=client -o yaml | kubectl apply -f -

2. Validate the rotation. To validate, follow these steps:

a. Verify that all pods are active in the following namespaces:

nai-system
nai-admin
envoy-gateway-system

b. Validate the new models. c. Validate the endpoints.

Endpoints

Status

Failed

On the

page, if the

column displays

for an endpoint, with the message Unable to

pull runtime image with provided credentials, follow these steps:


Last updated Oct 08, 2026