VMware Private AI Services Release Notes
This document contains the following sections
Introduction
VMware Private AI Services provides features and capabilities to assist you with the deployment of AI workloads.
VMware Private AI Services 2.1 | 20 MAR 2026, 12 MAY 2026 VMware Private AI Services 2.0.89 | 11 SEP 2025 Check for additions and updates to these release notes. |
Overview of Private AI Services
- Model Runtime and API GatewayYour IT organization can now run models as a service on our model runtime, which provides initial support for vLLM, Infinity, and llama.cpp to run your inference and embedding models. These models run behind an ML API Gateway, so your users or AI applications can interact directly with the models using API, while your IT team can transparently update model versions and horizontally scale up and down based on demand, without impacting the end user's workloads.
- Data Indexing and RetrievalYou can connect unstructured data in Word documents, PDF files, and other formats, to a vector database. You can have this data automatically parsed, embeddings added to the vector database, and data refreshed on a regular cadence.
- MCP Tool Gallery.Available in Private AI Services 2.1 and later.You can extend the capabilities of your agents by using MCP tools available on one or more MCP servers. Data Indexing and Retrieval is integrated in Private AI Services as an MCP tool, allowing agents to decide whether to retrieve content from a knowledge base and what search term to use.
- Agent BuilderApplication developers can easily compose a model with specific prompting and instructions, and knowledge bases created and maintained by Data Indexing and Retrieval. In a user-friendly UI, developers see what is available to them to compose and can then create an agent. They can easily iterate with that agent within an in-app playground where they can see results and experiment with different models and prompt instructions to see what works best for their use case. They can save their agent and use it as the back end for their AI-powered applications.
- Observability.Available in Private AI Services 2.1 and later.You can perform monitoring and diagnostics of models, used in agents, to track the health, quality, and behavior of AI applications in observability platforms like Grafana.
Private AI Services Releases
Private AI Services is installed as a package, separately from the VMware Private AI Foundation with NVIDIA core functionality. A Private AI Services version is available at the release date of VMware Private AI Foundation with NVIDIA. After VMware Private AI Foundation with NVIDIA is released, you can install or update to newer versions of Private AI Services without the need to update VMware Private AI Foundation with NVIDIA.
Installation and Upgrade
You can download the YAML file for installing the Private AI Services on a GPU-accelerated Supervisor instance from the
Private AI Services
download area of the Broadcom Support Portal.For instructions on installing the Private AI Services service and making it availability in VCF Automation organizations, see Install Private AI Services on the Supervisor of the GPU-Accelerated Workload Domain and Activate Private AI Services in an Organization Namespace.
For information on updating the Private AI Services service in the Supervisor instance assigned to the VCF Automation organization, see Upgrade a Supervisor Service to a Newer Version.
License Information
Private AI Services releases are available under a VMware Private AI Foundation with NVIDIA license. See VMware Private AI Foundation with NVIDIA Guide.
Private AI Services 2.1
Release Information
Release Element | Description |
|---|---|
Release date |
|
Compatible VCF version |
|
What's New
Available for VCF 9.1. Updated Model Endpoint UI
End users can now use an extended set of configuration settings when running model endpoints in the VCF Automation UI.
Available for VCF 9.1. UI-Driven Self-Service Enablement
In VCF 9.1, organization administrators can enable and manage Private AI Services within namespaces by using the VCF Automation UI. End users have a seamless and intuitive UI experience that supports the end-to-end configuration, deployment, and lifecycle management of Private AI Services.
Support for disconnected (air-gapped) environments
Private AI Services 2.1 introduces the artifact mirroring tool (AMT), enabling VI administrators to install and run full private AI capabilities, including NVIDIA GPU-powered model endpoints and agents, within secure, air-gapped environments.
Integrated observability and analytics
You can use the observability framework of Private AI Services for health and performance insights across the entire AI workload footprint, from inference engines and GPU utilization to knowledge base indexing and agents.
Integrated LLM tracing
You can use OpenTelemetry Collector platforms to trace all interactions between end-users and LLMs, agents, and knowledge bases.
Model Context Protocol (MCP) integration
By using MCP, MLOps engineers can connect agents to a wide array of external data sources and tools by using an industry-standard interface. This version of Private AI Services introduces a tool gallery for centralized MCP server management and tool enablement, and exposes knowledge bases over MCP so that AI application developers can build context-aware agents.
Support for CPU-based inference and embeddings
You can use llama.cpp as an inference engine for running GPU and CPU-based LLMs.
Data indexing and retrieval for Google native formats
You can upload Google Docs, Google Sheets, and Google Slides in knowledge bases.
Supported VKS Versions
This version of Private AI Services uses VMware vSphere Kubernetes release (VKr) 1.33 clusters with builtin-generic-v3.2.0 ClusterClass supported with VMware vSphere Kubernetes Service (VKS) 3.5.0 and later.
Verify that your VKr content library contains images for Ubuntu 24.04.
Consider upgrading to VKS 3.5.0 and later. Otherwise, the VKS clusters provisioned by Private AI Services for model endpoints will remain unreconciled.
Supported NVIDIA GPU Operator Versions
By default, this version of Private AI Services uses NVIDIA GPU Operator 25.10.1. For supported driver versions, see the NVIDIA GPU Operator documentation.
If you are upgrading Private AI Services in your VCF environment, you might need to evaluate upgrading the NVIDIA software for ESX host to ensure that the right vGPU drivers are installed on the ESX host and the guest VM. By default, NVIDIA GPU Operator 25.10.1 uses the GPU driver v580.x series on the guest OS side which might be incompatible with earlier NVIDIA software on the ESX host. For compatibility information, see the NVIDIA vGPU software driver documentation. If you provision new model endpoints in Private AI Services 2.1.0, the Model Runtime will report driver version mismatches.
Supported PostgreSQL Versions
This version of Private AI Services supports PostgreSQL 16.8 with pgvector extension 0.8.0.
Supported Inference Engine Versions
This version of Private AI Services is delivered with the following inference engines:
- vLLM v0.11.2 (completions and embeddings)
- llama.cpp ver. b7739 (completions and embeddings)
- Infinity 0.0.76 (embeddings only)
You can also override the version of the inference engine by specifying a custom engineImage in the model endpoint YAML specification. See Deploy Completion or Embedding Model Endpoints.
Supported File Types and Data Sources
The Data Retrieval and Indexing component of Private AI Services supports retrieving files from specific types of data sources and extracting text content from the specific file types.
Data Retrieval and Indexing Element | Supported Options | |
|---|---|---|
File types |
| |
Data source types | Microsoft SharePoint |
|
Google Drive | Google Drive API v3, using Google service accounts for authentication. | |
Atlassian Confluence |
| |
S3 Compatible Server | S3 API v4, using Access Key ID/Secret Key credentials for authentication | |
Caveats and Limitations
You can run up to 15 model endpoint replicas per namespace
Each model endpoint replica consumes a /24 CIDR block in your network address space, and the default network assignments for the VKS cluster used by Private AI Services is /20. As a result, you will run out of network address space after 15 model endpoints. To run more endpoints, set the
vks.candidatePodCIDRs
property in the configuration of the Private AI Services Supervisor Service to allocate a larger network range.- Log in to the vCenter instance for the management domain athttps://<vcenter_fqdn>/uias .
- In the vSphere Client side panel, clickSupervisor Managementand click theServicestab.
- In the Private AI Services card, select Actions > Manage Service.
- In the manage wizard, select the Supervisor instance and click Next.
- On theReviewpage, in theYAML Service Configtext area, add the following YAML code and clickFinish.vks: candidatePodCIDRs: - <CIDR-1> - <CIDR-2> - <CIDR-N>
Additional configuration is required for agents with models that do not support tool calling.
To use a RAG agent with data indexing and retrieval and with a model that does not support MCP tool calling or an agent for non-chat completions (legacy behavior), set the
x-pais-force-static-tool-execution
metadata header to true
in the agent configuration. See Pre-Load Tool Results in the Context of an Agent.In a disconnected environment, the NVIDIA GPU Operator can be deployed only from an OCI registry without authentication.
In the VKr version in VCF 9.0.2, the bundled kapp-controller v0.58.1 contains a known upstream issue where Helm v3.18.0 fails to pull charts from OCI registries that require authentication.
Resolved Issues
Private AI Services internally provisions a VKS Cluster that uses IP addresses within CIDR blocks 192.168.144.0/20 and 10.96.0.0/20. If those ranges are used by other elements of your virtual infrastructure, for example, Load Balancers, Private AI Services fails to start
Private AI Services internally uses mutual TLS (mTLS) between the microservices running in a user’s namespace. These certificates expire 3 months after the initial installation.
Known Issues
In Private AI Services version 2.1 with GPU Operator v25.10.1, ModelRuntime pods using GPUs might fail to start.
Kubernetes pods using NVIDIA GPUs fail to start with the following error that appears in the pod events.
failed to create containerd container: CDI device injection failed: unresolvable CDI devices
Customize the GPU Operator Helm chart values to deactive the Container Device Interface (CDI) by performing the following steps.
- Get an API token for access to the namespace in VCF Automation where Private AI Services is available.
- Click your account in the top right corner and selectUser Settings > My Account.
- On theAPI Tokenstab, create a token and save it on your machine.
- On a machine running kubectl and the VCF Consumption CLI, create a ConfigMap with the following YAML code in a file, for example, calledhelm-values.yaml.See Installing and Using VCF CLI in the VCF documentation.apiVersion: v1 kind: ConfigMap metadata: name: helm-values data: values.yaml: | cdi: enabled: false toolkit: version: "v1.18.2" env: - name: NVIDIA_CONTAINER_RUNTIME_MODE value: "legacy" - name: CDI_ENABLED value: "false"
- Create a context for the VCF Automation namespace where the Private AI Services instance is available.vcf context create<vcf_automation_context_name>--endpoint<vcf_automation_fqdn_or_ip>--ca-certificate<vcf_automation_certificate>.pem--api-token<vcf_automation_api_token>--type cci
- Log in to the VCF Automation namespace where Private AI Services is running by using the VCF Consumption CLI.vcf context use<vcf_automation_context_name>:<pais_namespace_name>:<project_name>
- Save the ConfigMap in the namespace by running the following command.kubectl apply -f helm-values.yaml
- Prepare YAML code to modify thePAISConfigurationCR in the namespace to reference the customized Helm chart, for example, in apaisconfiguration-update.yamlfile.spec: nvidiaConfig: gpuOperatorOverridesRef: name: helm-values
- Apply the update to the PAISConfiguration CR.kubectl patch PAISConfiguration default --patch-file paisconfiguration-update.yaml --type='merge'
After upgrading to Private AI Services 2.1, a running model endpoint might fail to re-deploy because of insufficient memory.
Because of increased VRAM requirements in the vLLM engine version in Private AI Services 2.1, some model endpoints might fail to reconcile with
RuntimeError: CUDA out of memory occurred with warming up sample with <number> dummy requests
.- Deploy a new model endpoint with the same routing name but with CLI engine arguments in thespec.inferenceServerCustomizationsection of the model endpoint YAML code that are adjusted to accommodate the new memory overhead. For information about the CLI arguments, see the vLLM documentation.
- Decommission the original model endpoint.
LLM traces from Private AI Services are not visible in the OpenTelemtry Collector
If the configuration of the target OpenTelemetry Collector in the
spec.observability.llmTraces
property of the Private AI Services activation YAML file is incorrect, an error is not reported. Distributed tracing uses only OpenTelemetry traceparent
propagation solely between Private AI Services components.You can only verify that the configuration is successful by seeing
instrument_code_for_tracing
spans in the OpenTelemetry Collector.If the
instrument_code_for_tracing
spans are not visible in the OpenTelemetry Collector, check the logs after the setup_tracing
lines for the Private AI Services activation by running the following command in the context of the VCF Automation namespace.kubectl logs -l "pais.vmware.com/component in (api,rex-worker,rex-mcp-server)" -c main --prefix --tail=-1 | grep setup_tracing -A 5
Expected downtime during upgrade from version 2.0.x to 2.1 of Private AI Services
The upgrade process deletes and re-creates the VKS cluster hosting model endpoints causing downtime while nodes are being re-created and models downloaded again.
Workaround: None
Private AI Services 2.0.89
Release Information
Release Element | Description |
|---|---|
Release date | 11 SEP 2025 |
Compatible VCF version | 9.0 |
What's New
Improved Supervisor Service Security and Compatibility
The Private AI Services deliverables are now signed with Broadcom certificates.
Generating Support Bundles for Private AI Services
You can gather diagnostic information for Broadcom technical support about Private AI Services operation. For information on how to generate and export a support bundle, see Broadcom knowledge base article 408731.
Supported VKS Versions
This version of Private AI Services uses VMware vSphere Kubernetes release (VKr) 1.32 clusters with VMware vSphere Kubernetes Service (VKS) ClusterClass builtin-generic-v3.2.0.
Supported GPU Operator Versions
By default, this version of Private AI Services uses NVIDIA GPU Operator 24.9.0. For supported driver versions, see the NVIDIA GPU Operator documentation.
Supported PostgreSQL Versions
This version of Private AI Services supports PostgreSQL 16.8 with pgvector extension 0.8.0.
Supported Inference Server Versions
This version of Private AI Services supports the vLLM inference server 0.6.5 for generating completions and the Infinity inference server 0.0.43 for generating embeddings. Both servers are installed with supported default configurations.
Supported File Types and Data Sources
The Data Retrieval and Indexing component of Private AI Services supports retrieving files from specific types of data sources and extracting text content from the specific file types.
Data Retrieval and Indexing Element | Supported Options | |
|---|---|---|
File types |
| |
Data source types | Microsoft SharePoint |
|
Google Drive | Google Drive API v3, using Google service accounts for authentication. | |
Atlassian Confluence |
| |
S3 Compatible Server | S3 API v4, using Access Key ID/Secret Key credentials for authentication | |
Known Issues
Expected downtime during upgrade from version 2.0.51 to 2.0.89 of Private AI Services
Model endpoints with single replicas, which is the default configuration, cannot maintain availability during cluster node upgrades, which replace nodes one at a time so that you do not have to over-provision GPUs.
Workaround: None
Private AI Services internally provisions a VKS Cluster that uses IP addresses within CIDR blocks 192.168.144.0/20 and 10.96.0.0/20. If those ranges are used by other elements of your virtual infrastructure, for example, Load Balancers, Private AI Services fails to start
Workaround: None.
Private AI Services internally uses mutual TLS (mTLS) between the microservices running in a user’s namespace. These certificates expire 3 months after the initial installation.
Workaround: Delete your PAISConfiguration and ModelEndpoints, and re-apply them but keep your PostgreSQL database so that you do not lose your data. This operation introduces a short outage.
Private AI Services 2.0.51
Release Information
Release Element | Description |
|---|---|
Release date | 17 JUN 2025 |
Compatible VCF version | 9.0 |
Supported VKS Versions
This version of Private AI Services uses VMware vSphere Kubernetes release (VKr) 1.32 clusters with VMware vSphere Kubernetes Service (VKS) ClusterClass builtin-generic-v3.2.0.
Supported NVIDIA GPU Operator Versions
By default, this version of Private AI Services uses NVIDIA GPU Operator 24.9.0. For supported driver versions, see the NVIDIA GPU Operator documentation.
Supported PostgreSQL Versions
This version of Private AI Services supports PostgreSQL 16.8 with pgvector extension 0.8.0.
Supported Inference Server Versions
This version of Private AI Services supports the vLLM inference server 0.6.5 for generating completions and the Infinity inference server 0.0.43 for generating embeddings. Both servers are installed with supported default configurations.
Supported File Types and Data Sources
The Data Retrieval and Indexing component of Private AI Services supports retrieving files from specific types of data sources and extracting text content from the specific file types.
Data Retrieval and Indexing Element | Supported Options | |
|---|---|---|
File types |
| |
Data source types | Microsoft SharePoint |
|
Google Drive | Google Drive API v3, using Google service accounts for authentication. | |
Atlassian Confluence |
| |
S3 Compatible Server | S3 API v4, using Access Key ID/Secret Key credentials for authentication | |
Known Issues Across Private AI Services Versions
Data Retrieval and Indexing reports a General Crawling Error instead of a Credentials Error when indexing a knowledge base if the provided credentials expire after the initial configuration of the data source
Data source credentials can be validated by the Priavet AI Services API and UI upon data source creation. If the provided credentials are invalid or do not grant access to the data source backend, a relevant error message appears.
However, if credentials expire or if access to the data source backend is revoked after data source creation, the indexing process fails with a General Crawling Error instead of reporting an error about credentials becoming invalid.
Workaround: Validate the credentials by using the data source creation workflow. If the data source credentials have become invalid, update the data source with the new ones.
Data Retrieval and Indexing does not validate if data source credentials for Google Drive grant permissions to read a folder
The Google Drive API does not explicitly reject API requests for listing files in a folder if the requesting client does not have permissions on the folder. Instead, the API grants access but returns an empty list.
This means that the validation of data source credentials passes successfully if the credentials are valid but do not grant access to any files.
As a side effect, if the permissions granted to credentials used by a knowledge base are changed after a knowledge base indexing, any subsequent indexing no-longer sees the files from the earlier indexing, and the indexer assumes that files are removed from the folder, and it subsequently removes documents from the knowledge base.
Workaround: Update the credentials used by the data source or ensure that the credentials in use grant access to the expected folders. If an indexing is started and documents get removed, start another indexing after fixing the credentials to ensure that all files are re-added to the knowledge base.
Documentation
Examine the VMware Private AI Foundation with NVIDIA Guide for an overview and how-to instructions on running models, creating knowledge bases, and creating agents.