VCF 9.1 Private AI Foundation

Private AI Sizer

Size your GPU infrastructure for LLM inference on VMware Cloud Foundation. All calculations follow VCF 9.1 design recommendations.

GPU Selection

Cluster Configuration

1= 8 total GPUs

LLM Model

Layers
32
KV Heads
8
Head Dim
128
Request a Model

Quantization

Context Length

1281,024 tokens4,096
912
Max Concurrent Requests
225
Tokens/sec
17 ms
TTFT
8
Data Parallel
Replicas

GPU Memory Breakdown

KV Cache
Free
Weights: 3%KV Cache: 79%Free: 18%
Model Weights
14.9 GB
KV Cache / Request
512.00 MB
Available for KV
57.1 GB
Total GPU Memory
640.0 GB
Usable (vLLM 90%)
576.0 GB
KV Cache / Token
128.00 KB

Parallelism Strategy

Min GPUs for Model
1
TP Degree
None (fits 1 GPU)
Replicas (Data Parallel)
8
Total GPUs
8

Performance Estimates

Generation Throughput
225 tok/s
Target (Interactive)
≥ 30 tok/s
Time-to-First-Token
16.6 ms
Target (Interactive)
≤ 200 ms
Throughput estimated per VCF 9.1 PAIF-ACC-RCMD-002: generation_time = model_bytes / memory_bandwidth. TTFT per PAIF-ACC-RCMD-003: prompt_tokens x (params x 2 / GPU_TFLOPS). Real-world performance varies with batching and load.

Host DRAM & Compute Requirements

Per Host
VM RAM (DRAM)
1.3 TB
2× GPU memory
Server RAM
1.6 TB
2.5× GPU memory
vCPUs
64
8 per GPU
Per VCF 9.1: VM RAM 1-2× GPU memory (PAIF-ACC-RCMD-008) to support vLLM runtime, model loading buffers, and KV cache overhead. Server RAM 2-3× (PAIF-ACC-RCMD-009) for hypervisor overhead. vCPUs 4-8 per GPU (PAIF-ACC-RCMD-011) for tokenization and request handling.

Methodology

All sizing formulas follow the VMware Cloud Foundation 9.1 Private AI Compute Detailed Design (PAIF-ACC-RCMD-001 through PAIF-ACC-RCMD-011). The 90% GPU memory utilization reflects the vLLM runtime default (gpu_memory_utilization=0.90), not a VCF specification.

Throughput and TTFT are theoretical single-request estimates. Production performance varies based on batching, concurrent load, model implementation, and vGPU vs. DirectPath I/O configuration. For production deployments, benchmark with your actual workloads.