Red Hat delivers peak performance on Kubernetes and CPUs in MLPerf Inference v6.1
Red Hat is proud to announce our results from the industry-standard MLPerf Inference v6.1 benchmark. This submission builds on our track record across recent rounds: In v5.1, we demonstrated cost-effective Llama-3.1-8B-FP8 inference with vLLM on NVIDIA H100 and L40S GPUs, and in v6.0 we delivered results across Qwen3-VL, Whisper, and gpt-oss-120b on NVIDIA H200 and B200 and AMD Instinct MI350 GPUs, including the first Kubernetes-based submission.Our v6.1 results highlight 3 things: peak performance on Kubernetes, 1 inference engine (vLLM) spanning…
Quelle: Originalartikel öffnen




