NVIDIA announced a deeper collaboration with Amazon Web Services (AWS) that strengthens the foundation for running AI in production at scale[1]. Built on three pillars, namely the launch of the new GPU-equipped "Amazon EC2 G7" instances, the standardization of GPU-accelerated vector search, and a performance certification for training, the effort aims to make it easier for enterprises to move AI from proof-of-concept to real operations[1][2].
New "Amazon EC2 G7" GPU Instances
At the center of this announcement are the new "Amazon EC2 G7" instances, powered by NVIDIA's "RTX PRO 4500 Blackwell Server Edition" GPUs[1][2]. AWS becomes the first major cloud provider to support these server-class Blackwell GPUs[2]. Availability began on June 18, 2026, initially in two regions, US East (Ohio) and US West (Oregon)[2].
On performance, the company says G7 delivers up to 4.6x AI inference and up to 2.1x graphics processing compared with the previous "G6" generation[1][2]. Each instance can pack up to 8 GPUs, with a total of 256GB of GPU memory, up to 700Gbps of EFA (Elastic Fabric Adapter) networking, and up to 7.6TB of local NVMe SSD storage[1]. Customers can choose configurations with 1, 2, 4, or 8 GPUs, and a bare-metal option (using a physical server directly, without a virtualization layer) is coming soon, letting teams right-size infrastructure rather than over-provision[1]. The instances use custom Intel Xeon 6 processors[2].
The intended uses are broad. In addition to AI inference, a single instance type can handle high-resolution video and image processing, CAD (computer-aided design), virtual desktops, gaming, and spatial computing[1]. For data analytics, using the "NVIDIA cuDF" library for Apache Spark on Amazon EMR brings GPU acceleration to analytics pipelines[1]. G7 is accessible through AWS deep learning machine images and containers, as well as Amazon EMR, EKS, and ECS, with support for Amazon SageMaker AI coming soon[1].
Making GPU-Accelerated Vector Search the Default
The second pillar is an overhaul of the "Amazon OpenSearch" search platform. In the next generation of "OpenSearch Serverless," GPU-based vector indexing powered by NVIDIA's open-source "cuVS" library becomes the default compute choice for all vector collections[1][3]. Vector search converts text, images, and other content into arrays of numbers (vectors) and finds the closest matches by meaning, a mechanism now widely used as the foundation of generative AI.
This shift matters most for use cases such as retrieval-augmented generation (RAG, a technique that improves answer accuracy by referencing external data), semantic search, recommendation systems, and autonomous agentic AI[1]. According to NVIDIA, compared with building indexes on CPUs alone, vector indexing becomes up to 10x faster at a quarter of the cost, and billion-scale vector databases can be built in under an hour[1][3]. The key change is that GPU-powered vector search, which used to be treated as a specialized optimization project, becomes a standard AWS capability anyone can use[1].
A Performance "Seal of Approval" for Training: Exemplar Cloud Status on GB300
The third pillar is a certification for training workloads. AWS achieved NVIDIA's "Exemplar Cloud" status for training on NVIDIA's top-tier "GB300" system[1]. This means the AWS environment meets the rigorous performance thresholds that NVIDIA defines against its reference architecture (a benchmark configuration used as a yardstick for performance)[1].
NVIDIA positions the certification as the result of deep co-engineering between the two companies[1]. Through the "Exemplar Clouds" initiative, developers and AI leaders can more easily judge whether the cloud infrastructure used for large-scale training delivers consistent, high performance, helping them improve total cost of ownership (TCO) and move more efficiently from planning to production[1].
Summary
This announcement from NVIDIA and AWS lifts every layer of the AI infrastructure stack at once, through three points: the new "EC2 G7" GPU instances, the standardization of GPU-accelerated vector search via cuVS, and Exemplar Cloud status on GB300. The common goal is to provide a foundation that performs at production scale without adding operational burden, making it easier for enterprises to move from experimenting with AI to actually using it. As generative AI spreads from pilots into everyday work, the question is how far the cloud side can lower the barriers to adoption. That said, the figures presented are all based on measurements by NVIDIA and AWS themselves, so the real-world impact depends on validation in each individual workload.
Source: https://blogs.nvidia.com/blog/nvidia-aws-ai-production-scale/
