Amazon EC2 G7 launches: NVIDIA Blackwell brings 4.6x AI inference and 10x vector search
Cloud AI inference and vector search are getting faster at the same time. On June 23, 2026, NVIDIA and AWS launched the Blackwell-based Amazon EC2 G7 instances. G7 uses NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs to deliver up to 4.6x AI inference over the prior G6 generation, while Amazon OpenSearch Serverless uses NVIDIA cuVS to index vectors up to 10x faster at a quarter of the cost. ASAP summarizes the announcement from the primary source.
Specs first: what a single G7 packs
Amazon EC2 G7 instances use NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs to deliver up to 4.6x AI inference and up to 2.1x graphics over the prior G6 generation. A single instance packs up to 8 GPUs, 256GB of total GPU memory, 700 Gbps EFA networking, and a 7.6TB NVMe SSD. It runs on AWS Deep Learning AMIs and Containers, EMR, EKS, and ECS, with SageMaker AI support coming soon.
How to read the 4.6x and 2.1x
The two multipliers describe different things. The 4.6x applies to AI inference and the 2.1x to graphics workloads, and both are maximums measured against the prior G6 generation. A maximum is a ceiling reached under specific conditions, so it is normal for real workloads to land below it. Rather than deciding on adoption from a single multiplier, it is more practical to first check whether the 256GB of total memory across 8 GPUs and the 700 Gbps of bandwidth fit your model size and batch scale. It is also worth noting, as context for reading the numbers, that RTX PRO 4500 is a Server Edition tier rather than NVIDIA's top-end data-center line.
cuVS ships as a default in OpenSearch Serverless
Amazon OpenSearch Serverless makes GPU-accelerated vector indexing a default capability through NVIDIA cuVS, running up to 10x faster than CPU at a quarter of the cost. NVIDIA and AWS say billion-scale vector databases become practical to build in under an hour. The change cuts both the cost and the time of indexing for RAG and search services at once.
Why this matters for practitioners
The implications for teams standing up RAG are concrete. Large-scale vector indexing has meant either spinning up your own GPU cluster or tolerating indexing times measured in days. With cuVS shipping as a default in OpenSearch Serverless, you can run indexing on the managed service alone without separate GPU infrastructure, and the quarter-of-the-cost figure versus CPU pulls forward the break-even point on document pipelines that were shelved over indexing cost. The 10x is a CPU-relative maximum and should be treated as such, but faster indexing also means you can refresh documents on a shorter cycle, which is felt most in internal search or customer-support bots where freshness matters.
Down to GB300 training: AWS earns NVIDIA Exemplar Cloud status
AWS earned NVIDIA Exemplar Cloud certification for NVIDIA GB300 training workloads. Exemplar status means AWS meets NVIDIA's reference-architecture performance benchmarks. Verification now spans not just inference GPUs but large-scale training infrastructure as well.
The point of bundling it into one stack
Taken separately, these are three pieces: one GPU instance, one search library, one certification. But the reason NVIDIA and AWS announced them together on the same day is to line up G7 for inference, cuVS for search, and GB300 for training inside a single cloud stack. It is a configuration designed so that users never leave the cloud as they move from inference to search to training. The 4.6x and 10x figures are maximums versus the prior generation and CPU, so real workloads may differ, and that puts the weight of this announcement on unified operation rather than on any single multiplier.
Source: ASAP summary of NVIDIA's blog "NVIDIA and AWS Collaborate to Bring AI to Production at Scale" (June 23, 2026; Amazon EC2 G7 with RTX PRO 4500 Blackwell Server Edition, up to 4.6x AI inference and up to 2.1x graphics over G6, up to 8 GPUs, 256GB, 700 Gbps EFA, 7.6TB NVMe, Amazon OpenSearch Serverless with NVIDIA cuVS for up to 10x vector indexing at a quarter of the cost, billion-scale vector databases in under an hour, and AWS NVIDIA GB300 Exemplar Cloud certification).

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr