NVIDIA Reveals Vera CPU Architecture With 88 Olympus Cores and High Bandwidth LPDDR5X Memory

news
NVIDIA Reveals Vera CPU Architecture With 88 Olympus Cores and High Bandwidth LPDDR5X Memory

NVIDIA has shared a closer look at its upcoming Vera data centre processor, revealing new information about the custom Olympus CPU core, memory design and expected performance.

Vera is being developed for the CPU work surrounding modern artificial intelligence systems. It will launch as part of NVIDIA’s Vera Rubin platform, where it will operate alongside Rubin accelerators in large AI servers.

The processor uses 88 custom Olympus cores based on the Arm architecture. Each core can handle two tasks through NVIDIA’s Spatial Multithreading technology. NVIDIA says this design is intended to maintain predictable per core performance when many workloads are running at the same time.

The company is placing particular emphasis on AI agents, which regularly move between GPU inference and CPU based tasks. An agent may use the CPU for code execution, data processing, searches, orchestration and tool calls before returning to the accelerator. These repeated transitions make CPU response time and memory access important to the overall performance of the system.

Vera featureTechnical detail
CPU cores88 custom Olympus cores
ArchitectureArm compatible design
MemoryUp to 1.5TB of LPDDR5X
Memory bandwidthUp to 1.2TB per second
MultithreadingTwo tasks per core
Platform pairingNVIDIA Rubin accelerators
Main workloadsAI agents, reinforcement learning, analytics and orchestration

NVIDIA uses a monolithic design to reduce internal latency

Vera uses a monolithic processor design instead of dividing its CPU cores across several compute chiplets. NVIDIA argues that chiplet based server processors can introduce additional latency when information must travel between separate dies.

The company refers to this overhead as a chiplet tax. According to NVIDIA, moving data across multiple compute dies can reduce effective memory bandwidth and make communication between cores less efficient.

oplus_2097152

Vera connects its cores, cache, memory and input and output hardware through the second generation NVIDIA Scalable Coherency Fabric. NVIDIA claims this fabric can provide up to three times the core to core bandwidth of competing chiplet based designs.

These figures are company claims and will need to be tested independently after systems become available. A monolithic design can improve communication latency, but it may also be more difficult and expensive to manufacture as the chip becomes larger.

Olympus focuses on strong performance from each core

The Olympus core uses an out of order execution design with branch prediction, register renaming, execution schedulers and dedicated load and store hardware.

NVIDIA’s early architecture diagram also shows a wide instruction frontend capable of decoding as many as 10 instructions during each cycle. A wider frontend can allow the processor to examine and execute more work at once, although real performance will depend on the complete architecture, software and workload.

NVIDIA describes Vera as delivering maximum single threaded performance at scale. The company is not only targeting short tests that use a small number of cores. It says the processor is designed to preserve strong per core performance while all 88 cores are handling active workloads.

The company claims Olympus offers 50 percent higher instructions per cycle than the Grace CPU architecture. NVIDIA has also presented a broader claim of up to twice the performance of current x86 server processors, though results will vary by workload and system configuration.

LPDDR5X memory provides up to 1.2TB per second of bandwidth

Vera will support up to 1.5TB of LPDDR5X memory and deliver as much as 1.2TB per second of memory bandwidth.

The memory will use compact SOCAMM modules instead of conventional DDR5 server DIMMs. NVIDIA says Vera can provide lower memory latency and greater bandwidth per core than current x86 server platforms while using less power.

High memory bandwidth is useful for AI infrastructure because the CPU may need to manage thousands of software environments, process large datasets and coordinate frequent transfers between the processor and accelerators.

NVIDIA has reported performance improvements across several early customer workloads. The claimed gains include faster sandbox execution, lower streaming latency and higher performance in selected agent simulation tests. These comparisons use specific applications rather than a common set of independent CPU benchmarks, so they do not yet provide a complete picture of Vera’s performance.

Vera represents NVIDIA’s effort to design the complete AI server platform, including the CPU, accelerators, networking, memory and software. Its real position against AMD EPYC and Intel Xeon processors will become clearer when independent testing examines performance, power use, software compatibility and cost across a wider range of data centre workloads.

Discover: News

Discussion (0)

Be the first to comment.