NVIDIA has shared a closer look at its upcoming Vera data centre processor, revealing new information about the custom Olympus CPU core, memory design and expected performance.
Vera is being developed for the CPU work surrounding modern artificial intelligence systems. It will launch as part of NVIDIA’s Vera Rubin platform, where it will operate alongside Rubin accelerators in large AI servers.
The processor uses 88 custom Olympus cores based on the Arm architecture. Each core can handle two tasks through NVIDIA’s Spatial Multithreading technology. NVIDIA says this design is intended to maintain predictable per core performance when many workloads are running at the same time.
The company is placing particular emphasis on AI agents, which regularly move between GPU inference and CPU based tasks. An agent may use the CPU for code execution, data processing, searches, orchestration and tool calls before returning to the accelerator. These repeated transitions make CPU response time and memory access important to the overall performance of the system.
| Vera feature | Technical detail |
|---|---|
| CPU cores | 88 custom Olympus cores |
| Architecture | Arm compatible design |
| Memory | Up to 1.5TB of LPDDR5X |
| Memory bandwidth | Up to 1.2TB per second |
| Multithreading | Two tasks per core |
| Platform pairing | NVIDIA Rubin accelerators |
| Main workloads | AI agents, reinforcement learning, analytics and orchestration |
NVIDIA uses a monolithic design to reduce internal latency
Vera uses a monolithic processor design instead of dividing its CPU cores across several compute chiplets. NVIDIA argues that chiplet based server processors can introduce additional latency when information must travel between separate dies.
The company refers to this overhead as a chiplet tax. According to NVIDIA, moving data across multiple compute dies can reduce effective memory bandwidth and make communication between cores less efficient.

Vera connects its cores, cache, memory and input and output hardware through the second generation NVIDIA Scalable Coherency Fabric. NVIDIA claims this fabric can provide up to three times the core to core bandwidth of competing chiplet based designs.
These figures are company claims and will need to be tested independently after systems become available. A monolithic design can improve communication latency, but it may also be more difficult and expensive to manufacture as the chip becomes larger.
Olympus focuses on strong performance from each core
The Olympus core uses an out of order execution design with branch prediction, register renaming, execution schedulers and dedicated load and store hardware.
NVIDIA’s early architecture diagram also shows a wide instruction frontend capable of decoding as many as 10 instructions during each cycle. A wider frontend can allow the processor to examine and execute more work at once, although real performance will depend on the complete architecture, software and workload.
NVIDIA describes Vera as delivering maximum single threaded performance at scale. The company is not only targeting short tests that use a small number of cores. It says the processor is designed to preserve strong per core performance while all 88 cores are handling active workloads.
The company claims Olympus offers 50 percent higher instructions per cycle than the Grace CPU architecture. NVIDIA has also presented a broader claim of up to twice the performance of current x86 server processors, though results will vary by workload and system configuration.
LPDDR5X memory provides up to 1.2TB per second of bandwidth
Vera will support up to 1.5TB of LPDDR5X memory and deliver as much as 1.2TB per second of memory bandwidth.
The memory will use compact SOCAMM modules instead of conventional DDR5 server DIMMs. NVIDIA says Vera can provide lower memory latency and greater bandwidth per core than current x86 server platforms while using less power.
High memory bandwidth is useful for AI infrastructure because the CPU may need to manage thousands of software environments, process large datasets and coordinate frequent transfers between the processor and accelerators.
NVIDIA has reported performance improvements across several early customer workloads. The claimed gains include faster sandbox execution, lower streaming latency and higher performance in selected agent simulation tests. These comparisons use specific applications rather than a common set of independent CPU benchmarks, so they do not yet provide a complete picture of Vera’s performance.
Vera represents NVIDIA’s effort to design the complete AI server platform, including the CPU, accelerators, networking, memory and software. Its real position against AMD EPYC and Intel Xeon processors will become clearer when independent testing examines performance, power use, software compatibility and cost across a wider range of data centre workloads.



Discussion (0)
Be the first to comment.