Intel has detailed Crescent Island, a new AI inference accelerator built around its Xe3P architecture and designed specifically for agentic AI workloads.
The accelerator combines up to 32 Xe3P cores with as much as 480GB of LPDDR5X memory, giving Intel a different approach from competing data center GPUs that rely heavily on HBM. The company is targeting lower cost, reduced power consumption and large memory capacity for inference focused deployments.
Intel plans to use Crescent Island in air cooled data center systems and workstations, with customer sampling targeted for the second half of 2026.
Intel Crescent Island specifications
| Feature | Detail |
|---|---|
| Architecture | Xe3P |
| Xe cores | 32 |
| Compute slices | 4 |
| Xe cores per slice | 8 |
| Vector Engines | 256 |
| XMX engines | 256 |
| Maximum memory | 480GB LPDDR5X |
| Intel reference card memory | 160GB LPDDR5X |
| L1 cache and SLM | 16MB total |
| Unified L2 cache | 32MB |
| TDP | 350W for air cooled PCIe model |
| Main workload | AI inference |
| Sampling target | Second half of 2026 |
Each Xe3P core contains eight Vector Engines and eight XMX engines. Across the complete Crescent Island GPU, that produces 256 Vector Engines and 256 XMX units.
Intel has also expanded numerical format support. Xe3P includes FP8 and FP4 capability in the Vector Engines, while the architecture supports formats ranging from FP4 through FP64.
LPDDR5X is central to Intel’s strategy
One of the biggest differences between Crescent Island and many competing AI accelerators is memory.
Instead of using HBM, Intel is relying on LPDDR5X.
The company says this choice provides higher density and lower power consumption while helping reduce system cost. It is also intended to make large memory capacities more practical for inference workloads that depend on long context windows and large KV caches.

Intel’s own reference PCIe card will carry 160GB of LPDDR5X, while system partners will be able to build configurations with as much as 480GB.
For comparison, the supplied information lists AMD Instinct MI450X at up to 432GB of HBM4 and NVIDIA Vera Rubin at up to 288GB of HBM4. These figures describe capacity rather than bandwidth or overall performance, so they should not be treated as direct measures of accelerator speed.
Crescent Island targets a 350W power envelope
Intel says the air cooled PCIe version of Crescent Island will operate at a 350W TDP.
That is part of the company’s focus on performance per watt and total cost of ownership rather than maximum possible accelerator performance.
The LPDDR5X memory subsystem also supports a densely packed channel layout, which Intel says improves bandwidth while keeping power consumption lower than an HBM based approach.
Crescent Island is therefore being positioned for organizations that need large memory pools and sustained inference performance without moving to very high power liquid cooled systems.
Graphics hardware has been removed
Crescent Island is not intended to serve as a conventional graphics card.
Intel has removed traditional graphics and 3D functionality so more die area can be devoted to general purpose GPU and AI compute.
That makes the accelerator a dedicated inference product rather than a gaming or visualization GPU.
The design is optimized around workloads such as language models, reasoning systems, multimodal AI and diffusion models.
Software will support heterogeneous AI infrastructure
Intel is also emphasizing software compatibility.
Crescent Island is designed to support frameworks including vLLM, SGLang, llm d and NVIDIA Dynamo.
The platform will support KV cache aware routing and offloading, which can help distribute inference workloads across larger systems.
Intel also says agents should be able to operate across heterogeneous infrastructure without requiring changes to agent code.
That is an important part of Crescent Island’s positioning because Intel is not only trying to compete through hardware. It is also attempting to make its accelerators easier to integrate into AI environments that may already contain GPUs from multiple vendors.
Intel is focusing on inference rather than training
Crescent Island is clearly designed around AI inference.
Intel highlights larger context lengths, high memory capacity, KV cache efficiency and strong performance per watt as core strengths.
The accelerator is also described as being optimized for prefill, a compute intensive stage of language model inference that processes the initial prompt before token generation begins.
That makes the design particularly relevant for organizations offering tokens as a service or running large numbers of AI agents.
Intel is targeting customer sampling in the second half of 2026, so more details about real world performance, pricing and partner systems should become available as Crescent Island moves closer to deployment.



Discussion (0)
Be the first to comment.