Intel Arc Pro B70 Outperforms Nvidia RTX 5090D in AI Inference Throughput with 2,320 Tokens/s
A GUNNIR benchmark reveals Intel's Arc Pro B70 GPU hits 2320.76 tokens per second in DeepSeek R1 inference, beating Nvidia's RTX 5090D by up to 8.6% at high concurrency.
In a fresh round of AI inference testing, Intel's Arc Pro B70 professional GPU has emerged as a surprising performance leader, outpacing Nvidia's flagship RTX 5090D in token throughput under heavy concurrent loads.
The benchmarks, conducted by graphics card manufacturer GUNNIR and reported by Wccftech on July 8, compared four-card configurations of the Intel Arc Pro B70 32GB, Nvidia RTX 5090D 32GB, and RTX 4090D 24GB. The workload was the DeepSeek R1-Distill Qwen 32B FP16 model, with a fixed sequence length of 128 tokens for both input and output, and the key metric was token-per-second throughput across concurrency levels from 1 to 512.
At low concurrency (below 32 simultaneous requests), the RTX 5090D held a clear lead, while the RTX 4090D ranked second. The Arc Pro B70 initially matched the 4090D but began to pull ahead as concurrency increased. At 32 and 64 requests, the B70 first caught up with and then surpassed the RTX 4090D.
The real turning point came at higher concurrency. With 128 simultaneous requests, the Arc Pro B70 delivered 8.6% more tokens per second than the RTX 5090D and a striking 34.2% more than the RTX 4090D. At 256 requests, the B70's advantage widened to 7.5% over the 5090D and 48.7% over the 4090D. Similar gains were observed at the maximum 512-request concurrency level.
The peak recorded throughput for the Intel card reached 2320.76 tokens per second — a new high for this class of hardware. While the full test report does not include latency or power consumption data, the results suggest that Intel's Arc Pro lineup is particularly well optimized for high‑density inference deployments, where the ability to juggle many simultaneous requests can trump raw single‑stream speed.
The tests underscore a shifting competitive landscape in the AI accelerator market, as Intel continues to refine its data-center and professional GPU architectures. More details, including frame-by‑frame concurrency charts, are available in the original report.