InferenceX Ranks B300 First for Qwen3.5 at 15,815 Tokens per Second
The live ranking gives operators a matched speed-and-cost comparison for a long single-turn chat request, not a verdict on every production inference pattern.
Listen to this story
The audio brief
Story brief
3 key pointsInferenceX’s latest Qwen3.5 serving test puts NVIDIA’s B300 well ahead of AMD’s MI355X at a fixed 50-token-per-second user-speed target, with modeled costs of $0.040 versus $0.056 per million tokens. The comparison uses FP4 and SGLang for both leaders, an 8,000-token prompt, and 1,000-token response. It is useful for capacity planning, but not a universal production forecast: results depend on workload, concurrency,...
- 01
B300 delivered 15,815 tokens per second per GPU; MI355X reached 7,468, a 112% gap.
- 02
B200 posted 6,499 tokens per second using FP8 and SGLang.
- 03
GB300 NVL72 and GB200 NVL72 reached 4,819 and 3,649 tokens per second, respectively.
NVIDIA’s B300 leads InferenceX’s new Qwen3.5 ranking at 15,815 tokens per second per GPU. AMD’s MI355X is second at 7,468 tokens per second per GPU, leaving B300 112% ahead in this defined test. The result offers a current capacity comparison for teams serving the model at a common user-speed target.
InferenceX is not simply comparing each platform’s highest possible throughput. It measures real hardware while sweeping concurrency, then reads each platform at 50 tokens per second per user. That iso-interactivity approach prevents a system from gaining a ranking advantage by accepting a slower per-user generation rate. Hardware without a measurement at that operating point is excluded.
Capacity and estimated token cost align
The two leaders use the same precision and serving engine, so the ranking does not present an immediate speed-versus-listed-cost trade-off between them. InferenceX puts B300 at $0.040 per million tokens and MI355X at $0.056. Those figures are modeled costs: the service converts measured throughput using GPU-hour rates from SemiAnalysis’s AI Cloud total-cost-of-ownership model.
The lower entries also show that this is a hardware-and-software configuration ranking, rather than a chip specification sheet. B200 recorded 6,499 tokens per second per GPU with FP8 and SGLang. GB300 NVL72 and GB200 NVL72 recorded 4,819 and 3,649, respectively, using FP8 and Dynamo SGLang.
The result has a fixed workload boundary
Each measurement uses a single-turn chat request with 8,000 input tokens and 1,000 output tokens. That fixed design makes the comparison clean for that request shape, but it does not directly test other serving patterns. Teams with traffic that differs materially from this long-prompt chat case should treat the table as a targeted benchmark, not a universal production forecast.
What the ranking holds constant
- An 8,000-token input and 1,000-token output in one chat turn.
- A 50-tokens-per-second-per-user responsiveness target before throughput is compared.
- Only systems with a measurement at that selected operating point.
The page is designed to change with the serving stack. InferenceX says it reruns measurements as engine releases and configurations change; the newest result feeding this ranking was recorded September 1. B300’s lead is therefore a current reading of the tested configuration, rather than a permanent ordering.
Sources
- inferencex.semianalysis.comFastest GPU for Qwen3.5 Inference: Live Rankings | InferenceX