Running AI on mixed hardware for speed and affordability IBM Research 23 June 2026 at 18:00 Researchers show that serving AI models with llm-d can boost inference speeds by up to 5 times and double throughput β all while using heterogeneous GPUs.