How llm-d makes the most of the hardware you already have IBM Research 8 September 2026 at 12:00 IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.