inference

Cerebras launches ultrafast mode for OpenAI’s GPT-5.6 Sol

Cerebras Systems announced Thursday that it is powering a new "Ultrafast" service tier in the OpenAI API for GPT-5.6 Sol, delivering up to 750 output tokens per second — up to 14 times faster than Standard processing with no loss…

AMD to buy startup that hardwires AI models into chips

AMD Advanced Micro Devices, Inc. announced Thursday it has reached a definitive agreement to acquire Toronto-based Taalas, a startup that hardwires AI model weights directly into silicon to deliver inference speeds far beyond conventional GPUs. Financial terms were not disclosed.

OpenAI finds optimization that halves inference costs

OpenAI engineers earlier this month developed an optimization that reduces inference costs by more than half for the models it has been applied to, according to a report by The Information. The breakthrough, which stems from squeezing more efficiency out…