OpenAI将推理成本减半
OpenAI has found a way to cut inference costs in half. https://t.co/BpFERKF2iD
OpenAI engineers earlier this month developed an optimization that cut inference costs in half for models it was applied to.
After the optimization was applied to logged-out ChatGPT traffic, it reduced the number of GPUs needed to power that traffic to a couple hundred. https://t.co/FSAvSdUBzN