OpenAI’s “Spring Update,” livestreamed from its San Francisco offices on May 13, 2024, introduced GPT-4o (“o” for “omni”) by demonstrating it live rather than describing it. A researcher held up a sheet of paper with a handwritten equation and talked through it with the model in real time, in a back-and-forth rather than a query-and-response. In a separate segment, presenters had the model translate a live conversation between English and Italian, and asked it to read emotion from a face on camera.
The pitch wasn’t a capability score or a benchmark chart. The whole presentation was built around latency and tone: a voice fast and unhalting enough that talking to it stopped feeling like operating software and started resembling a conversation. That demo became the reference point the rest of the industry was measured against for what an AI voice assistant was now expected to do, which meant OpenAI had, in one livestream, set the terms other companies would spend the following year trying to match.