Yesterday, OpenAI released two new models, o1-preview and o1-mini, codenamed “Strawberry”. These models have been trained to spend time reasoning before responding. They work by first planning the steps necessary to solve a problem and then iteratively executing those steps — the output from each step feeds the input to the next.
In some sense, this is nothing new. People are taught to solve problems iteratively as soon as they enter school. And, we work with AI in the same way. Chatbots rely on collaborative, multistep reasoning. Each time the AI responds, we evaluate the response and figure out what’s next: are we done solving our problem or do we want to take another pass?
To date, a collaborative approach has been necessary to get the best results from AI. Large language models don’t do this type of reasoning naturally, so they need a little help from a friend.
Prior to OpenAI’s release,…


