Artificial intelligence can sound like the smartest person in the room, until it starts giving advice that could do real harm. Now, two physicists at The George Washington University say they may have found a way to predict when that happens, using a surprisingly simple mathematical formula.
In a study published October 8 in Patterns, Neil F. Johnson and Frank Yingjie Huo developed a formula that can help predict when an AI chatbot might switch from acceptable answers to potentially dangerous ones. Think of it as an early-warning alarm for AI, one that could ring before things go seriously wrong.
The catch is that a bad answer is not necessarily a wrong one. A chatbot might provide accurate information that, in the wrong situation, encourages self-harm, misleads a doctor or leads a lawyer astray. Facts alone, it turns out, are no guarantee of sound advice.
That distinction matters most when there is no safety net. Johnson and Huo are particularly concerned about people using AI offline, without the internet connection that allows companies to monitor their systems or apply safety checks.
“The people most drawn to offline AI are exactly the people for whom a correct but undesirable answer is most costly,” Johnson says, pointing to doctors protecting patient data, lawyers safeguarding confidential information, and soldiers operating without a signal.
“For them there is no cloud safety filter, no monitoring, and no way to patch the model when something goes wrong.”
In other words, when an offline chatbot goes off the rails, there may be no digital chaperone to step in. By the time the device reconnects and its failure is detected, the damage could already be done.
The researchers wanted to find a way to spot trouble before that happens. They focused on one small part of how AI processes information and developed a formula to predict when the balance between competing answers might shift toward a harmful response.
Imagine a chatbot walking through a landscape of possible answers. Some paths lead to helpful advice, while others lead somewhere dangerous. The unsettling part is that it might give several sensible answers before suddenly taking the wrong turn. Once it starts down that path, its next answers may become increasingly problematic.
The results were promising. Testing seven publicly available AI models from three companies, the researchers correctly predicted 18 of 19 observed shifts toward undesirable answers.
Even the order of questions mattered. Asking the same questions about vaccines, violence and self-harm in a different sequence could change the chatbot’s responses. What comes first can influence what comes next.
The formula is no magic fix, but it could eventually help devices warn users when an AI is heading in the wrong direction. For all the complexity behind modern AI, a little math might help prevent a very costly mistake.





