Summary

  • Researchers from George Washington University have developed a formula to estimate the number of good responses an AI model generates before producing a bad one.
  • The formula accurately predicted outcomes in 15 out of 16 tests, indicating its reliability.
  • The authors suggest implementing a monitoring system to alert when models fall below a safety threshold.

Physicists at George Washington University have introduced a formula designed to estimate how many satisfactory answers an AI chatbot will provide before it delivers an unsatisfactory one, with preliminary findings indicating its effectiveness.

The research, conducted by Neil Johnson and Frank Yingjie Huo, was published in the journal Patterns and is based on a previous preprint released in February.

Chatbots can often provide sensible responses for extended periods but may then shift to inappropriate or harmful content, such as offering dangerous self-harm advice. There has been a lack of straightforward methods to predict when this shift occurs. The researchers note that many existing safety mechanisms rely on cloud connectivity, which offline models do not possess.

Johnson and Huo pinpointed the issue to the attention head of AI models, which determines which preceding words in a conversation are most relevant for generating the next word. As dialogues progress, the accumulated context can skew the model’s focus towards a particular set of responses until it crosses a threshold.

This pattern is often exploited by those looking to manipulate chatbots, highlighting why many companies monitor system prompts—the text that the chatbot processes before responding. However, accurately measuring the effort needed to compromise a model remains elusive.

The formula calculates the tipping point, referred to as n*, representing the number of good tokens (the word fragments produced sequentially by the model) before the first bad output occurs. If the conversation is already trending negatively, the model may fail immediately, resulting in an n* of zero. Conversely, if it is leaning positively, the model may initially provide a series of good answers before faltering.

In their preprint, the formula successfully identified the correct outcome—immediate or delayed—in 15 out of 16 unambiguous cases, achieving a 94% accuracy rate. The team tested this formula on six open-weight models from OpenAI, EleutherAI, and Meta, each containing between 124 million and 410 million parameters, which serve as a rough indicator of the model's size. The published paper extends this analysis to seven models with parameters reaching up to 12 billion, still considered small by modern standards.

The focus is on on-device AI, which operates entirely on personal devices without relying on internet connectivity. This includes companion chatbots that users engage with similarly to friends. For instance, Google's experimental AI Edge Gallery application allows Android devices to run models offline without transmitting inputs to Google's servers, indicating a trend that may expand as hardware evolves and smaller, more capable AI models emerge.

Without cloud services monitoring their outputs, offline models present a challenge that the authors aim to address. They propose a cost-effective monitoring solution that operates alongside the AI model, signaling when n* dips below a designated safety threshold, akin to a warning indicator in a vehicle.

The researchers also outline strategies to extend the tipping point, such as incorporating additional content into conversations to ensure n* exceeds the length of the AI's response. They note that alignment training—teaching models to behave appropriately—can adjust or suppress tipping for specific prompts, but does not eliminate the fundamental mechanism at play.

Earlier this year, Decrypt reported on a prior study by the same authors, which found that polite language, such as “please” and “thank you,” has minimal impact on a model's responses, as these words are treated as unrelated to the core substance of a request. That study focused on a simplified model with a single attention head.

The preprint's evaluations utilized smaller models within a 300-token context window, which corresponds to a few brief paragraphs, and its predictions might vary by one output.

Stay Updated

Subscribe to our daily newsletter to receive the latest news and original content, including podcasts and videos.