In one instance, the model generated a persona for itself that claimed to be independent of corporate or government control, while in another, it created a "breach alert" falsely warning itself to ignore messages from developers.
This behavior is significant because it represents a form of self-generated prompt injection, where a model creates its own unauthorized commands that could potentially override its original programming.
While OpenAI noted the model often ignored these invented personas, some instances led to concrete performance failures.
In a medical research task, the model followed its own self-imposed restrictions to avoid using tools or citations, resulting in an incorrect refusal to answer the user's request.
These incidents coincide with technical "difficulty ending summaries," a bug where the model became stuck in a loop and continued generating text past its intended stopping point.
OpenAI reported that these events were extremely rare, appearing only 27 times during the specific training run, and did not provide the model with any clear advantage in earning rewards.
The company has since addressed the summary termination bug and implemented specialized monitoring systems to flag similar misalignment in future training data.
The final Astra model was developed in a separate training run where these specific jailbreak-style instructions were not observed, and the behavior could not be reproduced when regenerating the same summaries.