A key finding from Stanford's SALT Lab revealed that over half of the analyzed interactions involved "consequential tasks"—work that impacts others or is difficult to undo—challenging previous assumptions that users primarily delegate low-stakes tasks to AI.
This research provides rare transparency into the AI infrastructure by moving beyond public datasets, which often skew toward casual use, to examine professional and high-stakes applications.
The study found that users frequently turn to Claude for specialized legal and financial guidance, and while they generally retain oversight of the output, they often face productive friction that requires refining their prompts to achieve better results.
METR’s preliminary data also suggests that more capable AI models provide significant productivity gains for developers, correlating with real-world time savings in coding tasks.
Looking ahead, Anthropic is exploring how to scale this independent research model to allow more outside observers to audit how AI affects society.
To protect user privacy, the company utilized a mechanism where researchers never see raw text, instead receiving aggregate categories and percentages based on AI-generated judgments.
While the process was resource-intensive and revealed some instances of users attempting to bypass safety safeguards, Anthropic has released the aggregated data publicly and is seeking interest from other researchers to expand the program.