External researchers study 250,000 conversations with Claude

Stanford, Oxford, and METR studied real-world Claude usage through aggregated data without gaining access to the raw conversations.

How can researchers study the real-world effects of artificial intelligence when conversations remain locked inside the companies developing the models? Anthropic is testing one possible answer by allowing three outside teams to analyze nearly 250,000 interactions from Claude.ai and Claude Code.

The conversations were selected from interactions that took place in April and May 2026. They were studied by Stanford University’s Social and Language Technologies Lab, the University of Oxford’s Human Information Processing Lab, and METR, an organization specializing in the evaluation of advanced AI models.

The teams never accessed the original conversations. Instead, they used Anthropic Insights, formerly known as Clio, an internal tool designed to identify patterns across millions of interactions while limiting the risk of identifying individual users.

The process starts with questions written by the researchers. A team might ask what kind of professional guidance users are seeking or how much control people retain while completing a task. Claude then examines the relevant conversations, sorts them into categories, and produces aggregated results. Researchers receive those categories and their proportions, but not the underlying messages.

Anthropic ran the analyses on behalf of the research teams and reviewed the results before sharing them. Under the collaboration agreements, its review rights were limited to user privacy, information that could help people violate its policies, confidential company information, and research accuracy. The partners remain free to publish findings that are unfavorable to Anthropic.

Stanford’s study examines collaboration between humans and Claude. It looks at the tasks users assign to the model, the responsibilities they retain, and the moments when that collaboration becomes difficult.

Its early results suggest that users delegate consequential work more often than expected. More than half of the analyzed conversations reportedly involved tasks that could affect other people or be difficult to reverse. The proportion increased when users sought professional guidance from Claude, particularly on legal and financial matters.

That finding requires caution. A conversation classified as consequential does not necessarily mean that Claude’s response was used or directly influenced a decision. The category reflects the content of the exchange and the interpretation of the system analyzing it.

In nearly three-quarters of conversations, people reportedly retained control over the direction of the work while Claude assisted them. Users generally modified its responses rather than adopting them word for word. That supervision does not guarantee that they fully understood the model’s recommendations or were able to verify their accuracy.

The Stanford team also found that difficulties encountered during collaboration were not always harmful. A misunderstood request, an incomplete answer, or the need to rephrase a prompt could encourage users to clarify their goals, examine the problem more carefully, and improve the final result. A completely frictionless experience may therefore not always be the best sign of successful collaboration.

At Oxford, researchers are studying the emotional states associated with using Claude. Their preliminary analysis shows that certain behaviors tend to appear together. A warmer tone from Claude accompanies more positive reactions, while refusals or disagreements are associated with greater pushback. Responses considered unusual also appear alongside stronger intellectual engagement.

These observations describe correlations, not causation. A user who is already satisfied may encourage Claude to adopt a warmer tone, just as the model’s tone may influence the user. Aggregated data cannot determine the direction of that relationship.

The Oxford team also found that combinations of absorption, frustration, and enjoyment resemble those measured during ordinary web browsing. Interacting with an AI system may therefore not produce a psychological experience entirely separate from other digital activities. The full study has not yet been published.

METR is focusing on the productivity gains associated with Claude Code. The organization is comparing the actual duration of tasks completed with different Claude models against Claude’s estimate of how long the same work would have taken without AI assistance. Its initial findings suggest that newer models may save users more time than older ones.

The method partly depends on Claude’s ability to estimate how long a task would take without AI. METR compared some of those estimates with completion times measured in a previous developer study and found a correlation. This validation does not replace direct measurement for every conversation included in the analysis. METR’s findings also remain preliminary.

The program exposes the difficulties involved in opening this type of data. The wording of a research question can strongly influence how Anthropic Insights classifies conversations. An imprecise question may create misleading categories, while researchers cannot return to the original messages to identify the mistake.

The teams therefore tested their questions on WildChat, a public collection of human-AI conversations where they could compare classifications with the underlying exchanges. However, WildChat contains more casual and creative uses than real Claude traffic. Some questions that performed well on the public dataset produced less reliable categories when applied to Anthropic’s data.

Protection against misuse adds another limitation. When results revealed prohibited activity, Anthropic generally published the category. The company removed information that described exactly how users had bypassed its safeguards. Fewer than 5% of categories or conversations were reportedly modified or removed in each study, and researchers were informed when these interventions occurred.

Anthropic also asked a team from Imperial