Tech
Openai Regularly Publish Reports on How People Use Products Like Claude and Chatgpt, But Researchers Say These Companies Only Release

There is no independent source to corroborate it, says Anka Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab. Reuel is co-lead of a new research project called the AI Observatory, a public platform that aggregates and analyzes real AI conversations with popular models like Claude and Gemini, collected with users' consent through seven existing datasets. The project aims to provide independent information for researchers and policymakers to assess how people use generative AI.
"Stakeholders are currently making highly consequential decisions about AI's benefits and risks based on very limited data, Reuel says. The AI Observatory found that AI use differs significantly across models and has changed over time, revealing many more sensitive behaviors than are captured in reports from major AI companies, which tend to focus on work rather than personal use. Anthropic's Economic Index, one of the most widely cited sources of AI usage data, focuses on work- and productivity-related uses of Claude, filtering out other conversations.
When the Observatory team applied Anthropic's methods to their dataset, they found that nearly half of the conversations—48%—would have been filtered out. Those filtered conversations were more likely to include health and relationships (44.2% versus 31.2% in Anthropic's analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). OpenAI's 2025 report on ChatGPT use similarly found that only 30% of consumer use was related to work.
David Widder, an assistant professor at UT-Austin's School of Information who researches how people interact with AI systems and is not involved with the AI Observatory, says having the Observatory's "bird's eye view analysis rather than sectioned off into a separate report" helps researchers understand the different uses more consistently. The datasets the Observatory examined include conversations from 2023 to 2025, finding differences in how people used AI and how platforms responded. Conversations within WildChat, one of the largest datasets included, got longer and more elaborate over time, with increases in prompt tokens, response tokens, and conversation turns.
There was also significantly more small talk over time, suggesting AI companionship was increasing, while the assistants' self-disclosure that they were chatbots decreased. Exchanges labeled as sensitive use—including sexual harassment and hate speech—dropped, suggesting platforms were deploying more effective safeguards. AI use also looked significantly different depending on the model.
Users ranged in topics, interaction styles, conversation structures, and the likelihood and type of sensitive use cases. For example, people used Grok and Gemini more frequently for information retrieval, with Grok especially popular for news and politics but also where misinformation tended to concentrate. (xAI did not respond to a request for comment.) Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance.
There were even differences among versions of the same model: people had shorter conversations with ChatGPT powered by GPT-3.5 and longer, more iterative ones with GPT-4o. "No single company report tells the whole story, says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel. Company reports, however, didn't tend to capture these nuances across or even within their own models.
Source: MIT Tech Review
Most read in this category
Loading article…