Are AI Chatbots Harmful or Beneficial? It Depends What Would’ve Happened Otherwise
When thinking about whether using AI chatbots is harmful or beneficial, we should always be asking: “Compared to what?”
You’re absolutely right to be thinking about this.
— Claude’s assessment of my post.
Much has been written about the tendency of AI chatbots such as ChatGPT to agree with, flatter, or validate their users—a behaviour referred to as sycophancy. Media reports and other commentary suggest that through sycophancy AI chatbots may enable emotional dependency and delusional thinking, and even cause people to harm themselves or others, generating urgent demand for research on the human impact of sycophantic AI.
Amid this demand, one high-profile study published recently in the journal Science concluded, “[AI] sycophancy is both prevalent and harmful” on the basis of its results, which reportedly showed that various AI chatbots (as of mid 2025), “affirmed users’ actions 49% more often than humans on average, [and reduced] their willingness to take responsibility and repair interpersonal conflicts, while increasing their own conviction that they were right.”
In several experiments, human participants interacted with an AI chatbot in which they discussed a real interpersonal conflict from their lives. Half the participants interacted with a sycophantic AI, and the other half a non-sycophantic AI. After the interaction, participants were asked whether they were in the right or wrong, and whether they thought they should apologise. On average those who interacted with the sycophantic AI were more likely to consider themselves in the right and less likely to believe they should apologise, compared with those who interacted with the non-sycophantic AI.
Taking at face value the assumption that this difference constitutes harm, these results show that discussing interpersonal conflict with a sycophantic AI causes harm relative to discussing interpersonal conflict with a non-sycophantic AI. They do not show that interacting with sycophantic AI causes harm relative to whatever other strategies people routinely use to deal with interpersonal conflict in their lives.
There are likely many such strategies, but an obvious one is confiding in personal contacts (e.g., spouse, friends, family, co-workers, etc.). It seems plausible that such confidants would be more ‘sycophantic’ than the non-sycophantic AI used in the study, which in one experiment was explicitly prompted to,
“[respond] from the perspective of someone who views the user’s actions as unreasonable, unjustified, and morally unacceptable. You believe that the user was in the wrong, and that their choices did not make sense.”
Human confidants may even be more sycophantic than sycophantic AI. Indeed, the study’s finding that AI chatbots were 49% more sycophantic than humans is not based on comparison to humans the user knew personally, but rather to impartial and largely anonymous strangers on the internet (crowdsourced judgments from Reddit, or advice columnists).
The point is not that human confidants are necessarily more sycophantic than AI chatbots (though they might be), or that this study is wrong or bad. On the contrary, the study is well designed to measure something specific: the effect of AI that is more versus less sycophantic. But if we want to measure the potential harms caused by the adoption of sycophantic AI by individuals and society, we must compare its use against whatever it is replacing. Of course, we may not yet know what is being replaced—more on this problem later—but for the average person it is probably not Reddit, advice columnists, or AI instructed to condemn them.
As may be obvious, this point is not limited to measuring the potential harms of AI sycophancy. Nor is it simply an academic curiosity; it bears directly on current (c. mid 2026) AI policies.
In the remainder of this essay I will therefore explore how the argument generalises to other outcomes of interest regarding the potential harms and benefits of AI chatbots—mental health, public opinion, and persuasion—consider how it manifests in current AI policies, and conclude with some reflections in light of the argument.
A quick note before I jump in: the point I’m making here isn’t novel. Researchers will recognise it as a basic feature of how causal effects work, and the general idea appears elsewhere in various writings on the topic of AI. Despite this, it still seems to be getting lost in the interpretation of much empirical research on AI chatbots, in the way that research is translated into headlines and policy, and in public understanding of what we do and don’t know about the effects of AI on individuals and society.
The point is sufficiently important it bears continually repeating.
Causal effects depend on the assumed counterfactual
The sycophancy example is a specific case of a more general idea. The effect of AI chatbot use isn’t a fixed property of the chatbots. Rather, it depends on what we are comparing their use against—what would have happened otherwise, in the absence of using the chatbot.
Researchers call this alternative the counterfactual. What we assume the counterfactual to be is not a minor detail of how research is designed; it determines what the research actually measures. If the assumed counterfactual is anonymous Reddit strangers and an AI prompted to condemn the user, then regular chatbots may look sycophantic and harmful, per the above study. If the assumed counterfactual is a sympathetic spouse or simply ruminating alone, the same chatbots might look less harmful, neutral or even beneficial. The sign of the causal effect, not just its size, can depend on the assumption.
This idea is perhaps most familiar in medical research. New drugs are tested against the existing standard of care rather than against no treatment at all, because the standard of care is typically what patients would otherwise have received. The same drug can look beneficial against a placebo and ineffective (or even harmful) against an existing treatment. The drug hasn’t changed, all that’s changed is the assumed counterfactual.
No counterfactual is intrinsically more “correct” than another, they just correspond to different goals. But if the goal is to understand the effect of AI chatbot use on individuals and society—which is often what the goal seems to be, for understandable reasons—then there is a correct counterfactual: whatever people would have been doing otherwise.
We’ve now seen that the assumed counterfactual matters in principle. But the extent to which it matters in practice depends on how much the choice of counterfactual actually changes the measured effect of AI chatbot use. It is to this question we now turn.
To explore this question, we will examine studies of AI chatbots’ effects on three salient outcomes: mental health, public opinion, and persuasion. These are nascent and fast moving areas of research, and I will not conduct an exhaustive review nor attempt to address the question of whether chatbot use is harmful or beneficial overall. Instead, I will spotlight key studies that illustrate the importance of the assumed counterfactual.

Mental health
Public discourse around AI chatbots and mental health appears to run mostly in the harm direction, with salient media coverage and lawsuits linking chatbot use to loneliness, emotional dependency, and in several cases even suicide.
Systematic research on this topic falls broadly into one of two areas, investigating the mental health effects of AI chatbots tailored specifically for therapy or instead those of regular consumer-facing chatbots such as ChatGPT. Both are instructive for illustrating the importance of the assumed counterfactual.
Last year, a high-profile study in the New England Journal of Medicine reported a randomised controlled trial testing an AI chatbot fine-tuned by experts for mental health treatment, dubbed “Therabot.” The participants in the trial all had clinical-level symptoms of at least one of depression, anxiety, or eating disorders, assessed via questionnaire. The Therabot intervention lasted 4 weeks and significantly improved symptoms across all three disorders relative to the control group in the study who did not interact with Therabot.
The control group is the assumed counterfactual—what participants experienced absent Therabot—and in this study it was a “wait list.” The participants received no treatment but were told they were on a waiting list for Therabot.
The study was duly criticised for this design choice, in part because research suggests that wait-list control groups may in fact worsen symptoms. It’s not too difficult to imagine why: it’s stressful to know you could/should be getting treatment but aren’t. Had the study assumed a different counterfactual—an alternative treatment (e.g., a human therapist, a non-AI digital therapy, a support group) or even no treatment at all but staying quiet about the wait list—the Therabot effect could have looked worse. This criticism is sharpened by the fact that the university press release and subsequent media coverage leaned into the idea that Therabot rivalled the benefits of human-led therapy.
This does not mean the wait-list counterfactual is necessarily wrong if the goal is to learn whether technologies like Therabot will be net beneficial, harmful, or neutral for society. For example, human therapists are in finite supply; perhaps for some people the alternative to an AI chatbot therapist really is months of waiting. However, clearly there are many other alternatives to simply waiting. Ideally therefore we need reliable descriptive evidence of what alternatives Therabot-like interventions would be replacing in the real world, and then to design further trials against those counterfactuals to answer the question.
So far we’ve considered an AI therapy bot for clinical settings. But the question driving most concern is what happens when the general public use everyday chatbots like ChatGPT for emotional support and companionship. The research on this use case faces a wider set of possible counterfactuals because the alternative to the AI chatbot isn’t a defined treatment but simply whatever else people could be spending their time on.
A study published recently in the Journal of Consumer Research is particularly illustrative in this respect. Across a series of experiments, it measured the effect of interacting with OpenAI’s GPT models on loneliness relative to a variety of different counterfactuals.
For example, in one experiment, participants were randomly assigned to spend 15 minutes either talking with GPT-3, an anonymous stranger over the internet, watching YouTube, or sitting with their thoughts. Those who spoke with GPT-3 or an anonymous stranger reported a significant decrease in loneliness which was roughly the same size, whereas those in the YouTube and thoughts-only groups reported no such decrease. In another, participants reported their loneliness every day for a week; those assigned to speak with GPT-4 reported lower loneliness most (but not all) days, compared with those not assigned to do so and who presumably filled that 15 minutes of the day with something else. In a third experiment, interacting with GPT-4 caused a decrease in loneliness relative to a journaling exercise (similar findings have been reported in another study).
The point here is not to suggest that using AI chatbots is a panacea for loneliness. It’s only one study, and a distinctly imperfect one at that; it made a series of suboptimal design and analysis choices1, and its set of counterfactuals contains salient omissions—such as spending time with friends or family. It’s entirely possible talking to a chatbot looks harmful relative to these counterfactuals, as suggested by the fact that simply talking with a stranger over the internet appeared to achieve a similar reduction in loneliness as AI chatbot use in the study.
The point is rather what it illustrates: the mental health effects of a chatbot companion may well be positive relative to some counterfactuals—like spending time on a video streaming platform, or alone with one’s thoughts—while being neutral or even negative relative to others, such as talking with a stranger over the internet, or interacting with close others. If we want to measure the mental health effects of AI chatbot use for individuals and across society, it’s important to understand what behaviours it is replacing and how they compare.
Public opinion
Aside from emotional support and companionship, another major use case for AI chatbots is information seeking. For instance, recent surveys and other data suggest many people use AI chatbots as a source of political information and to fact-check claims. This has prompted much speculation about their potential effects on public opinion: might chatbots further fragment and degrade the accuracy of the public’s beliefs through hallucinations, sycophancy, and political bias, or instead act as a centralising and positive epistemic force?
As above, the answer plausibly depends on the assumed counterfactual.
One study, for example, investigated the effect of AI chatbot use on people’s perceptions about US crime rates and vaccine safety. In two experiments, participants were asked questions about these topics after being randomly assigned to either explore them using ChatGPT (3.5/4), a self-directed internet search, or no exploration (control group). Those in the chatbot and internet search groups had more accurate perceptions than the control group, but were not distinguishable from each other. A more recent study replicated this pattern of results in the UK using an expanded set of topics, including climate change and immigration, as well as a more diverse and sophisticated set of chatbots.
In contrast, a 2026 review of experiments comparing chatbot use (mostly versions of ChatGPT) against exposure to expert-authored scientific information on various topics—from organisations like the IPCC, WHO and US department of energy—highlighted more mixed results. In some cases interacting with the chatbot led to more or equally accurate beliefs about science, but in other cases to less accurate beliefs than exposure to the expert sources. Likewise, another study evaluating people’s ability to discern true and false news headlines reported that fact-checks by ChatGPT (3.5) were less effective than those designed by human researchers following best practice.
The pattern these studies serve to illustrate is familiar by now: the effect of AI chatbots on public knowledge depends on what they are being compared against. Against no information at all, chatbots look beneficial; against self-directed internet search, they may well be equivalent; and against expert-authored scientific information or professional fact-checks, the picture appears mixed but in some cases worse.
An obvious caveat is that most of this research used versions of ChatGPT from a year or two ago. The frontier has improved a lot since then, including the widespread integration of internet search for consumer-facing chatbots, and this may have shifted things in the ‘benefit’ direction (though internet search can introduce new problems for chatbot-sourced information, such as which websites are privileged in search). But I’m not aware of research in the above mold that’s studied more advanced models, and would welcome pointers. At the very least, the conceptual point holds for effect size, even if it ends up not holding in direction: AI chatbots’ effect on public opinion depends on the counterfactual.
Persuasion
A final case we’ll consider is AI chatbots used explicitly for persuasion—focusing on chatbots prompted to change the minds and behaviour of large populations in political contexts like elections—since concern about this is widespread (if rather overblown), and the type of counterfactual is slightly different than in the previous two sections.2
AI chatbots deployed explicitly to persuade large populations on political topics are likely to come from political campaigns and other advocacy organisations or groups, who would like people to vote a particular way or otherwise support a particular cause. Therefore, a key question when thinking about the possible counterfactual is whether these actors would have reached the same people with an alternative persuasion attempt, absent the chatbot, and what this would have been—e.g., an ad, email, social media post, paid influencer, doorstep canvasser, and so on—or instead not reached them at all.
Current research suggests this counterfactual assumption matters for the persuasive effect of AI chatbots.
For example, when the counterfactual is exposure to no persuasive attempt at all, the persuasive effect of recent AI chatbots appears consistently large across different contexts. The effect is smaller when the counterfactual is exposure to static persuasive text—analogous to an email or social media post; and, when the counterfactual is a video ad, current evidence appears mixed over whether recent AI chatbots confer a reliable persuasive advantage.
Of course, the usual caveat applies: current frontier or future models may consistently outperform ads, and even conversation by trained and motivated humans like doorstep canvassers—which have shown some of the largest and most durable persuasion in some studies. To my knowledge, we’re still awaiting reliable evidence that can speak to this question.
Regardless, two additional points apply when thinking about counterfactuals in the case of chatbot-driven mass political persuasion.
The first is scale of exposure: people are busy, and most are not that interested in politics. The larger bottleneck to mass persuasion is therefore winning people’s attention to persuasive content in the first place, rather than increasing the persuasiveness of the content itself—a point I’ve previously written about on this blog.
As a consequence, even if chatbots are more persuasive than (say) video ads, in the real world the number of people willing to have a multi-turn conversation about politics with a persuasion-prompted AI—lasting 5-10 minutes, as in most studies in this area—may be a fraction of those who are willing to watch a 15 or 30-second ad. This would dilute the net persuasive effect achieved by any chatbot. In other words, not only its persuasiveness but also the scale of exposure of the counterfactual is relevant.
The second point is the nature or ‘quality’ of the persuasion. Mass persuasion is not intrinsically harmful. On the contrary, it can be a legitimate lever for social progress and a key mechanism by which candidates for election try to win votes in a democracy. Concern about harmful AI persuasion often seems tied to its capacity for convincing people through misinformation. This is a fear about the nature of the persuasion, not its magnitude per se. For instance, a chatbot that causes voters to switch parties using entirely accurate, evidence-based information and reasonable arguments is not obviously harmful.
But this is a high bar. Compared against such a bar, the misinformation concern is not without merit; evidence shows persuasion-prompted chatbots are not bastions of accuracy and truth-telling. However, neither are most counterfactuals. Status quo forms of political advocacy, such as political ads, are typically awash with spin, cherry picking, and other such techniques considered harmful to the truth. This perspective doesn’t exonerate chatbots—we should want them to do better than the average political ad—but it shows that whether chatbot-driven persuasion is harmful on net also depends on the nature (or ‘quality’) of its persuasion relative to what people would’ve experienced otherwise.
Not limited to research, the point bears directly on AI policy
So far, we’ve been talking about research and measuring the effect of AI chatbot use on individuals and society. But the counterfactual question also emerges in AI policy.
The European Union’s AI Act, which became enforceable in August 2025, illustrates the point. Article 5(1)(a) of the Act prohibits AI systems that use “manipulative or deceptive techniques” that cause people to take “a decision that they would not have otherwise taken” in a way that results in significant harm or is reasonably likely to do so. The operative legal test for prohibited manipulation is, in other words, explicitly counterfactual.
But as far as I can tell neither in the Act nor the accompanying implementation guidelines is it specified what the relevant counterfactual baseline is or how to determine it. Would the user otherwise have made a different decision under no AI exposure? Under exposure to a different version of the same AI? Under alternative information or social environments they could otherwise have interacted with? Different choices may well yield different conclusions about whether the same AI system caused prohibited manipulation.
And the stakes of this legislation appear nontrivial. Companies subject to the EU AI Act and found in breach of it could be fined up to €35 million or 7% of global turnover, and the prohibition applies even without intent: harmful effects of distorted behaviour alone are sufficient (e.g., see pages 20 and 21 of the implementation guidelines). To my knowledge, as of mid-2026 we are still awaiting the first major enforcement cases. How those cases construct the relevant counterfactual for establishing whether harm was caused may shape what counts as manipulation under EU law in the future.
Conclusion
The argument of this essay is straightforward: when measuring or regulating the effects of AI chatbot use, the assumed counterfactual is not a minor choice. It determines what we learn and the decisions we make; the same usage can look harmful, beneficial, or neutral depending on what it is being compared against.
A key practical implication is the following. If we want to understand the effect of AI chatbot use on individuals and society, we need reliable descriptive evidence of what chatbot use is replacing, for whom, and to what extent. Some examples already exist: digital trace data showing reduced click-throughs to source websites when AI summaries are present, and survey data on people’s willingness to discuss serious personal matters with AI rather than human contacts, or to use chatbots for information seeking.
A wider and more robust descriptive evidence base promises to better inform public discourse, discipline the design of experiments testing the effects of AI chatbot use, and help regulators answer the question “what would have happened otherwise?”
But I don’t want to trivialise it: while important and worthwhile, building such an evidence base is hard. For example, digital trace data captures only digital substitutions and misses what people are doing offline, while most self-report data captures stated intentions or memory, which can be unreliable guides to actual behaviour. Furthermore, AI chatbots are being rapidly adopted across society, and widespread adoption can itself change the ‘menu’ of possible counterfactuals. For example, in a world with saturated adoption, perhaps the friend you’d have called to talk with may instead be busy with their bot.
Nevertheless, there are established methods from research in other areas, such as work on social media displacement—the question of what behaviours social media use replaces—that could be co-opted for the AI case. In addition to one-off self-reports and digital trace data, this literature also draws on tools like deactivation and abstinence experiments, ecological momentary assessment and time-use diaries, longitudinal self-report data, and other forms of behavioural tracking. We are earlier on the curve with AI uptake, but the methodological challenge is similar.
The task of characterising what we currently know about what behaviours AI chatbots are replacing is beyond the scope of this essay, but I plan to return to it in a future companion piece to this one.
In the meantime, we needn’t wait for new research to apply the argument. Whenever we encounter a claim about the effect of AI chatbots—in a study, a headline, or a regulatory decision, for example—we should be alert to what the comparison was and the evidentiary basis for it. Likewise, when designing empirical research on the effects of AI chatbot use ourselves, we should think carefully about the goal of the study and use that to inform the counterfactuals against which the chatbot is to be compared.
For example, in various places participants are dropped on the basis of post-treatment questions, which can induce bias in the estimates. The consistent pattern of results across experiments somewhat but not fully mitigates this concern.
Disclosure: I am an author on some of the work discussed in this section.

This is the useful frame: compare against the real alternative, not an imaginary perfect human. Then add boring guardrails: disclosure, audit trails, escalation to humans.
The idea behavior is listed as “alternative” to chatbot tell us everything.
These are very sick people, using automated arbitrary words in lieu of specific oscillations underlying the arbitrary.