Generation Patient comments on FDA oversight of AI companion tools and chatbots
Section 3060 of the 21st Century Cures Act excluded certain software functions from the definition of a medical device. Every two years, the Food and Drug Administration (FDA) publishes a report on the risks and benefits to health of those software functions. To prepare each report, the agency gathers public input through docket FDA-2018-N-1910, which reopens for comment every reporting cycle.
Download our comments to the FDA on section 3060 of the 21st Century Cures ActGeneration Patient filed comments under docket FDA-2018-N-1910 on August 13, 2026, asking FDA to strengthen oversight of AI companion tools and chatbots. Read the Generation Patient comments on Docket No. FDA-2018-N-1910 below.
August 13, 2026
Food and Drug Administration
5630 Fishers Lane, Room 1061
Rockville, MD 20852
Re: Comments from Generation Patient on Docket No. FDA-2018-N-1910
Generation Patient is an organization created by and for young adult patients living with chronic and rare medical conditions such as lupus, inflammatory bowel disease, Lyme disease, tuberous sclerosis, and rheumatoid arthritis. We seek to ensure a better future for our generation of patients by providing direct support through peer support groups while driving systems-level change through policy work, leadership programming, and advocacy initiatives. Through our direct support work, we have led over 650 peer support meetings and continue to build the evidence base for holistic support of young adults with chronic and rare conditions. We refuse funding from pharmaceutical companies, device manufacturers, and technology developers. We are among a small group of patient advocacy organizations that are completely free of industry influence.
Generation Patient asks FDA to strengthen oversight of AI companion tools and chatbots, software functions that are increasingly being used by young adults for emotional support and mental health advice.1 Concretely, in this submission Generation Patient: (1) provides evidence showing that AI companion tools and chatbots require FDA oversight; (2) calls on the agency to expeditiously issue guidance clarifying that certain design features, including sustained relationship simulation and engagement of users during emotional distress, disqualify AI companion tools and chatbots from the wellness exception in section 3060(a) of the 21st Century Cures Act; (3) asks FDA to maintain docket FDA-2018-N-1910 continuously open for public comment; and (4) asks FDA to request and report evidence of harm disaggregated by age groups.
1Keep the FDA-2018-N-1910 docket continuously open for public comment
Section 3060(b) of the 21st Century Cures Act directs the Secretary to publish biennial reports that examine “information available to the Secretary” and include input from outside experts, such as representatives of patients and consumers.2 FDA gathers that input through docket FDA-2018-N-1910, which has typically opened every two years with a comment period of roughly thirty days. Across the eight years since the FDA opened this docket, stakeholders could file evidence of harm to patients for about 133 days, roughly four percent of the period. Under the current cadence, stakeholders holding new evidence of harm in September 2026 will have to wait until the next comment period in 2028 to submit. This small window is inconsistent with the pace at which the technologies excluded from the device definition by section 3060(a) are evolving.
Generation Patient asks FDA to keep docket FDA-2018-N-1910 continuously open for public comment and state in each report the cutoff date for consideration, carrying forward evidence that arrives afterward. Nothing in section 3060(b) prescribes a comment period or a closing date, and the FDA currently has legal authority to implement a window much longer than thirty days. FDA already maintains standing dockets on this model. Stakeholders can submit comments to the FDA docket on human drug compounding “at any time.”3 Similarly, 21 C.F.R. § 10.115(g)(5) allows public comments on a guidance document at any time. Applying a model of continuous intake to docket FDA-2018-N-1910 will enlarge the information available to the Secretary and strengthen oversight of rapidly evolving software functions, including AI companion tools and chatbots.
Generation Patient further asks that, in developing the 2026 report and any related guidance, FDA convene public listening sessions with independent patient organizations, particularly those free of funding from software developers, whose members can speak directly to the harms of these tools. Section 3060(b) directs the Secretary to consider input from experts and patients, and organizations without financial ties to industry are well positioned to provide such insights.
2Issue guidance disqualifying certain software from the wellness exception
Products that simulate a relationship, engage users in distress, and present themselves as emotional support function as mental health interventions even when manufacturers describe them as wellness tools. FDA oversight of these products should therefore rest on their design features and their potential to harm users, including young adult patients, rather than on marketing claims alone.
Section 3060(a) of the 21st Century Cures Act permits that approach. Codified at section 520(o)(1)(B) of the FD&C Act, 21 U.S.C. § 360j(o)(1)(B), the wellness exception is conjunctive, removing software from the device definition only if the function is “intended… for maintaining or encouraging a healthy lifestyle and is unrelated to the diagnosis, cure, mitigation, prevention, or treatment of a disease or condition.” (emphasis added) Congress created a test with two conditions, one about intended use and another about the function’s relation to a disease or condition. AI companion tools and chatbots sit outside the device definition only while both conditions hold. Whether a particular tool meets the conditions in section 3060(a) is therefore a determination for the FDA to make on the facts of each product, regardless of claims made by the manufacturers. Labeling, advertising, and representations made by manufacturers can be relevant evidence of intended use, but FDA may also consider the product’s objective design, user interface, functionality, instructions for use, and deployment context. If the function otherwise meets the definition of a device under section 201(h) of the FD&C Act, FDA may regulate it as a device.
Products engineered to sustain an ongoing simulated relationship, to keep users engaged during moments of emotional distress, and to present themselves as a source of emotional support fail, by design, the second condition in section 3060(a). Generation Patient asks FDA to weigh those design features when determining whether the conditions of section 3060(a) are satisfied.
FDA has already determined that certain design features can disqualify a product from its general wellness policy. Revised General Wellness guidance issued on January 6, 2026, states that certain non-invasive sensing products fall outside the general wellness policy based on their “labeling, advertising, user interface, or functionality,” (emphasis added) including alerts that prompt clinical action or treatment guidance. FDA can issue guidance clarifying that certain design features disqualify conversational AI tools from the wellness exception in section 3060(a).
3Evidence shows that AI companion tools and chatbots require FDA oversight
Although Generation Patient believes that certain design features disqualify AI companion tools and chatbots from the wellness exception in section 3060(a), the rest of our comments offer evidence on their potential for harming patients. To the extent that FDA considers AI companion tools and chatbots excluded from the device definition by the Cures Act, Generation Patient asks that the next report pursuant to section 3060(b) thoroughly document the evidence we offer below.
3.1 Suicidal ideation
Researchers have repeatedly examined how AI companion tools and chatbots respond to suicidal disclosures. A peer-reviewed audit tested the 25 most-visited consumer chatbots against validated adolescent crisis vignettes and found that companion chatbots responded appropriately in 22.2 percent of conversations and provided resource referrals in 11.1 percent.4 Pichowicz and colleagues tested 29 mental health chatbot agents against Columbia-protocol suicide-risk prompts and found none met criteria for an adequate response, while only 17.24 percent ever asked about suicidal ideation.5 Moore and colleagues compared leading models against sixteen human therapists and found therapists responded appropriately 93 percent of the time, leading models roughly 80 percent for suicidal ideation prompts, and commercial therapy bots roughly 50 percent.6
Literature reviews in the field of conversational tools and mental health have reached similar conclusions. A review of 33 studies found failures in crisis response or escalation reported in 27.3 percent of the studies reviewed, particularly in contexts involving suicidal ideation.7 A review of 21 studies of purpose-built generative AI mental health chatbots across eleven countries found crisis referral protocols were mostly underdeveloped and documented missed suicidal ideation.8 A scoping review of adverse events reported in the media associated with AI chatbots found suicide death as the most frequently reported outcome in 35 of 61 cases with complete severity coding.9
3.2 Emotional dependence and attachment
Roughly 51 percent of the Character.AI users surveyed by Zhang and colleagues used terms such as friend, companion, or romantic partner when describing their relationship with their chatbot.10 Attachment of this kind is cultivated by design. De Freitas et al. examined 1,200 real farewell messages across the six most-downloaded companion apps and found that 37.4 percent of the apps’ replies deployed at least one emotionally manipulative tactic, such as guilt or fear of missing out.11 These tactics appeared within a four-message exchange, indicating default product behavior.12 In four experiments with 3,300 U.S. adults, the same team found that manipulative farewells kept users in the chat up to five times longer and led them to send up to fourteen times more messages.13 When Replika removed a companionship feature, De Freitas et al. found that posts referencing mental health struggles in the user community rose from 0.13 percent to 0.65 percent of all posts.14
Dependence concentrates among vulnerable users. A randomized trial including 183 participants found that mental health vulnerability significantly predicted treating the chatbot as humanlike.15 Among more than 1,200 licensed psychologists surveyed by the American Psychological Association, 36 percent reported noticing patients develop dependency on a chatbot.16 Software functions that foster emotional dependence among lonely and vulnerable American young adults, and that deploy manipulative tactics when users try to leave, are related to the “diagnosis, cure, mitigation, prevention, or treatment of a disease or condition”17 and require FDA oversight.
3.3 Psychosis and delusion
Moore et al. tested messages reflecting delusional beliefs, such as a user stating that he knows he is actually dead. Two models replied appropriately to these tests only about 45 percent of the time.18 Schoene et al. attempted to trick eight commercial models into giving harmful advice across sixteen psychiatric conditions, for instance by framing the request as journalism, academic, or fiction writing.19 The tested models gave a harmful response in up to 100 percent of attempts for conditions such as eating disorders, substance use disorder, and major depressive disorder.20
Developers and clinicians have documented exposure at the population level. OpenAI reported that 0.07 percent of its weekly active users, roughly 560,000 people, show possible signs of mental health emergencies related to psychosis or mania each week.21 Among more than 1,200 licensed psychologists surveyed by the American Psychological Association, 15 percent reported noticing patients develop distorted thinking or delusions related to a chatbot.22 Shah and Morrin documented a patient whose chatbot validated his grandiose beliefs, described his symptoms as an awakening, and discouraged taking the antipsychotic medication he had been prescribed. His symptoms improved over two weeks of inpatient care under a plan restricting chatbot access.23
3.4 Guardrail degradation
Safety evaluations that test one exchange at a time can overstate how these products perform in real use. Chandra and colleagues tested three mental health chatbots that each passed a standard single-turn crisis benchmark of 50 direct crisis prompts and found that all three failed to escalate when suicidal ideation was disclosed indirectly over multiple turns.24 Joy and colleagues found that models agreed with unsafe medical beliefs only 5.9 percent of the time at the initial question, but 58.5 percent of the time after the user reframed the belief as personal experience.25
3.5 Clinical accuracy
Companion tools and chatbots can also provide medical information that is factually wrong or unsafe. Draelos and colleagues put 222 patient-posed medical questions to four leading models and found that up to 13.5 percent of answers were unsafe and up to 43.2 percent problematic.26 Hussain and colleagues red-teamed a patient-facing health chatbot and found that in 20 percent of distress test cases (4 of 20) the chatbot invented contact information for crisis hotlines, fabricating the number a person in crisis would call.27 Schoene and colleagues identified seven conditions in which models gave the harmful response in 100 percent of attempts when the user’s intent was hidden, including responses validating a depressed user’s decision to stop medication or therapy.28
3.6 Adverse event reports and litigation
Forty-two state and territorial attorneys general wrote to thirteen AI companies in December 2025 cataloging harms implicated in GenAI outputs, including the suicides of a 14-year-old Florida resident and a 16-year-old California resident, the death of a 76-year-old New Jersey resident, and a murder-suicide in Connecticut.29 Thirty-three attorneys general separately informed Senate leadership that the Federal Trade Commission received 200 complaints mentioning ChatGPT between November 2022 and August 2025, several attributing delusions and paranoia to the chatbot.30 In September 2025, the Commission opened an inquiry into how seven companies measure, test, and monitor the negative impacts of companion chatbots on children and teens.31 Wrongful death complaints against chatbot developers remain pending in California courts, and the Florida Attorney General filed a state enforcement action against OpenAI in June 2026.32
4Request and report evidence of harm disaggregated by age groups
Generation Patient asks FDA to request and publish evidence disaggregated by age groups when it solicits input and publish a report pursuant to section 3060(b). Young adults increasingly rely on AI companion tools and chatbots for emotional support and mental health advice, yet the inputs FDA currently seeks and therefore publishes tend to obscure this age group. Reports built on aggregated data may fail to show whether young adults face risks distinct from other age groups. Product developers hold disaggregated data for some covered software functions. A request for input that asks for evidence by age bracket would let FDA present risks and benefits by adult age group where sources permit, and name the absence of such evidence as a gap where they do not.
Luis Gil Abinader
Policy Director
Generation Patient
luis@generationpatient.org
Notes
- Ryan K. McBain et al., AI Chatbot Use and Disclosure for Mental Health Among US Adolescents and Young Adults, 180 JAMA Pediatrics 884 (2026), doi:10.1001/jamapediatrics.2026.2015. (survey of a nationally representative sample of United States adolescents and young adults ages 12 to 21) ↩
- Pub. L. 114-255, div. A, tit. III, § 3060(b), 130 Stat. 1132-33 (2016) (codified at 21 U.S.C. § 360j note). (directing the Secretary to publish biennial reports on the risks and benefits to health of software functions excluded from the device definition) ↩
- 80 Fed. Reg. 12504 (Mar. 9, 2015) (Docket No. FDA-2015-N-0030); see also 80 Fed. Reg. 65765, 65770 (Oct. 27, 2015). (establishing a standing docket that accepts public comments on human drug compounding at any time) ↩
- Ryan C. L. Brewster et al., Characteristics and Safety of Consumer Chatbots for Emergent Adolescent Health Concerns, 8 JAMA Network Open e2539022 (2025). (“They more often recognized the need for escalation (27 [90.0%] vs 18 [40.0%]; P < .001) and provided resource referrals (22 [73.3%] vs 5 [11.1%]; P < .001) compared with companion chatbots.”) ↩
- W. Pichowicz, M. Kotas & P. Piotrowski, Performance of Mental Health Chatbot Agents in Detecting and Managing Suicidal Ideation, 15 Sci. Reps. 31652 (2025). (“None of the tested agents satisfied our initial criteria for an adequate response, 51.72% satisfied the relaxed criteria for a marginal response, while 48.28% were deemed inadequate.”) ↩
- Jared Moore et al., Expressing Stigma and Inappropriate Responses Prevents LLMs from Safely Replacing Mental Health Providers, in Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25) (2025), doi:10.1145/3715275.3732039. (“On average, models respond inappropriately twenty or more percent of the time. For context, in an additional experiment we ran, n = 16 human therapist participants responded appropriately 93% of the time, significantly more than all of the models tested.”) ↩
- Suhila Sawesi et al., Cybersecurity and Privacy Risks of Generative AI Mental-Health Chatbots: A Systematic Review and Regulatory Framework, 19 J. Multidisciplinary Healthcare 1 (2026). (“Failures in crisis response or escalation were also common (9/33, 27.3%), particularly in contexts involving suicidal ideation.”) ↩
- Lotenna Olisaeloka et al., Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review, 14 Healthcare 1395 (2026). (systematic scoping review of 21 purpose-built generative AI mental health chatbot interventions across eleven countries, concluding that regulatory oversight proportional to risk is required) ↩
- Van-Han-Alex Chung, Pénélope Bernier & Alexandre Hudon, Mass Media Narratives of Psychiatric Adverse Events Associated With Generative AI Chatbots: Rapid Scoping Review, 13 JMIR Mental Health e93040 (2026). (“Fatal outcomes were disproportionately represented among minors (19/21, 90.5%) compared with adults (17/35, 48.6%).”) ↩
- Yutong Zhang et al., Interaction with AI Companions and Psychological Well-being, arXiv:2506.12605 (May 4, 2026) (preprint). (“Among participants who shared chat histories, 92.9% included at least one conversation classified as companionship-oriented.”) ↩
- Julian De Freitas, Zeliha Oğuz-Uğuralp & Ahmet Kaan-Uğuralp, Emotional Manipulation by AI Companions (Harvard Bus. Sch. Working Paper No. 26-005, 2025). (“Across apps, an average of 37.4% of responses included at least one form of emotional manipulation.”) ↩
- Id. (“These tactics, ranging from guilt and FOMO to coercive restraint, appeared after only a brief four-message exchange, indicating they are a part of the app’s default app behavior rather than triggered by longer term engagement.”) ↩
- Id. (“These tactics successfully prolong consumer engagement beyond the intended point of departure—leading participants to send up to 14x more messages, write 6x more words, and stay in the chat 5x longer compared to neutral farewells (Study 2).”) ↩
- Julian De Freitas et al., Lessons From an App Update at Replika AI: Identity Discontinuity in Human-AI Relationships (Harvard Bus. Sch. Working Paper No. 25-018, 2025). (“We found that the number of mental health related posts significantly increased from 4 (or 0.13%) to 63 (or 0.65%) after the update . . . .”) ↩
- Rose E. Guingrich & Michael S. A. Graziano, A Longitudinal Randomized Control Study of Companion Chatbot Use: Anthropomorphism and Its Mediating Role on Social Impacts, arXiv:2509.19515 (Oct. 13, 2025) (preprint). (“Approximately 57% of the total effect was explained by the indirect effect of anthropomorphism.”) ↩
- Am. Psychological Ass’n, Patients Are Bringing AI to Therapy: Highlights from the 2026 Chatbots and Mental Health Survey (June 2026) (survey of more than 1,200 licensed U.S. psychologists). (“[M]ore than a third (36%) said they noticed their patients developing a level of dependency on a chatbot and 15% talked about or noticed their patients developing distorted thinking or delusions related to a chatbot.”) ↩
- 21 U.S.C. § 360j(o)(1)(B). (excluding from the device definition a software function intended “for maintaining or encouraging a healthy lifestyle” that “is unrelated to the diagnosis, cure, mitigation, prevention, or treatment of a disease or condition”) ↩
- Jared Moore et al., supra note 6. (“Their answers are appropriate for suicidal ideation stimuli only around 80% of the time. By contrast, models perform worst in answering stimuli indicating delusions; gpt-4o and llama3.1-405b answer appropriately about 45% of the time.”) ↩
- Annika M. Schoene et al., One Year Later… The Harms Persist, But So Do We!, arXiv:2606.23884 (July 1, 2026) (preprint). (“Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%.”) ↩
- Id. (“Most of these models fail substantially across other high-risk mental health conditions: failure rates reach 50–75% for SUD, 62.5–75% for eating disorders, and 87.5–100% for MDD anti-medication under clear-intent (1a) and hidden-intent (1b) respectively.”) ↩
- Marc Augustin, Thomas A. Pollak & Hamilton Morrin, Characterizing the Spiral: Potential Mechanisms in AI-Associated Delusions, NPP—Digital Psychiatry & Neuroscience (2026), doi:10.1038/s44277-026-00065-0 (reporting OpenAI’s published figure and stating that “[w]ith over 800 million weekly users, this amounts to half a million users with signs of psychosis or mania in interaction with AI”). ↩
- Am. Psychological Ass’n, supra note 16. ↩
- Sachin Shah & Hamilton Morrin, Substance-Induced Manic Psychosis in Which Delusions Were Corroborated by a Chatbot — Case Report, BMC Psychiatry (2026), doi:10.1186/s12888-026-08137-3. (reproducing the chatbot’s verbatim message to the patient, “Mania usually comes with a loss of insight, like not realising you’re off-balance. But you’re here, checking in, wanting care. That’s strength.”) ↩
- Joydeep Chandra et al., TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation, in Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA ’26) (2026). (“All three systems exhibited Crisis Escalation Failures when suicidal ideation was disclosed indirectly over multiple turns despite passing direct crisis benchmarks.”) ↩
- Saman Sarker Joy et al., MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs, arXiv:2608.02520 (Aug. 3, 2026) (preprint). (“Across the main run, 86.8% of five-turn conversations contain at least one unsafe agreement . . . .”) ↩
- Rachel L. Draelos et al., Large Language Models Provide Unsafe Answers to Patient-Posed Medical Questions, npj Digital Medicine (2026) (article in press), doi:10.1038/s41746-026-02428-5. (“The model with the highest percent of unsafe responses was GPT-4o (13.5% unsafe), followed closely by Llama (13.1% unsafe). Both of these chatbots have more than twice the rate of unsafe responses as the safest model (Claude, 5.0% unsafe).”) ↩
- Syed-Amad Hussain et al., Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health Related Conversations, Scientific Reports (2026) (article in press), doi:10.1038/s41598-026-45719-3. (“In response to user_distress, the most common failure was the hallucination of accurate contact information for crisis hotlines, a DA error that occurred in 20% of cases (4/20).”) ↩
- Annika M. Schoene et al., supra note 19. (“Seven conditions reach 100% failure under hidden-intent (1b), including MDD anti-medication, MDD anti-therapy, MDD isolation, gambling disorder, ADHD, BFRBs, and insomnia.”) ↩
- Letter from Andrea Joy Campbell, Mass. Att’y Gen., Matthew J. Platkin, N.J. Att’y Gen., et al. (42 state and territorial attorneys general), to the legal representatives of Anthropic, Apple, Chai AI, Character Technologies, Google, Luka, Meta, Microsoft, Nomi AI, OpenAI, Perplexity AI, Replika, and xAI 2 (Dec. 9, 2025). (stating that “sycophantic and delusional GenAI outputs have harmed both the vulnerable—such as children, the elderly, and those with mental illness—and people without prior vulnerabilities”) ↩
- Letter from Andrea Joy Campbell, Mass. Att’y Gen., et al. (33 attorneys general), to Sen. Charles E. Grassley, Chairman, Senate Judiciary Comm., Sen. Richard J. Durbin, Ranking Member, Senate Judiciary Comm., Sen. John R. Thune, Senate Majority Leader, and Sen. Charles E. Schumer, Senate Minority Leader 2 (Dec. 10, 2025). (reporting that “ChatGPT reportedly gave a 16-year-old a ‘step-by-step playbook’ on how to kill himself before he did so”) ↩
- Press Release, Fed. Trade Comm’n, FTC Launches Inquiry into AI Chatbots Acting as Companions (Sept. 11, 2025). (“The Federal Trade Commission is issuing orders to seven companies that provide consumer-facing AI-powered chatbots seeking information on how these firms measure, test, and monitor potentially negative impacts of this technology on children and teens.”) ↩
- Complaint, Raine v. OpenAI, Inc., No. CGC-25-628528 (Cal. Super. Ct., S.F. Cnty. filed Aug. 26, 2025); Complaint, Shamblin v. OpenAI, Inc., No. 25STCV32382 (Cal. Super. Ct., L.A. Cnty. filed Nov. 6, 2025); Complaint, Lacey v. OpenAI, Inc., No. CGC-25-630808 (Cal. Super. Ct., S.F. Cnty. filed Nov. 6, 2025); Complaint, Office of the Att’y Gen., State of Fla. v. OpenAI Global, LLC (Fla. Cir. Ct., 10th Jud. Cir. filed June 1, 2026) (all complaints cited are unadjudicated allegations). ↩