AI Can Help Heal Romantic Distress
Most mental-health chatbots ask for weeks of sessions and then lose the people who needed them. A preregistered trial of overit, a breakup chatbot built at Cambridge and TUM, found that one 20-minute conversation was enough to move distress scores, and the gap was still there a month later.
Talk to Psychologist Browse AI chatbots
What’s new
Thomas Menzel (Technical University of Munich and University of Cambridge), with Michel Schimpf and Thomas Bohné (University of Cambridge), built overit: a mobile chatbot that walks someone through a single, structured conversation about a romantic breakup. In a preregistered randomized controlled trial of 254 adults in the US and UK, people who used the app felt substantially better after that one session than people who only filled out surveys.
The paper is on arXiv as Menzel, Schimpf & Bohné, 2026. The trial was preregistered on OSF (10.17605/OSF.IO/43MJ9).
Key insight
Breakups keep hurting long after the relationship ends. Distress tracks with depressive symptoms, intrusive thoughts and sleep disruption (Field et al., 2009; Perilloux & Buss, 2008). A lot of that persistence is cognitive: people lock onto self-limiting beliefs about themselves (“I was abandoned because I am not enough”) or the world (“nobody could want me”). Negative cognitions after a breakup predict grief and depression independently of how long ago it happened (Boelen & Reijntjes, 2009).
Cognitive reappraisal is the usual lever: surface a distressing interpretation, weigh it, and replace it with one that fits the facts better (Gross, 1998; Beck et al., 1979). Breakup-specific lab work already shows that reframing the ex-partner or the relationship can cut attachment-related distress (Langeslag & Sanchez, 2018).
Memory reconsolidation theory is the stronger claim. When an emotional memory is reactivated, it becomes briefly labile. If, in that window, the person encounters an experience that contradicts what the old learning predicts, the memory can update rather than merely being argued with (Schiller et al., 2010; Nader, 2015; Sevenster, Beckers & Kindt, 2013). Clinical accounts of the Therapeutic Reconsolidation Process spell out the sequence: reactivate the painful schema, juxtapose a contradictory meaning, then let the old feeling lose its grip (Ecker, Ticic & Hulley, 2012; Ecker, 2015; Lane et al., 2015).
If that sequence is real, a chatbot that elicits a self-limiting belief and then steers the user toward a counterfactual (“You did the best you could, but you were failed by someone you trusted”) could produce a lasting drop in distress from a single sitting. That is the bet overit tests.
How it works
Participants completed a baseline survey: Breakup Distress Scale score, time since the split, attachment (ECR-S; Wei et al., 2007), and the ex-partner’s first name. Then they talked to a Flutter app whose backend called Claude Sonnet 4.5 (claude-sonnet-4-5-20250929).
The conversation has four phases, mixing the Therapeutic Reconsolidation Process with classic cognitive restructuring (Hofmann et al., 2012):
- Context gathering. Open questions about what happened and how it landed. Listen for limiting beliefs.
- Belief exploration. Identify at least one core self-limiting belief. Challenge it. The session cannot leave this phase until that belief is named.
- Counterfactual generation. Offer alternative readings. “What if…” questions, not pep talks.
- Integration and closure. Ask what shifted, how they feel now, then end cleanly.
Sessions cap at 18 turns, typically about 20 minutes, text or voice (Whisper via Groq). Phase changes are mostly turn-counted, with a hard gate that the core belief must be identified before phase three.
Each user message triggers two model calls, not one. That split is the interesting engineering:
- A generation call (temperature 1.0) produces the reply the user sees. It gets only the current phase’s instructions, plus survey data and history.
- An evaluation call (temperature 0) scores the last three turns against five binary milestones: belief identified, belief challenged, counterfactual considered, new insight articulated, natural closure reached.
LLMs get worse at following instructions as the constraint list grows (Liu et al., 2024 lost-in-the-middle; Jiang et al., 2024, FollowBench). They also sycophantically agree with the user, which is the wrong bias for therapy (Sharma et al., 2024; Perez et al., 2023; Ouyang et al., 2022). Separating “where are we in the protocol” from “say something empathic” keeps the model from collapsing into comfort. Unstructured therapy bots have already been shown to reinforce unsafe beliefs in clinical scenarios (Au Yeung et al., 2025).
Results
254 adults recruited via Prolific, mean age 36.4, 70% female, mean 17.8 months since the breakup. 121 assigned to the chatbot, 133 to survey-only control. Primary outcome: 16-item Breakup Distress Scale (range 16–64; Field et al., 2009) at 7 days. 171 people completed that follow-up.
| Chatbot | Control | |
|---|---|---|
| Baseline BDS | 35.3 | 36.0 |
| Day 7 BDS (completers) | 26.6 | 32.2 |
| Mean drop | 9.23 | 3.68 |
| Month 1 BDS | 26.0 | 29.0 |
| Reported a sudden insight | 61.7% | 19.3% |
The preregistered mixed model found a time-by-condition interaction of B = −5.36, p < .001, completer d = −0.70. At the exploratory one-month wave the advantage was smaller but still there (B = −2.92, p = .017). Both groups kept improving after day 7; the chatbot group had already arrived somewhere the controls were still walking toward.
Insight partially mediated the effect (32.7% of the total). One participant wrote: “I had been carrying the blame for my failed relationship instead of recognizing that I did everything the best I could and I was failed by someone I trusted.” People who reported that kind of shift were the ones who felt better.
Post-hoc, the effect looked larger in men (−14.7 vs −3.3 points) than in women (−7.4 vs −3.8). Cell sizes were small. Preregistered moderators — attachment anxiety, avoidance, time since breakup, still being in contact — did not fire.
Why the effect size is interesting
d = 0.70 from one session is large relative to nearby literature, with the usual caveats about inert controls:
- A meta-analysis of 30 mHealth trials with a reappraisal component pooled at SMD = 0.34 (Morello et al., 2023).
- An umbrella review of 415 single-session intervention trials pooled at SMD = −0.25 (Schleider et al., 2025).
- Woebot, the early CBT chatbot, reduced depression over two weeks versus an information control (Fitzpatrick, Darcy & Vierhile, 2017).
- Therabot, the first RCT of a fully generative therapy bot, needed four weeks of use to move depression, anxiety and eating-disorder risk (Heinz et al., 2025).
Breakup distress is a narrower, maybe more plastic target than major depression. A single trial against surveys-only will inflate the number. Still: most digital mental-health products die of attrition. Eysenbach called it the law of attrition in 2005; multi-session programs routinely lose more than half their users (Eysenbach, 2005; Christensen, Griffiths & Farrer, 2009). A session you finish in 20 minutes does not have a dropout problem.
People already take emotional questions to LLMs, and they rate the answers as empathic (Ayers et al., 2023). Empathy without a protocol is not therapy. The dual-call design is a template for any goal-directed agent that has to challenge the user instead of agreeing with them.
Caveats
The control arm was assessment-only, so you cannot separate the reappraisal sequence from “someone listened.” Treatment participants also saw their distress score in-app; controls did not. The one-month wave was not preregistered and went only to people who finished the day-0 survey. The sample is Prolific iPhone users, mostly White, mean 18 months post-breakup, not the acute first week. Self-report only. The authors say the next experiment needs an active supportive-chatbot control and a reactivation-timing test if you actually want to claim reconsolidation rather than ordinary cognitive restructuring.
This is a research summary, not treatment. If you are in crisis, use local emergency services or a crisis line. A chatbot is not a clinician.
Talk it through on Netwrck
overit is a research app, not a product here. If you want a private conversation in that neighborhood, these bots are already on the site. They are not the overit protocol. They will listen; they will not run a four-phase reconsolidation script.
Psychologist
Clinical-style listener. Empathy, reflective statements, room to unpack what happened without performing it for a friend.
Are you feeling okay?
CBT-flavored. Built to sit with anxiety, low mood, and the stories people tell themselves after a loss.
Life Coach
Forward motion rather than rumination. Useful once the belief has shifted and the question is what to do on Monday.
Or search the full catalogue: AI chatbots. Generate the kind of images in this post on the art generator.
Economics, briefly
A human therapist hour is expensive and scarce. A 20-minute Claude session that moves a validated distress scale by seven-tenths of a standard deviation is cheap even at frontier rates. The interesting product question is not “can a bot comfort you” — they already overdo that — it is whether a dual-call agent can keep the protocol honest long enough for the belief to actually move.
References
- Menzel, T., Schimpf, M., & Bohné, T. (2026). Can AI help you get over your breakup? arXiv:2605.03261. Preregistration: OSF 43MJ9.
- Ecker, B., Ticic, R., & Hulley, L. (2012). Unlocking the emotional brain. Routledge. doi:10.4324/9780203804377.
- Ecker, B. (2015). Memory reconsolidation understood and misunderstood. International Journal of Neuropsychotherapy. doi:10.12744/tnpt.2015.00002-00046.
- Schiller, D., et al. (2010). Preventing the return of fear in humans using reconsolidation update mechanisms. Nature. doi:10.1038/nature08637.
- Nader, K. (2015). Reconsolidation and the dynamic nature of memory. Cold Spring Harbor Perspectives in Biology. doi:10.1101/cshperspect.a021782.
- Sevenster, D., Beckers, T., & Kindt, M. (2013). Prediction error governs pharmacologically induced amnesia for learned fear. Science. doi:10.3389/fnbeh.2013.00041.
- Lane, R. D., et al. (2015). Memory reconsolidation, emotional arousal, and the process of change in psychotherapy. Behavioral and Brain Sciences. doi:10.1017/S0140525X14000041.
- Field, T., Diego, M., Pelaez, M., Deeds, O., & Delgado, J. (2009). Breakup distress in university students. Adolescence, 44(176), 705–727. PubMed.
- Boelen, P. A., & Reijntjes, A. (2009). Negative cognitions in emotional problems following romantic relationship break-ups. Stress and Health. doi:10.1002/smi.1219.
- Langeslag, S. J. E., & Sanchez, M. E. (2018). Down-regulation of love feelings after a romantic break-up. Journal of Experimental Psychology: General. doi:10.1037/xge0000366.
- Sharma, M., et al. (2024). Towards understanding sycophancy in language models. arXiv:2310.13548.
- Perez, E., et al. (2023). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.
- Ayers, J. W., et al. (2023). Comparing physician and chatbot responses to patient questions. JAMA Internal Medicine. doi:10.1001/jamainternmed.2023.1838.
- Fitzpatrick, K. K., Darcy, A., & Vierhile, M. (2017). Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot). JMIR Mental Health. doi:10.2196/mental.7785.
- Heinz, M. V., et al. (2025). Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI. doi:10.1056/AIoa2400802.
- Morello, K., et al. (2023). Cognitive reappraisal in mHealth interventions to foster mental health in adults. Frontiers in Digital Health. doi:10.3389/fdgth.2023.1253390.
- Eysenbach, G. (2005). The law of attrition. Journal of Medical Internet Research. doi:10.2196/jmir.7.1.e11.
- Christensen, H., Griffiths, K. M., & Farrer, L. (2009). Adherence in internet interventions for anxiety and depression. JMIR. doi:10.2196/jmir.1194.
- Wei, M., Russell, D. W., Mallinckrodt, B., & Vogel, D. L. (2007). The Experiences in Close Relationship Scale (ECR)-Short Form. Journal of Personality Assessment.
- Au Yeung, J., et al. (2025). The psychogenic machine: simulating AI psychosis, delusion reinforcement and harm enablement. arXiv:2509.10970.
Netwrck