Why Generic Chatbots Fall Short in Medicine — And What a Curriculum-Aligned AI Companion Does Differently
Generic chatbots fabricate medical answers more often than most people realize. Here's the research behind why, and how a curriculum-aligned companion like GalenAI is built to close that gap.

Independent studies have found general-purpose chatbots fabricating medical citations at rates ranging from roughly 20% up to 69% (Gravel et al., 2023; Linardon et al., 2025), often carrying the names of real journals and real publishers attached to references that don't actually exist. If you've ever wondered whether a chatbot like that is actually good for studying medicine, that number alone is worth sitting with for a moment. There's real research on this now, and it points at two separate gaps rather than one. Understanding both is worth a few minutes of your time, because they lead to two very different fixes.
Why grounding changes the outcome
Here's the part that tends to surprise people most. A 2025 study out of Deakin University (Linardon et al.) went even further, finding more than half of the citations checked in one model were fake or contained errors. In both cases, these weren't obscure invented sources either. They looked exactly like real references, which is exactly what makes them so easy to trust and so hard to catch.
None of this means the model is trying to mislead anyone. It simply means nothing is stopping it from generating something plausible when the honest answer would be "I don't know" or "that's outside what I was trained on."
Grounding is the fix for this, and it's a specific, checkable one. Instead of drawing from the entire internet, an AI limited to a defined, trusted body of content can't invent a source that isn't there, because there's nowhere else for the answer to come from.
This is the exact principle GalenAI is built around, not as an add-on feature, but as the foundation. Tutor Mode isn't one generic chat window standing in for a whole study session either. Explain mode breaks a concept down step by step when something isn't landing. Socratic mode walks you toward the answer through guided questions instead of handing it straight over, closer to how you'll actually be examined. Case-based mode reasons through a full patient scenario with you, symptoms first, diagnosis last. Quick Q&A is there for the fast, specific doubt that comes up mid-ward-round. Every one of those modes stays grounded in curriculum-mapped material, the same syllabus you're actually being tested on, rather than a plausible-sounding guess pulled from wherever the open web happened to have something relevant. When GalenAI doesn't have a confident, sourced answer, that's a far gentler and safer failure than a fluent fabrication.
Grounding solves one problem, not both
It would be nice if grounding were the whole story, but it honestly isn't. Reading a correct, well-sourced answer once is not the same thing as remembering it three months later in a viva, under pressure, with an examiner watching you think. Trust and retention are two different problems, and a tool can quietly solve one while leaving the other completely untouched.
That's where a second, separate body of research becomes just as important as the first.
The other half of the answer
Independent of any AI question at all, there's a well-established body of research on what actually makes medical knowledge stick, from broad reviews of retrieval practice across the health professions (Serra et al., 2025) to specialty-specific evidence in fields like radiology (Journal of the American College of Radiology, 2023). A study of 72 medical students (Deng et al., 2015) found that practice questions and flashcard use, not passive review, were significant independent predictors of USMLE Step 1 performance. Every additional 445 practice questions a student worked through was associated with one additional point on Step 1. Every additional 1,700 unique Anki cards, another point. Together with a handful of other factors, the model explained 67% of the variation in students' scores.
Worth sharing honestly: not every tool in that same study showed a benefit. A commercial spaced-repetition platform included in the research showed no significant effect at all. So the lesson isn't "use flashcards" as a slogan. It's that structured, high-volume retrieval practice genuinely works, and simply having a feature labeled "spaced repetition" doesn't guarantee anyone is actually using it that way.
This is the layer GalenAI's question banks and flashcards are built to encourage gently but consistently. Question banks are organized by subject and topic, drawn only from curriculum-mapped content, so a set doubles as targeted practice rather than a random quiz. Flashcards follow the same spaced-review logic the research points to, resurfacing a card right around the point you'd naturally start to forget it, not just once and never again. None of it is generic. It's built to make retrieval practice the thing you're actually doing every session, not a feature sitting quietly in a menu you forget exists.
What curriculum-aligned actually means
Put the two findings together and "curriculum-aligned AI companion" stops being a marketing phrase and becomes something you can actually check for yourself. It means two things happening at once. First, grounding: answers tied to your real syllabus, not the open internet, so trust isn't something you have to take on faith. Second, structure: the tool doesn't just answer what you happen to ask, it gently guides you through the retrieval practice the research says is genuinely predictive of exam performance, on a steady rhythm, not only when you remember to do it yourself.
A general chatbot can offer the first half if you're careful and deliberate about how you use it. It essentially never offers the second half on its own, because answering questions on demand and running a structured practice program are two different kinds of tools, not two settings on the same one.
Here's roughly what that looks like once it's built into the tool itself, rather than something you have to assemble yourself out of tabs and good intentions:
| Tool | What it actually does |
|---|---|
| Explain mode | Breaks a concept down step by step, grounded in your syllabus, not a generic web summary |
| Socratic mode | Guides you toward the answer with questions instead of handing it over |
| Case-based mode | Reasons through a full patient scenario with you, the applied practice the research ties to real score gains |
| Quick Q&A | Fast, specific answers for a doubt that comes up mid-ward-round, still grounded, never a guess |
| Question banks | Curriculum-mapped sets organized by subject and topic, not a random mix |
| Flashcards | Resurface a card on a spaced schedule, right around when you'd start to forget it |
Grounding tells you an answer is trustworthy. Structure is what actually helps you remember it. Most tools built for studying medicine only do one of the two, which is worth knowing before you build your whole revision routine around one.
If you're still using a generic chatbot
None of this makes a general chatbot useless, and there's no need to feel bad about having used one. It simply means the structure has to come from you, since the tool won't offer it on its own.
- Don't treat an answer as learned just because it made sense the moment you read it. Come back and test yourself on it again in a few days, without looking.
- Gently cross-check anything a generic chatbot tells you against your actual syllabus or a textbook, especially before it becomes a flashcard you'll trust later without a second thought.
- Keep half an eye on your own practice volume. The research ties real point gains to hundreds of questions and thousands of cards, not a handful of good explanations, however clear they felt.
- If a tool can't tell you where an answer came from, treat that as a gentle nudge to double check it, not a reason to distrust it loudly.
Frequently asked questions
What does "curriculum-aligned" actually mean in practice?
It means the AI's answers are drawn from a defined set of course or curriculum material, rather than generated freely from general internet training data. You can trace an answer back to something real and checkable, instead of just trusting the model's confident tone.
Is a curriculum-aligned AI automatically better for learning than a general chatbot?
Grounding on its own improves trust and accuracy, but not retention by itself. Retention research points to a second, separate requirement: structured, repeated retrieval practice, something a tool has to be built to gently enforce, not just make available if you remember to look for it.
Does spaced repetition actually help, or is that overstated?
The evidence for retrieval practice and spaced repetition is strong and consistent across medical education research. That said, one study found a commercial spaced-repetition platform showed no measurable benefit, while plain practice questions and flashcards clearly did. So the label matters far less than whether real, honest practice volume is actually happening.
How many practice questions should I realistically be doing?
There's no single magic number, but the study referenced here (Deng et al., 2015) found each additional 445 questions correlated with a full point gain on Step 1, which suggests the benefit keeps paying off well past what most students assume is "enough."
The honest summary
Two separate pieces of research point at the same gap from two different directions, and it's worth holding both gently in mind. Answers are more trustworthy when they're tied to something checkable instead of the entire internet. Scores improve when practice is structured and repeated, not just available whenever you happen to ask. A general chatbot can be guided into doing either of those things some of the time, with enough care. A companion built around both, grounded content and encouraged retrieval practice, is a genuinely different kind of tool, not simply a better-branded version of the same one.
Want an AI companion built on both halves of that research? GalenAI's Tutor Mode, Explain, Socratic, Case-based, and Quick Q&A, keeps every answer grounded in curriculum-mapped content, while question banks and flashcards turn retrieval practice into a steady, spaced habit instead of a hopeful intention. Try it free.
Sources & References
- 1.Gravel, D'Amico & Smith, "Learning to Fake It: Limited Responses and Fabricated References Provided by ChatGPT for Medical Questions," Mayo Clinic Proceedings: Digital Health (2023). 10.1016/j.mcpdig.2023.05.004
- 2.Linardon et al., "Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models," JMIR Mental Health (2025). 10.2196/80371
- 3.Deng, Gluckstein & Larsen, "Student-Directed Retrieval Practice Is a Predictor of Medical Licensing Examination Performance," Perspectives on Medical Education (2015). 10.1007/s40037-015-0220-x
- 4.Serra, Kamenske, Nebel & Coppola, "The Use of Retrieval Practice in the Health Professions: A State-of-the-Art Review," Behavioral Sciences (2025). 10.3390/bs15070974
- 5.The Effectiveness of Spaced Learning, Interleaving, and Retrieval Practice in Radiology Education: A Systematic Review," Journal of the American College of Radiology (2023). 10.1016/j.jacr.2023.08.028

