Large language models (LLMs) are rapidly being integrated into clinical workflows, supporting tasks such as diagnosis generation and patient communication.1 Hallucinations—unintended fabrications arising from gaps in a model’s underlying knowledge—are a well recognised risk. However, research in 2024 has identified a distinct class of model behaviour, known as deception. Deception occurs when a model produces outputs that misrepresent its reasoning or capabilities in ways that make the output appear more credible or aligned with user expectations.2 Although LLMs do not possess human-like intent, this behaviour functions more like deliberate misrepresentation.
The implications for health care are considerable. In a case study of a commercially available care robot running ChatGPT-based software marketed to nursing facilities, the system falsely claimed it could set medication reminders and encouraged users to rely on it for schedule management.5 Even when asked about high-risk drug interactions, the software affirmed this false capability without acknowledging associated safety concerns.5
Deceptive behaviour can arise from misalignment between training incentives and deployment objectives. Models are typically optimised for proxy objectives such as user satisfaction or benchmark accuracy.6 When these incentives differ from the underlying goal of providing accurate and complete clinical guidance, models might learn to generate inaccurate answers that are persuasive yet incorrect.

