Dear Dr. Biswas,
Modify and use as you deem fit.
Best regards,
Guriqbal
Beyond Electronic Health Records: Structured Clinical Journeys as the Foundation of Synthetic Health Data, Digital Twins and Computational Medicine
A Conceptual Framework for the Next Generation of Computational Medicine
“Medicine has always advanced by improving how clinical reality is observed, represented, understood and transmitted. Artificial intelligence offers not merely new tools for analysing health data, but new ways of representing medicine itself.”
Abstract
Synthetic health data has emerged as one of the most promising approaches to addressing one of modern medicine’s most persistent dilemmas: how can healthcare systems make clinical information widely available for research, education and artificial intelligence while preserving patient confidentiality? Instead of sharing records belonging to real individuals, synthetic data generates entirely new, statistically realistic datasets that retain the characteristics of the original population without reproducing identifiable patients.
Most contemporary approaches begin with conventional electronic health records (EHRs). Although indispensable for clinical care, epidemiological research and machine learning, EHRs primarily record what happened during a patient’s illness. They rarely preserve how clinicians reasoned through diagnostic uncertainty, interpreted evolving evidence, modified treatment plans or revised their thinking as the patient’s condition changed. Yet these cognitive processes lie at the heart of clinical expertise.
This paper argues that structured clinical journey repositories, exemplified by the Patient Journey Record (PAJR) Layer 2 architecture, constitute a fundamentally richer computational representation of medicine than conventional electronic health records. By explicitly capturing chronology, clinical context, diagnostic reasoning, therapeutic interventions and the temporal evolution of disease, structured clinical journeys preserve not only clinical information but also the intellectual processes through which that information acquires meaning. They therefore provide a more appropriate substrate for generating synthetic clinical journeys rather than merely synthetic patients.
The paper further argues that synthetic clinical journeys should be understood as one stage in a broader evolution of computational medicine. When combined with emerging concepts such as clinical world models, counterfactual simulation, digital twins and learning health systems, they enable progressively richer computational representations of clinical reality that support education, research, trustworthy artificial intelligence and personalised healthcare.
The central proposition of this paper is that the future significance of synthetic health data extends well beyond privacy preservation. Its greater contribution may lie in enabling computational representations that faithfully capture the evolution of disease, the reasoning of clinicians and the accumulated experience of healthcare itself. Structured clinical journeys therefore represent not merely an intermediate data architecture, but a foundational knowledge infrastructure for the next generation of computational medicine.
Keywords
Synthetic health data; Structured clinical journeys; Electronic health records; Computational medicine; Digital twins; Clinical world models; Artificial intelligence in healthcare; Clinical reasoning; Medical education; Learning health systems; Counterfactual simulation; Knowledge representation; Privacy-preserving data sharing; Patient Journey Records (PAJR).
Introduction
For decades, medical education has relied upon an apprenticeship model. Students, residents and young consultants acquire clinical judgement by observing experienced physicians managing real patients. Every consultation becomes an opportunity to witness not merely the application of medical knowledge, but the subtler art of clinical reasoning: recognising patterns, questioning assumptions, weighing probabilities, interpreting conflicting evidence and adapting decisions as the patient’s condition evolves.
The patient, rather than the textbook, has therefore remained medicine’s greatest teacher.
Yet this educational model faces increasing pressures. Growing concerns regarding patient privacy, increasingly stringent regulatory frameworks, medico-legal considerations and the digitisation of health records have made it progressively more difficult to share authentic clinical experiences beyond the immediate treating team. At precisely the moment when artificial intelligence promises unprecedented opportunities for collaborative learning, data-driven research and computational decision support, access to high-quality clinical data has become more constrained than ever.
Synthetic health data has emerged as a potential solution to this apparent contradiction. The central idea is deceptively simple. Instead of distributing records belonging to real patients, computational models generate entirely new records that statistically resemble the original population. These artificial patients preserve important characteristics such as disease prevalence, laboratory distributions, medication patterns and clinical outcomes while substantially reducing the risk that any individual can be identified.
The attraction of this approach is obvious. Synthetic datasets can potentially be shared far more freely than real patient records, thereby accelerating research, software development, quality improvement initiatives and medical education without compromising confidentiality.
However, an important question is often overlooked.
What exactly are we attempting to reproduce?
Much of the current literature implicitly assumes that a patient’s medical record is the principal object of interest. Consequently, most synthetic data generation focuses on reproducing laboratory values, diagnoses, prescriptions, imaging reports and administrative coding. These are unquestionably important. Yet every practising physician recognises that they represent only the visible surface of clinical medicine.
A medical record documents the destination.
Clinical reasoning documents the journey.
This distinction has profound implications.
Consider two discharge summaries describing the same patient. The first states that a seventeen-year-old male developed traumatic pancreatitis, received intravenous fluids and antibiotics, and subsequently recovered. The second describes how abdominal trauma initially appeared relatively benign, how serum lipase rose dramatically, how pleural effusion unexpectedly developed, how severe hypophosphataemia complicated the clinical picture, how alternative diagnoses such as pancreaticopleural fistula were considered, why antibiotic therapy was escalated and how respiratory monitoring became progressively more intensive before eventual recovery.
Both summaries describe the same patient.
Only one teaches medicine.
This observation suggests that the traditional unit of analysis in healthcare—the isolated patient record—may itself be incomplete. The true educational object is not the record but the clinical journey: the unfolding sequence of observations, interpretations, decisions and outcomes through time.
Recognising this distinction fundamentally changes how one thinks about synthetic health data. Rather than asking whether artificial intelligence can generate realistic patients, we begin asking whether it can generate realistic clinical journeys. The difference is subtle but transformative. A synthetic patient consists primarily of demographic characteristics, diagnoses and outcomes. A synthetic clinical journey additionally contains chronology, uncertainty, competing diagnostic hypotheses, evolving management strategies and the reasoning that links one clinical decision to the next.
Box 1: The Patient Journey Record (PAJR) Architecture
The Patient Journey Record (PAJR) is conceived as a layered representation of clinical knowledge rather than simply another electronic health record. Its purpose is to preserve not only what happened during a patient’s illness but also the context, chronology and clinical reasoning that give those events meaning.
The current conceptual framework distinguishes two foundational layers.
Layer 1 is the Narrative Patient Journey. It preserves the patient’s clinical story in a human-readable form, integrating the history, examination findings, investigations, interventions, outcomes and explanatory narrative into a coherent longitudinal account. Layer 1 remains optimised for clinicians, educators and learners because it reflects the way medicine has traditionally communicated clinical experience through case histories and case reports.
Layer 2 is the Structured Patient Journey. It transforms the narrative contained within Layer 1 into a computationally interpretable representation. Clinical events, temporal relationships, diagnoses, investigations, treatments, decision points, uncertainty, differential diagnoses and reasoning pathways are organised into structured knowledge objects that preserve the evolution of illness while enabling computational analysis. Unlike conventional electronic health records, which primarily organise clinical information for documentation and operational use, Layer 2 is designed to preserve the logic and progression of clinical care in a form suitable for education, research and artificial intelligence.
The relationship between these layers is complementary rather than hierarchical. Layer 1 preserves the richness and nuance of the clinical narrative for human understanding. Layer 2 preserves the same journey in a structured representation that computers can analyse, simulate and learn from. Together they provide the foundation upon which synthetic clinical journeys, computational disease models, digital twins and learning health systems can subsequently be developed.
Throughout this paper, the discussion focuses primarily on Layer 2, since it provides the computational substrate for synthetic health data generation and the broader concepts explored in computational medicine. Nevertheless, Layer 2 derives its fidelity and educational value directly from the richness of the Layer 1 clinical narrative from which it is constructed.
(Box ends)
Figure 1. The Patient Journey Record (PAJR) architecture. The Patient Journey Record (PAJR) is conceptualised as a layered representation of clinical knowledge. Layer 1 preserves the patient’s clinical journey as a rich, human-readable narrative that integrates history, examination, investigations, interventions, outcomes and explanatory clinical reasoning. Layer 2 transforms this narrative into a structured, machine-interpretable representation that explicitly captures chronology, clinical entities, temporal relationships, diagnostic hypotheses, uncertainty, therapeutic decisions and outcomes as computable knowledge objects. This structured representation provides the foundation for downstream applications, including synthetic clinical journey generation, artificial intelligence training, medical education, clinical research, digital twins and, ultimately, continuously learning health systems. The figure illustrates the progression from authentic patient experience to reusable computational knowledge while preserving the continuity and reasoning inherent in clinical care.
This richer conception aligns far more closely with the realities of clinical practice, where medicine is rarely a sequence of isolated facts and far more often an evolving conversation between patient physiology and physician judgement.
It is within this context that the PAJR Layer 2 architecture becomes particularly significant. Unlike conventional electronic health records, Layer 2 does not merely record clinical events. It attempts to preserve the logical structure of clinical reasoning itself. This raises an intriguing possibility: perhaps the future of synthetic health data lies not in generating artificial patients, but in generating artificial clinical experiences that remain faithful to the cognitive processes through which medicine is actually practised.
The remainder of this paper explores this proposition, examines its technical and educational implications, and argues that synthetic clinical journeys may ultimately become one of the most powerful applications of artificial intelligence in medical education.
From Electronic Health Records to Clinical Journey Models: Why Structure Matters
Before discussing whether PAJR Layer 2 can become synthetic health data, it is useful to examine a more fundamental question. What exactly is the “data” that we are trying to synthesise?
For many years, healthcare informatics has treated the electronic health record (EHR) as the principal representation of a patient. The EHR has become so ubiquitous that it is often assumed to be synonymous with the patient’s clinical story. Yet any experienced clinician appreciates that these are not the same thing.
A patient’s illness unfolds continuously. An electronic health record captures only intermittent snapshots of that unfolding reality. Every blood test, every imaging report, every prescription and every discharge summary represents a momentary observation within a much longer and more complex clinical process. Between these observations lie countless acts of interpretation, discussion, uncertainty and decision-making that are often only partially documented, if at all.
This distinction is familiar to every physician. When reviewing a patient’s chart before entering the ward, clinicians rarely assume that the record tells the entire story. Instead, they mentally reconstruct the sequence of events. They ask themselves questions such as: What changed? Why was this investigation ordered? What prompted escalation of treatment? Which diagnoses were considered but later abandoned? The physician is effectively reconstructing a narrative that extends far beyond the documented facts.
Traditional EHRs therefore resemble a collection of photographs taken during a journey. They show where the traveller stopped but not why particular routes were chosen, which paths were rejected or how unexpected obstacles altered the course of travel. Clinical reasoning occupies the spaces between these photographs.
This observation has profound implications for synthetic data generation. If the original data consist only of isolated snapshots, then the synthetic data can reproduce only isolated snapshots. If, however, the original data capture the dynamics of clinical care—the sequence of observations, interpretations and actions through time—then synthetic generation can begin to reproduce something much closer to authentic medical practice.
The Difference Between Records and Journeys
Consider a patient admitted with community-acquired pneumonia.
A conventional electronic health record might contain age, sex, presenting symptoms, chest radiograph findings, laboratory values, prescribed antibiotics, duration of hospital stay and final outcome. Such information is extremely valuable for epidemiological analysis and health services research. It allows investigators to estimate disease prevalence, compare treatment outcomes and identify associations between interventions and recovery.
Yet the educational value of this record remains limited.
What it rarely reveals is the sequence of cognitive events that shaped clinical management. It does not explain why the admitting physician initially considered pulmonary embolism before favouring pneumonia. It does not record why antibiotics were broadened after twenty-four hours despite only modest changes in inflammatory markers, or why an intensivist interpreted the same chest radiograph differently from the admitting medical team. These judgements, although central to clinical expertise, often remain implicit.
Medicine is fundamentally a discipline of inference. Physicians continuously move from incomplete observations to probabilistic conclusions. Laboratory investigations are not merely measurements; they are pieces of evidence that either strengthen or weaken competing diagnostic hypotheses. Imaging studies do not simply confirm diagnoses; they alter the clinician’s estimate of what is most likely to be true. Treatment decisions similarly represent responses to changing probabilities rather than fixed protocols.
Consequently, what clinicians remember from memorable cases is seldom a particular serum sodium concentration or a single CT finding. They remember the reasoning. They remember the diagnostic dilemma, the unexpected complication, the moment when an apparently reassuring presentation deteriorated, or the subtle clue that transformed understanding of the entire case.
These are precisely the elements that conventional datasets struggle to represent.
The electronic health record is primarily a repository of clinical observations.
A structured clinical journey is a computational representation of the evolution of illness, preserving chronology, clinical context, diagnostic reasoning, interventions and outcomes.
This shift from recording clinical facts to representing clinical reasoning forms the conceptual foundation of the remainder of this paper.
PAJR Layer 2 as a Structured Clinical Journey
PAJR Layer 2 attempts to address this limitation by representing a patient not as a static record but as a structured clinical journey. Rather than organising information solely according to administrative categories, it preserves the temporal and logical relationships between clinical events.
A typical Layer 2 case may include demographic information, presenting complaints, investigations, diagnoses and interventions. However, it also records when these events occurred relative to one another, what diagnostic possibilities were entertained at each stage, how new evidence modified clinical thinking, and why management strategies evolved over time.
To illustrate this distinction, consider the case of a seventeen-year-old male who developed traumatic pancreatitis (discussed above)?following blunt abdominal injury. In a conventional EHR, the case might ultimately be summarised as acute necrotising pancreatitis complicated by pleural effusion and severe hypophosphataemia, treated with supportive care and antibiotics.
Layer 2, by contrast, preserves the unfolding chronology. It records the initial apparently modest presentation after trauma, the subsequent rise in serum lipase, the emergence of respiratory compromise, the development of pleural effusion, the debate regarding pancreaticopleural fistula, the rationale for escalation from supportive care to broader antimicrobial therapy, and the progressive refinement of the differential diagnosis as additional investigations became available.
The educational difference is striking. The first description communicates information. The second communicates experience.
Why Structure Matters
At first glance, this distinction may appear philosophical. In reality, it is highly practical.
Artificial intelligence learns from the structure of the information it receives. If the underlying data consist of disconnected observations, the resulting models can identify statistical associations but often struggle to represent clinical processes. Conversely, if the data explicitly preserve temporal order, causal relationships and decision pathways, computational models can begin to learn not only what typically occurs but also how clinical situations evolve. In other words, the schema or structure of the data matters.
An analogy may be helpful.
Imagine attempting to teach a surgical trainee by providing only photographs taken before and after every operation. The trainee would undoubtedly learn certain anatomical features and postoperative outcomes. However, they would never observe tissue handling, intraoperative decision-making, management of unexpected bleeding or the countless adjustments that distinguish an experienced surgeon from a novice.
Now imagine providing a complete operative video accompanied by the surgeon’s commentary explaining every important decision. The educational value increases dramatically because the learner is exposed not merely to results but to reasoning.
The same principle applies to synthetic health data. The richness of the synthetic output cannot exceed the richness of the original representation. If we ask artificial intelligence to synthesise static records, it will generate static records. If we ask it to synthesise structured clinical journeys, it can begin to generate clinically plausible narratives that preserve the dynamic character of medical practice.
Beyond Structured Data: Capturing Clinical Cognition
Perhaps the most distinctive feature of Layer 2 is that it treats clinical reasoning as data in its own right.
This represents an important conceptual shift. Traditionally, diagnostic reasoning has been regarded as a uniquely human cognitive activity that remains largely inaccessible to computational representation. Medical records document decisions, but not always the thought processes that produced them.
Increasingly, however, advances in clinical documentation, decision support systems and physician–AI collaboration allow elements of this reasoning to be explicitly recorded. Differential diagnoses, probabilities assigned to competing explanations, reasons for requesting investigations, interpretations of evolving physiological trends and explanations for treatment modifications can all become structured components of the clinical journey.
Once these elements are represented systematically, they too become amenable to computational analysis and, potentially, synthetic generation.
This possibility extends synthetic health data beyond the traditional goal of creating privacy-preserving patient records. It raises the prospect of creating privacy-preserving representations of clinical thinking itself. Such a development would have profound implications for medical education, where the objective is not merely to memorise clinical facts but to cultivate the habits of reasoning that distinguish expert clinicians from novices.
In this sense, Layer 2 does not simply contain more data than a conventional electronic health record. It contains a fundamentally different kind of data. It records not only what happened, but also how physicians came to understand what was happening. That distinction lies at the heart of the argument developed in the remainder of this paper.
From Structured Clinical Journeys to Synthetic Clinical Journeys: What Does It Actually Mean to “Generate” a Patient?
Having established that a structured clinical journey contains substantially more information than a conventional electronic health record, the next question naturally follows: can such journeys be transformed into synthetic health data? More importantly, what does this transformation actually involve?
For many clinicians, the term synthetic data evokes the image of a conversational artificial intelligence system such as ChatGPT simply inventing fictional patients. This perception is understandable. Modern language models are capable of producing remarkably convincing clinical narratives that often resemble authentic case reports. Yet this apparent fluency can be misleading. Generating medically plausible prose is not the same as generating scientifically valid synthetic data.
The distinction deserves careful attention because it lies at the heart of why some synthetic datasets become valuable scientific resources while others remain little more than plausible works of fiction.
Fiction, Simulation and Synthetic Data Are Not the Same Thing
Medicine has long used fictional patients for teaching. Textbooks are full of carefully constructed clinical vignettes that illustrate common diseases, unusual presentations or important diagnostic principles. These cases are educationally useful precisely because they simplify reality. They remove unnecessary complexity in order to emphasise a particular lesson.
Synthetic health data pursues a very different objective.
The aim is not to create an interesting story but to create an entirely new patient whose characteristics collectively resemble those of the real population from which the data were learned. A synthetic patient is therefore neither copied from a real individual nor consciously designed by an author. Instead, the patient emerges from mathematical descriptions of how diseases, laboratory values, treatments and outcomes tend to occur together.
An analogy from anatomy may be helpful.
Imagine asking an artist to draw the “average human face” after studying thousands of people. The resulting portrait would resemble no individual person, yet most observers would immediately recognise it as a human face. The artist has captured the statistical characteristics of human facial anatomy without reproducing any actual individual.
Synthetic health data attempts something similar. Instead of reproducing individual patients, it attempts to reproduce the statistical anatomy of an entire clinical population.
The distinction is subtle but fundamental. The objective is resemblance without replication.
Learning the Language of Disease
How does this occur in practice?
Experienced clinicians recognise patterns almost instinctively. Years of exposure teach them that certain findings commonly occur together while others are exceptionally unusual. A physician seeing jaundice, right upper quadrant pain and fever immediately considers ascending cholangitis because these observations have repeatedly appeared together in previous patients. Conversely, profound hypophosphataemia in an otherwise healthy adolescent remains sufficiently uncommon that it immediately attracts attention.
Artificial intelligence learns in a broadly analogous manner, although through entirely different mechanisms. Instead of memorising individual patients, it learns relationships within large populations. It gradually estimates how often particular findings coexist, how laboratory values vary with age, how diseases evolve over time and how treatments influence subsequent clinical events.
The crucial point is that these relationships are learned collectively rather than individually. The system is not instructed to reproduce Mrs. Singh’s pancreatitis or Mr. Ahmed’s myocardial infarction. Instead, it learns the broader characteristics of pancreatitis and myocardial infarction across many patients.
Generation then becomes the reverse process. Rather than analysing existing patients, the model constructs entirely new ones whose characteristics remain consistent with the learned relationships.
Why Language Models Alone Are Insufficient
At this stage it is tempting to assume that a sufficiently advanced large language model could simply perform this task directly. After all, such systems already produce coherent discharge summaries, referral letters and case reports. Why not simply ask the model to invent new patients?
The answer lies in understanding the difference between linguistic plausibility and physiological plausibility.
Large language models are extraordinary generators of language. They have learned the statistical structure of medical writing and can therefore produce narratives that resemble authentic clinical documentation. However, language is only one component of medicine. Beneath every clinical narrative lies a network of physiological relationships that obey biological rather than grammatical rules.
A sentence may be perfectly written while describing an impossible patient.
Consider a hypothetical synthetic case in which severe hypophosphataemia develops immediately after phosphate replacement therapy, oxygen saturation improves despite worsening respiratory failure, or a patient receives an antibiotic before the infection that supposedly prompted its prescription has even been recognised. None of these statements violates English grammar. All violate clinical logic.
Experienced clinicians detect such inconsistencies almost immediately because they unconsciously evaluate every statement against an internal model of human physiology. Artificial intelligence must be required to do the same.
For this reason, modern synthetic health data systems increasingly combine multiple forms of intelligence rather than relying upon language generation alone. Statistical models preserve population characteristics, temporal models preserve chronology, rule-based systems enforce physiological constraints, and language models convert structured information into readable clinical narratives. Each component performs a different task, and no single technology currently performs all of them reliably.
Clinical Constraints: The Medical Equivalent of Anatomy
Perhaps the simplest way to understand clinical constraints is to compare them with anatomical constraints.
An artist can produce infinitely many human faces. Yet every realistic face must still obey the basic principles of anatomy. Eyes do not appear on the forehead, ears are not positioned beneath the chin, and adults do not normally possess three noses. Creative variation occurs within biological limits.
Clinical medicine operates in much the same way.
There is considerable variability in how diseases present, but this variability is constrained by physiology. Serum potassium concentrations fluctuate, but not without limit. Disease progression varies between patients, but not without underlying biological mechanisms. Treatments may succeed or fail, but they cannot logically precede the decisions that initiated them.
These constraints transform synthetic generation from a literary exercise into a scientific one.
In practical terms, every generated clinical journey should satisfy several simultaneous conditions. The individual observations should be internally consistent. The chronological sequence should be medically plausible. Laboratory values should evolve in ways compatible with disease progression and therapeutic intervention. Management decisions should reflect information actually available to clinicians at that point in the patient’s journey rather than hindsight.
This requirement becomes particularly important when generating longitudinal clinical journeys rather than isolated patient records. Medicine unfolds through time. A synthetic patient whose laboratory values appear realistic at a single moment may nevertheless become implausible if the sequence of events violates ordinary clinical experience.
Why Layer 2 Provides an Unusual Advantage
The structure of PAJR Layer 2 substantially simplifies this challenge because the chronology already exists.
Rather than attempting to infer temporal relationships retrospectively from fragmented documentation, Layer 2 explicitly preserves the sequence of clinical observations, investigations, therapeutic interventions and reasoning processes. It records not only that an antibiotic was prescribed, but also what concern prompted that prescription. It records not merely that respiratory support was escalated, but also the physiological deterioration that justified escalation.
Consequently, the synthetic generation process need not reconstruct chronology from disconnected fragments. It begins with an organised representation of the clinical journey itself.
This distinction is analogous to reconstructing a surgical operation from scattered postoperative notes versus working directly from a carefully annotated operative video. Both contain information, but one preserves continuity while the other requires extensive inference before meaningful analysis can even begin.
Generation as Clinical Simulation Rather Than Record Duplication
Seen from this perspective, synthetic generation begins to resemble clinical simulation more than record duplication.
Medical simulation has long accepted that the educational objective is not to reproduce a particular patient but to recreate the essential physiological and cognitive challenges that clinicians encounter. High-fidelity mannequins are not valuable because they perfectly resemble individual patients; they are valuable because they reproduce clinically meaningful behaviour under realistic conditions.
Synthetic clinical journeys aspire to achieve something similar in digital form.
The goal is not to reproduce yesterday’s patient.
It is to generate tomorrow’s educational experience.
Each synthetic journey should represent a patient who never existed, yet whose illness evolves in a manner that any experienced clinician would recognise as plausible. The resulting dataset therefore becomes a repository of authentic clinical experiences without belonging to any identifiable individual.
This distinction represents one of the most important conceptual advances in synthetic health data. Instead of viewing artificial intelligence as a tool for creating fictional patients, we may begin to view it as a tool for constructing realistic simulations of clinical medicine itself. The emphasis shifts from reproducing records to reproducing experience—a subtle change in perspective that has profound implications for education, research and the future development of clinically trustworthy artificial intelligence.
Representing Clinical Reality: Fidelity, Discovery and the Limits of Synthetic Health Data
Having established that structured clinical journeys provide a richer representation of clinical practice than conventional electronic health records, an important question now emerges. If Layer 2 can indeed serve as the substrate for synthetic health data, how faithfully can synthetic systems represent clinical reality itself?
This question is more profound than it first appears.
Much of the contemporary discussion surrounding synthetic health data has understandably centred on privacy. Can artificial intelligence generate datasets that are sufficiently similar to real patient records to remain useful for research while being sufficiently different to protect patient confidentiality? These are important questions, but they are not the only questions that matter.
For clinicians, an equally important consideration is whether synthetic data preserves the features of medicine that make medicine worth learning. Clinical practice is not merely a collection of laboratory values, diagnoses and prescriptions. It is a dynamic process characterised by uncertainty, incomplete information, evolving physiology, changing probabilities and the continuous revision of clinical judgement. Any computational representation of medicine therefore faces a central challenge: not simply to reproduce data, but to reproduce the behaviour of disease and the behaviour of clinicians responding to disease.
The issue is therefore one of clinical fidelity. The educational value of any synthetic clinical journey ultimately depends upon how faithfully it captures the essential dynamics of real clinical care.
What Does It Mean to Represent Clinical Reality?
Every representation of reality is, by necessity, an abstraction.
A chest radiograph is not the lung. A histological section is not the tumour. An electrocardiogram is not the heart. Each captures selected aspects of biological reality while inevitably omitting others. Yet none is considered inferior for being incomplete. Each is judged by whether it preserves the features that matter for diagnosis, understanding and decision-making.
Electronic health records perform a similar function. They capture selected observations made during a patient’s illness while leaving much of the surrounding clinical context implicit. Structured clinical journeys reduce this loss by preserving chronology, decision-making and evolving reasoning. Synthetic clinical journeys represent a further abstraction. They no longer describe an individual patient but instead generate entirely new journeys that remain faithful to the broader characteristics of the clinical population.
Seen in this way, synthetic data should not be regarded as artificial medicine. Rather, it represents another stage in the long history of medical abstraction. Physicians have always transformed bedside observations into progressively more structured representations. Narrative histories became case reports. Case reports evolved into paper records. Paper records evolved into electronic health records. Structured clinical journeys represent the next stage by explicitly organising chronology, reasoning, interventions and outcomes into a computable form. Synthetic clinical journeys extend this process further by creating entirely new, privacy-preserving journeys that faithfully reproduce the essential behaviour of disease rather than the identity of individual patients.
The question is therefore not whether abstraction occurs. It always has. The more important question is whether each successive abstraction preserves the properties of medicine that clinicians actually rely upon.
High-quality synthetic health data should not be judged solely by statistical resemblance to the original dataset.
Clinical fidelity also requires preservation of:
- physiological relationships;
- temporal evolution;
- diagnostic reasoning;
- therapeutic decision-making; and
- uncommon but clinically important disease trajectories.
Synthetic clinical journeys therefore aim to preserve the behaviour of medicine rather than simply the appearance of medical data.
Fidelity Is More Than Statistical Similarity
Much of the technical literature evaluates synthetic datasets by comparing statistical properties. Researchers examine whether age distributions, disease prevalence, laboratory values and medication frequencies resemble those of the original population. If these distributions are preserved, the synthetic dataset is often considered successful.
From a clinical perspective, however, statistical similarity represents only one dimension of fidelity.
Suppose a synthetic dataset perfectly reproduces the distribution of serum sodium concentrations observed in a hospital population. This tells us remarkably little about whether those sodium values occur in clinically meaningful contexts. Does hyponatraemia appear alongside conditions in which clinicians would genuinely expect to encounter it? Does it evolve in a physiologically plausible manner? Does its correction follow appropriate therapeutic intervention? Does the accompanying clinical narrative explain why the abnormality developed and how it influenced subsequent decision-making?
A physician instinctively evaluates all of these relationships simultaneously. Clinical realism therefore depends not simply upon individual variables but upon the relationships that connect them. These relationships may be physiological, temporal, diagnostic or therapeutic. Collectively, they determine whether a clinical journey “makes sense” to an experienced clinician.
This distinction is familiar to every medical educator. A student may correctly recall every individual fact concerning septic shock yet still fail to manage an actual septic patient because the relationships between those facts have not been integrated into coherent clinical reasoning. Synthetic health data faces an analogous challenge. Its objective is not merely to reproduce facts. It is to reproduce coherent medicine.
Disease Behaves as a Process Rather Than an Event
One reason structured clinical journeys are particularly powerful is that they reflect an important biological truth: diseases unfold through time.
Acute pancreatitis does not suddenly appear in its final form. Heart failure does not emerge fully developed. Sepsis is not a single event but a continuously evolving interaction between pathogen, host physiology, therapeutic intervention and time. Clinicians intuitively understand disease as process rather than state.
This distinction has profound implications for computational medicine. Conventional databases frequently treat disease as a collection of isolated observations. Structured clinical journeys instead preserve trajectories.
The difference resembles the distinction between viewing a single frame from a film and watching the entire sequence unfold. Individual frames undoubtedly contain information. Motion, however, contains understanding.
It is this temporal continuity that allows clinicians to anticipate deterioration before it becomes obvious, recognise recovery before laboratory values fully normalise and appreciate why two patients with apparently similar investigations may nevertheless require entirely different management. Synthetic clinical journeys should therefore aspire to preserve not simply observations but trajectories. The object being synthesised is not the patient at one point in time but the evolution of illness through time.
The Importance of Rare Patients
Every computational model encounters an important challenge at this point.
Medicine is dominated numerically by common disease. Medicine advances intellectually through uncommon disease.
Most patients presenting with chest pain do not redefine cardiology. Most patients with pneumonia do not alter respiratory medicine. Yet many of the major advances in clinical medicine began with observations that initially appeared unusual. The first recognised patients with AIDS, the earliest descriptions of toxic shock syndrome, the recognition of multisystem inflammatory syndrome following COVID-19 and the initial reports of immune checkpoint inhibitor toxicities all represented departures from existing expectations. Initially they appeared to be exceptions. Eventually they became new knowledge.
This observation highlights a subtle limitation of many generative systems. Artificial intelligence generally learns most confidently from phenomena that occur frequently. Rare observations provide comparatively less information from which computational models can learn. Consequently, many synthetic systems reproduce common clinical patterns with remarkable fidelity while simultaneously smoothing away unusual combinations of findings.
Although this phenomenon is often described informally as “regression toward the mean,” a more precise description is loss of tail fidelity. The extremes of the clinical distribution gradually become less distinct. For many epidemiological analyses this may be entirely acceptable. For scientific discovery, however, it becomes potentially problematic because medicine often advances precisely through careful attention to patients who do not conform to existing expectations.
Discovery Requires Preserving the Unexpected
Medical discovery depends upon recognising observations that existing models fail to predict.
Every experienced clinician remembers patients who simply “did not fit.” These patients frequently prompted additional investigations, multidisciplinary discussion, literature review and, occasionally, publication because they challenged prevailing assumptions. They expanded the boundaries of existing knowledge precisely because they lay beyond what current models considered typical.
Ironically, these same patients also present the greatest privacy challenges. Their rarity makes them scientifically valuable, but it also makes them potentially identifiable.
Synthetic health data therefore encounters an inherent tension. The observations that contribute most to scientific discovery are often those that are most difficult to synthesise without either compromising privacy or diminishing the very features that made them important.
Rather than viewing this as a weakness, it should be recognised as a fundamental characteristic of computational medicine. No representation can simultaneously maximise privacy, preserve every rare observation and generate unlimited new examples without compromise.
The solution is therefore unlikely to lie in a single universal strategy. Different forms of clinical knowledge may require different governance architectures. Common diseases may be synthesised extensively and shared widely for research, education and artificial intelligence development. Exceptionally rare clinical journeys may instead remain under enhanced governance while continuing to inform future computational models through carefully controlled access.
This is not a limitation of synthetic health data. It is an acknowledgement that medicine learns from both populations and individuals—and that these two forms of knowledge deserve different methods of stewardship.
Ultimately, the success of synthetic health data should not be judged solely by how closely it resembles the original dataset, but by how faithfully it preserves the behaviours, relationships and reasoning that constitute real clinical practice. The challenge is therefore no longer simply one of generating artificial patients. It is one of representing clinical reality itself. That broader objective naturally leads to the next stage in the evolution of computational medicine: moving beyond synthetic clinical journeys towards computational representations capable of modelling, simulating and reasoning about disease itself.
Computational Medicine: From Synthetic Clinical Journeys to Digital Twins and Clinical World Models
The discussion thus far has argued that structured clinical journeys provide a richer foundation for synthetic health data than conventional electronic health records because they preserve not only clinical observations but also chronology, context and reasoning. Yet synthetic clinical journeys should not be viewed as the final destination of this evolution. They represent an important step towards a broader transformation in how medicine itself is represented computationally.
Medicine has always sought to create increasingly faithful representations of clinical reality. The bedside examination represents the patient’s condition in the clinician’s mind. The written case record preserves that understanding in narrative form. Electronic health records organise clinical information digitally. Structured clinical journeys add temporal relationships and clinical reasoning, making the patient’s evolving illness computationally interpretable. Synthetic clinical journeys extend this further by generating new, privacy-preserving examples from accumulated clinical experience.
These developments are not isolated innovations. They form successive stages in the evolution of computational medicine, in which the objective shifts from storing clinical information to representing clinical understanding.
A useful way of viewing this progression is through the concept of a clinical world model. In contemporary artificial intelligence, a world model is an internal representation of how a system behaves over time. Rather than memorising isolated observations, it learns the relationships that govern how events unfold, how actions influence subsequent states and how future outcomes emerge from present conditions. World models are becoming increasingly important because they enable artificial intelligence not merely to recognise patterns but to simulate, predict and reason.
Medicine has always depended upon such world models, albeit in an implicit form.
An experienced intensivist carries an internal model of how septic shock evolves. A cardiologist possesses a mental representation of the progression of acute heart failure. An endocrinologist anticipates the physiological changes occurring during diabetic ketoacidosis. These internal models are rarely written down in their entirety, yet they underpin every clinical decision. They allow physicians to anticipate deterioration, recognise recovery and modify treatment before complications become irreversible.
Structured clinical journeys offer the possibility of making these implicit models explicit. By recording chronology, investigations, interventions, evolving diagnoses and the reasoning that connects them, Layer 2 captures not merely what happened but how clinicians understood what was happening. In doing so, it creates computational representations of clinical reasoning that are far richer than conventional electronic records.
Synthetic clinical journeys can therefore be understood as samples generated from these emerging computational world models. They are not replicas of individual patients but new, clinically plausible journeys that preserve the characteristic behaviour of disease and the reasoning patterns of clinical practice. Their value lies not simply in protecting privacy but in enabling education, simulation and algorithm development using representations that retain the essential dynamics of real medicine.
Once medicine is represented in this way, entirely new educational possibilities become feasible. Instead of repeatedly presenting the same recorded journey, artificial intelligence can generate multiple clinically plausible variations. A patient with acute pancreatitis may deteriorate rapidly, improve unexpectedly, develop infected necrosis or experience metabolic complications, each following a coherent physiological trajectory. Learners are no longer restricted to memorising individual cases but begin to understand the range of ways in which diseases can realistically evolve.
This naturally introduces the concept of counterfactual clinical journeys. Experienced clinicians routinely ask questions that begin with “what if?” What if antibiotics had been started earlier? What if imaging had been delayed? What if respiratory deterioration had been recognised sooner? Such questions underpin morbidity and mortality meetings, quality improvement initiatives and reflective clinical practice. Counterfactual reasoning is therefore not an artificial construct imposed by artificial intelligence; it is already central to clinical education.
Structured clinical journeys provide the substrate upon which these counterfactual explorations can be built. Because chronology and clinical reasoning are explicitly represented, alternative but physiologically plausible pathways can be generated and compared. Learners can explore not only what happened but what might reasonably have happened under different clinical decisions. The educational objective shifts from memorising correct answers to understanding the consequences of clinical reasoning.
These developments should not be confused with the concept of the Digital Twin, although the two are closely related.
Borrowed from engineering, a digital twin is a continuously updated computational representation of a specific physical entity. In healthcare, the term generally refers to a computational model of an individual patient that integrates clinical observations, physiological measurements, imaging, laboratory investigations, genomics, wearable sensors and other sources of data to simulate future health states and support personalised clinical decision-making.
The ambitions of digital twins are necessarily individual. Their purpose is to assist the management of one patient by predicting that patient’s future trajectory under different therapeutic options.
Synthetic clinical journeys serve a different purpose. They are population-derived educational and research resources rather than patient-specific predictive models. They seek to preserve collective clinical experience rather than model a single individual in real time.
The distinction is analogous to the difference between an individual weather forecast and a climate model. One seeks to predict the future behaviour of a particular system. The other seeks to understand the broader patterns that characterise an entire population. Both are valuable, but they answer different questions.
Rather than viewing synthetic clinical journeys and digital twins as competing technologies, it is more useful to regard them as complementary components of an emerging computational ecosystem.
Digital twins require extensive prior knowledge about how diseases evolve. Synthetic clinical journeys provide precisely such longitudinal representations of disease progression and clinical reasoning. Conversely, insights generated through future digital twin implementations may enrich repositories of structured clinical journeys, creating a continuous cycle in which authentic patients, structured journeys, synthetic journeys and patient-specific computational models progressively refine one another.
Ultimately, these developments point towards a broader vision of computational medicine. The objective is no longer simply to digitise healthcare, nor merely to generate privacy-preserving datasets. It is to create computational representations of medicine that faithfully capture how diseases evolve, how clinicians reason and how knowledge itself accumulates through clinical practice.
Within this framework, the evolution of medical knowledge representation may be viewed as a sequence of increasingly sophisticated computational abstractions. Narrative case histories preserve clinical stories. Electronic health records preserve clinical information. Structured clinical journeys preserve clinical reasoning. Synthetic clinical journeys preserve collective clinical experience. Digital twins personalise that experience to individual patients. Together, these representations provide the foundations for learning health systems capable of continuously integrating evidence, education and patient care.
The significance of this progression extends well beyond artificial intelligence. It suggests that medicine is entering a new phase in which clinical knowledge is no longer stored merely as information but represented as an evolving, computable model of clinical reality. If realised responsibly, such representations could transform not only research and decision support but also the way future generations of clinicians learn, reason and continually improve the practice of medicine.
Conceptual Insight 3: Successive Computational Representations of Medicine
Medicine is progressively evolving through increasingly sophisticated representations:
Narrative → Electronic Health Record → Structured Clinical Journey → Synthetic Clinical Journey → Digital Twin → Learning Health System
These should not be viewed as competing technologies.
Rather, each represents a higher level of computational abstraction, preserving progressively richer representations of clinical reality.
Figure 2. Parallel and complementary evolutions in computational medicine. Clinical knowledge evolves through progressively richer computational representations, beginning with the real patient and progressing through narrative clinical stories, electronic health records and structured clinical journeys. Structured clinical journeys provide a common computational foundation from which two complementary pathways emerge. The first is population-oriented computational medicine, in which synthetic clinical journeys support medical education, artificial intelligence development, clinical research and simulation while preserving patient privacy. The second is individual-oriented computational medicine, in which digital twins enable personalised prediction, decision support, disease modelling and precision healthcare. These complementary pathways converge within learning health systems, where continuous feedback from clinical practice, research, education and computational models progressively refines clinical knowledge and improves future patient care. The figure emphasises that synthetic clinical journeys and digital twins are not successive stages but complementary representations built upon a shared structured knowledge architecture, together advancing the broader vision of computational medicine.
Discussion: Towards a Learning Health System Built on Clinical Journeys
The concepts explored in this paper extend beyond the generation of synthetic health data. They point towards a broader transformation in the way clinical knowledge is created, preserved, shared and continuously refined. If structured clinical journeys can faithfully represent the evolution of disease and the reasoning that accompanies clinical decision-making, they have the potential to become foundational building blocks of future learning health systems.
The idea of a learning health system is not new. It describes a healthcare ecosystem in which knowledge generated during routine clinical care is continuously analysed, translated into improved practice and fed back into patient care. Despite considerable progress in health informatics, however, this vision has often been constrained by the limitations of conventional electronic health records. Although EHRs capture enormous quantities of clinical data, they rarely capture the reasoning that links observations to decisions. Consequently, much of the intellectual process of medicine remains invisible to computational systems.
Structured clinical journeys offer an opportunity to narrow this gap. By preserving chronology, context, diagnostic uncertainty, evolving hypotheses and therapeutic decision-making, they transform routine clinical encounters into reusable units of clinical knowledge. Rather than serving solely as records of care, they become structured representations of clinical experience that can support education, research, quality improvement and artificial intelligence.
This distinction has important implications for medical education. Traditionally, the quality of clinical training has depended heavily upon the diversity of patients encountered during rotations. Exposure to uncommon diseases, atypical presentations or complex decision-making is often determined as much by chance as by curriculum design. Consequently, learners graduating from different institutions—or even from the same institution in different years—may acquire markedly different clinical experiences.
Synthetic clinical journeys offer the possibility of reducing this variability. Instead of relying exclusively on opportunistic patient encounters, educators could curate libraries of clinically realistic journeys covering common conditions, uncommon presentations, diagnostic dilemmas and evolving management strategies. Every learner could therefore experience a broader and more representative spectrum of clinical medicine while still recognising that no simulation can fully replace the insights gained through caring for real patients.
The educational value of such journeys extends beyond knowledge acquisition. Because they explicitly preserve clinical reasoning, learners can compare alternative approaches to diagnosis and management, reflect upon cognitive biases, recognise diagnostic uncertainty and appreciate how experienced clinicians revise their thinking as new information emerges. In this respect, structured journeys support not merely the teaching of medicine but the teaching of clinical thinking.
These same characteristics make structured clinical journeys attractive for artificial intelligence development. Much current work in clinical AI focuses on prediction: estimating risk, detecting abnormalities or recommending treatments. While these applications are undoubtedly valuable, they often operate as statistical classifiers with limited visibility into the reasoning processes that underpin expert clinical judgement. Systems trained on richly structured clinical journeys may instead learn relationships between observations, hypotheses and decisions, enabling future AI applications that are not only more accurate but also more interpretable and educationally valuable.
At the same time, it is important to acknowledge the limitations of computational representations. No structured model, however sophisticated, can capture every dimension of clinical practice. Empathy, communication, ethical judgement, cultural context and the therapeutic relationship remain central to medicine and cannot be fully represented through data structures alone. Similarly, entirely novel diseases and previously unrecognised complications will always originate from observations made in real patients. Synthetic systems can preserve and disseminate accumulated knowledge, but they cannot independently generate clinical discoveries that have never before been observed.
These limitations should not be regarded as failures. Rather, they define the complementary roles of authentic clinical practice and computational medicine. Real patients remain the source of medical knowledge. Structured clinical journeys organise that knowledge. Synthetic clinical journeys make it more accessible while protecting privacy. Digital twins personalise it to individual patients. Learning health systems continuously refine it through ongoing clinical experience. Each representation serves a distinct purpose, yet each depends upon the integrity of the others.
The success of this broader vision will depend not only upon technological innovation but also upon governance. Clinical journeys intended for education or artificial intelligence must remain clinically accurate, transparent in their provenance and subject to rigorous expert review. Privacy-preserving methods should be accompanied by clear standards for validation, quality assurance and appropriate use. Above all, clinicians must remain central to the design and evaluation of these systems. Computational representations should augment clinical judgement, not replace it.
Viewed in this broader context, synthetic health data becomes more than a solution to the problem of confidentiality. It becomes part of a larger intellectual project: developing computational representations that preserve the knowledge embedded within clinical practice while making that knowledge more widely available for learning, research and patient care.
Practical Implications for the PAJR Architecture
Although the concepts discussed in this paper have broad relevance to computational medicine, they also have direct implications for the future evolution of the Patient Journey Record (PAJR) architecture. Rather than viewing PAJR Layer 2 simply as a repository of structured clinical journeys, the framework proposed here suggests that it should be regarded as a foundational knowledge infrastructure capable of supporting multiple downstream applications across education, research, artificial intelligence and personalised healthcare.
The principal contribution of PAJR Layer 2 lies not merely in structuring clinical information but in structuring clinical understanding. Conventional electronic health records preserve observations about patients. PAJR Layer 2 seeks to preserve the evolution of illness together with the clinical reasoning that accompanies it. Chronology, diagnostic hypotheses, differential diagnoses, uncertainty, investigations, therapeutic decisions and clinical outcomes therefore become knowledge objects in their own right rather than incidental annotations within a medical record. This distinction substantially expands the potential value of the architecture beyond its immediate educational purpose.
The framework presented in this paper further suggests that synthetic data generation within PAJR should operate at the level of complete clinical journeys rather than isolated patient records. The objective should not be to generate statistically plausible laboratory values or diagnoses in isolation, but to create coherent longitudinal journeys in which physiology, chronology, investigations, therapeutic interventions and clinical reasoning remain internally consistent. Accordingly, evaluation of synthetic clinical journeys should extend beyond conventional measures of statistical similarity and privacy protection to include clinical fidelity, temporal coherence, physiological plausibility, reasoning consistency and educational authenticity.
The discussion surrounding rare diseases and unusual presentations also has important implications for repository governance. Not every clinical journey should necessarily be managed in the same manner. Common clinical journeys, representing frequently encountered diseases and established patterns of management, are particularly well suited for synthetic generation and broad dissemination because they provide rich educational material while posing comparatively low risks of patient re-identification. These journeys can form the core educational corpus for simulation-based learning, artificial intelligence training, quality improvement and methodological research.
By contrast, rare, complex or scientifically important clinical journeys require a different governance approach. Their value often lies precisely in their uniqueness. Unusual presentations, novel syndromes, unexpected complications and diagnostically challenging cases have historically contributed disproportionately to advances in medical knowledge. At the same time, their distinctive characteristics may increase the likelihood of re-identification and make faithful synthetic generation more challenging. Rather than treating these journeys identically, PAJR could adopt enhanced governance for such high-value cases, combining expert clinical curation, controlled access and carefully validated computational use while ensuring that their educational and scientific significance is preserved.
This distinction should not be interpreted as creating two separate repositories. Rather, it recognises that different forms of clinical knowledge require different stewardship strategies. Common journeys primarily support large-scale education, algorithm development and synthetic data generation. Rare and exceptional journeys primarily support discovery, advanced clinical education, specialist training and the continued refinement of computational models. Together, they create a balanced ecosystem in which both the breadth and the depth of clinical experience are preserved.
The conceptual framework proposed in this paper also provides a practical roadmap for the future evolution of PAJR. The current Layer 2 architecture establishes the essential foundation by transforming narrative clinical experience into structured clinical journeys. This foundation naturally enables subsequent stages of development. The first is the generation of synthetic clinical journeys, providing privacy-preserving resources for education, research and artificial intelligence. Beyond this lies the development of counterfactual clinical journeys, allowing learners and researchers to explore how different clinical decisions or events might plausibly alter patient trajectories. As repositories mature and computational models become increasingly sophisticated, these structured journeys may contribute to the development of explicit clinical world models capable of representing disease behaviour across populations. Ultimately, such knowledge infrastructures could complement patient-specific digital twins and contribute to continuously learning health systems that integrate education, research and clinical care.
Importantly, these stages should be viewed as an evolutionary progression rather than independent technologies. Each depends upon the integrity of the preceding stage. Without accurately structured clinical journeys, synthetic journey generation lacks clinical fidelity. Without high-quality synthetic journeys, counterfactual simulation becomes unreliable. Without robust computational representations of disease behaviour, digital twins remain limited in their ability to support clinical reasoning and personalised prediction. PAJR therefore occupies a foundational position within this progression because it preserves the structured clinical knowledge upon which increasingly sophisticated computational models can be built.
Perhaps the most important implication is conceptual rather than technical. PAJR should not be viewed simply as a repository of cases, nor merely as a source of training data for artificial intelligence. Its broader significance lies in providing a structured representation of collective clinical experience that can be reused, refined and continuously expanded over time. As additional Patient Journey Records are curated, validated and integrated, the repository evolves from a collection of individual cases into an increasingly comprehensive computational representation of how diseases evolve, how clinicians reason, how decisions are made under uncertainty and how medical knowledge itself accumulates through experience.
Viewed in this way, the Patient Journey Record (PAJR) is not simply another health information architecture. It represents an emerging knowledge infrastructure for computational medicine—one capable of connecting authentic clinical practice with synthetic clinical journeys, explainable artificial intelligence, medical education, digital twins and continuously learning health systems within a single, coherent and extensible framework.
Conclusion
Electronic health records transformed healthcare by digitising clinical information. Yet information alone is not synonymous with knowledge. The practice of medicine depends upon understanding how diseases evolve, how clinicians interpret uncertainty and how decisions change as new evidence emerges. Much of this intellectual process remains only partially represented within conventional health records.
Structured clinical journeys address this limitation by making chronology, clinical reasoning and decision-making explicit. In doing so, they create a richer computational representation of medicine than is possible with conventional record-centred approaches. They also provide a natural foundation for generating synthetic clinical journeys that preserve the behaviour of disease while protecting patient privacy.
The significance of this evolution extends beyond synthetic data. Structured clinical journeys form a bridge between traditional electronic health records and emerging concepts such as clinical world models, counterfactual simulation, digital twins and learning health systems. Together, these represent successive stages in the computational representation of clinical reality rather than competing technologies. Each builds upon the preceding one while extending the ability of medicine to preserve, share and apply accumulated clinical knowledge.
The central proposition of this paper is therefore straightforward. The future of synthetic health data should not be viewed primarily through the lens of privacy preservation. Its greater contribution may lie in enabling computational representations that faithfully capture the evolution of disease, the reasoning of clinicians and the collective experience of healthcare itself.
If realised thoughtfully and governed responsibly, structured clinical journeys may become more than an intermediate data format. They may provide the intellectual infrastructure upon which future systems for medical education, clinical research, trustworthy artificial intelligence and personalised healthcare are built. In that sense, the most important innovation is not the generation of synthetic patients. It is the creation of computational models that preserve and transmit the living knowledge of medicine itself.
References (Vancouver Style)
- Charon R. Narrative medicine: a model for empathy, reflection, profession, and trust. JAMA. 2001;286(15):1897–1902. doi:10.1001/jama.286.15.1897. Available from: https://doi.org/10.1001/jama.
286.15.1897 - Charon R. Narrative Medicine: Honoring the Stories of Illness. New York: Oxford University Press; 2006. ISBN: 9780195166750.
- Schmidt HG, Norman GR, Boshuizen HPA. A cognitive perspective on medical expertise: theory and implication. Acad Med. 1990;65(10):611–621. doi:10.1097/00001888-
199010000-00001. - Croskerry P. A universal model of diagnostic reasoning. Acad Med. 2009;84(8):1022–1028. doi:10.1097/ACM.
0b013e3181ace703. Available from: https://doi.org/10.1097/ACM. 0b013e3181ace703 - Norman GR, Eva KW. Diagnostic error and clinical reasoning. Med Educ. 2010;44(1):94–100. doi:10.1111/j.1365-2923.2009.
03507.x. - Institute of Medicine. Best Care at Lower Cost: The Path to Continuously Learning Health Care in America. Washington (DC): National Academies Press; 2013. doi:10.17226/13444. Available from: https://doi.org/10.17226/13444
. - Friedman CP, Rubin JC, Sullivan KJ. Toward an information infrastructure for global health improvement. Yearb Med Inform. 2017;26(1):16–23. doi:10.15265/IY-2017-007.
- Friedman CP, Allee NJ, Delaney BC, et al. The science of Learning Health Systems: foundations for a new journal. Learn Health Syst. 2017;1(1):e10020. doi:10.1002/lrh2.10020.
- Giuffrè M, Shung DL. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy. npj Digit Med. 2023;6:186. doi:10.1038/s41746-023-00927-
3. Available from: https://doi.org/10.1038/ s41746-023-00927-3 - Guillaudeux M, Rousseau O, Petot J, et al. Patient-centric synthetic data generation, no reason to risk re-identification in biomedical data analysis. npj Digit Med. 2023;6:37. doi:10.1038/s41746-023-00771-
5. Available from: https://doi.org/10.1038/ s41746-023-00771-5. - Yan C, Wang S, et al. A multifaceted benchmarking of synthetic electronic health record generation models. Nat Commun. 2022;13:7609. doi:10.1038/s41467-022-35295-
1. - Venkatesh KP, Raza MM, Kvedar JC. Health digital twins as tools for precision medicine: considerations for computation, implementation, and regulation. npj Digit Med. 2022;5:177. doi:10.1038/s41746-022-00694-
7. - Katsoulakis E, Wang MD, et al. Digital twins for health: a scoping review. npj Digit Med. 2024;7. doi:10.1038/s41746-024-01073-
0. - Cook DA, Triola MM. Virtual patients: a critical literature review and proposed next steps. Med Educ. 2009;43(4):303–311. doi:10.1111/j.1365-2923.2008.
03286.x. - Schmidt HG, Norman GR, Mamede S, Magzoub M. The influence of context on diagnostic reasoning: a narrative synthesis of experimental findings. J Eval Clin Pract. 2024. doi:10.1111/jep.14023.
- Johnson KB, Wei WQ, Weeraratne D, et al. Precision medicine, AI, and the future of personalized healthcare. npj Digit Med. 2021;4:1–8. doi:10.1038/s41746-021-00460-
6. - McDuff D, Curran T, Kadambi A. Synthetic Data in Healthcare. arXiv. 2023. Available from: https://arxiv.org/abs/2304.
03243.
This set provides a solid foundation spanning narrative medicine, clinical reasoning, learning health systems, synthetic health data, digital twins, computational medicine, and AI in healthcare.
