Video walkthrough

Methodological issues of the electronic health records’ use in the context of epidemiological investigations, in light of missing data: a review of the recent literature

Student 6:43 CC AI

paperi.ai
0:00 / 0:00

Thomas Tsiampalis, Demosthenes B. Panagiotakos

Electronic health records can improve prevention, treatment, research, and even pandemic response—but the review finds that missing data are often handled inconsistently, leaving important uncertainty behind.

Abstract

Background Electronic health records (EHRs) are widely accepted to enhance the health care quality, patient monitoring, and early prevention of various diseases, even when there is incomplete or missing information in them. Aim The present review sought to investigate the impact of EHR implementation on healthcare quality and medical decision in the context of epidemiological investigations, considering missing or incomplete data. Methods Google scholar, Medline (via PubMed) and Scopus databases were searched for studies investigating the impact of EHR implementation on healthcare quality and medical decision, as well as for studies investigating the way of dealing with missing data, and their impact on medical decision and the development process of prediction models. Electronic searches were carried out up to 2022. Results EHRs were shown that they constitute an increasingly important tool for both physicians, decision makers and patients, which can improve national healthcare systems both for the convenience of patients and doctors, while they improve the quality of health care as well as they can also be used in order to save money. As far as the missing data handling techniques is concerned, several investigators have already tried to propose the best possible methodology, yet there is no wide consensus and acceptance in the scientific community, while there are also crucial gaps which should be addressed. Conclusions Through the present thorough investigation, the importance of the EHRs’ implementation in clinical practice was established, while at the same time the gap of knowledge regarding the missing data handling techniques was also pointed out.

Transcript

Electronic health records can improve prevention, treatment, research, and even pandemic response—but the review finds that missing data are often handled inconsistently, leaving important uncertainty behind. Electronic health records are widely accepted to enhance health care quality, patient monitoring, and early prevention of various diseases, even when they contain incomplete or missing information.

Electronic health records remain useful for improving health care quality, monitoring patients, and preventing disease early, even when their data are incomplete or missing. The review presents the challenges faced during the use of EHRs for epidemiological investigations in the context of missing data.

It also discusses the most frequent statistical methodologies used for handling missing information and confronting that obstacle to derive valid conclusions. The review was conducted according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses, known as PRISMA.

Included designs were case studies, cohort studies, cross-sectional studies, retrospective case-control studies, prospective cohort studies, and cluster-randomized controlled trials published in English. Systematic reviews and meta-analyses were excluded, although they assisted in retrieving articles not allocated in the search process.

The electronic and manual searches initially identified one thousand nine hundred seventy-two references: three hundred thirteen from PubMed, five hundred nineteen from Scopus, and one thousand one hundred forty from Google Scholar. A total of seventeen studies were included in the narrative review, and those studies were divided into two categories.

Table two traces the review’s study-selection process: the authors identified one thousand nine hundred seventy-two records across PubMed, Scopus, and Google Scholar, removed twenty duplicates, and screened one thousand nine hundred fifty-two titles and abstracts. After excluding one thousand eight hundred ninety-seven irrelevant records and thirty-eight that could not be retrieved, seventeen studies remained.

These were organized into two categories: EHRs and improvement of medical quality and health system, with eight studies, and missing data in the context of EHRs, with nine. Electronic health records can assist in both the prevention and the treatment of disease.

In one example, EHR data supported rules for diagnosis coding of chronic kidney disease at the hospital of Saint Etienne. In bacteremia, EHRs supported prognosis and early diagnosis, while Machine Learning techniques predicted blood-culture results for timely administration of the correct treatment and reduced medical costs.

Other examples included identifying common treatment pathways, predicting multi-type major adverse cardiovascular events, and improving smoking-status documentation and counseling assistance to smokers. Table three synthesizes studies showing how electronic health records can support medical quality and health-system performance.

The reported applications range from using medical-image selfies for clinical insights, to directing referrals, checking chronic kidney disease diagnosis codes, predicting bacteremia, and identifying treatment pathways. It also describes benefits such as improved access to charts, recommended care, appropriate testing, and smoking-status documentation, illustrating that EHRs support both clinical decision-making and system organization.

The reviewed literature contains a variety of approaches toward managing missing EHR data, but only fifty-eight of ninety studies, or sixty-four percent, addressed missing data before analysis. Some simple approaches select sub-datasets containing complete information or use stratified mean imputation, while advanced methods interpolate longitudinal variables under limited conditions.

Few studies used informative observations, where the presence of a variable is meaningful for possibly missing values. A deep learning unsupervised method for imputing missing values in patient records significantly reduced imputation biases under various scenarios when compared with four other imputation techniques.

Figure one maps the main imputation approaches identified in the review for electronic health records. Around the central topic, it shows mean-value replacement combined with an autoencoder, deep-learning unsupervised methods, statistical methods for continuous measures, complete-information sub-datasets, data-driven approaches moving from denser to sparser EHRs, and the use of informative observations.

This matters because the review emphasizes that missingness in routinely collected health data can be widespread and informative, with no broadly accepted best method. Important EHR benefits include easy access to computerized records and the elimination of poor penmanship, a widespread and significant obstacle in the medical world.

The release of EHR data to patients through smart apps was associated with annual savings of approximately two million euros for the hospital and one million euros for patients. EHR use can substantially reduce redundant medical tests and the need to mail hard copies of test results to different providers.

Electronically stored data increases data availability, improves the ability to conduct research, and facilitates identification of evidence-based best health practices. Although EHRs have drawbacks when used solely as data sources for studies informing public health decisions, they contain crucial data elements that help with pandemic response.

The review has limitations because there is no well-established metric to evaluate EHR performance in clinical practice. Therefore, no quantitative assessment of EHR cost-effectiveness in medical decision making could be performed. No pooled analysis or quality assessment of the reviewed studies was performed because it was outside the scope of the work and, in many cases, was not feasible.

EHRs are valuable for clinical care and epidemiological research, but missing-data methods remain diverse and lack wide scientific consensus, while the review itself cannot quantitatively compare performance or cost-effectiveness.

A derivative work by Paperi · AI-generated script, voice and captions · pages and figures unaltered

Made with Paperi.

Drop in a research PDF — get a narrated video walkthrough like this one, with highlights that follow the narration. Free to start.

Try it with your paper →

More papers

Hybrid Threshold Denoising Framework Using Singular Value Decomposition for Side-Channel Analysis Preprocessing 3:54

Hybrid Threshold Denoising Framework Using Singular Value Decomposition for Side-Channel Analysis Preprocessing

A device can reveal its encryption key without saying a word. Tiny electrical traces carry clues about the secret, but noise can bury them before anyone can read them.

A high-efficiency elementary network of interchangeable superconducting qubit devices 3:33

A high-efficiency elementary network of interchangeable superconducting qubit devices

What if a quantum computer did not have to be built as one giant, delicate object? This experiment shows that separate quantum devices can be connected by a cable, unplugged, and still exchange information with about one percent loss.

Resting-state occipito-frontal alpha connectome is linked to differential word learning ability in adult learners 3:14

Resting-state occipito-frontal alpha connectome is linked to differential word learning ability in adult learners

Why can one adult learn a new word almost effortlessly while another struggles? This study suggests the answer may partly lie in how two distant parts of the resting brain keep time together.

All 2 papers in Computer Science →