TY - JOUR AU - Joseph , Nora AU - Lindblad , Ida AU - Zaker , Sara AU - Elfversson , Sharareh AU - Albinzon , Maria AU - Ødegård , Øyvind AU - Hantler , Li AU - Hellström , Per M. PY - 2022/01/27 Y2 - 2024/03/28 TI - Automated data extraction of electronic medical records: Validity of data mining to construct research databases for eligibility in gastroenterological clinical trials JF - Upsala Journal of Medical Sciences JA - ujms VL - 127 IS - 1 SE - Original Articles DO - 10.48101/ujms.v127.8260 UR - https://ujms.net/index.php/ujms/article/view/8260 SP - AB - Background: Electronic medical records (EMRs) are adopted for storing patient-related healthcare information. Using data mining techniques, it is possible to make use of and derive benefit from this massive amount of data effectively. We aimed to evaluate validity of data extracted by the Customized eXtraction Program (CXP).Methods: The CXP extracts and structures data in rapid standardised processes. The CXP was programmed to extract TNFα-native active ulcerative colitis (UC) patients from EMRs using defined International Classification of Disease-10 (ICD-10) codes. Extracted data were read in parallel with manual assessment of the EMR to compare with CXP-extracted data.Results: From the complete EMR set, 2,802 patients with code K51 (UC) were extracted. Then, CXP extracted 332 patients according to inclusion and exclusion criteria. Of these, 97.5% were correctly identified, resulting in a final set of 320 cases eligible for the study. When comparing CXP-extracted data against manually assessed EMRs, the recovery rate was 95.6–101.1% over the years with 96.1% weighted average sensitivity.Conclusion: Utilisation of the CXP software can be considered as an effective way to extract relevant EMR data without significant errors. Hence, by extracting from EMRs, CXP accurately identifies patients and has the capacity to facilitate research studies and clinical trials by finding patients with the requested code as well as funnel down itemised individuals according to specified inclusion and exclusion criteria. Beyond this, medical procedures and laboratory data can rapidly be retrieved from the EMRs to create tailored databases of extracted material for immediate use in clinical trials. ER -