Menu
OR
OR
Use of natural language processing to extract uterine weight from pathology reports

Obstetrics & Gynecology

...

9 October, 2026

Int J Gynaecol Obstet. 2026 Oct 9. doi: 10.1002/ijgo.71384. Online ahead of print.

ABSTRACT

INTRODUCTION: For patients requiring hysterectomy, uterine weight is an important factor affecting surgical complexity and surgical outcomes. However, uterine weight is not readily available for analysis in administrative databases. In the present study, we use natural language processing to extract uterine weight from narrative pathology text.

METHODS: Pathology text was obtained for 5000 patients sampled ±3 days of hysterectomy from hospital administrative datasets. The gross pathology subsections that describe the submitted tissues were retained for analysis. Manual annotation was performed to create a gold standard dataset. The texts were pre-processed, tokenized (split into words and punctuation), and tagged with parts-of-speech (e.g., noun, verb, number, adverb). Numeric tokens were retained for analysis. Feature engineering (e.g., variable construction and selection) included creating subspecimen flags, three tokens before and after the numeric tokens, weight units, and the prefix "weigh." Models included classification and regression trees (CART), extreme gradient boosting (XGB), and CatBoost, trained on 75%, validated on 15%, and tested in a 10% hold-out dataset. Performance on a real-world application was assessed for all hysterectomies performed over a 4-year period.

RESULTS: Trained on all numeric tokens using features determined a priori, CART had 99.1% precision and 100% recall, with eight false positives. XGB performed similarly (99.3% precision; 99.9% recall). Performance was best using CatBoost's ordered target encoding and all values of the previous and subsequent three tokens (99.8% precision; 100% recall). Applied to the test dataset, CART slightly outperformed CatBoost with two fewer false positives. Summing all weights in the test set reports (n = 474), the extracted total specimen weight was a mean 9.4 g higher than the truth, owing predominantly to duplicate weights (22/37 discrepant reports). Application in the real-world dataset (n = 57 546) identified a small number of reports with data quality issues specific to low-weight specimens that should be excluded (≤10 g) or subjected to manual adjudication (>10 to ≤15 g).

CONCLUSION: Natural language processing followed by simple CART can accurately extract total uterine weight from narrative pathology reports. Further work is needed to extract specific anatomic weights.

PMID:42852604 | DOI:10.1002/ijgo.71384

Read Full Article

Journal Source :

International Journal of Gynecology & Obstetrics

© 2026 Imedsource, All Rights Reserved

Sustain the Knowledge Stream!

For unrestricted access to content and a seamless experience,

Continue to Login

Don't have an account?

We’d love to get to know you better!

Complete your profile to help us deliver a more personalized experience.