Maximum entropy modeling for mining patient medication status from free text.

Serguei V. Pakhomov, Alexander Ruggieri, Christopher G. Chute

Research output: Contribution to journalArticle

Abstract

Using a classification scheme of patient medication status we sought to recognize and categorize medications mentioned in the unrestricted text of clinical documents generated in clinical practice. The categories refer to the patient's status with respect to the medication such as discontinuation, start or initiation, and continuation of a given medication. This categorization is performed with a machine learning technique, Maximum Entropy (ME), that is well suited to incorporating heterogeneous sources of information necessary for classifying patient's medication status. We use hand labeled training data to generate ME models and test 5 different training feature sets. Our results show that the most optimal feature set includes a combination of the following: two words preceding and following the mention of the drug, the subject of the sentence in which the drug mention occurs, the 2 words following the subject, and a binary feature vector of lexicalized semantic cues indicative of medication status or its change. The average predictive power of a model trained on these features is approximately 89%.

Original languageEnglish (US)
Pages (from-to)587-591
Number of pages5
JournalProceedings / AMIA ... Annual Symposium. AMIA Symposium
StatePublished - 2002
Externally publishedYes

ASJC Scopus subject areas

  • Medicine(all)

Fingerprint Dive into the research topics of 'Maximum entropy modeling for mining patient medication status from free text.'. Together they form a unique fingerprint.

  • Cite this