2018
pdf
bib
abs
Guess Me if You Can: Acronym Disambiguation for Enterprises
Yang Li
|
Bo Zhao
|
Ariel Fuxman
|
Fangbo Tao
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Acronyms are abbreviations formed from the initial components of words or phrases. In enterprises, people often use acronyms to make communications more efficient. However, acronyms could be difficult to understand for people who are not familiar with the subject matter (new employees, etc.), thereby affecting productivity. To alleviate such troubles, we study how to automatically resolve the true meanings of acronyms in a given context. Acronym disambiguation for enterprises is challenging for several reasons. First, acronyms may be highly ambiguous since an acronym used in the enterprise could have multiple internal and external meanings. Second, there are usually no comprehensive knowledge bases such as Wikipedia available in enterprises. Finally, the system should be generic to work for any enterprise. In this work we propose an end-to-end framework to tackle all these challenges. The framework takes the enterprise corpus as input and produces a high-quality acronym disambiguation system as output. Our disambiguation models are trained via distant supervised learning, without requiring any manually labeled training examples. Therefore, our proposed framework can be deployed to any enterprise to support high-quality acronym disambiguation. Experimental results on real world data justified the effectiveness of our system.
2017
pdf
bib
abs
Identifying Semantically Deviating Outlier Documents
Honglei Zhuang
|
Chi Wang
|
Fangbo Tao
|
Lance Kaplan
|
Jiawei Han
Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing
A document outlier is a document that substantially deviates in semantics from the majority ones in a corpus. Automatic identification of document outliers can be valuable in many applications, such as screening health records for medical mistakes. In this paper, we study the problem of mining semantically deviating document outliers in a given corpus. We develop a generative model to identify frequent and characteristic semantic regions in the word embedding space to represent the given corpus, and a robust outlierness measure which is resistant to noisy content in documents. Experiments conducted on two real-world textual data sets show that our method can achieve an up to 135% improvement over baselines in terms of recall at top-1% of the outlier ranking.
pdf
bib
Life-iNet: A Structured Network-Based Knowledge Exploration and Analytics System for Life Sciences
Xiang Ren
|
Jiaming Shen
|
Meng Qu
|
Xuan Wang
|
Zeqiu Wu
|
Qi Zhu
|
Meng Jiang
|
Fangbo Tao
|
Saurabh Sinha
|
David Liem
|
Peipei Ping
|
Richard Weinshilboum
|
Jiawei Han
Proceedings of ACL 2017, System Demonstrations
2016
pdf
bib
Cross-media Event Extraction and Recommendation
Di Lu
|
Clare Voss
|
Fangbo Tao
|
Xiang Ren
|
Rachel Guan
|
Rostyslav Korolov
|
Tongtao Zhang
|
Dongang Wang
|
Hongzhi Li
|
Taylor Cassidy
|
Heng Ji
|
Shih-fu Chang
|
Jiawei Han
|
William Wallace
|
James Hendler
|
Mei Si
|
Lance Kaplan
Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations