Neuerscheinungen 2011Stand: 2020-01-07 |
Schnellsuche
ISBN/Stichwort/Autor
|
Herderstraße 10 10625 Berlin Tel.: 030 315 714 16 Fax 030 315 714 14 info@buchspektrum.de |
Pavan Hemdev
Information Extraction: A Smart Calendar Application
Using NLP, Computational Linguistics, Machine Learning and Information Retrieval Techniques
2011. 84 S.
Verlag/Jahr: VDM VERLAG DR. MÜLLER 2011
ISBN: 3-639-35305-6 (3639353056)
Neue ISBN: 978-3-639-35305-1 (9783639353051)
Preis und Lieferzeit: Bitte klicken
The amount of information available on the web and other electronic formats is increasing at a rapid rate. Moreover, e-mails are now becoming the preferred mode of communication. This thesis investigates various Information Extraction techniques (Tokenization, POS Tagger, Chunker, NER, Co-reference Resolution) and develops a system that inferences calendar appointments from a user´s e-mail account. More specifically, the system identifies the subject, date and time of an appointment and upon user confirmation enters it into a calendar service. It makes use of an intelligent user feedback mechanism that helps tailor the system towards individual users. A novel approach adopted towards constructing rules to identify entities in the absence of a domain relevant corpus, reinstates the importance of a rule-based approach towards building a Named Entity Recognizer. It allows the system to be easily extended and helps identify unseen patterns without much domain expertise. Finally, the thesis tries to provide a data format that could be used in future systems, paving the way for a world in which devices could truly communicate with each other.
Pavan Hemdev is a Technology Enthusiast with a Msc in Computer Science from Oxford University. Currently, he is engaged in launching his first start-up in the mobile space back home in Mumbai, India. Having participated in social organisations, he would one day like to set up a technology school in India.