Semantic MediaWiki (SMW) is a free extension of MediaWiki that helps to search, organise, tag, browse, evaluate, and share the wiki's content. While traditional wikis contain only texts which computers can neither understand nor evaluate, SMW adds semantic annotations that bring the power of the Semantic Web to the wiki.
MuNPEx is a multi-lingual noun phrase (NP) extraction component developed for the GATE architecture, implemented in JAPE. It currently supports English, German, French, and Spanish (in beta).
MuNPEx requires a part-of-speech (POS) tagger to work and can additionally use detected named entities (NEs) to improve chunking performance. Please read the documentation (or source code) for more details.
Our goal is to develop a probabilistic knowledge base that mirrors the content of the web. We are developing a system that uses semi-supervised learning methods to learn to extract symbolic knowledge from unstructured text and HTML. We are exploring methods of continous learning, where our system runs 24x7, continuously learning to read better, and continuously extracting facts from the web.
Webstemmer is a web crawler and HTML layout analyzer that automatically extracts main text of a news site without having banners, ads and/or navigation links mixed up
M. Mintz, S. Bills, R. Snow, и D. Jurafsky. Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, стр. 1003--1011. Suntec, Singapore, Association for Computational Linguistics, (августа 2009)
M. Vargas-Vera, и D. Celjuska. WI '04: Proceedings of the 2004 IEEE/WIC/ACM International Conference on Web Intelligence, стр. 615--618. Washington, DC, USA, IEEE Computer Society, (2004)
R. Swan, и J. Allan. CIKM '99: Proceedings of the eighth international conference on Information and knowledge management, стр. 38--45. New York, NY, USA, ACM, (1999)
E. Riloff, C. Schafer, и D. Yarowsky. Proceedings of the 19th international conference on Computational linguistics, стр. 1--7. Morristown, NJ, USA, Association for Computational Linguistics, (2002)
E. Riloff, и R. Jones. AAAI '99/IAAI '99: Proceedings of the sixteenth national conference on Artificial intelligence and the eleventh Innovative applications of artificial intelligence conference innovative applications of artificial intelligence, стр. 474--479. Menlo Park, CA, USA, American Association for Artificial Intelligence, (1999)
Y. Li, R. Krishnamurthy, S. Raghavan, S. Vaithyanathan, и H. Jagadish. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, стр. 21--30. Honolulu, Hawaii, Association for Computational Linguistics, (октября 2008)