Our goal is to develop a probabilistic knowledge base that mirrors the content of the web. We are developing a system that uses semi-supervised learning methods to learn to extract symbolic knowledge from unstructured text and HTML. We are exploring methods of continous learning, where our system runs 24x7, continuously learning to read better, and continuously extracting facts from the web.
M. Vargas-Vera, and D. Celjuska. WI '04: Proceedings of the 2004 IEEE/WIC/ACM International Conference on Web Intelligence, page 615--618. Washington, DC, USA, IEEE Computer Society, (2004)