A Focused Crawler Pluging for Heritrix Based on Language Modelling
Topicrawler is a plugin for Heritrix
curl -fsSL https://raw.githubusercontent.com/tudarmstadt-lt/topicrawler/master/install.sh | bash -e
Please use the following citation if you use the topicrawler for your project. Also we would like to hear from you if you use or plan to use the crawler.
Steffen Remus and Chris Biemann (2016): Domain-Specific Corpus Expansion with Focused Webcrawling. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), Protorož, Slovenia (bib)
@InProceedings{REMUS16.316,
author = {Steffen Remus and Chris Biemann},
title = {Domain-Specific Corpus Expansion with Focused Webcrawling},
booktitle = {Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)},
year = {2016},
month = {may},
date = {23-28},
location = {Portorož, Slovenia},
editor = {Nicoletta Calzolari (Conference Chair) and Khalid Choukri and Thierry Declerck and Marko Grobelnik and Bente Maegaard and Joseph Mariani and Asuncion Moreno and Jan Odijk and Stelios Piperidis},
publisher = {European Language Resources Association (ELRA)},
address = {Paris, France},
isbn = {978-2-9517408-9-1},
language = {english}
}
Having trouble? Check out our documentation and don't hesitate to contact us.