A Proposed Architecture for Continuous Web Monitoring Through Online Crawling of Blogs
This addresses the need for timely information for psychologists, marketers, and political analysts, but it is incremental as it builds on existing focused crawling techniques.
The authors tackled the problem of continuous web monitoring by proposing an architecture for online crawling of blogs, using a focused crawler and a weighted graph based on key phrases to fetch and analyze data, achieving continuous monitoring of the Web space.
Getting informed of what is registered in the Web space on time, can greatly help the psychologists, marketers and political analysts to familiarize, analyse, make decision and act correctly based on the society`s different needs. The great volume of information in the Web space hinders us to continuously online investigate the whole space of the Web. Focusing on the considered blogs limits our working domain and makes the online crawling in the Web space possible. In this article, an architecture is offered which continuously online crawls the related blogs, using focused crawler, and investigates and analyses the obtained data. The online fetching is done based on the latest announcements of the ping server machines. A weighted graph is formed based on targeting the important key phrases, so that a focused crawler can do the fetching of the complete texts of the related Web pages, based on the weighted graph.