Title: TEXT: Automatic Template Extraction from Heterogeneous Web Pages is developed in J2EE (Jsp with Mysql) by Mirror Technologies Pvt Ltd -- Vadapalani, Chennai.
Domain: Data Mining
Key Features:
1) World Wide Web is the most useful source of information. In order to achieve high productivity of publishing, the web pages in many websites are automatically populated by using the common templates with contents.
2) Moreover, our proposed algorithms are fully automated and robust without requiring many parameters. Experimental results with real life data sets confirm the effectiveness of our algorithms. Proposed an algorithm to extract a template using not only structural information, but also visual lay out information.
Algorithms:
RTDM: We implemented RTDM since it is the related work having the most similar problem formulation with us. It requires a training data set and the similarity threshold to decide the number of templates.
TEXT-MDL: It is the naive agglomerative clustering algorithm with the approximate entropy model introduced . It requires no input parameter.
TEXT-HASH: It is the agglomerative clustering algorithm with MinHash signatures discussed . It requires an input parameter which is the length of MinHash signature.
TEXT-MAX: It is the clustering algorithm with both MinHash signatures and Heuristic to reduce the search space. It requires the length of the signature as an input parameter.
Visit http://www.learnbenchindia.com/
For more details contact:
LEARN BENCH INDIA
No 73, South Sivan kovil Street,
Vadapalani, Chennai, Tamil Nadu.
Telephone: +91-44-42048874.
Phone: 9381948474, 9382948474
E-Mail: [email protected]