Publication:
A tree learning approach to web document sectional hierarchy extraction

Loading...
Thumbnail Image

Date

Institution Authors

Organizational Units

Authors

Pembe, F.Canan

Göngör, Tunga

Advisor

item.page.editor

Editor

Department

Journal Title

Journal ISSN

Volume Title

Publisher

DOI

Research Projects

Organizational Units

Journal Issue

Abstract

There is an increasing availability of documents in electronic form due to the widespread use of the Internet. Hypertext Markup Language (HTML) which is mostly concerned with the presentation of documents is still the most commonly used format on the Web, despite the appearance of semantically richer markup languages such as XML. Effective processing of Web documents has several uses such as the display of content on small-screen devices and summarization. In this paper, we investigate the problem of identifying the sectional hierarchy of a given HTML document together with the headings in the document. We propose and evaluate a learning approach suitable to tree representation based on Support Vector Machines.

Description

Journal or Series

ISSN

ISBN

978-989-674-021-4

Rights

Attribution-NonCommercial-NoDerivs 3.0 United States

Citation

Endorsement

Review

Supplemented By

Referenced By

Creative Commons license

Except where otherwise noted, this item's license is described as Attribution-NonCommercial-NoDerivs 3.0 United States

Related Patent

Related Goal

11
Görüntülenme
0
İndirme
Google Scholar
Scholar'da Ara ↗
Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.