About This Resource

A library for extracting the main content and metadata from web pages. It removes common page clutter and produces cleaner HTML for reading, indexing, or conversion into other formats.