Skip to main content

Trafilatura Core playground

Paste HTML, choose settings, and select Clean HTML. The website sends the input to its API and displays cleaned HTML and diagnostics. Copy or download the result; Reset clears it and restores defaults. For a first example, see Getting started.

The optional Source URL supplies metadata and relative-image context. It is never fetched.

Cleaning controls

  • Boilerplate — precision selects more narrowly, recall retains more, and balanced is the default. Keep skips extraction and cleans the whole document.
  • Images — include, exclude, alt-text placeholders without source URLs, or resolved URLs. Resolution uses Source URL or the document's base URL; otherwise relative URLs remain relative.
  • Links — exclude removes the link while keeping its text.
  • Tables — exclude also removes cell text.
  • Comments — detected user-comment sections, not HTML comments.

All four content controls default to include, within extraction and cleaning rules. Image fallback attributes must survive extraction; keep retains more source context.

Generated examples

Select npm CLI, npm library, or Python library for code matching your settings. Empty or whitespace-only input, or HTML longer than 600 characters after trimming, uses the complete “Reading saved pages” sample from the README. Input of up to 600 characters is embedded after trimming surrounding whitespace. Your selected settings apply in either case.

Install and run through the npm CLI, npm library, or Python guide.

Updated: September 26, 2026