Skip to main content

About Trafilatura Core

Trafilatura Core is an open-source extraction engine based on Trafilatura, implemented in TypeScript and Python. Both implementations extract main content by removing boilerplate from supplied HTML. The Core is this single task: HTML goes in, cleaned HTML comes out. The optional Source URL provides metadata and image-resolution context, never a fetch request.

Use the TypeScript library or CLI, Python library, or playground. For crawling and browser rendering, use Markdownee, whose TypeScript and Python versions use the matching Core implementation. For Markdown conversion, use a separate tool such as Turndown.

Lineage and boundaries

The TypeScript extraction engine is a direct port of Python Trafilatura v2.2.0's fast=True extraction path. Trafilatura is the original Python implementation by Adrien Barbaresi; go-trafilatura, by Markus Mobius, served as a DOM translation aid. Metadata extraction also derives from Python Trafilatura.

The fast path excludes upstream's two extractor-comparison stages, so results can differ from its fast=False default. A parity suite compares selected text against the pinned upstream path in precision, balanced, and recall modes. The native Python library translates the TypeScript extraction and metadata logic, with native adapters whose parsing and output differences are documented separately.

Keep mode skips extraction and applies whole-document cleanup. Content controls still operate within the cleaning rules. Output is compact HTML; if presentation formatting fails, cleaned unformatted HTML is returned with a warning.

Publisher

Glueo, s.r.o., a Prague-based software development company, operates Trafilatura Core.

Updated: October 3, 2026