Skip to main content

Trafilatura Core Python library

Requires Python 3.10 or newer:

pip install trafilaturacore

Save this as clean.py:

from trafilaturacore import clean

result = clean(
    "<nav>Home</nav><p>Saved article text.</p>",
    boilerplate="keep",
)
print(result.html)
python clean.py

It prints <p>Saved article text.</p>. Use the default balanced mode to extract main content from a full article. Python provides library APIs only.

Options and results

clean() accepts strings or UTF-8 bytes and returns a CleanResult with html, messages, and optional metadata. declared_page_type mirrors the metadata's declaredPageType; this page declaration does not steer extraction.

OptionValues and behavior
boilerplateprecision, balanced (default), recall, keep
image_handlinginclude (default), exclude, alt-text, resolved-url
link_handling, table_handling, comment_handlinginclude (default) or exclude
urlAbsolute HTTP(S) metadata/image context; never fetched
configValidated cleaning dictionary
max_input_bytesPositive UTF-8 byte limit; default 10 MiB, hard ceiling 64 MiB

See the content controls. Alt text falls back through a single-image figure's caption, ARIA text, and title; explicit empty alt removes decorative images. Resolved URLs support lazy sources and srcset. Extraction may remove attributes before these transforms.

Async use

aclean() accepts the same options and performs cleaning in a worker thread:

import asyncio
from trafilaturacore import aclean

async def main():
    result = await aclean("<p>Saved article text.</p>", boilerplate="keep")
    print(result.html)

asyncio.run(main())

Cancelling the await does not stop an already running worker. Use process isolation if your application needs hard deadlines.

Custom cleaning and limits

The JSON keys are allowedTags, allowedAttributes, allowedClasses, nonTextTags, transformTags, and selfClosing. Supplied fields replace native defaults; omitted fields retain them. Attribute/class allowlists accept shell-style wildcards. selfClosing accepts standard HTML void elements only.

Policies still remove scripts, embedding elements, SVG/MathML, stylesheet elements, event handlers, refresh metadata, and disallowed URL schemes. Inline styles retain limited presentation properties and reject CSS URLs, expressions, comments, and escapes. Remote links and images may remain. Sanitize untrusted output for its rendering context and apply a Content Security Policy.

Invalid options raise TypeError or ValueError. Resource failures raise ResourceLimitError with code ERR_TRAFILATURACORE_RESOURCE_LIMIT, without falling back. Ordinary failed/empty extraction can fall back to whole-document cleanup with diagnostics. Native limits include 128 nesting levels, 100,000 source nodes, 256 attributes per tag, 200,000 aggregate attributes, 50,000 expanded table cells, and 32 MiB intermediate/output HTML.

Language differences

Python uses lxml/nh3 plus native date and URL adapters. Parser repair, fragment serialization, custom-config defaults, CSS preservation, diagnostics, and resource limits can differ from TypeScript. htmldate recognizes additional date formats; Courlan handles URL/domain normalization. Metadata image strings are cleaned, and the first raw OpenGraph type declaration is retained.

Dependencies install separately and may need build prerequisites where platform wheels are unavailable. See the npm library for TypeScript or the playground to generate Python code.

Updated: September 25, 2026