Trafilatura Core Python library
Requires Python 3.10 or newer:
pip install trafilaturacore
Save this as clean.py:
from trafilaturacore import clean
result = clean(
"<nav>Home</nav><p>Saved article text.</p>",
boilerplate="keep",
)
print(result.html)
python clean.py
It prints <p>Saved article text.</p>. Use the default balanced mode to extract
main content from a full article. Python provides library APIs only.
Options and results
clean() accepts strings or UTF-8 bytes and returns a CleanResult with html,
messages, and optional metadata. declared_page_type mirrors the metadata's
declaredPageType; this page declaration does not steer extraction.
| Option | Values and behavior |
|---|---|
boilerplate | precision, balanced (default), recall, keep |
image_handling | include (default), exclude, alt-text, resolved-url |
link_handling, table_handling, comment_handling | include (default) or exclude |
url | Absolute HTTP(S) metadata/image context; never fetched |
config | Validated cleaning dictionary |
max_input_bytes | Positive UTF-8 byte limit; default 10 MiB, hard ceiling 64 MiB |
See the content controls. Alt text falls back through a single-image figure's caption, ARIA text, and title; explicit empty alt removes decorative images. Resolved URLs support lazy sources and srcset. Extraction may remove attributes before these transforms.
Async use
aclean() accepts the same options and performs cleaning in a worker thread:
import asyncio
from trafilaturacore import aclean
async def main():
result = await aclean("<p>Saved article text.</p>", boilerplate="keep")
print(result.html)
asyncio.run(main())
Cancelling the await does not stop an already running worker. Use process isolation if your application needs hard deadlines.
Custom cleaning and limits
The JSON keys are allowedTags, allowedAttributes, allowedClasses,
nonTextTags, transformTags, and selfClosing. Supplied fields replace
native defaults; omitted fields retain them. Attribute/class allowlists accept
shell-style wildcards. selfClosing accepts standard HTML void elements only.
Policies still remove scripts, embedding elements, SVG/MathML, stylesheet elements, event handlers, refresh metadata, and disallowed URL schemes. Inline styles retain limited presentation properties and reject CSS URLs, expressions, comments, and escapes. Remote links and images may remain. Sanitize untrusted output for its rendering context and apply a Content Security Policy.
Invalid options raise TypeError or ValueError. Resource failures raise
ResourceLimitError with code ERR_TRAFILATURACORE_RESOURCE_LIMIT, without
falling back. Ordinary failed/empty extraction can fall back to whole-document
cleanup with diagnostics. Native limits include 128 nesting levels, 100,000 source
nodes, 256 attributes per tag, 200,000 aggregate attributes, 50,000 expanded table
cells, and 32 MiB intermediate/output HTML.
Language differences
Python uses lxml/nh3 plus native date and URL adapters. Parser repair, fragment serialization, custom-config defaults, CSS preservation, diagnostics, and resource limits can differ from TypeScript. htmldate recognizes additional date formats; Courlan handles URL/domain normalization. Metadata image strings are cleaned, and the first raw OpenGraph type declaration is retained.
Dependencies install separately and may need build prerequisites where platform wheels are unavailable. See the npm library for TypeScript or the playground to generate Python code.
Updated: September 25, 2026