Skip to main content

Trafilatura Core vs Mozilla Readability: how extraction works

Mozilla Readability centers extraction on an Arc90-origin model: score candidate containers, select a promising region, and expand it. Trafilatura and Trafilatura Core organize extraction around ordered structural rules, content pruning, and recovery stages. Both approaches use heuristics; both can retry. The difference is how they organize the search for main content.12

Core is our open-source extraction engine based on Trafilatura, implemented in TypeScript and native Python. Its TypeScript engine ports Python Trafilatura's fast=True path; its maintained native Python implementation translates that TypeScript port, with documented language differences.34 That boundary matters: Core retains internal recovery mechanisms but omits upstream's external readability-lxml and jusText comparisons.5

The implementation details below refer to pinned source snapshots of Mozilla Readability 0.6.0, Python Trafilatura 2.2.0, and Trafilatura Core 0.8.6, rechecked on October 7, 2026. Source links identify the revisions used for this algorithm comparison.165

Readability scores and expands candidates

Mozilla's library powers Firefox Reader View and carries Arc90's original Readability attribution.7 The pinned algorithm includes more than that early design.

After pruning and normalizing the document, Readability scores text blocks using length and punctuation, propagates scores to ancestors, applies tag and class/id weights, and discounts link-heavy candidates. It can promote a shared ancestor of strong candidates and append qualifying siblings. Selection isn't confined to one original container.1

A result below charThreshold triggers another attempt with filters relaxed in sequence: unlikely-node removal, class weighting, then conditional cleaning. If those attempts remain short, Readability takes the longest nonempty result. Calling it single-pass would be wrong; its retries remain centered on candidate scoring.1

Trafilatura recovers in stages

Trafilatura searches for likely content through ordered structural rules. These recognize signals such as <article>, itemprop="articleBody", and article or main-content naming patterns.8 Inside those regions, discard rules and repeated link-density checks prune boilerplate. Dedicated handlers process paragraphs, lists, quotations, code, and tables. If extraction remains short, a recovery pass searches beyond the initially selected region while avoiding text already captured.2

The surrounding pipeline adds other ways to recover content:

  • Alternative extractors. With fast=False, Trafilatura can compare its result with its bundled Python readability_lxml fork and conditionally try jusText. This isn't a call to Mozilla's JavaScript package. The comparison uses length and structural tests, not only an empty-result check.9
  • Baseline rescue. Outside precision mode, an undersized result can trigger broader extraction from embedded body data, article elements, paragraph-like elements, or eventually the document's text.106
  • Recall escalation. Balanced mode can retry a nonempty extraction that covers unusually little of the page with more permissive recall settings, accepting the replacement only when it meets size and improvement checks.6

Those stages address different failure cases. A structural rule may find the intended article; recovery may reclaim a missed block; baseline rescue may return text when the ordinary extractor found too little. Broader selection also creates more opportunity to retain boilerplate. More recovered text isn't automatically more correct text.

What Trafilatura Core retains

Core's TypeScript engine ports Python Trafilatura's fast=True path, with go-trafilatura used as a DOM translation aid. It doesn't reproduce the default Python toolkit's external extractor comparisons.3 Its retained path includes the structural extractor and recovery pass, baseline rescue outside precision mode, and automatic recall escalation from balanced mode.5

ExtractorMain selectionRecovery beyond the initial selection
Mozilla ReadabilityCandidate scores, ancestor selection, sibling expansionRetries with filters relaxed
Python TrafilaturaStructural rules and content pruningInternal recovery; optional external comparisons; mode-dependent baseline and recall stages
Trafilatura CorePort of Trafilatura's fast extraction pathInternal recovery and mode-dependent baseline/recall stages; no external comparisons

These are differences in available mechanisms, not an accuracy ranking.165

Core's mode gates are specific. Precision skips baseline rescue and automatic recall escalation. Recall applies permissive extraction rules directly. Balanced can escalate when a nonempty result is under 3,000 characters and below 20% of the page's text; the recall candidate must reach the minimum extraction size and be more than 1.5 times as long as the existing result to replace it.5 A fallback isn't promised on every page or in every mode.

Settings and practical checks

Core exposes the extraction policy through precision, balanced, and recall; keep bypasses main-content extraction and cleans the whole document. Images, links, tables, and comments have separate controls.11 Python Trafilatura exposes favor_precision and favor_recall alongside its content controls.12 Readability also has configuration, including candidate count, character threshold, class preservation, and a link-density modifier.7 Core's named policies make the intended retention tradeoff explicit, but configurability is secondary to the extraction design.

For a workload mixing articles, documentation, and pages with scattered text, I'd test the recovery paths directly: check whether a missing section returns when selection broadens, and whether navigation or recommendations return with it. Save the HTML, mark the wanted passages, and compare retained content and unwanted content separately. The architecture tells you which mechanisms exist; representative pages tell you whether they help your workload.

Core accepts supplied HTML and returns cleaned HTML. Its optional source URL provides context; it doesn't fetch or render a page, and Markdown conversion belongs in a separate step.3 Its native Python library runs in Python and has no product CLI; parsing, serialization, and resource limits can differ from TypeScript.4 Readability accepts a DOM document and returns article HTML, plain text, and metadata; Node.js callers supply a DOM implementation.7 Keep this integration boundary in mind when testing: feeding one extractor browser-rendered HTML and the other an earlier server response would change the input as well as the algorithm.

Try the Core playground with a saved page, or use the npm library and Python library guides. If you render extracted HTML from untrusted pages, sanitize it for the output context and apply a Content Security Policy. Extraction is not a security boundary.117

Citations

  1. Mozilla: Readability 0.6.0 extraction implementation. Retrieved October 7, 2026 ↩ ↩2 ↩3 ↩4 ↩5

  2. Adrien Barbaresi: Trafilatura 2.2.0 main extractor and recovery. Retrieved October 7, 2026 ↩ ↩2

  3. Trafilatura Core: About Trafilatura Core and implementation lineage. Retrieved October 7, 2026 ↩ ↩2 ↩3

  4. Trafilatura Core: Native Python library and language differences and Python library guide. Retrieved October 7, 2026 ↩ ↩2

  5. Trafilatura Core: Core 0.8.6 extraction, baseline rescue, and recall escalation and extraction thresholds. Retrieved October 7, 2026 ↩ ↩2 ↩3 ↩4 ↩5

  6. Adrien Barbaresi: Trafilatura 2.2.0 extraction pipeline. Retrieved October 7, 2026 ↩ ↩2 ↩3 ↩4

  7. Mozilla: Readability 0.6.0 API, attribution, and security guidance. Retrieved October 7, 2026 ↩ ↩2 ↩3 ↩4

  8. Adrien Barbaresi: Trafilatura 2.2.0 structural selection rules. Retrieved October 7, 2026 ↩

  9. Adrien Barbaresi: Trafilatura 2.2.0 external extractor comparisons. Retrieved October 7, 2026 ↩

  10. Adrien Barbaresi: Trafilatura 2.2.0 baseline extraction. Retrieved October 7, 2026 ↩

  11. Trafilatura Core: npm library guide. Retrieved October 7, 2026 ↩ ↩2

  12. Adrien Barbaresi: Trafilatura 2.2.0 extraction options and option handling. Retrieved October 7, 2026 ↩

Updated: October 7, 2026