# Extracting clean content and values

Sometimes what you want out of a page is the page itself, without navigation,
ads or footers, as the input of a search index, a summarizer or a
retrieval-augmented generation (RAG) pipeline. See Clean content.

Sometimes you know which parts of the page you want, you have already
[selected](selectors.md) them, and what you have is a string that is
not a value yet: a price with a currency symbol on it, a date written in
Spanish, a phone number written however its owner felt like. See
Parsing values.

## Clean content

[Trafilatura](https://trafilatura.readthedocs.io/en/latest/) turns HTML into text or markdown, dropping boilerplate. Call it
on [`response.text`](request-response.md) from a callback:

```python
import scrapy
import trafilatura


class ContentSpider(scrapy.Spider):
    name = "content"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["https://quotes.toscrape.com/"]

    def parse(self, response):
        yield {
            "url": response.url,
            "content": trafilatura.extract(response.text, output_format="markdown"),
        }
        yield from response.follow_all(css="a", callback=self.parse)
```

Save that as `content.py` and write the whole crawl to a JSON Lines
file:

```shell
scrapy runspider content.py -O content.jsonl
```

Title, author, date and site name come from a second call, which roughly
triples the time spent on each page, so ask for them only if you need them:

```python
metadata = trafilatura.extract_metadata(response.text)
yield {
    "url": response.url,
    "title": metadata.title,
    "date": metadata.date,
    "content": trafilatura.extract(response.text, output_format="markdown"),
}
```

### Where the content is not prose

Trafilatura is tuned for articles, and it keeps what it is confident about.
On a product page or a listing, where most of the page is not prose, it often
returns a fraction of what you were after, or nothing at all.

There the answer is to [select](selectors.md) the parts you want.
For a part that is itself rich content, such as a product description or the
body of a post, [clear-html](https://github.com/zytedata/clear-html) normalizes the selected node into clean HTML or
text, keeping embedded images and videos:

```pycon
>>> from clear_html import clean_node, cleaned_node_to_html
>>> from clear_html import cleaned_node_to_text
>>> node = clean_node(response.css("article")[0].root, response.url)
>>> cleaned_node_to_html(node)
'<article>\n\n<p>Hi <strong>there</strong></p>\n\n</article>'
>>> cleaned_node_to_text(node)
'Hi there'
```

### Pages you cannot download as they are

If the content is not in the HTML that Scrapy downloads, it is loaded
dynamically; see [Selecting dynamically-loaded content](dynamic-content.md).

If the website answers your requests with an error page or a challenge
instead of the content, see [Avoiding getting banned](practices.md).

[Zyte API](https://docs.zyte.com/zyte-api/get-started.html) covers both, and through [scrapy-zyte-api](https://github.com/scrapy-plugins/scrapy-zyte-api) it can also return
the content already structured: its article extraction gives you the text of
an article and a simplified HTML version of its body, which [markdownify](https://github.com/matthewwithanm/python-markdownify)
turns into markdown:

```pycon
>>> from markdownify import markdownify
>>> article = response.raw_api_response["article"]
>>> markdownify(article["articleBodyHtml"])
'Hi **there**'
```

## Parsing values

The following libraries turn a selected string into a value. They also work
as [input processors](loaders.md "Input and Output processors").

- [extruct](https://github.com/scrapinghub/extruct) reads JSON-LD, microdata, RDFa, Open Graph and Dublin Core out of
  a page, which is often the cheapest source of a title, an author or a
  date:
  ```pycon
  >>> import extruct
  >>> data = extruct.extract(response.text, base_url=response.url)
  >>> data["opengraph"][0]["properties"]
  [('og:title', 'Hi there')]
  ```
- [dateparser](https://github.com/scrapinghub/dateparser) reads a date written in prose, in any of the languages it
  supports:
  ```pycon
  >>> import dateparser
  >>> dateparser.parse("12 de octubre de 2025")
  datetime.datetime(2025, 10, 12, 0, 0)
  ```
- [price-parser](https://github.com/scrapinghub/price-parser) separates the amount of a price from its currency:
  ```pycon
  >>> from price_parser import Price
  >>> Price.fromstring("1.199,00 €")
  Price(amount=Decimal('1199.00'), currency='€')
  ```

  For numbers written out in words, use [number-parser](https://github.com/scrapinghub/number-parser) instead.
- [zyte-parsers](https://github.com/zytedata/zyte-parsers) parses fields out of a selected node: breadcrumbs, GTIN,
  rating, review count, brand and, building on [price-parser](https://github.com/scrapinghub/price-parser), price.
  ```pycon
  >>> from zyte_parsers import extract_breadcrumbs
  >>> extract_breadcrumbs(response.css("nav")[0], base_url=response.url)
  (Breadcrumb(name='Books', url='https://example.com/books'),)
  ```
- [phonenumbers](https://github.com/daviddrysdale/python-phonenumbers) turns a phone number as written into E.164, given the
  region it belongs to:
  ```pycon
  >>> import phonenumbers
  >>> number = phonenumbers.parse("0664 123 4567", "AT")
  >>> phonenumbers.format_number(number, phonenumbers.PhoneNumberFormat.E164)
  '+436641234567'
  ```
- [ftfy](https://github.com/rspeer/python-ftfy) repairs text that reached you as mojibake:
  ```pycon
  >>> import ftfy
  >>> ftfy.fix_text("The Mona Lisa doesnâ€™t have eyebrows.")
  "The Mona Lisa doesn't have eyebrows."
  ```
- [py3langid](https://github.com/adbar/py3langid) tells you which language a text is in:
  ```pycon
  >>> import py3langid
  >>> py3langid.classify("Dieser Text ist auf Deutsch geschrieben.")
  ('de', -154.48843383789062)
  ```

To parse a value out of JavaScript code in the page, see
[Parsing JavaScript code](dynamic-content.md).
