Skip to content

Parsing HTML

from_html() turns markup back into the same tree the constructors build, so a parsed page can be searched, edited, re-rendered, or serialized.

uv add 'django-div[parse]'
from django_div import from_html

page = from_html('<div class="card"><a href="/x">Click</a></div>')

type(page)                      # <class 'django_div.Div'>
page.find("a").attrs["href"]    # '/x'
page.text                       # 'Click'

parse() and from_html()

parse() always returns a list of top-level items. from_html() unwraps the single-root case, which is what you usually want.

from django_div import parse

parse("<p>a</p><p>b</p>")       # [P(...), P(...)]
from_html("<p>a</p><p>b</p>")   # [P(...), P(...)]
from_html("<p>a</p>")           # P(...)

Searching

page.text                          # all text in the subtree, unescaped
page.find("a")                     # first matching descendant, or None
page.find("a", class_="external")  # match on attributes too
page.find_all("a")                 # every matching descendant
page.iter_find("a")                # the same, lazily
page.walk()                        # every node, depth first, including text

Attribute names use the same Python spellings as the constructors, so class_="external" matches class="external".

for heading in page.find_all("h2"):
    print(heading.text)

Pass a predicate as the first argument for class tokens or other conditions:

page.find(lambda node: node.has_class("card"))
page.find_all(lambda node: node.tag in {"h2", "h3"})
page.iter_find(lambda node: node.has_class("external"), target="_blank")

All three search methods visit descendant tags in depth-first document order, excluding the root. They traverse fragments but never pass fragments or text nodes to predicates. Keyword attributes still match by exact equality and filter candidates before the predicate runs. find() stops at the first match. Predicates should inspect nodes without modifying the tree.

Editing

Parsed trees are ordinary models. Mutate attrs, append to children, then render.

page = from_html(response.text)

for link in page.find_all("a", target="_blank"):
    link.attrs["rel"] = "noopener"

print(page)

Transforming a copy

transform(visitor) copies the tree and visits every node, including text, comments, fragments, and the root. Children are visited in document order before their parent, so parents receive their already-transformed children.

from django_div import Tag


def remove_scripts(node):
    if isinstance(node, Tag) and node.tag == "script":
        return None
    return node


cleaned = page.transform(remove_scripts)

Return the node to keep it, another HtmlItem to replace it, or None to remove it with its subtree. Return Fragment(node.children) to unwrap an element, or a fragment of new elements to replace one node with siblings. Returned replacements are not traversed again. Descendants of a removed parent have already been visited. Removing the root returns None.

Callbacks receive copies, with independent child lists and attribute dictionaries and preserved subclasses. Nested attribute values remain shared, so replace a style/class mapping instead of editing it in place. A node appearing twice in the source is copied separately for each occurrence. A replacement supplied by the callback is used as-is; returning an external node shares that object. Exceptions propagate without changing the source unless the callback itself mutates shared or external objects.

Traversal is iterative and supports deep trees. The callback must return an HtmlItem or None; wrap text in Text and siblings in Fragment. This is structural editing, not HTML sanitization. See the tree filtering recipe.

Choosing a parser

Parsing goes through BeautifulSoup, which is a front end over several backends. django-div picks the best one installed (lxml, then html5lib, then the standard library), because they are not equivalent:

Input html.parser lxml html5lib
<p>one<p>two <p>one<p>two</p></p> <p>one</p><p>two</p> <p>one</p><p>two</p>
<ul><li>a<li>b</ul> nested <li> correct correct
46 KB document 28.8 ms 17.8 ms 54.2 ms

The standard library parser nests implicit closes instead of closing them, which quietly produces a wrong tree, and it isn't the fastest either. It remains the fallback because it needs no install.

Override per call when you need to:

from_html(markup, parser="html5lib")

best_parser() reports what would be chosen.

Fragments stay fragments

lxml and html5lib wrap a fragment in an invented <html><body> skeleton. That would turn a parse-edit-render round trip into a rewrite, so django-div strips wrappers the source did not ask for.

from_html("<p>hi</p>")    # P(...), not Html(Body(P(...)))

If the source really does contain <html>, <head>, or <body>, those are kept.

Whitespace

Whitespace between two elements is a word break in the rendered page, so it is preserved, but collapsed to a single space, since the source indentation itself carries no meaning.

from_html("<p><b>a</b> <b>b</b></p>")
# <p><b>a</b> <b>b</b></p>          the space matters

from_html("<div>\n    <p>one</p>\n    <p>two</p>\n</div>")
# <div><p>one</p> <p>two</p></div>  indentation collapsed

Inside pre and textarea, whitespace is significant and kept verbatim.

Comments and doctypes

Comments survive as Comment items and doctypes as Doctype items, so a whole document round-trips:

parse("<!DOCTYPE html><!--note--><p>hi</p>")
# [Doctype(...), Comment(...), P(...)]

Doctype.content is what follows <!DOCTYPE: "html" for modern documents, or the full legacy identifier string when parsing old markup.

Limits

Parsing is a tree conversion, not a browser. A few things to expect:

  • Character references are resolved by the parser, so &amp; comes back as & and is re-escaped on render. The output is equivalent, not identical.
  • Attribute order follows the source, but duplicate attributes are collapsed by the parser before django-div sees them.
  • Malformed markup is repaired according to whichever backend is in use, so round-tripping broken HTML gives you the repaired version.
  • Attribute names come back lowercased, because that is what an HTML parser does. This matters for inline SVG and MathML only, where names are case-sensitive:
from_html('<svg viewBox="0 0 24 24"></svg>').attrs
# {'viewbox': '0 0 24 24'}   a browser ignores this spelling

Building SVG keeps the case. Reading it back does not, so repair the names yourself if you must round-trip inline SVG.

Class tokens and readable text

Search attributes use exact equality. To find elements carrying a class among several tokens, pass a predicate to the lazy iterator:

links = page.iter_find(lambda node: node.tag == "a" and node.has_class("external"))
page.get_text(" ", strip=True)  # trim text nodes and join with spaces

link.classes gives the token list for string, iterable, or mapping class values. has_class() reflects edits to attrs["class"] immediately.

Working with multiple roots

Use a Fragment when you want multiple roots to act as one renderable tree:

from django_div import Fragment, parse

content = Fragment(parse("<h1>Title</h1><p>Body</p>"))
content.find("p").text          # 'Body'
content.get_text(" ", strip=True)  # 'Title Body'
str(content)                   # '<h1>Title</h1><p>Body</p>'

Wrapping does not change the parser's handling of whitespace or its return types. Fragment boundaries are Python tree structure and have no HTML marker, so parsing rendered HTML does not restore those boundaries; JSON does.