Parsing HTML¶
from_html() turns markup back into the same tree the constructors build, so
a parsed page can be searched, edited, re-rendered, or serialized.
from django_div import from_html
page = from_html('<div class="card"><a href="/x">Click</a></div>')
type(page) # <class 'django_div.Div'>
page.find("a").attrs["href"] # '/x'
page.text # 'Click'
parse() and from_html()¶
parse() always returns a list of top-level items. from_html() unwraps the
single-root case, which is what you usually want.
from django_div import parse
parse("<p>a</p><p>b</p>") # [P(...), P(...)]
from_html("<p>a</p><p>b</p>") # [P(...), P(...)]
from_html("<p>a</p>") # P(...)
Searching¶
page.text # all text in the subtree, unescaped
page.find("a") # first matching descendant, or None
page.find("a", class_="external") # match on attributes too
page.find_all("a") # every matching descendant
page.iter_find("a") # the same, lazily
page.walk() # every node, depth first, including text
Attribute names use the same Python spellings as the constructors, so
class_="external" matches class="external".
Pass a predicate as the first argument for class tokens or other conditions:
page.find(lambda node: node.has_class("card"))
page.find_all(lambda node: node.tag in {"h2", "h3"})
page.iter_find(lambda node: node.has_class("external"), target="_blank")
All three search methods visit descendant tags in depth-first document order,
excluding the root. They traverse fragments but never pass fragments or text
nodes to predicates. Keyword attributes still match by exact equality and
filter candidates before the predicate runs. find() stops at the first
match. Predicates should inspect nodes without modifying the tree.
Editing¶
Parsed trees are ordinary models. Mutate attrs, append to children, then
render.
page = from_html(response.text)
for link in page.find_all("a", target="_blank"):
link.attrs["rel"] = "noopener"
print(page)
Transforming a copy¶
transform(visitor) copies the tree and visits every node, including text,
comments, fragments, and the root. Children are visited in document order
before their parent, so parents receive their already-transformed children.
from django_div import Tag
def remove_scripts(node):
if isinstance(node, Tag) and node.tag == "script":
return None
return node
cleaned = page.transform(remove_scripts)
Return the node to keep it, another HtmlItem to replace it, or None to
remove it with its subtree. Return Fragment(node.children) to unwrap an
element, or a fragment of new elements to replace one node with siblings.
Returned replacements are not traversed again. Descendants of a removed
parent have already been visited. Removing the root returns None.
Callbacks receive copies, with independent child lists and attribute dictionaries and preserved subclasses. Nested attribute values remain shared, so replace a style/class mapping instead of editing it in place. A node appearing twice in the source is copied separately for each occurrence. A replacement supplied by the callback is used as-is; returning an external node shares that object. Exceptions propagate without changing the source unless the callback itself mutates shared or external objects.
Traversal is iterative and supports deep trees. The callback must return an
HtmlItem or None; wrap text in Text and siblings in Fragment.
This is structural editing, not HTML sanitization. See the
tree filtering recipe.
Choosing a parser¶
Parsing goes through BeautifulSoup, which is a front end over several
backends. django-div picks the best one installed (lxml, then html5lib,
then the standard library), because they are not equivalent:
| Input | html.parser |
lxml |
html5lib |
|---|---|---|---|
<p>one<p>two |
<p>one<p>two</p></p> |
<p>one</p><p>two</p> |
<p>one</p><p>two</p> |
<ul><li>a<li>b</ul> |
nested <li> |
correct | correct |
| 46 KB document | 28.8 ms | 17.8 ms | 54.2 ms |
The standard library parser nests implicit closes instead of closing them, which quietly produces a wrong tree, and it isn't the fastest either. It remains the fallback because it needs no install.
Override per call when you need to:
best_parser() reports what would be chosen.
Fragments stay fragments¶
lxml and html5lib wrap a fragment in an invented <html><body> skeleton.
That would turn a parse-edit-render round trip into a rewrite, so django-div
strips wrappers the source did not ask for.
If the source really does contain <html>, <head>, or <body>, those are
kept.
Whitespace¶
Whitespace between two elements is a word break in the rendered page, so it is preserved, but collapsed to a single space, since the source indentation itself carries no meaning.
from_html("<p><b>a</b> <b>b</b></p>")
# <p><b>a</b> <b>b</b></p> the space matters
from_html("<div>\n <p>one</p>\n <p>two</p>\n</div>")
# <div><p>one</p> <p>two</p></div> indentation collapsed
Inside pre and textarea, whitespace is significant and kept verbatim.
Comments and doctypes¶
Comments survive as Comment items and doctypes as Doctype items, so a
whole document round-trips:
Doctype.content is what follows <!DOCTYPE: "html" for modern
documents, or the full legacy identifier string when parsing old markup.
Limits¶
Parsing is a tree conversion, not a browser. A few things to expect:
- Character references are resolved by the parser, so
&comes back as&and is re-escaped on render. The output is equivalent, not identical. - Attribute order follows the source, but duplicate attributes are collapsed by the parser before django-div sees them.
- Malformed markup is repaired according to whichever backend is in use, so round-tripping broken HTML gives you the repaired version.
- Attribute names come back lowercased, because that is what an HTML parser does. This matters for inline SVG and MathML only, where names are case-sensitive:
from_html('<svg viewBox="0 0 24 24"></svg>').attrs
# {'viewbox': '0 0 24 24'} a browser ignores this spelling
Building SVG keeps the case. Reading it back does not, so repair the names yourself if you must round-trip inline SVG.
Class tokens and readable text¶
Search attributes use exact equality. To find elements carrying a class among several tokens, pass a predicate to the lazy iterator:
links = page.iter_find(lambda node: node.tag == "a" and node.has_class("external"))
page.get_text(" ", strip=True) # trim text nodes and join with spaces
link.classes gives the token list for string, iterable, or mapping class
values. has_class() reflects edits to attrs["class"] immediately.
Working with multiple roots¶
Use a Fragment when you want multiple roots to act as one renderable tree:
from django_div import Fragment, parse
content = Fragment(parse("<h1>Title</h1><p>Body</p>"))
content.find("p").text # 'Body'
content.get_text(" ", strip=True) # 'Title Body'
str(content) # '<h1>Title</h1><p>Body</p>'
Wrapping does not change the parser's handling of whitespace or its return types. Fragment boundaries are Python tree structure and have no HTML marker, so parsing rendered HTML does not restore those boundaries; JSON does.