Cookbook¶
Practical recipes, none of them Django-specific. The Python recipes are covered by
tests/test_cookbook.py; browser-side JavaScript illustrates how to use the
rendered markup.
For Django-specific recipes, see the Django cookbook.
Building¶
A reusable component¶
A component is a function. There is no registry, no base class, and no special syntax. Composition is a function call.
def card(title, *body, href=None):
heading = A(title, href=href) if href else title
return Div(H2(heading), Div(*body, class_="card-body"), class_="card")
card("Hello", P("Body"), href="/x")
Taking *body and passing it through keeps the caller's syntax identical to
a built-in element's.
Reusable button variants¶
Use with_attrs() to derive variants from a configured element:
from django_div import Button
base = Button("Save", class_="btn", type="submit")
primary = base.with_attrs(class_="btn primary", disabled=True)
base still has class="btn" and no disabled attribute. Classes are
replaced as a whole, so include any base classes you want to keep. Pass
None to omit an attribute on render. The attribute dictionary and child
list are copied, while existing child objects and nested attribute values
remain shared.
A disclosure with matching ARIA state¶
Use the same Python flag to describe whether a panel is expanded and whether its content is hidden:
from django_div import Button, Div, Fragment
def disclosure(content, *, expanded=False):
return Fragment(
Button(
"Details", type="button", aria_controls="details",
aria_expanded=expanded,
),
Div(content, id="details", hidden=not expanded),
)
disclosure("More information")
<button type="button" aria-controls="details" aria-expanded="false">Details</button><div id="details" hidden>More information</div>
ARIA booleans render as explicit "true" or "false"; HTML booleans such
as hidden render bare when true and disappear when false. This recipe
sets the initial state. If JavaScript toggles the panel, update both
aria-expanded and hidden. Give each panel a unique ID when rendering
multiple disclosures.
A whole document¶
A doctype isn't an element, so it has its own item class. Doctype()
renders the HTML5 doctype, and a document is a list:
def document(title, *body, lang="en"):
return [
Doctype(),
Html(Head(Meta(charset="utf-8"), Title(title)), Body(*body), lang=lang),
]
def render(items):
return "".join(str(item) for item in items)
render(document("Home", H1("Hi")))
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8" /><title>Home</title></head><body><h1>Hi</h1></body></html>
HTML templates for JavaScript¶
Use Template to hold HTML that JavaScript can clone and insert later.
Build its contents with elements as usual:
from django_div import Div, Span, Template
template = Template(
Div(Span("Hello"), class_="card"),
id="card-template",
)
For an existing trusted HTML string, wrap the markup in Raw:
from django_div import Raw, Template
template = Template(
Raw(content='<div class="card"><span>Hello</span></div>'),
id="card-template",
)
Both produce the same HTML when rendered with str(template):
Once the template is in the document, JavaScript can insert its contents:
const template = document.querySelector("#card-template");
document.body.append(template.content.cloneNode(true));
Ordinary string children are still escaped inside Template. Use Raw
only for trusted markup, never untrusted user input. Placeholders such as
{{ name }} remain literal; django-div does not evaluate them.
A table from data¶
def data_table(rows, columns):
return Table(
Thead(Tr(Th(column) for column in columns)),
Tbody(Tr(Td(row[column]) for column in columns) for row in rows),
)
data_table([{"name": "Ana", "age": 33}], ["name", "age"])
<table><thead><tr><th>name</th><th>age</th></tr></thead><tbody><tr><td>Ana</td><td>33</td></tr></tbody></table>
Nested generators work because each one is flattened as it is consumed.
Navigation with an active item¶
A class mapping turns a condition into a class, and an all-false mapping
drops the attribute rather than emitting class="".
def nav(links, current):
return Nav(
Ul(Li(A(label, href=url, class_={"active": url == current}))
for label, url in links)
)
nav([("Home", "/"), ("Docs", "/docs/")], "/docs/")
A generator can't sit beside a keyword argument
This is a Python rule, not a django-div one:
Add brackets and it's fine:
Custom elements and web components¶
tag_class() registers the result globally, in TAG_CLASSES, so parsing
produces that class too. That is the point, but it means the registry grows at
runtime. BUILTIN_TAGS is the fixed set this library ships. For a one-off
that shouldn't be registered, Tag("my-widget", ...) skips it.
Passing data to JavaScript¶
Use JsonScript to pass data to JavaScript. It escapes characters that could
close the script element while preserving the original JSON values:
Read the data after the element is in the document:
JsonScript accepts dictionaries, lists, and nested Pydantic models.
Models use their field aliases and JSON-compatible values; None is
preserved as JSON null. No Django dependency is required.
Use this helper for data instead of interpolating values into executable
Script content. Plain HTML escaping is unsuitable inside a script element.
Open Graph metadata¶
Group page metadata in a reusable function. Fragment emits sibling meta
elements without adding a wrapper. Use page_type to avoid shadowing
Python's built-in type; the metadata property remains og:type.
from django_div import Fragment, Head, Meta, Title
def open_graph(*, title, url, image, description, page_type="website"):
return Fragment(
Meta(property="og:title", content=title),
Meta(property="og:type", content=page_type),
Meta(property="og:url", content=url),
Meta(property="og:image", content=image),
Meta(property="og:description", content=description),
)
head = Head(
Title("Introducing django-div"),
open_graph(
title="Introducing django-div",
url="https://example.com/posts/django-div/",
image="https://example.com/images/django-div.png",
description="Build and parse HTML in Python.",
page_type="article",
),
)
str(head) produces the following HTML (line breaks added for readability):
<head>
<title>Introducing django-div</title>
<meta property="og:title" content="Introducing django-div" />
<meta property="og:type" content="article" />
<meta property="og:url" content="https://example.com/posts/django-div/" />
<meta property="og:image" content="https://example.com/images/django-div.png" />
<meta property="og:description" content="Build and parse HTML in Python." />
</head>
Use absolute URLs for the page and image. Values are escaped as ordinary
attributes, so no Raw wrapper is needed. Omit page_type to use the default
"website". Extend the function with metadata your project needs, such as
og:site_name or og:image:alt. See the Open Graph protocol.
JSON-LD from a Pydantic model¶
Schema.org markup is JSON in a <script>, and a <script> is raw text: the
parser reads to the closing tag without decoding entities, so a name or
description containing </script> would close the tag early and the rest of
the payload would land on the page as live markup. JsonLd writes the
dangerous characters as the \uXXXX escapes JSON already understands, which
changes no data and makes that impossible.
Any Pydantic model works; nothing has to inherit from anything in
django-div. @context and @type are not Python names, so declare them as
aliases, which is also how you reach schema.org's camelCase properties like
sameAs:
from pydantic import BaseModel, ConfigDict, Field
from django_div import Head, JsonLd, Title
class Organization(BaseModel):
model_config = ConfigDict(populate_by_name=True)
context: str = Field("https://schema.org", alias="@context")
type: str = Field("Organization", alias="@type")
name: str
url: str
logo: str | None = None
same_as: list[str] | None = Field(None, alias="sameAs")
head = Head(
Title("Acme"),
JsonLd(Organization(name="Acme", url="https://acme.example")),
)
Models are dumped by alias with None dropped, so logo and same_as never
render. A JSON-LD null is not a value, so dropping it is what the format
means anyway. To dump on other terms, call model_dump() yourself and pass
the dict.
If you write a lot of these, the two constant fields are worth a base class of your own -- but that is your vocabulary to shape, not something django-div should decide for you.
Pass a list to emit several objects at once, and read one back out of a page
with json.loads(tag.text):
import json
from django_div import from_html
page = from_html(str(head))
data = json.loads(page.find("script", type="application/ld+json").text)
SVG icons¶
SVG elements aren't generated, because SVG's <text> would collide with the
Text model. Build the ones you need:
Svg = tag_class("svg")
Use = tag_class("use")
def icon(name, size=16):
return Svg(
Use(href=f"/static/icons.svg#{name}"),
width=size, height=size, aria_hidden="true",
)
Inline SVG shapes¶
SVG attribute names are case-sensitive, and a name without an underscore
passes through untouched, so viewBox renders as viewBox. Underscores
still become hyphens, which is what stroke_width needs:
Svg = tag_class("svg")
Circle = tag_class("circle")
Path = tag_class("path")
check = Svg(
Circle(cx="12", cy="12", r="10", fill="none"),
Path(d="M8 12l3 3 5-6"),
viewBox="0 0 24 24", width=24, height=24,
stroke="currentColor", stroke_width="2", aria_hidden="true",
)
<svg viewBox="0 0 24 24" width="24" height="24" stroke="currentColor" stroke-width="2" aria-hidden="true"><circle cx="12" cy="12" r="10" fill="none"></circle><path d="M8 12l3 3 5-6"></path></svg>
Shapes render as a full open and close pair, never self-closed, because only
HTML void elements self-close. A browser accepts <path></path> in inline
SVG.
Parsing SVG loses the attribute case
Building SVG keeps viewBox. Reading it back does not: an HTML parser
folds every attribute name to lowercase, so from_html returns
viewbox, which a browser ignores. The same applies to
preserveAspectRatio, gradientUnits, and the other camelCase names.
Treat SVG as write-only, or repair the names yourself after a parse.
XML, not just HTML¶
Tag doesn't care whether a name is HTML, so feeds and other XML work:
Note
Only HTML void elements self-close, and HTML escaping rules are applied. For heavy XML work a dedicated library is a better fit.
Parsing¶
These need the parse extra.
Extract every link¶
Find links by class token¶
Pass a predicate when a class can appear alongside other classes. Keyword attributes narrow the results using their existing exact-match behavior:
from django_div import from_html
page = from_html(
'<div><a class="external featured" target="_blank" href="/one">One</a>'
'<a class="external" href="/two">Two</a></div>'
)
links = page.find_all(
lambda node: node.tag == "a" and node.has_class("external"),
target="_blank",
)
[link.attrs["href"] for link in links]
# ['/one']
find() takes the same predicate and stops at the first match;
iter_find() yields matches lazily. All three search descendant tags in
document order, excluding the root. To compare a complete class attribute
instead, continue using find_all("a", class_="external").
Make relative URLs absolute¶
from urllib.parse import urljoin
def absolutize(tree, base):
for link in tree.find_all("a"):
if "href" in link.attrs:
link.attrs["href"] = urljoin(base, link.attrs["href"])
return tree
absolutize(from_html('<div><a href="/a">A</a></div>'), "https://example.test/")
Build a table of contents¶
def table_of_contents(tree, levels=("h2", "h3")):
return Ul(
Li(A(heading.text, href="#" + heading.attrs["id"]))
for heading in tree.find_all()
if heading.tag in levels and "id" in heading.attrs
)
Parsing and building in the same expression is the point: the input is HTML and so is the output.
Scrape a table into dicts¶
def cells(row):
return [cell.text.strip() for cell in row.find_all() if cell.tag in {"th", "td"}]
def table_to_dicts(table):
rows = table.find_all("tr")
headers = cells(rows[0])
return [dict(zip(headers, cells(row))) for row in rows[1:]]
Keep only certain elements¶
Transform a copy to remove unwanted nodes and unwrap other containers.
Children are processed before parents; returning a fragment preserves
already-filtered children without their original wrapper. The source remains
unchanged. Import Fragment, Tag, Text, and parse from django_div.
KEEP = {"p", "b", "i", "em", "strong", "a", "ul", "ol", "li", "code", "br"}
KEEP_ATTRS = {"a": {"href", "title"}}
DROP_ENTIRELY = {"script", "style"}
def keep_node(item):
if isinstance(item, (Text, Fragment)):
return item
if isinstance(item, Tag):
if item.tag in DROP_ENTIRELY:
return None
if item.tag not in KEEP:
return Fragment(item.children)
item.attrs = {
name: value for name, value in item.attrs.items()
if name in KEEP_ATTRS.get(item.tag, set())
}
return item
return None
def keep_only(items):
return Fragment(items).transform(keep_node)
items = parse('<p onclick="evil()">ok <script>alert(1)</script><b>b</b></p>')
str(keep_only(items))
# <p>ok <b>b</b></p>
This is not a sanitizer
It is a shape filter for content you already trust: trimming a CMS
export, normalizing pasted markup. It is not an XSS defense. Real
sanitization has to handle javascript: URLs, CSS escapes, mutation
XSS, and namespace confusion. For untrusted input use
nh3 or
bleach.
Note that DROP_ENTIRELY exists because unwrapping a <script> would
keep its code as visible text.
Readable text¶
.text concatenates, matching the DOM's textContent, so adjacent blocks
run together:
For search indexing or summaries, join trimmed text nodes:
The separator goes between every text node, including inline elements;
choose it for the output you need. .text keeps its original behavior.
Pretty-print a tree¶
Rendering has no indentation, by design. When you want it for debugging:
def pretty(item, indent=0):
pad = " " * indent
if not isinstance(item, Tag):
text = str(item).strip()
return [pad + text] if text else []
if item.is_void:
return [pad + str(item)]
opening = str(item).split(">", 1)[0] + ">"
lines = [pad + opening]
for child in item.children:
lines += pretty(child, indent + 1)
return [*lines, pad + f"</{item.tag}>"]
Markdown¶
The same tree renders as Markdown, and Markdown reads back in. Those recipes (changelog generation, link hardening, code-block extraction, document merging) have their own page: the Markdown cookbook.
Serializing¶
Cache a parsed page¶
Parsing is the expensive part. Serialize once, reload cheaply:
tree = from_html(response.text)
cache.set("page", tree.model_dump_json())
tree = Tag.model_validate_json(cache.get("page"))
Element classes survive the round trip, so find_all() and friends still
work on the way back.
Compare two pages structurally¶
Models compare by value, so equality ignores nothing that matters and nothing that doesn't:
from_html("<div><p>x</p></div>") == from_html("<div><p>x</p></div>") # True
from_html("<div><p>x</p></div>") == from_html("<div><p>y</p></div>") # False
Assert on structure, not strings¶
The most useful thing from_html() does in a test suite is let you stop
matching substrings:
def test_search_form():
page = from_html(response.text)
assert page.find("input", name="q") is not None
assert page.find("button").text == "Go"
That survives reformatting, attribute reordering, and added wrappers, all of
which break assert '<input name="q">' in html.