Skip to content

Cookbook

Practical recipes, none of them Django-specific. The Python recipes are covered by tests/test_cookbook.py; browser-side JavaScript illustrates how to use the rendered markup.

For Django-specific recipes, see the Django cookbook.

Building

A reusable component

A component is a function. There is no registry, no base class, and no special syntax. Composition is a function call.

def card(title, *body, href=None):
    heading = A(title, href=href) if href else title
    return Div(H2(heading), Div(*body, class_="card-body"), class_="card")

card("Hello", P("Body"), href="/x")
<div class="card"><h2><a href="/x">Hello</a></h2><div class="card-body"><p>Body</p></div></div>

Taking *body and passing it through keeps the caller's syntax identical to a built-in element's.

Reusable button variants

Use with_attrs() to derive variants from a configured element:

from django_div import Button

base = Button("Save", class_="btn", type="submit")
primary = base.with_attrs(class_="btn primary", disabled=True)
<button class="btn primary" type="submit" disabled>Save</button>

base still has class="btn" and no disabled attribute. Classes are replaced as a whole, so include any base classes you want to keep. Pass None to omit an attribute on render. The attribute dictionary and child list are copied, while existing child objects and nested attribute values remain shared.

A disclosure with matching ARIA state

Use the same Python flag to describe whether a panel is expanded and whether its content is hidden:

from django_div import Button, Div, Fragment


def disclosure(content, *, expanded=False):
    return Fragment(
        Button(
            "Details", type="button", aria_controls="details",
            aria_expanded=expanded,
        ),
        Div(content, id="details", hidden=not expanded),
    )


disclosure("More information")
<button type="button" aria-controls="details" aria-expanded="false">Details</button><div id="details" hidden>More information</div>

ARIA booleans render as explicit "true" or "false"; HTML booleans such as hidden render bare when true and disappear when false. This recipe sets the initial state. If JavaScript toggles the panel, update both aria-expanded and hidden. Give each panel a unique ID when rendering multiple disclosures.

A whole document

A doctype isn't an element, so it has its own item class. Doctype() renders the HTML5 doctype, and a document is a list:

def document(title, *body, lang="en"):
    return [
        Doctype(),
        Html(Head(Meta(charset="utf-8"), Title(title)), Body(*body), lang=lang),
    ]

def render(items):
    return "".join(str(item) for item in items)

render(document("Home", H1("Hi")))
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8" /><title>Home</title></head><body><h1>Hi</h1></body></html>

HTML templates for JavaScript

Use Template to hold HTML that JavaScript can clone and insert later. Build its contents with elements as usual:

from django_div import Div, Span, Template

template = Template(
    Div(Span("Hello"), class_="card"),
    id="card-template",
)

For an existing trusted HTML string, wrap the markup in Raw:

from django_div import Raw, Template

template = Template(
    Raw(content='<div class="card"><span>Hello</span></div>'),
    id="card-template",
)

Both produce the same HTML when rendered with str(template):

<template id="card-template"><div class="card"><span>Hello</span></div></template>

Once the template is in the document, JavaScript can insert its contents:

const template = document.querySelector("#card-template");
document.body.append(template.content.cloneNode(true));

Ordinary string children are still escaped inside Template. Use Raw only for trusted markup, never untrusted user input. Placeholders such as {{ name }} remain literal; django-div does not evaluate them.

A table from data

def data_table(rows, columns):
    return Table(
        Thead(Tr(Th(column) for column in columns)),
        Tbody(Tr(Td(row[column]) for column in columns) for row in rows),
    )

data_table([{"name": "Ana", "age": 33}], ["name", "age"])
<table><thead><tr><th>name</th><th>age</th></tr></thead><tbody><tr><td>Ana</td><td>33</td></tr></tbody></table>

Nested generators work because each one is flattened as it is consumed.

A class mapping turns a condition into a class, and an all-false mapping drops the attribute rather than emitting class="".

def nav(links, current):
    return Nav(
        Ul(Li(A(label, href=url, class_={"active": url == current}))
           for label, url in links)
    )

nav([("Home", "/"), ("Docs", "/docs/")], "/docs/")
<nav><ul><li><a href="/">Home</a></li><li><a href="/docs/" class="active">Docs</a></li></ul></nav>

A generator can't sit beside a keyword argument

This is a Python rule, not a django-div one:

Ul(Li(x) for x in items, class_="errors")   # SyntaxError

Add brackets and it's fine:

Ul([Li(x) for x in items], class_="errors")

Custom elements and web components

MyWidget = tag_class("my-widget")
MyWidget("hi", data_state="ready")
<my-widget data-state="ready">hi</my-widget>

tag_class() registers the result globally, in TAG_CLASSES, so parsing produces that class too. That is the point, but it means the registry grows at runtime. BUILTIN_TAGS is the fixed set this library ships. For a one-off that shouldn't be registered, Tag("my-widget", ...) skips it.

Passing data to JavaScript

Use JsonScript to pass data to JavaScript. It escapes characters that could close the script element while preserving the original JSON values:

from django_div import JsonScript

JsonScript({"a": "</script>"}, id="config")
<script type="application/json" id="config">{"a":"\u003c/script\u003e"}</script>

Read the data after the element is in the document:

const config = JSON.parse(document.getElementById("config").textContent);

JsonScript accepts dictionaries, lists, and nested Pydantic models. Models use their field aliases and JSON-compatible values; None is preserved as JSON null. No Django dependency is required.

Use this helper for data instead of interpolating values into executable Script content. Plain HTML escaping is unsuitable inside a script element.

Open Graph metadata

Group page metadata in a reusable function. Fragment emits sibling meta elements without adding a wrapper. Use page_type to avoid shadowing Python's built-in type; the metadata property remains og:type.

from django_div import Fragment, Head, Meta, Title


def open_graph(*, title, url, image, description, page_type="website"):
    return Fragment(
        Meta(property="og:title", content=title),
        Meta(property="og:type", content=page_type),
        Meta(property="og:url", content=url),
        Meta(property="og:image", content=image),
        Meta(property="og:description", content=description),
    )


head = Head(
    Title("Introducing django-div"),
    open_graph(
        title="Introducing django-div",
        url="https://example.com/posts/django-div/",
        image="https://example.com/images/django-div.png",
        description="Build and parse HTML in Python.",
        page_type="article",
    ),
)

str(head) produces the following HTML (line breaks added for readability):

<head>
  <title>Introducing django-div</title>
  <meta property="og:title" content="Introducing django-div" />
  <meta property="og:type" content="article" />
  <meta property="og:url" content="https://example.com/posts/django-div/" />
  <meta property="og:image" content="https://example.com/images/django-div.png" />
  <meta property="og:description" content="Build and parse HTML in Python." />
</head>

Use absolute URLs for the page and image. Values are escaped as ordinary attributes, so no Raw wrapper is needed. Omit page_type to use the default "website". Extend the function with metadata your project needs, such as og:site_name or og:image:alt. See the Open Graph protocol.

JSON-LD from a Pydantic model

Schema.org markup is JSON in a <script>, and a <script> is raw text: the parser reads to the closing tag without decoding entities, so a name or description containing </script> would close the tag early and the rest of the payload would land on the page as live markup. JsonLd writes the dangerous characters as the \uXXXX escapes JSON already understands, which changes no data and makes that impossible.

Any Pydantic model works; nothing has to inherit from anything in django-div. @context and @type are not Python names, so declare them as aliases, which is also how you reach schema.org's camelCase properties like sameAs:

from pydantic import BaseModel, ConfigDict, Field

from django_div import Head, JsonLd, Title


class Organization(BaseModel):
    model_config = ConfigDict(populate_by_name=True)

    context: str = Field("https://schema.org", alias="@context")
    type: str = Field("Organization", alias="@type")
    name: str
    url: str
    logo: str | None = None
    same_as: list[str] | None = Field(None, alias="sameAs")


head = Head(
    Title("Acme"),
    JsonLd(Organization(name="Acme", url="https://acme.example")),
)

Models are dumped by alias with None dropped, so logo and same_as never render. A JSON-LD null is not a value, so dropping it is what the format means anyway. To dump on other terms, call model_dump() yourself and pass the dict.

If you write a lot of these, the two constant fields are worth a base class of your own -- but that is your vocabulary to shape, not something django-div should decide for you.

Pass a list to emit several objects at once, and read one back out of a page with json.loads(tag.text):

import json

from django_div import from_html

page = from_html(str(head))
data = json.loads(page.find("script", type="application/ld+json").text)

SVG icons

SVG elements aren't generated, because SVG's <text> would collide with the Text model. Build the ones you need:

Svg = tag_class("svg")
Use = tag_class("use")

def icon(name, size=16):
    return Svg(
        Use(href=f"/static/icons.svg#{name}"),
        width=size, height=size, aria_hidden="true",
    )
<svg width="16" height="16" aria-hidden="true"><use href="/static/icons.svg#check"></use></svg>

Inline SVG shapes

SVG attribute names are case-sensitive, and a name without an underscore passes through untouched, so viewBox renders as viewBox. Underscores still become hyphens, which is what stroke_width needs:

Svg = tag_class("svg")
Circle = tag_class("circle")
Path = tag_class("path")

check = Svg(
    Circle(cx="12", cy="12", r="10", fill="none"),
    Path(d="M8 12l3 3 5-6"),
    viewBox="0 0 24 24", width=24, height=24,
    stroke="currentColor", stroke_width="2", aria_hidden="true",
)
<svg viewBox="0 0 24 24" width="24" height="24" stroke="currentColor" stroke-width="2" aria-hidden="true"><circle cx="12" cy="12" r="10" fill="none"></circle><path d="M8 12l3 3 5-6"></path></svg>

Shapes render as a full open and close pair, never self-closed, because only HTML void elements self-close. A browser accepts <path></path> in inline SVG.

Parsing SVG loses the attribute case

Building SVG keeps viewBox. Reading it back does not: an HTML parser folds every attribute name to lowercase, so from_html returns viewbox, which a browser ignores. The same applies to preserveAspectRatio, gradientUnits, and the other camelCase names. Treat SVG as write-only, or repair the names yourself after a parse.

XML, not just HTML

Tag doesn't care whether a name is HTML, so feeds and other XML work:

Tag("rss",
    Tag("channel", Tag("title", "News"), Tag("item", Tag("title", "First"))),
    version="2.0")
<rss version="2.0"><channel><title>News</title><item><title>First</title></item></channel></rss>

Note

Only HTML void elements self-close, and HTML escaping rules are applied. For heavy XML work a dedicated library is a better fit.

Parsing

These need the parse extra.

page = from_html(markup)
[(a.text, a.attrs["href"]) for a in page.find_all("a")]
[("A", "/a"), ("B", "/b")]

Pass a predicate when a class can appear alongside other classes. Keyword attributes narrow the results using their existing exact-match behavior:

from django_div import from_html

page = from_html(
    '<div><a class="external featured" target="_blank" href="/one">One</a>'
    '<a class="external" href="/two">Two</a></div>'
)
links = page.find_all(
    lambda node: node.tag == "a" and node.has_class("external"),
    target="_blank",
)
[link.attrs["href"] for link in links]
# ['/one']

find() takes the same predicate and stops at the first match; iter_find() yields matches lazily. All three search descendant tags in document order, excluding the root. To compare a complete class attribute instead, continue using find_all("a", class_="external").

Make relative URLs absolute

from urllib.parse import urljoin

def absolutize(tree, base):
    for link in tree.find_all("a"):
        if "href" in link.attrs:
            link.attrs["href"] = urljoin(base, link.attrs["href"])
    return tree

absolutize(from_html('<div><a href="/a">A</a></div>'), "https://example.test/")
<div><a href="https://example.test/a">A</a></div>

Build a table of contents

def table_of_contents(tree, levels=("h2", "h3")):
    return Ul(
        Li(A(heading.text, href="#" + heading.attrs["id"]))
        for heading in tree.find_all()
        if heading.tag in levels and "id" in heading.attrs
    )
<ul><li><a href="#a">A</a></li><li><a href="#b">B</a></li></ul>

Parsing and building in the same expression is the point: the input is HTML and so is the output.

Scrape a table into dicts

def cells(row):
    return [cell.text.strip() for cell in row.find_all() if cell.tag in {"th", "td"}]

def table_to_dicts(table):
    rows = table.find_all("tr")
    headers = cells(rows[0])
    return [dict(zip(headers, cells(row))) for row in rows[1:]]
[{"name": "Ana", "age": "33"}]

Keep only certain elements

Transform a copy to remove unwanted nodes and unwrap other containers. Children are processed before parents; returning a fragment preserves already-filtered children without their original wrapper. The source remains unchanged. Import Fragment, Tag, Text, and parse from django_div.

KEEP = {"p", "b", "i", "em", "strong", "a", "ul", "ol", "li", "code", "br"}
KEEP_ATTRS = {"a": {"href", "title"}}
DROP_ENTIRELY = {"script", "style"}

def keep_node(item):
    if isinstance(item, (Text, Fragment)):
        return item
    if isinstance(item, Tag):
        if item.tag in DROP_ENTIRELY:
            return None
        if item.tag not in KEEP:
            return Fragment(item.children)
        item.attrs = {
            name: value for name, value in item.attrs.items()
            if name in KEEP_ATTRS.get(item.tag, set())
        }
        return item
    return None


def keep_only(items):
    return Fragment(items).transform(keep_node)
items = parse('<p onclick="evil()">ok <script>alert(1)</script><b>b</b></p>')
str(keep_only(items))
# <p>ok <b>b</b></p>

This is not a sanitizer

It is a shape filter for content you already trust: trimming a CMS export, normalizing pasted markup. It is not an XSS defense. Real sanitization has to handle javascript: URLs, CSS escapes, mutation XSS, and namespace confusion. For untrusted input use nh3 or bleach.

Note that DROP_ENTIRELY exists because unwrapping a <script> would keep its code as visible text.

Readable text

.text concatenates, matching the DOM's textContent, so adjacent blocks run together:

tree = from_html("<article><h1>Title</h1><p>one two</p></article>")
tree.text          # 'Titleone two'

For search indexing or summaries, join trimmed text nodes:

tree.get_text(" ", strip=True)   # 'Title one two'

The separator goes between every text node, including inline elements; choose it for the output you need. .text keeps its original behavior.

Pretty-print a tree

Rendering has no indentation, by design. When you want it for debugging:

def pretty(item, indent=0):
    pad = "  " * indent
    if not isinstance(item, Tag):
        text = str(item).strip()
        return [pad + text] if text else []
    if item.is_void:
        return [pad + str(item)]
    opening = str(item).split(">", 1)[0] + ">"
    lines = [pad + opening]
    for child in item.children:
        lines += pretty(child, indent + 1)
    return [*lines, pad + f"</{item.tag}>"]
<div>
  <p>
    hi
  </p>
</div>

Markdown

The same tree renders as Markdown, and Markdown reads back in. Those recipes (changelog generation, link hardening, code-block extraction, document merging) have their own page: the Markdown cookbook.

Serializing

Cache a parsed page

Parsing is the expensive part. Serialize once, reload cheaply:

tree = from_html(response.text)
cache.set("page", tree.model_dump_json())

tree = Tag.model_validate_json(cache.get("page"))

Element classes survive the round trip, so find_all() and friends still work on the way back.

Compare two pages structurally

Models compare by value, so equality ignores nothing that matters and nothing that doesn't:

from_html("<div><p>x</p></div>") == from_html("<div><p>x</p></div>")   # True
from_html("<div><p>x</p></div>") == from_html("<div><p>y</p></div>")   # False

Assert on structure, not strings

The most useful thing from_html() does in a test suite is let you stop matching substrings:

def test_search_form():
    page = from_html(response.text)
    assert page.find("input", name="q") is not None
    assert page.find("button").text == "Go"

That survives reformatting, attribute reordering, and added wrappers, all of which break assert '<input name="q">' in html.