# django-div
> Build and parse HTML in Python with Pydantic models.
---
# django-div
Build and parse HTML in Python with Pydantic models.
```python
from django_div import A, Div, P
print(Div(P("Hello, World!"), A("Click", href="/x"), class_="card"))
```
```html
```
Children are positional, attributes are keyword arguments. Text is escaped,
void elements self-close, and Python attribute spellings map onto HTML ones.
## Why this exists
Most HTML-in-Python libraries build markup and stop there. Here the tree is a
Pydantic model, which means the same objects can go in four directions:
- **Build** Compose elements as Python values, with escaping handled for you. [ Building HTML](building/)
- **Parse** Read existing markup back into the same tree, then search and edit it. [ Parsing HTML](parsing/)
- **Serialize** Round-trip a page through JSON with its element classes intact. [ Serializing](serializing/)
- **Render in Django** Components as templates, with escaping that works both ways. [ Django](django/)
## Install
With uv, create a virtual environment if needed and install the package:
```console
uv venv
uv pip install django-div
```
With pip, install into your active virtual environment:
```console
python -m pip install django-div
```
For uv-managed projects, `uv add` records the dependency in `pyproject.toml`.
The same extras below also work with `uv pip install` and `python -m pip install`:
```console
uv add django-div # building only
uv add 'django-div[parse]' # plus from_html()/parse(), via bs4 + lxml
uv add 'django-div[html5]' # spec-exact parsing, ~3x slower than lxml
uv add 'django-div[markdown]' # plus from_markdown(), via markdown-it-py
```
django-div needs Python 3.12 or later.
Pydantic is the only hard dependency. Parsers are optional and their imports
are guarded, so building HTML pulls in nothing else. Django is optional too:
only `django_div.django` imports it.
## A tour in one page
```python
from django_div import Div, H1, Li, P, Span, Ul, from_html
# Nesting, attributes, escaping
Div(H1("Title"), P("a < b"), class_="page")
#
# None and False children drop out, so inline conditionals work
Div("Hello", user and Span(user.name))
# Collections flatten, so comprehensions splat in
Ul(Li(item) for item in items)
# Parse markup into the same kind of tree
page = from_html('')
page.find("a").attrs["href"] # '/x'
page.text # 'Click'
# Edit it and render it back out
for link in page.find_all("a", target="_blank"):
link.attrs["rel"] = "noopener"
print(page)
```
## Where to next
The four cards above cover building, parsing, serializing, and Django. Beyond
those:
- [Markdown](markdown/): render the tree as Markdown, read Markdown in
- Cookbooks: [HTML](cookbook/) · [Django](django-cookbook/) ·
[Markdown](markdown-cookbook/)
- [API reference](reference/): every function, class, and constant
- [Contributing](contributing/): setup, conventions, tests
## Reusable elements and browser data
Derive element variants with `with_attrs()`, render explicit ARIA states
from Python booleans, and pass browser data with `JsonScript`.
The [cookbook](cookbook/) includes button variants, disclosures,
HTML templates, and JSON data recipes.
Search by class tokens or custom conditions with predicates, and use
`transform()` to replace, remove, or unwrap nodes on a copied tree.
See [searching](parsing/#searching) and
[transforming a copy](parsing/#transforming-a-copy) for traversal and
copy semantics.
## llms.txt
This documentation is available in the [llms.txt](https://llmstxt.org/)
format, a Markdown convention suited to LLMs and AI coding assistants.
Two files are published:
- [`llms.txt`](https://django-div.readthedocs.io/en/latest/llms.txt): a short
description of the project plus links to each section. The structure is
described [here](https://llmstxt.org/#format).
- [`llms-full.txt`](https://django-div.readthedocs.io/en/latest/llms-full.txt):
the same index with the content of every page inlined.
Every page is also published as Markdown alongside its HTML, so you can point
an assistant at a single section rather than the whole corpus. Append `.md` to
the page name:
```text
https://django-div.readthedocs.io/en/latest/building.md
https://django-div.readthedocs.io/en/latest/django.md
```
## Where it came from
django-div started as five throwaway scripts trying to answer one question:
how do you make `Div("hello", class_="x")` work when Pydantic wants keyword
fields? Those experiments are still in the git history at the
`Baseline: original django-div demo experiments` commit.
## Prior art
[htpy](https://htpy.dev), [dominate](https://github.com/Knio/dominate), and
[django-components](https://github.com/django-components/django-components)
cover adjacent ground, and htpy in particular landed on a very similar
constructor shape. django-div's angle is the Pydantic model underneath: the
same objects parse, validate, and serialize.
---
# Building HTML
Every HTML element has a class, named after the tag with a capital letter:
`Div`, `P`, `H1`, `Textarea`. The set covers all 138 elements MDN tracks --
the 113 current ones in the
[WHATWG living standard](https://html.spec.whatwg.org/multipage/indices.html#elements-3)
plus 19 deprecated and 6 experimental, which warn when you build one --
each documented on [MDN](https://developer.mozilla.org/en-US/docs/Web/HTML/Element).
Children are positional arguments, attributes are keyword arguments.
```python
from django_div import A, Div, H1, P
Div(
H1("Welcome"),
P("A paragraph with ", A("a link", href="/docs"), "."),
class_="page",
)
```
```html
Welcome A paragraph with a link .
```
Rendering happens on `str()`, so `print(tag)`, f-strings, and
`"".join(...)` all work directly.
## Attributes
Python spellings are translated to HTML ones: a trailing underscore is
dropped, and remaining underscores become hyphens.
| Python | HTML |
| --- | --- |
| `class_="card"` | `class="card"` |
| `for_="email"` | `for="email"` |
| `data_test_id="hero"` | `data-test-id="hero"` |
| `aria_label="Close"` | `aria-label="Close"` |
| `hx_get="/rows"` | `hx-get="/rows"` |
| `http_equiv="refresh"` | `http-equiv="refresh"` |
### Reserved words
The trailing underscore also covers every Python keyword. Add one underscore
to the end of the keyword, and the rule removes it again on render.
| Python | HTML | Where it applies |
| --- | --- | --- |
| `class_` | `class` | every element |
| `for_` | `for` | ``, `` |
| `as_` | `as` | ` ` |
| `async_` | `async` | `")
# <script>alert(1)</script>
Div(title='he said "hi"')
#
```
To emit markup you have already vetted, wrap it in `Raw`:
```python
from django_div import Raw
Div(Raw(content="bold ")) # bold
```
Raw is a loaded gun
`Raw` bypasses escaping completely. Never build one from user input.
Anything implementing the `__html__` protocol (a Django `SafeString`, a
`markupsafe.Markup`, a rendered form) is trusted and passes through
unescaped, so interop with other libraries needs no special handling.
### Comments
`Comment` renders an HTML comment. A comment is not an escaping context, so
its content is not HTML-escaped. Instead the three comment-syntax rules are
neutralized on render, because each one would otherwise let content escape
the comment:
```python
Comment(content="note") #
Comment(content="a--b") #
Comment(content=">boom") #
Comment(content="x-") #
```
- `--` inside a comment would end it early, so it becomes `- -`.
- HTML5 reads `` and `` as *complete* comments, so a leading `>`
or `->` would leave the rest of the content parsing as live markup. For
untrusted content, that is script execution. A space is padded in front.
- A trailing `-` would produce the invalid ``; a space is padded.
The padding is the whole defense: unlike a `
Style("a > b { color: red }") #
```
Because nothing escapes a raw text element except its own closing tag,
content containing that closing tag would break out of the element, and for
a `"))
# ValueError: ` block instead.
### Preformatted elements
`pre` and `textarea` have significant whitespace. This only affects parsing (see
[Parsing HTML](../parsing/)), since building preserves exactly what you
pass.
## Elements not in the list
`Tag` takes the name as its first argument, and handles anything, including
custom elements and web components:
```python
from django_div import Tag
Tag("my-widget", "hi", data_state="ready")
# hi
```
To get a reusable class instead, use `tag_class()`. It registers the result,
so parsing will produce that class too:
```python
from django_div import tag_class
MyWidget = tag_class("my-widget")
MyWidget("hi", data_state="ready")
```
---
# Parsing HTML
`from_html()` turns markup back into the same tree the constructors build, so
a parsed page can be searched, edited, re-rendered, or serialized.
```console
uv add 'django-div[parse]'
```
```python
from django_div import from_html
page = from_html('')
type(page) #
page.find("a").attrs["href"] # '/x'
page.text # 'Click'
```
## `parse()` and `from_html()`
`parse()` always returns a list of top-level items. `from_html()` unwraps the
single-root case, which is what you usually want.
```python
from django_div import parse
parse("a
b
") # [P(...), P(...)]
from_html("a
b
") # [P(...), P(...)]
from_html("a
") # P(...)
```
## Searching
```python
page.text # all text in the subtree, unescaped
page.find("a") # first matching descendant, or None
page.find("a", class_="external") # match on attributes too
page.find_all("a") # every matching descendant
page.iter_find("a") # the same, lazily
page.walk() # every node, depth first, including text
```
Attribute names use the same Python spellings as the constructors, so
`class_="external"` matches `class="external"`.
```python
for heading in page.find_all("h2"):
print(heading.text)
```
Pass a predicate as the first argument for class tokens or other conditions:
```python
page.find(lambda node: node.has_class("card"))
page.find_all(lambda node: node.tag in {"h2", "h3"})
page.iter_find(lambda node: node.has_class("external"), target="_blank")
```
All three search methods visit descendant tags in depth-first document order,
excluding the root. They traverse fragments but never pass fragments or text
nodes to predicates. Keyword attributes still match by exact equality and
filter candidates before the predicate runs. `find()` stops at the first
match. Predicates should inspect nodes without modifying the tree.
## Editing
Parsed trees are ordinary models. Mutate `attrs`, append to `children`, then
render.
```python
page = from_html(response.text)
for link in page.find_all("a", target="_blank"):
link.attrs["rel"] = "noopener"
print(page)
```
### Transforming a copy
`transform(visitor)` copies the tree and visits every node, including text,
comments, fragments, and the root. Children are visited in document order
before their parent, so parents receive their already-transformed children.
```python
from django_div import Tag
def remove_scripts(node):
if isinstance(node, Tag) and node.tag == "script":
return None
return node
cleaned = page.transform(remove_scripts)
```
Return the node to keep it, another `HtmlItem` to replace it, or `None` to
remove it with its subtree. Return `Fragment(node.children)` to unwrap an
element, or a fragment of new elements to replace one node with siblings.
Returned replacements are not traversed again. Descendants of a removed
parent have already been visited. Removing the root returns `None`.
Callbacks receive copies, with independent child lists and attribute
dictionaries and preserved subclasses. Nested attribute values remain
shared, so replace a style/class mapping instead of editing it in place.
A node appearing twice in the source is copied separately for each occurrence.
A replacement supplied by the callback is used as-is; returning an external
node shares that object. Exceptions propagate without changing the source
unless the callback itself mutates shared or external objects.
Traversal is iterative and supports deep trees. The callback must return an
`HtmlItem` or `None`; wrap text in `Text` and siblings in `Fragment`.
This is structural editing, not HTML sanitization. See the
[tree filtering recipe](../cookbook/#keep-only-certain-elements).
## Choosing a parser
Parsing goes through BeautifulSoup, which is a front end over several
backends. django-div picks the best one installed (`lxml`, then `html5lib`,
then the standard library), because they are not equivalent:
| Input | `html.parser` | `lxml` | `html5lib` |
| --- | --- | --- | --- |
| `one
two` | `
one
two
` | `one
two
` | `one
two
` |
| `` | nested `` | correct | correct |
| 46 KB document | 28.8 ms | **17.8 ms** | 54.2 ms |
The standard library parser nests implicit closes instead of closing them,
which quietly produces a wrong tree, and it isn't the fastest either. It
remains the fallback because it needs no install.
Override per call when you need to:
```python
from_html(markup, parser="html5lib")
```
`best_parser()` reports what would be chosen.
## Fragments stay fragments
lxml and html5lib wrap a fragment in an invented `` skeleton.
That would turn a parse-edit-render round trip into a rewrite, so django-div
strips wrappers the source did not ask for.
```python
from_html("hi
") # P(...), not Html(Body(P(...)))
```
If the source really does contain ``, ``, or ``, those are
kept.
## Whitespace
Whitespace between two elements is a word break in the rendered page, so it
is preserved, but collapsed to a single space, since the source indentation
itself carries no meaning.
```python
from_html("a b
")
# a b
the space matters
from_html("")
# indentation collapsed
```
Inside `pre` and `textarea`, whitespace is significant and kept verbatim.
## Comments and doctypes
Comments survive as `Comment` items and doctypes as `Doctype` items, so a
whole document round-trips:
```python
parse("hi
")
# [Doctype(...), Comment(...), P(...)]
```
`Doctype.content` is what follows `').attrs
# {'viewbox': '0 0 24 24'} a browser ignores this spelling
```
Building SVG keeps the case. Reading it back does not, so repair the names
yourself if you must round-trip inline SVG.
## Class tokens and readable text
Search attributes use exact equality. To find elements carrying a class
among several tokens, pass a predicate to the lazy iterator:
```python
links = page.iter_find(lambda node: node.tag == "a" and node.has_class("external"))
page.get_text(" ", strip=True) # trim text nodes and join with spaces
```
`link.classes` gives the token list for string, iterable, or mapping class
values. `has_class()` reflects edits to `attrs["class"]` immediately.
## Working with multiple roots
Use a `Fragment` when you want multiple roots to act as one renderable tree:
```python
from django_div import Fragment, parse
content = Fragment(parse("Title Body
"))
content.find("p").text # 'Body'
content.get_text(" ", strip=True) # 'Title Body'
str(content) # 'Title Body
'
```
Wrapping does not change the parser's handling of whitespace or its return
types. Fragment boundaries are Python tree structure and have no HTML marker,
so parsing rendered HTML does not restore those boundaries; JSON does.
---
# Serializing
Every item is a Pydantic model, so a tree dumps to a dict or JSON and loads
back with its element classes intact.
```python
from django_div import Div, P, Tag
tree = Div(P("hi"), class_="card")
payload = tree.model_dump_json()
restored = Tag.model_validate_json(payload)
str(restored) == str(tree) # True
type(restored.children[0]) #
```
## The shape
```python
Div(P("x"), class_="card").model_dump()
```
```python
{
"type": "tag",
"tag": "div",
"attrs": {"class": "card"},
"children": [
{
"type": "tag",
"tag": "p",
"attrs": {},
"children": [{"type": "text", "content": "x"}],
},
],
}
```
Two fields carry the type information:
`type`
: Distinguishes the leaf kinds, which otherwise look identical: `Text`,
`Raw`, and `Comment` all hold a single `content` string.
`tag`
: Names the element, so `Div` comes back as `Div` rather than a generic
`Tag`.
Attribute keys are stored in their **HTML** spelling, normalized when the
element is built. A hand-built tree and a parsed one are therefore
structurally identical.
```python
Div(class_="card").attrs == {"class": "card"}
```
## Loading
`Tag.model_validate()` and `Tag.model_validate_json()` dispatch on `tag`, so
you can load a tree without knowing what its root is.
```python
Tag.model_validate(payload) # Div, P, whatever the root was
Div.model_validate(payload) # forces Div, no dispatch
```
Unknown tags fall back to the generic `Tag`, which still renders correctly.
Registering the element first, with `tag_class()`, gets you the class back
instead.
## What this is good for
Serialization is the reason the models are Pydantic rather than plain
classes. It buys a few things that string-building alone doesn't:
- **Caching a rendered tree** as JSON, then reloading and patching part of it
without re-parsing HTML.
- **Sending markup across a boundary** (a queue, an API) as structured data
that can be validated on arrival rather than trusted as a string.
- **Diffing two versions of a page** structurally instead of textually.
- **Storing user-authored content** in a form you can inspect and constrain,
rather than sanitizing HTML strings after the fact.
## Combining with parsing
Parsing and serializing compose, which is the whole point:
```python
from django_div import Tag, from_html
tree = from_html(response.text) # HTML -> models
payload = tree.model_dump_json() # models -> JSON
tree = Tag.model_validate_json(payload) # JSON -> models
print(tree) # models -> HTML
```
---
# Markdown
The tree renders to Markdown as well as HTML, and Markdown reads back into
the same kind of tree. Every example on this page is executed by
`tests/test_markdown.py`.
```python
from django_div import Div, H1, P, from_html
from django_div.markdown import from_markdown, to_markdown
to_markdown(Div(H1("Title"), P("Body text."))) # '# Title\n\nBody text.'
from_markdown("# Title") # H1(...)
```
`to_markdown()` needs nothing extra; `from_markdown()` needs the `markdown`
extra:
```console
uv add 'django-div[markdown]'
```
## Writing Markdown
`to_markdown()` takes one item, a list of items (what `parse()` and
`from_markdown()` return), or anything in between:
```python
to_markdown(H1("Title")) # '# Title'
to_markdown(Div(H1("Title"), P("Body"))) # blocks joined by blank lines
to_markdown([H1("A"), P("b")]) # '# A\n\nb'
```
Because parsing produces the same tree, `from_html` composes with it into an
HTML-to-Markdown converter:
```python
to_markdown(from_html("Title Body with a link .
"))
# '# Title\n\nBody with [a link](/x).'
```
### What maps to what
| HTML | Markdown |
| --- | --- |
| `h1`–`h6` | `#` … `######` |
| `p` | paragraph |
| `em` / `i`, `strong` / `b` | `*x*`, `**x**` |
| `s` / `del` | `~~x~~` |
| `code`, `pre` | `` `x` ``, fenced block (language from `class="language-*"`) |
| `a`, `img` | `[text](href)`, ``; `title` included when present |
| `ul` / `ol` / `li` | `-` / `1.` items, nesting indented, `start=` honored |
| `blockquote` | `>` prefixed lines |
| `table` | GFM pipe table; see [Tables](#tables) |
| `dl` / `dt` / `dd` | definition list; the term, then `: definition` attached |
| `hr`, `br` | `---`, backslash hard break |
| `div`, `section`, … | invisible; children render as blocks |
| `span`, `mark`, … | invisible; children flow through inline |
| `script`, `style`, `head`, … | dropped |
| everything else | falls back to its HTML, which Markdown permits |
The tables driving this (`INLINE_WRAPPERS`, `CONTAINER_TAGS`,
`TRANSPARENT_TAGS`, `DROP_TAGS`, `HEADING_TAGS`) are module constants, so
teaching the renderer a new element is one dict or set entry.
#### Element tables
These three sets decide what happens to an element with no Markdown
equivalent. Everything absent from all of them falls back to its HTML.
| Constant | Effect | Elements |
| --- | --- | --- |
| `CONTAINER_TAGS` | rendered as blocks, the element disappears | `article`, `aside`, `body`, `details`, `dialog`, `div`, `fieldset`, `figure`, `footer`, `form`, `header`, `hgroup`, `html`, `main`, `menu`, `nav`, `search`, `section` |
| `TRANSPARENT_TAGS` | children flow through inline, no HTML fallback | `abbr`, `bdi`, `bdo`, `cite`, `data`, `dfn`, `kbd`, `label`, `mark`, `output`, `q`, `rp`, `rt`, `ruby`, `samp`, `slot`, `small`, `span`, `sub`, `sup`, `time`, `u`, `var` |
| `DROP_TAGS` | removed with their content | `head`, `link`, `meta`, `script`, `style`, `template`, `title` |
`INLINE_WRAPPERS` maps an inline element to a symmetric Markdown wrapper:
| Tag | Renders `x` as |
| --- | --- |
| `b` | `**x**` |
| `del` | `~~x~~` |
| `em` | `*x*` |
| `i` | `*x*` |
| `ins` | `x` (unwrapped) |
| `s` | `~~x~~` |
| `strong` | `**x**` |
`HEADING_TAGS` maps a heading to its prefix:
| Tag | Prefix |
| --- | --- |
| `h1` | `#` |
| `h2` | `##` |
| `h3` | `###` |
| `h4` | `####` |
| `h5` | `#####` |
| `h6` | `######` |
### Tables
Tables get the fullest treatment, because they are where HTML-to-Markdown
conversions usually fall apart:
```python
Table(
Caption("Prices"),
Thead(Tr(Th("Item"), Th("Cost", style={"text_align": "right"}))),
Tbody(Tr(Td("Apple"), Td("1"))),
)
```
```markdown
Prices
| Item | Cost |
| --- | --: |
| Apple | 1 |
```
- **Alignment** comes from `text-align` styles (string or mapping) or the
legacy `align` attribute, and round-trips: `from_markdown` keeps the
alignment markdown-it records, and `to_markdown` emits it back as
`:--` / `:-:` / `--:`.
- **`thead` / `tbody` / `tfoot` render in that order**, matching how HTML
displays them, even when the source declares them differently.
- **A caption** becomes a paragraph above the table, because GFM has no caption.
- **A headerless table** gets an empty header row, since GFM requires one;
the first data row is not promoted.
- **Hard breaks and block content in cells** flatten to ` `, which GFM
permits, so a cell can never split its own row.
- **Nested tables** stay inside their cell as HTML, exactly once.
- **`colspan`/`rowspan` have no GFM form**, so such tables fall back to
their HTML rather than silently misplacing data.
### Fences and code
Content containing backticks can't break out of its own code span or fence, because
the marker grows past it:
`````python
to_markdown(Pre("a ``` b")) # '````\na ``` b\n````'
to_markdown(P(Code("uses ` tick"))) # '`` uses ` tick ``'
`````
### Lossy on purpose
Markdown has no home for `class`, `id`, `data-*`, or most other attributes,
so they are dropped. Text is emitted verbatim, not escaped, so content that
looks like Markdown syntax will be treated as Markdown by whatever renders
the output. Treat `to_markdown()` as a conversion, not an encoding:
round-trips preserve structure, not bytes.
## Reading Markdown
`from_markdown()` returns the same kind of tree as `from_html()`: typed
element classes, not a foreign AST. It deliberately contains no Markdown
parser: markdown-it-py renders CommonMark plus GFM tables and strikethrough,
and the HTML comes back through `parse()`.
```python
tree = from_markdown("# Title\n\nBody text.")
tree[0].tag # 'h1'
tree[0].text # 'Title'
```
A single-root document unwraps to the item itself, like `from_html()`;
anything else is a list.
### The tree is the point
Everything that works on a parsed HTML tree works on a parsed Markdown
document, including searching, editing, and serializing:
```python
doc = from_markdown("# Guide\n\nSee [the docs](/docs) and [the api](/api).")
[(a.text, a.attrs["href"]) for a in doc[1].find_all("a")]
# [('the docs', '/docs'), ('the api', '/api')]
from_markdown("# Title").model_dump() # {'tag': 'h1', ...}
```
And because `to_markdown()` accepts the same tree back, Markdown documents
can be edited *structurally*, with no regexes over source text:
```python
doc = from_markdown("# Title\n\n## Section\n\nBody.")
for item in doc:
if item.tag in ("h1", "h2"):
item.tag = f"h{int(item.tag[1]) + 1}" # demote one level
to_markdown(doc) # '## Title\n\n### Section\n\nBody.'
```
### Fidelity
Reading then writing is stable: a second round trip reproduces the first,
and fenced code keeps its language and content exactly:
````python
doc = from_markdown("```python\nif a < b:\n go()\n```")
doc.find("code").attrs["class"] # 'language-python'
to_markdown(doc) # the same fence back, byte for byte
````
Alignment in tables survives the loop too. See [Tables](#tables).
---
# Django
Django is optional. Only `django_div.django` imports it, so installing
django-div in a non-Django project pulls in nothing.
## Escaping, both directions
Rendering escapes text and attribute values, so output is safe markup by
construction. A tag can go straight into a template context with no `|safe`:
```django
{{ card }}
```
Why `__html__` alone wasn't enough
Django's template engine calls `str()` on any value that isn't already a
string *before* it looks for `__html__`. A plain `str` return would have
been escaped, so `__str__` itself reports the result as safe.
Interop runs the other way too. Anything carrying `__html__` (a
`SafeString`, a `markupsafe.Markup`, a rendered form) passes through a tag
unescaped, while ordinary strings are still escaped:
```python
Div(mark_safe("bold ")) # bold
Div("bold ") # <b>bold</b>
Div(form.as_p()) # the form's own markup, intact
```
Lazy objects resolve correctly, so translations work:
```python
Div(gettext_lazy("Hello")) # Hello
```
### Component return values
Plain-string component results are escaped, just like text children.
For example, returning `"Hello "` produces
`<b>Hello</b>`. Return elements for structured HTML, or use
`Raw(content=trusted_html)` for intentional, trusted markup. Objects with
`__html__`, including Django `SafeString`, retain their trusted status.
This changes the previous behavior that treated all component string
results as safe HTML. Existing components returning HTML strings should
return elements or explicitly trusted markup instead.
## Components as templates
`DjangoDivTemplates` is a template backend whose templates are Python
callables. Register it alongside your existing engines:
settings.py
```python
TEMPLATES = [
{
"BACKEND": "django_div.django.DjangoDivTemplates",
"NAME": "django_div",
"DIRS": [],
"APP_DIRS": False,
"OPTIONS": {
"context_processors": [
"django.template.context_processors.request",
"django.contrib.auth.context_processors.auth",
],
},
},
{
"BACKEND": "django.template.backends.django.DjangoTemplates",
# ... your usual configuration, untouched
},
]
```
A component is a callable, usually returning an `HtmlItem`, addressed by its
dotted path. Plain-string results are escaped; explicitly trusted markup is
preserved as described above:
myapp/components.py
```python
from django_div import H1, Div, Li, P, Ul
def home(title, items, **context):
return Div(
H1(title),
Ul(Li(item) for item in items),
class_="page",
)
```
myapp/views.py
```python
from django.shortcuts import render
def home_view(request):
return render(request, "myapp.components.home", {
"title": "Hi",
"items": ["one", "two"],
})
```
Because it is a normal backend, `render()`, `render_to_string()`,
`get_template()`, and the generic class-based views all work unchanged.
### Context handling
A component receives the context as keyword arguments:
- Declare `**kwargs` and you get the whole context.
- Name your parameters and you get only those.
```python
def greet(name="world"): # ignores request, user, csrf_token, ...
return P(f"Hello, {name}")
```
That second case matters. Context processors add `request`, `user`, `perms`,
and more to every render; without the filtering, adding one would break every
component signature at once.
### Incremental adoption
A component that can't be resolved raises `TemplateDoesNotExist`, so Django's
loader falls through to the next engine. Ordinary `.html` templates keep
working, and you can convert one view at a time.
Template names are import paths
`"myapp.components.home"` is imported and called. Never build a template
name from user input. That is an arbitrary-import primitive.
## Without the template layer
For views that build their own markup:
```python
from django_div import H1, Div, Form, Input
from django_div.django import as_response, csrf_input
def index(request):
return as_response(Div(H1("Hi")))
def search(request):
return as_response(
Form(
csrf_input(request),
Input(name="q", placeholder="Search"),
method="post",
)
)
```
`as_response()` passes extra keyword arguments to `HttpResponse`, so
`status=` and `content_type=` work as usual.
`csrf_input()` produces the hidden token field using Django's own machinery,
so it stays correct if that changes.
## Rendering a tag yourself
`render()` returns a `SafeString` when Django is installed, and a plain `str`
otherwise:
```python
Div("hi").render()
```
## What this is not
django-div does not replace Django's template language, and does not try to.
There is no inheritance, no `{% block %}`, and no partial loading. A
component is a function, so composition is function calls and default
arguments instead.
If you want template inheritance, keep those pages in Django templates and
use django-div for the parts that benefit from being Python.
## Components returning siblings
Return a `Fragment` to render siblings without an extra wrapper:
```python
from django_div import Fragment, H1, P
def content(title):
return Fragment(H1(title), P("Body"))
```
It works with the template backend and `as_response()` just like a tag.
The backend caches signature metadata for ordinary functions, up to 256
entries. It still calls the component and builds its output for each render.
---
# Cookbook
Practical recipes, none of them Django-specific. The Python recipes are covered by
`tests/test_cookbook.py`; browser-side JavaScript illustrates how to use the
rendered markup.
For Django-specific recipes, see the [Django cookbook](../django-cookbook/).
## Building
### A reusable component
A component is a function. There is no registry, no base class, and no
special syntax. Composition is a function call.
```python
def card(title, *body, href=None):
heading = A(title, href=href) if href else title
return Div(H2(heading), Div(*body, class_="card-body"), class_="card")
card("Hello", P("Body"), href="/x")
```
```html
```
Taking `*body` and passing it through keeps the caller's syntax identical to
a built-in element's.
### Reusable button variants
Use `with_attrs()` to derive variants from a configured element:
```python
from django_div import Button
base = Button("Save", class_="btn", type="submit")
primary = base.with_attrs(class_="btn primary", disabled=True)
```
```html
Save
```
`base` still has `class="btn"` and no `disabled` attribute. Classes are
replaced as a whole, so include any base classes you want to keep. Pass
`None` to omit an attribute on render. The attribute dictionary and child
list are copied, while existing child objects and nested attribute values
remain shared.
### A disclosure with matching ARIA state
Use the same Python flag to describe whether a panel is expanded and
whether its content is hidden:
```python
from django_div import Button, Div, Fragment
def disclosure(content, *, expanded=False):
return Fragment(
Button(
"Details", type="button", aria_controls="details",
aria_expanded=expanded,
),
Div(content, id="details", hidden=not expanded),
)
disclosure("More information")
```
```html
Details More information
```
ARIA booleans render as explicit `"true"` or `"false"`; HTML booleans such
as `hidden` render bare when true and disappear when false. This recipe
sets the initial state. If JavaScript toggles the panel, update both
`aria-expanded` and `hidden`. Give each panel a unique ID when rendering
multiple disclosures.
### A whole document
A doctype isn't an element, so it has its own item class. `Doctype()`
renders the HTML5 doctype, and a document is a list:
```python
def document(title, *body, lang="en"):
return [
Doctype(),
Html(Head(Meta(charset="utf-8"), Title(title)), Body(*body), lang=lang),
]
def render(items):
return "".join(str(item) for item in items)
render(document("Home", H1("Hi")))
```
```html
Home Hi
```
### HTML templates for JavaScript
Use `Template` to hold HTML that JavaScript can clone and insert later.
Build its contents with elements as usual:
```python
from django_div import Div, Span, Template
template = Template(
Div(Span("Hello"), class_="card"),
id="card-template",
)
```
For an existing trusted HTML string, wrap the markup in `Raw`:
```python
from django_div import Raw, Template
template = Template(
Raw(content='Hello
'),
id="card-template",
)
```
Both produce the same HTML when rendered with `str(template)`:
```html
Hello
```
Once the template is in the document, JavaScript can insert its contents:
```javascript
const template = document.querySelector("#card-template");
document.body.append(template.content.cloneNode(true));
```
Ordinary string children are still escaped inside `Template`. Use `Raw`
only for trusted markup, never untrusted user input. Placeholders such as
`{{ name }}` remain literal; django-div does not evaluate them.
### A table from data
```python
def data_table(rows, columns):
return Table(
Thead(Tr(Th(column) for column in columns)),
Tbody(Tr(Td(row[column]) for column in columns) for row in rows),
)
data_table([{"name": "Ana", "age": 33}], ["name", "age"])
```
```html
```
Nested generators work because each one is flattened as it is consumed.
### Navigation with an active item
A `class` mapping turns a condition into a class, and an all-false mapping
drops the attribute rather than emitting `class=""`.
```python
def nav(links, current):
return Nav(
Ul(Li(A(label, href=url, class_={"active": url == current}))
for label, url in links)
)
nav([("Home", "/"), ("Docs", "/docs/")], "/docs/")
```
```html
```
A generator can't sit beside a keyword argument
This is a Python rule, not a django-div one:
```python
Ul(Li(x) for x in items, class_="errors") # SyntaxError
```
Add brackets and it's fine:
```python
Ul([Li(x) for x in items], class_="errors")
```
### Custom elements and web components
```python
MyWidget = tag_class("my-widget")
MyWidget("hi", data_state="ready")
```
```html
hi
```
`tag_class()` registers the result **globally**, in `TAG_CLASSES`, so parsing
produces that class too. That is the point, but it means the registry grows at
runtime. `BUILTIN_TAGS` is the fixed set this library ships. For a one-off
that shouldn't be registered, `Tag("my-widget", ...)` skips it.
### Passing data to JavaScript
Use `JsonScript` to pass data to JavaScript. It escapes characters that could
close the script element while preserving the original JSON values:
```python
from django_div import JsonScript
JsonScript({"a": ""}, id="config")
```
```html
```
Read the data after the element is in the document:
```javascript
const config = JSON.parse(document.getElementById("config").textContent);
```
`JsonScript` accepts dictionaries, lists, and nested Pydantic models.
Models use their field aliases and JSON-compatible values; `None` is
preserved as JSON `null`. No Django dependency is required.
Use this helper for data instead of interpolating values into executable
`Script` content. Plain HTML escaping is unsuitable inside a script element.
### Open Graph metadata
Group page metadata in a reusable function. `Fragment` emits sibling meta
elements without adding a wrapper. Use `page_type` to avoid shadowing
Python's built-in `type`; the metadata property remains `og:type`.
```python
from django_div import Fragment, Head, Meta, Title
def open_graph(*, title, url, image, description, page_type="website"):
return Fragment(
Meta(property="og:title", content=title),
Meta(property="og:type", content=page_type),
Meta(property="og:url", content=url),
Meta(property="og:image", content=image),
Meta(property="og:description", content=description),
)
head = Head(
Title("Introducing django-div"),
open_graph(
title="Introducing django-div",
url="https://example.com/posts/django-div/",
image="https://example.com/images/django-div.png",
description="Build and parse HTML in Python.",
page_type="article",
),
)
```
`str(head)` produces the following HTML (line breaks added for readability):
```html
Introducing django-div
```
Use absolute URLs for the page and image. Values are escaped as ordinary
attributes, so no `Raw` wrapper is needed. Omit `page_type` to use the default
`"website"`. Extend the function with metadata your project needs, such as
`og:site_name` or `og:image:alt`. See the [Open Graph protocol](https://ogp.me/).
### JSON-LD from a Pydantic model
Schema.org markup is JSON in a `` would close the tag early and the rest of
the payload would land on the page as live markup. `JsonLd` writes the
dangerous characters as the `\uXXXX` escapes JSON already understands, which
changes no data and makes that impossible.
Any Pydantic model works; nothing has to inherit from anything in
django-div. `@context` and `@type` are not Python names, so declare them as
aliases, which is also how you reach schema.org's camelCase properties like
`sameAs`:
```python
from pydantic import BaseModel, ConfigDict, Field
from django_div import Head, JsonLd, Title
class Organization(BaseModel):
model_config = ConfigDict(populate_by_name=True)
context: str = Field("https://schema.org", alias="@context")
type: str = Field("Organization", alias="@type")
name: str
url: str
logo: str | None = None
same_as: list[str] | None = Field(None, alias="sameAs")
head = Head(
Title("Acme"),
JsonLd(Organization(name="Acme", url="https://acme.example")),
)
```
Models are dumped by alias with `None` dropped, so `logo` and `same_as` never
render. A JSON-LD null is not a value, so dropping it is what the format
means anyway. To dump on other terms, call `model_dump()` yourself and pass
the dict.
If you write a lot of these, the two constant fields are worth a base class
of your own -- but that is your vocabulary to shape, not something django-div
should decide for you.
Pass a list to emit several objects at once, and read one back out of a page
with `json.loads(tag.text)`:
```python
import json
from django_div import from_html
page = from_html(str(head))
data = json.loads(page.find("script", type="application/ld+json").text)
```
### SVG icons
SVG elements aren't generated, because SVG's `` would collide with the
`Text` model. Build the ones you need:
```python
Svg = tag_class("svg")
Use = tag_class("use")
def icon(name, size=16):
return Svg(
Use(href=f"/static/icons.svg#{name}"),
width=size, height=size, aria_hidden="true",
)
```
```html
```
### Inline SVG shapes
SVG attribute names are case-sensitive, and a name without an underscore
passes through untouched, so `viewBox` renders as `viewBox`. Underscores
still become hyphens, which is what `stroke_width` needs:
```python
Svg = tag_class("svg")
Circle = tag_class("circle")
Path = tag_class("path")
check = Svg(
Circle(cx="12", cy="12", r="10", fill="none"),
Path(d="M8 12l3 3 5-6"),
viewBox="0 0 24 24", width=24, height=24,
stroke="currentColor", stroke_width="2", aria_hidden="true",
)
```
```html
```
Shapes render as a full open and close pair, never self-closed, because only
HTML void elements self-close. A browser accepts ` ` in inline
SVG.
Parsing SVG loses the attribute case
Building SVG keeps `viewBox`. Reading it back does not: an HTML parser
folds every attribute name to lowercase, so `from_html` returns
`viewbox`, which a browser ignores. The same applies to
`preserveAspectRatio`, `gradientUnits`, and the other camelCase names.
Treat SVG as write-only, or repair the names yourself after a parse.
### XML, not just HTML
`Tag` doesn't care whether a name is HTML, so feeds and other XML work:
```python
Tag("rss",
Tag("channel", Tag("title", "News"), Tag("item", Tag("title", "First"))),
version="2.0")
```
```html
News First
```
Note
Only HTML void elements self-close, and HTML escaping rules are applied.
For heavy XML work a dedicated library is a better fit.
## Parsing
These need the `parse` extra.
### Extract every link
```python
page = from_html(markup)
[(a.text, a.attrs["href"]) for a in page.find_all("a")]
```
```python
[("A", "/a"), ("B", "/b")]
```
### Find links by class token
Pass a predicate when a class can appear alongside other classes. Keyword
attributes narrow the results using their existing exact-match behavior:
```python
from django_div import from_html
page = from_html(
''
)
links = page.find_all(
lambda node: node.tag == "a" and node.has_class("external"),
target="_blank",
)
[link.attrs["href"] for link in links]
# ['/one']
```
`find()` takes the same predicate and stops at the first match;
`iter_find()` yields matches lazily. All three search descendant tags in
document order, excluding the root. To compare a complete class attribute
instead, continue using `find_all("a", class_="external")`.
### Make relative URLs absolute
```python
from urllib.parse import urljoin
def absolutize(tree, base):
for link in tree.find_all("a"):
if "href" in link.attrs:
link.attrs["href"] = urljoin(base, link.attrs["href"])
return tree
absolutize(from_html(''), "https://example.test/")
```
```html
```
### Build a table of contents
```python
def table_of_contents(tree, levels=("h2", "h3")):
return Ul(
Li(A(heading.text, href="#" + heading.attrs["id"]))
for heading in tree.find_all()
if heading.tag in levels and "id" in heading.attrs
)
```
```html
```
Parsing and building in the same expression is the point: the input is HTML
and so is the output.
### Scrape a table into dicts
```python
def cells(row):
return [cell.text.strip() for cell in row.find_all() if cell.tag in {"th", "td"}]
def table_to_dicts(table):
rows = table.find_all("tr")
headers = cells(rows[0])
return [dict(zip(headers, cells(row))) for row in rows[1:]]
```
```python
[{"name": "Ana", "age": "33"}]
```
### Keep only certain elements
Transform a copy to remove unwanted nodes and unwrap other containers.
Children are processed before parents; returning a fragment preserves
already-filtered children without their original wrapper. The source remains
unchanged. Import `Fragment`, `Tag`, `Text`, and `parse` from `django_div`.
```python
KEEP = {"p", "b", "i", "em", "strong", "a", "ul", "ol", "li", "code", "br"}
KEEP_ATTRS = {"a": {"href", "title"}}
DROP_ENTIRELY = {"script", "style"}
def keep_node(item):
if isinstance(item, (Text, Fragment)):
return item
if isinstance(item, Tag):
if item.tag in DROP_ENTIRELY:
return None
if item.tag not in KEEP:
return Fragment(item.children)
item.attrs = {
name: value for name, value in item.attrs.items()
if name in KEEP_ATTRS.get(item.tag, set())
}
return item
return None
def keep_only(items):
return Fragment(items).transform(keep_node)
```
```python
items = parse('ok b
')
str(keep_only(items))
# ok b
```
This is not a sanitizer
It is a shape filter for content you already trust: trimming a CMS
export, normalizing pasted markup. It is **not** an XSS defense. Real
sanitization has to handle `javascript:` URLs, CSS escapes, mutation
XSS, and namespace confusion. For untrusted input use
[nh3](https://pypi.org/project/nh3/) or
[bleach](https://pypi.org/project/bleach/).
Note that `DROP_ENTIRELY` exists because unwrapping a `