expand abbreviations from a definition line

`*[TERM]: expansion` on its own line, PHP Markdown Extra's form, and every
whole-word use in the document becomes <abbr title="…">.

A transformer rather than an inline parser, because a definition may appear
after the use it explains and a parser only ever sees what it has already read.
A block parser rather than a pattern found later, because the line has to stop
being content — an author would notice that going wrong before anything else.

Whole-word matching is the part that would have bitten: without it a definition
of HTML quietly rewrites HTMLish and xHTML too, so both have cases. Code spans,
autolinks, raw HTML and an already-expanded term are skipped, the longest
definition wins where two could match, and the expansion is escaped into the
attribute so a quoted phrase cannot end it.

Definitions are document-scoped. A term defined in a page does not reach an
included fragment, which is parsed on its own bytes exactly as footnotes are —
stated in content-model.md rather than left to be discovered.

The nesting gate caught firstMatch four levels deep; the inner search is its own
function now, which reads better than it did before the warning.

core 2794/2800, ext 1483/2000, 34 gates green, 0 warnings.
This commit is contained in:
Claude Opus 5
2026-08-01 21:25:44 +06:00
committed by bdeshi
parent 53e9ef473a
commit 661fb469e8
10 changed files with 393 additions and 26 deletions
+6 -1
View File
@@ -320,13 +320,18 @@ them is a render transform, and the list lives in `cmd/khosra/wire.go` where fea
| Subscript | `H~2~O` |
| Superscript | `10^6^` |
| Highlight | `==marked==` |
| Abbreviations | `*[HTML]: HyperText Markup Language` on its own line; every whole-word use expands |
| Heading ids | automatic, from the heading's text — the anchor a table of contents needs |
The last four are khosra's own (ADR-0061), because goldmark's strikethrough claims a single tilde as well as
The last five are khosra's own (ADR-0061, ADR-0062), because goldmark's strikethrough claims a single tilde as well as
a double and would read `H~2~O` as struck text. **A subscript or superscript may not contain a space** — it
holds a formula, not a phrase — which is what keeps `x^2 + y^2` prose. Its content is taken literally, so
there is no emphasis inside one.
An abbreviation's definition may sit anywhere in the file, before or after the uses it explains, and renders
nothing itself. It applies to that document only: a term defined in a page does not reach an included
fragment, which is parsed on its own bytes exactly as footnotes are.
Deliberately absent, so their absence is a decision rather than an oversight: **task lists** (a note-taking
affordance, not a publishing one), **linkify** (it rewrites an author's plain text into markup, which is the
line ADR-0034 draws), **CJK line breaking** (wrong script family — it does nothing for Bengali), and the
+17
View File
@@ -989,3 +989,20 @@ subscript, which is a limit worth stating rather than discovering; and khosra no
implementation it used to get from upstream.
Revisit if: a mark wants markup inside a single run — which needs recursive inline parsing and is a
different mechanism, not a wider table.
## ADR-0062 — Abbreviations are a definition line plus a replacement pass
Date: 2026-08-01 · Status: accepted (extends ADR-0061's package with a second mechanism)
Decision: `*[TERM]: expansion` on its own line defines an abbreviation, renders nothing itself, and every
whole-word occurrence of TERM in that document becomes `<abbr title="expansion">`. A block parser claims the
line; an AST transformer does the replacing. Code spans, autolinks, raw HTML and an already-expanded term
are skipped. The longest defined term wins where two could match.
Why: PHP Markdown Extra's form, so it is a syntax authors already know. It is a transformer rather than an
inline parser because a definition may appear *after* the use it explains — a parser only ever sees what it
has already read — and a block parser rather than a pattern found later because the line has to stop being
content, which is the part an author would notice going wrong. Whole-word matching is what keeps `HTMLish`
and `xHTML` intact; without it a definition quietly rewrites every substring on the page.
Consequence: cheap — one more mark from the same package, and the definition can sit anywhere in the file.
Expensive — the replacement walks every text node in the document, which is work proportional to the page
rather than to the number of definitions; and definitions are document-scoped, so a term defined in a page
does not reach an included fragment, which is parsed on its own bytes (ADR-0038) exactly as footnotes are.
Revisit if: definitions want to be site-wide, which is a settings-cascade question and not this mechanism.
+2 -2
View File
@@ -26,7 +26,7 @@ table owns.
| `internal/render/chrome.go` | the engine's own words: phrase table, month names, digits, and the `t`/`num`/`day` template funcs (ADR-0034) |
| `internal/render/templates/` | reference theme, complete: `base.html` (shell, navigation, language links, feed and OpenGraph), `page.html` (bundle, sequence, tags, extras), `list.html`, `extras.html`, `shortcodes.html`, `theme.css` (ADR-0026, ADR-0049) |
| `internal/ext/shortcodes/` | first feature: `::name{key=value}` block parser and node renderer, rendering through a theme fragment (ADR-0036). `figure`, `gallery`, `include`, plus the derivative pass and remembered picture inspection (ADR-0042, ADR-0044). `FootnotePrefix` namespaces an included file's footnote ids (ADR-0058). Directive syntax since ADR-0059 |
| `internal/ext/notation/` | the inline marks CommonMark lacks: `~sub~`, `^sup^`, `==mark==`, and `~~strike~~`, which it owns so a single tilde can mean subscript (ADR-0061) |
| `internal/ext/notation/` | the inline marks CommonMark lacks: `~sub~`, `^sup^`, `==mark==`, and `~~strike~~`, which it owns so a single tilde can mean subscript (ADR-0061). `abbr.go` adds `*[TERM]:` definitions and the pass that expands them (ADR-0062) |
| `internal/ext/scaffold/` | writes one draft directory bundle into a site root through `os.Root`: never an overwrite |
| `internal/ext/watch/` | polls `content/` and `templates/` on an interval it is given, ignores editor droppings, and reports a settled change (ADR-0022, ADR-0048, ADR-0056). `site.yaml` is deliberately not fingerprinted (ADR-0055) |
| `internal/ext/check/` | third feature: validates a site root — what the engine worked around, broken internal links, missing titles and alt text, mixed series ordering, and calls left in the retired shortcode form (ADR-0059) |
@@ -48,7 +48,7 @@ section-narrowed, sequence navigation and a series archive on any nested bundle,
URL, generated derivatives under `/derived/`, Atom feeds per site,
section and tag, a bundle's extras as a browsable tree, plus `/robots.txt` and `/sitemap.xml`.
Markdown is CommonMark plus tables, footnotes, definition lists, heading ids and the inline marks
(`~sub~`, `^sup^`, `==mark==`, `~~strike~~`), and nothing else (ADR-0058, ADR-0061); HTML an author writes renders, because the site root is trusted (ADR-0060). Chrome text, dates and digits render in English or Bengali; authored text is untouched but for typographic
(`~sub~`, `^sup^`, `==mark==`, `~~strike~~`, `*[TERM]:`), and nothing else (ADR-0058, ADR-0061, ADR-0062); HTML an author writes renders, because the site root is trusted (ADR-0060). Chrome text, dates and digits render in English or Bengali; authored text is untouched but for typographic
smoothing (ADR-0034); line breaking is left to CSS (ADR-0045). This repo holds engine source only — the site root is external and passed with
`khosra check` validates a site root and exits non-zero on anything that makes it wrong; `khosra new`
scaffolds a draft bundle into one. A running server notices changes under `content/` and `templates/` by
+46 -19
View File
@@ -113,31 +113,58 @@ check.go 223 · doc.go 8
- check.go:172 func asset(fsys fs.FS, trimmed string, site *content.Site) bool
- check.go:200 func mixedOrdering(bundles []content.Bundle, site *content.Site) []Finding
## internal/ext/notation — 159 lines + 88 test
## internal/ext/notation — 411 lines + 149 test
doc.go 8 · notation.go 151
abbr.go 246 · doc.go 8 · notation.go 157
- abbr.go:17 var defined = parser.NewContextKey()
- abbr.go:21 var definitionKind = ast.NewNodeKind("AbbrDefinition")
- abbr.go:23 type definition struct{ ast.BaseBlock }
- abbr.go:25 func (n *definition) Kind() ast.NodeKind { return definitionKind }
- abbr.go:27 func (n *definition) Dump(source []byte, level int) { ast.DumpHelper(n, source, level, nil, nil) }
- abbr.go:30 var abbrKind = ast.NewNodeKind("Abbr")
- abbr.go:32 type abbr struct
- abbr.go:37 func (n *abbr) Kind() ast.NodeKind { return abbrKind }
- abbr.go:39 func (n *abbr) Dump(source []byte, level int) { ast.DumpHelper(n, source, level, nil, nil) }
- abbr.go:45 type definitionParser struct{}
- abbr.go:47 func (definitionParser) Trigger() []byte { return []byte{'*'} }
- abbr.go:49 func (definitionParser) Open(parent ast.Node, reader text.Reader, pc parser.Context) (ast.Node, parser.State)
- abbr.go:65 func (definitionParser) Continue(node ast.Node, reader text.Reader, pc parser.Context) parser.State
- abbr.go:69 func (definitionParser) Close(node ast.Node, reader text.Reader, pc parser.Context) {}
- abbr.go:71 func (definitionParser) CanInterruptParagraph() bool { return true }
- abbr.go:73 func (definitionParser) CanAcceptIndentedLine() bool { return false }
- abbr.go:77 func parseDefinition(line string) (term, expansion string, ok bool)
- abbr.go:101 type expand struct{}
- abbr.go:103 func (expand) Transform(doc *ast.Document, reader text.Reader, pc parser.Context)
- abbr.go:136 func replace(node *ast.Text, source []byte, terms map[string]string, ordered []string)
- abbr.go:172 func slice(seg text.Segment, from, to int) text.Segment
- abbr.go:180 func firstMatch(body string, ordered []string) (string, int)
- abbr.go:193 func earliest(body, term string) int
- abbr.go:209 func whole(body string, index, length int) bool
- abbr.go:226 func renderDefinition(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error)
- abbr.go:232 func renderAbbr(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error)
- abbr.go:243 func (nodeRenderer) registerAbbr(reg renderer.NodeRendererFuncRegisterer)
- notation.go:17 var marks = []mark{
- notation.go:27 type mark struct
- notation.go:33 func New() goldmark.Extender { return extension{} }
- notation.go:35 type extension struct{}
- notation.go:41 func (extension) Extend(md goldmark.Markdown)
- notation.go:51 var kind = ast.NewNodeKind("Notation")
- notation.go:54 type node struct
- notation.go:59 func (n *node) Kind() ast.NodeKind { return kind }
- notation.go:61 func (n *node) Dump(source []byte, level int) { ast.DumpHelper(n, source, level, nil, nil) }
- notation.go:65 type processor struct{ tag string }
- notation.go:67 func (p processor) IsDelimiter(b byte) bool { return b == '~' || b == '=' }
- notation.go:69 func (p processor) CanOpenCloser(opener, closer *parser.Delimiter) bool
- notation.go:73 func (p processor) OnMatch(consumes int) ast.Node { return &node{tag p.tag} }
- notation.go:76 type inlineParser struct{ m mark }
- notation.go:78 func (p inlineParser) Trigger() []byte { return []byte{p.m.char} }
- notation.go:80 func (p inlineParser) Parse(parent ast.Node, block text.Reader, pc parser.Context) ast.Node
- notation.go:89 func (p inlineParser) doubled(block text.Reader, line []byte, segment text.Segment, pc parser.Context) ast.Node
- notation.go:108 func (p inlineParser) single(block text.Reader, line []byte, segment text.Segment) ast.Node
- notation.go:137 type nodeRenderer struct{}
- notation.go:139 func (nodeRenderer) RegisterFuncs(reg renderer.NodeRendererFuncRegisterer)
- notation.go:143 func render(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error)
- notation.go:56 var kind = ast.NewNodeKind("Notation")
- notation.go:59 type node struct
- notation.go:64 func (n *node) Kind() ast.NodeKind { return kind }
- notation.go:66 func (n *node) Dump(source []byte, level int) { ast.DumpHelper(n, source, level, nil, nil) }
- notation.go:70 type processor struct{ tag string }
- notation.go:72 func (p processor) IsDelimiter(b byte) bool { return b == '~' || b == '=' }
- notation.go:74 func (p processor) CanOpenCloser(opener, closer *parser.Delimiter) bool
- notation.go:78 func (p processor) OnMatch(consumes int) ast.Node { return &node{tag p.tag} }
- notation.go:81 type inlineParser struct{ m mark }
- notation.go:83 func (p inlineParser) Trigger() []byte { return []byte{p.m.char} }
- notation.go:85 func (p inlineParser) Parse(parent ast.Node, block text.Reader, pc parser.Context) ast.Node
- notation.go:94 func (p inlineParser) doubled(block text.Reader, line []byte, segment text.Segment, pc parser.Context) ast.Node
- notation.go:113 func (p inlineParser) single(block text.Reader, line []byte, segment text.Segment) ast.Node
- notation.go:142 type nodeRenderer struct{}
- notation.go:144 func (r nodeRenderer) RegisterFuncs(reg renderer.NodeRendererFuncRegisterer)
- notation.go:149 func render(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error)
## internal/ext/scaffold — 102 lines + 89 test
@@ -254,7 +281,7 @@ chrome.go 110 · render.go 473 · view.go 130
- view.go:111 type Item struct
- view.go:122 type Alternate struct
## internal/web — 734 lines + 1605 test
## internal/web — 734 lines + 1608 test
asset.go 58 · discover.go 71 · extras.go 93 · feed.go 125 · resolve.go 170 · web.go 217
+1 -1
View File
@@ -152,7 +152,7 @@ not insert characters into an author's prose to influence line breaking.
`.HTML` is Markdown output, so the dialect decides which elements a theme has to be ready to style
(`content-model.md`). Beyond CommonMark's own, since ADR-0058: `<table>` with `<thead>`/`<tbody>`,
`<dl>`/`<dt>`/`<dd>`, `<del>`, `<sub>`, `<sup>`, `<mark>`, and goldmark's footnote markup — `<sup id="fnref:N">` in the text and a
`<dl>`/`<dt>`/`<dd>`, `<del>`, `<sub>`, `<sup>`, `<mark>`, `<abbr title=…>`, and goldmark's footnote markup — `<sup id="fnref:N">` in the text and a
`<div class="footnotes">` carrying an `<ol>` of notes. Headings arrive with an `id`.
None of it is optional and none of it is configurable: a theme that styles none of these still renders a
@@ -19,7 +19,9 @@ The page has a footnote of its own[^page], a table, a definition list and ~~a st
The notation marks cover what prose needs most often: the gauge measures H~2~O to about 10^-3^ m, and the
==important== column is the second. Where there is no mark for something, HTML does it: the reading is taken
with <kbd>Shift</kbd> held. Content in the site root is the author's own, so it renders rather than being
dropped.
dropped. Readings go to the NIWA archive, and the definition below expands every use of it.
*[NIWA]: National Institute of Water and Atmospheric Research
| Gauge | Reading | Note |
|-------|---------|------|
+246
View File
@@ -0,0 +1,246 @@
package notation
import (
"sort"
"strings"
"unicode"
"unicode/utf8"
"github.com/yuin/goldmark/ast"
"github.com/yuin/goldmark/parser"
"github.com/yuin/goldmark/renderer"
"github.com/yuin/goldmark/text"
"github.com/yuin/goldmark/util"
)
// defined holds this document's abbreviations, keyed by the term as written.
var defined = parser.NewContextKey()
// definitionKind marks the line that declared one. It renders nothing: a definition is a note to the parser,
// not content.
var definitionKind = ast.NewNodeKind("AbbrDefinition")
type definition struct{ ast.BaseBlock }
func (n *definition) Kind() ast.NodeKind { return definitionKind }
func (n *definition) Dump(source []byte, level int) { ast.DumpHelper(n, source, level, nil, nil) }
// abbrKind is one expanded occurrence.
var abbrKind = ast.NewNodeKind("Abbr")
type abbr struct {
ast.BaseInline
title string
}
func (n *abbr) Kind() ast.NodeKind { return abbrKind }
func (n *abbr) Dump(source []byte, level int) { ast.DumpHelper(n, source, level, nil, nil) }
// definitionParser claims a whole line of the form `*[TERM]: expansion`.
//
// Its own block parser rather than a pattern found later, because the line has to stop being content: left to
// the paragraph parser it would render as prose, which is what the author is least expecting.
type definitionParser struct{}
func (definitionParser) Trigger() []byte { return []byte{'*'} }
func (definitionParser) Open(parent ast.Node, reader text.Reader, pc parser.Context) (ast.Node, parser.State) {
line, seg := reader.PeekLine()
term, expansion, ok := parseDefinition(string(line))
if !ok {
return nil, parser.NoChildren
}
terms, _ := pc.Get(defined).(map[string]string)
if terms == nil {
terms = map[string]string{}
pc.Set(defined, terms)
}
terms[term] = expansion
reader.Advance(seg.Len() - 1)
return &definition{}, parser.NoChildren
}
func (definitionParser) Continue(node ast.Node, reader text.Reader, pc parser.Context) parser.State {
return parser.Close
}
func (definitionParser) Close(node ast.Node, reader text.Reader, pc parser.Context) {}
func (definitionParser) CanInterruptParagraph() bool { return true }
func (definitionParser) CanAcceptIndentedLine() bool { return false }
// parseDefinition reads `*[TERM]: expansion`. A space after the asterisk makes it a list item instead, which
// is why the bracket must follow immediately.
func parseDefinition(line string) (term, expansion string, ok bool) {
rest, found := strings.CutPrefix(strings.TrimRight(line, "\r\n"), "*[")
if !found {
return "", "", false
}
term, rest, found = strings.Cut(rest, "]")
if !found || term == "" {
return "", "", false
}
expansion, found = strings.CutPrefix(rest, ":")
if !found {
return "", "", false
}
expansion = strings.TrimSpace(expansion)
if expansion == "" {
return "", "", false
}
return term, expansion, true
}
// expand replaces every defined term in the document's text with an abbreviation.
//
// A transformer rather than an inline parser, because a definition may appear after the use it explains, and
// an inline parser only ever sees what it has already read.
type expand struct{}
func (expand) Transform(doc *ast.Document, reader text.Reader, pc parser.Context) {
terms, _ := pc.Get(defined).(map[string]string)
if len(terms) == 0 {
return
}
// Longest first, so a definition of "HTML5" is not shadowed by one of "HTML".
ordered := make([]string, 0, len(terms))
for term := range terms {
ordered = append(ordered, term)
}
sort.Slice(ordered, func(i, j int) bool { return len(ordered[i]) > len(ordered[j]) })
source := reader.Source()
var texts []*ast.Text
_ = ast.Walk(doc, func(n ast.Node, entering bool) (ast.WalkStatus, error) {
if !entering {
return ast.WalkContinue, nil
}
switch n.Kind() {
case ast.KindCodeSpan, ast.KindRawHTML, ast.KindAutoLink, abbrKind:
// An author's literal text, a tag, a URL, and an expansion already made: none of them is prose.
return ast.WalkSkipChildren, nil
case ast.KindText:
texts = append(texts, n.(*ast.Text))
}
return ast.WalkContinue, nil
})
for _, node := range texts {
replace(node, source, terms, ordered)
}
}
// replace swaps one text node for the alternating run of text and abbreviations it contains.
func replace(node *ast.Text, source []byte, terms map[string]string, ordered []string) {
parent := node.Parent()
if parent == nil {
return
}
seg := node.Segment
body := string(source[seg.Start:seg.Stop])
at := 0
var built []ast.Node
for at < len(body) {
term, index := firstMatch(body[at:], ordered)
if term == "" {
break
}
start := at + index
if start > at {
built = append(built, ast.NewTextSegment(slice(seg, at, start)))
}
marked := &abbr{title: terms[term]}
marked.AppendChild(marked, ast.NewTextSegment(slice(seg, start, start+len(term))))
built = append(built, marked)
at = start + len(term)
}
if len(built) == 0 {
return
}
if at < len(body) {
built = append(built, ast.NewTextSegment(slice(seg, at, len(body))))
}
for _, n := range built {
parent.InsertBefore(parent, node, n)
}
parent.RemoveChild(parent, node)
}
// slice is the segment covering body[from:to].
func slice(seg text.Segment, from, to int) text.Segment {
out := seg
out.Start, out.Stop = seg.Start+from, seg.Start+to
return out
}
// firstMatch finds the earliest whole-word occurrence of any term. `ordered` is longest first and the
// comparison is strict, so two terms starting at the same place resolve to the longer one.
func firstMatch(body string, ordered []string) (string, int) {
best, at := "", -1
for _, term := range ordered {
index := earliest(body, term)
if index >= 0 && (at < 0 || index < at) {
best, at = term, index
}
}
return best, at
}
// earliest is the first whole-word occurrence of one term, or -1. Its own function so firstMatch stays flat:
// skipping a match that turned out to sit inside a longer word is a loop, not a condition.
func earliest(body, term string) int {
for from := 0; from+len(term) <= len(body); {
index := strings.Index(body[from:], term)
if index < 0 {
return -1
}
index += from
if whole(body, index, len(term)) {
return index
}
from = index + 1
}
return -1
}
// whole reports whether the match at index stands alone rather than sitting inside a longer word.
func whole(body string, index, length int) bool {
if index > 0 {
before, _ := utf8.DecodeLastRuneInString(body[:index])
if unicode.IsLetter(before) || unicode.IsDigit(before) {
return false
}
}
if end := index + length; end < len(body) {
after, _ := utf8.DecodeRuneInString(body[end:])
if unicode.IsLetter(after) || unicode.IsDigit(after) {
return false
}
}
return true
}
// renderDefinition writes nothing: the line was a note to the parser.
func renderDefinition(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error) {
return ast.WalkSkipChildren, nil
}
// renderAbbr writes the element. The title is authored text going into an attribute, so it is escaped —
// correctness rather than trust: an apostrophe or a quote in an expansion would otherwise end the attribute.
func renderAbbr(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error) {
if !entering {
_, _ = w.WriteString("</abbr>")
return ast.WalkContinue, nil
}
_, _ = w.WriteString(`<abbr title="`)
_, _ = w.Write(util.EscapeHTML([]byte(n.(*abbr).title)))
_, _ = w.WriteString(`">`)
return ast.WalkContinue, nil
}
func (nodeRenderer) registerAbbr(reg renderer.NodeRendererFuncRegisterer) {
reg.Register(definitionKind, renderDefinition)
reg.Register(abbrKind, renderAbbr)
}
+8 -2
View File
@@ -43,7 +43,12 @@ func (extension) Extend(md goldmark.Markdown) {
for _, m := range marks {
inline = append(inline, util.Prioritized(inlineParser{m}, 500))
}
md.Parser().AddOptions(parser.WithInlineParsers(inline...))
md.Parser().AddOptions(
parser.WithInlineParsers(inline...),
// Above the paragraph parser, so a definition line stops being content (ADR-0062).
parser.WithBlockParsers(util.Prioritized(definitionParser{}, 99)),
parser.WithASTTransformers(util.Prioritized(expand{}, 200)),
)
md.Renderer().AddOptions(renderer.WithNodeRenderers(util.Prioritized(nodeRenderer{}, 500)))
}
@@ -136,8 +141,9 @@ func (p inlineParser) single(block text.Reader, line []byte, segment text.Segmen
// anything an author supplied, so it is written directly.
type nodeRenderer struct{}
func (nodeRenderer) RegisterFuncs(reg renderer.NodeRendererFuncRegisterer) {
func (r nodeRenderer) RegisterFuncs(reg renderer.NodeRendererFuncRegisterer) {
reg.Register(kind, render)
r.registerAbbr(reg)
}
func render(w util.BufWriter, source []byte, n ast.Node, entering bool) (ast.WalkStatus, error) {
+61
View File
@@ -86,3 +86,64 @@ func TestCodeSpansAreUntouched(t *testing.T) {
t.Errorf("a code span must survive verbatim:\n%s", got)
}
}
func TestAbbreviationsExpandWhereverTheyAppear(t *testing.T) {
got := html(t, "The HTML spec.\n\n*[HTML]: HyperText Markup Language\n\nHTML again.\n")
if n := strings.Count(got, `<abbr title="HyperText Markup Language">HTML</abbr>`); n != 2 {
t.Errorf("both uses should expand, got %d:\n%s", n, got)
}
// A definition may follow the use it explains, which is why this is a transformer and not a parser.
if strings.Contains(got, "*[HTML]") || strings.Contains(got, "HyperText Markup Language<") {
t.Errorf("the definition line is a note to the parser, not content:\n%s", got)
}
}
func TestAnAbbreviationIsAWholeWord(t *testing.T) {
got := html(t, "HTML and HTMLish and xHTML.\n\n*[HTML]: HyperText Markup Language\n")
if strings.Count(got, "<abbr") != 1 {
t.Errorf("only the standalone use should expand:\n%s", got)
}
for _, keep := range []string{"HTMLish", "xHTML"} {
if !strings.Contains(got, keep) {
t.Errorf("%q should survive untouched:\n%s", keep, got)
}
}
}
func TestTheLongestDefinitionWins(t *testing.T) {
got := html(t, "HTML5 and HTML.\n\n*[HTML]: Markup\n*[HTML5]: The fifth one\n")
if !strings.Contains(got, `<abbr title="The fifth one">HTML5</abbr>`) {
t.Errorf("HTML5 should not be shadowed by HTML:\n%s", got)
}
if !strings.Contains(got, `<abbr title="Markup">HTML</abbr>`) {
t.Errorf("the shorter term should still expand on its own:\n%s", got)
}
}
func TestCodeAndUrlsAreNotExpanded(t *testing.T) {
got := html(t, "Use `HTML` here and <https://HTML.example/> there, but HTML in prose.\n\n*[HTML]: Markup\n")
if strings.Count(got, "<abbr") != 1 {
t.Errorf("only the prose use should expand:\n%s", got)
}
if !strings.Contains(got, "<code>HTML</code>") {
t.Errorf("a code span is literal text:\n%s", got)
}
}
func TestAnExpansionCannotBreakItsAttribute(t *testing.T) {
got := html(t, `Term T here.`+"\n\n"+`*[T]: a "quoted" thing`+"\n")
if strings.Contains(got, `title="a "quoted" thing"`) {
t.Errorf("the attribute was broken by the expansion:\n%s", got)
}
if !strings.Contains(got, "&quot;quoted&quot;") {
t.Errorf("the expansion should be escaped into the attribute:\n%s", got)
}
}
// A list item starts with an asterisk too, so the definition form must not eat one.
func TestAListItemIsNotADefinition(t *testing.T) {
got := html(t, "* [a link](/x/): still a list\n")
if !strings.Contains(got, "<li>") {
t.Errorf("a list item must survive:\n%s", got)
}
}
+3
View File
@@ -124,6 +124,9 @@ var exampleFeatures = []featureCase{
{what: "notation marks become elements, and a single tilde is a subscript not a strike", path: "/writing/notes-on-water/", code: 200,
expect: []string{"H<sub>2</sub>O", "10<sup>-3</sup>", "<mark>important</mark>", "<del>a struck phrase</del>"},
absent: []string{"<del>2</del>"}},
{what: "an abbreviation expands and its definition line renders nothing", path: "/writing/notes-on-water/", code: 200,
expect: []string{`<abbr title="National Institute of Water and Atmospheric Research">NIWA</abbr>`},
absent: []string{"*[NIWA]"}},
{what: "the dialect renders tables, definition lists and strikethrough", path: "/writing/notes-on-water/", code: 200,
expect: []string{"<table>", "<th>Gauge</th>", "<dl>", "<dt>Monsoon</dt>", "<del>a struck phrase</del>"}},
{what: "a fragment's footnote ids are namespaced, so the page's own keep working", path: "/writing/notes-on-water/", code: 200,