Add the EPUB edition: converter, class-based audit, shared front matter

tools/make_epub.py converts the same src/ the PDF is set from, running
polish.py first, with original page numbers as page-list metadata, endnotes
gathered in one linked Notes section, and the colophon and errata included.
It raises on any macro it does not declare. tools/epub_audit.py checks by
class of fault (LaTeX residue, escaping, empty blocks, links, images, XML,
content, typography drift from the PDF); make epub runs both.

\byedition{PDF}{EPUB} lets the colophon carry the sentences that are true of
only one edition; the errata introduction moves into src/errata.tex so both
editions print one copy. reprocheck.py now accepts an .epub and skips
environment parameters that are layout, not copy.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-29 14:36:18 +06:00
co-authored by Claude Sonnet 5.5
parent 2ec880d1c4
commit 4424cc9887
8 changed files with 1472 additions and 10 deletions
+3
View File
@@ -23,3 +23,6 @@ src/main.tex
# Python bytecode, and per-session scratch work (backups, probes, OCR re-scan).
__pycache__/
.scratch/
# The reflowable edition, like the PDF, is remade by `make epub`.
ross-1988-retypeset.epub
+5 -1
View File
@@ -11,7 +11,7 @@ SOURCES := build.sh tools/polish.py tools/gen_seal.py tools/gen_cover_mosaic.py
src/errata.tex $(wildcard src/pages/*.tex)
.DEFAULT_GOAL := pdf
.PHONY: pdf build image prep plate cover check reprocheck verify clean distclean shell help
.PHONY: pdf build image prep plate cover epub check reprocheck verify clean distclean shell help
# The PDF is a build product and is not in git. This remakes it from src/ and
# plates/, neither of which needs the ProQuest scan — only prep and plate do.
@@ -43,6 +43,10 @@ reprocheck: ## every token of src/pages must reach the built PDF
# so make would skip the rebuild and verify would pass against stale output.
verify: check build reprocheck ## the full gate: check, typeset unconditionally, then prove nothing was dropped
epub: ## build the reflowable EPUB (needs the PDF built, for its cover), then audit it
$(RUN) python3 tools/make_epub.py
$(RUN) python3 tools/epub_audit.py
cover: ## rebuild the cover mosaic from plates/ and fonts/
$(RUN) python3 tools/gen_cover_mosaic.py --force
+3 -3
View File
@@ -5,19 +5,19 @@
\begingroup\small\setstretch{1.2}\setlength{\parindent}{0pt}\setlength{\parskip}{0.6em}
This is a re-typeset, searchable edition of Fiona G.\,E. Ross, \emph{The Evolution of the Printed Bengali Character from 1778 to 1978} (Ph.D. thesis, School of Oriental and African Studies, University of London, 1988), made in 2026 from the ProQuest scan of 431 leaves, ProQuest number 10731406. Every page was transcribed from the page image; the scan's OCR text layer was not used as a source.
The text is reflowed and the original pagination kept: a number in the outer margin marks where each page of the 1988 thesis begins, and the running head gives the page range, so that the Contents, the List of Plates and the author's own cross-references still refer to the original numbering. The footnotes are set as endnotes, grouped by chapter and keeping their numbers.
The text is reflowed and the original pagination kept: \byedition{a number in the outer margin marks where each page of the 1988 thesis begins, and the running head gives the page range}{a small number at the right-hand end of the line marks where each page of the 1988 thesis begins, and the same numbers make up the reader's page list}, so that the Contents, the List of Plates and the author's own cross-references still refer to the original numbering. The footnotes are set as endnotes, grouped by chapter and keeping their numbers.
The 178 plates are reproduced from the scan at its own 300 dpi, cropped clear of the page number, caption and scanner margins, with the captions re-set. The type specimens the author sets into the run of her own sentences are likewise cut from the scan rather than retyped, since it is their letterforms that the argument concerns.
Spelling and punctuation stand as printed, inconsistencies included. The author's own slips are corrected and listed in the Errata; slips inside quoted matter are left as printed, since they may belong to the source quoted. Corrections she made by hand in the scanned copy are adopted; the library ownership stamps are omitted.
Set by XeLaTeX in XCharter, an extension of Matthew Carter's Charter, with Liberation Sans for the margin marks and running heads and Tiro Bangla for Bengali. The cover is set in EB Garamond, its opening line in XCharter. The mark below is generated from a hash of this edition's transcribed text and changes whenever that text does: \texttt{\sealseed}.
\byedition{Set by XeLaTeX in XCharter, an extension of Matthew Carter's Charter, with Liberation Sans for the margin marks and running heads and Tiro Bangla for Bengali.}{The text face is the reader's to choose; Bengali is set in Tiro Bangla, which is embedded.} The cover is set in EB Garamond, its opening line in XCharter. The mark below is generated from a hash of this edition's transcribed text and changes whenever that text does: \texttt{\sealseed}.
\begin{center}
\begin{tikzpicture}[scale=0.62]\sealbody\end{tikzpicture}
\end{center}
The ProQuest notice that precedes the title page in the scan is given overleaf.\par
The ProQuest notice that precedes the title page in the scan is given \byedition{overleaf}{below}.\par
\endgroup
+7
View File
@@ -1,4 +1,11 @@
% Errata: the original's own slips, corrected in this edition.
% The introductory note lives here, not in \printerrata, so that the PDF and the
% EPUB set the same words from one place.
{\small Slips in the original typescript that have been corrected in this
edition. The page number is the original's and links to the passage; the
reading as typed is given first. Corrections the author made by hand in the
scanned copy are adopted silently and are not listed here.\par}\medskip
% One \erratumline{original page}{as printed}{corrected} per \erratum{}{} in
% src/pages/; tools/check.py enforces the correspondence. Keep in page order.
\erratumline{10}{Navarnārī}{Navanārī}
+5 -4
View File
@@ -282,6 +282,11 @@
% src/errata.tex (enforced by tools/check.py), which \printerrata sets out.
\newwrite\erratafile
\newcommand{\erratum}[2]{#1\write\erratafile{#2 -> #1}}
% \byedition{PDF wording}{EPUB wording}: a sentence of the shared front or back
% matter that is only true of one edition -- the colophon's "running head", its
% typefaces, "overleaf". LaTeX sets the first; tools/make_epub.py the second. One
% source, so the sentences both editions share cannot drift apart.
\newcommand{\byedition}[2]{#1}
\newcommand{\erratumline}[3]{\noindent\makebox[0.6in][l]{p.~\pg{#1}}%
\parbox[t]{\dimexpr\textwidth-0.6in\relax}{\raggedright
reads `#2'; corrected here to `#3'}\par\smallskip}
@@ -290,10 +295,6 @@
\fancyhead[R]{\small Errata}\fancyfoot[C]{\small\thepage}}
\newcommand{\printerrata}{\clearpage\pagestyle{errata}\section*{Errata}%
\pdfbookmark[0]{Errata}{sec:errata}%
{\small Slips in the original typescript that have been corrected in this
edition. The page number is the original's and links to the passage; the
reading as typed is given first. Corrections the author made by hand in the
scanned copy are adopted silently and are not listed here.\par}\medskip
\input{src/errata.tex}}
% ---- plates -----------------------------------------------------------------
+162
View File
@@ -0,0 +1,162 @@
#!/usr/bin/env python3
r"""Audit the EPUB by CLASS of fault, not by instance. `make epub` runs it and
it exits non-zero on any failure.
Each check targets a class of fault the rendering sweep of 2026-09-26 found at
least one instance of. They are written against the class, so a new instance
of an old fault fails here even if it looks nothing like the first one:
residue any LaTeX syntax reaching the reader -- a backslash, a brace, or
a bracketed length. (First instance: "\\[0.6em]" printed as text.)
escaping generated markup re-escaped into visible text. (First: <div>,
<tr> showing as text after a second conversion pass.)
empty a block element holding nothing, or only page markers -- a stray
blank line. (First: 31 empty <p> around page breaks.)
links a fragment link resolving to no id in the book. (First: 199
Contents and Plates links pointing into the wrong file.)
images an <img> without width and height, or with a missing file.
xml any document that is not well-formed.
content any token of the transcription that does not reach the EPUB --
the class that holds silently dropped macros, dropped glyphs and
invisible page numbers alike. Runs the project's reprocheck.
typography the EPUB's typographic pass drifting from the PDF's: every range
en dash, tie, cross-reference link and unbreakable specimen that
tools/polish.py produces must reach the EPUB.
Empty table cells are deliberately NOT a fault: the Scheme of Transliteration's
last row is half-filled in the source itself.
"""
import glob
import html
import os
import posixpath
import re
import subprocess
import sys
import xml.dom.minidom as md
import zipfile
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
os.chdir(ROOT)
sys.path.insert(0, os.path.join(ROOT, 'tools'))
EPUB = sys.argv[1] if len(sys.argv) > 1 else 'ross-1988-retypeset.epub'
z = zipfile.ZipFile(EPUB)
X = {n: z.read(n).decode() for n in sorted(z.namelist()) if n.endswith('.xhtml')}
# every document a reader reads -- not the navigation, not the cover image page
CH = {n: d for n, d in X.items() if not re.search(r'/(nav|cover)\.xhtml$', n)}
def visible(d):
d = re.sub(r'<head>.*?</head>', '', d, flags=re.S)
return html.unescape(re.sub(r'<[^>]+>', ' ', d)) # tags -> space: cells never merge
def scan(pattern, docs, use_visible=True):
hits = []
for n, d in docs.items():
t = visible(d) if use_visible else d
for m in re.finditer(pattern, t):
ctx = t[max(0, m.start() - 40):m.end() + 40].replace('\n', ' ')
hits.append('%s: ...%s...' % (n.split('/')[-1], ctx))
return hits
results = []
def check(cls, what, hits):
results.append((cls, what, hits))
# ---- residue: LaTeX syntax in what the reader sees
check('residue', 'backslash', scan(r'\\', CH))
check('residue', 'brace', scan(r'[{}]', CH))
check('residue', 'bracketed length', scan(
r'\[\s*-?\d*\.?\d+\s*(em|ex|pt|pc|in|mm|cm|bp|sp)\s*\](\{[^}]*\})?', CH))
# ---- escaping: generated markup turned into text
check('escaping', 'escaped tag', scan(r'&lt;/?[a-zA-Z]', X, use_visible=False))
check('escaping', 'unrestored placeholder', scan(r'@@BLOCK\d+@@', X, use_visible=False))
# ---- empty: a block with nothing, or only page markers, in it
MARK = r'<span epub:type="pagebreak"[^>]*>[^<]*</span>'
check('empty', 'empty or marker-only block', scan(
r'<(p|div|blockquote|section|li|figcaption)(\s[^>]*)?>(\s|%s)*</\1>' % MARK,
CH, use_visible=False))
# ---- links
ids = {n: set(re.findall(r'\bid="([^"]+)"', d)) for n, d in X.items()}
bad = []
for n, d in X.items():
for h in re.findall(r'href="([^"]+)"', d):
if h.startswith(('http:', 'https:', 'mailto:')) or h.endswith('.css'):
continue
p, _, f = h.partition('#')
t = posixpath.normpath(posixpath.join(posixpath.dirname(n), p)) if p else n
if t not in X or (f and f not in ids[t]):
bad.append('%s -> %s' % (n.split('/')[-1], h))
check('links', 'fragment resolving nowhere', bad)
# ---- images
names = set(z.namelist())
check('images', 'img without width+height',
['%s: %s' % (n.split('/')[-1], t[:90]) for n, d in X.items()
for t in re.findall(r'<img [^>]*>', d) if not ('width=' in t and 'height=' in t)])
check('images', 'img file missing',
['%s: %s' % (n.split('/')[-1], s) for n, d in X.items()
for s in re.findall(r'src="([^"]+)"', d)
if posixpath.normpath(posixpath.join(posixpath.dirname(n), s)) not in names])
# ---- xml
mal = []
for n in z.namelist():
if n.endswith(('.xhtml', '.opf', '.xml', '.ncx')):
try:
md.parseString(z.read(n))
except Exception as e: # noqa: BLE001
mal.append('%s: %s' % (n, str(e)[:70]))
check('xml', 'not well-formed', mal)
# ---- content: every source token must reach the EPUB (the project's own gate)
r = subprocess.run([sys.executable, 'tools/reprocheck.py', EPUB],
capture_output=True, text=True)
first = r.stdout.splitlines()[0] if r.stdout else r.stderr.strip()
m = re.search(r'(\d+) tokens short', first)
check('content', 'tokens lost (reprocheck)',
[] if m and m.group(1) == '0' else [first] + r.stdout.splitlines()[1:6])
# ---- typography: the EPUB carries every transform polish.py makes
import polish # noqa: E402
# Expected: everything the PDF sets, the way build.sh sets it -- the pages
# through polish.py, the colophon and errata raw. (The colophon's ProQuest
# address carries an en dash of its own.)
pol = re.sub(r'(?m)(?<!\\)%.*$', '',
''.join(polish.polish(open(f).read())
for f in sorted(glob.glob('src/pages/p*.tex')))
+ open('src/colophon.tex').read() + open('src/errata.tex').read())
body = ''.join(CH.values())
text = ''.join(visible(d) for d in CH.values())
want = {
'range en dashes': (len(re.findall(r'(?<!-)--(?!-)', pol)), text.count('–')),
'cross-reference links': (pol.count('\\xref{'), body.count('class="xref"')),
'unbreakable specimens': (pol.count('\\mbox{'), body.count('class="nowrap"')),
}
drift = ['%s: polish.py makes %d, EPUB has %d' % (k, a, b)
for k, (a, b) in want.items() if a != b]
ties = pol.count('~')
if text.count(' ') < ties:
drift.append('ties: polish.py makes %d, EPUB has %d no-break spaces'
% (ties, text.count(' ')))
check('typography', 'drift from the PDF pass', drift)
# ---- report
w = max(len(c) + len(t) for c, t, _ in results) + 3
failed = 0
for cls, what, hits in results:
ok = not hits
failed += not ok
print('[%s] %-*s %4d%s' % (' ok ' if ok else 'FAIL', w, cls + ': ' + what, len(hits),
'' if ok else ' e.g. ' + hits[0][:150]))
print('\n%d failing check(s)' % failed)
sys.exit(1 if failed else 0)
+1244
View File
File diff suppressed because it is too large Load Diff
+43 -2
View File
@@ -9,6 +9,7 @@ is by bag, not sequence, because endnotes and the Contents move text about.
Found the biblist \hangafter=1 digit-swallowing bug (2026-09-14).
python3 tools/reprocheck.py [ross-1988-retypeset.pdf]
python3 tools/reprocheck.py ross-1988-retypeset.epub
"""
import re, sys, glob, subprocess, unicodedata
from collections import Counter
@@ -31,6 +32,9 @@ DROP = {"ig", "fig", "includegraphics", "label", "index", "vspace", "hspace",
"thispagestyle", "pagestyle", "setstretch", "addmargin", "addvspace",
"vskip", "hskip", "rule", "makebox", "raisebox", "resizebox", "XeTeXglyph"}
# environment -> (optional, mandatory) parameters that follow \begin{env}
ENV_PARAMS = {"addmargin": (1, 1), "tabular": (0, 1), "minipage": (1, 1)}
def argspans(s, i):
"""yield (start, end) of consecutive brace groups beginning at i"""
out = []
@@ -80,6 +84,22 @@ def detex(s):
spans, after = argspans(s, j)
if name in ("begin", "end"): # env name and column spec are not copy
spans = []
# An environment's own parameters follow its name, after the
# {env} group -- \begin{addmargin}[0.47in]{0.5in}. They are layout,
# so skip exactly what each environment declares; left in, they
# were counted as copy ("0", "47", "in") that no edition prints.
env = re.match(r"\s*\{(\w+)\}", s[j:])
if name == "begin" and env and env.group(1) in ENV_PARAMS:
after = j + env.end()
nopt, nman = ENV_PARAMS[env.group(1)]
for _ in range(nopt):
k = re.match(r"\s*\[[^\]]*\]", s[after:])
if k:
after += k.end()
if nman:
_sp, after = argspans(s, after)
if len(_sp) > nman: # keep only the declared count
after = _sp[nman - 1][1] + 1
if name == "XeTeXglyph": # \XeTeXglyph 824 - unbraced slot id
k = re.match(r"\s*[0-9]+", s[after:])
if k:
@@ -115,8 +135,29 @@ for f in files:
src.append(detex(open(f, encoding="utf-8").read()))
A = tokens("\n".join(src))
pdftxt = subprocess.run(["pdftotext", "-layout", PDF, "-"],
capture_output=True, text=True).stdout
def epub_text(path):
"""Visible text of the EPUB's chapter documents. Tags become spaces, so
table cells and adjacent elements never merge into one token. The
navigation document is left out on purpose: it repeats every heading, and
counting it would hide a heading lost from the chapter itself."""
import zipfile, html as _html
z = zipfile.ZipFile(path)
parts = []
for n in sorted(z.namelist()):
# the chapters, and the one notes document the endnotes now live in;
# not the colophon or errata, which are not transcription
if re.search(r"/(ch\d+|notes)\.xhtml$", n):
d = re.sub(r"<head>.*?</head>", "", z.read(n).decode(), flags=re.S)
parts.append(_html.unescape(re.sub(r"<[^>]+>", " ", d)))
return "\n".join(parts)
# The same check serves both editions: every token of the transcription must
# reach the output, whichever output it is.
if PDF.endswith(".epub"):
pdftxt = epub_text(PDF)
else:
pdftxt = subprocess.run(["pdftotext", "-layout", PDF, "-"],
capture_output=True, text=True).stdout
B = tokens(pdftxt)
where = {}