Add the EPUB edition: converter, class-based audit, shared front matter
tools/make_epub.py converts the same src/ the PDF is set from, running
polish.py first, with original page numbers as page-list metadata, endnotes
gathered in one linked Notes section, and the colophon and errata included.
It raises on any macro it does not declare. tools/epub_audit.py checks by
class of fault (LaTeX residue, escaping, empty blocks, links, images, XML,
content, typography drift from the PDF); make epub runs both.
\byedition{PDF}{EPUB} lets the colophon carry the sentences that are true of
only one edition; the errata introduction moves into src/errata.tex so both
editions print one copy. reprocheck.py now accepts an .epub and skips
environment parameters that are layout, not copy.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -23,3 +23,6 @@ src/main.tex
|
||||
# Python bytecode, and per-session scratch work (backups, probes, OCR re-scan).
|
||||
__pycache__/
|
||||
.scratch/
|
||||
|
||||
# The reflowable edition, like the PDF, is remade by `make epub`.
|
||||
ross-1988-retypeset.epub
|
||||
|
||||
@@ -11,7 +11,7 @@ SOURCES := build.sh tools/polish.py tools/gen_seal.py tools/gen_cover_mosaic.py
|
||||
src/errata.tex $(wildcard src/pages/*.tex)
|
||||
|
||||
.DEFAULT_GOAL := pdf
|
||||
.PHONY: pdf build image prep plate cover check reprocheck verify clean distclean shell help
|
||||
.PHONY: pdf build image prep plate cover epub check reprocheck verify clean distclean shell help
|
||||
|
||||
# The PDF is a build product and is not in git. This remakes it from src/ and
|
||||
# plates/, neither of which needs the ProQuest scan — only prep and plate do.
|
||||
@@ -43,6 +43,10 @@ reprocheck: ## every token of src/pages must reach the built PDF
|
||||
# so make would skip the rebuild and verify would pass against stale output.
|
||||
verify: check build reprocheck ## the full gate: check, typeset unconditionally, then prove nothing was dropped
|
||||
|
||||
epub: ## build the reflowable EPUB (needs the PDF built, for its cover), then audit it
|
||||
$(RUN) python3 tools/make_epub.py
|
||||
$(RUN) python3 tools/epub_audit.py
|
||||
|
||||
cover: ## rebuild the cover mosaic from plates/ and fonts/
|
||||
$(RUN) python3 tools/gen_cover_mosaic.py --force
|
||||
|
||||
|
||||
+3
-3
@@ -5,19 +5,19 @@
|
||||
\begingroup\small\setstretch{1.2}\setlength{\parindent}{0pt}\setlength{\parskip}{0.6em}
|
||||
This is a re-typeset, searchable edition of Fiona G.\,E. Ross, \emph{The Evolution of the Printed Bengali Character from 1778 to 1978} (Ph.D. thesis, School of Oriental and African Studies, University of London, 1988), made in 2026 from the ProQuest scan of 431 leaves, ProQuest number 10731406. Every page was transcribed from the page image; the scan's OCR text layer was not used as a source.
|
||||
|
||||
The text is reflowed and the original pagination kept: a number in the outer margin marks where each page of the 1988 thesis begins, and the running head gives the page range, so that the Contents, the List of Plates and the author's own cross-references still refer to the original numbering. The footnotes are set as endnotes, grouped by chapter and keeping their numbers.
|
||||
The text is reflowed and the original pagination kept: \byedition{a number in the outer margin marks where each page of the 1988 thesis begins, and the running head gives the page range}{a small number at the right-hand end of the line marks where each page of the 1988 thesis begins, and the same numbers make up the reader's page list}, so that the Contents, the List of Plates and the author's own cross-references still refer to the original numbering. The footnotes are set as endnotes, grouped by chapter and keeping their numbers.
|
||||
|
||||
The 178 plates are reproduced from the scan at its own 300 dpi, cropped clear of the page number, caption and scanner margins, with the captions re-set. The type specimens the author sets into the run of her own sentences are likewise cut from the scan rather than retyped, since it is their letterforms that the argument concerns.
|
||||
|
||||
Spelling and punctuation stand as printed, inconsistencies included. The author's own slips are corrected and listed in the Errata; slips inside quoted matter are left as printed, since they may belong to the source quoted. Corrections she made by hand in the scanned copy are adopted; the library ownership stamps are omitted.
|
||||
|
||||
Set by XeLaTeX in XCharter, an extension of Matthew Carter's Charter, with Liberation Sans for the margin marks and running heads and Tiro Bangla for Bengali. The cover is set in EB Garamond, its opening line in XCharter. The mark below is generated from a hash of this edition's transcribed text and changes whenever that text does: \texttt{\sealseed}.
|
||||
\byedition{Set by XeLaTeX in XCharter, an extension of Matthew Carter's Charter, with Liberation Sans for the margin marks and running heads and Tiro Bangla for Bengali.}{The text face is the reader's to choose; Bengali is set in Tiro Bangla, which is embedded.} The cover is set in EB Garamond, its opening line in XCharter. The mark below is generated from a hash of this edition's transcribed text and changes whenever that text does: \texttt{\sealseed}.
|
||||
|
||||
\begin{center}
|
||||
\begin{tikzpicture}[scale=0.62]\sealbody\end{tikzpicture}
|
||||
\end{center}
|
||||
|
||||
The ProQuest notice that precedes the title page in the scan is given overleaf.\par
|
||||
The ProQuest notice that precedes the title page in the scan is given \byedition{overleaf}{below}.\par
|
||||
|
||||
\endgroup
|
||||
|
||||
|
||||
@@ -1,4 +1,11 @@
|
||||
% Errata: the original's own slips, corrected in this edition.
|
||||
% The introductory note lives here, not in \printerrata, so that the PDF and the
|
||||
% EPUB set the same words from one place.
|
||||
{\small Slips in the original typescript that have been corrected in this
|
||||
edition. The page number is the original's and links to the passage; the
|
||||
reading as typed is given first. Corrections the author made by hand in the
|
||||
scanned copy are adopted silently and are not listed here.\par}\medskip
|
||||
|
||||
% One \erratumline{original page}{as printed}{corrected} per \erratum{}{} in
|
||||
% src/pages/; tools/check.py enforces the correspondence. Keep in page order.
|
||||
\erratumline{10}{Navarnārī}{Navanārī}
|
||||
|
||||
+5
-4
@@ -282,6 +282,11 @@
|
||||
% src/errata.tex (enforced by tools/check.py), which \printerrata sets out.
|
||||
\newwrite\erratafile
|
||||
\newcommand{\erratum}[2]{#1\write\erratafile{#2 -> #1}}
|
||||
% \byedition{PDF wording}{EPUB wording}: a sentence of the shared front or back
|
||||
% matter that is only true of one edition -- the colophon's "running head", its
|
||||
% typefaces, "overleaf". LaTeX sets the first; tools/make_epub.py the second. One
|
||||
% source, so the sentences both editions share cannot drift apart.
|
||||
\newcommand{\byedition}[2]{#1}
|
||||
\newcommand{\erratumline}[3]{\noindent\makebox[0.6in][l]{p.~\pg{#1}}%
|
||||
\parbox[t]{\dimexpr\textwidth-0.6in\relax}{\raggedright
|
||||
reads `#2'; corrected here to `#3'}\par\smallskip}
|
||||
@@ -290,10 +295,6 @@
|
||||
\fancyhead[R]{\small Errata}\fancyfoot[C]{\small\thepage}}
|
||||
\newcommand{\printerrata}{\clearpage\pagestyle{errata}\section*{Errata}%
|
||||
\pdfbookmark[0]{Errata}{sec:errata}%
|
||||
{\small Slips in the original typescript that have been corrected in this
|
||||
edition. The page number is the original's and links to the passage; the
|
||||
reading as typed is given first. Corrections the author made by hand in the
|
||||
scanned copy are adopted silently and are not listed here.\par}\medskip
|
||||
\input{src/errata.tex}}
|
||||
|
||||
% ---- plates -----------------------------------------------------------------
|
||||
|
||||
@@ -0,0 +1,162 @@
|
||||
#!/usr/bin/env python3
|
||||
r"""Audit the EPUB by CLASS of fault, not by instance. `make epub` runs it and
|
||||
it exits non-zero on any failure.
|
||||
|
||||
Each check targets a class of fault the rendering sweep of 2026-09-26 found at
|
||||
least one instance of. They are written against the class, so a new instance
|
||||
of an old fault fails here even if it looks nothing like the first one:
|
||||
|
||||
residue any LaTeX syntax reaching the reader -- a backslash, a brace, or
|
||||
a bracketed length. (First instance: "\\[0.6em]" printed as text.)
|
||||
escaping generated markup re-escaped into visible text. (First: <div>,
|
||||
<tr> showing as text after a second conversion pass.)
|
||||
empty a block element holding nothing, or only page markers -- a stray
|
||||
blank line. (First: 31 empty <p> around page breaks.)
|
||||
links a fragment link resolving to no id in the book. (First: 199
|
||||
Contents and Plates links pointing into the wrong file.)
|
||||
images an <img> without width and height, or with a missing file.
|
||||
xml any document that is not well-formed.
|
||||
content any token of the transcription that does not reach the EPUB --
|
||||
the class that holds silently dropped macros, dropped glyphs and
|
||||
invisible page numbers alike. Runs the project's reprocheck.
|
||||
typography the EPUB's typographic pass drifting from the PDF's: every range
|
||||
en dash, tie, cross-reference link and unbreakable specimen that
|
||||
tools/polish.py produces must reach the EPUB.
|
||||
|
||||
Empty table cells are deliberately NOT a fault: the Scheme of Transliteration's
|
||||
last row is half-filled in the source itself.
|
||||
"""
|
||||
import glob
|
||||
import html
|
||||
import os
|
||||
import posixpath
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
import xml.dom.minidom as md
|
||||
import zipfile
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
os.chdir(ROOT)
|
||||
sys.path.insert(0, os.path.join(ROOT, 'tools'))
|
||||
EPUB = sys.argv[1] if len(sys.argv) > 1 else 'ross-1988-retypeset.epub'
|
||||
|
||||
z = zipfile.ZipFile(EPUB)
|
||||
X = {n: z.read(n).decode() for n in sorted(z.namelist()) if n.endswith('.xhtml')}
|
||||
# every document a reader reads -- not the navigation, not the cover image page
|
||||
CH = {n: d for n, d in X.items() if not re.search(r'/(nav|cover)\.xhtml$', n)}
|
||||
|
||||
|
||||
def visible(d):
|
||||
d = re.sub(r'<head>.*?</head>', '', d, flags=re.S)
|
||||
return html.unescape(re.sub(r'<[^>]+>', ' ', d)) # tags -> space: cells never merge
|
||||
|
||||
|
||||
def scan(pattern, docs, use_visible=True):
|
||||
hits = []
|
||||
for n, d in docs.items():
|
||||
t = visible(d) if use_visible else d
|
||||
for m in re.finditer(pattern, t):
|
||||
ctx = t[max(0, m.start() - 40):m.end() + 40].replace('\n', ' ')
|
||||
hits.append('%s: ...%s...' % (n.split('/')[-1], ctx))
|
||||
return hits
|
||||
|
||||
|
||||
results = []
|
||||
|
||||
|
||||
def check(cls, what, hits):
|
||||
results.append((cls, what, hits))
|
||||
|
||||
|
||||
# ---- residue: LaTeX syntax in what the reader sees
|
||||
check('residue', 'backslash', scan(r'\\', CH))
|
||||
check('residue', 'brace', scan(r'[{}]', CH))
|
||||
check('residue', 'bracketed length', scan(
|
||||
r'\[\s*-?\d*\.?\d+\s*(em|ex|pt|pc|in|mm|cm|bp|sp)\s*\](\{[^}]*\})?', CH))
|
||||
|
||||
# ---- escaping: generated markup turned into text
|
||||
check('escaping', 'escaped tag', scan(r'</?[a-zA-Z]', X, use_visible=False))
|
||||
check('escaping', 'unrestored placeholder', scan(r'@@BLOCK\d+@@', X, use_visible=False))
|
||||
|
||||
# ---- empty: a block with nothing, or only page markers, in it
|
||||
MARK = r'<span epub:type="pagebreak"[^>]*>[^<]*</span>'
|
||||
check('empty', 'empty or marker-only block', scan(
|
||||
r'<(p|div|blockquote|section|li|figcaption)(\s[^>]*)?>(\s|%s)*</\1>' % MARK,
|
||||
CH, use_visible=False))
|
||||
|
||||
# ---- links
|
||||
ids = {n: set(re.findall(r'\bid="([^"]+)"', d)) for n, d in X.items()}
|
||||
bad = []
|
||||
for n, d in X.items():
|
||||
for h in re.findall(r'href="([^"]+)"', d):
|
||||
if h.startswith(('http:', 'https:', 'mailto:')) or h.endswith('.css'):
|
||||
continue
|
||||
p, _, f = h.partition('#')
|
||||
t = posixpath.normpath(posixpath.join(posixpath.dirname(n), p)) if p else n
|
||||
if t not in X or (f and f not in ids[t]):
|
||||
bad.append('%s -> %s' % (n.split('/')[-1], h))
|
||||
check('links', 'fragment resolving nowhere', bad)
|
||||
|
||||
# ---- images
|
||||
names = set(z.namelist())
|
||||
check('images', 'img without width+height',
|
||||
['%s: %s' % (n.split('/')[-1], t[:90]) for n, d in X.items()
|
||||
for t in re.findall(r'<img [^>]*>', d) if not ('width=' in t and 'height=' in t)])
|
||||
check('images', 'img file missing',
|
||||
['%s: %s' % (n.split('/')[-1], s) for n, d in X.items()
|
||||
for s in re.findall(r'src="([^"]+)"', d)
|
||||
if posixpath.normpath(posixpath.join(posixpath.dirname(n), s)) not in names])
|
||||
|
||||
# ---- xml
|
||||
mal = []
|
||||
for n in z.namelist():
|
||||
if n.endswith(('.xhtml', '.opf', '.xml', '.ncx')):
|
||||
try:
|
||||
md.parseString(z.read(n))
|
||||
except Exception as e: # noqa: BLE001
|
||||
mal.append('%s: %s' % (n, str(e)[:70]))
|
||||
check('xml', 'not well-formed', mal)
|
||||
|
||||
# ---- content: every source token must reach the EPUB (the project's own gate)
|
||||
r = subprocess.run([sys.executable, 'tools/reprocheck.py', EPUB],
|
||||
capture_output=True, text=True)
|
||||
first = r.stdout.splitlines()[0] if r.stdout else r.stderr.strip()
|
||||
m = re.search(r'(\d+) tokens short', first)
|
||||
check('content', 'tokens lost (reprocheck)',
|
||||
[] if m and m.group(1) == '0' else [first] + r.stdout.splitlines()[1:6])
|
||||
|
||||
# ---- typography: the EPUB carries every transform polish.py makes
|
||||
import polish # noqa: E402
|
||||
# Expected: everything the PDF sets, the way build.sh sets it -- the pages
|
||||
# through polish.py, the colophon and errata raw. (The colophon's ProQuest
|
||||
# address carries an en dash of its own.)
|
||||
pol = re.sub(r'(?m)(?<!\\)%.*$', '',
|
||||
''.join(polish.polish(open(f).read())
|
||||
for f in sorted(glob.glob('src/pages/p*.tex')))
|
||||
+ open('src/colophon.tex').read() + open('src/errata.tex').read())
|
||||
body = ''.join(CH.values())
|
||||
text = ''.join(visible(d) for d in CH.values())
|
||||
want = {
|
||||
'range en dashes': (len(re.findall(r'(?<!-)--(?!-)', pol)), text.count('–')),
|
||||
'cross-reference links': (pol.count('\\xref{'), body.count('class="xref"')),
|
||||
'unbreakable specimens': (pol.count('\\mbox{'), body.count('class="nowrap"')),
|
||||
}
|
||||
drift = ['%s: polish.py makes %d, EPUB has %d' % (k, a, b)
|
||||
for k, (a, b) in want.items() if a != b]
|
||||
ties = pol.count('~')
|
||||
if text.count(' ') < ties:
|
||||
drift.append('ties: polish.py makes %d, EPUB has %d no-break spaces'
|
||||
% (ties, text.count(' ')))
|
||||
check('typography', 'drift from the PDF pass', drift)
|
||||
|
||||
# ---- report
|
||||
w = max(len(c) + len(t) for c, t, _ in results) + 3
|
||||
failed = 0
|
||||
for cls, what, hits in results:
|
||||
ok = not hits
|
||||
failed += not ok
|
||||
print('[%s] %-*s %4d%s' % (' ok ' if ok else 'FAIL', w, cls + ': ' + what, len(hits),
|
||||
'' if ok else ' e.g. ' + hits[0][:150]))
|
||||
print('\n%d failing check(s)' % failed)
|
||||
sys.exit(1 if failed else 0)
|
||||
+1244
File diff suppressed because it is too large
Load Diff
+43
-2
@@ -9,6 +9,7 @@ is by bag, not sequence, because endnotes and the Contents move text about.
|
||||
Found the biblist \hangafter=1 digit-swallowing bug (2026-09-14).
|
||||
|
||||
python3 tools/reprocheck.py [ross-1988-retypeset.pdf]
|
||||
python3 tools/reprocheck.py ross-1988-retypeset.epub
|
||||
"""
|
||||
import re, sys, glob, subprocess, unicodedata
|
||||
from collections import Counter
|
||||
@@ -31,6 +32,9 @@ DROP = {"ig", "fig", "includegraphics", "label", "index", "vspace", "hspace",
|
||||
"thispagestyle", "pagestyle", "setstretch", "addmargin", "addvspace",
|
||||
"vskip", "hskip", "rule", "makebox", "raisebox", "resizebox", "XeTeXglyph"}
|
||||
|
||||
# environment -> (optional, mandatory) parameters that follow \begin{env}
|
||||
ENV_PARAMS = {"addmargin": (1, 1), "tabular": (0, 1), "minipage": (1, 1)}
|
||||
|
||||
def argspans(s, i):
|
||||
"""yield (start, end) of consecutive brace groups beginning at i"""
|
||||
out = []
|
||||
@@ -80,6 +84,22 @@ def detex(s):
|
||||
spans, after = argspans(s, j)
|
||||
if name in ("begin", "end"): # env name and column spec are not copy
|
||||
spans = []
|
||||
# An environment's own parameters follow its name, after the
|
||||
# {env} group -- \begin{addmargin}[0.47in]{0.5in}. They are layout,
|
||||
# so skip exactly what each environment declares; left in, they
|
||||
# were counted as copy ("0", "47", "in") that no edition prints.
|
||||
env = re.match(r"\s*\{(\w+)\}", s[j:])
|
||||
if name == "begin" and env and env.group(1) in ENV_PARAMS:
|
||||
after = j + env.end()
|
||||
nopt, nman = ENV_PARAMS[env.group(1)]
|
||||
for _ in range(nopt):
|
||||
k = re.match(r"\s*\[[^\]]*\]", s[after:])
|
||||
if k:
|
||||
after += k.end()
|
||||
if nman:
|
||||
_sp, after = argspans(s, after)
|
||||
if len(_sp) > nman: # keep only the declared count
|
||||
after = _sp[nman - 1][1] + 1
|
||||
if name == "XeTeXglyph": # \XeTeXglyph 824 - unbraced slot id
|
||||
k = re.match(r"\s*[0-9]+", s[after:])
|
||||
if k:
|
||||
@@ -115,8 +135,29 @@ for f in files:
|
||||
src.append(detex(open(f, encoding="utf-8").read()))
|
||||
A = tokens("\n".join(src))
|
||||
|
||||
pdftxt = subprocess.run(["pdftotext", "-layout", PDF, "-"],
|
||||
capture_output=True, text=True).stdout
|
||||
def epub_text(path):
|
||||
"""Visible text of the EPUB's chapter documents. Tags become spaces, so
|
||||
table cells and adjacent elements never merge into one token. The
|
||||
navigation document is left out on purpose: it repeats every heading, and
|
||||
counting it would hide a heading lost from the chapter itself."""
|
||||
import zipfile, html as _html
|
||||
z = zipfile.ZipFile(path)
|
||||
parts = []
|
||||
for n in sorted(z.namelist()):
|
||||
# the chapters, and the one notes document the endnotes now live in;
|
||||
# not the colophon or errata, which are not transcription
|
||||
if re.search(r"/(ch\d+|notes)\.xhtml$", n):
|
||||
d = re.sub(r"<head>.*?</head>", "", z.read(n).decode(), flags=re.S)
|
||||
parts.append(_html.unescape(re.sub(r"<[^>]+>", " ", d)))
|
||||
return "\n".join(parts)
|
||||
|
||||
# The same check serves both editions: every token of the transcription must
|
||||
# reach the output, whichever output it is.
|
||||
if PDF.endswith(".epub"):
|
||||
pdftxt = epub_text(PDF)
|
||||
else:
|
||||
pdftxt = subprocess.run(["pdftotext", "-layout", PDF, "-"],
|
||||
capture_output=True, text=True).stdout
|
||||
B = tokens(pdftxt)
|
||||
|
||||
where = {}
|
||||
|
||||
Reference in New Issue
Block a user