tools/make_epub.py converts the same src/ the PDF is set from, running
polish.py first, with original page numbers as page-list metadata, endnotes
gathered in one linked Notes section, and the colophon and errata included.
It raises on any macro it does not declare. tools/epub_audit.py checks by
class of fault (LaTeX residue, escaping, empty blocks, links, images, XML,
content, typography drift from the PDF); make epub runs both.
\byedition{PDF}{EPUB} lets the colophon carry the sentences that are true of
only one edition; the errata introduction moves into src/errata.tex so both
editions print one copy. reprocheck.py now accepts an .epub and skips
environment parameters that are layout, not copy.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The cover is a full-bleed mosaic of historical Bengali types filling the word
ব্যঞ্জণ, produced by tools/gen_cover_mosaic.py from plates/ and fonts/ alone
(build.sh remakes it if missing). It is static, so ebook readers can cache it.
The mark on the About this edition page is drawn by tools/gen_seal.py, seeded
from a SHA-256 of src/ (never the build's own outputs), so it changes exactly
when the transcription does. The colophon is shortened to name the cover
typefaces and the mark's mechanism, and drops the trailing ProQuest note.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The scan is a bilevel 300 dpi JBIG2 layer plus lower-resolution colour, so
storing every plate as 8-bit PNG only added size. tools/optimize_plates.py
writes copies under work/plates-opt (bilevel to generic-mode, lossless JBIG2;
genuine halftone to JPEG q88) and repoints work/pages at them; plates/ and
src/ are untouched. The Dockerfile builds jbig2enc and adds qpdf, ghostscript
and img2pdf.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two independent OCR engines (tesseract, ABBYY) were compared word for word
with src/pages, then every difference was settled on the page image.
Three silent normalisations of the author's slips become errata (51 -> 54):
precusor (p.79), inaugaration (p.294), and the handwritten marginal note's
twelth (p.58). Two punctuation slips of the transcription are corrected to
the scan: p.56 'design:' (was ';') and p.230 'Gottschall.' (was ',').
QUESTIONS.md logs the twelve handwritten corrections that were adopted but
never recorded, and the scope of the re-scan (diacritics and italics were
not covered).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The line was 78 characters at 12pt, above the 60-75 band — not the "~73"
I had recorded. That figure came from a font-metric estimate (measure
divided by the unweighted mean lowercase advance, counting no spaces) and
was simply wrong. Measured directly off the built PDF — 79 full lines,
glyphs plus inter-word gaps — 12pt ran 78 and 13pt runs 73.
So this is not restoring a lost measure, it is fixing a line that was too
long, and raising the size is the better of the two fixes: narrowing the
measure to 5.7in would have shortened the line while leaving the type
small. (Sammay proposed the size bump.)
Set through fontspec Scale=1.0833 over the 12pt class, deliberately: it
scales every size in the family, including the plate captions, but leaves
\baselineskip class-derived, so \setstretch{1.59} still yields the 22.9pt
line measured off the original. That leading is anchored to the inline
specimens and must not drift with the type size.
Verified after the build:
characters per line 78 -> 73 (target band 60-75)
pages 455 -> 479
underfull boxes 84 -> 71
Bengali/Latin ratio 2.08 -> 2.11 Scale=MatchLowercase is applied
after the main font's Scale, so
Tiro Bangla needed no adjustment
specimens still clear of the lines above and below
overfull 0, font warnings 0, reprocheck 0 tokens short
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CLAUDE.md requires re-auditing the cross-reference rule after any change
to it, and polish.py changed twice this session. tools/linkaudit.py does
it: three classes checked mechanically, the fourth printed to be read.
dangling targets 0
plate page mismatch (listed vs set) 0 (all 178)
chapter page mismatch 0
`p. N` self-references 25 all read, all sound
The finding: an elided range end is not a plate number of its own.
plate_sub ran re.sub(r"\d+") over the whole range, so "pls. 146-8" linked
its 8 to plate 8 (orig. p.35) and "pls. 142-5" linked its 5 to plate 5
(p.23). Three links, all pointing at the wrong plate — invisible to every
other check, since the targets exist and the text is correct. The end is
now expanded against the leading digits of the number it is elided
against: 146-8 -> plate 148 (p.334), 142-5 -> plate 145 (p.329), both
confirmed against the List of Plates and against where \plateop sets
them.
Worth noting what the audit also confirmed: in two footnotes the rule
discriminates a citation page from a self-reference in the same sentence
— "(London, 1962), p. 43; also see below, p. 63" links only the second.
NOT YET BUILT: the Docker daemon is down, so make verify has not run
against this change. polish.py was run on the host and the audit re-run
against its output, but the PDF does not yet carry the fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It still announced itself as end-of-session-4 and had never heard of
XCharter, which matters because CLAUDE.md tells the next session to read
it first. Records the font switch, the small-caps removal, the link
styling, the nine diacritic corrections, and the next four items in the
order I would take them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Includes what remains unchecked and why the remaining yield looks low,
plus the limitation both nets share: they key on inconsistency, so a name
transcribed the same wrong way throughout passes silently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Abbreviations are now plain uppercase throughout (Sammay). The pass
small-capped an explicit whitelist, which could only ever be partial: on
the Abbreviations and Conventions page BFBS, LMS and MLCo stood
unconverted in one column beside small-capped BL, BMS, EIC, IOL, OUP,
SOAS and SPG, and "MS EUR 30" split a single shelfmark across two styles.
Completing the whitelist was the obvious fix and the wrong one. The
original is a typescript — a typewriter cannot set small caps — so the
source defines none, and rule 3 leaves every acronym as the plain
uppercase it prints. smallcaps() is now a documented no-op.
p0431: Rtusamhāra -> Ṛtusaṃhāra, verified at 700 dpi (dot under the R,
dot under the m). It was catchable because it broke the Bibliography's
own pattern: the author gives a title in full transliteration and its
anglicised form plain — "Ānanda Bājāra Patrikā (Ananda Bazar Patrika)" —
so a Bibliography entry with the marks stripped is an anomaly, while most
of the remaining flags there are her distinction rather than our error.
Makefile: verify now depends on `build`, not `pdf`. `pdf` is
timestamp-conditional, and a source edited while a build is running
leaves a PDF newer than the source it does not contain — make then skips
the rebuild and verify passes against stale output. That happened here,
and reprocheck caught it correctly twice while I twice misread it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ScholaX read as too decorative (Sammay), and the font metrics agree: of
the faces considered it sets the smallest lowercase relative to its
capitals — x-height/cap 0.645 against XCharter's 0.717 — so its caps and
ascenders dominate while the lowercase does the reading.
Chosen on measured grounds, not taste alone. All 35 non-ASCII characters
this text uses are covered; the dots-below transliteration set
(ḍ ḥ ṃ ṅ ṇ Ṛ ṛ ṣ ṭ) disqualifies ETbb, Domitian, STEP and Tempora, which
are otherwise reasonable book faces. Real small caps and oldstyle figures
confirmed present, so the acronym setting is unaffected. XCharter ships
with TeX Live, so no font is added. Erewhon was the runner-up until the
metrics showed its x-height is smaller than the current face, which would
have moved readability the wrong way.
Bengali needed no manual adjustment: Scale=MatchLowercase re-derives Tiro
Bangla from the main font's x-height, and the Bengali-to-Latin height
ratio measured off the Scheme of Transliteration held at 2.12 -> 2.08
across the switch, within the ±1px noise of a 300 dpi read.
455 pages, down from 470. Underfull boxes 123 -> 84: XCharter's narrower
set gives the line-breaker more room at the same measure. Zero overfull,
zero font warnings, reprocheck 0 tokens short.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Eight diacritic corrections at seven sites, each verified against the page
image at 900 dpi:
p0049 Nastaʿlīq -> Nasta`līq the mark is a quote, not a ʿayn (Sammay)
p0072 Cálidás -> Cálidās
p0256 Rajākṛṣṇa -> Rājākṛṣṇa
p0260 Īsvaracandra -> Īśvaracandra (x2)
p0430 Iśvaracandra -> Īśvaracandra
p0430 Gītā -> Gīta (x2)
p0430 Munṣī -> Munśī
Found by two nets over the transcription: a rare-character inventory
(U+02BF appeared exactly once in 431 pages — the tell) and a check for
names that appear with more than one accenting, 84 occurrences of which
~32 are now checked. Neither net catches a name mis-transcribed the SAME
way everywhere, so this is a partial sweep, not a clean bill.
The French accents are all handwritten additions in the scan — pen-drawn
cedillas and acutes — applied inconsistently by the author and reproduced
per instance. All of that cluster checks out.
Clickable regions now carry a 0.4pt hairline rule in RGB(95,125,165) on
the Contents/List-of-Plates page numbers and on resolved cross-references;
endnote superscripts and whole-line entries stay unmarked. polish.py emits
\xref for the former and \hyperlink for the latter so the two differ.
Also: plate 1's List-of-Plates entry uses \plx, which build_maps() did not
scan, so plate 1 was the one plate of 178 whose "pl. 1" references never
linked.
build.sh takes an atomic mkdir lock: two builds sharing work/ corrupt
main.aux and produce an error that names the wrong file entirely.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A finite \parfillskip was tried at 0.75\textwidth (no effect) and then,
after deriving the threshold, at 0.15\textwidth (real effect, bad
trade). Badness is 100(G/X)^3 against tolerance 4000, so the threshold
is G ~ 3.4(X + 24pt): on a 430pt measure any X at or above
0.25\textwidth prescribes a threshold longer than the measure and can
never fire. The only values in range are 0.15 and 0.10, and 0.10 bites
harder — so the lever is effectively binary.
At 0.15 it did what it promised and cost more than it bought:
paragraph-final mean 161.7 -> 142.7, p90 319 -> 267, max 398 -> 341
paragraph-internal median 22.8 -> 35.7pt, >3em short 29.4% -> 48.0%
underfull boxes 123 -> 1902
There are far more interior lines than final ones. Back to ragged2e's
default, with the derivation kept in the preamble so this is not
rediscovered by experiment a third time.
Net of all five levers: 5 fixed a real conflict, 2 and 4 are harmless
and near-inert here (the ragged right margin is never reached, and only
1.6% of lines begin with a protrudable character), 1 and 3 were
reverted as measurably harmful. Rag is within ~1% of where it started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
\platefullmeasure resets \parshape for a plate's own paragraph, so a
plate keeps the full text measure even where the original interrupts a
displayed quotation with one. Verified in the built PDF:
plate 50 (p0130) 319.2 x 484.3 pt height-limited, unchanged
plate 51 (p0131) 430.5 x 356.0 pt full measure (was 372.2)
plate 97 (p0218) 430.2 x 274.8 pt full measure (was 372.2)
plate 165 (p0387) 351.1 x 484.5 pt height-limited, unchanged
Line-breaking levers, measured with work/rag.py over 102
paragraph-internal lines on nine text pages:
1. permitted rag 2em -> 4em: made every statistic WORSE (median rag
22.8 -> 24.8pt, sd 79.6 -> 89.1, lines >3em short 29.4 -> 35.3%)
and emptied the log of underfull warnings, which was the goalposts
moving rather than a gain. Reverted; the preamble says not to try
it again without measuring.
2. \linepenalty=30: no measurable change on this sample. Kept.
3. finite \RaggedRightParfillskip: inert as set — 0.75\textwidth is
322.5pt of stretch, past the 90th percentile of last-line rag, so
it never incurs badness. Kept, pending a value that bites.
4. protrusion factor 1100 -> 1500: below the resolution of a
line-ending metric; it changes how even the rag looks, not where
lines break. Kept, unquantified.
5. \emergencystretch was set to 3em and silently overridden to 1em
lower down; the override is gone. No visual effect (zero overfull
either way), but the conflict was real.
Also fixed: \RaggedRight assigns \rightskip, \parfillskip AND \parindent
from its own parameters at every invocation, so setting those directly
does nothing. All three are now set as \RaggedRight* before the call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It is assembled from src/preamble.tex, src/colophon.tex and work/pages
each build, so it churned in the diff and carried no information the
tracked sources do not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four plates overhung the measure and displaced into the right margin.
On those pages the original interrupts a displayed quotation with a
plate (p0129/30/31/32, p0217/18, p0386/87), so \plateop ran inside
extract's \addmargin, where the paragraph is set by \parshape to
\linewidth — while the minipage was built at \textwidth. The quotation
indents 0.5in + 0.3in, so the overhang was exactly 0.8in, which is also
\marginparwidth; that coincidence made it look like a marginnote fault
for a long time. \plateop, \plate and \fig now build at \linewidth,
correct in any context and unchanged on the other 174 plates.
Also: \mdseries on the three \bnsignfont cells of the Scheme of
Transliteration, which inherited the table's \bfseries and asked for a
bold Tiro Bangla that does not exist.
The build log is now clean — no overfull boxes, no font warnings. The
remaining underfulls are the documented cost of ragged right with
hyphenation off.
Docs: PLAN.md's Type row carried four stale values (11pt, 5.2in measure,
leading 1.22, hyphenation on) and now records the measured leading and
its derivation; CLAUDE.md's macro table documents \chapstart's optional
short title and \toclnp; QUESTIONS.md splits the quotation slips by
whether the source can be consulted at all — five are unpublished
archive material and are a record, not a task.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A searchable, re-typeset edition of Fiona Ross, "The Evolution of the
Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988),
transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178
plates and 410 inline type specimens cut from the scan; 51 errata.
Tracked: the transcription (src/pages), the preamble and its typographic
decisions, the cut images (plates/ — not reliably regenerable, the crop
specs for the inline cuts were never scripted), tools, and the four
working documents.
Not tracked: the built PDF, which `make` remakes from src/ and plates/;
the ProQuest scan under source/, which is third-party and needed only by
`make prep` and `make plate`; scans/ and work/, both regenerable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>