8 Commits
Author SHA1 Message Date
bdeshiandClaude Sonnet 5.5 4424cc9887 Add the EPUB edition: converter, class-based audit, shared front matter
tools/make_epub.py converts the same src/ the PDF is set from, running
polish.py first, with original page numbers as page-list metadata, endnotes
gathered in one linked Notes section, and the colophon and errata included.
It raises on any macro it does not declare. tools/epub_audit.py checks by
class of fault (LaTeX residue, escaping, empty blocks, links, images, XML,
content, typography drift from the PDF); make epub runs both.

\byedition{PDF}{EPUB} lets the colophon carry the sentences that are true of
only one edition; the errata introduction moves into src/errata.tex so both
editions print one copy. reprocheck.py now accepts an .epub and skips
environment parameters that are layout, not copy.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-09-29 14:36:18 +06:00
bdeshiandClaude Sonnet 5.5 2ec880d1c4 Add a generated cover and a content-hashed version mark
The cover is a full-bleed mosaic of historical Bengali types filling the word
ব্যঞ্জণ, produced by tools/gen_cover_mosaic.py from plates/ and fonts/ alone
(build.sh remakes it if missing). It is static, so ebook readers can cache it.

The mark on the About this edition page is drawn by tools/gen_seal.py, seeded
from a SHA-256 of src/ (never the build's own outputs), so it changes exactly
when the transcription does. The colophon is shortened to name the cover
typefaces and the mark's mechanism, and drops the trailing ProQuest note.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-09-29 14:36:18 +06:00
bdeshiandClaude Sonnet 5.5 d48ce1f50c Re-encode plates to lossless JBIG2 / JPEG at build time (115 MB -> 12 MB)
The scan is a bilevel 300 dpi JBIG2 layer plus lower-resolution colour, so
storing every plate as 8-bit PNG only added size. tools/optimize_plates.py
writes copies under work/plates-opt (bilevel to generic-mode, lossless JBIG2;
genuine halftone to JPEG q88) and repoints work/pages at them; plates/ and
src/ are untouched. The Dockerfile builds jbig2enc and adds qpdf, ghostscript
and img2pdf.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-09-29 14:36:18 +06:00
bdeshiandClaude Sonnet 5.5 8857e5a008 Ignore bytecode and .scratch; stop tracking the one committed .pyc
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-09-29 14:36:18 +06:00
bdeshiandClaude Opus 5 e602e7e631 Link audit: fix three plate links from elided ranges
CLAUDE.md requires re-auditing the cross-reference rule after any change
to it, and polish.py changed twice this session. tools/linkaudit.py does
it: three classes checked mechanically, the fourth printed to be read.

  dangling targets                      0
  plate page mismatch (listed vs set)   0   (all 178)
  chapter page mismatch                 0
  `p. N` self-references               25   all read, all sound

The finding: an elided range end is not a plate number of its own.
plate_sub ran re.sub(r"\d+") over the whole range, so "pls. 146-8" linked
its 8 to plate 8 (orig. p.35) and "pls. 142-5" linked its 5 to plate 5
(p.23). Three links, all pointing at the wrong plate — invisible to every
other check, since the targets exist and the text is correct. The end is
now expanded against the leading digits of the number it is elided
against: 146-8 -> plate 148 (p.334), 142-5 -> plate 145 (p.329), both
confirmed against the List of Plates and against where \plateop sets
them.

Worth noting what the audit also confirmed: in two footnotes the rule
discriminates a citation page from a self-reference in the same sentence
— "(London, 1962), p. 43; also see below, p. 63" links only the second.

NOT YET BUILT: the Docker daemon is down, so make verify has not run
against this change. polish.py was run on the host and the audit re-run
against its output, but the PDF does not yet carry the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:00:13 +06:00
bdeshiandClaude Opus 5 5f15c89a20 Drop small caps for acronyms; ninth diacritic fix; verify forces a build
Abbreviations are now plain uppercase throughout (Sammay). The pass
small-capped an explicit whitelist, which could only ever be partial: on
the Abbreviations and Conventions page BFBS, LMS and MLCo stood
unconverted in one column beside small-capped BL, BMS, EIC, IOL, OUP,
SOAS and SPG, and "MS EUR 30" split a single shelfmark across two styles.

Completing the whitelist was the obvious fix and the wrong one. The
original is a typescript — a typewriter cannot set small caps — so the
source defines none, and rule 3 leaves every acronym as the plain
uppercase it prints. smallcaps() is now a documented no-op.

p0431: Rtusamhāra -> Ṛtusaṃhāra, verified at 700 dpi (dot under the R,
dot under the m). It was catchable because it broke the Bibliography's
own pattern: the author gives a title in full transliteration and its
anglicised form plain — "Ānanda Bājāra Patrikā (Ananda Bazar Patrika)" —
so a Bibliography entry with the marks stripped is an anomaly, while most
of the remaining flags there are her distinction rather than our error.

Makefile: verify now depends on `build`, not `pdf`. `pdf` is
timestamp-conditional, and a source edited while a build is running
leaves a PDF newer than the source it does not contain — make then skips
the rebuild and verify passes against stale output. That happened here,
and reprocheck caught it correctly twice while I twice misread it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 02:12:42 +06:00
bdeshiandClaude Opus 5 05c41040f7 Diacritic corrections, link styling, and a build lock
Eight diacritic corrections at seven sites, each verified against the page
image at 900 dpi:

  p0049  Nastaʿlīq -> Nasta`līq   the mark is a quote, not a ʿayn (Sammay)
  p0072  Cálidás   -> Cálidās
  p0256  Rajākṛṣṇa -> Rājākṛṣṇa
  p0260  Īsvaracandra -> Īśvaracandra  (x2)
  p0430  Iśvaracandra -> Īśvaracandra
  p0430  Gītā -> Gīta  (x2)
  p0430  Munṣī -> Munśī

Found by two nets over the transcription: a rare-character inventory
(U+02BF appeared exactly once in 431 pages — the tell) and a check for
names that appear with more than one accenting, 84 occurrences of which
~32 are now checked. Neither net catches a name mis-transcribed the SAME
way everywhere, so this is a partial sweep, not a clean bill.

The French accents are all handwritten additions in the scan — pen-drawn
cedillas and acutes — applied inconsistently by the author and reproduced
per instance. All of that cluster checks out.

Clickable regions now carry a 0.4pt hairline rule in RGB(95,125,165) on
the Contents/List-of-Plates page numbers and on resolved cross-references;
endnote superscripts and whole-line entries stay unmarked. polish.py emits
\xref for the former and \hyperlink for the latter so the two differ.

Also: plate 1's List-of-Plates entry uses \plx, which build_maps() did not
scan, so plate 1 was the one plate of 178 whose "pl. 1" references never
linked.

build.sh takes an atomic mkdir lock: two builds sharing work/ corrupt
main.aux and produce an error that names the wrong file entirely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 00:50:04 +06:00
bdeshiandClaude Opus 5 82849eb94d Initial commit: Ross 1988 re-typeset edition
A searchable, re-typeset edition of Fiona Ross, "The Evolution of the
Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988),
transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178
plates and 410 inline type specimens cut from the scan; 51 errata.

Tracked: the transcription (src/pages), the preamble and its typographic
decisions, the cut images (plates/ — not reliably regenerable, the crop
specs for the inline cuts were never scripted), tools, and the four
working documents.

Not tracked: the built PDF, which `make` remakes from src/ and plates/;
the ProQuest scan under source/, which is third-party and needed only by
`make prep` and `make plate`; scans/ and work/, both regenerable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 19:03:25 +06:00