A searchable, re-typeset edition of Fiona Ross, "The Evolution of the Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988), transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178 plates and 410 inline type specimens cut from the scan; 51 errata. Tracked: the transcription (src/pages), the preamble and its typographic decisions, the cut images (plates/ — not reliably regenerable, the crop specs for the inline cuts were never scripted), tools, and the four working documents. Not tracked: the built PDF, which `make` remakes from src/ and plates/; the ProQuest scan under source/, which is third-party and needed only by `make prep` and `make plate`; scans/ and work/, both regenerable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4.0 KiB
4.0 KiB
Ross 1988 thesis — re-typeset edition: plan (v2, confirmed 2026-09-13)
Source: 10731406.pdf — Fiona G. E. Ross, The Evolution of the Printed Bengali
Character from 1778 to 1978, Ph.D., SOAS, 1988. ProQuest scan, 431 PDF pages,
300 dpi bitonal (a few plates in colour), ABBYY text layer (NOT used as source).
Confirmed design
| Aspect | Decision |
|---|---|
| Layout | Reflowed text. \origpage{N} marks where original page N begins: bold sans number in the outer margin + thin bar in the line + page.N hyper-target. Running head shows the original page range on each edition page; edition page number in the footer |
| Notes | Footnotes become endnotes (enotez) in a final Notes section, grouped "Notes to ", keeping original numbers via \fn{n}{text}; marks link to notes, notes link back |
| Type | 11pt TeX Gyre ScholaX (Century Schoolbook) on a 5.2in measure (≈70 characters a line); Liberation Sans for marks and heads; leading 1.22; ragged right (ragged2e \RaggedRight, hyphenation on); hanging punctuation (microtype protrusion, dashes hang far less); widows and orphans forbidden with \raggedbottom. Tiro Bangla for Bengali that shows the modern script (the Scheme of Transliteration). Bengali that shows a historical typeface is never typed: it is cut from the scan by tools/crop_inline.py and set with \ig at native size |
| Structure | \chapstart{Title} = new page + centred bold head + PDF bookmark + notes group. One .tex per original page in src/pages/pNNNN.tex (NNNN = PDF page); \chapstart must come before \origpage in a file |
| Plates | Cropped from the 300 dpi scan by tools/crop_plate.py (scanner bars, page number, caption stripped; gutter shadow swept; deskewed when the page is out of square — the one step that resamples), embedded lossless at native resolution. \plateop{N}{no}{img}{caption} for a plate that begins page N; \plate{no}{img}{caption} otherwise. Per-page crop recipe recorded in the manifest notes column |
| Colophon | src/colophon.tex = leaf 2: description of the conversion + ProQuest notice text. Blank stamped leaf (PDF p.3) dropped and noted there |
| Editorial | Original spelling and typography stand. Handwritten corrections in the scan are adopted silently and logged in QUESTIONS.md. The author's own slips are corrected via \erratum{corrected}{as printed} and listed in the Errata section. Slips inside quoted matter are never corrected — mark \qslip{as printed} and table them in QUESTIONS.md so they can be checked against the source. A grammatical slip stands, with an editorial [sic]. All confirmed by Sammay, 2026-09-13. The three typographic intentions the typewriter could not realise — en-dashes for numeric ranges, true quotation marks, small caps for institutional acronyms — are applied at set time by tools/polish.py; the transcription itself stays as typed (Sammay, 2026-09-13) |
| Uncertain | \unsure{...} → unsure.log → QUESTIONS.md; never silently guessed |
| Contents | Contents and List-of-Plates entries are whole-line hyperlinks to the original page target (page.N) |
| Versioning | Git (from 2026-09-14, Sammay). The built PDF is not tracked and not archived — make regenerates it from src/ and plates/. The ProQuest scan (source/) is untracked too, being third-party; only make prep and make plate need it. Superseded design drafts stay in older-versions/ as text |
Workflow per batch
tools/prep.sh A B→scans/hi-NNN.jpg(130 dpi) +work/ocr/pNNNN.txt(tesseract scaffold).- Read every page image. For clean prose pages
tools/ocr2tex.py N printed [--head T]gives a draft; proofread word-by-word against the image, add\emph,\fn, diacritics. - Plates:
python3 tools/crop_plate.py NNNNthen a one-line\plateopfile. - Update
manifest.tsv;./build.sh; spot-check.
Remaining
Introduction pp.17–30 (PDF 19–32) → Part I from p.31 (PDF 33) … Bibliography p.419 (PDF 421) → end (PDF 431). ~245 text + ~170 plate pages: ~10–12 sessions.