Files
tepbc-ross-retyped/PLAN.md
T
bdeshiandClaude Opus 5 5f15c89a20 Drop small caps for acronyms; ninth diacritic fix; verify forces a build
Abbreviations are now plain uppercase throughout (Sammay). The pass
small-capped an explicit whitelist, which could only ever be partial: on
the Abbreviations and Conventions page BFBS, LMS and MLCo stood
unconverted in one column beside small-capped BL, BMS, EIC, IOL, OUP,
SOAS and SPG, and "MS EUR 30" split a single shelfmark across two styles.

Completing the whitelist was the obvious fix and the wrong one. The
original is a typescript — a typewriter cannot set small caps — so the
source defines none, and rule 3 leaves every acronym as the plain
uppercase it prints. smallcaps() is now a documented no-op.

p0431: Rtusamhāra -> Ṛtusaṃhāra, verified at 700 dpi (dot under the R,
dot under the m). It was catchable because it broke the Bibliography's
own pattern: the author gives a title in full transliteration and its
anglicised form plain — "Ānanda Bājāra Patrikā (Ananda Bazar Patrika)" —
so a Bibliography entry with the marks stripped is an anomaly, while most
of the remaining flags there are her distinction rather than our error.

Makefile: verify now depends on `build`, not `pdf`. `pdf` is
timestamp-conditional, and a source edited while a build is running
leaves a PDF newer than the source it does not contain — make then skips
the rebuild and verify passes against stale output. That happened here,
and reprocheck caught it correctly twice while I twice misread it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 02:12:42 +06:00

5.4 KiB
Raw Blame History

Ross 1988 thesis — re-typeset edition: plan (v2, confirmed 2026-09-13)

Source: 10731406.pdf — Fiona G. E. Ross, The Evolution of the Printed Bengali Character from 1778 to 1978, Ph.D., SOAS, 1988. ProQuest scan, 431 PDF pages, 300 dpi bitonal (a few plates in colour), ABBYY text layer (NOT used as source).

Confirmed design

Aspect Decision
Layout Reflowed text. \origpage{N} marks where original page N begins: bold sans number in the outer margin + thin bar in the line + page.N hyper-target. Running head shows the original page range on each edition page; edition page number in the footer
Notes Footnotes become endnotes (enotez) in a final Notes section, grouped "Notes to ", keeping original numbers via \fn{n}{text}; marks link to notes, notes link back
Type 12pt XCharter (Matthew Carter's Charter, extended; TeX Live, loaded by path) on a 5.95in measure — about 73 characters a line. Replaced TeX Gyre ScholaX on 2026-09-15 (Sammay): measured off the metrics, ScholaX had the smallest lowercase relative to its capitals of the faces considered, x-height/cap 0.645 against XCharter's 0.717, so its caps and ascenders dominated the page. XCharter covers all 35 non-ASCII characters here (the dots-below set defeats ETbb, Domitian, STEP, Tempora) with real small caps and oldstyle figures; Liberation Sans for marks and heads. Leading 1.59 — a 22.9pt line, which is the original's own, measured off the scan at 300 dpi by autocorrelating the row-ink profile; the inline specimens set it, not the text. Inline specimens at native size (\igscale = 72.27/300), the 74% of the line the original gives them. First line indented 1.6em, \parskip 0.15em (ragged2e resets \parindent from \RaggedRightParindent, so both are set). Ragged right and unhyphenated (\RaggedRight, \hyphenpenalty=10000, \lefthyphenmin=62) — the typescript breaks only at spaces and at hyphens it actually typed. Hanging punctuation (microtype protrusion, dashes hang far less); widows and orphans forbidden with \raggedbottom. Tiro Bangla for Bengali that shows the modern script (the Scheme of Transliteration). Bengali that shows a historical typeface is never typed: it is cut from the scan by tools/crop_inline.py and set with \ig
Structure \chapstart{Title} = new page + centred bold head + PDF bookmark + notes group. One .tex per original page in src/pages/pNNNN.tex (NNNN = PDF page); \chapstart must come before \origpage in a file
Plates Cropped from the 300 dpi scan by tools/crop_plate.py (scanner bars, page number, caption stripped; gutter shadow swept; deskewed when the page is out of square — the one step that resamples), embedded lossless at native resolution. \plateop{N}{no}{img}{caption} for a plate that begins page N; \plate{no}{img}{caption} otherwise. Per-page crop recipe recorded in the manifest notes column
Colophon src/colophon.tex = leaf 2: description of the conversion + ProQuest notice text. Blank stamped leaf (PDF p.3) dropped and noted there
Editorial Original spelling and typography stand. Handwritten corrections in the scan are adopted silently and logged in QUESTIONS.md. The author's own slips are corrected via \erratum{corrected}{as printed} and listed in the Errata section. Slips inside quoted matter are never corrected — mark \qslip{as printed} and table them in QUESTIONS.md so they can be checked against the source. A grammatical slip in the author's own prose stands exactly as printed and unmarked (Sammay, 2026-09-14, reversing the editorial [sic] first added on 2026-09-13): the seven [sic]s in this edition are all hers, and an eighth in her own notation would have been indistinguishable from them. The edition now makes no interpolation into the text at all. Other decisions confirmed by Sammay, 2026-09-13. The typographic intentions the typewriter could not realise — en-dashes for numeric ranges, true quotation marks — are applied at set time by tools/polish.py. Small caps for acronyms were removed on 2026-09-15 (Sammay): a typescript cannot set small caps, so the source defines none, and a partial whitelist left the Abbreviations page visibly inconsistent; the transcription itself stays as typed (Sammay, 2026-09-13)
Uncertain \unsure{...} → unsure.log → QUESTIONS.md; never silently guessed
Contents Contents and List-of-Plates entries are whole-line hyperlinks to the original page target (page.N)
Versioning Git (from 2026-09-14, Sammay). The built PDF is not tracked and not archived — make regenerates it from src/ and plates/. The ProQuest scan (source/) is untracked too, being third-party; only make prep and make plate need it. Superseded design drafts stay in older-versions/ as text

Workflow per batch

  1. tools/prep.sh A B → scans/hi-NNN.jpg (130 dpi) + work/ocr/pNNNN.txt (tesseract scaffold).
  2. Read every page image. For clean prose pages tools/ocr2tex.py N printed [--head T] gives a draft; proofread word-by-word against the image, add \emph, \fn, diacritics.
  3. Plates: python3 tools/crop_plate.py NNNN then a one-line \plateop file.
  4. Update manifest.tsv; ./build.sh; spot-check.

Remaining

Introduction pp.17–30 (PDF 19–32) → Part I from p.31 (PDF 33) … Bibliography p.419 (PDF 421) → end (PDF 431). ~245 text + ~170 plate pages: ~10–12 sessions.