open - the door

truecopy/open

It is the only way from bytes to rows. It enforces the caps and the deadline, and it releases the engine. Its refusal is one your application can say in its own language.

import { openDocument, UnreadableDocument, DEFAULT_LIMITS } from 'truecopy/open';

const document = await openDocument(file, options?);

What lives here forces, and it forces for a reason no interface can match: there is no other route in.

file is whatever was dropped, and the door sorts it in three:

  • a PDF goes through pdf.js;
  • a Word document is read as a Word document, below;
  • an OpenDocument text file is read on the same grid, below;
  • text is read as text: a CSV, a TSV, a table pasted out of a viewer, an OCR dump. This path loads no engine at all, and it picks its own ruler: see documentFromText. A file that declares UTF-16 with a byte order mark is decoded as UTF-16, because the mark is the only thing a plain text file says about itself.

Anything else - an image, a program, a spreadsheet, an archive that is neither a .docx nor a .odt - is refused by name, and that refusal is the half that matters. Until 2.0.6 there were only two paths and everything that was not a PDF fell down the text one, so a .docx came back as eleven hundred rows of mojibake, confidently, with no doubt raised anywhere: the exact failure this library exists to prevent, met on a real corpus of French collective agreements.

document.origin says which path ran. Its five values are ORIGINS - 'pdf', 'text', 'image', 'docx', 'odt' - and Origin is the type read off that list rather than written out a second time.

Default Why it exists
maximumBytes 20 MB past this it is not the document expected: refuse rather than freeze the tab
maximumPages 40 a document with thousands is not a statement
deadlineMilliseconds 30 000 a read that never returns leaves a screen with no button and no way out

No byte leaves the process.

keepPage, because a cap cuts at the wrong end

maximumPages cuts at the end and only there, which is the wrong end of a long document: an annual report prints its portfolio from page 313 to page 1427, and reading its first forty pages reads a cover, a letter and an auditor’s opinion.

await openDocument(file, { keepPage: (n) => n >= 313, maximumPages: 200 });

Opening the engine on a page is what a reading costs - measured at 99% of it on a 381-page report - so a page nobody wants is better not opened than opened and dropped.

The pages that are kept keep the number the document gives them: a finding about page 313 says 313, whether or not page 312 was opened. maximumPages then counts the pages opened, so it still bounds the work.

superscripts, because a raised ordinal is not a line

A superscript sits a few points above the line it belongs to - more than the tolerance that groups items into rows - so it becomes a row of its own, and rows come out top to bottom. The mark is therefore emitted before the words it belongs to. An AMF filing reads er Resultat du 1 semestre 2026; with the option it reads Resultat du 1er semestre 2026.

await openDocument(file, { superscripts: true });

A second shape is a column defect rather than a row one: the er sits inside the tolerance, and the cut puts a boundary between 1 and it, giving 1 ersemestre. The same option covers both, because the mark is taken on the assembled row.

It is off unless asked for. It moves row text, and a reader whose patterns learned the split form keeps reading it until its own bench says otherwise. Widening the row tolerance instead would weld genuinely neighbouring lines together, which is why this is recognition and not a looser threshold: what separates a superscript from a line of its own is the glyph height, which PositionedItem now carries whenever the engine measured one. The field is left off when no height is reported, so a paste, a CSV, a .docx and any caller comparing an exact item are untouched.

A Word document is read as one

const document = await openDocument(new File([bytes], 'agreement.docx'));
document.origin; // 'docx'

What is read is what the file already declares: paragraphs, and table cells on their own grid. Nothing votes on a boundary here - a .docx writes its rows and cells down, unlike a page, where the cut has to be worked out - so the ruler is the field’s index, the one a CSV is read with. A cell left empty stays an empty field in its own column, because closing that gap is how the third value of a row ends up under the second header. A cell spanning several columns holds the place of all of them.

One page, always. Word paginates when it renders, on the fonts and the paper of whoever opens it, so numbering pages off the breaks stored in the file would put a page in a citation that the next reader cannot find.

No dependency: a ZIP is inflated with DecompressionStream, which Node 20 and every browser already have. An entry is bounded as it inflates rather than on the size it declares, so a small archive holding a very large part is refused rather than handed to a tab.

An OpenDocument file is read on the same grid

const document = await openDocument(new File([bytes], 'agreement.odt'));
document.origin; // 'odt'

Since 2.0.7. Everything the section above says holds here word for word - paragraphs and rows of cells, the index ruler, one page - and that is the claim worth making: what a reader gets does not depend on which editor wrote the file.

MEASURED before it existed, on the French collective agreements the DILA publishes: 4 767 agreements out of 395 581 came back without a citation for no reason other than their being a .odt. They were refused by name, correctly, and refused all the same.

Three things it will not print, each because printing them would put in a citation something the page does not show: a tracked deletion, which the format keeps in the file; an annotation; and the title or description a drawing carries for a screen reader. A footnote body is left out for the same family of reasons, and that one is a known limit rather than a rule.

Known limit, and deliberate: a cell merged down a column reads as empty on the rows that continue it. Word writes the value once, so repeating it would put a value on rows where the document prints none.

The refusal is named, so the sentence is yours

try {
  await openDocument(file);
} catch (error) {
  if (error instanceof UnreadableDocument) {
    // 'empty' | 'too-big' | 'no-text' | 'too-slow'
    // 'binary' | 'no-engine' | 'not-opened'
    showInYourOwnWords(error.reason);
  }
}

binary is the one that names what the file is rather than what is missing from it: bytes that are not text and not a document this library opens.

The library used to ship English sentences and nothing else, which broke its own rule: it names the rule that broke, your application writes the sentence in its own voice. An application speaking anything but English had to keep a copy of this whole file to say “this file is over 20 MB”.

reason is for the code; message is there for whoever has no application to write the sentence for them.

The pdf.js worker

// Vite
import workerSrc from 'pdfjs-dist/legacy/build/pdf.worker.mjs?url';
await openDocument(file, { workerSrc });

In a browser this is required: pdf.js refuses to start without it. In Node you can leave it out and pdf.js runs inline: slower on a long document, correct everywhere, and it is what makes the whole chain testable outside a browser.

truecopy deliberately does not resolve that URL. Vite wants ?url, webpack wants new URL(…, import.meta.url), a plain page wants a path it can serve. Picking one would lock every caller into that bundler for the sake of one line.

Bringing your own engine

import * as pdfjs from 'pdfjs-dist'; // the modern build
await openDocument(file, { pdfjs });

The default is the legacy build, which runs in Node and keeps the chain from real bytes to a parsed row testable. The modern one is smaller by over a hundred kilobytes brotli, which decides it for anything under a byte budget. Neither is right for both, so neither is chosen for you.

PdfEngine is structural: both builds satisfy it without knowing the type exists, and so would another engine.

The font pack, and what it does not repair

// Vite, both resolved by you, like workerSrc
await openDocument(file, {
  standardFontDataUrl: '/standard_fonts/',
  cMapUrl: '/cmaps/',
  cMapPacked: true
});

standardFontDataUrl is where pdf.js finds its standard font pack; cMapUrl and cMapPacked are its character maps, cMapPacked saying the pack is the compressed one that ships in pdfjs-dist. All three are passed through untouched: pdf.js reads them, this library does not.

What they are not is a repair for garbled text, and that deserves saying because the warning invites the mistake. Measured on 125 real annual reports: pdf.js prints Ensure that the standardFontDataUrl API parameter is provided on every one of them, three come back with genuinely broken text - difOciqe e defa..abqe where the page prints a French sentence - and passing the pack changes not one character of any of them.

The corruption is a subsetted font whose ToUnicode map is absent or wrong: the glyphs are embedded, their mapping to characters is not, and a pack of standard fonts has nothing to say about a custom one. So these silence a warning and serve a document that really does use an unmapped standard font. They do not rescue a document like those three, and nothing here can: that text is lost at the source.

Also here

  • withDeadline(promise, ms): bound any read in time. The timer is always cleared, even when the read wins the race: a forgotten timer keeps the process awake and fails tests long after they passed.
  • positionedItems(items): the one pure step between a PDF engine and this library’s data, exported so a test driving a real engine over real bytes can prove the chain joins up.

Everything that turns those items into rows and columns lives in layout, which needs no engine, no bytes and no clock.