kit - six rules, in your own test suite

truecopy/kit

An interface is dodged with a return null. An assertion is not. Drop the conformance kit into your gate with a corpus of your own documents.

selfCheck() { return null } compiles, passes review and ships, and that is exactly what one of the two readers this library came from had, until a check like this one measured it from the outside.

import { checkContract, contractReport, failures, pdfWithText } from 'truecopy/kit';

it('holds the reading contract', async () => {
  const results = await checkContract(
    myReader,
    [
      { name: 'balances', document: aSoundDocument, expected: 'read' },
      { name: 'does not balance', document: aWrongDocument, expected: 'needs-review' }
    ],
    {
      referencePdf: pdfWithText([{ word: '2018', x: 50, y: 700 }]),
      open: openLikeTheApp,
      foreign: [{ name: 'a payslip', document: aPayslip }]
    }
  );

  expect(failures(results)).toEqual([]);
  writeFileSync('contract.txt', contractReport(results));
});

The six rules

  1. The verdict the corpus announces is the verdict returned.
  2. A reading that contradicts its document never comes back as sound.
  3. Everything read is reviewable by the person.
  4. The chain from bytes to records really runs, on a real PDF.
  5. A document without substance is refused, not silently returned empty.
  6. A document of another kind is refused.

Not one of them knows your domain.

Your door, not the kit’s

open is required, and it is your function. Both readers that gave rise to this already had an extractor of their own, and a kit that imposes one stops measuring the project’s reader and starts measuring its own, the one thing a conformance suite must never do.

Rule 5 has a default you can replace

A document without substance is one carrying nothing checkable, and the kit builds its own with documentWithoutSubstance(). Hand in withoutSubstance when yours is more honest: a scan that came back as one blank page, an export whose rows are all headers. What rule 5 asks is that such a document be refused rather than silently returned empty, and the closer the document is to the one your readers actually meet, the more that answer is worth.

Rule 6 needs a corpus only you have

A foreign document often carries the shape of what you seek (years, numbers, amounts) without being it. A payslip carries the vocabulary of retirement. An invoice carries dates and totals. Only you know what your reader has to be told apart from, so foreign comes from you.

This is the rule nobody writes and every reader lacks.

A real PDF, not a stub

pdfWithText(words) builds one: a real content stream, a real cross-reference table. Longer than a stub, and that is the point: what a reader must be able to do is open the file the person downloaded, and a test that fakes the PDF engine proves none of it.

pdfWithPages(pages) does the same over several pages, which is the only way to exercise what appears across them: a blank page in the middle, or two pages cut into a different number of columns.

A real Word document, for the same reason

import { docxWithText, docxWithBody } from 'truecopy/kit';

const file = new File(
  [await docxWithText(['A paragraph', ['Year', 'Amount']])],
  'agreement.docx'
);

docxWithText(blocks) builds a real archive with a real CRC: a string is a paragraph, an array of strings is a table row, and those are the two shapes a Word document has. Spelling the XML out for either of them tests the fixture more than it tests the reader.

docxWithBody(bodyXml) takes the markup instead, for the cases where the markup is the case: a merged cell, an empty one, a shape the document builder has no vocabulary for.

Both take compress, and it is not decoration. Word writes deflated parts and this builder writes stored ones, so the two paths through the archive reader are two different code paths, and a corpus that never sets it only ever exercises one.

And an OpenDocument one, on the same grid

import { odtWithText, odtWithBody } from 'truecopy/kit';

const file = new File(
  [await odtWithText(['A paragraph', ['Year', 'Amount']])],
  'agreement.odt'
);

odtWithText(blocks) and odtWithBody(bodyXml) take exactly what their Word counterparts take, and compress behaves the same way. That is the point of them being a pair: the two formats are read on one model, so a fixture written for one is a fixture written for the other, and a reader that drifts apart on them shows it here rather than on a corpus.

Reach for odtWithBody where the markup is the case. OpenDocument keeps things a page does not show - a tracked deletion, an annotation, the description a drawing carries for a screen reader - and the reader refuses to print all three. Writing that markup by hand is how you hold it to that.

A report you commit

writeFileSync('contract.txt', contractReport(results));

Pact’s idea, and the reason it outlived the assertion helpers that did the same work: the contract lives outside both sides, as an artifact you can version, review and diff. A green suite says the rules held today; a committed report says what held, and a pull request that moves a line makes the change argue for itself.

The counts are in it on purpose. A reading that quietly falls from twenty-six records to twenty-four still passes every rule, and no assertion anywhere is going to notice. The diff is.

Nothing here touches the filesystem: this library runs in a browser, and one that imports fs stops doing that. You write the string where you want it.