records - which rows belong to the same record

truecopy/records

A record often occupies two or three printed rows. Every mechanism here works on the printed row, so one row per record undercounts a real document by three to five times, without a word.

import { recordsFrom, spineWidthOf } from 'truecopy/records';

const { records, loose, findings } = recordsFrom(rows, options?);
Field
records RecordBlock[]: the spine row, and every rows index it carries
loose number[]: rows that belong to no record, returned rather than lost
findings RecordFinding[]: spine-not-sharp or no-spine, with its numbers

In a great many real documents a record occupies two or three printed rows: a street on one line, the figures on the next, the postcode on a third. Read one row per record, a property schedule of sixty-five holdings comes back as two hundred and forty-nine, and nothing says so.

What joins two rows

A table has a spine: the rows that carry its full width. A narrower row joins the nearest spine it fits beside, every column it fills being a column that spine leaves empty.

That second half is the whole difference from the rule everyone writes first, and a count cannot see it. Attaching each fragment to its nearest spine returns the same number of records on both measured documents, and swallows twenty-three page-furniture rows into them on one: a page header sits one row from a spine, so proximity takes it. It cannot fill a column the spine leaves empty, so compatibility refuses it.

spineWidthOf(rows); // the widest row, less one

The widest row less one, because a schedule prints a zero as a blank: a perfectly complete row is one cell short of the widest, and requiring the exact maximum rejects half the real records.

Nothing is ever grouped away

Every row comes back either in a record or in loose. A cover page, a letterhead, a page number: a row this mechanism could not place is exactly what a caller may need to look at, and a list that quietly loses a quarter of a document is the plausible-but-wrong reading argued against everywhere else here.

A table where a record is a row returns one record per row and changes nothing, which is the case it was measured against.

When it says it cannot tell

if (findings.some((f) => f.code === 'spine-not-sharp')) {
  // a caller who knows its document passes the width it knows
  recordsFrom(rows, { spineWidth: 8 });
}

spine-not-sharp is a real property of some documents rather than a threshold left untuned. A record that leaves one column empty - a fee of zero, a tax that does not apply - is exactly as wide as a rich fragment, and no width separates them. Measured on a real property schedule: sixty-one records at one threshold, seventy-one at the next, and the judged answer is sixty-five with neither reachable.

no-spine says no row reaches the full width at all, and every row comes back on its own.

Two options, and both exist because the caller may know something this library does not:

Option
spineWidth how many filled cells a complete row carries. Left out, measured
reach how far above or below its spine a fragment may sit. Two by default

Wider does not read more, it reads the neighbour: the row three above a spine is usually the row below the previous spine.

Hand it a page at a time

A whole document handed over flat has its width set by the widest row printed anywhere in it. Measured on a real annual report: 3115 rows, a widest row of eight, and seven records, where the same document read page by page gives eleven hundred.

const { pages } = await readTable(file);
const perPage = pages.map((rows) => recordsFrom(rows));

pages is what readTable hands back beside the flat rows, and this is the second reading that needs it.