records - which rows belong to the same record
truecopy/records
A record often occupies two or three printed rows. Every mechanism here works on the printed row, so one row per record undercounts a real document by three to five times, without a word.
import { recordsFrom, spineWidthOf } from 'truecopy/records';
const { records, loose, findings } = recordsFrom(rows, options?);
| Field | |
|---|---|
records |
RecordBlock[]: the spine row, and every rows index it carries |
loose |
number[]: rows that belong to no record, returned rather than lost |
findings |
RecordFinding[]: spine-not-sharp or no-spine, with its numbers |
In a great many real documents a record occupies two or three printed rows: a street on one line, the figures on the next, the postcode on a third. Read one row per record, a property schedule of sixty-five holdings comes back as two hundred and forty-nine, and nothing says so.
What joins two rows
A table has a spine: the rows that carry its full width. A narrower row joins the nearest spine it fits beside, every column it fills being a column that spine leaves empty.
That second half is the whole difference from the rule everyone writes first, and a count cannot see it. Attaching each fragment to its nearest spine returns the same number of records on both measured documents, and swallows twenty-three page-furniture rows into them on one: a page header sits one row from a spine, so proximity takes it. It cannot fill a column the spine leaves empty, so compatibility refuses it.
spineWidthOf(rows); // the widest row, less one
The widest row less one, because a schedule prints a zero as a blank: a perfectly complete row is one cell short of the widest, and requiring the exact maximum rejects half the real records.
Nothing is ever grouped away
Every row comes back either in a record or in loose. A cover page, a letterhead, a page number: a row this mechanism could not place is exactly what a caller may need to look at, and a list that quietly loses a quarter of a document is the plausible-but-wrong reading argued against everywhere else here.
A table where a record is a row returns one record per row and changes nothing, which is the case it was measured against.
When it says it cannot tell
if (findings.some((f) => f.code === 'spine-not-sharp')) {
// a caller who knows its document passes the width it knows
recordsFrom(rows, { spineWidth: 8 });
}
spine-not-sharp is a real property of some documents rather than a threshold left untuned. A record that leaves one column empty - a fee of zero, a tax that does not apply - is exactly as wide as a rich fragment, and no width separates them. Measured on a real property schedule: sixty-one records at one threshold, seventy-one at the next, and the judged answer is sixty-five with neither reachable.
no-spine says no row reaches the full width at all, and every row comes back on its own.
Two options, and both exist because the caller may know something this library does not:
| Option | |
|---|---|
spineWidth |
how many filled cells a complete row carries. Left out, measured |
reach |
how far above or below its spine a fragment may sit. Two by default |
Wider does not read more, it reads the neighbour: the row three above a spine is usually the row below the previous spine.
Hand it a page at a time
A whole document handed over flat has its width set by the widest row printed anywhere in it. Measured on a real annual report: 3115 rows, a widest row of eight, and seven records, where the same document read page by page gives eleven hundred.
const { pages } = await readTable(file);
const perPage = pages.map((rows) => recordsFrom(rows));
pages is what readTable hands back beside the flat rows, and this is the second reading that needs it.