On this page· 3
Specification
- Revision 6.0
- file format 6
- Whole specification on one page
Merge takes two packs a and b and a strategy, and produces one pack.
Preconditions:
- Each input MUST pass verify steps 1 to 8 (§9) before merging.
- Each input appendix MUST have container version byte 5 (§7.2), else
appendix_legacy_version. - An input appendix holding a tensor catalog section MUST be rejected with
merge_conflict: tensor catalogs encode section indexes internally and cannot be re-indexed. - An input section whose
section_typeis not in the CSDT registry MUST be rejected withmerge_conflict.
Appendix layer:
- Lift every section of
aandbas the tuple(section_type, format_version, payload_class, flags, record_count, record_stride, dtype bytes, align_log2, data bytes). - The output pool is all lifted tuples, sorted ascending by that tuple in that field order (integers numerically, byte strings lexicographically), with exact duplicates removed.
sections_dedupedcounts the removed duplicates. - Each input section's new index is its position in the pool. Every
localbinding target (§5.4) in that input's records is rewritten to the new section index;record_indexis unchanged, because sections move whole. - The pool is written as a new version-5 container. An empty pool means no appendix.
Record layer, applied to a's records and then b's, one record at a time, keyed by dedup key (§5.3):
payloadrecords from the inputs are dropped; one newpayloadrecord describing the output appendix is added at the end. Its members areembedded(true when the output has an appendix),csdt_file_checksum(the output container'sfile_checksum, else 0),sections(one{index, section_type, record_count}per pool entry, withoutdtype) andrefs(theref_idof everycsdt_refrecord in the output). All four are emitted even when empty, so an empty pool with nocsdt_refgives{"k":"payload","v":1,"embedded":false,"csdt_file_checksum":0,"sections":[],"refs":[]}.- New key: insert (
records_inserted). - Existing record equal to the incoming one (value equality after parsing): keep, count
records_skipped. - Both
ent:- existing is a stub and incoming is not: take incoming, count
stubs_resolvedandrecords_updated; - incoming is a stub and existing is not: keep existing (
records_skipped); - otherwise a conflict (
conflicts), resolved by the strategy:higher_saliencetakes incoming only if itssalienceis strictly greater (missing counts as 0.0);latesttakes incoming;unionkeeps existing;manualkeeps existing and records the conflict as deferred. A conflict countsconflicts; when the strategy takes the incoming record it also countsrecords_updated, and when it keeps the existing one,records_skipped.
- existing is a stub and incoming is not: take incoming, count
- Any other kind:
latesttakes incoming (records_updated); every other strategy keeps existing (records_skipped). - The output records are written in sort key order (§5.3) with the rules of §2 to §6.
Properties:
- For inputs whose shared dedup keys carry equal records,
merge(a, b)andmerge(b, a)produce byte-identical files. With conflicts, the result depends on argument order underlatestandhigher_salienceties. - Merge treats
mutrecords like any opaque record (dedup by the padded LSN inid). Combining two mutation logs with merge is a producer error (§5.2.1), not something merge detects. - The merge report keys are
records_inserted,records_updated,records_skipped,stubs_resolved,conflicts,sections_merged,sections_deduped.
Diff compares two packs a (before) and b (after).
- Index each side by dedup key (§5.3). A key that occurs more than once on either side is listed in
duplicate_keysand is not otherwise compared. added: keys only inb.removed: keys only ina.changed: keys in both whose records differ in identity form (§10.3). Each entry carries the key and both record lines.- Appendix: sections are matched by section index; a section is
added,removed, orchangedwhen itssection_typeor its descriptor CRC32C differs. - Every output list is sorted by sort key (records) or by section index (sections).
Diff never reports a difference that exists only in member order, whitespace or number spelling, because it compares identity forms.
Shard produces a new pack holding a connected subset of one pack's graph. The Sharder conformance level is provisional in revision 6.0; this section is its contract.
Graph view of a pack: nodes are ent records by id; directed edges are rel records src → tgt labelled kind; an entity has an embedding when an embref record names it.
Shard specs:
entity_centric{seed, max_depth, max_fanout?, relationship_filter?, direction, phase_coherent}: breadth-first fromseedtomax_depthhops, following edges indirection(outgoing,incoming,both), at mostmax_fanoutneighbours per node when set, and only edges whosekindis inrelationship_filterwhen set.search_result{entity_ids, context_depth, max_fanout?}: the listed entities plus their neighbourhood tocontext_depth.predicate{...}: entities whose fields satisfy the predicate.cluster{cluster_id, include_bridges}: MUST fail withshard_unsupported_specuntil a cluster field exists in the record layer. It MUST NOT return an empty result instead.union[specs]: the union of the member results.
Output pack:
- the selected
entrecords, plus a stubent(stub: true) for each edge endpoint outside the selection when the config asks for stubs; - the
relrecords whose endpoints are both in the output; - a
hyperrecord only when every member is selected; embrefrecords of selected entities, with the appendix rewritten to hold only the referenced rows and the bindings remapped;- every other appendix section listed in
dropped_sections, never dropped silently; - a
metarecord whosecreated_atis supplied by the caller. Two runs with the same input, spec andcreated_atMUST produce byte-identical packs.
A shard whose quality score is below a configured minimum fails with shard_quality_below_threshold.
© 2020-2026 Cintile Inc. All Rights Reserved.
Anyone may implement this format. Copying or republishing the text of this specification requires permission from Cintile Inc.