Specification
PLXI specification · revision 6.0 · file format 6 · plxi.org
All sections of revision 6.0 on one page. Ids are the same as on the section pages.
| Document | PLXI file format, version 6: specification |
| Revision | 6.0 (2026-10-01; amended 2026-10-06) |
| Applies to | file format 6 |
This document defines the PLXI version 6 file format (.plxi), its transport encoding (.plxi.gz), the record layer carried in its text body, and the behaviour of the operations defined on it: verify, merge, diff and shard. It is written so that a reader or writer can be implemented in any language without reading Rust.
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.
Other conventions:
- Byte offsets are zero-based. "Absolute offset" means an offset from the first byte of the
.plxifile.Lis the file length in bytes. - A range
a..bof bytes or offsets starts ataand ends beforeb: bytes 0..4 are bytes 0, 1, 2 and 3. - All multi-byte integers in binary structures are little-endian and unsigned unless stated otherwise.
- Grammar is written in ABNF [RFC5234]. ABNF quoted strings are case-insensitive, so every case-sensitive literal below is written with
%xbyte values. - "JSON" means RFC 8259 JSON [RFC8259]. "UTF-8" is as defined in RFC 3629 [RFC3629].
- SHA-256 is as defined in FIPS 180-4 [FIPS180-4]. CRC32C is the Castagnoli CRC, polynomial
0x1EDC6F41[RFC3720].
A .plxi file is a text section followed, optionally, by a binary appendix, followed by a fixed binary footer:
offset 0 header line: 256 bytes of header content padded with
spaces, then LF (257 bytes)
offset 257 body: zero or more JSONL record lines, each ending in LF
marker line "===PLXI_BINARY_APPENDIX===" LF
(only when an appendix exists)
T = text_section_size: first byte after the text section
padding: 0x00 bytes up to the next multiple of 64
(only when an appendix exists)
A = csdt_offset: embedded CSDT container, csdt_size bytes
(only when an appendix exists)
L - 96 footer: 96 bytes
L end of fileRules:
- The header line MUST be exactly 257 bytes: 256 bytes of header content and padding (§3) followed by one LF (
0x0A) at offset 256. The body therefore always starts at offset 257. - The footer MUST occupy the last 96 bytes of the file. A file shorter than 257 + 96 = 353 bytes is not a PLXI v6 file (
invalid_header). - Without an appendix, the footer MUST follow the last body line directly:
T = L - 96,csdt_offset = 0,csdt_size = 0, and no marker line is present. - With an appendix (
csdt_size > 0):- The marker line
===PLXI_BINARY_APPENDIX===followed by LF MUST follow the last record line.Tis the offset of the byte after that LF. - Padding of
(64 - T mod 64) mod 64bytes, each0x00, MUST follow, so thatcsdt_offset = Trounded up to a multiple of 64. - The embedded container MUST start at
csdt_offsetand MUST end atL - 96:csdt_offset + csdt_size = L - 96.
- The marker line
csdt_offsetMUST be 0 whenevercsdt_sizeis 0.
The 64-byte placement exists so that the embedded CSDT container's own 64-byte alignment guarantees hold when the whole .plxi file is memory-mapped at a page-aligned base.
header-line = header-content padding LF ; exactly 257 octets
header-content = magic SP version SP records SP csdt-offset SP csdt-size SP
sha-field
magic = %x50.4C.58.49 ; "PLXI"
version = %x76 dec
; "v" then the format version, "v6"
records = %x72.65.63.6F.72.64.73.3D dec ; "records=" N
csdt-offset = %x63.73.64.74.5F.6F.66.66.73.65.74.3D dec
; "csdt_offset=" O
csdt-size = %x63.73.64.74.5F.73.69.7A.65.3D dec
; "csdt_size=" C
sha-field = %x73.68.61.32.35.36.3D [sha-hex]
; "sha256=" H or "sha256=" (placeholder)
sha-hex = 64lhex
lhex = %x30-39 / %x61-66
; lowercase hexadecimal only
dec = %x30 / (%x31-39 *19%x30-39)
; decimal, no sign, no leading zeros, value < 2^64
padding = *SP
; space-fill so header-content + padding = 256 octets
SP = %x20
LF = %x0A
Each token is separated from the next by exactly one SP. There are exactly six tokens.
versionMUST be6. A header whose version token parses as any other number MUST be rejected withunsupported_version. Versions 1 to 5 are retired numbering systems and have no readers.recordsis the number of record lines in the body (§4).csdt-offsetandcsdt-sizelocate the appendix (§2). Whencsdt-sizeis nonzero,csdt-offsetMUST be a multiple of 64.sha-fieldcarries the lowercase hex SHA-256 digest defined in §6.
A writer first emits a placeholder header and, when its output is seekable, overwrites it in place with the final header after the body and appendix are written. The fixed 256-byte width makes the rewrite possible without moving any other byte.
- A final header MUST carry a 64-digit
sha-hex, and itsrecords,csdt-offset,csdt-sizeand digest MUST equal the footer'srecord_count,csdt_offset,csdt_sizeandsha256. - A placeholder header is exactly
PLXI v6 records=0 csdt_offset=0 csdt_size=0 sha256=(empty digest) padded to 256 bytes. A writer whose output cannot seek (a gzip stream) leaves it in place. For a file with a placeholder header, the footer is the only source of the counts, offsets and digest. - A header that is neither final nor placeholder (for example an empty digest with a nonzero
records) MUST be rejected withheader_footer_mismatch.
The header line is not covered by the SHA-256 digest (§6). Everything a reader trusts from it is therefore cross-checked against the footer.
- The body is a sequence of lines, each one JSON object [RFC8259] encoded as UTF-8 and terminated by one LF.
- A writer MUST NOT emit empty lines, CR characters outside JSON strings, or a line containing a raw LF inside a JSON value (JSON strings escape LF as
\n). - A reader MUST reject a body that is not valid UTF-8 with
invalid_utf8. - A reader MUST skip empty lines and MUST strip one trailing CR before parsing a line. Skipped lines do not count as records.
- The marker line (§2) is not a record. It MUST appear only as the last line of the text section and only when an appendix exists; anywhere else it is
invalid_record. recordsin the header andrecord_countin the footer count the record lines.- A writer MUST NOT emit a JSON object with duplicate member names. A reader that meets one MUST either reject the line with
json, or keep the last value for the name; the reference reader keeps the last value (serde_json behaviour, measured in §10.4). Conformance tests do not exercise duplicates.
Every record line is a JSON object with:
k(REQUIRED): a JSON string, the record kind. A line without a stringkMUST be rejected withinvalid_record.v(OPTIONAL): a non-negative JSON integer, the per-kind format version. Absent means 1.- the kind's own members (§5).
A record is typed when its kind is in the typed registry (§5.1) and its v is at most the highest version this reader supports for that kind. Every typed kind is at version 1. Otherwise the record is opaque.
- A reader MUST parse a typed record against its schema (§5.1). A schema failure (missing REQUIRED member, wrong JSON type) MUST be reported as
json. - Members the schema does not name MUST be preserved in the record's
extramap and re-emitted after the declared members. - A writer MUST emit a typed record as:
kfirst, thenv, then the declared members in schema order, thenextramembers in the order they were read. OPTIONAL members whose value is absent are omitted, except those marked "always emitted" in §5.1. - Re-emitting a typed record normalises its values: a 32-bit float member (for example
salience,confidence) is re-emitted as the shortest decimal that round-trips the 32-bit value, and whitespace, escapes and number spelling follow §10.2. A typed round trip therefore preserves the value but MAY change the bytes.
A record MUST be kept opaque, not rejected, when either:
- its kind is not in the typed registry, including the reserved kinds of §5.2 and any future kind; or
- its kind is typed but its
vis higher than this reader supports.
For an opaque record, a reader MUST keep the whole JSON object, including k and v, and a writer MUST re-emit it with the same members, the same values, and the members in the order they were read (§10.1). This is the opaque round-trip rule. It guarantees JSON value equality and member order; it does not guarantee byte equality, because escapes and number spellings are re-emitted in the form of §10.2 (a JSON integer outside the 64-bit range, for example, becomes a binary64 number; §10.4).
Opaque records take part in sorting, merge and diff through the identity rules of §5.3.
The rank is the canonical sort position. The primary id is the record's identity within its kind; pad20(n) means n in decimal, left-padded with 0 to 20 digits. The field list is the v1 schema; ? marks an OPTIONAL member, all others are REQUIRED. Every kind also carries v? and extra.
| Rank | k |
Primary id | v1 members |
|---|---|---|---|
| 0 | meta |
created_at "|" source |
created_at, source, tool?, description?, ngdb_generation? (u64) |
| 1 | payload |
empty string | embedded (bool), csdt_file_checksum? (u32, always emitted, default 0), sections? (array of {index u32, section_type u16, record_count u64, dtype? string}, always emitted, default []), refs? (array of string, always emitted, default []) |
| 2 | csdt_ref |
ref_id |
ref_id, path, shard_set_id? (u32), file_checksum (u32), source_hash, section_type (u16), section_index (u32), record_range? ([u64, u64]) |
| 3 | ent |
id |
id, t, name?, salience? (f32), stub? (bool, always emitted, default false), properties?, prov? |
| 4 | rel |
src "|" tgt "|" kind |
src, tgt, kind, properties?, prov? |
| 5 | hyper |
id |
id, members (array of string), kind, properties? |
| 6 | embref |
entity_id "|" target-key |
entity_id, target (§5.4), annotations? |
| 7 | prov |
subject "|" method "|" pad20(unix_secs) |
subject, agent, method, unix_secs (u64), inputs?, notes? |
| 8 | grounded_answer |
answer_id |
Application kind. |
| 9 | citation |
source_answer_id "|" entity_id "|" pad20(span.start) |
Application kind. |
| 10 | clause |
clause_id |
Application kind. |
| 11 | knob_set |
scope "|" key "|" pad20(unix_secs) |
Application kind. |
| 12 | discovery |
gid |
Application kind. |
| 13 | promo |
content_hash "|" pad20(order) |
Application kind. |
| 14 | codebook_manifest |
codebook_id |
Application kind. |
| 15 | session |
session_id |
Application kind. |
| 16 | sev |
pad20(step) "|" t |
Application kind. |
| 17 | session_end |
pad20(end_unix) |
Application kind. |
| 18 | outcome |
identity form (§10.3) of outcome |
Application kind. |
Members without a stated type are JSON strings; u32, u64 are non-negative JSON integers in range.
Changing a typed kind's schema incompatibly requires raising its v (§12.2). Adding an OPTIONAL member at the same v is compatible, because older readers keep it in extra.
These kind strings are reserved, but have no typed parser in this revision. Readers MUST treat them as opaque (§4.4).
mut
Ten further kind names are reserved and are specified separately. A reader keeps any record whose kind it does not know as an opaque record (§4.4).
A reserved kind is promoted to the typed registry only by a spec revision that adds it to §5.1 with the same v1 schema.
Kind names that begin with x- belong to applications: no revision of this specification adds one to the typed kinds (§5.1) or the reserved kinds (§5.2).
mut is one entry of an ordered mutation log. The schema below is the wire form of mut v1.
| Member | Type | Rule |
|---|---|---|
k |
string | "mut" |
v |
integer | 1 |
id |
string | the LSN in decimal, left-padded with 0 to 20 digits. MUST equal pad20(lsn) |
gen |
u64 | publication generation |
lsn |
u64 | log sequence number within gen |
op |
string | one of put_record, delete_record, or one of the op names reserved for future use |
rk |
string, OPTIONAL | kind of the affected record. REQUIRED for put_record and delete_record |
rid |
string, OPTIONAL | primary id (§5.1) of the affected record. REQUIRED for put_record and delete_record |
rec |
object, OPTIONAL | put_record only: the complete affected record as its own JSONL object, including its k. Its kind and primary id MUST equal rk and rid. delete_record MUST NOT carry rec |
| other | any | members of the reserved ops, preserved. None of them may be named id |
Emitted member order (under §10.1): k, v, id, gen, lsn, op, rk, rid, rec, then other members.
Rules the member table cannot express:
- One file holds one
gen. LSNs restart whengenchanges, so two generations in one file would shareids and merge would drop records. - LSNs within a
genare gap-free and are applied in LSN order. Becausemutis opaque, it sorts after every typed kind, byid, which is LSN order. - A mutation log is ordered, not a set. Producers MUST NOT combine logs with merge (§8).
rk/rididentify records, neverid: an opaque record's identity readsidfirst (§5.3), so an entity id there would collapse a put and a delete of the same record.
For every record:
- kind = the
kstring. - rank = §5.1 rank when
kis a typed kind string (this includes an opaque record whosekis a typed kind at a higherv), otherwise the maximum rank, placed after all typed kinds. - primary id = §5.1 for typed records. For an opaque record: the value of the first of
id,ref_id,session_id,answer_id,clause_idthat is present as a JSON string; if none is, the identity form (§10.3) of the whole object. - sort key =
(rank, kind, primary id), compared rank numerically, then kind and primary id by UTF-8 byte order. - dedup key =
(kind, primary id).
A writer MUST emit records in ascending sort key order. Records with equal sort keys keep their input order (the sort is stable). A writer SHOULD NOT emit two records with the same dedup key; readers MUST accept such files, and diff reports them (§8.2).
Note: an opaque record whose k is a typed kind at a higher v has the same dedup key as a v1 record with the same primary id. Merge treats the two as the same record (§8.1).
embref.target, session.trajectory, codebook_manifest.target and promo.curvature_refs[] are one of two JSON shapes:
{"local":{"section_index":S,"record_index":R}}: recordRof sectionSof this file's own appendix. target-key ="local|" pad10(S) "|" pad20(R).{"ext":{"ref_id":F,"record_index":R}}: recordRof the section named by thecsdt_refrecord whoseref_idisF.Rcounts from the section start, not from anyrecord_range. target-key ="ext|" F "|" pad20(R).
(pad10 pads to 10 digits.)
The digest D is SHA-256 over the file bytes from offset 257 up to, not including, offset L - 96. That range is, in order: the body, the marker line (when present), the zero padding (when present), and the embedded container (when present). The header line and the footer are excluded, so the header can be rewritten in place after streaming without changing the digest.
D is stored raw in footer bytes 64..96 and, in a final header, as 64 lowercase hex digits. A verifier MUST compare the computed digest with the footer (checksum_mismatch) and, for a final header, the header digest with the footer digest (header_footer_mismatch).
The embedded container carries its own CRC32C checks (§7.3); they are in addition to D, not a replacement.
The appendix is a complete, standalone CSDT container, byte for byte as the CSDT library writes it. PLXI defines no binary payload encoding of its own. The container is specified by the Cintilé container format [CSDT]; this section states only what a PLXI implementation needs.
- Every offset inside the container (section table offset, section data offsets) is relative to the container's first byte,
csdt_offset, not to the.plxifile. - Because
csdt_offsetis a multiple of 64 and the container's own alignment is 64, every aligned view the container promises is aligned in the file as well. - The container header is 128 bytes. Bytes 0 to 3 are
CSDT. Byte 4 is the generation discriminator for every header layout.
Not part of this publication.
Not part of this publication.
Not part of this publication.
Merge takes two packs a and b and a strategy, and produces one pack.
Preconditions:
- Each input MUST pass verify steps 1 to 8 (§9) before merging.
- Each input appendix MUST have container version byte 5 (§7.2), else
appendix_legacy_version. - An input appendix holding a tensor catalog section MUST be rejected with
merge_conflict: tensor catalogs encode section indexes internally and cannot be re-indexed. - An input section whose
section_typeis not in the CSDT registry MUST be rejected withmerge_conflict.
Appendix layer:
- Lift every section of
aandbas the tuple(section_type, format_version, payload_class, flags, record_count, record_stride, dtype bytes, align_log2, data bytes). - The output pool is all lifted tuples, sorted ascending by that tuple in that field order (integers numerically, byte strings lexicographically), with exact duplicates removed.
sections_dedupedcounts the removed duplicates. - Each input section's new index is its position in the pool. Every
localbinding target (§5.4) in that input's records is rewritten to the new section index;record_indexis unchanged, because sections move whole. - The pool is written as a new version-5 container. An empty pool means no appendix.
Record layer, applied to a's records and then b's, one record at a time, keyed by dedup key (§5.3):
payloadrecords from the inputs are dropped; one newpayloadrecord describing the output appendix is added at the end. Its members areembedded(true when the output has an appendix),csdt_file_checksum(the output container'sfile_checksum, else 0),sections(one{index, section_type, record_count}per pool entry, withoutdtype) andrefs(theref_idof everycsdt_refrecord in the output). All four are emitted even when empty, so an empty pool with nocsdt_refgives{"k":"payload","v":1,"embedded":false,"csdt_file_checksum":0,"sections":[],"refs":[]}.- New key: insert (
records_inserted). - Existing record equal to the incoming one (value equality after parsing): keep, count
records_skipped. - Both
ent:- existing is a stub and incoming is not: take incoming, count
stubs_resolvedandrecords_updated; - incoming is a stub and existing is not: keep existing (
records_skipped); - otherwise a conflict (
conflicts), resolved by the strategy:higher_saliencetakes incoming only if itssalienceis strictly greater (missing counts as 0.0);latesttakes incoming;unionkeeps existing;manualkeeps existing and records the conflict as deferred. A conflict countsconflicts; when the strategy takes the incoming record it also countsrecords_updated, and when it keeps the existing one,records_skipped.
- existing is a stub and incoming is not: take incoming, count
- Any other kind:
latesttakes incoming (records_updated); every other strategy keeps existing (records_skipped). - The output records are written in sort key order (§5.3) with the rules of §2 to §6.
Properties:
- For inputs whose shared dedup keys carry equal records,
merge(a, b)andmerge(b, a)produce byte-identical files. With conflicts, the result depends on argument order underlatestandhigher_salienceties. - Merge treats
mutrecords like any opaque record (dedup by the padded LSN inid). Combining two mutation logs with merge is a producer error (§5.2.1), not something merge detects. - The merge report keys are
records_inserted,records_updated,records_skipped,stubs_resolved,conflicts,sections_merged,sections_deduped.
Diff compares two packs a (before) and b (after).
- Index each side by dedup key (§5.3). A key that occurs more than once on either side is listed in
duplicate_keysand is not otherwise compared. added: keys only inb.removed: keys only ina.changed: keys in both whose records differ in identity form (§10.3). Each entry carries the key and both record lines.- Appendix: sections are matched by section index; a section is
added,removed, orchangedwhen itssection_typeor its descriptor CRC32C differs. - Every output list is sorted by sort key (records) or by section index (sections).
Diff never reports a difference that exists only in member order, whitespace or number spelling, because it compares identity forms.
Shard produces a new pack holding a connected subset of one pack's graph. The Sharder conformance level is provisional in revision 6.0; this section is its contract.
Graph view of a pack: nodes are ent records by id; directed edges are rel records src → tgt labelled kind; an entity has an embedding when an embref record names it.
Shard specs:
entity_centric{seed, max_depth, max_fanout?, relationship_filter?, direction, phase_coherent}: breadth-first fromseedtomax_depthhops, following edges indirection(outgoing,incoming,both), at mostmax_fanoutneighbours per node when set, and only edges whosekindis inrelationship_filterwhen set.search_result{entity_ids, context_depth, max_fanout?}: the listed entities plus their neighbourhood tocontext_depth.predicate{...}: entities whose fields satisfy the predicate.cluster{cluster_id, include_bridges}: MUST fail withshard_unsupported_specuntil a cluster field exists in the record layer. It MUST NOT return an empty result instead.union[specs]: the union of the member results.
Output pack:
- the selected
entrecords, plus a stubent(stub: true) for each edge endpoint outside the selection when the config asks for stubs; - the
relrecords whose endpoints are both in the output; - a
hyperrecord only when every member is selected; embrefrecords of selected entities, with the appendix rewritten to hold only the referenced rows and the bindings remapped;- every other appendix section listed in
dropped_sections, never dropped silently; - a
metarecord whosecreated_atis supplied by the caller. Two runs with the same input, spec andcreated_atMUST produce byte-identical packs.
A shard whose quality score is below a configured minimum fails with shard_quality_below_threshold.
verify checks a pack and returns a report with the keys sha256_ok, header_footer_agree, record_count{declared, actual}, appendix{present, file_checksum_ok, sections[{index, type, crc_ok}]}, embrefs{total, resolved, dangling, type_mismatch}. Opening a pack runs steps 1 to 8 by default; the caller may opt out explicitly.
Input: the bytes of a .plxi file. If they start with 1F 8B (gzip ID1, ID2 [RFC1952]), a reader MAY decompress them and continue with the result; .plxi.gz is a transport encoding only (§11).
- Size.
L >= 353, elseinvalid_header. - Header. Bytes 0..257 are ASCII, byte 256 is LF, and bytes 0..256 match
header-content padding(§3.1): elseinvalid_header. A version other than 6:unsupported_version. - Footer. Bytes
L-96 .. Lper the §10.0 table: magicPLXF, elseinvalid_footer; version 6, elseunsupported_version; reserved bytes 44..64 all zero, elseinvalid_footer. - Header against footer. Classify the header (§3.3). Final:
records,csdt_offset,csdt_size, digest equal the footer's, elseheader_footer_mismatch. Placeholder: continue with the footer's values. Neither:header_footer_mismatch. - Layout. With
T = text_section_size,A = csdt_offset,C = csdt_size, all integers from the footer, compared without overflow:257 <= T <= L - 96, elseinvalid_footer;C = 0:A = 0andT = L - 96, elseinvalid_footer;C > 0:A mod 64 = 0, elseappendix_alignment;A = Trounded up to a multiple of 64 andA + C = L - 96, elseinvalid_footer; the padding bytesT .. Aare all0x00, elseinvalid_footer.- The checks above are made in 64-bit arithmetic, before any conversion to an in-memory index, so a layout that fails them is
invalid_footeron every platform. Only a layout that passes them and still holds an offset or size that cannot be represented as an in-memory index on the platform (for example at or above 2^32 on a 32-bit target) fails, withlimit_exceeded; it MUST NOT be truncated.
- Digest. SHA-256 of bytes
257 .. L-96equals footer bytes 64..96, elsechecksum_mismatch. - Body. Bytes
257 .. Tare valid UTF-8, elseinvalid_utf8. WhenC > 0, the last line is the marker line, elseinvalid_record. Every other non-empty line parses as a record (§4), elsejsonorinvalid_record. A marker line elsewhere:invalid_record. - Count. The number of record lines equals the footer's
record_count, elserecord_count_mismatch. - Appendix (Reader-Appendix, only when
C > 0): open the container at bytesA .. A+Cwith the checks of §7.3; check every section's CRC32C; compare the footer'scsdt_file_checksumwith the container'sfile_checksum. - Bindings (Reader-Appendix): for each
localtarget (§5.4),section_indexis below the section count andrecord_indexis below that section'srecord_count; for eachexttarget, acsdt_refrecord with thatref_idexists in the pack. Each failure counts as dangling. Verify reports counts; a consumer that binds a dangling target fails withdangling_embref.- Type check under the Compact embedding profile (the entity-embedding binding that importers use): a
localtarget of anembrefis expected to name a section of typeCompact(0x0020) whoserecord_strideequals the Compact record size (320 bytes), with dtypeRECORD. Verify counts mismatches intype_mismatch; a consumer that binds one fails withsection_type_mismatch. The profile belongs to the consumer, not to the file format: other kinds of section are legalembreftargets.
- Type check under the Compact embedding profile (the entity-embedding binding that importers use): a
- 1Sizeinvalid_header
- 2Headerinvalid_headerunsupported_version
- 3Footerinvalid_footerunsupported_version
- 4Header against footerheader_footer_mismatch
- 5Layoutinvalid_footerappendix_alignmentlimit_exceeded
- 6Digestchecksum_mismatch
- 7Bodyinvalid_utf8invalid_recordjson
- 8Countrecord_count_mismatch
- 9Appendixappendix_invalidappendix_crc_mismatchonly when the pack has an appendix
- 10Bindingsdangling_embrefsection_type_mismatchcounts only; a consumer that binds a failing target reports
dangling_embreforsection_type_mismatch
Steps run in this order. A verifier that stops at the first failure MUST report the kind of that first failure, so every conforming implementation reports the same kind for the same file. Reports for steps it did not reach are absent.
| Offset | Size | Field | Type | Rule |
|---|---|---|---|---|
| 0 | 4 | magic | bytes | 50 4C 58 46 (PLXF) |
| 4 | 4 | version | u32 | 6 |
| 8 | 8 | record_count | u64 | number of record lines (§4.1) |
| 16 | 8 | text_section_size | u64 | T: absolute offset of the first byte after the text section: header line, body and, when an appendix exists, the marker line (§2) |
| 24 | 8 | csdt_offset | u64 | A (0 when no appendix) |
| 32 | 8 | csdt_size | u64 | C (0 when no appendix) |
| 40 | 4 | csdt_file_checksum | u32 | copy of the embedded container's file_checksum (CRC32C over its section table); 0 when no appendix |
| 44 | 20 | reserved | bytes | MUST be written as zero; readers MUST reject nonzero (invalid_footer) |
| 64 | 32 | sha256 | bytes | raw digest D (§6) |
The footer is the authoritative source for every count and offset, because it is always final (§3.3).
- A writer MUST emit typed records in the member order of §4.3 and opaque records, and every nested JSON object inside any record, with members in the order they were read or inserted.
- An implementation MUST NOT reorder members as a side effect of its build configuration. The reference implementation enables the
preserve_orderfeature ofserde_jsonin every build, so its default build and its--no-default-featuresbuild emit the same bytes. - A writer MUST emit a JSON text without insignificant whitespace, one record per line.
When a value is re-emitted, a writer MUST use these spellings (they are serde_json 1.0.149's, measured in §10.4):
- Strings:
"and\are escaped as\"and\\; U+0008, U+0009, U+000A, U+000C, U+000D as\b,\t,\n,\f,\r; other code points below U+0020 as\u00XXwith lowercase hex; everything else, including/, U+007F and non-ASCII, is written as raw UTF-8. A JSON string that decodes to a lone surrogate MUST be rejected when read (json). - Numbers written in the input as an integer token (no fraction, no exponent) whose value fits a signed or unsigned 64-bit integer are written as that integer, exactly.
- Every other number is an IEEE 754 binary64 value, written with the shortest digit string that reads back to the same value:
- when
1e-5 <= |x| < 1e16, in positional form, with.0appended when the value is integral (100.0,0.00001,1000000000000000.0); - otherwise in exponent form
d[.ddd]e+Nord[.ddd]e-N(1e+16,1e-6,1.8446744073709552e+19); - negative zero is written
-0.0; - when more than one shortest digit string reads back to the same value, the one serde_json 1.0.149 writes is the spelling. For example, the binary64 value
0x42e87faaebb9a0d4reads back from both215492859907334.62and215492859907334.63; serde_json writes215492859907334.62, so that is the spelling, and215492859907334.63(the Rust standard library's{}and{:e}digits) does not conform.
- when
The corpus carries a numbers case that pins these spellings; an implementation whose number formatter differs fails it. A second case, pos-numbers-d96, pins the tie rule above and exact reading: its input spells 0x42e87faaebb9a0d4 as 215492859907334.63 and as 215492859907334.62, both re-emitted as 215492859907334.62, and carries the literal 2.1549285990733466e14, which is the different value 0x42e87faaebb9a0d5 and is re-emitted as 215492859907334.66. Identity forms, dedup keys and conformance hashes use these spellings (§10.3).
When a number token is read, a reader MUST return the binary64 value nearest to the token's decimal value (round half to even). The reference reader is serde_json 1.0.149 with its float_roundtrip feature on. Together with the shortest-spelling rule above, this makes canonical JSON a fixed point: reading an emitted line or an identity form and emitting it again yields the same bytes. Rationale: serde_json's default reader can return a value one unit in the last place away from the nearest one; with it, 1,198 of a 20,001-value sweep changed spelling when read and re-emitted, and one value moved by one unit per cycle for nine cycles.
Dedup keys that fall back to a whole object, diff equality, conformance goldens and SHA comparisons between records use the PLXI identity form of a JSON value:
- object members sorted recursively by member name, comparing names as UTF-8 byte strings;
- array element order unchanged;
- no insignificant whitespace;
- strings and numbers spelled per §10.2.
The identity form of a record line is the identity form of the JSON value that its re-emitted line (§4.3, §4.4) parses to.
The identity form is not the JSON Canonicalization Scheme [RFC8785]. The differences, measured against serde_json = 1.0.149, without and with preserve_order:
| Input | PLXI identity form | RFC 8785 | Rule that differs |
|---|---|---|---|
member names "דּ" and "😀" |
U+FB33 first (UTF-8 EF.. < F0..) |
U+1F600 first (UTF-16 D83D < FB33) |
key order: UTF-8 bytes vs UTF-16 code units (RFC 8785 §3.2.3) |
1.0, 100.0, 1E2 |
1.0, 100.0, 100.0 |
1, 100, 100 |
integral binary64 keeps .0 |
-0.0 |
-0.0 |
0 |
negative zero |
18446744073709551615 |
18446744073709551615 |
18446744073709552000 |
64-bit integers stay exact |
1e-6 |
1e-6 |
0.000001 |
positional range lower bound (1e-5 here, 1e-6 in ECMAScript) |
1e16, 1e17 |
1e+16, 1e+17 |
10000000000000000, 100000000000000000 |
positional range upper bound (1e16 here, 1e21 in ECMAScript) |
18446744073709551616 |
1.8446744073709552e+19 |
18446744073709552000 |
same binary64 value, different spelling |
Strings, literals and the absence of whitespace agree in every case probed. Implementations MUST NOT substitute a JCS library for the identity form.
A .plxi.gz file is a gzip member [RFC1952] whose decompressed bytes are a complete .plxi file. It exists for shipping; a reader MUST decompress it fully before reading the appendix, and memory-mapped, zero-copy access is not available through it. A writer that streams into gzip cannot seek, so it leaves a placeholder header (§3.3).
The file format version is 6, stored in the header's version token and the footer's version field. Readers MUST reject every other value (unsupported_version). A change that an unmodified version-6 reader would misread requires version 7. There is no in-place compatibility path between file versions.
Each kind versions independently through v (§4.2).
- Adding an OPTIONAL member keeps
v; older readers preserve it inextra. - Removing a member, making one REQUIRED, or changing a member's type or meaning raises
v. Older readers then keep the record opaque (§4.4), and it still round-trips through them. - Adding a kind needs no version change; older readers keep it opaque.
This document is revision 6.0. A revision that only clarifies or adds OPTIONAL members, kinds or error kinds is 6.x. Report keys and error kind strings version together: adding one is a minor change with new conformance cases; removing one is a major change.
Not part of this publication.
An implementation claims one or more levels. Each level is proved by the corpus groups listed; the case ids are listed on the cases page.
| Level | Requirements | Corpus groups |
|---|---|---|
| Reader-Core | §2, §3, §4, §5.3, §6, §9 steps 1-8, §10.0-10.3 | positive, graph (no-appendix pack, empty pack, .plxi.gz), negative (framing, digest, count, UTF-8, header byte 80) |
| Reader-Appendix | Reader-Core plus §7, §9 steps 9-10 | graph (Compact appendix with F32/BF16/MXFP4/B2048 sections), negative (container offsets, align_log2, wrong section type, wrong stride, appendix CRC flip) |
| Writer | emits files that pass Reader-Core and Reader-Appendix verify; §4.3, §4.4, §5.3 order, §10.1, §10.2; container version byte 5 only | positive goldens (.golden.jsonl), numbers case |
| Merger | §8.1 | merge (one per strategy, the id dedup trap, appendix section union), negative (an appendix with legacy container version byte 3) |
| Differ | §8.2 | diff |
| Sharder (provisional) | §8.3 | shard |
The corpus has limits an implementation should know: it tests the structure and checksums of the embedded container, not the meaning of the data inside it.
Not part of this publication.
- [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, March 1997. https://www.rfc-editor.org/rfc/rfc2119
- [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, May 2017. https://www.rfc-editor.org/rfc/rfc8174
- [RFC5234] Crocker, D., Ed., and P. Overell, "Augmented BNF for Syntax Specifications: ABNF", STD 68, RFC 5234, January 2008. https://www.rfc-editor.org/rfc/rfc5234
- [RFC8259] Bray, T., Ed., "The JavaScript Object Notation (JSON) Data Interchange Format", RFC 8259, December 2017. https://www.rfc-editor.org/rfc/rfc8259
- [RFC3629] Yergeau, F., "UTF-8, a transformation format of ISO 10646", STD 63, RFC 3629, November 2003. https://www.rfc-editor.org/rfc/rfc3629
- [RFC8785] Rundgren, A., Jordan, B., and S. Erdtman, "JSON Canonicalization Scheme (JCS)", RFC 8785, June 2020. https://www.rfc-editor.org/rfc/rfc8785
- [RFC1952] Deutsch, P., "GZIP file format specification version 4.3", RFC 1952, May 1996. https://www.rfc-editor.org/rfc/rfc1952
- [RFC3720] Satran, J., et al., "Internet Small Computer Systems Interface (iSCSI)", RFC 3720, April 2004 (CRC32C, Castagnoli polynomial). https://www.rfc-editor.org/rfc/rfc3720
- [FIPS180-4] National Institute of Standards and Technology, "Secure Hash Standard (SHS)", FIPS PUB 180-4, August 2015. https://nvlpubs.nist.gov/nistpubs/FIPS/NIST.FIPS.180-4.pdf
- [CSDT] Cintilé, "Cintilé container format (CSDT)", CSDT v3, container version byte 5. Not published.
© 2020-2026 Cintile Inc. All Rights Reserved.
Anyone may implement this format. Copying or republishing the text of this specification requires permission from Cintile Inc.