Specification

PLXI specification · revision 6.0 · file format 6 · plxi.org

All sections of revision 6.0 on one page. Ids are the same as on the section pages.

PLXI file format, version 6: specification

Document PLXI file format, version 6: specification
Revision 6.0 (2026-10-01; amended 2026-10-06)
Applies to file format 6

1. Scope and conventions

This document defines the PLXI version 6 file format (.plxi), its transport encoding (.plxi.gz), the record layer carried in its text body, and the behaviour of the operations defined on it: verify, merge, diff and shard. It is written so that a reader or writer can be implemented in any language without reading Rust.

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.

Other conventions:

  • Byte offsets are zero-based. "Absolute offset" means an offset from the first byte of the .plxi file. L is the file length in bytes.
  • A range a..b of bytes or offsets starts at a and ends before b: bytes 0..4 are bytes 0, 1, 2 and 3.
  • All multi-byte integers in binary structures are little-endian and unsigned unless stated otherwise.
  • Grammar is written in ABNF [RFC5234]. ABNF quoted strings are case-insensitive, so every case-sensitive literal below is written with %x byte values.
  • "JSON" means RFC 8259 JSON [RFC8259]. "UTF-8" is as defined in RFC 3629 [RFC3629].
  • SHA-256 is as defined in FIPS 180-4 [FIPS180-4]. CRC32C is the Castagnoli CRC, polynomial 0x1EDC6F41 [RFC3720].

2. File layout

A .plxi file is a text section followed, optionally, by a binary appendix, followed by a fixed binary footer:

offset 0          header line: 256 bytes of header content padded with
                  spaces, then LF (257 bytes)
offset 257        body: zero or more JSONL record lines, each ending in LF
                  marker line "===PLXI_BINARY_APPENDIX===" LF
                  (only when an appendix exists)
T                 = text_section_size: first byte after the text section
                  padding: 0x00 bytes up to the next multiple of 64
                  (only when an appendix exists)
A                 = csdt_offset: embedded CSDT container, csdt_size bytes
                  (only when an appendix exists)
L - 96            footer: 96 bytes
L                 end of file
Two strips showing the parts of a pack in file order.First strip, a pack with an appendix: header line, 257 bytes; records, one JSON record per line; marker line; padding with zero bytes up to a multiple of 64; appendix, csdt_size bytes; footer, 96 bytes. Offsets under the strip: 0, 257, T, A, L − 96, L. A bracket over records, marker line, padding and appendix reads: the digest covers these bytes, 257 up to L − 96. Second strip, a pack without an appendix: header line, records, footer; T equals L − 96.The digest covers these bytes: 257 up to L − 96header line257 bytesrecordsone JSON recordper line ⋯marker lineonly with anappendixpaddingzero bytes up toa multiple of 64appendixcsdt_sizebytesfooter96 bytes0257TAL − 96LWithout an appendix: T = L − 96header line257 bytesrecordsone JSON recordper line ⋯footer96 bytes0257L − 96Lsolid: always presentdashed: only with an appendix
Figure 2-1. The parts of a pack, in file order. Not to scale; sizes are written in the cells. A dashed outline is present only when the pack has an appendix.First strip, a pack with an appendix: header line, 257 bytes; records, one JSON record per line; marker line; padding with zero bytes up to a multiple of 64; appendix, csdt_size bytes; footer, 96 bytes. Offsets under the strip: 0, 257, T, A, L − 96, L. A bracket over records, marker line, padding and appendix reads: the digest covers these bytes, 257 up to L − 96. Second strip, a pack without an appendix: header line, records, footer; T equals L − 96.

Rules:

  1. The header line MUST be exactly 257 bytes: 256 bytes of header content and padding (§3) followed by one LF (0x0A) at offset 256. The body therefore always starts at offset 257.
  2. The footer MUST occupy the last 96 bytes of the file. A file shorter than 257 + 96 = 353 bytes is not a PLXI v6 file (invalid_header).
  3. Without an appendix, the footer MUST follow the last body line directly: T = L - 96, csdt_offset = 0, csdt_size = 0, and no marker line is present.
  4. With an appendix (csdt_size > 0):
    • The marker line ===PLXI_BINARY_APPENDIX=== followed by LF MUST follow the last record line. T is the offset of the byte after that LF.
    • Padding of (64 - T mod 64) mod 64 bytes, each 0x00, MUST follow, so that csdt_offset = T rounded up to a multiple of 64.
    • The embedded container MUST start at csdt_offset and MUST end at L - 96: csdt_offset + csdt_size = L - 96.
  5. csdt_offset MUST be 0 whenever csdt_size is 0.

The 64-byte placement exists so that the embedded CSDT container's own 64-byte alignment guarantees hold when the whole .plxi file is memory-mapped at a page-aligned base.

3. Header line

3.1 Grammar

header-line     = header-content padding LF        ; exactly 257 octets
header-content  = magic SP version SP records SP csdt-offset SP csdt-size SP
                  sha-field
magic           = %x50.4C.58.49                    ; "PLXI"
version         = %x76 dec
                  ; "v" then the format version, "v6"
records         = %x72.65.63.6F.72.64.73.3D dec    ; "records=" N
csdt-offset     = %x63.73.64.74.5F.6F.66.66.73.65.74.3D dec
                  ; "csdt_offset=" O
csdt-size       = %x63.73.64.74.5F.73.69.7A.65.3D dec
                  ; "csdt_size=" C
sha-field       = %x73.68.61.32.35.36.3D [sha-hex]
                  ; "sha256=" H or "sha256=" (placeholder)
sha-hex         = 64lhex
lhex            = %x30-39 / %x61-66
                  ; lowercase hexadecimal only
dec             = %x30 / (%x31-39 *19%x30-39)
                  ; decimal, no sign, no leading zeros, value < 2^64
padding         = *SP
                  ; space-fill so header-content + padding = 256 octets
SP              = %x20
LF              = %x0A

Each token is separated from the next by exactly one SP. There are exactly six tokens.

3.2 Semantics

  • version MUST be 6. A header whose version token parses as any other number MUST be rejected with unsupported_version. Versions 1 to 5 are retired numbering systems and have no readers.
  • records is the number of record lines in the body (§4).
  • csdt-offset and csdt-size locate the appendix (§2). When csdt-size is nonzero, csdt-offset MUST be a multiple of 64.
  • sha-field carries the lowercase hex SHA-256 digest defined in §6.

3.3 Final and placeholder headers

A writer first emits a placeholder header and, when its output is seekable, overwrites it in place with the final header after the body and appendix are written. The fixed 256-byte width makes the rewrite possible without moving any other byte.

  • A final header MUST carry a 64-digit sha-hex, and its records, csdt-offset, csdt-size and digest MUST equal the footer's record_count, csdt_offset, csdt_size and sha256.
  • A placeholder header is exactly PLXI v6 records=0 csdt_offset=0 csdt_size=0 sha256= (empty digest) padded to 256 bytes. A writer whose output cannot seek (a gzip stream) leaves it in place. For a file with a placeholder header, the footer is the only source of the counts, offsets and digest.
  • A header that is neither final nor placeholder (for example an empty digest with a nonzero records) MUST be rejected with header_footer_mismatch.

The header line is not covered by the SHA-256 digest (§6). Everything a reader trusts from it is therefore cross-checked against the footer.

4. Body: JSONL records

4.1 Lines

  • The body is a sequence of lines, each one JSON object [RFC8259] encoded as UTF-8 and terminated by one LF.
  • A writer MUST NOT emit empty lines, CR characters outside JSON strings, or a line containing a raw LF inside a JSON value (JSON strings escape LF as \n).
  • A reader MUST reject a body that is not valid UTF-8 with invalid_utf8.
  • A reader MUST skip empty lines and MUST strip one trailing CR before parsing a line. Skipped lines do not count as records.
  • The marker line (§2) is not a record. It MUST appear only as the last line of the text section and only when an appendix exists; anywhere else it is invalid_record.
  • records in the header and record_count in the footer count the record lines.
  • A writer MUST NOT emit a JSON object with duplicate member names. A reader that meets one MUST either reject the line with json, or keep the last value for the name; the reference reader keeps the last value (serde_json behaviour, measured in §10.4). Conformance tests do not exercise duplicates.

4.2 The record envelope

Every record line is a JSON object with:

  • k (REQUIRED): a JSON string, the record kind. A line without a string k MUST be rejected with invalid_record.
  • v (OPTIONAL): a non-negative JSON integer, the per-kind format version. Absent means 1.
  • the kind's own members (§5).

A record is typed when its kind is in the typed registry (§5.1) and its v is at most the highest version this reader supports for that kind. Every typed kind is at version 1. Otherwise the record is opaque.

4.3 Typed records

  • A reader MUST parse a typed record against its schema (§5.1). A schema failure (missing REQUIRED member, wrong JSON type) MUST be reported as json.
  • Members the schema does not name MUST be preserved in the record's extra map and re-emitted after the declared members.
  • A writer MUST emit a typed record as: k first, then v, then the declared members in schema order, then extra members in the order they were read. OPTIONAL members whose value is absent are omitted, except those marked "always emitted" in §5.1.
  • Re-emitting a typed record normalises its values: a 32-bit float member (for example salience, confidence) is re-emitted as the shortest decimal that round-trips the 32-bit value, and whitespace, escapes and number spelling follow §10.2. A typed round trip therefore preserves the value but MAY change the bytes.

4.4 Opaque records (forward compatibility)

A record MUST be kept opaque, not rejected, when either:

  1. its kind is not in the typed registry, including the reserved kinds of §5.2 and any future kind; or
  2. its kind is typed but its v is higher than this reader supports.

For an opaque record, a reader MUST keep the whole JSON object, including k and v, and a writer MUST re-emit it with the same members, the same values, and the members in the order they were read (§10.1). This is the opaque round-trip rule. It guarantees JSON value equality and member order; it does not guarantee byte equality, because escapes and number spellings are re-emitted in the form of §10.2 (a JSON integer outside the 64-bit range, for example, becomes a binary64 number; §10.4).

Opaque records take part in sorting, merge and diff through the identity rules of §5.3.

5. Kind registry

5.1 Typed kinds (19)

The rank is the canonical sort position. The primary id is the record's identity within its kind; pad20(n) means n in decimal, left-padded with 0 to 20 digits. The field list is the v1 schema; ? marks an OPTIONAL member, all others are REQUIRED. Every kind also carries v? and extra.

Table 5-1. Typed kinds (19)
Rank k Primary id v1 members
0 meta created_at "|" source created_at, source, tool?, description?, ngdb_generation? (u64)
1 payload empty string embedded (bool), csdt_file_checksum? (u32, always emitted, default 0), sections? (array of {index u32, section_type u16, record_count u64, dtype? string}, always emitted, default []), refs? (array of string, always emitted, default [])
2 csdt_ref ref_id ref_id, path, shard_set_id? (u32), file_checksum (u32), source_hash, section_type (u16), section_index (u32), record_range? ([u64, u64])
3 ent id id, t, name?, salience? (f32), stub? (bool, always emitted, default false), properties?, prov?
4 rel src "|" tgt "|" kind src, tgt, kind, properties?, prov?
5 hyper id id, members (array of string), kind, properties?
6 embref entity_id "|" target-key entity_id, target (§5.4), annotations?
7 prov subject "|" method "|" pad20(unix_secs) subject, agent, method, unix_secs (u64), inputs?, notes?
8 grounded_answer answer_id Application kind.
9 citation source_answer_id "|" entity_id "|" pad20(span.start) Application kind.
10 clause clause_id Application kind.
11 knob_set scope "|" key "|" pad20(unix_secs) Application kind.
12 discovery gid Application kind.
13 promo content_hash "|" pad20(order) Application kind.
14 codebook_manifest codebook_id Application kind.
15 session session_id Application kind.
16 sev pad20(step) "|" t Application kind.
17 session_end pad20(end_unix) Application kind.
18 outcome identity form (§10.3) of outcome Application kind.

Members without a stated type are JSON strings; u32, u64 are non-negative JSON integers in range.

Changing a typed kind's schema incompatibly requires raising its v (§12.2). Adding an OPTIONAL member at the same v is compatible, because older readers keep it in extra.

5.2 Reserved kinds

These kind strings are reserved, but have no typed parser in this revision. Readers MUST treat them as opaque (§4.4).

mut

Ten further kind names are reserved and are specified separately. A reader keeps any record whose kind it does not know as an opaque record (§4.4).

A reserved kind is promoted to the typed registry only by a spec revision that adds it to §5.1 with the same v1 schema.

Kind names that begin with x- belong to applications: no revision of this specification adds one to the typed kinds (§5.1) or the reserved kinds (§5.2).

5.2.1 mut v1 (frozen)

mut is one entry of an ordered mutation log. The schema below is the wire form of mut v1.

Table 5-2. Members of a `mut` record, version 1 (10)
Member Type Rule
k string "mut"
v integer 1
id string the LSN in decimal, left-padded with 0 to 20 digits. MUST equal pad20(lsn)
gen u64 publication generation
lsn u64 log sequence number within gen
op string one of put_record, delete_record, or one of the op names reserved for future use
rk string, OPTIONAL kind of the affected record. REQUIRED for put_record and delete_record
rid string, OPTIONAL primary id (§5.1) of the affected record. REQUIRED for put_record and delete_record
rec object, OPTIONAL put_record only: the complete affected record as its own JSONL object, including its k. Its kind and primary id MUST equal rk and rid. delete_record MUST NOT carry rec
other any members of the reserved ops, preserved. None of them may be named id

Emitted member order (under §10.1): k, v, id, gen, lsn, op, rk, rid, rec, then other members.

Rules the member table cannot express:

  • One file holds one gen. LSNs restart when gen changes, so two generations in one file would share ids and merge would drop records.
  • LSNs within a gen are gap-free and are applied in LSN order. Because mut is opaque, it sorts after every typed kind, by id, which is LSN order.
  • A mutation log is ordered, not a set. Producers MUST NOT combine logs with merge (§8).
  • rk/rid identify records, never id: an opaque record's identity reads id first (§5.3), so an entity id there would collapse a put and a delete of the same record.

5.3 Identity, sort order and dedup key

For every record:

  • kind = the k string.
  • rank = §5.1 rank when k is a typed kind string (this includes an opaque record whose k is a typed kind at a higher v), otherwise the maximum rank, placed after all typed kinds.
  • primary id = §5.1 for typed records. For an opaque record: the value of the first of id, ref_id, session_id, answer_id, clause_id that is present as a JSON string; if none is, the identity form (§10.3) of the whole object.
  • sort key = (rank, kind, primary id), compared rank numerically, then kind and primary id by UTF-8 byte order.
  • dedup key = (kind, primary id).

A writer MUST emit records in ascending sort key order. Records with equal sort keys keep their input order (the sort is stable). A writer SHOULD NOT emit two records with the same dedup key; readers MUST accept such files, and diff reports them (§8.2).

Note: an opaque record whose k is a typed kind at a higher v has the same dedup key as a v1 record with the same primary id. Merge treats the two as the same record (§8.1).

5.4 Payload binding targets

embref.target, session.trajectory, codebook_manifest.target and promo.curvature_refs[] are one of two JSON shapes:

  • {"local":{"section_index":S,"record_index":R}}: record R of section S of this file's own appendix. target-key = "local|" pad10(S) "|" pad20(R).
  • {"ext":{"ref_id":F,"record_index":R}}: record R of the section named by the csdt_ref record whose ref_id is F. R counts from the section start, not from any record_range. target-key = "ext|" F "|" pad20(R).

(pad10 pads to 10 digits.)

6. Integrity: SHA-256 coverage

The digest D is SHA-256 over the file bytes from offset 257 up to, not including, offset L - 96. That range is, in order: the body, the marker line (when present), the zero padding (when present), and the embedded container (when present). The header line and the footer are excluded, so the header can be rewritten in place after streaming without changing the digest.

D is stored raw in footer bytes 64..96 and, in a final header, as 64 lowercase hex digits. A verifier MUST compare the computed digest with the footer (checksum_mismatch) and, for a final header, the header digest with the footer digest (header_footer_mismatch).

The embedded container carries its own CRC32C checks (§7.3); they are in addition to D, not a replacement.

7. Appendix: embedded CSDT container

7.1 What it is

The appendix is a complete, standalone CSDT container, byte for byte as the CSDT library writes it. PLXI defines no binary payload encoding of its own. The container is specified by the Cintilé container format [CSDT]; this section states only what a PLXI implementation needs.

  • Every offset inside the container (section table offset, section data offsets) is relative to the container's first byte, csdt_offset, not to the .plxi file.
  • Because csdt_offset is a multiple of 64 and the container's own alignment is 64, every aligned view the container promises is aligned in the file as well.
  • The container header is 128 bytes. Bytes 0 to 3 are CSDT. Byte 4 is the generation discriminator for every header layout.

7.2 Version bytes

Not part of this publication.

7.3 Container checks

Not part of this publication.

7.4 Section descriptor fields PLXI uses

Not part of this publication.

8. Operations

8.1 Merge

Merge takes two packs a and b and a strategy, and produces one pack.

Preconditions:

  1. Each input MUST pass verify steps 1 to 8 (§9) before merging.
  2. Each input appendix MUST have container version byte 5 (§7.2), else appendix_legacy_version.
  3. An input appendix holding a tensor catalog section MUST be rejected with merge_conflict: tensor catalogs encode section indexes internally and cannot be re-indexed.
  4. An input section whose section_type is not in the CSDT registry MUST be rejected with merge_conflict.

Appendix layer:

  1. Lift every section of a and b as the tuple (section_type, format_version, payload_class, flags, record_count, record_stride, dtype bytes, align_log2, data bytes).
  2. The output pool is all lifted tuples, sorted ascending by that tuple in that field order (integers numerically, byte strings lexicographically), with exact duplicates removed. sections_deduped counts the removed duplicates.
  3. Each input section's new index is its position in the pool. Every local binding target (§5.4) in that input's records is rewritten to the new section index; record_index is unchanged, because sections move whole.
  4. The pool is written as a new version-5 container. An empty pool means no appendix.

Record layer, applied to a's records and then b's, one record at a time, keyed by dedup key (§5.3):

  1. payload records from the inputs are dropped; one new payload record describing the output appendix is added at the end. Its members are embedded (true when the output has an appendix), csdt_file_checksum (the output container's file_checksum, else 0), sections (one {index, section_type, record_count} per pool entry, without dtype) and refs (the ref_id of every csdt_ref record in the output). All four are emitted even when empty, so an empty pool with no csdt_ref gives {"k":"payload","v":1,"embedded":false,"csdt_file_checksum":0,"sections":[],"refs":[]}.
  2. New key: insert (records_inserted).
  3. Existing record equal to the incoming one (value equality after parsing): keep, count records_skipped.
  4. Both ent:
    • existing is a stub and incoming is not: take incoming, count stubs_resolved and records_updated;
    • incoming is a stub and existing is not: keep existing (records_skipped);
    • otherwise a conflict (conflicts), resolved by the strategy: higher_salience takes incoming only if its salience is strictly greater (missing counts as 0.0); latest takes incoming; union keeps existing; manual keeps existing and records the conflict as deferred. A conflict counts conflicts; when the strategy takes the incoming record it also counts records_updated, and when it keeps the existing one, records_skipped.
  5. Any other kind: latest takes incoming (records_updated); every other strategy keeps existing (records_skipped).
  6. The output records are written in sort key order (§5.3) with the rules of §2 to §6.

Properties:

  • For inputs whose shared dedup keys carry equal records, merge(a, b) and merge(b, a) produce byte-identical files. With conflicts, the result depends on argument order under latest and higher_salience ties.
  • Merge treats mut records like any opaque record (dedup by the padded LSN in id). Combining two mutation logs with merge is a producer error (§5.2.1), not something merge detects.
  • The merge report keys are records_inserted, records_updated, records_skipped, stubs_resolved, conflicts, sections_merged, sections_deduped.

8.2 Diff

Diff compares two packs a (before) and b (after).

  1. Index each side by dedup key (§5.3). A key that occurs more than once on either side is listed in duplicate_keys and is not otherwise compared.
  2. added: keys only in b. removed: keys only in a.
  3. changed: keys in both whose records differ in identity form (§10.3). Each entry carries the key and both record lines.
  4. Appendix: sections are matched by section index; a section is added, removed, or changed when its section_type or its descriptor CRC32C differs.
  5. Every output list is sorted by sort key (records) or by section index (sections).

Diff never reports a difference that exists only in member order, whitespace or number spelling, because it compares identity forms.

8.3 Shard

Shard produces a new pack holding a connected subset of one pack's graph. The Sharder conformance level is provisional in revision 6.0; this section is its contract.

Graph view of a pack: nodes are ent records by id; directed edges are rel records src → tgt labelled kind; an entity has an embedding when an embref record names it.

Shard specs:

  • entity_centric{seed, max_depth, max_fanout?, relationship_filter?, direction, phase_coherent}: breadth-first from seed to max_depth hops, following edges in direction (outgoing, incoming, both), at most max_fanout neighbours per node when set, and only edges whose kind is in relationship_filter when set.
  • search_result{entity_ids, context_depth, max_fanout?}: the listed entities plus their neighbourhood to context_depth.
  • predicate{...}: entities whose fields satisfy the predicate.
  • cluster{cluster_id, include_bridges}: MUST fail with shard_unsupported_spec until a cluster field exists in the record layer. It MUST NOT return an empty result instead.
  • union[specs]: the union of the member results.

Output pack:

  • the selected ent records, plus a stub ent (stub: true) for each edge endpoint outside the selection when the config asks for stubs;
  • the rel records whose endpoints are both in the output;
  • a hyper record only when every member is selected;
  • embref records of selected entities, with the appendix rewritten to hold only the referenced rows and the bindings remapped;
  • every other appendix section listed in dropped_sections, never dropped silently;
  • a meta record whose created_at is supplied by the caller. Two runs with the same input, spec and created_at MUST produce byte-identical packs.

A shard whose quality score is below a configured minimum fails with shard_quality_below_threshold.

9. Verify algorithm

verify checks a pack and returns a report with the keys sha256_ok, header_footer_agree, record_count{declared, actual}, appendix{present, file_checksum_ok, sections[{index, type, crc_ok}]}, embrefs{total, resolved, dangling, type_mismatch}. Opening a pack runs steps 1 to 8 by default; the caller may opt out explicitly.

Input: the bytes of a .plxi file. If they start with 1F 8B (gzip ID1, ID2 [RFC1952]), a reader MAY decompress them and continue with the result; .plxi.gz is a transport encoding only (§11).

  1. Size. L >= 353, else invalid_header.
  2. Header. Bytes 0..257 are ASCII, byte 256 is LF, and bytes 0..256 match header-content padding (§3.1): else invalid_header. A version other than 6: unsupported_version.
  3. Footer. Bytes L-96 .. L per the §10.0 table: magic PLXF, else invalid_footer; version 6, else unsupported_version; reserved bytes 44..64 all zero, else invalid_footer.
  4. Header against footer. Classify the header (§3.3). Final: records, csdt_offset, csdt_size, digest equal the footer's, else header_footer_mismatch. Placeholder: continue with the footer's values. Neither: header_footer_mismatch.
  5. Layout. With T = text_section_size, A = csdt_offset, C = csdt_size, all integers from the footer, compared without overflow:
    • 257 <= T <= L - 96, else invalid_footer;
    • C = 0: A = 0 and T = L - 96, else invalid_footer;
    • C > 0: A mod 64 = 0, else appendix_alignment; A = T rounded up to a multiple of 64 and A + C = L - 96, else invalid_footer; the padding bytes T .. A are all 0x00, else invalid_footer.
    • The checks above are made in 64-bit arithmetic, before any conversion to an in-memory index, so a layout that fails them is invalid_footer on every platform. Only a layout that passes them and still holds an offset or size that cannot be represented as an in-memory index on the platform (for example at or above 2^32 on a 32-bit target) fails, with limit_exceeded; it MUST NOT be truncated.
  6. Digest. SHA-256 of bytes 257 .. L-96 equals footer bytes 64..96, else checksum_mismatch.
  7. Body. Bytes 257 .. T are valid UTF-8, else invalid_utf8. When C > 0, the last line is the marker line, else invalid_record. Every other non-empty line parses as a record (§4), else json or invalid_record. A marker line elsewhere: invalid_record.
  8. Count. The number of record lines equals the footer's record_count, else record_count_mismatch.
  9. Appendix (Reader-Appendix, only when C > 0): open the container at bytes A .. A+C with the checks of §7.3; check every section's CRC32C; compare the footer's csdt_file_checksum with the container's file_checksum.
  10. Bindings (Reader-Appendix): for each local target (§5.4), section_index is below the section count and record_index is below that section's record_count; for each ext target, a csdt_ref record with that ref_id exists in the pack. Each failure counts as dangling. Verify reports counts; a consumer that binds a dangling target fails with dangling_embref.
    • Type check under the Compact embedding profile (the entity-embedding binding that importers use): a local target of an embref is expected to name a section of type Compact (0x0020) whose record_stride equals the Compact record size (320 bytes), with dtype RECORD. Verify counts mismatches in type_mismatch; a consumer that binds one fails with section_type_mismatch. The profile belongs to the consumer, not to the file format: other kinds of section are legal embref targets.
Reader-Core 1–8Reader-Appendix 9–10
  1. 1Sizeinvalid_header
  2. 2Headerinvalid_headerunsupported_version
  3. 3Footerinvalid_footerunsupported_version
  4. 4Header against footerheader_footer_mismatch
  5. 5Layoutinvalid_footerappendix_alignmentlimit_exceeded
  6. 6Digestchecksum_mismatch
  7. 7Bodyinvalid_utf8invalid_recordjson
  8. 8Countrecord_count_mismatch
  9. 9Appendixappendix_invalidappendix_crc_mismatchonly when the pack has an appendix
  10. 10Bindingsdangling_embrefsection_type_mismatchcounts only; a consumer that binds a failing target reports dangling_embref or section_type_mismatch
Figure 9-1. The ten verify steps in the order they run, each with the error kinds it can report. Opening a pack runs steps 1 to 8 by default.A verifier that stops at the first failure reports the kind of that first failure.

Steps run in this order. A verifier that stops at the first failure MUST report the kind of that first failure, so every conforming implementation reports the same kind for the same file. Reports for steps it did not reach are absent.

10. Footer byte table and canonical JSON

10.0 Footer byte table (96 bytes)

Byte map of the 96-byte footer, nine fields.016324864801magic2version3record_count4text_section_size5csdt_offset6csdt_size78reserved9sha2567csdt_file_checksum08162432404856647280881magic2version3record_count4text_section_size5csdt_offset6csdt_size78reserved9sha2567csdt_file_checksum
Figure 10-1. The 96 footer bytes, one cell per byte, 16 bytes per row. Each outlined run is one field of Table 10-1; the number in the circle is the field’s row in that table.Figure 10-1. The 96 footer bytes, one cell per byte, 8 bytes per row. Each outlined run is one field of Table 10-1; the number in the circle is the field’s row in that table.In order: magic, 4 bytes at offset 0; version, 4 bytes at 4; record_count, 8 bytes at 8; text_section_size, 8 bytes at 16; csdt_offset, 8 bytes at 24; csdt_size, 8 bytes at 32; csdt_file_checksum, 4 bytes at 40; reserved, 20 bytes at 44; sha256, 32 bytes at 64.
Table 10-1. Footer fields (9 fields, 96 bytes)
Offset Size Field Type Rule

The footer is the authoritative source for every count and offset, because it is always final (§3.3).

10.1 Emitted bytes: one member order for every build

  • A writer MUST emit typed records in the member order of §4.3 and opaque records, and every nested JSON object inside any record, with members in the order they were read or inserted.
  • An implementation MUST NOT reorder members as a side effect of its build configuration. The reference implementation enables the preserve_order feature of serde_json in every build, so its default build and its --no-default-features build emit the same bytes.
  • A writer MUST emit a JSON text without insignificant whitespace, one record per line.

10.2 Value spelling

When a value is re-emitted, a writer MUST use these spellings (they are serde_json 1.0.149's, measured in §10.4):

  • Strings: " and \ are escaped as \" and \\; U+0008, U+0009, U+000A, U+000C, U+000D as \b, \t, \n, \f, \r; other code points below U+0020 as \u00XX with lowercase hex; everything else, including /, U+007F and non-ASCII, is written as raw UTF-8. A JSON string that decodes to a lone surrogate MUST be rejected when read (json).
  • Numbers written in the input as an integer token (no fraction, no exponent) whose value fits a signed or unsigned 64-bit integer are written as that integer, exactly.
  • Every other number is an IEEE 754 binary64 value, written with the shortest digit string that reads back to the same value:
    • when 1e-5 <= |x| < 1e16, in positional form, with .0 appended when the value is integral (100.0, 0.00001, 1000000000000000.0);
    • otherwise in exponent form d[.ddd]e+N or d[.ddd]e-N (1e+16, 1e-6, 1.8446744073709552e+19);
    • negative zero is written -0.0;
    • when more than one shortest digit string reads back to the same value, the one serde_json 1.0.149 writes is the spelling. For example, the binary64 value 0x42e87faaebb9a0d4 reads back from both 215492859907334.62 and 215492859907334.63; serde_json writes 215492859907334.62, so that is the spelling, and 215492859907334.63 (the Rust standard library's {} and {:e} digits) does not conform.

The corpus carries a numbers case that pins these spellings; an implementation whose number formatter differs fails it. A second case, pos-numbers-d96, pins the tie rule above and exact reading: its input spells 0x42e87faaebb9a0d4 as 215492859907334.63 and as 215492859907334.62, both re-emitted as 215492859907334.62, and carries the literal 2.1549285990733466e14, which is the different value 0x42e87faaebb9a0d5 and is re-emitted as 215492859907334.66. Identity forms, dedup keys and conformance hashes use these spellings (§10.3).

When a number token is read, a reader MUST return the binary64 value nearest to the token's decimal value (round half to even). The reference reader is serde_json 1.0.149 with its float_roundtrip feature on. Together with the shortest-spelling rule above, this makes canonical JSON a fixed point: reading an emitted line or an identity form and emitting it again yields the same bytes. Rationale: serde_json's default reader can return a value one unit in the last place away from the nearest one; with it, 1,198 of a 20,001-value sweep changed spelling when read and re-emitted, and one value moved by one unit per cycle for nine cycles.

10.3 Identity form

Dedup keys that fall back to a whole object, diff equality, conformance goldens and SHA comparisons between records use the PLXI identity form of a JSON value:

  • object members sorted recursively by member name, comparing names as UTF-8 byte strings;
  • array element order unchanged;
  • no insignificant whitespace;
  • strings and numbers spelled per §10.2.

The identity form of a record line is the identity form of the JSON value that its re-emitted line (§4.3, §4.4) parses to.

10.4 Relation to RFC 8785 (JCS)

The identity form is not the JSON Canonicalization Scheme [RFC8785]. The differences, measured against serde_json = 1.0.149, without and with preserve_order:

Table 10-2. Identity form compared with RFC 8785 (7)
Input PLXI identity form RFC 8785 Rule that differs
member names "דּ" and "😀" U+FB33 first (UTF-8 EF.. < F0..) U+1F600 first (UTF-16 D83D < FB33) key order: UTF-8 bytes vs UTF-16 code units (RFC 8785 §3.2.3)
1.0, 100.0, 1E2 1.0, 100.0, 100.0 1, 100, 100 integral binary64 keeps .0
-0.0 -0.0 0 negative zero
18446744073709551615 18446744073709551615 18446744073709552000 64-bit integers stay exact
1e-6 1e-6 0.000001 positional range lower bound (1e-5 here, 1e-6 in ECMAScript)
1e16, 1e17 1e+16, 1e+17 10000000000000000, 100000000000000000 positional range upper bound (1e16 here, 1e21 in ECMAScript)
18446744073709551616 1.8446744073709552e+19 18446744073709552000 same binary64 value, different spelling

Strings, literals and the absence of whitespace agree in every case probed. Implementations MUST NOT substitute a JCS library for the identity form.

11. Transport encoding: .plxi.gz

A .plxi.gz file is a gzip member [RFC1952] whose decompressed bytes are a complete .plxi file. It exists for shipping; a reader MUST decompress it fully before reading the appendix, and memory-mapped, zero-copy access is not available through it. A writer that streams into gzip cannot seek, so it leaves a placeholder header (§3.3).

12. Versioning

12.1 File format version

The file format version is 6, stored in the header's version token and the footer's version field. Readers MUST reject every other value (unsupported_version). A change that an unmodified version-6 reader would misread requires version 7. There is no in-place compatibility path between file versions.

12.2 Record versions

Each kind versions independently through v (§4.2).

  • Adding an OPTIONAL member keeps v; older readers preserve it in extra.
  • Removing a member, making one REQUIRED, or changing a member's type or meaning raises v. Older readers then keep the record opaque (§4.4), and it still round-trips through them.
  • Adding a kind needs no version change; older readers keep it opaque.

12.3 Spec revisions

This document is revision 6.0. A revision that only clarifies or adds OPTIONAL members, kinds or error kinds is 6.x. Report keys and error kind strings version together: adding one is a minor change with new conformance cases; removing one is a major change.

13. Differences from the reference implementation

Not part of this publication.

14. Conformance levels

An implementation claims one or more levels. Each level is proved by the corpus groups listed; the case ids are listed on the cases page.

Table 14-1. Conformance levels (6)
Level Requirements Corpus groups
Reader-Core §2, §3, §4, §5.3, §6, §9 steps 1-8, §10.0-10.3 positive, graph (no-appendix pack, empty pack, .plxi.gz), negative (framing, digest, count, UTF-8, header byte 80)
Reader-Appendix Reader-Core plus §7, §9 steps 9-10 graph (Compact appendix with F32/BF16/MXFP4/B2048 sections), negative (container offsets, align_log2, wrong section type, wrong stride, appendix CRC flip)
Writer emits files that pass Reader-Core and Reader-Appendix verify; §4.3, §4.4, §5.3 order, §10.1, §10.2; container version byte 5 only positive goldens (.golden.jsonl), numbers case
Merger §8.1 merge (one per strategy, the id dedup trap, appendix section union), negative (an appendix with legacy container version byte 3)
Differ §8.2 diff
Sharder (provisional) §8.3 shard

The corpus has limits an implementation should know: it tests the structure and checksums of the embedded container, not the meaning of the data inside it.

15. Reserved

Not part of this publication.

16. References

© 2020-2026 Cintile Inc. All Rights Reserved.

Anyone may implement this format. Copying or republishing the text of this specification requires permission from Cintile Inc.

Sections