On this page· 5
Specification
- Revision 6.0
- file format 6
- Whole specification on one page
| Offset | Size | Field | Type | Rule |
|---|---|---|---|---|
| 0 | 4 | magic | bytes | 50 4C 58 46 (PLXF) |
| 4 | 4 | version | u32 | 6 |
| 8 | 8 | record_count | u64 | number of record lines (§4.1) |
| 16 | 8 | text_section_size | u64 | T: absolute offset of the first byte after the text section: header line, body and, when an appendix exists, the marker line (§2) |
| 24 | 8 | csdt_offset | u64 | A (0 when no appendix) |
| 32 | 8 | csdt_size | u64 | C (0 when no appendix) |
| 40 | 4 | csdt_file_checksum | u32 | copy of the embedded container's file_checksum (CRC32C over its section table); 0 when no appendix |
| 44 | 20 | reserved | bytes | MUST be written as zero; readers MUST reject nonzero (invalid_footer) |
| 64 | 32 | sha256 | bytes | raw digest D (§6) |
The footer is the authoritative source for every count and offset, because it is always final (§3.3).
- A writer MUST emit typed records in the member order of §4.3 and opaque records, and every nested JSON object inside any record, with members in the order they were read or inserted.
- An implementation MUST NOT reorder members as a side effect of its build configuration. The reference implementation enables the
preserve_orderfeature ofserde_jsonin every build, so its default build and its--no-default-featuresbuild emit the same bytes. - A writer MUST emit a JSON text without insignificant whitespace, one record per line.
When a value is re-emitted, a writer MUST use these spellings (they are serde_json 1.0.149's, measured in §10.4):
- Strings:
"and\are escaped as\"and\\; U+0008, U+0009, U+000A, U+000C, U+000D as\b,\t,\n,\f,\r; other code points below U+0020 as\u00XXwith lowercase hex; everything else, including/, U+007F and non-ASCII, is written as raw UTF-8. A JSON string that decodes to a lone surrogate MUST be rejected when read (json). - Numbers written in the input as an integer token (no fraction, no exponent) whose value fits a signed or unsigned 64-bit integer are written as that integer, exactly.
- Every other number is an IEEE 754 binary64 value, written with the shortest digit string that reads back to the same value:
- when
1e-5 <= |x| < 1e16, in positional form, with.0appended when the value is integral (100.0,0.00001,1000000000000000.0); - otherwise in exponent form
d[.ddd]e+Nord[.ddd]e-N(1e+16,1e-6,1.8446744073709552e+19); - negative zero is written
-0.0; - when more than one shortest digit string reads back to the same value, the one serde_json 1.0.149 writes is the spelling. For example, the binary64 value
0x42e87faaebb9a0d4reads back from both215492859907334.62and215492859907334.63; serde_json writes215492859907334.62, so that is the spelling, and215492859907334.63(the Rust standard library's{}and{:e}digits) does not conform.
- when
The corpus carries a numbers case that pins these spellings; an implementation whose number formatter differs fails it. A second case, pos-numbers-d96, pins the tie rule above and exact reading: its input spells 0x42e87faaebb9a0d4 as 215492859907334.63 and as 215492859907334.62, both re-emitted as 215492859907334.62, and carries the literal 2.1549285990733466e14, which is the different value 0x42e87faaebb9a0d5 and is re-emitted as 215492859907334.66. Identity forms, dedup keys and conformance hashes use these spellings (§10.3).
When a number token is read, a reader MUST return the binary64 value nearest to the token's decimal value (round half to even). The reference reader is serde_json 1.0.149 with its float_roundtrip feature on. Together with the shortest-spelling rule above, this makes canonical JSON a fixed point: reading an emitted line or an identity form and emitting it again yields the same bytes. Rationale: serde_json's default reader can return a value one unit in the last place away from the nearest one; with it, 1,198 of a 20,001-value sweep changed spelling when read and re-emitted, and one value moved by one unit per cycle for nine cycles.
Dedup keys that fall back to a whole object, diff equality, conformance goldens and SHA comparisons between records use the PLXI identity form of a JSON value:
- object members sorted recursively by member name, comparing names as UTF-8 byte strings;
- array element order unchanged;
- no insignificant whitespace;
- strings and numbers spelled per §10.2.
The identity form of a record line is the identity form of the JSON value that its re-emitted line (§4.3, §4.4) parses to.
The identity form is not the JSON Canonicalization Scheme [RFC8785]. The differences, measured against serde_json = 1.0.149, without and with preserve_order:
| Input | PLXI identity form | RFC 8785 | Rule that differs |
|---|---|---|---|
member names "דּ" and "😀" |
U+FB33 first (UTF-8 EF.. < F0..) |
U+1F600 first (UTF-16 D83D < FB33) |
key order: UTF-8 bytes vs UTF-16 code units (RFC 8785 §3.2.3) |
1.0, 100.0, 1E2 |
1.0, 100.0, 100.0 |
1, 100, 100 |
integral binary64 keeps .0 |
-0.0 |
-0.0 |
0 |
negative zero |
18446744073709551615 |
18446744073709551615 |
18446744073709552000 |
64-bit integers stay exact |
1e-6 |
1e-6 |
0.000001 |
positional range lower bound (1e-5 here, 1e-6 in ECMAScript) |
1e16, 1e17 |
1e+16, 1e+17 |
10000000000000000, 100000000000000000 |
positional range upper bound (1e16 here, 1e21 in ECMAScript) |
18446744073709551616 |
1.8446744073709552e+19 |
18446744073709552000 |
same binary64 value, different spelling |
Strings, literals and the absence of whitespace agree in every case probed. Implementations MUST NOT substitute a JCS library for the identity form.
© 2020-2026 Cintile Inc. All Rights Reserved.
Anyone may implement this format. Copying or republishing the text of this specification requires permission from Cintile Inc.