Only this pageAll pages
Powered by GitBook
Couldn't generate the PDF for 113 pages, generation stopped at 100.
Extend with 50 more pages.
1 of 100

Internet Object

Loading...

Loading...

Internet Object

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Core Concepts

Loading...

Loading...

Structure and Syntax

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Definitions

Loading...

Loading...

Loading...

Loading...

Collections

Loading...

Loading...

Loading...

Loading...

Schema Definition Language

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Streaming

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Parsing & Errors

Loading...

Loading...

Loading...

Loading...

Loading...

Serialization

Loading...

Loading...

Loading...

Loading...

Loading...

Conformance

Loading...

Loading...

Versioning & Stability

Loading...

Loading...

Getting Started

A short, language-agnostic tour of Internet Object in pure IO.

This is a five-minute tour of Internet Object (IO) using the format itself — no programming language required.

1. A single object

The simplest document is one object. Fields are comma-separated; the header line names them:

name: string, age: int, email: email
---
John Doe, 30, [email protected]

Above the --- is the header (here, the schema); below it is the data. Because the schema fixes the field order, the data is just values — no repeated keys.

2. A schema and a collection

Define the schema once in the header with $schema, then stream many records, each beginning with ~:

~ $schema: { name: string, age: int, email: email, active: bool }
---
~ John Doe, 30, [email protected], T
~ Jane Doe, 25, [email protected], F

T/F are booleans. A collection of records shares one schema — compact and validated.

Fields can carry constraints. Invalid data is reported, not silently accepted:

Define a shape once and reference it with $:

Header keys without a prefix carry document metadata, kept separate from the data:

  • — how it compares to JSON and others

  • — header and data in depth

  • — the schema language

  • — records and streaming

3. Constraints

4. Nesting and reuse

5. Metadata

Where to next

Why Internet Object?
Internet Object Document
Internet Object Schema
Collections
~ $schema: { name: string, age: { int, min: 0, max: 120 } }
---
~ John, 30      # ✓
~ Mary, 200     # ✗ mismatched-max
~ $address: { street, city }
~ $schema: { name: string, address: $address }
---
~ John, { Main St, NYC }
~ count: 2
~ $schema: { name, age: int }
---
~ John, 30
~ Jane, 25

Advanced Schema Concepts

Conventions

How to read this specification — requirement keywords, examples, and error codes.

Requirement keywords

The key words MUST, MUST NOT, REQUIRED, SHALL, SHALL NOT, SHOULD, SHOULD NOT, RECOMMENDED, MAY, and OPTIONAL in this document are to be interpreted as described in RFC 2119 and RFC 8174, and only when they appear in all capitals.

They apply to every chapter, not only to Conformance Requirements. Where a chapter states a rule in ordinary prose — "a section name must be unique" — the requirement is the same; the capitals mark where the wording has been made precise, and their absence is not permission.

Normative and informative

Every page carries a status in its front matter:

Where the two disagree, the normative page wins. If you find such a disagreement, it is a defect in this specification, not a choice.

Examples are written in Internet Object and marked ```ruby, whose highlighting happens to suit the format. They are executable: a checker runs every complete example against the reference implementation on each change, so an example that contradicts the text fails the build rather than sitting quietly on the page.

Two annotations carry meaning inside an example:

Marker
Means

The cross means an error, never "not the form we are discussing". Where a line is legal but not the construct under discussion, the example says so in words instead — a distinction worth keeping, because most such lines are perfectly good values of some other kind.

A fenced block without a --- separator is a fragment: it illustrates shape and is not executed.

Every reported error carries a stable code. Codes are normative; the messages that accompany them are not, and may be reworded or translated freely. Tooling MUST branch on the code and MUST NOT parse the message.

How codes are named — and the closed vocabulary they draw from — is . The codes themselves are catalogued in .

Two pairs of words are easy to confuse, because each names a different axis:

All four combinations occur, and a document may hold them at once. See the .

status

Meaning

candidate

Normative. An implementation is measured against it. Still open to change before 1.0.

informative

Explanatory. Rationale, comparisons, history — nothing here constrains an implementation.

# ✗ <error-code>

this line is rejected, with that code

# → <value>

this line loads to that value

open / closed object

written without braces / with braces. A question of syntax.

strict / extensible schema

rejects undeclared members / accepts them (*). A question of validation.

Examples

Error codes

Terminology

See Also

Error Codes
Error Model
Glossary
Conformance Requirements
Error Codes
Glossary

Decimal

The decimal type — fixed-precision decimal numbers.

The decimal type validates an exact, fixed-precision decimal — for money and other values where binary floating point would lose accuracy. In data it is written with an m suffix: 123.45m.

For the literal syntax, see Decimal values.

TypeDef

A decimal MemberDef accepts only the options below. Any other key is invalid.

Option
Type
Description

Unlike the other numeric types, decimal has no format option — and this is deliberate, not an omission. A format selects among the literals that can express a value, and a decimal has only one: <digits>.<digits>m. Radix notations cannot express a fractional value, and (1.23e2m is invalid). With a single possible spelling there is nothing to select.

The m suffix is always written; without it the output would read back as a plain number.

precision and scale together give SQL-style DECIMAL(precision, scale) validation:

  • scale — the number of fractional digits MUST equal scale.

  • precision — the total significant digits MUST NOT exceed precision.

With neither precision nor scale, a decimal is compared by its exact value.

Resolution follows the :

  • ·

  • ·

type

string

The type name decimal. First positional value.

default

decimal

Value used when the member is omitted. Second positional value.

choices

array of decimal

Restricts the value to a fixed set.

price: { decimal, precision: 5, scale: 2 }
---
~ 123.45m    # ✓  (5 digits, 2 after the point)
rate: { decimal, scale: 2 }
---
~ 1.5m       # ✗ mismatched-scale  (1 fractional digit, scale requires 2)
amount?*: decimal
---
~ {}      # ✓ omitted → absent
~ N       # ✓ null
~ 9.99m   # ✓

No format option

Precision & scale

Optional, nullable & defaults

See Also

scientific notation is not part of the decimal literal
common rules
Decimal values
Numeric Types
BigInt
TypeDef
MemberDef

precision

int

Maximum total number of significant digits.

scale

int

Exact number of digits after the decimal point.

min

decimal

Minimum allowed value (inclusive).

max

decimal

Maximum allowed value (inclusive).

multipleOf

decimal

The value must be an exact multiple of this.

optional

bool

If true, the member may be omitted. Shorthand: ? suffix.

null

bool

If true, the member may be null. Shorthand: * suffix.

URL

The url type — a string validated as a URL.

url is a string shortcut (see String Types) whose value MUST be a valid URL. It shares the string MemberDef and adds URL-format validation.

Quote URL values. A URL contains : and /, which end an unquoted (open) string, so URLs must be written as quoted strings.

website: url
---
~ 'https://example.com'         # ✓
~ "https://example.com/p?q=1"   # ✓
~ 'not a url'                    # ✗ invalid-url

Restrict to a fixed set with choices:

homepage: { url, choices: ['https://a.com', 'https://b.com'] }
---
~ 'https://a.com'    # ✓

See Also

  • ·

String Types
Email
MemberDef

Parser Behavior & Recovery

Boundary-bounded syntax-error recovery and processing options.

A conformant processor SHOULD recover from errors and continue, so that a single bad record does not discard the rest of a document.

Syntax-error recovery is bounded by structure

On a syntax error, the parser skips tokens until the next boundary and resumes there. The boundaries are:

  • the record separator ~ (start of the next collection item), and

  • the section separator --- (start of the next section),

  • or end of input.

The malformed middle record is reported as an error; the records before and after it are still parsed.

Each record is validated independently. A validation error in one record does not stop validation of the others (see ).

A processor typically offers options that control recovery and output. Common ones:

  • continue-on-error — collect errors and keep going (recommended), versus failing on the first error.

  • skip-errors — omit error entries from the loaded result, returning only the records that succeeded.

Option names and exact semantics are implementation-defined; this section describes the behaviors a conformant processor is expected to provide.

When continuing past an error, a processor marks the failed record with an error placeholder in the result so consumers can tell which records succeeded and which did not. See .

  • ·

Strings

The three string forms — open, regular, and raw.

Strings represent sequences of Unicode code points. They carry textual data and preserve whitespace and formatting within their boundaries.

Internet Object supports three string forms, each with its own syntax and use cases:

stringValue = openString | regularString | rawString
Form
Description
Example

All three forms preserve whitespace and Unicode content as written.

  • Open string — simple, unstructured text with no leading or trailing whitespace and no structural characters.

  • Regular string — text that needs structural characters, leading/trailing whitespace, or escape sequences.

  • Raw string — text with many backslashes or quotes (file paths, regular expressions), where escaping would be cumbersome.

  • — schemas for strings

  • — the numeric forms

  • — all value types

Schema References

Reusable schemas and types referenced with $.

A reference (ref) is a $-prefixed definition in the header that names a reusable schema or type. You define it once and refer to it elsewhere as $name. Refs come in two forms:

  • Schema reference — names an object shape (a ).

  • Type reference — names a single constrained type (a MemberDef), e.g. a percentage.

Nulls

The null value — an explicit absence of a value.

A null represents the absence of a value — data that is missing, unknown, or intentionally empty. Null is a scalar value with a compact and a verbose form.

Token
Name
Description

Creating Collections

Creating collections, with or without a schema.

A collection is a sequence of records in the data section, each introduced by a tilde ~. You can create one with or without a schema — though a schema is recommended.

Without a schema, each record is parsed on its own and its values are mapped to positional indices. Records may differ in shape:

Define a schema in the header; every record is validated against it. Use a keyed reference (address: $address) to reuse a shape:

Reference a shape with a keyed member (address: $address). A bare $address in the schema is read as a field literally named $address, not as the referenced shape.

Unquoted; the simplest form; ends at a structural character or whitespace.

John Doe

Regular string

Quoted with single or double quotes; supports escaping.

"John Doe"

Raw string

Prefixed with r; quoted; backslashes are literal.

r'C:\path' or r"C:\path"

When to use each form

See Also

String Types
Numeric Values
Value Representations
Open string

Validation recovery is bounded by the object

Processing options

Error nodes

See Also

Collection Rules
Error Accumulation
Error Model
Error Accumulation
Syntax Errors
  • Collection · Collection Rules

  • Schema References

Simple collection (no schema)

Explicit collection (with schema)

See Also

The special ref
$schema
is the document's default schema.

Define an object shape once, reuse it across fields and schemas:

$schema: $person sets the default schema by reference. A ref can be used as a field's type (home: $address) or as an array's element type (tags: [$address]).

  • Refs are resolved after the entire header has been read, so order within the header is not significant — a ref MAY appear before the definition it targets. For readability, you SHOULD still define a ref before you use it.

  • A ref to a name that is never defined is an error (undefined-schema).

  • Reusing a ref many times keeps a document small and consistent.

See Error Handling in Definitions for the resolution errors.

A ref whose body is a single constrained type acts as a reusable type — your own named shortcut, the document-local counterpart of built-ins like uint8 or email:

Implementation status (beta). Type references are being added. Today a top-level $ definition is compiled as an object schema (a SchemaDef), so its braces are read as an object shape — a constrained-type body such as { number, min: 0, max: 100 } is not yet interpreted as a reusable number type. The forms above show the target syntax. Schema references (object shapes) work today.

  • Definitions · Variables

  • Object (SchemaDef)

  • Internet Object Schema

SchemaDef

Schema references

Resolution rules

Type references

See Also

null

Verbose null

the verbose null keyword

The compact and verbose forms are equivalent; the compact form is recommended for terse data.

Null is the absence of a value, distinct from an empty string or an empty array:

Null keywords are case-sensitive and spelled exactly. Any other token is not an error — it is parsed as an open string, so it is not null:

To store one of these as text, that is exactly what happens. To express the absence of a value, write N or null.

  • Value Representations — all value types

  • Booleans — true and false

  • Optional & nullable members — ? and * in schemas

N

Compact null

Syntax

Structural elements

null in compact form

Valid forms

Null versus empty

Not null

See Also

~ $schema: { name: string, age: int }
---
~ John, 28              # parsed
~ Bad, { unclosed      # syntax error here; parser skips to the next ~
~ Bob, 35              # parsed — recovery resumed at this record
---
~ Ironman, 20, Male, { Bond Street, New York, NY }
~ Spiderman, 25, Male, { Duke Street, New York, NY }, cool
~ $address: { street, city, state }
~ $schema: { name: string, age: { int, min: 18 }, address: $address }
---
~ Ironman, 20, { Bond Street, New York, NY }      # ✓
~ Wonderwoman, 25, { Z Street, San Francisco, CA } # ✓
~ $address: { street: string, city: string }
~ $person: { name: string, home: $address, office?: $address }
~ $schema: $person
---
~ John, { Main St, NYC }, { 5th Ave, NYC }
~ Jane, { Oak Ave, LA }
# A reusable "percent" type and "short text" type
~ $percent: { number, min: 0, max: 100 }
~ $shortText: { string, maxLen: 40 }
~ $schema: { name: $shortText, score: $percent }
null        = compactNull | verboseNull
compactNull = "N"
verboseNull = "null"
---
N, null
N        # null — no value
""       # an empty string (a value)
[]       # an empty array (a value)
n                    # open string "n", not null
NULL                 # open string "NULL", not null
Null                 # open string "Null", not null
nil                  # open string "nil", not null
undefined            # open string "undefined", not null

Validation Model

The parse, validate, load, and stringify pipeline.

Processing an Internet Object document is defined as a pipeline of four stages. Each stage has a clear input and output, so implementations behave consistently.

text ──parse──▶ document tree ──validate──▶ checked tree ──load──▶ values
                                                              ◀─stringify── values

Parse

Input: UTF-8 text. Output: a document tree (header, sections, records, values).

Parsing checks only syntax — that the text is well-formed. It does not consult any schema. Syntax errors are produced here (see Error Model).

Validate

Input: the document tree + a schema. Output: the same tree, with each value checked.

Validation applies the schema: types, constraints (min, maxLen, pattern, choices, …), optionality, and nullability. Validation errors are produced here. With no schema, data is accepted structurally and mapped to positional keys.

An implementation will usually offer two ways in: validating a document read from text, and validating values the host language already holds — an object from an API response, a row from a database.

These are two routes to one stage, not two stages. Validation is defined on the logical value. For the same schema and the same logical value, both routes MUST reach the same outcome: the same accept-or-reject decision, the same error codes, in the same order.

Two values are the same logical value when they hold the same members with the same names and the same typed contents — regardless of spelling. Text ~ Alice, 15 under {name: string, age: int} is the same logical value as the native {name: "Alice", age: 15}, because positional binding is part of reading the text, not part of validating it.

This is worth stating because the two routes are commonly written as separate code, each walking its own kind of input. Nothing forces them to stay in step, and a divergence is close to undetectable from inside a single implementation: each route has its own tests, and both pass. It surfaces only when the same data is sent both ways, or when a second implementation reads the specification and builds one validator — at which point the specification can only describe one of the two behaviors, and every user of the other one is affected.

Testing this. Run the conformance corpus's validation cases through every entry point the implementation offers, asserting both produce the same codes. Sampling a handful of cases is not enough: the routes agree on the common shapes by construction, so the disagreements live in exactly the cases nobody thinks to pick.

Input: the validated tree. Output: in-memory values.

Loading converts checked values into their final representations (numbers, booleans, dates, byte data, nested objects/arrays), applying defaults for omitted fields.

The inverse of the pipeline: in-memory values are serialized back to Internet Object text, honoring schema hints such as a number's format or a string's quote style. A value that is loaded and then stringified SHOULD round-trip to an equivalent document.

Composition & Reuse

Composing and reusing schemas through references.

Large schemas are built by composing smaller, named pieces. Define a shape once in the header as a $ reference and reuse it wherever it's needed — across fields, arrays, and other schemas.

Reuse a shape across fields

~ $address: { street, city }
~ $schema: { name: string, home: $address, office?: $address }
---
~ John, { Main St, NYC }, { 5th Ave, NYC }
~ Jane, { Oak Ave, LA }

Compose schemas from other schemas

A reference can be used inside another reference, building larger shapes from smaller ones:

~ $address: { street, city }
~ $person: { name: string, address: $address }
~ $schema: { lead: $person, members: [$person] }
---
~ { Ann, { Main St, NYC } }, [{ Bob, { Oak Ave, LA } }, { Cy, { 5th Ave, NYC } }]

Here members is an array whose element type is the $person schema.

Set the default schema by reference

$schema may itself be a reference:

~ $address: { street, city }
~ $person: { name: string, home: $address }
~ $schema: $person
---
~ John, { Main St, NYC }

Guidance

  • For readability, define a shape before you reference it. Order within the header is not significant — references resolve after the whole header is read (see ).

  • Reuse keeps documents consistent and small; change a shape once, everywhere updates.

  • ·

Overview

Streaming — an incremental, record-oriented transport over the Internet Object data model.

Streaming is Internet Object consumed incrementally. A producer frames records onto a byte or text stream, and a consumer reads them back one logical record at a time as the bytes arrive — without waiting for the whole document. It is not a separate format or a second parser: it is the same data model, the same grammar, and the same validation, delivered over time.

Streaming is part of the format, not an add-on. It adds only three things — framing, transport coordination, and an emission envelope around each record. It defines no new type, no new validation rule, and no new serialization behavior. Those all come from the core specification, unchanged.

This chapter is the language-neutral, normative contract for Internet Object streaming. It governs every implementation in every language and on every platform. It uses the requirement keywords MUST, MUST NOT, SHOULD, SHOULD NOT, and MAY as defined in RFC 2119 and RFC 8174. A conformant implementation satisfies every MUST and MUST NOT.

Streaming is an incremental, record-oriented transport. A producer frames Internet Object records onto a stream; a consumer reads them back one logical record at a time as bytes arrive.

Bool

The bool type — true/false values.

The bool type validates a boolean value. In data it is written compactly as T/F or verbosely as true/false.

For the boolean value syntax, see .

A bool MemberDef accepts only the options below. Any other key is invalid.

Why Internet Object?

Why choose Internet Object over JSON, CSV, YAML, and binary formats.

Internet Object is a text-based, schema-first data format for interchange over the internet. It keeps JSON's readability while removing its biggest costs: repeated keys, no schema, no comments, and no native streaming.

The same data, in JSON and in IO:

The keys are stated once in the schema, so each record carries only its values. For collections of similar objects this is dramatically smaller — closer to CSV's density, but with types, nesting, and validation.

Need
JSON
CSV
YAML
Internet Object

Case Sensitivity Rules

Case sensitivity for keys, keywords, and type names.

Internet Object is case-sensitive throughout. Keys, keywords, and type names MUST be written with exact casing.

Member keys are distinct by case — Name and name are two different fields:

The literal keywords MUST be written exactly as defined. Their accepted forms are:

Meaning
Accepted
Not accepted

See Also

Schema References
Schema References
Object (SchemaDef)
Array

Entry points

Load

Stringify

See Also

Conformance Requirements
Parsing & Errors
   IO text ──parse──▶ document tree ──┐
                                      ├──validate──▶ same outcome, either way
   native values ─────────────────────┘
The governing requirement: streaming MUST behave like a record protocol, not a raw chunk parser. Transport chunk boundaries carry no meaning — splitting or coalescing the bytes MUST NOT change the records a consumer sees. What the consumer observes is a sequence of records, each either a successfully parsed value or a recoverable error, in wire order.

The protocol defines two roles:

  • the reader — consumes a stream and emits one item per logical data record;

  • the writer — frames records onto a stream.

It also places obligations on adapters (transport bridges) and transports. All of these are abstract roles, not the API of any one library — see Readers & Writers.

Internet Object core is the single authority for data semantics. The streaming protocol is subordinate to it and inclusive of it.

  • Subordinate on semantics. Streaming MUST NOT redefine, reinterpret, or override any Internet Object semantics — what a type means, how values coerce, whether a value validates, how default, optional, null, and choices resolve, how strict and extensible schemas behave, how values serialize, or what an error's identity is.

  • Inclusive, not bolt-on. Streaming MUST reuse core. It MUST NOT fork, shadow, or partially re-implement core parsing, schema resolution, validation, or serialization.

Everything streaming adds is around the core result: framing before it, transport beneath it, and an envelope after it.

A single test enforces the relationship above. It is the heart of the protocol:

For the same record text and the same definitions state, a streamed record MUST produce the same parsed record value, and the same error identity, that the non-streaming core path (parse → schema processing → validation) produces for an equivalent one-record document.

If a behavior cannot be derived from "run core over this record's text," it is out of scope for streaming. The protocol references the core type, schema, validation, and serialization rules; it never restates them. When you need those rules, follow the links to the relevant chapter — for example Validation Model, Internet Object Schema, and Error Model.

These terms have precise meanings throughout this chapter:

Term
Meaning

Logical record

One Internet Object collection record, introduced on the wire by ~. The unit the reader emits.

Header

The definitions block at the start of a stream, before the first ---.

Section

Page
Covers

The on-the-wire grammar, the mandatory --- terminator, control frames, and UTF-8 encoding.

The two-kind item model, record indexing, and degenerate inputs.

This is Streaming Protocol v1. The two-kind item model, the framing rules, and the error model are frozen for v1. Additive, optional metadata MAY be introduced without a version bump; breaking changes require a new major version. The protocol is versioned independently of any implementation, and an implementation declares which protocol version it implements. See the Version History.

  • Collection — the record sequence streaming transports

  • Data Streaming — how collections motivate streaming

  • Validation Model — the parse → validate → load pipeline streaming reuses

  • Error Model — the core error identities streaming preserves

What streaming is

Relationship to the core specification

The equivalence rule

Terminology

How this chapter is organized

Versioning

See Also

Option
Type
Description

type

string

The type name bool. First positional value.

default

bool

Value used when the member is omitted. Second positional value.

Optional with a default:

Nullable:

Input
Result

T/true/F/false

the boolean

any other token

expected-boolean error

N, key nullable (*)

  • Booleans (value syntax)

  • TypeDef · MemberDef

TypeDef

Booleans
active: bool
---
~ T        # ✓ true
~ false    # ✓
~ yes      # ✗ expected-boolean
verified?: { bool, T }    # optional; defaults to true when omitted
---
~ {}        # ✓ verified resolves to T
~ F         # ✓ verified is F
flag*: bool               # nullable (value may be N)
---
~ N         # ✓ null
~ T         # ✓

Examples

Optional, nullable & defaults

See Also

Human-readable

✓

✓

✓

✓

No repeated keys

✗

✓

✗

✓

  • Smaller payloads — keys live in the schema, not in every record.

  • Validation built in — types and constraints travel with the data; bad values are reported with precise errors.

  • Comments — annotate documents inline with #.

  • Collections & streaming — emit and consume records one at a time.

  • Richer types — int/uint/decimal/bigint, date/time/datetime, binary, and reusable named types.

  • JSON-compatible where it counts — quoted-key object syntax is accepted, easing migration. See .

For one-off, schema-less, small payloads — or where ubiquitous tooling matters most — JSON is perfectly adequate. Internet Object pays off when you have many similar records, want validation, or care about size and streaming.

  • Getting Started · Objectives

  • JSON Compatibility

[
  { "name": "John Doe", "age": 30, "email": "[email protected]", "active": true },
  { "name": "Jane Doe", "age": 25, "email": "[email protected]", "active": false }
]
~ $schema: { name: string, age: int, email: email, active: bool }
---
~ John Doe, 30, [email protected], T
~ Jane Doe, 25, [email protected], F

The core idea

How it compares

What you gain

When JSON is still fine

See Also

true

T, true

t, True, TRUE

false

F, false

f, False

null

N, null

n, Null

Built-in type names are lowercase: string, int, uint8, bool, datetime, decimal, and so on. String or INT are not recognized types.

  • Booleans · Nulls

  • Structural Characters & Separators

Name: string, name: string
---
{ Name: Alice, name: alice }

Keys

Keywords

active: bool
---
~ T        # ✓
~ true     # ✓
~ t        # ✗ expected-boolean

Type names

See Also

Conformance Requirements

MUST/SHOULD/MAY duties of parsers, validators, and serializers.

This section states the duties of a conformant implementation. Internet Object is language-independent; these requirements describe behavior, not any particular API.

Requirement keywords are defined once, for the whole specification, in Conventions — they are not local to this page, and a rule stated in ordinary prose elsewhere is no weaker for it.

All implementations

  • MUST accept input encoded as UTF-8.

  • MUST treat the format as case-sensitive (keys, keywords, type names).

  • MUST recognize the structural characters and keywords exactly as defined.

  • MUST report every error with a code from the registry in , named by the rule in , together with the position in the source. An implementation MUST NOT invent a code, assemble one at runtime, or report an error without one.

  • MUST NOT accept a prefix of a malformed construct and discard the remainder — a truncated value that parses is worse than a rejected one, because nothing reports it.

  • MUST build a document tree according to the .

  • SHOULD recover from a syntax error by skipping to the next boundary (~ or ---) and continuing, rather than aborting the whole document.

  • MUST validate data against the schema: types, constraints, optionality, nullability.

  • MUST recognize the closed set of built-in types and their allowed options (each type's ).

  • MUST reject a value that violates its type or constraints, distinguishing a type failure from a constraint failure as defines.

  • MUST validate each record independently; one invalid record MUST NOT invalidate others.

  • MUST produce output that re-parses to equivalent data, and MUST produce output that parses without error — see .

  • MUST preserve each value's type, not merely its printed form, and MUST quote any string or key that would otherwise read back differently — see .

  • MUST write a member's name whenever no schema in scope can recover it, and MUST NOT repeat a name a schema already carries — see .

The section is normative for all of the above.

  • The specification carries its own version (currently 1.0 Draft).

  • Implementations carry their own versions independently and SHOULD declare which specification version they conform to (e.g. "implements Internet Object 1.0").

The official TypeScript/JavaScript implementation, , serves as a reference implementation. Where this specification and an implementation disagree during the draft period, the discrepancy is tracked and resolved case by case; the specification is the intended source of truth as it stabilizes.

  • — requirement keywords, and how examples are marked

  • ·

  • ·

Error Handling in Definitions

Errors that arise from header definitions and references.

Definitions are resolved after the entire header has been read, and references are checked again as data is validated. Two errors are specific to definitions:

Condition
Error code
Cause

Reference to an undefined schema or type

undefined-schema

Error codes are stable; messages and positions may vary between implementations. Branch on the code, not the message.

A $ reference must name a schema or type defined in the header. An undefined name fails with undefined-schema:

A @ reference must name a variable defined in the header. An undefined name fails with undefined-variable:

Because definitions resolve only after the whole header is read, order within the header is not significant — a reference MAY appear before the definition it targets. The following resolves even though $address is defined after the schema that uses it:

For readability you SHOULD still define a reference before you use it; doing so reads top to bottom and makes the dependency obvious.

  • · ·

  • — the full error catalogue

Schema Representation

How a schema is written and how data is mapped to it — open/closed, positional/keyed, and the default schema.

A schema is written in the same object syntax as data. This page covers how to write one and how a document's data is mapped to it. For the building blocks of a schema, see the ; for types and constraints, see and .

"Open" and "closed" here describe braces, not behaviour: an open object is written without them, a closed one with them. Whether a schema accepts undeclared fields is a separate question, and this specification calls that strict or extensible — see .

A top-level schema may be written in open form, without surrounding braces:

A nested object member must be enclosed in braces, because the braces are what mark the value as an object:

A member with no type defaults to any, so name, age, address is shorthand for three

Union Types (anyOf)

Members that accept more than one type, via anyOf.

When a field must accept values of more than one type, use the anyOf constraint on the any type (see ). A value is valid if it matches any one of the listed alternatives.

A value matching none of the alternatives is rejected:

Each alternative may be a full MemberDef (with constraints) or an object shape, not just a bare type name:

  • Order alternatives from most specific to least specific.

optional

bool

If true, the member may be omitted. Shorthand: ? suffix on the key.

null

bool

If true, the member may be null. Shorthand: * suffix on the key.

null

N, not nullable

forbidden-null error

omitted, default set

the default

omitted, optional (?), no default

absent

omitted, required

missing-value error

not-a-number

NaN

nan, NAN

infinity

Inf, -Inf

inf, Infinity

Built-in schema & validation

✗

✗

✗

✓

Nested / structured data

✓

✗

✓

✓

Comments

✗

✗

✓

✓

Streaming of records

✗

✓

✗

✓

Precise numerics (bigint, decimal)

✗

✗

✗

✓

Dates/times & binary as first-class

✗

✗

partial

✓

JSON Compatibility
  • MUST NOT invent new built-in type names; document-local types are declared with $ references.

  • MUST produce the same outcome for the same logical value, whatever route it arrived by — the same accept-or-reject decision, the same error codes, in the same order. Validation is defined on the value, not on how it was delivered. See Entry points.

  • MUST NOT drop a member, and MUST NOT infer a schema the document does not carry.
  • SHOULD honor schema serialization hints (e.g. number format, string quote style).

  • A conformant parser

    A conformant validator

    A conformant serializer

    Versioning

    Reference implementation

    See Also

    Error Model
    Error Codes
    grammar
    TypeDef
    Error Codes
    Round-Trip Guarantees
    Value Formatting
    Key Emission
    Serialization
    internet-object
    Conventions
    Error Codes
    Error Model
    Validation Model
    Formal Grammar

    A run of records sharing one schema context, introduced by a --- control frame.

    Control frame

    A header-definition block or a section marker (--- / --- $Schema). Control frames are never emitted as data items.

    Stream item

    The envelope the reader emits per logical record (see Stream Items).

    Frame

    A contiguous span of stream text the reader buffers and resolves as a unit — one record, or the whole header.

    Default schema context

    The active schema used to validate records that carry no explicit schema selector.

    Atomic header resolution, preloaded definitions, precedence, and schema switching.

    Streaming Error Model

    Error categories, recoverable-versus-fatal disposition, and stream-absolute positions.

    Readers & Writers

    Reader and writer obligations, lifecycle, adapters, backpressure, and conformance.

    Wire Format & Framing
    Stream Items
    Schema & State

    $name is used but no $name is defined anywhere in the header

    Reference to an undefined variable

    undefined-variable

    @name is used but no @name is defined anywhere in the header

    Undefined schema reference

    Undefined variable reference

    Reference order

    See Also

    Definitions
    Schema References
    Variables
    Error Model

    Prefer anyOf over a bare any when the set of acceptable types is known — it keeps validation meaningful.

    • Any · MemberDef

    • Schema References

    Constrained and structured alternatives

    Guidance

    Any

    See Also

    ~ $schema: { name: string, home: $address }
    ---
    ~ John, { Main St, NYC }    # ✗ undefined-schema — $address is never defined
    ~ $schema: { name: string, isActive: bool }
    ---
    ~ John, @active             # ✗ undefined-variable — @active is never defined
    ~ $schema: { name: string, home: $address }
    ~ $address: { street, city }
    ---
    ~ John, { Main St, NYC }    # ✓
    id: { any, anyOf: [string, int] }
    ---
    ~ 42      # ✓ matches int
    ~ abc     # ✓ matches string
    flag: { any, anyOf: [bool, int] }
    ---
    ~ T       # ✓
    ~ hello   # ✗ matches neither
    value: { any, anyOf: [ { int, multipleOf: 5 }, { int, multipleOf: 3 } ] }
    ---
    ~ 10      # ✓ multiple of 5
    ~ 9       # ✓ multiple of 3
    any
    members. See the
    type.

    Every schema member has a name. A data object can supply its values positionally (matched to the schema's members in order) or by key:

    Within one object, positional values MUST come before any keyed values. Once a keyed value appears, every later value MUST also be keyed. A positional value after a keyed one fails:

    Prefer positional data only when all members are required and the order is unambiguous; otherwise use keyed values for clarity and resilience to schema changes.

    A member's type can be a nested object schema or an array. Use { … } for an object and [ … ] for an array; an array's element type goes inside the brackets:

    Here tags is an array of strings and skills is an array of objects. See Array and Object (SchemaDef) for details.

    The default schema is the header definition named $schema. It is applied to every record in the data section:

    The name is case-sensitive: $Schema or $mySchema is an ordinary reference, not the default. A document with named references but no $schema has no default schema.

    Define a shape once with a $ reference and reuse it by name. Reference a shape with a keyed member (home: $address); a bare $address in a schema is read as a member literally named $address, not as the referenced shape:

    See Schema References for resolution rules and type references.

    When there is no $schema — an empty header, or a header with only metadata and references — each value is mapped to a positional index key ("0", "1", …):

    loads as:

    • Overview — the building blocks of a schema

    • Schema Data Types — the base types and shortcuts

    • MemberDef — constraints, optional, nullable, and defaults

    • Schema References — reusable $ schemas and types

    • — applying a schema across many records

    name, age, address
    ---
    John, 30, { Main St, NYC }
    name: string, address: { street: string, city: string }
    ---
    John, { Main St, NYC }

    Open and closed schema objects

    Overview
    Schema Data Types
    MemberDef
    Extensible & Dynamic Schemas
    ~ $schema: { name: string, age: int }
    ---
    ~ John, 30                # positional: maps to name, age
    ~ { name: Mary, age: 25 } # keyed: maps by name
    ~ $schema: { name: string, age: int }
    ---
    ~ { John, age: 30 }       # ✓ positional then keyed
    ~ { name: John, 30 }      # ✗ unexpected-positional-member
    ~ $schema: {
        name: string,
        address: { street: string, city: string, state: string },
        tags: [string],
        skills: [{ title: string, level: int }]
    }
    ---
    ~ Jane Doe, { X Street, New York, NY }, [lead, mentor], [{ Coding, 5 }, { Design, 3 }]
    ~ $schema: { name: string, age: int }
    ---
    ~ John, 30
    ~ Mary, 25
    ~ $address: { street: string, city: string, state: string }
    ~ $schema: { name: string, home: $address, work?: $address }
    ---
    ~ John Doe, { Main St, NYC, NY }, { 5th Ave, NYC, NY }
    ~ Jane Doe, { Oak Ave, LA, CA }
    ---
    John Doe, 25, { X-street, California, US }
    { "0": "John Doe", "1": 25, "2": { "0": "X-street", "1": "California", "2": "US" } }

    How data maps to a schema

    Mixing positional and keyed values

    Nested objects and arrays

    Choosing the schema

    The default schema

    Reusable named schemas

    No schema

    See Also

    any

    Decimal

    Fixed-precision decimal values for exact, financial-grade arithmetic.

    A Decimal is a fixed-precision decimal value for cases that demand exact numbers — especially financial calculations, where floating-point approximation can introduce errors. A Decimal stores an exact value with a defined precision and scale.

    Unlike a standard floating-point Number, a Decimal does not approximate: 0.1m is exactly one-tenth, not the nearest binary fraction.

    Syntax

    A Decimal is written as a number with the m suffix. It requires a leading digit and, if a decimal point is present, at least one digit after it. Scientific notation is not supported.

    decimal = ["-" | "+"] digit+ [ "." digit+ ] "m"
    
    digit = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9"

    Structural characters

    Symbol
    Name
    Unicode
    Description
    • 123.45m — fractional decimal

    • 123m — integer decimal

    • 0.001m — leading zeros preserved

    Each Decimal carries a precision (the total number of significant digits) and a scale (the number of digits after the decimal point):

    A schema can constrain these; see .

    A Decimal requires a leading digit and, with a decimal point, a trailing digit. Scientific notation is rejected:

    A doubled suffix is not an error. 123.45mm is an , by the same rule that makes 12mm one: nothing in it announced a decimal, because the m that would have done so is followed by more text. See .

    A plain 123.45 (no m) is a valid Number, not a Decimal — the m suffix is what selects fixed precision.

    Internet Object preserves:

    • Exact decimal precision and scale

    • Syntactic fidelity as written, except that an explicit + sign is not preserved

    It does not interpret:

    • Rounding behavior for operations

    • Currency or unit semantics

    Those semantics belong to the schema, the validator, or the application.

    • — all numeric forms

    • — standard floating-point numbers

    • — arbitrary-precision integers

    Stream Items

    The two-kind stream-item model the reader emits, record indexing rules, and degenerate inputs.

    A reader emits a sequence of stream items, in wire order. Each item corresponds to exactly one logical data record and is exactly one of two kinds: a record (success) or a record-error (recoverable failure). This page defines that abstract model. The concrete representation — object shape, field names, a sum type, an iterator of results — is chosen by each platform, but every implementation MUST preserve the model described here.

    Two kinds, never more. Implementations MUST NOT introduce additional item kinds in v1. The discriminant names the protocol event — record versus record-error — not the type of the payload. New metadata MAY be added only as optional, additive fields that do not change the meaning of the fields below.

    A record item represents one successfully parsed and validated record. It carries:

    Error Accumulation

    Accumulating per-object validation errors and per-region syntax errors.

    Rather than stopping at the first problem, a conformant processor accumulates errors and returns them together, alongside whatever data parsed successfully. This gives authors a full picture in one pass.

    Each record is validated independently and may contribute zero, one, or many errors. A failing record is marked as an error; the others are unaffected:

    The result contains two valid records and one error entry — not a single fatal failure.

    Syntax errors are accumulated per recovered region (between boundaries). One unparsable record yields one error, and parsing resumes at the next ~ or ---.

    Because errors are accumulated rather than thrown, the loaded result includes the records that succeeded. Consumers can render valid data and surface the error list side by side (for example, editor markers at each error's position).

    When two sections share a name, the duplicate is automatically renamed

    Syntax Errors

    Common syntax errors and how the parser recovers.

    A syntax error is a problem in the shape of the text — an unbalanced brace, a missing comma, an unterminated string — detected while tokenizing or parsing, before any schema validation. (Errors about values — wrong type, out of range — are validation errors; see .)

    The unclosed { raises expected-closing-bracket. (The predicate unterminated- is reserved for constructs the tokenizer closes, such as a quoted string; a bracket is closed by the parser.)

    Values are comma-separated. Where a comma is omitted between two values, the text between them is a single , and an open string may contain spaces — so the result is one well-formed value, and there is nothing for a parser to reject:

    A conformant parser MUST NOT report an error here, and MUST NOT

    Data Streaming

    How collections enable streaming of records.

    Because a collection is a sequence of independent records, it is naturally streamable: a producer can emit records over time and a consumer can process each as it arrives, without waiting for the whole document.

    Further records for the same collection can be sent later in additional batches; a processor merges them into the same collection. Each record is validated on its own, so a malformed record does not interrupt the stream.

    The full protocol lives in its own chapter. Framing, the stream-item model, schema and state, the error model, and reader and writer obligations are specified normatively in the chapter. This page only shows why a collection is streamable; the contract is there.

    • — the normative, platform-agnostic streaming protocol

    Email

    The email type — a string validated as an email address.

    email is a string shortcut (see ) whose value MUST be a valid email address. It shares the string (choices, pattern, minLen, …) and adds email-format validation.

    Restrict to a fixed set with choices:

    Abstract

    A text-based, schema-first, document-oriented, streamable data interchange format.

    Internet Object is a data interchange format designed for modern web communication over the internet. This specification introduces Internet Object as a text-based, schema-first, document-oriented, and streamable format that prioritizes human readability and language independence. By separating the schema from the data, it serializes structured data compactly for efficient transmission between servers and clients across the web, while preserving the clarity and approachability of a plain-text format.

    • — a guided walkthrough

    • ·

    Manifesto

    A declaration of the convictions behind Internet Object — why it exists and what it refuses to compromise.

    Data is the substance of the Internet, and we still move it wastefully.

    JSON proved that humans must be able to read their data — and then it stayed wasteful, lossy, and schemaless by default. The binary formats proved that the wire must be small — and then they made data unreadable and forked their schemas into a separate toolchain. Both were half-right. We accepted the choice between them for twenty years. We don't have to anymore.

    Internet Object begins from a refusal: we will not choose between readable and efficient, between expressive and exact, between human and machine. A format for the Internet must serve all of them, at the scale the Internet actually runs.

    These are the convictions it is built on.

    1. The wire is not free. A format used everywhere, constantly, has no right to be wasteful. Every key repeated on every record is bytes paid for again and again — in storage, in bandwidth, in energy, in latency. We declare structure once and never repeat it.

    2. Structure is not data. Keys, types, and constraints belong to the schema. A record carries only its values. Conflating the two was the inherited mistake; separating them is what makes data compact, validated, and self-describing all at once.

    Collection Rules

    See Also

    Introducing Internet Object
    Objectives
    Why Internet Object?
    3. The model is the contract, not the bytes. A schema and its data are one truth, validated one way. Whether that truth is serialized as human-readable text or as a compact binary form, it means the same thing and validates the same way. The encoding is a choice; the meaning never changes.

    4. Data must tell the truth. A decimal is exact. A bigint is whole. A date is a date, and binary is binary. No silent coercion, no precision lost in transit. Where other formats shrug, Internet Object is precise.

    5. Meaning belongs in the format. Validation is not a second language in a second file. The shape of data — its types, its constraints — is declared with the data, in the same syntax as the data, and enforced as a property of the format itself.

    6. The Internet is a stream, not a file. Data arrives over time, in volume, between parties who often already agree on its shape. Documents, collections, and streaming are first-class — not features bolted onto a format that assumed everything fits in memory at once.

    7. Open to all. Locked to none. Internet Object is an openly specified contract on the wire, not a library in a language or a product from a vendor. Anyone can implement it; no one can lock you into it. Data you store today will still be understood long after the tools that wrote it are gone.


    Internet Object is a promise: that data can be readable and small, expressive and validated, precise and fast — bound to no language and no vendor.

    Not a better library. A better default.

    • Abstract — the technical statement of what Internet Object is

    • The Poetic Principles — the same values in verse

    • Objectives — the concrete goals the format is designed to meet

    • Why Internet Object? — how it compares to JSON, CSV, and YAML

    See Also

    Separates the integer and fractional parts

    -

    Minus sign

    U+002D

    Negative value

    -789.01m — negative decimal
  • 0m, 0.0m — zero, with and without a scale

  • m

    Decimal suffix

    U+006D

    Marks the value as a Decimal

    0–9

    Digits

    Multiple

    Decimal digits

    .

    Decimal point

    Valid forms

    Precision and scale

    Invalid forms

    Preservation of structure

    See Also

    Decimal
    open string
    A number, or a word that begins with a digit?
    Numeric Values
    Number
    BigInt

    U+002E

    Collection · Collection Rules

  • Creating Collections

  • ~ $schema: { name: string, address: { street, city, state }, active: bool }
    ---
    ~ John Doe, { Red Street, Phoenix, AZ }, T
    ~ Alex, { Carnival Street, San Francisco, CA }, T

    See Also

    Streaming
    Streaming
    ·
  • MemberDef

  • userEmail: email
    ---
    ~ [email protected]    # ✓
    ~ notanemail          # ✗ invalid-email

    See Also

    String Types
    MemberDef
    String Types
    URL
    so the rest of the document still loads. A recovering parser
    MUST NOT
    drop a section, and
    MUST NOT
    let one overwrite another: either would lose data with nothing to show for it.

    The renaming rule, stated exactly, because two implementations that disagree here produce differently-named sections from the same document:

    On encountering a section whose name is already in use, append _2 to the original name. If that name is also in use, try _3, then _4, and so on, until an unused name is found. The counter is per name, not per document, and it counts names already taken — including names a later section spelled out for itself.

    So a document with three users sections yields users, users_2, users_3; and a document with sections named a, b, a, b yields a, b, a_2, b_2 — not a_2, b_3.

    Because the rule counts names already taken, an explicit name cannot be silently displaced:

    This applies to sections that carry no name of their own, too: they take the default name data, so three unnamed sections become data, data_2, data_3.

    The document is still invalid: section names must be unique, and the error is reported alongside the recovered data. The error is duplicate-section-name — a structural fault, not a lexical one, since every character in the document is valid.

    • Error Codes · Error Model · Parser Behavior & Recovery

    • Collection Rules

    Per-object validation errors

    Per-region syntax errors

    Partial output

    Duplicate section names

    See Also

    insert the missing separators. This is the one fault the format cannot diagnose for you, and it is the direct cost of having open strings at all: the same property that lets
    New York
    be written without quotes is what makes
    101 Thomas 25
    a single value.

    The rule stops at values. A missing separator before a key is still an error, because a key cannot follow an unseparated value:

    A quoted string with no closing quote raises unterminated-string. This applies to every quoted form, including the annotated ones (r'...', b'...', dt'...'):

    On a syntax error the parser skips ahead to the next boundary — a record separator ~ or a section separator --- (or end of file) — records the error, and resumes. So one malformed record does not prevent later records from being parsed. See Parser Behavior & Recovery.

    • Error Model

    • Parser Behavior & Recovery

    • Comments

    Common syntax errors

    Unbalanced brackets

    Missing separators merge values

    Error Model
    open string

    Unterminated string

    Recovery is bounded by structure

    See Also

    ---
    123.45m, 123m, 0.001m, -789.01m, 0m, 0.0m
    123.45m              # precision 5, scale 2
    0.000123m            # precision 6, scale 6
    ---
    1.23e2m              # ✗ scientific notation is not supported for Decimal
    .45m                 # ✗ invalid-decimal — missing leading digit (use 0.45m)
    123.m                # ✗ invalid-decimal — missing trailing digit (use 123.0m or 123m)
    ---
    123.45mm             # → "123.45mm" — a string; write "123.45m" for the decimal
    companyEmail: { email, choices: [[email protected], [email protected]] }
    ---
    ~ [email protected]       # ✓
    ~ [email protected]      # ✗ mismatched-choice
    ~ $schema: { name: string, age: { int, max: 25 } }
    ---
    ~ James, 20    # ✓
    ~ Alex, 30     # ✗ mismatched-max
    ~ Bob, 22      # ✓
    --- users        → users
    --- users_2      → users_2   (written that way by the author)
    --- users        → users_3   (skips users_2, which is taken)
    pt: { object, schema: { x: int } }
    ---
    { 1                      # ✗ expected-closing-bracket  (the '{' is never closed)
    ---
    ~ 101 Thomas 25      # → "101 Thomas 25" — one value, not three
    {John age: 25 gender: M}   # ✗ unexpected-token
    "John Doe            # ✗ unterminated-string
    r'C:\path           # ✗ unterminated-string
    Field
    Meaning

    kind

    record (success).

    recordIndex

    The zero-based position of this record in the stream (see below).

    schemaName

    Present only when an explicit schema selector applied; otherwise absent. See .

    value

    The complete parsed record value — identical to what the non-streaming core path produces for that record.

    A record item MUST carry a complete record value. It MUST NOT carry an error.

    A record-error item represents one record that failed recoverably. Iteration continues to the next record. It carries:

    Field
    Meaning

    kind

    record-error (recoverable failure).

    recordIndex

    The zero-based position of this record — counted exactly as a successful one.

    schemaName

    A record-error item MUST carry no value. An item MUST NOT carry both a value and a recoverable error — the two kinds are mutually exclusive.

    recordIndex is the stable identity of a record's position in the stream. Its rules are strict:

    • It is zero-based and counts logical data records — not chunks, lines, or sections.

    • It MUST increment for both kinds. A failed record consumes an index exactly as a successful one does, so indices are dense and gap-free.

    • It is stream-global. It MUST NOT reset on a schema switch or on a bare ---.

    • Items MUST be emitted in wire order: the item with index n before the item with index n+1.

    Consider a stream whose second record is malformed:

    The reader emits three items in order: a record at index 0, a record-error at index 1, and a record at index 2. The bad record still consumes index 1; the third record is 2, not 1. The reader MUST NOT leak partial fragments of the broken record before its error item.

    The reader's behavior on trivial inputs follows directly from the framing rules and the equivalence rule — these cases are specified so every implementation agrees:

    • An empty source (zero bytes) MUST emit zero items and complete normally.

    • A whitespace-only or blank-line-only source MUST emit zero items and complete normally.

    • A header-only stream (definitions, then ---, then end of stream with no records) MUST emit zero items and complete normally.

    • A trailing bare --- with no following records MUST NOT emit an item and MUST NOT error.

    • A bare ~ with no payload (or ~ followed only by whitespace) is delegated to core like any other record. Under an active schema it yields whatever core produces — typically a validation failure, hence one record-error — and with no active schema it yields core's empty record. Streaming MUST NOT special-case it into a stream-only error, because that would diverge from the non-streaming parse of the same input.

    • Schema & State — when schemaName is present and how selectors work

    • Streaming Error Model — what a record-error item's error preserves

    • Wire Format & Framing — how ~ records and --- frames are read off the wire

    • — the record sequence each item maps to

    The record item

    ---
    ~ { id: 1 }
    ~ { BROKEN
    ~ { id: 2 }

    The record-error item

    Record indexing

    Degenerate and empty inputs

    See Also

    Collection Rules

    Validation rules for collections — schema-less records, empty records, errors.

    Records without a schema

    If no schema is defined, each record may have a different shape, and values are mapped to positional indices (0, 1, 2, …):

    ---
    ~ John Doe, 20, female
    ~ true, false
    ~ marketing, 123, { Z Street, Los Angeles, CA }

    The first record loads as { "0": "John Doe", "1": 20, "2": "female" }, and so on.

    It is good practice to define a schema even though collections allow schema-less records.

    Empty records

    A record consisting of just ~ is an empty object ({}). It is valid only if every field in the schema is optional and/or nullable:

    If any field is required, an empty record fails:

    Each record is validated independently. A failing record is marked as an error; the rest are unaffected:

    A conformant processor SHOULD collect per-record errors and continue, rather than stopping at the first failure.

    • ·

    Overview

    Overview of the parsing pipeline and the two error classes.

    Turning Internet Object text into validated data happens in stages:

    1. Tokenize — split the text into tokens (values, separators, structural characters).

    2. Parse — assemble tokens into a document tree (header, sections, records, values).

    3. Validate — check the data against the schema.

    4. Load — produce the final in-memory values.

    Errors fall into two classes, matching the stages that produce them:

    Class
    Stage
    Example

    The distinction matters because the two classes recover differently:

    • Syntax errors are bounded by structure — the parser skips to the next boundary (~ or ---) and continues.

    • Validation errors are bounded by the object — each record is validated on its own and may report zero, one, or many errors, without affecting other records.

    Whichever class it belongs to, every error reports the same three things: a stable code, a human-readable message, and the position in the source. Only the code is part of this specification — messages may be reworded or translated — so tooling branches on the code.

    • — how every code is named, and the closed vocabulary it draws from

    • — the catalogue of codes, by class

    • — how recovery works; processing options

    • — collecting many errors and partial output

    Structural Elements

    The characters and tokens that structure and delimit an Internet Object document.

    The Internet Object format uses a small set of structural characters, literals, and other special characters to structure and delimit data. Working together with objects, strings, arrays, numbers, and whitespace, these elements compose the format's grammar and let documents express complex, flexible data structures.

    Categories

    • Structural Characters & Separators — core syntax characters that organize data

    • Literals — predefined constant values (booleans, null, special numbers)

    • — functional modifiers for variables, schemas, and values

    • — recognized Unicode whitespace and its handling

    • — data types and how values are written

    • — comment syntax and usage

    • — character encoding and Unicode support

    Definitions

    The header's definition section — metadata, variables, and references.

    Besides the schema, an Internet Object document's header can hold definitions. A definition is a key–value pair on its own line, introduced by a tilde ~:

    There are three kinds of definition, distinguished by the key's prefix:

    Prefix
    Kind
    Purpose

    See Also

    Other Special Characters
    Whitespace & Indentation
    Value Representations
    Comments
    Encoding

    Same rule as the record item: present only when an explicit selector applied.

    error

    The failure, preserving the core error identity (category + code). See Streaming Error Model.

    Collection
    Schema & State

    Syntax error

    tokenize / parse

    unbalanced {, missing comma, unterminated string

    Validation error

    validate

    In this section

    See Also

    Error Codes
    Error Model
    Parser Behavior & Recovery
    Error Accumulation
    Syntax Errors
    Collection Rules

    wrong type, out of range, missing required field

    Independent validation & error handling

    See Also

    Collection
    Creating Collections
    Parser Behavior & Recovery
    ~ $schema: { name?*: string, age?*: { int, max: 25 } }
    ---
    ~ John, 25     # ✓
    ~ William      # ✓ (age omitted)
    ~              # ✓ (empty object; all fields optional/nullable)
    ~ $schema: { name: string, age?*: { int, max: 25 } }
    ---
    ~ John, 25     # ✓
    ~              # ✗ missing-value — name is required
    ~ $schema: { name: string, age: { int, max: 25 } }
    ---
    ~ James, 20    # ✓
    ~ Alex, 30     # ✗ mismatched-max — age exceeds 25
    ~ Bob, 22      # ✓

    Metadata

    Document-level data (paging, status, …). Surfaces in the output header.

    @

    Value variable

    A reusable value referenced as @name. See .

    $

    Reference (ref)

    A reusable schema or type referenced as $name. See .

    The special key $schema names the document's default schema.

    Bare keys carry document metadata. They appear under a header in the loaded result, separate from the data:

    @ defines a value variable; $ defines a reusable schema (a ref). Both are then used by name:

    A document may be header-only. If there is no data, the header still ends with the --- separator.

    • Header — where definitions live in a document

    • Variables — value variables (@)

    • Schema References — schema and type refs ($)

    ~ key: value

    (none)

    ~ pageSize: 10
    ~ success: T
    ~ $schema: { name: string }
    ---
    ~ John
    ~ Jane
    ~ @active: T
    ~ $address: { street, city }
    ~ $schema: { name: string, addr: $address, isActive: bool }
    ---
    ~ John, { Main St, NYC }, @active
    ~ Jane, { Oak Ave, LA }, @active

    Metadata

    Variables and references

    See Also

    Date and Time

    Temporal values — dates, times, and date-times as annotated strings.

    Temporal values are written as annotated strings in ISO 8601-compatible formats, with a prefix that selects the kind: d for a date, t for a time, and dt for a combined date-time. A parser converts them to native date/time objects on deserialization.

    The content between the quotes must be valid for its kind.

    Syntax

    temporalValue = dateValue | timeValue | dateTimeValue
    dateValue     = "d"  ("'" dateContent "'"     | '"' dateContent '"')
    timeValue     = "t"  ("'" timeContent "'"     | '"' timeContent '"')
    dateTimeValue = "dt" ("'" dateTimeContent "'" | '"' dateTimeContent '"')
    
    dateContent     = yearPart [monthPart [dayPart]]
    timeContent     = hourPart [minutePart [secondPart [millisecondPart]]]
    dateTimeContent = dateContent ["T" timeContent] [timeZone]
    
    yearPart        = digit digit digit digit
    monthPart       = ["-"] ( "0" digit | "1" ("0" | "1" | "2") )
    dayPart         = ["-"] ( "0" digit | ("1" | "2") digit | "3" ("0" | "1") )
    hourPart        = [":"] ( ("0" | "1") digit | "2" ("0" | "1" | "2" | "3") )
    minutePart      = [":"] ("0" | "1" | "2" | "3" | "4" | "5") digit
    secondPart      = [":"] ("0" | "1" | "2" | "3" | "4" | "5") digit
    millisecondPart = "." digit digit digit
    timeZone        = "Z" | ("+" | "-") hourPart [minutePart]

    Structural characters

    Symbol
    Name
    Unicode
    Description
    Kind
    With separators
    Without separators
    Partial forms
    Defaults

    Hours use the 24-hour clock (00–23). A date-only value is treated as UTC midnight; a time-only value uses a reference date and is treated as UTC.

    Out-of-range or malformed temporals produce the stable error code invalid-datetime:

    • Explicit UTC — dt'2024-03-20T14:30:45Z' is UTC.

    • Explicit offset — dt'2024-03-20T14:30:45+05:30' carries that offset; its UTC equivalent is 2024-03-20T09:00:45Z.

    • No timezone — a date-time with no zone is treated as UTC.

    The reference implementation has known gaps in temporal parsing; a conformant parser SHOULD behave as specified above, not as the current implementation does:

    • Overflow dates are normalized rather than rejected: d'2024-02-30' becomes 2024-03-01 instead of raising invalid-datetime.

    • Over-long fractions are truncated: a millisecond field with more than three digits is truncated to three rather than rejected.

    • — all value types

    • — the numeric forms

    • — date/time schemas

    TypeDef

    TypeDef — the fixed option contract that every MemberDef of a type is validated against.

    A TypeDef is the fixed option contract for a built-in type. It defines exactly which options a MemberDef of that type may use, the order of its positional values, and the type of each option. A TypeDef is itself written in Internet Object object syntax, so the same rules that apply to objects apply to it.

    TypeDef vs. MemberDef. A TypeDef is the spec-defined contract for a type; a MemberDef is what a schema author writes for one member. Every MemberDef is validated against the TypeDef of its declared type.

    Positional and keyed options

    A TypeDef fixes the meaning of each positional value. For every type the order is:

    1. type — the type name (number, int16, string, …)

    2. default — the value used when the member is omitted

    3. choices — the allowed set of values

    Any further options are written as keyed entries (min: 0, pattern: …) in any order after the positional ones. So all of these are valid number MemberDefs:

    A TypeDef written out looks like an ordinary object whose members are the options (most of them optional). Illustratively, the number TypeDef is shaped like this:

    Each type page lists its own TypeDef in full; that table is the authoritative source of the options a type accepts. optional and null are the two every type defines — usually written with the ? / * shorthand on the member name, and as keyed options when the name is quoted (see ). The null key is written quoted, since a bare null is the null keyword.

    A MemberDef may use only the options its type's TypeDef defines. An unknown option is rejected with unknown-member:

    This is what makes options portable: because the contract is fixed, the same MemberDef validates identically in every conformant implementation.

    An option is either a constraint (it can reject a value) or a presentation option (it only selects how a value is written, and never rejects anything). Which is which is fixed per option, not per implementation — see .

    A schema author cannot change or extend a built-in type's TypeDef — the option set is defined by this specification. To build a reusable, document-local type (a base type plus preset constraints), define a type reference in the header instead; see .

    • — how authors write a member's type and constraints

    • — every type's TypeDef table

    • · — example TypeDefs

    Binary

    The binary type — byte data written as base64.

    The binary type validates a sequence of bytes. In data, binary is written as a base64 literal with a b prefix: b'SGVsbG8='. (Base64 is the encoding; binary is the type.)

    For the literal value syntax, see Binary values.

    TypeDef

    A binary MemberDef accepts the options below.

    Option
    Type
    Description

    Binary is under development. The binary schema type is being added to the reference implementation (it is not yet registered), and the b'…' value literal is not yet accepted by the parser. The design above is the agreed target; examples are not yet executable and are excluded from the example verifier.

    • ·

    Internet Object 1.0

    Thin, schema-first and robust data-interchange object format for Internet

    Internet Object is a text-based, schema-first, document-oriented data interchange format — JSON's readability without its repeated keys, missing schema, absent comments, or lack of streaming. This site is the official specification and serves as the format's reference documentation.

    Status. This is the 1.0 Draft of the specification, published alongside the public beta of the reference implementation. The specification and the implementation are still converging, so some pages note behavior that is ahead of the current implementation. Implementations version independently and declare the specification version they conform to.

    Start here

    • Manifesto — the convictions behind the format, and why it exists

    • — how it compares to JSON, CSV, and YAML

    • — a five-minute tour in pure Internet Object

    • — the structure of a document

    • — types, constraints, and validation

    • — what a conformant implementation must do

    Field
    Value

    Header

    The header section — schemas, definitions, variables, and metadata.

    The header sits at the beginning of an Internet Object document and defines the schema and definitions for the data that follows. It carries the metadata, context, variables, and schema references needed to interpret the data consistently. By stating this information once, up front, the header keeps the data section compact and unambiguous.

    A schema defines the structure and meaning of the data in a document. When the header contains only a schema — with no other definitions — that schema is the document's default schema. It describes the shape of the data while keeping the structure separate from the data itself, which makes the data more compact and easier to process.

    This header declares five members:

    1. name — an untyped member, typically a string.

    Objectives

    The design goals that shape the Internet Object format.

    The Internet Object serialization format aims to redefine data interchange on the internet by addressing key challenges and limitations present in existing formats.

    The inception of Internet Object began as a side project aimed at addressing limitations observed in the JSON format. Over time, it evolved into an independent research endeavor focused on tackling data-transfer challenges such as size, schema validation, data streaming, header and metadata support, and more. The design of the Internet Object format revolves around the following key objectives:

    To optimize the format for internet wire transfer, Internet Object MUST be conceived and developed without being excessively influenced by existing mechanisms. However, it MAY draw inspiration from other formats as needed.

    Internet Object documents SHOULD be text-based, human-friendly, and easy to work with. Developers SHOULD be able to write these documents using plain-text IDEs without needing frameworks, libraries, or utilities.

    To ensure a small footprint, the Internet Object format SHOULD separate data and schema, allowing data to be sent alone over the network.

    To uphold data integrity during wire transfer, the Internet Object format SHOULD prioritize a schema-first approach.

    Embracing a comprehensive document-oriented approach, the Internet Object format SHOULD facilitate the bundling of essential components—including records, data, definitions, schemas, and comments—within a single document. This approach helps keep related information together, improving organization and maintainability.

    The Poetic Principles

    This poem encapsulates the core guiding principles that shape the design and objectives of the Internet Object format.

    This poem distils the foundational principles of Internet Object into a few memorable verses. It is an informative companion to the specification: the lines below restate, in artistic form, the design values explained throughout these pages — small size, readability, the separation of data and definitions, the independence of records, and a healthy distrust of unvalidated input.

    Size holds weight, in bytes confined, Small prevails, large left behind.

    Simplicity shines over complexity's shroud, Readability echoes, accurate and loud.

    Reusability births productivity's rise, Verbosity's burden efficiency defies.

    Data, definitions, separate ways, Together they clutter, apart they amaze.

    Headers and data, distinctions drawn, Confusion dissolves, clarity's dawn.

    Errors and statuses, data's divide, Their entanglement brings chaos inside.

    Internet Object Schema
    Variables
    Schema References

    Internet Object MUST support complex data types so that large numbers and complex data structures can be serialized and deserialized efficiently for the wire.

    The Internet Object format SHOULD support streaming of independent records, allowing for efficient and continuous data transfer. The failure of a single record MUST NOT affect the processing of other records.

    The Internet Object format SHOULD work seamlessly across platforms, operating systems, and programming languages to ensure broad adoption and versatility.

    By supporting inline comments, the Internet Object format allows users to document schemas and definitions directly within the data itself. This feature enhances readability and maintainability.

    To increase adaptability, Internet Object SHOULD promote reusability through references and variables. This capability enables customization of data structures and more effective data manipulation.

    • Abstract — the format in one paragraph

    • Introducing Internet Object · Why Internet Object?

    • The Poetic Principles — the same goals in verse

    Uninfluenced development

    Human friendly

    Minimal footprint

    Schema first

    Document oriented

    Complex data types

    Streaming friendly

    Platform and language independence

    Comments

    Reusability

    See Also

    Two lone records, states unswayed, No interference, connections unmade.

    Trust not the sender, vigilance displayed, Expect the unanticipated, foundations laid.

    Surprises, enchanting, yet beware, Not all of them good, handle with care.

    The verses echo the format's objectives:

    • Small over large — compact payloads; keys live in the schema, not in every record.

    • Readability and simplicity — plain text that is easy to read and to write by hand.

    • Separation of data and definitions — the header holds schema and metadata; the data stays clean below it.

    • Record independence — records do not depend on one another, so one bad record never breaks the rest.

    • Distrust the sender — validate incoming data and expect the unexpected.

    • Objectives — the design goals stated plainly

    • Abstract · Introducing Internet Object

    Poem

    The principles behind the verses

    See Also

    Combined date-time value

    ' "

    Quotes

    U+0027 U+0022

    Enclose the temporal content

    -

    Hyphen

    U+002D

    Date separator (optional); also a negative offset

    :

    Colon

    U+003A

    Time separator (optional)

    .

    Period

    U+002E

    Millisecond separator

    T

    Letter T

    U+0054

    Separates the date and time

    Z

    Letter Z

    U+005A

    UTC designator

    +

    Plus

    U+002B

    Positive timezone offset

    Time

    HH:mm:ss.SSS

    HHmmss.SSS

    HH:mm:ss, HH:mm, HH

    missing parts → 00

    DateTime

    date T time [zone]

    —

    any date form + optional time

    missing time → 00:00:00.000; missing zone → UTC

    Valid offset range — offsets run from -12:00 to +14:00; both ±HH:mm and ±HHMM are accepted on input.

  • On serialization, an implementation SHOULD emit offsets in the ±HH:mm form and round-trip explicit timezone information unchanged.

  • d

    Date prefix

    U+0064

    Date-only value

    t

    Time prefix

    U+0074

    Time-only value

    dt

    DateTime prefix

    Date

    YYYY-MM-DD

    YYYYMMDD

    YYYY-MM, YYYY

    Valid forms

    Dates — d'…'

    Times — t'…'

    Date-times — dt'…'

    A combined example

    Format specifications

    Invalid forms

    Timezone handling

    Implementation status (beta)

    See Also

    Value Representations
    Numeric Values
    Date and Time

    —

    missing month/day → 01

    Validation against the TypeDef

    TypeDefs are fixed

    See Also

    MemberDef
    Constraints and presentation
    Schema References
    MemberDef
    Schema Data Types
    Numeric Types
    String Types

    array of binary

    Restricts the value to a fixed set.

    len

    int ≥ 0

    Exact length in bytes.

    minLen

    int ≥ 0

    Minimum length in bytes.

    maxLen

    int ≥ 0

    Maximum length in bytes.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix.

    null

    bool

    If true, the member may be null. Shorthand: * suffix.

    type

    string

    The type name binary. First positional value.

    default

    binary

    Value used when the member is omitted.

    Example

    Implementation status (beta)

    See Also

    Binary values
    TypeDef
    MemberDef

    choices

    age: int — an explicitly typed member that must hold an integer.

  • address — an untyped member that may hold a string or a nested object.

  • isActive? — the ? suffix marks the member as optional; it may be omitted from the data.

  • remark — an untyped member, typically a free-text note.

  • Alongside the structure, the schema records type annotations and optionality, which sharpens validation and documents the data model in one place. For the full schema syntax, see the Internet Object schema.

    Definitions are key-value pairs declared in the header to hold metadata, variables, reusable schemas, and other shared values. Each definition is written on its own line, prefixed with ~.

    Here the header mixes response metadata with schema definitions instead of using a default schema. The metadata records the page size (pageSize), the current page (currentPage), and the total record count (recordCount). It also defines a reusable address shape ($address) with the members street, city, and state, and a top-level schema ($schema) that references it. The $schema key is reserved: it names the default schema applied to the data section.

    For the full treatment of metadata, value variables (@), and references ($) — including how they are resolved — see the Definitions chapter.

    • Data Sections — what follows the --- separator

    • Definitions — variables and schema references in depth

    • Internet Object Schema — the schema language

    Default schema

    Definitions

    See Also

    d'2024-03-20'            # full date
    d'2024-03'              # year and month (day defaults to 01)
    d'2024'                # year only (month and day default to 01)
    d'20240320'            # without separators
    d"2024-12-31"          # double quotes
    t'14:30:45.123'         # with milliseconds
    t'14:30:45'             # hour, minute, second
    t'14:30'                # hour and minute (second defaults to 00)
    t'14'                   # hour only
    t'143045'               # without separators
    dt'2024-03-20T14:30:45.123Z'   # full, UTC
    dt'2024-03-20T14:30Z'          # without seconds
    dt'2024-03-20T14:30:45+05:30'  # with a timezone offset
    dt'2024-03-20'                 # date only (time defaults to 00:00:00.000)
    dt"2024-12-31T23:59:59.999Z"   # double quotes
    ---
    d'2024-03-20', t'14:30:45.123', dt'2024-03-20T14:30Z'
    ---
    d'2024-13-20'           # ✗ invalid-date
    t'25:00:00'             # ✗ invalid-time — hour out of range
    t'12:60:00'             # ✗ invalid-time — minute out of range
    dt'2024-03-20 14:30:00' # ✗ invalid-datetime — missing the T separator
    dt'2024-03-20T14:30+25:00'  # ✗ invalid-datetime — timezone offset out of range
    ~ $schema: {
        a: { number, 20 },                        # type + default
        b: { int16, 1, [1, 2, 3] },               # type + default + choices
        c: { number, 50, min: 10, max: 99 },      # default + keyed options
        d: { number, 10, [5, 10, 15], min: 5 }    # default + choices + keyed option
    }
    ---
    ~                       # all omitted → defaults a 20, b 1, c 50, d 10
    ~ 25, 3, 60, 15         # a 25, b 3, c 60, d 15
    type?       : string,        # a number-family name (see Numeric Types)
    default?    : number,
    choices?    : [number],
    min?        : number,
    max?        : number,
    multipleOf? : number,
    format?     : { string, choices: [decimal, hex, octal, binary, scientific] },
    optional?   : { bool, F },
    null?       : { bool, F }
    age: { number, minimum: 10 }
    ---
    42                       # ✗ unknown-member — number has no option "minimum" (use min)
    avatar: { binary, maxLen: 65536 }
    ---
    ~ b'SGVsbG8='
    name, age: int, address, isActive?, remark
    ---
    ~ pageSize: 1
    ~ currentPage: 1
    ~ recordCount: 4
    ~ $address: {street, city, state}
    ~ $schema: {name, age, $address}
    ---

    Status

    Work-in-Progress Draft

    Last updated

    2026-09-07

    Website

    Docs

    License

    Spec: · Examples: CC0 · Libraries: Apache-2.0

    Author and Researcher

    Mohamed Aamir Maniar at ManiarTech® Lab

    Contact

    [email protected]

    Version

    Why Internet Object?
    Getting Started
    Internet Object Document
    Schema Definition Language
    Conformance Requirements

    1.0 Draft

    Raw Strings

    Raw strings — literal strings where backslashes are not escapes.

    A raw string is a sequence of Unicode code points prefixed with a lower-case r and enclosed in single quotes (', U+0027) or double quotes (", U+0022). Raw strings suit text with many backslashes, quotes, or structural characters — file paths or regular expressions, for example. They process no escape sequences except the enclosing quote, which is written by doubling it inside the string.

    Raw strings are scalar values. They preserve all content as written, including whitespace, newlines, and Unicode characters.

    Syntax

    A raw string is prefixed with a lower-case r and enclosed in single or double quotes. The prefix is case-sensitive, like everything else in Internet Object: R'...' is not a raw string, and reports unknown-annotation. The same holds for every annotation prefix — b, d, t, dt. The only special rule is that the enclosing quote, when it appears inside the string, must be written as two consecutive enclosing quotes.

    Symbol
    Name
    Unicode
    Description

    Backslash is literal. The reverse solidus (\, U+005C) is always a literal character in a raw string — there is no backslash escaping.

    Examples of valid raw strings:

    • Whitespace — leading, trailing, and internal whitespace are preserved.

    • No escaping — no escape sequences are processed except doubling the enclosing quote.

    • Multiline — newline and carriage-return characters are preserved.

    Comments are not allowed inside raw strings, but may appear outside or between values, per the format's comment rules.

    To hold a quote of the same kind, double it: r'Jonas D''costa' and r"He said, ""Hello!""".

    Without quotes there is no annotation, so nothing marks the text as raw. That is not an error; it is simply read by the ordinary rules — here as a keyed member, because of the colon:

    Internet Object preserves:

    • All Unicode code points and whitespace as written

    • The doubled-quote convention for an embedded enclosing quote

    It does not interpret or enforce:

    • Application-specific constraints

    • Any escaping beyond doubled enclosing quotes

    • — the three string forms

    • — the unquoted form

    • — quoted strings with escaping

    Binary

    Binary values written as Base64 byte strings.

    A binary value is a sequence of raw bytes carried as text. It is written as a Base64 byte string: the prefix b followed by Base64 content in single or double quotes. Binary values suit images, encrypted content, cryptographic keys, or any arbitrary byte sequence in an otherwise text-based document.

    The content between the quotes is Base64 per RFC 4648. Base64 is the encoding; binary is the value type.

    Implementation status (beta). Binary literals are not yet available in the reference implementation: b'…' and b"…" currently raise a syntax error (unexpected-token), and no binary schema type is registered. This page documents the intended design; the examples below are illustrative and are not yet executable. Track progress in the .

    A binary value is prefixed with b and enclosed in single or double quotes; the content must be valid Base64.

    Symbol
    Name
    Unicode
    Description

    Without quotes there is no annotation at all, so nothing announces a binary literal and the text is an ordinary :

    • Whitespace — leading and trailing whitespace around the quotes is ignored; whitespace inside the Base64 content is not allowed.

    • Prefix case — the prefix MUST be lower-case b; the Base64 content is case-sensitive.

    • Padding — standard = padding is required for correct decoding.

    Internet Object does not interpret the structure of the decoded bytes — any format, compression, or application meaning is the application's concern.

    • — all value types

    • — the binary schema type

    • — text values (open, regular, raw)

    Encoding

    Character encoding — UTF-8 is mandatory; Unicode, BOM, and line endings.

    The Internet Object format uses UTF-8 as its default and mandatory encoding for all text. This ensures reliable interchange across platforms, systems, and programming languages.

    UTF-8 requirement

    Every conformant implementation MUST support UTF-8. UTF-8 is chosen because it is:

    • Universal — supported by virtually all modern systems and languages.

    • ASCII-compatible — the ASCII range (0–127) is encoded identically.

    • Complete — it can represent every Unicode character.

    • Byte-order independent — no endianness concerns, unlike UTF-16 or UTF-32.

    • Self-synchronizing — corruption of one character does not derail later parsing.

    UTF-8 is mandatory; an implementation MAY additionally support other encodings for specific needs.

    Encoding
    Support
    Notes

    UTF-8 is the baseline. If another encoding fits your situation, convert to or from UTF-8 at the boundary. Because UTF-8 is the only mandatory encoding, every parser and serializer must handle it.

    Internet Object supports the full Unicode character set through UTF-8:

    • Basic Multilingual Plane — U+0000 to U+FFFF.

    • Supplementary planes — U+10000 to U+10FFFF.

    • Control characters — handled per the Unicode standard; in strings they should be escaped.

    For normalization, NFC (Normalization Form Canonical Composed) is the recommended form. An implementation should normalize consistently when comparing strings; the internal storage form is unconstrained.

    • A UTF-8 BOM is the byte sequence EF BB BF (U+FEFF) at the start of a document.

    • A parser treats a leading BOM as whitespace and ignores it, so a BOM never causes a parse error.

    • A BOM is optional and not recommended for UTF-8; if you use one, do so consistently.

    All common line-ending conventions are accepted and treated equivalently:

    • Unix/Linux — LF (\n)

    • Windows — CRLF (\r\n)

    • Classic Mac — CR (\r)

    Mixed line endings within one document are handled gracefully.

    A conformant parser SHOULD:

    1. Accept UTF-8 input and skip a leading BOM if present.

    2. Report a clear error for invalid UTF-8 byte sequences and reject overlong encodings.

    3. Handle UTF-16 surrogate pairs correctly when decoding \u escape sequences.

    A conformant serializer SHOULD:

    1. Always emit valid UTF-8.

    2. Be consistent about including or omitting a BOM for the target system.

    3. Emit escape sequences for control characters where needed.

    • — recognized whitespace characters

    • — string representation and escaping

    • — comment syntax and Unicode support

    Booleans

    Boolean values — true and false, in compact and verbose forms.

    A boolean is a logical value, either true or false. Booleans are scalar values used for flags, binary states, and conditions.

    Each value has a compact and a verbose form, letting you trade brevity for explicitness.

    Syntax

    boolean        = compactBoolean | verboseBoolean
    compactBoolean = "T" | "F"
    verboseBoolean = "true" | "false"

    Structural elements

    Token
    Name
    Description

    T

    The compact and verbose forms are equivalent; the compact form is recommended for terse data.

    Boolean keywords are case-sensitive and spelled exactly. Any other token is not an error — it is parsed as a different value type, so it is not a boolean:

    Under a bool schema, a non-boolean value fails validation with expected-boolean. Without a schema, the values above are simply kept as their parsed type (string or number).

    • — all value types

    • — the absence of a value

    • — the boolean schema type

    Variables

    Value variables — reusable values referenced with @.

    A value variable is a reusable value defined in the header with an @-prefixed key and used anywhere a value is expected by writing @name. Variables reduce repetition, shrink payloads, and let you keep sensitive values in one place.

    @ is for values. For reusable schemas and types, use $ references — see Schema References.

    Defining and using

    ~ @active: T
    ~ $schema: { name: string, isActive: bool }
    ---
    ~ John, @active
    ~ Jane, @active

    In schema constraints

    A variable can supply a constraint value, such as a choices list:

    ~ @r: red
    ~ @g: green
    ~ @b: blue
    ~ $schema: { name: string, color: { string, choices: [@r, @g, @b] } }
    ---
    ~ John, red

    Use cases

    Reduce size and repetition

    Define a value once and reference it many times:

    Variables let you isolate secrets/keys in the header instead of scattering them through the data:

    • ·

    Version History

    How specification changes are recorded; the change log begins at the first 1.0 release.

    This page records notable changes to the specification, newest first, once versioned releases begin. Implementations (such as io-js2) version independently under their own SemVer and declare which specification version they conform to.

    1.0 Draft — in progress

    The specification is in its 1.0 Draft: it is still being authored, most features are at Candidate maturity, and it may change without a version bump until 1.0 is finalized.

    There are no released versions yet — the first entry here will be 1.0, recorded when the draft is finalized. Until then:

    • for the current maturity of each feature, see Feature Status;

    • for how versions and stability work, see the ;

    • for planned direction, see the .

    • · ·

    Internet Object Document

    The two-part structure of an Internet Object document — header and data.

    Internet Object is a document-oriented format built on a clear separation between a header and data. This mirrors how HTTP and MIME keep headers apart from the message body: the header describes the payload, and the data carries it.

    The header is optional and, when present, holds schemas and definitions. The data section begins with the --- separator. That separator is the boundary between the two parts: when a header is present, --- is required to mark where it ends and the data begins.

    Document shapes

    A document can take one of a few shapes depending on whether it carries a header, data, or both.

    Full document

    A document with both a header and a data section is a full document. The header declares the schema; the data conforms to it.

    name, age: int, address: {street, city, state}, active
    ---
    John Doe, 25, {Bond Street, New York, NY}, T

    Data-only document

    When the schema is not needed — or is already known to the recipient — a document can carry data alone. With a single object, the leading --- is optional:

    A collection of records can be written without a separator at all; each record begins with ~:

    Sometimes a request yields no rows — for example, a query that returns result metadata but an empty result set. The header carries the metadata, and the --- separator marks an empty data section:

    A single document can hold multiple data sections, letting related datasets travel together. Each section starts with its own --- separator and names the schema it uses:

    • — schemas, definitions, and metadata

    • — separators, objects, and collections

    • — why the split exists

    Wire Format & Framing

    The on-the-wire framing of a streamed document — the mandatory terminator, control frames, and encoding.

    The streaming wire format is the existing Internet Object document grammar, consumed incrementally. The markers ~, ---, and the header-definition grammar are core constructs; streaming reuses them as its framing layer and does not define them. This page specifies the framing obligations of a streamed document — what a producer must put on the wire and what a consumer may rely on — not the grammar itself. For what the markers mean, see and .

    Chunk boundaries are not semantic. Transport may split or coalesce the byte stream however it likes. Splitting or coalescing chunks MUST NOT change the records a reader emits. Framing is determined by the markers below, never by where a packet happens to end.

    ~ is the only normative data-record marker. A writer MUST frame every logical data record with ~

    Arrays

    Array value syntax — ordered, comma-separated collections of values.

    An array is an ordered collection of values enclosed in square brackets. Arrays are containers used to express lists, sequences, and multi-dimensional tabular structures.

    Each value in an array may be:

    • A scalar (string, number, boolean, null), or

    • A structured value (an object or another array).

    Arrays are syntactically compact, support nesting, and avoid ambiguity by requiring every value to be present — trailing and elided elements are not allowed.

    An array begins with

    Introducing Internet Object

    A guided walkthrough of Internet Object and how it compares to JSON.

    Internet Object (IO) is a document-oriented data serialization format designed to optimize data transmission over networks. This specification introduces IO as an alternative to existing formats such as JSON, offering a structured approach to data representation and exchange.

    The fundamental structure of IO is an ordered collection of values, analogous to CSV (Comma-Separated Values) but with extended capabilities. These capabilities include support for nested objects, arrays, and inline keys, providing enhanced expressiveness and flexibility.

    • Document-oriented design: In contrast to value-oriented formats, IO adopts a document-centric approach, facilitating the separation of data from definitions to enhance clarity and maintainability.

    • Ordered collection with extended functionality: IO's core structure maintains an ordered collection of values while supporting complex data structures such as nested objects and arrays.

    Comments

    Single-line comments for annotating documents.

    Internet Object supports single-line comments for documenting and annotating data. A comment starts with a hash sign (#) and runs to the end of the line.

    • Start character — the hash sign (#, U+0023).

    • Scope — a single line only.

    See Also

    Versioning Policy
    Roadmap
    Versioning Policy
    Feature Status
    Roadmap
    https://internetobject.org
    https://github.com/maniartech/InternetObject-specs
    CC BY-ND 4.0

    UTF-32

    Optional

    Fixed width; larger files

    ASCII

    Optional

    A compatible subset (basic characters only)

    ISO-8859-1

    Optional

    Legacy Latin-1 support

    UTF-8

    Mandatory

    Default and required everywhere

    UTF-16

    Optional

    Alternative encodings

    Unicode support

    Byte order mark (BOM)

    Line endings

    Implementation guidance

    See Also

    Whitespace & Indentation
    Strings
    Comments

    Useful where the platform is natively UTF-16

    Single quote

    U+0027

    Encloses the string; doubled inside to represent itself

    "

    Double quote

    U+0022

    Encloses the string; doubled inside to represent itself

    (space, tab, etc.)

    Whitespace

    Multiple

    Preserved as written

    Any

    Any Unicode code point

    Multiple

    Allowed, except an unescaped enclosing quote

    r

    Raw prefix

    U+0072

    Marks the string as raw

    Structural characters

    Valid forms

    Optional behaviors

    Comments

    Invalid forms

    Preservation of structure

    See Also

    Strings
    Open Strings
    Regular Strings

    '

    Single quote

    U+0027

    Encloses the Base64 content

    "

    Double quote

    U+0022

    Encloses the Base64 content

    A–Z, a–z, 0–9

    Base64 alphabet

    —

    Base64 data characters

    +, /

    Base64 alphabet

    U+002B, U+002F

    Base64 data characters

    =

    Padding

    U+003D

    Base64 padding

    Decoding — a parser decodes the content into a byte sequence (commonly a byte array or buffer) and preserves the exact bytes; invalid Base64 is a parse error.

    b

    Byte prefix

    U+0062

    Marks the value as a Base64 byte string

    Syntax

    Structural characters

    Valid forms

    Invalid forms

    Behavior

    See Also

    Roadmap
    open string
    Value Representations
    Binary
    Strings

    '

    Compact true

    true in compact form

    F

    Compact false

    false in compact form

    true

    Verbose true

    the verbose true keyword

    false

    Verbose false

    the verbose false keyword

    Valid forms

    Not a boolean

    See Also

    Value Representations
    Nulls
    Bool

    Keep sensitive values together

    See Also

    Definitions
    Schema References

    Header-only document

    Document with multiple sections

    See Also

    Header
    Data Sections
    Document-Oriented Nature
  • Schema-first approach: IO emphasizes schema-first design to ensure data consistency and predictability. While schemas are optional, their inclusion significantly enhances data integrity and validation.

  • Concise syntax: The syntax of IO is optimized for readability and efficiency, minimizing data size without compromising clarity.

  • Metadata integration: IO documents can incorporate metadata, variables, and multiple schemas within the header section, providing comprehensive context for the data.

  • The following example demonstrates a basic IO document structure:

    This structure illustrates IO's concise syntax and inherent schema support. For comparison, an equivalent JSON representation would be:

    IO supports collections and various data types, as demonstrated in the following example:

    This example illustrates several key features:

    • Explicit data type definitions in the schema (string, int, bool)

    • Nested object structures (address)

    • Collection of objects denoted by the tilde (~) prefix

    • Correspondence between the order of values and the schema definition

    The equivalent JSON representation would be:

    This comparison demonstrates IO's capacity to represent structured data collections efficiently, offering a compact and readable format while maintaining an ordered structure.

    In many scenarios, it is beneficial to define schemas separately from the data. This approach allows for schema reuse, versioning, and easier maintenance. Here is an example of a separate schema followed by a document using that schema:

    First, the schema, defined on its own (for example, in a file named person.io):

    Then, a document that carries metadata in its header and a collection of records below:

    In this example:

    • The schema is defined separately, potentially in a file named person.io.

    • The document references the schema URL in its metadata.

    • The document includes additional metadata such as record count and pagination information.

    • The collection contains multiple records, each prefixed with ~.

    • Each record follows the structure defined in the schema, including an array of skills.

    This structure allows for efficient data transmission, as the schema only needs to be sent once and can be cached by the receiving system. It also facilitates updates to the schema without necessarily changing the data format.

    Internet Object represents a significant advancement in data serialization technology. By combining the simplicity of ordered collections with the robustness of schema-based validation, Internet Object offers a powerful yet accessible solution for modern data exchange needs. Its key strengths include:

    1. Efficiency in data transmission and storage

    2. Clarity through its schema-first approach and document-oriented design

    3. Flexibility in handling various data structures and types

    4. Compatibility with existing JSON-based systems

    These attributes make Internet Object suitable for a wide range of applications, from web-based and networked environments to data storage and interchange in diverse domains such as IoT, cloud computing, and enterprise systems.

    The subsequent sections of this specification provide comprehensive details on Internet Object's syntax, schema definition language, supported data types, and advanced features. This information enables developers, system architects, and data engineers to fully leverage the capabilities of Internet Object in their projects and applications.

    • Getting Started — a five-minute tour in pure Internet Object

    • Why Internet Object? — how it compares to JSON, CSV, and YAML

    • Internet Object Document — the structure in depth

    Core structure

    Key features

    Illustrative examples

    Basic IO document structure

    IO document with collections and data types

    Advanced examples

    Separate schema and document with collection

    Conclusion

    See Also

    Placement — anywhere in the document.

  • Content — everything after # on the same line is ignored by the parser.

  • A comment can stand on its own line or trail a value:

    • A comment can appear on any line, standalone or trailing a value.

    • A comment supports full Unicode text.

    • A comment cannot span multiple lines.

    • No escaping is needed inside a comment.

    • A # inside a quoted string is literal text, not a comment.

    • Be clear and concise — use direct language.

    • Explain why, not what — focus on reasoning, not the obvious.

    • Keep comments current — update them when the data or schema changes.

    • Stay consistent — keep a uniform commenting style across documents.

    • Internet Object Document — overall document structure

    • Encoding — Unicode support in text content

    Syntax

    Examples

    Comment placement

    Rules

    Best practices

    See Also

    rawString = "r" (singleQuotedRaw | doubleQuotedRaw)
    singleQuotedRaw = "'" { character | doubleSingleQuote } "'"
    doubleQuotedRaw = '"' { character | doubleDoubleQuote } '"'
    character = any Unicode code point except the enclosing quote
    doubleSingleQuote = "''" (a single quote inside a single-quoted raw string)
    doubleDoubleQuote = '""' (a double quote inside a double-quoted raw string)
    r'C:\program files\example\app.exe'
    r"C:\program files\example\app.exe"
    r'^(19|20)\d\d([- /.])(0[1-9]|1[012])\2(0[1-9]|[12][0-9]|3[01])$'
    r"^(19|20)\d\d([- /.])(0[1-9]|1[012])\2(0[1-9]|[12][0-9]|3[01])$"
    r'जॉन डो'
    r"Can contain unicode characters 😃"
    r'Jonas D''costa'        # a single quote inside, written as two single quotes
    r"He said, ""Hello!"""   # a double quote inside, written as two double quotes
    r'Jonas D'costa'             # ✗ unexpected-token — unescaped quote
    r"He said, "Hello!""         # ✗ unexpected-token — unescaped quote
    r'Unclosed string            # ✗ unterminated-string — no closing quote
    ---
    rC:\program files\app.exe    # → a member named `rC`, not a raw string
    binaryValue = "b" (singleQuotedBase64 | doubleQuotedBase64)
    singleQuotedBase64 = "'" base64Content "'"
    doubleQuotedBase64 = '"' base64Content '"'
    base64Content   = { base64Character }
    base64Character = "A"…"Z" | "a"…"z" | "0"…"9" | "+" | "/" | "="
    b'SGVsbG8gV29ybGQ='        # "Hello World"
    b"SGVsbG8gV29ybGQ="        # same, with double quotes
    b'QWxhZGRpbjpvcGVuIHNlc2FtZQ=='   # "Aladdin:open sesame"
    b'TWFu'                    # "Man"  (no padding needed)
    b'TWE='                    # "Ma"   (one pad)
    b'TQ=='                    # "M"    (two pads)
    b''                        # empty byte string
    b'SGVsbG8 gV29ybGQ='       # ✗ invalid-binary — space within the Base64 content
    b'SGVsbG8@V29ybGQ='        # ✗ invalid-binary — invalid character '@'
    B'SGVsbG8gV29ybGQ='        # ✗ unknown-annotation — the prefix must be lower-case `b`
    b'SGVsbG8=                 # ✗ unterminated-string — no closing quote
    ---
    bSGVsbG8=                  # → "bSGVsbG8=" — a string, not binary
    ---
    T, F, true, false
    t                    # open string "t", not true
    TRUE                 # open string "TRUE", not true
    True                 # open string "True", not true
    1                    # the number 1, not true
    0                    # the number 0, not false
    ~ @co: 'ACME Corporation'
    ~ $schema: { name: string, employer: string }
    ---
    ~ John, @co
    ~ Jane, @co
    ~ @key: 'sk_live_8f3a9c2b'
    ~ $schema: { account: string, apiKey: string }
    ---
    ~ acct_001, @key
    ---
    John Doe, 25, {Bond Street, New York, NY}, T
    ~ John Doe, 25, {Bond Street, New York, NY}
    ~ Jane Doe, 48, {Malibu Point 10880, Malibu, CA}
    ~ recordCount: 0
    ~ pageSize: 10
    ~ currentPage: 1
    ~ nextPage: N
    ~ prevPage: N
    ---
    ~ $address: {street, city, state, zip}
    ~ $person: {firstName, lastName, age, gender}
    --- $person
    ~ John, Doe, 25, M
    ~ Jane, Doe, 22, F
    --- $address
    ~ Bond Street, New York, NY, 500001
    ~ George Street, New York, NY, 500002
    name, age, active, address: {street, city}
    ---
    John Doe, 25, T, {Bond Street, New York}
    {
      "name": "John Doe",
      "age": 25,
      "active": true,
      "address": {
        "street": "Bond Street",
        "city": "New York"
      }
    }
    name:string, age:int, active:bool, address: {street:string, city:string}
    ---
    ~ John Doe, 25, T, {Bond Street, New York}
    ~ Jane Doe, 20, T, {Main Street, San Francisco}
    [
      {
        "name": "John Doe",
        "age": 25,
        "active": true,
        "address": {
          "street": "Bond Street",
          "city": "New York"
        }
      },
      {
        "name": "Jane Doe",
        "age": 20,
        "active": true,
        "address": {
          "street": "Main Street",
          "city": "San Francisco"
        }
      }
    ]
    # Person schema
    name:string, age:int, active:bool, address: {street:string, city:string}, skills:[string]
    ~ schemaUrl: "https://example.com/schemas/person.io"
    ~ recordCount: 3
    ~ page: 1
    ~ totalPages: 1
    ---
    ~ John Doe, 25, T, {Bond Street, New York}, [JavaScript, Python]
    ~ Jane Doe, 30, F, {Main Street, San Francisco}, [Java, C++, Rust]
    ~ Bob Smith, 28, T, {Park Avenue, Chicago}, [Ruby, Go]
    # Internet Object document: personnel records
    
    # Address schema definition
    ~ $address: {street:string, zip:{string, maxLen:5}, city:string}
    
    # Person schema definition
    ~ $schema: {
        name:string,               # individual's full name
        age:int,                   # age in years
        homeAddress?: $address,    # optional home address
        officeAddress?: $address   # optional office address
    }
    ---
    # Personnel records
    ~ John Doe, 25, {Queens, "50010", NewYork}, {Bond Street, "50001", NewYork}
    ~ Jane Doe, 20, {Queens, "50010", NewYork}, {Bond Street, "50001", NewYork}
    {
        # person details
        name: John Doe, # inline comment
        age: 30,        # another inline comment
    
        # contact information
        contact: {
            email: '[email protected]',
            phone: '+1-555-0123'
        }
    }
    , and the reader emits exactly one item per
    ~
    -introduced record. Quoted multiline values remain part of the same logical record — a newline inside a quoted string does not start a new record.

    This stream carries two logical records. Each ~ line is one record; the reader emits one item for each.

    A stream MAY begin with header definitions, written in the core header-definition grammar. To make the boundary between header and data unambiguous as bytes arrive, the terminator is mandatory:

    • A conforming writer MUST emit an explicit --- (or --- $Schema) before the first data record, even when the header is empty. An empty header serializes to exactly ---.

    • This terminator is what opens the data section. It makes the first token unambiguous to the reader:

      • a stream beginning with --- has no header (or an empty one), and its data MAY stream immediately;

      • a stream beginning with ~ is a header (definitions block) that the reader MUST buffer until the terminating ---.

    Only the first --- is load-bearing for header-versus-data separation. Within the data section, records use ~ alone; any later --- is an ordinary schema switch, not a second header boundary.

    The smallest possible conforming stream with data is therefore a bare terminator followed by records:

    A control frame is structural input that is not a data record: a header-definition block, or a section marker. Control frames are never emitted as data items.

    • --- $Schema selects the schema context for the records that follow.

    • A bare --- resets the active section to the default schema context.

    • A header-definition block (everything before the first ---) establishes shared stream state.

    Because control frames carry state rather than data, the reader applies their effect but emits nothing for them. The detailed rules for schema selection and definition state are in Schema & State.

    A document that begins directly with ~ data and contains no --- is the ordinary non-streaming collection form. A reader MAY accept it, so that a non-streaming document stays equivalent under the reader. However:

    • This form cannot be emitted incrementally. The reader must buffer it to end of stream to determine that it was data and not an unterminated header.

    • A writer MUST NOT emit this form. A writer always emits the --- terminator (the section above), which removes the ambiguity.

    In short: readers tolerate the legacy form for compatibility; writers never produce it.

    Streaming decodes the wire the same way the core format does, with the additional obligation that decoding state survives chunk boundaries.

    • Byte sources MUST be decoded as UTF-8, preserving multibyte decoder state across chunk boundaries so that a code point split across two chunks decodes correctly.

    • A leading UTF-8 byte-order mark (EF BB BF) at the very start of the stream MUST be stripped. A BOM-like sequence anywhere else is ordinary content and MUST NOT be stripped.

    • Newlines MUST be normalized for framing: \r\n and a lone \r are treated as \n. Record framing MUST NOT depend on the producer's newline convention.

    • Text sources are already-decoded text and MUST NOT be re-decoded as bytes.

    • Newline normalization is a framing concern only. It MUST NOT alter how core interprets bytes inside a quoted value.

    For the full character-encoding rules of the format, see Encoding.

    • Stream Items — what the reader emits for each ~ record

    • Schema & State — how --- and --- $Schema select schema context

    • Internet Object Document — the header and data structure streaming reuses

    • — the format's character-encoding rules

    Records and the data marker

    Internet Object Document
    Collection
    ~ $schema: { name: string, role: string }
    ---
    ~ Alice, admin
    ~ Bob, guest
    ---
    ~ Alice
    ~ Bob

    The mandatory header terminator

    Control frames

    The legacy headerless form

    Encoding

    See Also

    [
    and ends with
    ]
    , containing zero or more comma-separated values.

    Here, value is any valid Internet Object value, as defined in Value Representations. The syntax and behavior of each value type (strings, numbers, booleans, objects, arrays, null) is defined in its own page.

    Symbol
    Name
    Unicode
    Description

    [

    Open square bracket

    U+005B

    Starts an array

    Whitespace is permitted around elements and structural characters for readability.

    All forms with an equivalent value structure are interpreted identically.

    An empty array is written as a pair of brackets with nothing between them:

    This is a valid array with no elements.

    Arrays may contain other arrays, allowing arbitrarily deep structures.

    Comments are allowed around and within arrays, following the format's general comment syntax.

    Comments must not break value boundaries, and must not appear inside strings or object keys.

    A comma separates values, so every comma needs a value on each side of it:

    Missing separators are a different matter, and not an error:

    An open string may contain spaces, so a b c is a single well-formed value and there is nothing for the parser to object to. See Syntax Errors for why this is the one place the format cannot help you.

    Internet Object preserves:

    • Value order

    • Whitespace (non-significant in interpretation)

    • Syntactic fidelity (as written)

    It does not interpret:

    • The meaning of order

    • Whether values must be unique

    Those semantics belong to the schema, the validator, or the application.

    • Value Representations — all value types

    • Objects — the other structured value

    • Array — schemas for arrays

    Syntax

    array = "[" [ value *("," value) ] "]"
    []                       # Empty array
    [apple, banana, cherry]  # String values
    [1, 2, 3]                # Number values
    [T, F, N]                # Boolean and null values
    [{x:1}, {y:2}]           # Array of objects
    [1, [2, 3], [4, [5, 6]]] # Nested arrays
    [[1,2],[3,4]]            # 2D array
    [ a , b , c ]   # Valid
    []
    [1, [2, 3], [[4]]]
    [
      1, 2,  # inline comment
      3
    ]
    [a, b, ]     # ✗ unexpected-token — trailing comma
    [a,,c]       # ✗ unexpected-token — elided value
    [ , ]        # ✗ unexpected-token — no value at all
    [,a]         # ✗ unexpected-token — nothing before the comma
    ---
    [a b c]      # → ["a b c"] — one open string, not three values
    [a, b]       # ✓ valid
    [a, null, c] # ✓ use null for a missing value

    Structural characters

    Valid forms

    Optional behaviors

    Whitespace and formatting

    Empty array

    Nesting

    Comments

    Invalid forms

    Corrected versions

    Preservation of structure

    See Also

    Error Codes

    How every Internet Object error code is named, and the closed vocabulary it draws from.

    Every error an Internet Object processor reports carries a stable error code: a lowercase, hyphenated identifier such as expected-string or mismatched-max. Codes are part of the specification. Messages are not — they may be reworded, translated, or made more helpful at any time, so tooling MUST branch on the code and MUST NOT parse the message.

    This page defines how codes are named. The codes themselves are catalogued in Error Model.

    The rule

    <predicate>-<subject> — the predicate first, drawn from the closed vocabulary below; then the subject it applies to.

    expected-string          the predicate is `expected`, the subject is `string`
    mismatched-max           the predicate is `mismatched`, the subject is `max`
    unterminated-string      the predicate is `unterminated`, the subject is `string`
    duplicate-section-name   subjects may be multi-word

    Putting the predicate first is what makes a missing code visible. The predicate vocabulary is small and fixed; the set of subjects is large. Reading a predicate's codes as one list turns a gap into something you can see:

    Read as one list, a type with no code is a hole you can see. That is not a hypothetical: reading this list is how expected-date and expected-time were found missing. One code, expected-datetime, had been serving all three temporal types, so a date member given a string reported a type the schema never mentioned.

    One hole remains, deliberately. binary is a base type with no expected-binary, because no implementation registers binary as a schema type yet — a code that nothing can emit is a promise the registry cannot keep, so it lands with the type (see ).

    A code describing a DOCUMENT may use only these thirteen predicates. The set is closed: adding one is a change to this specification, not a decision an implementation may take on its own. That is what stops synonyms creeping back in — before this rule existed, "the value is the wrong type" was spelled four different ways.

    The rule governs codes about the document — its text, its values, its schema. Conditions of the transport are a separate namespace, stream-, defined in : a buffer limit or an aborted connection is a fact about the delivery, not about the data, and forcing it into a predicate (exceeded-stream-buffer) obeys the letter of the rule while describing the wrong thing.

    Predicate
    Means
    Example

    The subject is not always the same kind of thing, and which one it is follows a single rule:

    A type problem names the type. A constraint problem names the constraint.

    Both values are rejected by the same member, but for different reasons and with different fixes. The first is not an integer at all. The second is a perfectly good integer that broke a rule the schema author wrote — and the code names max, so the reader knows which line of the schema to look at.

    This is why constraint codes name the keyword rather than the type: the type is always recoverable from the error's position and the schema, but the failed constraint is not recoverable from the value at all.

    Every value constraint keyword has exactly one code:

    Keyword
    Code

    Two MemberDef keywords are deliberately absent, because they do not constrain a value's content: optional: false is a presence rule and reports missing-value, and null: false is a nullability rule and reports forbidden-null.

    Each of these separates two conditions that look alike and need different fixes.

    expected- is a type problem; missing- is a presence problem.

    expected- means this is not that type at all. invalid- means it is that type, and it is malformed.

    The same split applies to every type whose literal carries a marker, and the marker is what makes the literal recognizable in the first place:

    Marker
    Claims
    Not that type at all
    That type, malformed

    Read as a grid, a missing cell is visible at a glance — which is how expected-date, expected-time, invalid-date and invalid-time were each found absent, with one code doing the work of three. The single blank is expected-binary, and it is : no implementation registers binary as a schema type yet, so nothing could emit it.

    Note that the subject is always the type, never the encoding or the notation: invalid-binary, not invalid-base64; invalid-number, not invalid-hex.

    Both mean "the number is too big", and they need opposite fixes.

    out-of-range- appears only where the limit is intrinsic to the type and the author declared nothing. Anything the author wrote is mismatched-.

    The subject here is the type family, not the exact type: the value above overflows int8, and the code says integer. That is the one place the subject is deliberately broader than the fault, and it is a trade the family earns — a code per sized type (out-of-range-int8, -int16, -uint32, …) would multiply the vocabulary to describe a difference the error's position already carries.

    A typo and a reserved name must not report the same code: one is fixed by correcting the spelling, the other by choosing a different type until a later version of the specification supports it.

    • Never assembled at runtime. A code built from a value — say, from the declared type name — produces identifiers that appear in no registry, that no implementation can be checked against, and that a consumer cannot know to expect.

    • Never an implementation limit. A code meaning "this library has not built that yet" cannot be part of a specification: another implementation may support the feature and would then be wrong for not reporting an error. Genuine implementation limits are reported outside this vocabulary, and a conformance suite skips such cases rather than expecting a code.

    • Never absent. Every reported error carries a code. An error that reaches a caller without one cannot be branched on and renders as a blank in tooling, which is no better than reporting nothing at all.

    • — the catalogue of codes, by class

    • — reporting many errors in one pass

    Open Strings

    Open strings — the unquoted string form.

    An open string is the simplest string form: a sequence of Unicode code points with no enclosing quotes. Open strings suit simple, unstructured text that does not begin or end with whitespace and does not require escaping of special or structural characters.

    Open strings are scalar values. They preserve all internal whitespace and Unicode content, but cannot start or end with whitespace.

    Syntax

    An open string begins with any non-whitespace code point and ends at the first whitespace or structural character, or at the end of the document.

    openString    = nonWhitespace (codepoint)*
    nonWhitespace = any Unicode code point except whitespace
    codepoint     = any Unicode code point except a structural character or document end

    Structural characters

    Symbol
    Name
    Unicode
    Description

    A quote character may not appear in an open string — not at the start, not in the middle, not at the end. The run before a quote is read as an annotation name, because that is exactly the shape of an : r'…', b"…", dt'…', d'…', t'…'. The two cannot both be true, and the annotation wins.

    So an apostrophe in ordinary text is not writable open, however natural it looks:

    Written
    Read as

    Quoting is the escape, and the only one: "don't stop" is a regular string and carries the apostrophe with no escaping, because a single quote needs none inside double quotes. A writer quotes such a value automatically; see .

    Examples of valid open strings:

    Multiple open strings in an object:

    A multiline open string (no escaping required):

    An open string may begin with a digit, and many everyday values do — measurements like 12mm, times like 3pm, and part codes and identifiers like 013ABSD. Digits followed by letters is ordinary text, not a malformed number.

    The one exception is a base prefix. 0x, 0o and 0b announce hexadecimal, octal and binary, so a run that begins with one and does not decode is a failed number rather than a string:

    Quoting settles it either way: "0x123FG" is a string, unambiguously. See .

    • Whitespace — an open string cannot start or end with whitespace, but preserves all internal whitespace.

    • No escaping — character escaping is not processed; quotes and other characters appear as-is.

    • Multiline — an open string can span multiple lines as long as no structural character interrupts it.

    Comments are not allowed inside open strings, but may appear outside or between values, per the format's comment rules.

    Neither of these is an error. Each is a perfectly good value — just not an open string, which is what makes them worth showing: an open string is defined by what it may not begin with.

    Internet Object preserves:

    • All Unicode code points and internal whitespace as written

    • The unquoted, open form of the string

    It does not interpret or enforce:

    • Escaping or encoding

    • Leading or trailing whitespace (which is disallowed)

    • Application-specific constraints

    • — the three string forms

    • — quoted strings with escaping

    • — literal strings without escape processing

    MemberDef

    MemberDef — defining one member's type, constraints, and optional/nullable/default behavior.

    A MemberDef (member definition) defines a single member of a schema: its type and the constraints on its value. It is an Internet Object object whose first value is the type, optionally followed by a default and choices, then keyed options. Every MemberDef is validated against its type's TypeDef, so only the options that type defines are allowed.

    Writing a MemberDef

    The first value is the type; a default and choices may follow positionally; all other options are keyed:

    ~ $schema: {
        age:    { number, min: 0, max: 120 },        # type + keyed options
        level:  { int16, 1, [1, 2, 3] },             # type + default + choices
        name:   { string, pattern: "^[A-Za-z]+$" },  # type + keyed option
        tags:   { array, of: string, minLen: 1 }     # container type + options
    }
    ---
    ~ 30, 2, John, [a, b]

    A member can also be just a type (age: int) or just a name (age, which defaults to any). The positions and options each type accepts are listed in its Schema Data Types page.

    Validation against the TypeDef

    A MemberDef may use only the options its type defines. An unknown option is rejected with unknown-member:

    MemberDef options come in two kinds, and the distinction is normative — an implementation that confuses them will accept or reject documents that others do not.

    Kind
    Answers
    Examples
    Rejects a value?
    Changes how it is written?

    A presentation option is write-only. It tells a writer which spelling to emit; it MUST NOT restrict what a reader accepts:

    Both records hold 255, and a writer emits 0xff for both. This follows from the data model: 0xff and 255 parse to the same value, so the notation is not part of the value and cannot be validated. It is also not recoverable — a value built in code carries no notation at all, so a rule enforcing one would apply when a document is parsed and silently not apply when the same value is loaded from a host object.

    To require a particular written form, constrain the text itself. Use a string with a pattern ({ string, pattern: "^0x[0-9a-fA-F]{6}$" }) — then the form is the value. If the real requirement is a range, say so with the type (uint8); if it is a fixed set, use choices.

    The type name carries a third thing: what the value is. email and url are strings with extra validation; date and time are temporal values narrower than datetime. Unlike format, a type name does constrain — and for date / time it also selects the written form, because the value kind and its literal are the same decision.

    • Optional (? on the key) — the member may be omitted: age?: { number, min: 0 }.

    • Nullable (* on the key) — the value may be null: age*: number.

    The suffixes are shorthand: optional and null are ordinary MemberDef options, so score?*: number and score: { number, optional: T, "null": T } declare the same member.

    The null option key is written quoted. A bare null: is the null keyword, so { number, null: T } is rejected with invalid-key; write "null": T (or r'null': T). optional: needs no quoting.

    The suffixes are part of the name token, not separate syntax, so they can follow only a bare name. A name that has to be quoted — because it holds a comma, a colon, a space, or begins with a digit — cannot carry them, and writes the options instead:

    In a schema that mixes both, only the quoted name is affected:

    Because the suffix belongs to the name token, a quoted name is literal: "a?" is a member named a?, not an optional member named a. Quoting says this name, exactly.

    Both use { … }, which makes them easy to confuse, but they serve different purposes:

    • A MemberDef validates one value — it has a type and constraints.

    • A SchemaDef describes an object's shape — a map of field names to types or MemberDefs.

    The parser tells them apart by the first entry:

    1. Is the first value a known type? → it is a MemberDef.

    2. Otherwise (the first entry is a key: … pair, or a bare field name) → it is a SchemaDef.

    Here address declares fields (street, city), while age declares a type with a range.

    For a nested object, write the shape inline ({ … }). The explicit { object, schema: { … } } form is equivalent and documented on :

    • — the allowed options for each type

    • — how schemas are built

    • — declaring object shapes

    • — reusing shapes with $

    Whitespace & Indentation

    Recognized whitespace characters and how the parser treats them.

    In the Internet Object format, whitespace is any character with a Unicode code point less than or equal to U+0020 (the range U+0000 to U+0020). This range covers both non-printable control characters and common whitespace such as the horizontal tab (U+0009), newline (U+000A), vertical tab (U+000B), form feed (U+000C), carriage return (U+000D), and space (U+0020).

    Because the format is not whitespace-sensitive, indentation carries no meaning: it is ordinary whitespace between tokens. You may indent objects, arrays, and definitions freely for readability without changing how a document parses.

    Beyond the U+0000–U+0020 range, the format also treats characters in the Unicode whitespace category as whitespace, such as the non-breaking space (U+00A0), em space (U+2003), and en space (U+2002). Recognizing these makes the format easier to work with in languages that use non-Latin scripts, such as Arabic, Chinese, or Japanese.

    The format also recognizes the zero-width non-breaking space (U+FEFF) as whitespace. This character is often used as a byte order mark (BOM) in Unicode-encoded documents.

    The following table lists the valid whitespace characters:

    Code points
    Description
    Notes
    • Whitespace insensitive — the parser ignores whitespace surrounding values and structural elements.

    • Preserved inside strings — whitespace within a value or string is kept exactly as written.

    • Recognized by code point — whitespace is identified by Unicode code point, per the table above.

    • Reserved — these whitespace characters

    • Aid readability — use spaces and tabs to format a document so it is easy to read.

    • Avoid clutter — excessive whitespace adds no meaning and reduces readability.

    • Stay consistent — apply whitespace uniformly across a document for easier maintenance.

    • Watch for invisible characters — zero-width and other invisible spaces can slip into keys or values unnoticed; avoid pasting them in.

    • — Unicode handling and document encoding

    • — whitespace handling inside strings

    BigInt

    Arbitrary-precision integer values for very large whole numbers.

    A BigInt is an arbitrary-precision integer — a whole number with no upper or lower size limit. It suits values that exceed the safe range of a standard Number, such as cryptographic quantities, large identifiers, and high-volume counters.

    A standard Number is exact only for integers within roughly ±2^53−1 (about ±9 quadrillion). A BigInt stays exact at any magnitude.

    Syntax

    A BigInt is written as an integer with the n suffix:

    bigint = ["-" | "+"] (decimalBigInt | binaryBigInt | octalBigInt | hexBigInt)
    
    decimalBigInt = digit+ "n"
    binaryBigInt  = "0b" binaryDigit+ "n"
    octalBigInt   = "0o" octalDigit+ "n"
    hexBigInt     = "0x" hexDigit+ "n"
    
    digit       = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9"
    binaryDigit = "0" | "1"
    octalDigit  = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7"
    hexDigit    = digit | "A" | "B" | "C" | "D" | "E" | "F" | "a" | "b" | "c" | "d" | "e" | "f"

    Structural characters

    Symbol
    Name
    Unicode
    Description

    A BigInt can also be written in binary, octal, or hexadecimal — each still ending in n. The following are all equal to 42n:

    These are genuine syntax errors:

    A BigInt holds whole numbers only, so a fractional BigInt is an error:

    Lenient fallbacks. A malformed suffix does not raise an error — it falls back to an open string. 123nn parses as the text "123nn", and n123 as the text "n123". A conformant parser SHOULD instead reject these; they are tracked as implementation issues.

    Internet Object preserves:

    • The chosen notation (decimal, binary, octal, hex)

    • Exact integer precision at any magnitude

    • Syntactic fidelity as written, except that an explicit + sign is not preserved

    It does not interpret:

    • Mathematical relationships between values

    • Domain-specific constraints on large integers

    Those semantics belong to the schema, the validator, or the application.

    • — all numeric forms

    • — standard floating-point numbers

    • — fixed-precision decimal arithmetic

    Open & Dynamic Schemas

    Strict and extensible schemas — accepting extra fields with the * marker.

    A schema is strict by default: a record MUST contain only the declared fields. Adding a * marker makes it extensible, and an extensible schema accepts fields beyond those declared.

    A note on the words. This specification says strict and extensible for schemas, and keeps open and closed for the object syntax — an open object is written without braces, a closed one with them. The two axes are unrelated, and one document can hold all four combinations, so a single pair of words for both would make sentences that cannot be read. Readers arriving from JSON Schema should map strict to additionalProperties: false.

    Strict by default

    Extra values in a record are rejected unless the schema opts in with *:

    Place * after the declared fields to accept extras. Positional extras are keyed by index; keyed extras keep their names:

    The * above is bare, and that is what makes it the wildcard. Quoted, it is an ordinary member name — the same rule that governs the ? and * : quoting says this name, exactly.

    Data is free to use * as a key; JSON-sourced configuration routinely does:

    A member named * does not make its schema extensible, and a wildcard is not a member: it never appears among the schema's member names, and a schema may carry both at once.

    *: <type> constrains every extra field; *: { <type>, …constraints } adds constraints:

    A wildcard may reference a SchemaDef, which makes {*: $ref} the natural schema for dictionary/map-shaped data — data whose keys are values in their own right (IDs, codes, locale tags) rather than field names. The keys stay in the data section; the wildcard types every value:

    The wildcard forms, in full:

    Form
    Meaning

    Declared members and the wildcard compose — { id: string, *: number } requires id and lets any other member be a number.

    When a single field must accept more than one type, use anyOf (see ):

    Only a root object may be written open. The header of a schema, and a record in the data, may both omit their braces — that is the ordinary form. Every child object MUST be braced.

    There is no braceless nested object to get wrong, because dropping the braces does not produce a partial object — it produces a different declaration entirely:

    The same holds in the data. A child object without braces is not a short object, it is extra members in the parent, and a strict schema rejects them as unknown-member.

    • ·

    Date and Time

    The datetime, date, and time types.

    Internet Object has three temporal types, each with its own literal value:

    Type
    Literal
    Captures

    date

    d'2024-03-20'

    calendar date

    For the literal value syntax in detail, see .

    These types share one TypeDef. A MemberDef accepts only the options below.

    Option
    Type
    Description

    There is deliberately no format option. For the numeric and string types a format selects a spelling for one unchanging value, but a date is not a datetime — the three temporal types are genuinely different, so the choice lives in the type name, which both constrains the value and selects its literal. See .

    A writer therefore emits the literal matching the declared type: date → d"…", time → t"…", datetime → dt"…". Where no schema applies, the kind is inferred from the value instead — see .

    Input
    Result
    • ·

    Round-Trip Guarantees

    What serialization guarantees across a parse-write cycle, and what it does not.

    Round-tripping is the property that makes serialization testable: it turns "does this writer behave correctly?" into a question with a mechanical answer.

    The two invariants

    For any document x that parses without error, a conformant writer satisfies both:

    1. Value preservation. Writing a parsed value and parsing the result yields the same value:

    parse(write(parse(x)))  ==  parse(x)

    Equality is over the value model — types, member names, order, and nesting — not over the text.

    2. Output validity. A writer's output always parses, with no errors:

    parse(write(v))  succeeds

    The second does not follow from the first, and it is the one that catches most real defects: output that is nearly right — an unquoted key containing a colon, a string that loses its trailing space, an object missing its enclosure — fails here immediately.

    Idempotence

    Writing is a fixed point after the first pass. For a document that already carries a header:

    The first write normalizes; every later write changes nothing. An implementation whose output keeps changing across cycles has a defect, even if each individual output re-parses.

    • every member — a writer never drops one, keyed or keyless

    • each value's type, including bigint, decimal, datetime, and binary

    • member order, and the positions of absent optional members

    • names that a schema cannot recover

    Round-tripping is defined over the value model, not the source text. These are expected to change:

    Not preserved
    Why

    A consequence worth stating plainly: text equality is not the test. Comparing a writer's output byte-for-byte against its input is expected to fail, and is not a conformance signal. Compare parsed values.

    The invariants above are directly executable, which makes them the backbone of a writer's test suite:

    1. For every document in the conformance corpus, assert invariant 1 and invariant 2.

    2. Assert idempotence on the second write.

    3. Generate documents across the value and schema space and assert the same three properties.

    Generated round-trip testing is strongly recommended. Each of the writer defects listed as known gaps in is caught by invariant 2 alone.

    • ·

    NaN and Infinity

    The special numeric values NaN and Infinity.

    The special numeric values represent "Not a Number" and infinite quantities, following IEEE 754. They model the edge cases of numeric computation — an undefined result, or a value beyond the finite range. Only the Number form supports them; BigInt and Decimal do not.

    Syntax

    The special values are written as fixed keywords:

    specialValue  = nanValue | infinityValue
    nanValue      = "NaN"
    infinityValue = ["+" | "-"] "Inf"

    Structural elements

    Token
    Name
    Description

    NaN

    The special values are case-sensitive and spelled exactly. Any other spelling is not an error — it is simply parsed as an open string (text), so it is not the numeric special value:

    To store one of these as a number, write it exactly: NaN, Inf, -Inf, or +Inf. To store the word as text, quote it: "infinity".

    Internet Object preserves:

    • The exact special value written

    • The sign of negative infinity

    • Syntactic fidelity as written, except that an explicit + sign is not preserved

    It does not interpret:

    • The operations that produced these values

    • Comparison or equality semantics (for example, that NaN equals nothing, including itself)

    Those semantics belong to the application.

    • — standard floating-point numbers

    • — all numeric forms

    Overview

    Overview of serialization — turning values back into Internet Object text.

    Serialization is the reverse of parsing: it turns an in-memory value back into Internet Object text. A component that does this is a writer.

    Where parsing is permissive — it accepts every form the grammar allows — serialization is narrow. Many different texts parse to the same value, but a writer emits exactly one of them. That chosen form is the canonical output, and pinning it down is what lets two independent implementations agree.

    The core principle

    Internet Object exists to move data, not names. A schema is a contract shared by the endpoints; it does not have to travel with every payload. So the normal wire form is positional values — the receiver already knows the schema, and repeating field names on the wire is redundant. Removing that redundancy is the point of the format.

    A name therefore appears in the output only when it cannot be recovered any other way:

    A member is written positionally (bare value) when its name is recoverable from a schema in scope — a document header, a section schema, or a parent MemberDef. A name is written inline (key: value) only when it is not recoverable. A name is never both hoisted into a schema and repeated inline.

    Serialization happens at two levels, and they answer different questions:

    Level
    Produces
    Question it answers

    Only a document has a header. A writer MUST NOT infer a schema during serialization: if a document carries no schema, none is written, and the data is emitted schema-less.

    • MUST produce output that re-parses to an equivalent value — see .

    • MUST preserve each value's type, not merely its printed form — see .

    • MUST NOT drop a member. Every member present in the value appears in the output.

    • MUST NOT

    • — when a member is written bare and when it is written key: value

    • — how each scalar, key, and string is written

    • — records, enclosure, headers, and sections

    • — what is preserved, and what is deliberately not

    • — the input side of the same pipeline

    • — streaming writers delegate to these rules

    Versioning Policy

    How the Internet Object specification is versioned, and the stability tiers that govern each feature.

    This page defines how the Internet Object specification is versioned and how the stability of each feature is governed, so the specification can be published and evolve continuously without being a perpetual draft. It follows the model used by mature standards (CSS module levels, TC39 stages, Kubernetes alpha/beta/GA): a written policy plus a per-page status index, with stability tracked per page rather than as one global label. For where each page currently stands, see .

    • The format specification carries a major version (for example, "Internet Object 1.0"). A backward-incompatible change to the format requires a new major version (2.0).

    • Sub-protocols built on the format — for example — carry their own version (the Streaming Protocol is at v1) and advance on their own clock.

    Schema & State

    How a stream resolves definitions atomically, selects schemas, and applies precedence with preloaded state.

    A stream carries state: header definitions, a default schema context, and the explicit schema in effect for the current section. This page specifies how that state is established, how it is selected per record, and how it composes with definitions supplied before the stream begins. All of the semantics — what a definition means, how a schema resolves, how default, optional, null, and choices behave — belong to core; streaming only governs when state is resolved and which schema applies. See and .

    The header is the definitions block before the first ---. Definition references inside it are position-independent: a definition at any position may reference another regardless of order. Because of this:

    Readers & Writers

    Obligations of the reader and writer roles, plus adapters, transports, backpressure, and conformance.

    The protocol defines two roles — the writer that frames records onto a stream, and the reader that consumes a stream and emits one item per record — together with the obligations of the adapters and transports that connect them. These are abstract roles. A concrete implementation may expose them under any names and any platform idioms, but it MUST honor the duties below.

    A writer is responsible for canonical framing. It MUST:

    • Serialize through core. A writer MUST serialize record values using core's serializer. It MUST NOT introduce stream-only formatting for strings, defaults, nulls, arrays, or objects.

    Numeric Values

    The numeric value forms — Number, BigInt, and Decimal.

    Internet Object provides accurate numeric representation for everything from simple counting to exact financial arithmetic. It supports three numeric forms — Number, BigInt, and Decimal — each suited to a different requirement.

    • — a 64-bit IEEE 754 double-precision value, for general-purpose calculations and fractional values.

    • — an arbitrary-precision integer, for whole numbers beyond the 64-bit range.

    Document-Oriented Nature

    The document as the unit of exchange — header, data, and sections.

    Internet Object is document-oriented: the unit of exchange is a self-contained document, not a bare value or a loose row. A single document bundles three things that other formats usually keep apart — the schema that describes the data, the data itself, and any metadata about it — into one stream. A --- separator divides the document into two regions: a header and a data section.

    Everything above --- is the header (here, a count metadatum and the schema); everything below is the data (two records). The header is read once and governs all the data that follows.

    A document is two regions separated by a single ---:

    ]

    Close square bracket

    U+005D

    Ends an array

    ,

    Comma

    U+002C

    Separates values within an array

    Encoding
    invent a schema that the value did not carry.
  • SHOULD honor schema serialization hints (for example a number format, or a string quote style).

  • Value

    a data row — John, 30

    how is each value written?

    Document

    header + --- + data sections

    Two levels

    What a conformant writer must do

    In this section

    See Also

    Round-Trip Guarantees
    Value Formatting
    Key Emission
    Value Formatting
    Record & Document Output
    Round-Trip Guarantees
    Parsing & Errors
    Conformance Requirements
    Readers & Writers

    does the schema travel with the data?

    Emit the header at most once.
    The header (if any) is emitted before the first data record, and the writer MUST emit the
    ---
    terminator per
    — even for an empty header.
  • Switch schemas only when needed. A writer SHOULD emit a schema switch only when the effective schema actually changes.

  • Never emit invalid control frames. A writer MUST NOT emit midstream definition-mutation control frames in v1, nor unresolved or undefined schema switches.

  • Never emit the legacy headerless form. A writer always emits the --- terminator, so it never produces the buffer-to-end-of-stream legacy form described in Wire Format & Framing.

  • Write sequentially. Writer calls MUST be issued sequentially; the protocol does not define concurrent-write framing in v1.

  • A writer MAY forward pre-framed Internet Object text verbatim — a "raw forward" capability. When it does, the caller is responsible for correct framing, and the writer's automatic schema-switch tracking is no longer reliable for subsequent structured writes. After a raw forward, the next structured write MUST carry an explicit schema selector if the active schema may have changed.

    A reader consumes the stream incrementally and emits items in wire order. It MUST:

    • Process incrementally. A reader MUST process records incrementally and MUST NOT re-parse or re-materialize already-emitted records.

    • Stay bounded in memory. A reader's memory growth MUST be bounded by pending undecoded bytes, the current incomplete frame (one record, or the header), and minimal lookahead — not by total stream history.

    • Reuse compiled state. A reader MUST reuse accepted definitions and resolved schema objects across records; an unchanged schema context MUST NOT trigger repeated schema compilation.

    • Be single-consumption. A reader is single-consumption in v1. Consuming it more than once has undefined behavior.

    • Be lazy. A reader MUST advance the source only as the consumer requests the next item. This laziness is what provides read-side backpressure for pull-based sources.

    • Release on early termination. When the consumer stops early, the reader MUST release the underlying source (for example, release a stream lock) and discard buffered-but-unemitted bytes.

    • Support cancellation. A reader SHOULD support cooperative cancellation. When cancelled, iteration terminates fatally and the source is released; cancellation MUST NOT emit a record-error item.

    An adapter is a transport bridge — it moves bytes between a transport and the reader or writer. A transport is the underlying channel. Their obligations keep the record protocol intact end to end:

    • Adapters are transport bridges only. They MUST NOT bypass, replace, or fork core parsing, validation, or serialization.

    • Adapters MUST preserve record order and MUST NOT invent, merge, or discard logical records.

    • Adapters that consume bytes MUST preserve correctness across chunk boundaries (see the encoding rules in Wire Format & Framing).

    • Adapters MUST NOT downgrade a fatal control-state error into a record-error item.

    • A flow-controlled transport's write operation SHOULD resolve only after the transport has accepted the frame per that transport's backpressure model, and a writer MUST honor that backpressure.

    • A large same-schema stream MUST keep parsing incrementally — without recompiling the schema per record and without retaining emitted records.

    • Read-side backpressure is provided naturally by pull-based sources: a lazy reader only advances the source when the consumer asks for the next item.

    • Push-based sources may lack producer backpressure. An implementation that offers a push source MUST document this, and producers MUST apply their own flow control.

    An implementation is conformant if it satisfies every MUST and MUST NOT in this chapter and passes the shared, language-neutral conformance corpus that accompanies the protocol.

    Implementations SHOULD also verify the equivalence rule directly against their own core: streamed output MUST equal non-streamed core output for the same input and definitions, including across arbitrary chunk boundaries. The strongest form of this test feeds the same input split every possible way — whole, per line, per byte, and at random boundaries including mid-multibyte and mid-marker — and asserts that every split produces identical items. This single property catches the overwhelming majority of streaming defects, because it forces the framing layer to be invisible to the result.

    • Wire Format & Framing — the framing the writer produces and the reader consumes

    • Stream Items — the items a reader emits

    • Streaming Error Model — recoverable versus fatal disposition for the reader

    • Conformance Requirements — the format's broader conformance duties

    Writer obligations

    Raw forwarding

    Reader obligations and lifecycle

    Adapters and transports

    Performance and backpressure

    Conformance

    See Also

    Wire Format & Framing

    invalid-datetime

    missing-

    a mandatory thing is absent

    missing-value

    undefined-

    a referenced name has no definition

    undefined-schema

    unknown-

    not a member of an allowed set

    unknown-member

    reserved-

    a name this specification reserves for a future version

    reserved-type

    duplicate-

    appears more than once

    duplicate-member

    unexpected-

    appears where the grammar disallows it

    unexpected-token

    unterminated-

    an opened construct is never closed

    unterminated-string

    forbidden-

    present but explicitly disallowed

    forbidden-null

    out-of-range-

    a value does not fit the type's own range

    out-of-range-integer

    mismatched-

    violated a constraint the schema author declared

    mismatched-max

    empty-

    empty where content is required

    empty-memberdef

    choices

    mismatched-choice

    multipleOf

    mismatched-multiple-of

    precision / scale

    mismatched-precision / mismatched-scale

    anyOf

    mismatched-any-of

    decimal

    expected-decimal

    invalid-decimal

    n

    bigint

    expected-bigint

    invalid-bigint

    dt'…'

    datetime

    expected-datetime

    invalid-datetime

    d'…'

    date

    expected-date

    invalid-date

    t'…'

    time

    expected-time

    invalid-time

    b'…'

    binary

    (pending)

    invalid-binary

    expected-

    the required type or token is absent, or a different one was found

    expected-integer

    invalid-

    present and of the right kind, but malformed

    min / max

    mismatched-min / mismatched-max

    minLen / maxLen / len

    mismatched-min-len / mismatched-max-len / mismatched-len

    pattern

    0x 0o 0b

    number

    expected-number

    invalid-number

    The predicate vocabulary

    Choosing the subject

    Four distinctions worth learning

    expected- vs missing-

    expected- vs invalid-

    out-of-range- vs mismatched-

    undefined- vs unknown- vs reserved-

    What a code must never be

    See Also

    What a code must never be
    Streaming: Error Model
    deliberate
    Error Model
    Error Accumulation
    Conformance Requirements

    mismatched-pattern

    m

    Structural character (terminates the string)

    {

    Open curly bracket

    U+007B

    Structural character (terminates the string)

    }

    Close curly bracket

    U+007D

    Structural character (terminates the string)

    [

    Open square bracket

    U+005B

    Structural character (terminates the string)

    ]

    Close square bracket

    U+005D

    Structural character (terminates the string)

    "

    Double quote

    U+0022

    Not permitted anywhere in an open string — see below

    '

    Single quote

    U+0027

    Not permitted anywhere in an open string — see below

    r'raw'

    a raw string — the annotation this rule exists for

    (space, tab, etc.)

    Whitespace

    Multiple

    Terminates the string; cannot start or end it

    :

    Colon

    U+003A

    Structural character (terminates the string)

    ,

    Comma

    don't stop

    unknown-annotation — don is not an annotation

    o'clock

    unknown-annotation

    5'9

    Valid forms

    A quote ends the run, and starts an annotation

    Beginning with a digit

    Optional behaviors

    Comments

    Invalid forms

    Preservation of structure

    See Also

    annotated string
    Value formatting
    A number, or a word that begins with a digit?
    Strings
    Regular Strings
    Raw Strings

    U+002C

    unexpected-token

    Presentation

    How is this value written?

    format, encloser, escapeLines

    no

    yes

    Both (?*) — score?*: { number, min: 0 }.
  • Default — the second positional value (or keyed default:) supplies a value when the member is omitted: role?: { string, guest }.

  • Constraint

    Is this value allowed?

    min, max, choices, pattern, len, multipleOf

    yes

    Constraints and presentation

    Optional, nullable, and default

    Quoted names take the long form

    MemberDef vs. SchemaDef

    Nested shapes

    See Also

    Object (SchemaDef)
    TypeDef
    Overview
    Object (SchemaDef)
    Schema References

    no

    Space used in Ogham script.

    U+2000

    En quad

    Space equal to the width of the lowercase letter "n".

    U+2001

    Em quad

    Space equal to the width of the uppercase letter "M".

    U+2002

    En space

    Space equal to half the width of the em space.

    U+2003

    Em space

    Space equal to the width of the em space.

    U+2004

    Three-per-em space

    Space equal to one-third of an em space.

    U+2005

    Four-per-em space

    Space equal to one-quarter of an em space.

    U+2006

    Six-per-em space

    Space equal to one-sixth of an em space.

    U+2007

    Figure space

    Space equal to the width of a numeral.

    U+2008

    Punctuation space

    Space used for punctuation.

    U+2009

    Thin space

    Space narrower than the regular space.

    U+200A

    Hair space

    Very narrow space used for special purposes.

    U+2028

    Line separator

    Separates lines of text.

    U+2029

    Paragraph separator

    Separates paragraphs of text.

    U+202F

    Narrow no-break space

    Non-breaking space narrower than the regular space.

    U+205F

    Medium mathematical space

    Space used in mathematical notation.

    U+3000

    Ideographic space

    Space used in East Asian scripts.

    U+FEFF

    Byte order mark (BOM)

    Zero-width non-breaking space, often used as a BOM.

    MUST NOT
    appear within identifiers or keys.

    U+0000 to U+0020

    Space, line feed, carriage return, tab, etc.

    Any character with code point <= 0x20. Includes the ASCII space and control characters.

    U+1680

    EBNF definition

    Whitespace characters

    Rules

    Best practices

    See Also

    Encoding
    Strings

    Ogham space mark

    Negative value

    0b

    Binary prefix

    Multiple

    Begins a binary BigInt

    0o

    Octal prefix

    Multiple

    Begins an octal BigInt

    0x

    Hex prefix

    Multiple

    Begins a hexadecimal BigInt

    n

    BigInt suffix

    U+006E

    Marks the value as a BigInt

    0–9

    Digits

    Multiple

    Decimal digits

    -

    Minus sign

    Valid forms

    Decimal BigInt

    Alternative bases

    Invalid forms

    Preservation of structure

    See Also

    Numeric Values
    Number
    Decimal

    U+002D

    { *: [type] }

    every extra member must be an array of type

    { *: $ref }

    every extra member must match the referenced SchemaDef

    { * } (or * listed last)

    open — extra members allowed, any type

    { *: type }

    every extra member must match type

    { *: { … } }

    Allowing extra fields with *

    * bare is grammar; "*" quoted is a name

    Typing the extra fields

    Map-shaped objects

    Dynamic types with anyOf

    When to use curly braces

    See Also

    name suffixes
    Any
    Internet Object Schema
    Any
    Object (SchemaDef)

    every extra member must match the inline SchemaDef

    choices

    array

    Restricts the value to a fixed set.

    min

    datetime

    Earliest allowed value (inclusive).

    max

    datetime

    Latest allowed value (inclusive).

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix.

    null

    bool

    If true, the member may be null. Shorthand: * suffix.

    N, not nullable

    forbidden-null error

    omitted, optional (?)

    absent

    omitted, required

    missing-value error

    time

    t'14:30:45'

    time of day

    datetime

    dt'2024-03-20T14:30:45Z'

    date + time (ISO 8601)

    type

    string

    datetime, date, or time. First positional value.

    default

    datetime

    valid temporal literal

    the value

    below min / above max

    mismatched-min / mismatched-max error

    N, nullable (*)

    TypeDef

    Constraints

    min / max

    Optional, nullable & defaults

    See Also

    Date and Time
    Constraints and presentation
    Temporal kind
    Date and Time (value syntax)
    TypeDef
    MemberDef

    Value used when the member is omitted.

    null

    the document's sections, their names, and their schema bindings

    a redundant record enclosure

    normalized to the canonical form

    a name that the schema can recover

    omitted by design — that is the point of the format

    the notation of a number written without a schema

    0xff and 255 are one value; a schema-less number is written in decimal

    the temporal literal of a value written without a schema

    the kind is inferred from the instant — declare date / time to fix the spelling

    comments

    not part of the value model

    whitespace, indentation, line breaks

    insignificant

    the original quote style of a string

    What is preserved

    What is deliberately not preserved

    Testing a writer

    See Also

    Value Formatting
    Conformance Requirements
    Value Formatting
    Record & Document Output

    the writer picks the leanest valid form

    Not a Number

    An undefined or unrepresentable numeric result

    Inf

    Infinity

    A value beyond the finite range

    -

    Minus sign

    Negative infinity (-Inf)

    +

    Plus sign

    Explicit positive infinity (+Inf); not preserved

    Valid forms

    Not the special value

    Preservation of structure

    See Also

    Number
    Numeric Values
    The header MUST be buffered and resolved as a single atomic frame before any data record is processed. The header MUST NOT be resolved piecemeal. A reader buffers from the start of the stream up to the terminating ---, resolves the whole block at once, and only then begins emitting records.

    Definitions are header-only in v1. After the first logical data record begins, the header phase is over. There is no normative midstream definition-mutation syntax — a stream cannot add or change definitions once data has started.

    A record is validated under exactly one schema context:

    • --- $Name selects a named schema for the records that follow it.

    • A bare --- resets the active section to the default schema context.

    • A record that carries no explicit selector is validated through the active default schema context.

    The schemaName reported on a stream item reflects the explicit selector declared for that record's section, and when present MUST include the leading $ sigil:

    • If an explicit selector applied (for example $User), schemaName is that name.

    • If the record was validated only through the active default schema context with no explicit selector, schemaName MUST be absent. An implementation MUST NOT synthesize a name (such as $schema) merely because a default schema was active.

    A schema switch changes the parsing and serialization context; it does not by itself produce a stream item.

    Here the reader emits two record items: the first reports schemaName $User, the second reports $Order. The two --- control frames are applied but never emitted as items.

    A --- $Name selector MUST reference an already-defined schema. A switch to an unknown or invalid schema is not a recoverable per-record error — it is a fatal stream error, because the section's records can no longer be validated against a known shape. The error preserves its core identity (the undefined-schema validation error); see Streaming Error Model.

    If no default schema exists, a bare --- selects the schemaless default context rather than inventing a stream-only schema.

    A reader MAY be constructed with preloaded definitions and an optional fallback default schema context before any stream bytes are read. This is how out-of-band schemas — agreed once between a publisher and a subscriber — are supplied to the reader. If none are provided, the initial definitions state is empty.

    Precedence MUST match core's external-definitions behavior:

    • In-stream header definitions override matching preloaded keys.

    • An in-stream $schema overrides the fallback default schema context.

    • If the stream defines no $schema, the fallback default schema context remains active.

    This is what makes the shared, out-of-band deployment mode work: the wire can carry only data sections (or a header with metadata but no schema), while the reader validates against the schema it already holds. For the two deployment modes — embedded versus shared — see Schema-First Design.

    Header-defined definitions become shared stream state and apply to later records. A reader MUST reuse accepted definitions and resolved schema objects across records: an unchanged schema context MUST NOT trigger repeated schema compilation. All definition lookup, schema resolution, default handling, and member validation MUST be delegated to core — a reader MUST NOT embed its own copy of the rules for default, optional, null, choices, or extensible-schema handling.

    • Definitions — the header definition block streaming resolves atomically

    • Schema References — how $name references resolve

    • Schema-First Design — embedded versus shared (out-of-band) schemas

    • Streaming Error Model — why an unknown schema switch is fatal

    The header is resolved atomically

    Definitions
    Schema References

    Selecting a schema per record

    Unknown schema switches are fatal

    Preloaded definitions and precedence

    State reuse

    See Also

  • Header — information about the data: the schema, reusable definitions, and metadata.

  • Data — the values themselves: one object, or a collection of records.

  • The header is optional. The simplest document is just a value, with no header and no separator at all:

    As soon as you need a schema, definitions, or metadata, you add a header and close it with ---. A document may also be header-only (a header followed by --- with no data) — useful for sending a schema or configuration on its own.

    Each header entry sits on its own line, introduced by a tilde ~. The header carries three kinds of thing:

    • Schema — the shape and types of the data. The reserved key $schema names the document's default schema.

    • Definitions — reusable building blocks: value variables (@name) and references ($name) that the schema or data can point to.

    • Metadata — plain keys such as count, status, or paging fields. Metadata describes the payload and is surfaced separately from the data, not mixed into it.

    Because the header is parsed once and then applied to every record, the cost of describing the data is paid a single time, no matter how many records follow. See Definitions for the full header model.

    The data section holds either a single object or a collection of records, each record introduced by ~. The defining trait is what the records don't carry: since the field names and types live in the header, each record carries only values, not repeated keys.

    Compare this to repeating "name":, "age":, and "city": on every record, as a key-per-value format would. With no schema at all, the values are still accepted and mapped to positional keys (0, 1, 2, …). See Data Sections and Collection.

    A document is not limited to a single dataset. Additional --- separators introduce further sections, each able to name its own schema — so related datasets travel together in one document:

    Each section may carry a name, a schema, or both; an unnamed section takes the default name data. This makes one document a natural container for, say, a result set plus its lookup tables, or several record types from one API response. The precise rules for naming and selecting section schemas are in Data Sections.

    Because a document carries its own schema and metadata, it is self-describing: a receiver can understand and validate exactly what was sent, with no out-of-band agreement.

    Versus JSON. JSON transmits data but has no place for a schema or document-level metadata — the contract is shipped and versioned separately, and the two can drift apart. An Internet Object document keeps the contract and the data in one stream.

    Versus CSV. CSV has rows but no types, no nesting, and no metadata. Internet Object records are typed by the header and may nest objects and arrays, while staying just as compact row-to-row.

    For a fuller comparison, see Why Internet Object?.

    • Separation of concerns — structure, metadata, and data are stated in distinct regions, so each can be read and reasoned about on its own.

    • Compactness — keys and types are declared once in the header; records repeat values, not names.

    • Self-describing & portable — schema, metadata, and data move as one unit, so a document validates itself wherever it lands.

    • Streaming — once the header is read, records can be produced and consumed incrementally, one at a time, without waiting for the whole document.

    • Many datasets, one document — sections bundle related data without inventing an envelope format.

    • Internet Object Document — the document in depth

    • Header · Data Sections

    • Schema-First Design — the other half of the model

    • Why Internet Object? — how it compares to JSON, CSV, and YAML

    Anatomy of a document

    The header — information about the data

    The data — the values

    One document, many sections

    Self-contained and self-describing

    Why it matters

    See Also

    expected-string    expected-number    expected-integer   expected-decimal
    expected-bigint    expected-boolean   expected-object    expected-array
    expected-datetime  expected-date      expected-time
    ~ $schema: { age: { int, max: 120 } }
    ---
    ~ "thirty"     # ✗ expected-integer  — the TYPE is at fault
    ~ 200          # ✗ mismatched-max    — the CONSTRAINT is at fault
    ~ $schema: { name: string, age: int }
    ---
    ~ John, "thirty"   # ✗ expected-integer — age is present, but is not an integer
    ~ John             # ✗ missing-value    — age is absent altogether
    ~ $schema: { when: datetime }
    ---
    ~ "2024-03-20"       # ✗ expected-datetime — a plain string, not a datetime literal
    ~ dt"not-a-date"     # ✗ invalid-datetime  — a datetime literal that does not parse
    ~ $schema: { small: int8, capped: { int, max: 120 } }
    ---
    ~ 200, 100     # ✗ out-of-range-integer — 200 does not fit `int8`; widen the TYPE
    ~ 100, 200     # ✗ mismatched-max       — 200 breaks a declared `max`; fix the DATA
    ~ $schema: { a: $Missing }   # ✗ undefined-schema — referenced, never defined
    ---
    ~ 1
    ~ $schema: { a: strng }      # ✗ unknown-type — no such type; a typo
    ---
    ~ 1
    ~ $schema: { a: int64 }      # ✗ reserved-type — a real name, not usable in this version
    ---
    ~ 1
    John Doe
    जॉन डो
    Wow Great
    😃
    013ABSD
    12mm
    जॉन डो, Wow Great, 😃
    Lorem ipsum dolor sit amet consetetur sadipscing elitr sed
    diam nonumy eirmod.
    
    Tempor invidunt ut labore et dolore magna aliquyam erat
    sed diam voluptua
    ---
    013ABSD              # → an open string
    0x123FG              # ✗ invalid-number — announced hex, and G is not a hex digit
    ---
     John Doe      # leading space is trimmed → "John Doe"; to keep it, write " John Doe"
    ---
    "John Doe"     # a REGULAR string, not an open string
    age: { number, minimum: 10 }
    ---
    42                       # ✗ unknown-member — number has no option "minimum" (use min)
    ~ $schema: { flags: { uint16, format: hex } }
    ---
    ~ 255          # ✓ accepted — decimal input is fine
    ~ 0xff         # ✓ accepted — the same value
    ~ $schema: { name: string, role?: { string, guest }, nickname?*: string }
    ---
    ~ John                       # role → "guest"; nickname omitted
    ~ Mary, admin, N             # role "admin"; nickname null
    ~ $schema: { "a,b"?*: number }               # ✗ invalid-definition
    ---
    ~ 1
    ~ $schema: { "a,b": { number, optional: T, "null": T } }   # ✓ the same member, spelled out
    ---
    ~ 1
    ~ $schema: { name: string, "a,b": { number, optional: T } }
    ---
    ~ John, 1
    ~ Jane
    ~ $schema: {
        address: { street: string, city: string },   # SchemaDef — a nested object shape
        age:     { int, min: 0, max: 120 }            # MemberDef — a type with constraints
    }
    ---
    ~ { Main St, NYC }, 30
    ~ $schema: { name: string, meta: { author: string, version: { int, min: 1 } } }
    ---
    ~ John, { Jane, 2 }
    whitespace         = ascii_whitespace | unicode_whitespace ;
    
    ascii_whitespace   = ? any character with Unicode code point U+0000 to U+0020 ? ;
    unicode_whitespace = U+1680 | U+2000 | U+2001 | U+2002 | U+2003 | U+2004
                       | U+2005 | U+2006 | U+2007 | U+2008 | U+2009 | U+200A
                       | U+2028 | U+2029 | U+202F | U+205F | U+3000 | U+FEFF ;
    123n                 # positive BigInt
    -42n                 # negative BigInt
    0n                   # zero
    9007199254740992n    # beyond the safe Number range
    ---
    42n, 0x2An, 0b101010n, 0o52n
    0xn                  # ✗ invalid-number — missing hex digits
    0bn                  # ✗ invalid-number — missing binary digits
    ---
    123.45n              # ✗ a BigInt cannot have a fractional part (use Decimal)
    ~ $schema: { name: string, age: int }
    ---
    ~ John, 30                # ✓
    ~ Alex, 25, extra1        # ✗ unknown-member
    ~ $schema: { name: string, age: int, * }
    ---
    ~ John, 30                       # ✓
    ~ Alex, 25, Male, cool           # ✓ extras at index 2 and 3
    ~ { Mia, 28, role: dev }         # ✓ extra keyed field "role"
    ~ $schema: { name: string, * }         # extensible — accepts any extra field
    ~ $schema: { name: string, "*": int }  # CLOSED — declares a member called *
    ~ $schema: { "*": string, admin: string }
    ---
    ~ allow, deny            # the member named * holds "allow"
    ~ $schema: { name: string, *: string }
    ---
    ~ { John, role: dev }     # ✓
    ~ { Alex, code: 123 }     # ✗ expected-string — extra value must be a string
    ~ $schema: { name: string, *: { string, minLen: 4 } }
    ---
    ~ { John, dept: Sales }   # ✓
    ~ { Mia, id: "12" }       # ✗ mismatched-min-len — extra is shorter than 4
    ~ $question: { questionName: string, points: number }
    ~ $questions: { *: $question }        # map: ANY key, every value must match $question
    ~ $schema: { questions: $questions }
    ---
    { QID1: { Q2, 5 }, QID2: { Q1, 3 } }
    test: { any, anyOf: [string, number] }
    ---
    ~ One     # ✓
    ~ 1       # ✓
    ~ Two     # ✓
    name, age, address: { street, city, state }, isActive
    ---
    ~ Alice, 30, { Main St, NYC, NY }, T
    address: { street, city }    # `address` is an object with two members
    address: string              # `address` is a string; there is no nested object
    address: street              # ✗ unknown-type — `street` is read as a TYPE name, not a member
    created: datetime, birthday: date, opensAt: time
    ---
    ~ dt'2024-03-20T14:30:00Z', d'1990-05-01', t'09:00:00'   # ✓
    when: { datetime, min: dt'2024-01-01T00:00:00Z', max: dt'2024-12-31T00:00:00Z' }
    ---
    ~ dt'2024-06-01T00:00:00Z'    # ✓
    deletedAt?*: datetime   # optional + nullable
    ---
    ~ {}    # ✓ omitted → absent
    ~ N     # ✓ null
    write(parse(write(parse(x))))  ==  write(parse(x))
    NaN                  # Not a Number
    Inf                  # positive infinity
    -Inf                 # negative infinity
    +Inf                 # positive infinity (explicit sign, not preserved)
    nan                  # open string "nan", not NaN
    INF                  # open string "INF", not Inf
    infinity             # open string "infinity", not Inf
    -NaN                 # open string "-NaN" — NaN cannot be signed
    ~ $User: { name: string }
    ~ $Order: { id: int }
    --- $User
    ~ Alice
    --- $Order
    ~ 1001
    ~ count: 2
    ~ $schema: { name: string, age: int }
    ---
    ~ John, 30
    ~ Jane, 25
    John, 30
    ~ $schema: { name: string, age: int, city: string }
    ---
    ~ John, 30, Phoenix
    ~ Jane, 25, Dallas
    ~ $person: { name, age: int }
    ~ $address: { street, city }
    --- $person
    ~ John, 30
    ~ Jane, 25
    --- $address
    ~ Main St, NYC

    Implementations are not versioned by this document. They follow their own Semantic Versioning and declare which specification version they implement. The specification's clock and an implementation's release clock are independent.

    Every feature carries exactly one tier. Tiers are tracked per feature, not per specification release.

    Tier
    Meaning
    May change

    Stable

    Part of the frozen specification contract.

    Only in a new specification major

    Candidate

    Feature-complete and under review; intended to become Stable. Not yet part of the frozen contract.

    As analogues: Draft is close to TC39 Stage 1–2 or Kubernetes alpha; Candidate to Stage 3 (candidate) or beta; Stable to Stage 4 or GA.

    • A Stable feature MUST NOT change incompatibly except in a new specification major.

    • A Candidate feature is feature-complete and SHOULD be treated as near-final, but MAY still change — with a changelog notice — before graduating to Stable.

    • A Draft feature MAY change or be removed at any time and MUST be clearly marked.

    • A feature MUST be Deprecated for at least one major cycle before removal.

    • Each page declares its tier in a status: field in its front matter; the page is generated from those fields, so the dashboard never drifts from the pages. To change a status, edit the page's status: and regenerate.

    • Draft → Candidate: the design is complete and reviewed.

    • Candidate → Stable: behavior is final and consistent across implementations; graduation is announced in the changelog. Two further conditions apply, and both are deliberate:

      1. More than one implementation. "Consistent across implementations" cannot be established by the implementation the specification was written alongside. Until a second, independent implementation has been built from these pages and agrees, a promotion would be self-certification.

      2. A soak period of at least three months after the reference implementation is publicly released. A specification is a claim about what people will need; that claim is tested by use, not by review. Three months of real documents finds the assumptions no reviewer thought to question.

      Consequently no page is Stable today, and none is expected to be for some time. That is the policy working, not a backlog.

    • Stable → Deprecated → Removed: with a replacement and a target major.

    • An implementation MAY implement Candidate or Draft features but SHOULD mark them as such in its own API (for example, SemVer-exempt or @experimental).

    • A feature is normally promoted to Stable only once it is interoperably implemented and verified — for example, by a shared conformance suite. Until then it stays Candidate.

    • Implementations declare the specification version and which Candidate or Draft features they include.

    Declare a specification major (such as 1.0) final when:

    Specification changes — especially anything affecting a Stable feature, and every feature graduation — MUST be recorded in the Version History. Day-to-day finalization work is tracked in the Roadmap.

    • Feature Status — the current tier of each feature

    • Roadmap — the finalization plan

    • Version History — the specification changelog

    • Conformance Requirements — what a conformant implementation must do

    What is versioned

    Feature Status
    Streaming
    (proposed) → Draft → Candidate → Stable → Deprecated → Removed
                    │         │
                    └─────────┴── may still change while Draft or Candidate

    Stability tiers

    Core rules

    Graduation and deprecation lifecycle

    Relationship to implementations

    1.0 readiness checklist

    Changelog

    See Also

    Decimal — a fixed-precision decimal with exact arithmetic, for financial values and other cases that require precise decimal representation.

    Internet Object accepts several notations. The table distinguishes integer-only bases from fractional notations and gives a recommendation.

    Integer-only bases. Binary (base 2), octal (base 8), and hexadecimal (base 16) can represent only integers. For fractional values, use base-10 decimal or scientific notation.

    Notation
    Supported forms
    Recommendation

    Decimal integer (base 10)

    Number, BigInt

    BigInt for large integers; Number for general use

    Decimal fractional

    Number, Decimal

    Each form is identified by a distinct suffix (Number has none):

    Two senses of "decimal". Decimal (base 10) is the common numeral system used by all forms. Decimal (the form) is the fixed-precision type for exact arithmetic.

    Use case
    Recommended form
    Reason

    General calculations

    Number

    Standard performance and compatibility

    Financial amounts

    Decimal

    • Value Representations — all value types

    • NaN and Infinity — special numeric values (Number only)

    • Numeric Types — numeric schemas and constraints

    Numeric forms

    Number
    BigInt
    42          # Number  (standard floating-point)
    42n         # BigInt  (arbitrary-precision integer)
    42.50m      # Decimal (fixed-precision decimal)

    Number bases and notations

    Distinguishing the forms

    Choosing a form

    See Also

    Overview

    How Internet Object schemas describe the shape of data, and the pieces that make them up.

    An Internet Object schema describes the shape of the objects in a document: their members, the type of each value, and the constraints each value must satisfy. A schema is written in the same object syntax as the data it validates, so it is compact, readable, and easy to author by hand.

    Schema and data share one syntax. Unlike map-based schema languages (JSON Schema, XML Schema), an IO schema looks like the object it validates. There is no second grammar to learn.

    A schema is normally declared once in the header — as the default $schema or as a reusable $ reference — and applied to every record in the data section. Validating a value against a schema either succeeds or produces a stable error code.

    The shape of an object

    A schema is a comma-separated list of members. In its simplest form it is just a list of keys; each value then has the any type:

    Give a member a type by writing key: type:

    Add constraints by replacing the bare type with a MemberDef — a small object whose first value is the type and whose remaining entries are constraints:

    A value outside the constraint fails validation with a stable code:

    Members nest: a member's type can itself be an object schema or an array:

    A schema is assembled from a few orthogonal pieces. Each has its own reference page:

    Piece
    What it does
    Reference

    The exact placement rules — open vs. closed objects, keyed vs. positional members, and how the default $schema is chosen — are covered in .

    A member name can carry a marker that changes how a missing or null value is treated. These markers are first-class schema features, not informal conventions: a conformant validator MUST honor them.

    • Optional (?) — the member MAY be omitted from the data.

    • Nullable (*) — the member MAY be null (N).

    • Both (

    See for the full resolution rules and for accepting members beyond those listed.

    The shape of a member, informally:

    The complete, normative grammar for documents and schemas is in the appendix.

    A header defines a reusable $address and a default $schema; the data section is validated against it:

    Both records validate: John Doe supplies an address; Jane Doe omits the optional address.

    • — open/closed, keyed/positional, default schema

    • — the base types and shortcuts

    • — types, constraints, optional/nullable/default

    • — reusable $ schemas and types

    Data Sections

    The data section — section separators, objects, and collections.

    The data section is where the actual data of an Internet Object document resides. A document can have one or more data sections, each introduced by a separator line (---) and optionally labelled with a section name and schema. The data itself is either a single object or a collection of objects, giving a flexible yet structured way to represent information. The diagram below shows the shape of a data section.

    Internet Object document data section structure

    Structure overview

    Section separator line

    Each data section begins with a separator line (---) that divides the document into distinct sections. The separator can carry two optional elements:

    • Section name — identifies the section and its purpose.

    • Schema name — names the schema that constrains the section, prefixed with $.

    Separator line. The separator line MUST end with a newline (\n) or EOF (end of file).

    The separator can take several forms, from least to most detailed, each ending with a newline (\n) or EOF:

    • Without name and schema — the simplest form, just the separator (---).

    • With section name — the separator followed by a name (--- employee).

    • With section name and schema — a name and schema name, separated by a colon (--- employee : $employee).

    • What a section name may contain — a section name is a bare name: letters, marks, digits, - and _. It has no quoted form, so unlike a member name it cannot carry a colon, a comma, a space, or any other structural character. See below.

    • Omitting the section name — in a multi-section document, the section name may be omitted only once. When omitted, the name is derived from the associated schema (e.g. --- $employee implies the section name employee).

    The separator line is read to the end of the line, so a section name has no delimiters to mark where it stops. It is therefore a bare name and the only one in the format that cannot be quoted:

    This is narrower than a member name, which may be quoted and can then contain anything at all. The asymmetry is deliberate — a member name sits inside a record, where the quotes bound it, whereas a section name sits on a line of its own.

    Two obligations follow, one for each side.

    A reader MUST reject a name outside that set and report invalid-section-name. It must not accept a prefix and discard the rest: the production is anchored, so --- a,b: $x is an error rather than a section named a, and leading or trailing spaces are not part of the name and are not silently absorbed. Accepting a truncated name would change the data with no diagnostic — the one outcome the format does not tolerate.

    Known implementation gap. The reference implementation currently truncates rather than reporting: --- a,b: $x reads the name as a and surfaces only a downstream unexpected-token, and --- lead: $x silently drops the leading space. Tracked as .

    A writer MUST NOT emit a section name it cannot read back. When converting foreign data whose keys fall outside the set — code:en, a,b, a key with a leading space — the multi-section layout is simply unavailable, and the writer MUST fall back to a single section, where those keys become ordinary member names and may be quoted:

    See .

    The simplest form. It uses the default section name (data) and the document's default schema.

    Here the section name is employee. The schema is the document's default schema.

    Here both the name and schema are stated explicitly, as employee and $employee.

    Here only the schema is named. The section name is derived from the schema name (employee). If that name is already used elsewhere in the document, it is an error.

    After the separator line comes the data. It is either a single object or a collection of objects — the flexibility that lets the format carry many kinds of information efficiently.

    Objects are structured entities composed of key-value pairs. Each object is written within curly braces {} and may contain nested objects or other values, forming a hierarchy.

    Collections are lists of objects, allowing multiple records within one data section. Each object in a collection is written the same way as a standalone object but belongs to the broader collection. See for record syntax, type promotion, and validation rules.

    Unlike JSON, a data section may hold a value that is not an object — an array, or a bare scalar. Because a record binds values to names, such a value is promoted into a record under its positional key:

    The same promotion applies per row in a collection, so ~ [1, 2] is the record { "0": [1, 2] }. A writer converting foreign data MUST use this binding rather than invent a member name, or its output would decode differently from the identical document written by hand — see .

    A single object can follow the separator directly.

    A single-section document with no header or schema does not need a separator, so the example above can also be written as:

    A collection lists objects, each prefixed with ~ on its own line:

    A data section may be empty — just the separator line with no data.

    An entirely empty document needs no separator at all.

    A document can include multiple sections, each with its own data:

    Organized by separators and built from objects and collections, the data section offers a robust, flexible way to carry data — keeping documents clear, consistent, and effective across a wide range of applications.

    • — what precedes the --- separator

    • · — value syntax

    • — records and collection rules

    Record & Document Output

    Writing records, record enclosure, headers, and data sections.

    Key Emission and Value Formatting settle how the pieces are written. This page settles how they are assembled into records, sections, and a document.

    Records

    A record is one row of data — a ~ item in a collection, or the single object of an object section. Its members are written in schema order, separated by , .

    name: string, age: int
    ---
    ~ John, 30
    ~ Mary, 25

    Absent members hold their place

    A declared member that is absent, optional, and has no default is written as an empty position, so that later members are not read into the wrong slot:

    # schema: { a: string, b?: number, c: string }   value: a = p, c = q
    {p, , q}          # correct — c stays in slot 2
    {p, q}            # WRONG — q would be read as b

    Trailing empty positions carry no information and are trimmed.

    Record enclosure

    A record's own braces are optional: x, 4 and {x, 4} are the same record. But when a record consists of exactly one value and that value is an object, the braces become ambiguous — a reader may take them as the record's own enclosure rather than as the value.

    A writer MUST therefore enclose such a record explicitly:

    The full reading rule lives with the value syntax; see . A writer does not rely on that rule — it always emits the unambiguous form.

    A document is a header, a --- separator, and one or more data sections. Whether the header travels with the data is a writer choice:

    Header
    Output

    Header omitted — data only, and no separator:

    Header included:

    A writer MUST NOT infer a header the document does not carry. When a schema-less document is written with the header included, the header is empty but the --- separator is still emitted, so the first token of the data is unambiguous:

    The data then follows the no-schema rules in — every name is unrecoverable, so every name is written.

    A data section may hold a value that is not an object, and IO promotes it into a record under its positional key: --- followed by [1, 2, 3] decodes as { "0": [1, 2, 3] } ().

    A writer converting foreign data MUST bind such a value to that same positional member. Naming it anything else produces a document that decodes differently from the identical text written by hand:

    Both forms below parse; the second is wrong because it decodes differently. An invented member name (value: [number] … [1, 2, 3]) yields { value: [1, 2, 3] }, so the same data written by the library and by hand would disagree.

    An array whose items are records is a collection and needs no promotion: each record becomes a row. Promotion applies only where there are no names to bind to — an array of scalars, an array of arrays, or a bare scalar.

    A member name is quoted by the same rules as a data key (). Since the ? and * suffixes belong to the bare-name token, a writer that quotes a name MUST expand that member to the long MemberDef form:

    The member "a,b" is an optional, nullable number. The suffix form is not available to it:

    so a writer emits the long form instead:

    Bare names are unaffected — age?*: number is written as it stands. Appending a suffix to a quoted name produces a header the writer's own reader rejects with invalid-definition.

    Because the two parts are independent, a writer SHOULD expose them separately as well as combined: the schema can then be published, cached, or versioned on its own while records stay lean. Composing the header, a blank line, and the data reproduces the whole document exactly.

    A section is introduced by ---. A section may be named, schema-bound, or both:

    Form
    Meaning

    A multi-section document writes each section in order. A writer SHOULD separate the header and each named or schema-bound section with a blank line; blank lines are insignificant to a reader, so this affects legibility only.

    A section name is a bare name — letters, marks, digits, - and _ — and it is the one name in the format that cannot be quoted, because the separator line runs to the end of the line and nothing would bound it.

    That makes the multi-section layout unavailable for some data. When a key falls outside the set, a writer MUST NOT emit it as a section name and MUST fall back to a single section, where the same key is an ordinary member name and may be quoted:

    This is the general rule of applied to one construct: a writer must never emit text its own reader cannot read. A leading space is the case worth remembering — it is not part of the name, and a reader that absorbs it changes the data without reporting anything.

    • — record enclosure and the reading rule

    • ·

    BigInt

    The bigint type — arbitrary-precision integers.

    The bigint type validates an arbitrary-precision integer — values too large for the double-based number family (see Numeric Types). In data it is written with an n suffix: 123n, 0xFFn.

    For the literal syntax, see BigInt values.

    TypeDef

    A bigint MemberDef accepts only the options below. Any other key is invalid.

    Option
    Type
    Description

    Selects the base a writer uses. It is — any notation is still accepted as input.

    The base prefix and the n suffix are both part of the written literal, so the output re-parses as the same bigint:

    A bigint has no fractional part, so its scientific mantissa is an integer and its exponent is never negative: trailing zeros move into the exponent, and a value with none is written e0.

    Resolution follows the :

    • ·

    • ·

    Schema Data Types

    The Internet Object schema type system — base types, shortcuts, and TypeDefs.

    A schema constrains each member to a type. The type system is small by design: a handful of base types, plus a closed set of shortcuts — names that stand for a base type with preset constraints baked in. Every type is configured through its TypeDef, the fixed set of options it accepts.

    Base types

    Type
    Validates
    Reference

    A member written without a type defaults to any, so name, age declares two any members.

    date, time, and datetime are their own types, not subtypes of string. Earlier drafts described them as string-derived; they are temporal types with their own literal values. See .

    A shortcut is a built-in name equal to a base type plus preset constraints. It is not a new type — it is a convenient, validated configuration of a base type. A conformant validator MUST recognize every shortcut name.

    Base
    Shortcuts
    Each shortcut is…

    See for the full numeric family and ranges, and for email and url.

    Reserved. int64, uint64, float32, and float64 are reserved for future use and are not yet validated by the reference implementation.

    Each type defines a TypeDef — the exact set of options it accepts (for example, string accepts pattern, minLen, maxLen; number accepts min, max, multipleOf). Supplying a type together with chosen options produces a MemberDef, the definition of a single member:

    The first value in a MemberDef is the type; the second is the default; the third is choices. Remaining options are written as key: value pairs. An option a type does not define is rejected. Each type page lists its TypeDef in full.

    • — how schemas are built

    • — the option-contract model

    • — types, constraints, optional/nullable/default

    • · ·

    Any

    The any type — accepts any value, optionally constrained by anyOf or choices.

    The any type accepts a value of any type. It is the default when a field is declared without a type (name is the same as name: any). You can still narrow it with choices or, for a union of types, anyOf.

    a, b: any, c: { type: any }
    ---
    ~ hello, 42, T        # ✓ — anything goes

    Declaring alternatives (anyOf)

    anyOf lets a field accept any one of several types or MemberDefs — Internet Object's union type.

    id: { any, anyOf: [string, int] }
    ---
    ~ 42        # ✓ matches int
    ~ abc       # ✓ matches string
    flag: { any, anyOf: [bool, int] }
    ---
    ~ hello     # ✗ matches neither

    Each alternative may be a full MemberDef or a SchemaDef:

    value: { any, anyOf: [{ int, multipleOf: 5 }, { int, multipleOf: 3 }] }
    ---
    ~ 10        # ✓ multiple of 5
    ~ 9         # ✓ multiple of 3

    TypeDef

    An any MemberDef accepts only the options below.

    Option
    Type
    Description
    • ·

    Number

    Standard 64-bit IEEE 754 floating-point numbers.

    A number is a 64-bit double-precision floating-point value conforming to IEEE 754. Numbers are scalar values used to express integers, fractional values, and special numeric constants.

    Numbers support several representations: decimal, the alternative bases (binary, octal, hexadecimal), scientific notation, and the special values NaN and Inf.

    A number can be written in several forms:

    Symbol
    Name
    Unicode

    Array

    The array type — ordered, typed collections of values.

    The array type validates an ordered list of values. Unlike a scalar type, an array is a container: besides constraining the list itself (its length), you declare the type of its elements. An array with no element type accepts items of any type.

    For the array value syntax ([a, b, c]), see . This page covers the array schema type.

    There are two equivalent ways to declare what an array holds: the [ … ] shorthand and the keyed of: form.

    Other Special Characters

    Functional modifiers — variable, schema, optional, nullable, and sign characters.

    Special characters work alongside structural characters and literals to add functionality or context to an Internet Object document. Each has a specific semantic meaning and modifies the behavior of schemas, values, or parsing.

    Symbol
    Name
    Unicode
    Context
    Application

    Decimal for exact/financial values; Number otherwise

    Binary (base 2)

    Number, BigInt

    BigInt for large binary integers; Number otherwise

    Octal (base 8)

    Number, BigInt

    BigInt for large octal integers; Number otherwise

    Hexadecimal (base 16)

    Number, BigInt

    BigInt for large hex integers; Number otherwise

    Scientific notation

    Number, Decimal

    Decimal for precise values; Number otherwise

    Special values

    Number only

    NaN and Infinity for undefined/infinite results

    Exact precision, no rounding error

    Large counters or IDs

    BigInt

    No precision limit for integers

    Scientific notation

    Number

    Built-in floating-point support

    Cryptographic values

    BigInt

    Handles arbitrarily large integers

    With notice, before it graduates

    Draft

    Provisional; still evolving. Use at your own risk.

    At any time

    Deprecated

    Still specified; scheduled for removal.

    Removed in the next specification major

    Reserved

    Syntax or semantics reserved for future definition; not yet specified.

    May be defined at any time

    Informative

    A non-normative page (guides, rationale, appendices); carries no maturity guarantee.

    n/a

    Feature Status

    With only schema — the separator followed by just the schema name (--- $employee).

    Default section name and schema — if both the name and schema are omitted, the section name defaults to data and the document's default schema is used.

  • Unique section names — each section MUST have a unique name; duplicate names are not allowed. A document that repeats one is invalid, but a parser still recovers from it: the duplicate is renamed (data → data_2, users → users_2) and the error is reported, so no section is lost. See Duplicate section names.

  • Rules for section names and schemas

    Section names are bare names

    Examples of section separators

    Separator line without name and schema

    Separator line with a section name

    Separator line with a section name and schema

    Separator line with only a schema

    Data

    Objects

    Collections

    Root values that are not objects

    Examples of data

    Single object

    Collection of objects

    Empty data section

    Multi-section document example

    See Also

    Section names are bare names
    ISSUE-20
    Record & Document Output
    Collection
    Record & Document Output
    Header
    Objects
    Arrays
    Collections

    omitted

    data only, no separator — the schema is assumed known at the endpoint

    included

    header, ---, then the data

    ---

    an unnamed section using the default schema

    --- $Schema

    an unnamed section bound to a named schema

    --- name: $Schema

    Documents

    A root value that is not a record

    Member names in the header

    Header and data are separately addressable

    Sections

    See Also

    Record enclosure under schema validation
    Key Emission
    Data Sections
    Value Formatting
    Round-trip
    Objects
    Data Sections
    Header
    Collection

    a named section bound to a named schema

    choices

    array of bigint

    Restricts the value to a fixed set.

    min

    bigint

    Minimum allowed value (inclusive).

    max

    bigint

    Maximum allowed value (inclusive).

    multipleOf

    bigint

    The value must be an exact multiple of this.

    format

    string

    Presentation, write-only. Base used when writing: decimal (default), hex, octal, binary, scientific.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix.

    null

    bool

    If true, the member may be null. Shorthand: * suffix.

    0x124f80n

    octal

    0o377n

    0o4447600n

    binary

    0b11111111n

    0b100100100111110000000n

    scientific

    255e0n

    12e5n

    type

    string

    The type name bigint. First positional value.

    default

    bigint

    Value used when the member is omitted. Second positional value.

    format

    255n is written

    1200000n is written

    decimal (default)

    255n

    1200000n

    hex

    format

    Examples

    Optional, nullable & defaults

    See Also

    write-only
    common rules
    BigInt values
    Numeric Types
    Decimal
    TypeDef
    MemberDef

    0xffn

    string

    Text

    String Types

    number

    An IEEE-754 number

    Numeric Types

    bigint

    An arbitrary-precision integer (123n)

    BigInt

    decimal

    A fixed-precision decimal (123.45m)

    Decimal

    bool

    true / false

    Bool

    date, time, datetime

    Temporal values (d'…', t'…', dt'…')

    Date and Time

    binary

    Byte data, written as base64 (b'…')

    Binary

    object

    A structured shape (a SchemaDef)

    Object (SchemaDef)

    array

    An ordered list of values

    Array

    any

    Any value; the default when no type is given

    Any

    string

    email, url

    string with a built-in pattern

    number

    int, uint, int8, int16, int32, uint8 (byte), uint16, uint32

    Shortcuts

    TypeDef and MemberDef

    See Also

    Date and Time
    Numeric Types
    String Types
    Overview
    TypeDef
    MemberDef
    Numeric Types
    String Types
    Date and Time

    number restricted to whole values in a fixed range

    choices

    array

    Restricts the value to a fixed set (of any type).

    anyOf

    array of MemberDef/type

    The value must match one of these.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix.

    null

    bool

    If true, the member may be null. Shorthand: * suffix.

    type

    string

    The type name any.

    default

    any

    choices

    Optional, nullable & defaults

    See Also

    Union Types (anyOf)
    TypeDef
    MemberDef

    Value used when the member is omitted.

    At sign

    U+0040

    Variable

    Prefixed to a name, declares or references a variable

    $

    Dollar sign

    U+0024

    Schema

    Prefixed to a name, declares or references a schema

    ?

    Question mark

    U+003F

    Schema

    Suffixed to a bare member name, marks the member optional

    *

    Asterisk

    U+002A

    Schema

    Suffixed to a bare member name, marks the member nullable; bare on its own, makes a schema accept undeclared members. Quoted ("*") it is an ordinary name, in a schema or in data

    -

    Hyphen / minus

    U+002D

    Numeric

    Marks a negative value

    +

    Plus

    U+002B

    Numeric

    Marks a positive value

    • Context sensitive — a character's meaning depends on its position and context.

    • Variable prefix — @ prefixes variable declarations and references.

    • Schema prefix — $ prefixes schema definitions and references.

    • Schema suffixes — ? and * are suffixed to bare member names in a schema; a quoted name uses the keyed optional: and "null": options instead.

    • Bare versus quoted — a special character is only special when written bare. "*" is a member name, not the wildcard; "a?" is a name ending in ?, not an optional a.

    • Numeric prefixes — + and - prefix numeric values to indicate sign.

    • Case sensitive — all special characters are case-sensitive.

    • Reserved usage — these characters are reserved for their specific functions.

    • Definitions — variables and schema references

    • Numeric Values — numeric formatting and signs

    • Structural Elements — overview of all structural characters

    Special character set

    @

    Usage examples

    Variable references and schema definitions

    Schema modifiers

    Numeric signs

    Character rules

    See Also

    sectionName = ( letter | mark | digit | "-" | "_" )+
    # data: { "code:en": […], "a,b": […] }
    # WRONG - neither name survives the round trip
    --- code:en: $a
    --- a,b: $b
    
    # RIGHT - one section; the keys are member names, which may be quoted
    --- $schema
    { "code:en": […], "a,b": […] }
    ---
    ~ John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    ~ Jane Doe, 20, Male, {Duke Street, New York, NY}
    --- employee
    ~ John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    ~ Jane Doe, 20, Male, {Duke Street, New York, NY}
    --- employee : $employee
    ~ John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    ~ Jane Doe, 20, Male, {Duke Street, New York, NY}
    --- $employee
    ~ John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    ~ Jane Doe, 20, Male, {Duke Street, New York, NY}
    ---
    [1, 2, 3]                # the record { "0": [1, 2, 3] }
    ---
    42                       # the record { "0": 42 }
    ---
    John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    ---
    ~ John Doe, 25, Male, {Bond Street, New York, NY}, [agile, swift]
    ~ Jane Doe, 20, Male, {Duke Street, New York, NY}
    ---
    --- $library
    # Bookville Library
    City Central Library, "123 Library St, Bookville"
    
    --- $books
    ~ The Great Gatsby, "F. Scott Fitzgerald", 1234567890, T, [Fiction, Classic], 1925
    ~ "1984", George Orwell, 2345678901, F, [Fiction, Dystopian], 1949, { user123, d"2024-02-20"}
    
    --- subscribers: $users
    ~ user123, John Doe, Standard, [{2345678901, d"2024-01-20"}]
    ~ user456, Jane Smith, Premium, []
    # schema: { o1: object, o2?: object }   value: o1 = { key: val }
    {{key: val}}      # correct — outer braces are the record, inner are the value
    {key: val}        # ambiguous — leaves the reading to the first-key rule
    John, 30
    name: string, age: int
    ---
    John, 30
    ---
    a: 1
    # input [1, 2, 3] — REQUIRED: decodes as { "0": [1, 2, 3] }, as hand-written IO would
    "0": [number]
    ---
    [1, 2, 3]
    ~ $schema: { "a,b"?*: number }               # ✗ invalid-definition
    ---
    ~ 1
    ~ $schema: { "a,b": { number, optional: T, "null": T } }   # ✓ what a writer emits
    ---
    ~ 1
    ~ $accounting: {n: string, a: number}
    ~ $sale: {n: string, a: number}
    
    --- accounting: $accounting
    ~ John, 23
    
    --- sales: $sale
    ~ Sal, 27
    # data: { "code:en": […], "a,b": […] }
    # WRONG - `--- code:en: $a` reads back as a section named `code`, and the rest fails
    --- code:en: $a
    
    # RIGHT - one section; the keys are member names
    --- $schema
    { "code:en": […], "a,b": […] }
    mask: { bigint, format: hex }
    ---
    ~ 255n                  # ✓ accepted — written back as 0xffn
    ~ 0xffn                 # ✓ the same value
    id: bigint
    ---
    ~ 123n                  # ✓
    ~ 0xFFn                 # ✓ (255)
    ~ 99999999999999999999999999999n   # ✓
    big: { bigint, min: 100n }
    ---
    ~ 50n     # ✗ mismatched-min
    id?*: bigint    # optional + nullable
    ---
    ~ {}     # ✓ omitted → absent
    ~ N      # ✓ null
    ~ 7n     # ✓
    name:  { string, minLen: 1, maxLen: 100 },   # a string MemberDef
    score: { int, min: 0, max: 100 }             # a number MemberDef
    ---
    ~ John, 85
    pick: { any, choices: [1, One, T] }
    ---
    ~ One       # ✓
    ~ Two       # ✗ mismatched-choice
    note?*: any
    ---
    ~ {}        # ✓ omitted → absent
    ~ N         # ✓ null
    # Variable declarations
    ~ @r: red
    ~ @g: green
    ~ @b: blue
    # A schema using variables in an inline constraint
    ~ $schema: {
        name: string,
        email: email,
        joiningDt: date,
        color: {string, choices: [@r, @g, @b]}
    }
    ---
    # Data using variable references
    ~ John Doe, '[email protected]', d'2020-01-01', @r
    # Optional and nullable member declarations
    ~ $user: {
        name: string,          # Required member
        email?: string,        # Optional member (may be omitted)
        avatar*: string,       # Nullable member (may be null)
        metadata*?: object,    # Optional and nullable member
    
        # The suffixes attach to a bare name. A quoted name says the same with options.
        "code:en": { string, optional: T }
    }
    
    # A schema that accepts undeclared members
    ~ $flexible: {
        id: string,
        name: string,
        *                      # Accept additional, undeclared members
    }
    # Positive and negative numbers
    temperature: +23.5         # Explicit positive
    balance: -150.75           # Negative value
    elevation: +8848           # Positive integer
    debt: -5000                # Negative integer

    TypeDef

    The fixed set of options each type accepts

    Optional, nullable & defaults

    ?, *, and default values

    Open & dynamic schemas

    Allowing extra members with * and *: type

    References

    Reusable schemas and types ($name)

    Union types

    A value matching one of several types (anyOf)

    Composition & reuse

    Building large schemas from small ones

    ?*
    ) — the member may be omitted or
    null
    .
  • Default — the second value in a MemberDef supplies a value to use when the member is omitted (equivalently, the keyed default: option).

  • Error Model — the validation error catalogue

  • JSON Compatibility — mapping to and from JSON Schema

  • name, age, address
    ---
    John, 30, { Main St, NYC }
    name: string, age: int, isActive: bool
    ---
    John, 30, T
    score: { int, min: 0, max: 100 }
    ---
    85
    score: { int, min: 0, max: 100 }
    ---
    150          # ✗ mismatched-max — score is above max
    name: string, address: { street: string, city: string }
    ---
    John, { Main St, NYC }

    MemberDef

    One member: a type plus its constraints

    MemberDef

    Data types

    The base types and built-in shortcuts (string, int, email, …)

    ~ $schema: { name: string, email?: string, nickname?*: string, role?: { string, guest } }
    ---
    ~ John                          # email & nickname omitted, role defaults to "guest"
    ~ Mary, [email protected], N, admin    # nickname is null, role is "admin"
    schema    = member ( "," member )*
    member    = openMember | [ key modifier? ":" ] typeOrDef
    openMember = "*" [ ":" typeOrDef ]
    modifier  = "?" | "*" | "?*"
    typeOrDef = typeName | ref | memberDef | "{" schema "}" | "[" typeOrDef "]"
    memberDef = "{" ( typeName | ref ) ( "," constraint )* "}"
    ref       = "$" name
    ~ $address: { street: string, city: string, zip?: int }
    ~ $schema: {
        name: string,
        age?: int,
        email: { string, pattern: "^[^@]+@[^@]+$" },
        isActive: bool,
        address?: $address
    }
    ---
    ~ John Doe, 30, [email protected], T, { Bond Street, New York }
    ~ Jane Doe, 28, [email protected], F

    The building blocks

    Optional, nullable, and default members

    Schema grammar

    A complete example

    See Also

    Schema Representation
    MemberDef
    Open & Dynamic Schemas
    Formal Grammar
    Schema Representation
    Schema Data Types
    MemberDef
    Schema References

    Description

    0–9

    Digits

    Multiple

    Decimal digits

    .

    Decimal point

    U+002E

    Separates the integer and fractional parts

    -

    Minus sign

    Decimal numbers may be integers or fractional values, with an optional sign. A leading or trailing decimal point is allowed (.5 and 5.).

    Numbers can be written in binary, octal, or hexadecimal. The prefix is case-insensitive, and hex digits may be upper or lower case.

    Scientific notation uses e or E (case-insensitive) for the exponent, which may be signed.

    The same value can be written in several bases and notations:

    All five values above are 42.

    Open strings may begin with a digit, and people write such values constantly: 3pm, 12mm, 007th, part codes like 013ABSD, version strings like 1.2.3, addresses like 10.0.0.1. So a reader needs a rule for when a run of characters is a broken number rather than ordinary text — and it cannot be "it looks numeric", because most of those do.

    Two rules decide it.

    Rule 1 — all or nothing

    A run is a number only if the entire run is a valid number literal. If anything is left over, the whole run is an open string.

    Rule 2 — a marker is a claim

    The base prefixes 0x, 0o, 0b and the type suffixes m, n can only mean number. A run that carries one and is not a valid literal of that type is an error, not a string.

    Written
    Read as
    Which rule

    0xFF, 1.2, 12e5, 123.45m

    a number

    Rule 1 — the whole run is valid

    013ABSD, 12mm, 3pm

    open string

    Rule 1 is what keeps a partial parse from inventing a value. 1e is not a complete number, so it is the string "1e" — never the number 1, which is what an implementation produces if it keeps the part it managed to read and discards the rest. That is the failure the rule exists to prevent, and it needs no error to prevent it: text that stays text loses nothing.

    Rule 2 is why 0oz is rejected while 013ABSD is not. Nothing in 013ABSD says number; 0o says nothing else. Quoting is the escape hatch and is always available — "0oz" is simply a string — and a writer is required to quote any string that would otherwise read back as a broken literal, so a value that arrives as text leaves as text.

    And the runs that are not errors, because nothing in them claims to be a number:

    Internet Object preserves:

    • The chosen notation (decimal, binary, octal, hex, scientific)

    • Whitespace (non-significant)

    • Syntactic fidelity as written, except that an explicit + sign is not preserved

    It does not interpret:

    • Mathematical relationships between values

    • Precision beyond IEEE 754

    • Domain-specific numeric constraints

    Those semantics belong to the schema, the validator, or the application.

    • Numeric Values — all numeric forms and notations

    • BigInt — arbitrary-precision integers

    • Decimal — fixed-precision decimal arithmetic

    • NaN and Infinity — special numeric values

    • — numeric schemas

    number = ["-" | "+"] (
        decimalNumber
      | binaryNumber
      | octalNumber
      | hexNumber
      | scientificNumber
    ) | specialValue
    
    decimalNumber    = digit+ [ "." digit* ] | "." digit+
    binaryNumber     = "0b" binaryDigit+
    octalNumber      = "0o" octalDigit+
    hexNumber        = "0x" hexDigit+
    scientificNumber = ( digit+ [ "." digit* ] | "." digit+ ) ("e" | "E") ["-" | "+"] digit+
    specialValue     = "NaN" | "Inf" | "-Inf" | "+Inf"
    
    digit       = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9"
    binaryDigit = "0" | "1"
    octalDigit  = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7"
    hexDigit    = digit | "A" | "B" | "C" | "D" | "E" | "F" | "a" | "b" | "c" | "d" | "e" | "f"

    Syntax

    Structural characters

    # Integers
    42                   # integer
    -17                  # negative integer
    +17                  # positive integer (explicit sign, not preserved)
    
    # Fractional
    3.14159              # fractional
    -0.5                 # negative fractional
    .5                   # leading dot -> 0.5
    5.                   # trailing dot -> 5
    
    # Zero
    0
    +0
    -0
    # Binary (0b / 0B), digits 0-1
    0b1010               # 10
    0B1111               # 15
    -0b1010              # -10
    
    # Octal (0o / 0O), digits 0-7
    0o755                # 493
    0O644                # 420
    
    # Hexadecimal (0x / 0X), digits 0-9 A-F
    0xFF                 # 255
    0XDeadBeef           # 3735928559
    0xff                 # 255 (lower-case digits)
    1.23e4               # 1.23 × 10^4  = 12300
    1.23e-4              # 1.23 × 10^-4 = 0.000123
    -2.5e+3              # -2.5 × 10^3  = -2500
    5e3                  # 5 × 10^3     = 5000
    .5e2                 # 0.5 × 10^2   = 50
    6.022e23             # Avogadro's number
    ---
    42, 0x2A, 0b101010, 0o52, 4.2e1
    0x                   # ✗ invalid-number — claimed hex, gave no digits
    0b                   # ✗ invalid-number — claimed binary, gave no digits
    0b 1010              # ✗ invalid-number — the space does not rescue the claim
    0xGH                 # ✗ invalid-number — G and H are not hex digits
    0o89                 # ✗ invalid-number — 8 and 9 are not octal digits
    1.2.3                # → "1.2.3"      a version string
    10.0.0.1             # → "10.0.0.1"   an address
    1e                   # → "1e"         incomplete, so not a number
    013ABSD              # → "013ABSD"    a part code

    Valid forms

    Decimal numbers

    Alternative bases

    Scientific notation

    Equivalent forms

    A number, or a word that begins with a digit?

    Invalid forms

    Preservation of structure

    See Also

    The element type may be any type, an inline object shape, or a reference:

    The shorthand holds a type, not constraints. Use [ … ] for a bare element type ([string], [[int]], [{ name: string }]). To add list-level constraints such as len, switch to the keyed form: { array, of: string, len: 3 }. Writing constraints inside the brackets — { [string], len: 3 } — is not valid and raises a syntax error.

    An array MemberDef accepts only the options below. Any other key is invalid.

    Option
    Type
    Description

    type

    string

    The type name array.

    default

    array

    len precedence. When len is set, minLen and maxLen are ignored.

    Each element MUST satisfy the element type, or validation fails per item:

    Applies only when the member is omitted; must be a valid array.

    An element type may itself be an array, giving nested or fixed-size multidimensional arrays. Both forms below are valid:

    The suffixes attach to a bare member name. A name that has to be quoted uses the equivalent options — "a,b": { array, of: string, optional: T } — as described in MemberDef.

    Input
    Result

    value present, valid

    the array

    not an array

    expected-array error

    item count outside minLen/maxLen/len

    • The { [type], …constraints } combined form is not supported (use { array, of: type, … }).

    Adapted from the playground (all forms below parse):

    • Arrays (value syntax)

    • Object (SchemaDef) — for object element shapes

    • Schema References — reusable element types

    • TypeDef · MemberDef

    tags: array                  # any array (elements unconstrained)
    tags: []                     # same — any array
    tags: [string]               # array of strings (shorthand)
    tags: { array, of: string }  # array of strings (keyed form)

    Declaring the element type

    Arrays
    people:  [{ name, age, role }]    # array of objects (inline shape)
    authors: [$person]                # array of a referenced schema
    scores: [int]
    ---
    ~ [1, 2, 3]      # ✓
    ~ [1, two, 3]    # ✗ expected-integer on the second item
    top3: { array, of: string, len: 3 }
    ---
    ~ [a, b, c]      # ✓
    ~ [a, b]         # ✗ mismatched-len — must have exactly 3 items
    tags: { array, of: string, minLen: 1 }
    ---
    ~ []             # ✗ mismatched-min-len — must have at least 1 item
    ~ [a]            # ✓
    tags?: { array, of: string, default: [] }
    # Array of arrays of integers (shorthand)
    matrix: [[int]]
    ---
    ~ [[1, 2], [3, 4]]
    # Fixed 3×3 integer matrix (keyed form, with len)
    grid: { array, of: { array, of: int, len: 3 }, len: 3 }
    ---
    ~ [[1, 1, 1], [1, 1, 1], [1, 1, 1]]
    tags?: [string]    # optional (key may be missing)
    tags*: [string]    # nullable (value may be N)
    tags?*: [string]   # optional and nullable
    anything:    array,                              # any array
    strings:     [string],                           # array of strings
    min2strings: { array, of: { string, minLen: 2 }, minLen: 2 },
    objects:     { array, of: { name, age, role } }
    ---
    [1, two], [aa, bbb], [aaa, bbbb], [{ John Doe, 25, Student }, { Jane Doe, 30, Teacher }]

    TypeDef

    Constraints

    Element type (of)

    len / minLen / maxLen

    default

    Nested and multidimensional arrays

    Optional, nullable & defaults

    Implementation status (beta)

    Examples

    See Also

    Numeric Types

    The number type and its family of integer, unsigned, and float shortcuts.

    The number type validates a numeric value. It is the base of a small family of predefined shortcuts — int, uint, int8, byte, and so on — where each shortcut is simply number with a fixed set of constraints baked in. For example, int8 is number restricted to whole values in −128…127.

    Two related numeric types have their own pages: (arbitrary-precision integers, suffix n) and (fixed-precision decimals, suffix m). For how numbers are written (decimal, hex, octal, binary, scientific, NaN, Inf), see .

    Each name below is the base number type plus preset constraints. A conformant validator MUST recognize all of these names.

    Type
    Whole number?
    Range

    byte is an alias for uint8. A value outside a type's own range MUST be rejected with out-of-range-integer — the limit belongs to the type, so the fix is to widen the type. A value that breaks a min/max the schema declared is a different error; see .

    Whole-number rule. The int/uint/int8…32/uint8…32 shortcuts MUST reject values with a fractional part (int accepts 42, not 42.5). number and float accept any finite double.

    Reserved (not yet supported). int64, uint64, float32, and float64 are reserved for future use; the reference implementation currently rejects them.

    A number MemberDef accepts only the options below. Any other key is invalid.

    Option
    Type
    Description

    Inclusive bounds. A value below min is rejected with mismatched-min, above max with mismatched-max — the code names the keyword that rejected it, so the direction is never in doubt.

    These are distinct from out-of-range-integer, which means the value does not fit the type (int8 given 200, with no bound declared). One is fixed by changing the data, the other by widening the type.

    An explicit min/max replaces a shortcut's built-in bound rather than narrowing it — e.g. { int8, min: -200 } currently accepts −200. See Implementation status below.

    The value must be an exact multiple of the given number.

    Restricts the value to a fixed set. As the third positional value it may omit the key.

    Controls how the number is written on serialization. It is : it does not restrict which input notations are accepted (any is read).

    The base prefix is part of the written literal, and a sign precedes it (-0xff). A value with a fractional part has no radix literal, so a radix format does not apply to it and the decimal notation is written instead.

    NaN and Inf/-Inf are valid only for number/float. When combined with a numeric bound, the reference implementation currently coerces them to null rather than validating them — see Implementation status.

    How a member resolves (verified behavior):

    Input
    Result

    The suffixes are shorthand for the keyed options — c?*: number and c: { number, optional: T, "null": T } declare the same member. A member name that has to be quoted can only use the keyed form; see .

    A few behaviors on this page describe the agreed target; the reference implementation is catching up:

    • Whole-number enforcement for the int family is not yet applied (int currently accepts 3.14).

    • byte alias is being added (currently only uint8 is recognized).

    A schema mixing notations and family members (adapted from the playground):

    Any input notation works for any number type — 0x11, 0o21, 0b10001, and 17 all denote the same value.

    • — how numbers are written

    • · — related numeric types

    • · — how type options work

    String Types

    The string type and its email and url shortcuts.

    The string type validates text. It has two predefined shortcuts that are string with a built-in pattern: and .

    date, time, and datetime are not string subtypes — they are their own types with their own values. See .

    For how strings are written (open, quoted, raw), see .

    Value Formatting

    How each scalar, string, and key is written so that it reads back unchanged.

    A writer MUST preserve a value's type, not merely its printed appearance. 42 and 42n print similarly and mean different things; a string that looks like a number must not read back as a number. This page fixes the written form of every value kind.

    Type
    Written as
    Note

    Object (SchemaDef)

    The object type — structured key/value data described by a SchemaDef.

    The object type validates structured key/value data. Like an array, it is a container: you declare its shape — the set of fields and their types. That shape is called a SchemaDef.

    For the object value syntax ({ … }), see .

    A field may itself be any type, nested object, array, or .

    In a record with several fields, write nested objects in the open positional form (~ John, { 1, 2 }). A record written wholly as { … }

    Collection

    The structure of a collection — an ordered sequence of records in a data section.

    A collection is an ordered sequence of records within a data section of a document. Each record is an object, written on its own line and introduced by a tilde ~. Collections make it efficient to serialize, batch, and stream many objects — datasets, tables, event logs — in a concise, uniform form.

    Familiar parallels. A collection is conceptually similar to a dataset in CSV, a stream in JSON Lines, or a record array in Avro — but each record is a full Internet Object object.

    A collection is always part of a document, not a standalone document. A document has a header and a data section; a data section holds either a single object or a collection of one or more records. Records may be homogeneous (the same shape) or heterogeneous (each shaped differently); each record is independent, so a failure in one does not affect the rest.

    A collection is one or more records. Each record is a tilde ~ followed by an object:

    Streaming Error Model

    The streaming error model — error categories, recoverable-versus-fatal disposition, and stream-absolute positions.

    Streaming preserves the core error identity of everything semantic and defines its own errors only for the transport and lifecycle it owns. This page specifies the categories an error can have, the disposition that decides whether iteration continues or stops, and the positioning rules that make a streamed error indistinguishable from the non-streaming one. The core errors themselves are defined in ; streaming references them and does not restate them.

    Match on category and code, never on message. Every error carries a stable category and a stable string code. Message text is non-contractual and MAY be localized. Tooling MUST branch on the category and code, not on the human-readable message.

    Every error carries a category and a stable string code. The category MUST be derived from the originating error's class.

    A code has exactly one class, and the class a code belongs to is the group it appears under in the core . The two never disagree: a code in the catalogue's Syntax errors section is raised as a syntax error everywhere, and likewise for validation. That is a requirement on implementations, not an observation — where the two diverge, two conformant readers report different categories for the same input.

    Value used when the member is omitted.

    of

    type / MemberDef

    The element type. Equivalent to the [ … ] shorthand.

    len

    int ≥ 0

    Exact number of elements required.

    minLen

    int ≥ 0

    Minimum number of elements.

    maxLen

    int ≥ 0

    Maximum number of elements.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix on the key.

    null

    bool

    If true, the member may be null. Shorthand: * suffix on the key.

    mismatched-min-len / mismatched-max-len / mismatched-len error

    an item fails of

    the item type's own error, e.g. expected-string

    value is N (null), key is nullable (*)

    null

    value is N (null), key is not nullable

    forbidden-null error

    value omitted, default set

    the default

    value omitted, key optional (?), no default

    absent

    value omitted, required, no default

    missing-value error

    U+002D

    Negative value

    +

    Plus sign

    U+002B

    Explicit positive value (not preserved on output)

    e / E

    Exponent

    Multiple

    Scientific-notation exponent

    0b

    Binary prefix

    Multiple

    Begins a binary number

    0o

    Octal prefix

    Multiple

    Begins an octal number

    0x

    Hex prefix

    Multiple

    Begins a hexadecimal number

    Rule 1 — no marker, so nothing is claimed

    1.2.3, 10.0.0.1, 2024.01.15

    open string

    Rule 1 — likewise; a version is not a number

    1e, 1.23ee4, 5em

    open string

    Rule 1 — incomplete, so not a number at all

    0x123FG, 0b, 0oz

    invalid-number

    Rule 2 — 0x/0b/0o claimed a number

    .45m, 123.m

    invalid-decimal

    Rule 2 — m claimed a decimal

    12.3n

    invalid-bigint

    Rule 2 — n claimed a bigint

    Numeric Types
    Schema Data Types
    TypeDef
    MemberDef
    Open & Dynamic Schemas
    Schema References
    Union Types
    Composition & Reuse

    int

    yes

    unbounded integer

    uint

    yes

    integer ≥ 0

    int8

    yes

    −128 … 127

    uint8 / byte

    yes

    0 … 255

    int16

    yes

    −32 768 … 32 767

    uint16

    yes

    0 … 65 535

    int32

    yes

    −2 147 483 648 … 2 147 483 647

    uint32

    yes

    0 … 4 294 967 295

    choices

    array of number

    Restricts the value to a fixed set. Third positional value.

    min

    number

    Minimum allowed value (inclusive).

    max

    number

    Maximum allowed value (inclusive).

    multipleOf

    number

    The value must be an exact multiple of this.

    format

    string

    Presentation, write-only. Notation used when writing: decimal (default), hex, octal, binary, scientific.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix on the key.

    null

    bool

    If true, the member may be null. Shorthand: * suffix on the key.

    value is N (null), key is nullable (*)

    null

    value is N (null), key is not nullable

    forbidden-null error

    value omitted, default set

    the default

    value omitted, key optional (?), no default

    absent

    value omitted, required, no default

    missing-value error

    Explicit min/max can widen a shortcut's range (under review).
  • NaN/Inf under a bound resolve to null (under review).

  • number

    no

    IEEE-754 double

    float

    no

    type

    string

    Type name (number or any family member). First positional value.

    default

    number

    value present, valid

    the value

    value present, below min / above max

    mismatched-min / mismatched-max error

    value present, outside the type's range

    The number family

    TypeDef

    Constraints

    min / max

    multipleOf

    choices

    format

    Special values

    Optional, nullable & defaults

    Implementation status (beta)

    Examples

    See Also

    BigInt
    Decimal
    Numeric Values
    min / max
    write-only
    notation
    MemberDef
    Numeric Values
    BigInt
    Decimal
    TypeDef
    MemberDef

    IEEE-754 double

    Value used when the member is omitted. Second positional value.

    out-of-range-integer error

    age: int8
    ---
    ~ 200      # ✗ out-of-range-integer — int8 max is 127
    age: { number, min: 18, max: 25 }
    ---
    ~ 18     # ✓
    ~ 25     # ✓
    ~ 35     # ✗ mismatched-max
    rollNo: { number, multipleOf: 5 }
    ---
    ~ 10     # ✓
    ~ -10    # ✓
    ~ 12     # ✗ — must be a multiple of 5
    code: { number, choices: [234, 245, 456] }
    ---
    ~ 245    # ✓
    ~ 5      # ✗ mismatched-choice
    flags: { uint16, format: hex }      # written as 0xff
    mask:  { uint8,  format: binary }   # written as 0b11111111
    a?: { number, 7 }    # optional with default 7  →  omitted yields 7
    b*: number           # nullable  →  N yields null
    c?*: number          # optional and nullable
    ~ $row: { hex: uint8, oct: uint8, bin: uint8, dec: number, sci: { number, min: 999999999 } }
    --- rows: $row
    ~ 0x11, 0o2, 0b11, 10, 4.329e+10
    ~ 0x22, 0o3, 0b100, 20, 2.329e+20
    Type
    Is

    string

    any text

    email

    string validated against an email pattern

    url

    string validated against a URL pattern

    Quote values containing : or spaces (URLs, times-of-day, "Last, First"). An unquoted https://x.com is misread because the open string ends at :.

    A string MemberDef accepts only the options below. Any other key is invalid.

    Option
    Type
    Description

    type

    string

    string, email, or url. First positional value.

    default

    string

    len precedence. When len is set, minLen and maxLen are ignored.

    Length is measured in Unicode code points — NOT bytes, and NOT UTF-16 code units. "café" is 4 and "🙂" is 1, even though the first is 5 bytes in UTF-8 and the second is 2 UTF-16 units.

    This has to be stated because every language's default string length means something different, and three of the common answers are wrong here:

    Code points are the only unit that is a property of the text rather than of an encoding or a runtime, and Internet Object is UTF-8 on the wire, where UTF-16 units have no meaning at all.

    The failure is invisible in ordinary testing: "café" measures 4 under all three interpretations, so a wrong implementation passes every ASCII and Latin-1 test and only diverges on characters outside the Basic Multilingual Plane. The conformance corpus probes it directly (validation/strings-constraints.io).

    Language

    Idiom

    "🙂"

    Python

    len(s)

    1

    ✅

    A regular expression. Use a raw string (r'…') to avoid escaping backslashes.

    Quote choices that look like numbers or contain commas, e.g. ["19.02, 72.85"], so they are treated as strings.

    Input
    Result

    valid text

    the string

    fails a declared constraint

    mismatched-* error, named after the keyword — mismatched-min-len, mismatched-pattern, mismatched-choice, …

    malformed for a sub-format

    • Strings (value syntax)

    • Email · URL

    • Date and Time

    • TypeDef · MemberDef

    The string family

    Email
    URL
    Date and Time
    Strings
    contact: email, site: url
    ---
    ~ [email protected], 'https://example.com'    # ✓
    ~ notanemail, 'https://example.com' # ✗ invalid-email
    name: { string, minLen: 5, maxLen: 20 }
    ---
    ~ Ethan              # ✓
    ~ Alexandra Daddario # ✓
    ~ Leo                # ✗ mismatched-min-len
    ssn: { string, pattern: r'^[0-9]{3}-[0-9]{2}-[0-9]{4}$' }
    ---
    ~ '123-45-6789'   # ✓
    ~ '12345678'      # ✗ mismatched-pattern
    dept: { string, choices: [cs, mech, civil] }
    ---
    ~ cs     # ✓
    ~ art    # ✗ mismatched-choice
    nickname?*: { string, anonymous }   # optional + nullable, default "anonymous"
    ---
    ~ {}      # ✓ → "anonymous" (omitted, default applies)
    ~ N       # ✓ → null
    ~ John    # ✓ → "John"

    TypeDef

    Constraints

    minLen / maxLen / len

    pattern

    choices

    Optional, nullable & defaults

    See Also

    number

    42, 3.14, -7

    plain literal

    bigint

    42n

    the n suffix is required, or it reads back as a number

    decimal

    3.14m

    the m suffix is required, or it reads back as a number

    boolean

    T / F

    true / false are also valid

    null

    N

    null is also valid

    datetime

    dt"2024-03-20T14:30:00.000Z"

    also d"…" (date) and t"…" (time)

    binary

    b"SGVsbG8="

    base64 payload

    special numbers

    Inf, -Inf, NaN

    the literals the grammar defines

    A member may declare a presentation option that selects a spelling — format on the numeric and string types, plus encloser and escapeLines on strings. These are write-only: a writer MUST honor them, and a reader MUST NOT enforce them.

    A numeric value carrying a radix format (hex, octal, binary) is written in that base with its base prefix and its type suffix — 0xffn, not ff — so that the output still reads back as the same typed value. A format selects a base, never a type, and never licenses output that the member's own schema would reject. A sign is written before the base prefix (-0xff). A value with a fractional part has no radix literal, so a radix format does not apply to it and the decimal spelling is written instead.

    Where no schema is available, a writer chooses the spelling itself, under one rule: infer only what the value evidences. Any spelling that re-parses to an equal value is permitted; a writer MUST NOT guess a form the value does not support. A notation such as hex is not part of the value — 0xff and 255 parse identically — so a schema-less number is written in decimal.

    The three temporal types share one value — an instant — but not one literal. A writer MUST NOT normalize everything to dt"…"; it selects the literal in this order:

    1. A declared type wins. Under a schema, date writes d"…", time writes t"…", and datetime writes dt"…". The kind is a type name, not a format, so it is never inferred when it has been declared.

    2. Otherwise, infer from the value — under the same "infer only what the value evidences" rule above. An instant whose time component is entirely zero evidences a date; one whose date component is the time-only sentinel 1900-01-01 evidences a time; anything else is a datetime.

    Inference is value-preserving but not text-preserving: a datetime that happens to fall at midnight is written d"…", and one on 1900-01-01 is written t"…". Each re-parses to the very same instant, so round-trip holds — only the spelling may differ from the input text. Declare the type when the spelling matters.

    A string is written in the leanest form that reads back as the same string. A writer chooses among three forms:

    Form
    Written as
    When

    open

    John

    the default — no quoting needed

    regular

    "John"

    A string MUST be quoted when leaving it open would change how it reads back. There are two ways that happens, and both are the writer's responsibility: the text may read back as a different value, or it may fail to read back at all.

    That is the case when the string:

    • is empty;

    • looks like a number — 3.14, 007, -5, .5;

    • carries a base prefix and does not decode — 0x123FG, 0b. A base prefix , so a bare 0x123FG is read as invalid-number rather than as text. Quoting is what tells the reader this is a string, and it is the only thing that can. A run carrying no marker is not in this class however numeric it looks — 1e and 1.23ee4 read back as themselves and are written bare;

    • is a keyword — T, F, N, true, false, null;

    • looks like a date or time — 2024-03-20, 14:30:00;

    • contains a comma, or a structural character that would end the value;

    • contains a section separator (---), which would otherwise split the document;

    • has leading or trailing whitespace, which an open string loses on re-parse.

    Ordinary codes and identifiers need none of this. A writer emits them bare, because they read back as themselves:

    The test is precise on purpose: quote a string when the bare text would read back as a number (Rule 1) or as an error (Rule 2), and not otherwise. Quoting everything that begins with a digit is safe but wrong for a format whose output is meant to be lean — and it hides tokenizer defects, because a quoted value never exercises the path a bare one takes.

    A string containing characters that would need heavy escaping — a literal newline, many backslashes — is written as a raw string instead: r"line1 line2".

    Known gap. A string beginning with @ or $ currently cannot exist as data at all: the reference parser resolves it as a variable or schema reference even when quoted ("@ref", r'@ref'), and fails. The parser fix is tracked with Parsing & Errors; once representable, such strings MUST be written quoted.

    When Key Emission decides a key is written, the key itself follows the same principle: it MUST be quoted whenever writing it bare would not read back as that key.

    A key is written bare only if it is an identifier-like open string. It MUST be quoted when it:

    • is numeric — "5", "3.14";

    • is a keyword — "null", "true", "false", "N", "T", "F";

    • contains a colon, comma, brace, bracket, or quote — "a:b";

    • contains a section separator (---) — "a---b";

    • has leading or trailing whitespace.

    Keys drawn from foreign data — imported JSON, locale tags such as code:en — routinely contain colons. A writer that emits them bare produces output that cannot be re-parsed.

    • Arrays are written [ … ], elements formatted by these same rules. An array whose element type is a schema writes each element positionally.

    • Objects are written { … } and are always enclosed — only a top-level record may use the open form. See Record & Document Output.

    • An untyped object (a member declared object with no shape) has no schema to recover its member names from, so its members are written keyed.

    • Strings · Numeric Values

    • Binary · Date and Time

    • Key Emission — whether a key is written at all

    Type-preserving scalars

    b: 42n, d: 3.14m, t: T, f: F, n: N, dt: dt"2024-03-20T14:30:00.000Z"
    d"2024-03-20"                   # no schema -> written d"2024-03-20"
    t"14:30:00"                     # no schema -> written t"14:30:00"
    dt"2024-03-20T14:30:00.000Z"    # no schema -> written dt"2024-03-20T14:30:00.000Z"
    a: "  pad  ", b: "3.14", c: "T", d: "has, comma", e: "0x123FG"
    ---
    013ABSD, 12mm, 3pm, 1.2.3, 10.0.0.1, 007th, 1e, 1.23ee4
    "5": a, "null": b, "a:b": c, "true": d, "a---b": e, plain: f

    Temporal kind

    Strings

    Keys

    Containers

    See Also

    maps its values to the
    record's
    fields, not to one field's object.

    A field name is written bare when it is identifier-like, and quoted otherwise — when it contains a comma, a colon, a space, or begins with a digit, as names carried over from JSON routinely do:

    A quoted name is taken literally, so the ? and * suffixes cannot follow it. Optional, nullable & defaults below shows what a quoted field writes instead. The same literalness makes "*" an ordinary field name rather than the wildcard — see Open & Dynamic Schemas.

    An object MemberDef accepts only the options below.

    Option
    Type
    Description

    type

    string

    The type name object.

    default

    object

    Objects nest to any depth:

    An empty SchemaDef ({} or object) accepts any object. To allow extra fields beyond those declared, add * to the shape — see Open & Dynamic Schemas:

    A member declared as bare object (or {}) accepts any object value — any members, keyed or positional, at any depth. No structural validation is applied to its contents:

    Because no schema can recover member names for an untyped object, writers serialize its members keyed (key: value) — positional emission would be unrecoverable on re-parse.

    Input
    Result

    valid object

    the object

    field fails its type

    the field's error (e.g. expected-integer)

    N, nullable (*)

    Because the suffixes are part of the bare-name token, a quoted field name states the same two properties as keyed options — where schema: takes the reference that home?*: $address wrote after the colon:

    The two forms are equivalent. MemberDef gives the rule in full.

    An object-typed member — especially as a schema's first member — is what makes the record enclosure question visible. For a row written as a single closed object, whether it is read as the record itself or as a value for member 0 depends on the row's first key:

    A key the schema declares — the row is the record:

    A key it does not declare — the whole row is a value:

    Both readings are well-defined, but the intent is implicit. See Record enclosure under schema validation for the full rule and the best-practice forms ({{…}} or o1: {…}) that state it explicitly — writers always emit the enclosed form.

    • Objects (value syntax)

    • Open & Dynamic Schemas · Schema References

    • TypeDef · MemberDef

    addr: { street: string, city: string }   # inline SchemaDef
    meta: {}                                  # any object (no fixed shape)
    meta: object                              # same as {}
    home: $address                            # a referenced SchemaDef
    name: string, location: { x: int, y: int }
    ---
    ~ John, { 1, 2 }            # ✓ location = { x: 1, y: 2 }
    ~ John, { 1, two }          # ✗ expected-integer (y is not an int)

    Declaring the shape (SchemaDef)

    Objects
    reference
    ~ $schema: { name: string, "code:en": string, "a,b": number }
    ---
    ~ John, hello, 1
    ~ $address: { street: string, city: string }
    ~ $schema: { name: string, home: $address }
    ---
    ~ John, { Main St, NYC }     # ✓
    ~ $schema: { name: string, * }
    ---
    ~ John, extra1, extra2       # ✓ extra fields accepted
    ~ $schema: { metadata: object }
    ---
    ~ {any: thing, nested: {deeply: T}}     # ✓ accepted as-is
    ~ $address: { street: string, city: string }
    ~ $schema: { name: string, home?*: $address }
    ---
    ~ John, { Main St, NYC }     # ✓
    ~ Jane, N                    # ✓ home is null
    ~ $address: { street: string, city: string }
    ~ $schema: { name: string, "home,work": { object, schema: $address, optional: T } }
    ---
    ~ John, { Main St, NYC }     # ✓
    ~ Jane                       # ✓ omitted
    ~ $schema: { o1: object, o2?: object }
    ---
    {o1: {a: 1}}      # → o1 = { a: 1 }
    ~ $schema: { o1: object, o2?: object }
    ---
    {key: val}        # → o1 = { key: val }

    Field names

    TypeDef

    Nesting

    Open and dynamic objects

    Untyped objects

    Optional, nullable & defaults

    Interaction with record enclosure

    See Also

    Here object is an object as defined in the Objects specification, in either open or closed form. An absent body (a bare ~) is an empty record. A bare scalar or array is promoted to an object — see Type promotion.

    Symbol
    Name
    Unicode
    Role

    ~

    Tilde

    U+007E

    Begins a record

    A record (also called a collection item) is the top-level object immediately following a tilde in a collection. Every record is parsed as an object:

    • A record MUST be a valid object — either open form (comma-separated values) or closed form (enclosed in { }).

    • A bare ~ is an empty record and loads as an empty object ({}).

    If a record looks like a single scalar (number, string, boolean, null) or an array, it is promoted to an open object holding that value at positional index 0. Unnamed values in an open object always take positional keys (0, 1, 2, …):

    Every record is an object, whether it is written as a scalar, an array, or an explicit object. A schema later maps these values to named members.

    Open form is the most concise and is the recommended style:

    A record may be a closed object enclosed in { }:

    Quoted keys and standard JSON punctuation are accepted, so JSON-shaped records read naturally:

    Open and closed records may be mixed in one collection:

    • Whitespace around the tilde, commas, and braces is insignificant.

    • Records are usually separated by newlines, but any whitespace works.

    • A trailing comma inside an object is allowed and ignored.

    • Comments (# to end of line) may trail a record or stand alone; they are ignored.

    Some inputs are genuine syntax errors; others parse without error but not the way you might expect. Both are worth recognizing.

    A record (~) cannot follow a bare, non-collection object in the same section:

    An unterminated object or array, or stray tokens after a closed object, are also errors. (Each is reported against the record it appears in; the surrounding records are unaffected.)

    Missing separators do not raise an error — spaces never separate values, so a run of words collapses into a single open string:

    If a record loads with fewer members than you expect, look for a missing comma — spaces alone never separate values.

    Each record is parsed and validated on its own. If a record fails — a syntax error or a validation error — only that record is reported as an error; the records before and after it still load:

    A conformant processor SHOULD collect per-record errors and continue rather than stop at the first failure. See Collection Rules for schema validation, empty-record rules, and a worked example.

    Record order is preserved exactly as written. Whitespace and comments are insignificant and do not appear in the loaded result. How members are named, deduplicated, and mapped to fields is governed by the schema — or, without a schema, by positional index.

    • Objects — the object grammar a record follows

    • Creating Collections — collections with and without a schema

    • Collection Rules — validation, empty records, and error handling

    • Data Streaming — collections produced and consumed over time

    • — where a collection sits in a document

    • — validating records

    Syntax

    collection     = collectionItem+
    collectionItem = "~" [ object ]
    ---
    ~ 1                   # record: { "0": 1 }
    ~ true                # record: { "0": true }
    ~ [red, green, blue]  # record: { "0": [red, green, blue] }
    ~ John Doe            # record: { "0": "John Doe" }
    ~ {1, 2}              # record: { "0": 1, "1": 2 }
    ~                     # empty record: {}
    ~ name: John Doe      # record: { "name": "John Doe" }
    ---
    ~ 101, Thomas, 25, HR, {Bond Street, New York, NY}
    ~                                                    # empty record → {}
    ~ 102, George, 30, Sales, {Duke Street, New York, NY}
    ---
    ~ {Jane Doe, 20, f, N/A, [0xFF0000, 0x0000FF], F}
    ---
    ~ {"name": "John Doe", "address": {"street": "Main St", "city": "Seattle"}, "is_active": true}
    ~ {"name": "Eve", "age": 33, "location": {"city": "Dallas", "state": "TX"}, "is_active": false}
    ---
    ~ Dave, 40, m, {Main St, Seattle, WA}, [purple], T
    ~ {Eve, 33, f, {Elm St, Dallas, TX}, [orange], F}
    ---
    101, Thomas, 25        # a single (non-collection) object …
    ~ 102, George          # ✗ unexpected-token — a record cannot follow a bare object
    ~ {101, 25, HR} extra               # ✗ unexpected-token — tokens after a closed object
    ~ Alice, f, {Third St, NY, [green]  # ✗ expected-closing-bracket — object/array not closed
    ---
    ~ 101 Thomas 25 HR     # one value: the open string "101 Thomas 25 HR"
    ~ 101, 25 HR           # two values: 101 and the open string "25 HR"
    ~ John, 28, m, {Main St, LA}, [red], T          # loads
    ~ Jane, N/A, f, {Second St, LA}, [blue], F      # loads
    ~ Alice, f, {Third St, NY, [green], T           # ✗ expected-closing-bracket — object not closed
    ~ Bob, 35, m, {Fourth St, NY}, [yellow], T      # loads — unaffected by the error above

    Structural characters

    Records

    Type promotion

    Valid forms

    Open-object records (recommended)

    Closed-object records

    JSON-style records

    Mixed records

    Whitespace, commas, and comments

    Invalid forms

    Genuine errors

    Common mistakes

    Independent validation

    Preservation of order

    See Also

    Category
    Raised by

    syntax

    A tokenization or parsing failure — the catalogue's Syntax errors.

    validation

    A schema validation failure — the catalogue's Validation errors.

    stream

    The syntax, validation, and general categories and their codes are defined by Internet Object core; streaming preserves them unchanged. The stream category and its codes are defined here, because transport and lifecycle are streaming's own domain.

    Disposition — what happens to iteration — is distinct from category — what the error is. The same category can be recoverable in one place and fatal in another.

    A recoverable error is localized to one logical record. The default and normative behavior is to emit one record-error item and continue to the next record. These carry a core category (syntax, validation, or general). The rules:

    • Parse failures are boundary-based. When parsing fails for one record and recovery advances to the next ~, the next ---, or end of stream, that record produces one record-error item, using the primary parse error for that boundary.

    • Validation may find several problems, but the item carries one. Validation runs after a successful parse and MAY collect multiple errors for one record. The reader MUST still emit exactly one record-error item for that record; its public error is the first collected validation error in v1.

    • Core identity is preserved. The emitted error MUST carry the same category (derived from the core error class) and the same code that the non-streaming path produces. Streaming MUST NOT remap codes, collapse the category distinction, or invent a stream-local taxonomy for core errors.

    • Truncated input is a syntax error. If the source closes cleanly while a logical record is incomplete — a ~ frame began but end of stream arrived before the record could be fully parsed — the reader MUST emit one record-error item for the incomplete record, using the core parse error for truncated input (category syntax).

    • No partial fragments. The reader MUST NOT emit partial record fragments before or instead of an error.

    • Warnings are not errors. Non-fatal core warnings MUST NOT be promoted to record-error items in v1.

    A fatal error terminates iteration. The conditions are invalid control state (an invalid control frame, invalid header definitions, or an unknown schema switch), a source or transport failure, cancellation, or a buffer-limit overflow. A fatal error:

    • MUST terminate iteration. The platform signals this in its own idiom — an exception, a rejected promise, an error result — but iteration does not continue.

    • MUST NOT be emitted as a record-error item.

    • MAY carry a core category (for example, an unknown schema switch is fatal but preserves the core validation / undefined-schema identity) or the stream category.

    The streaming layer defines exactly these fatal codes in v1, all category stream:

    Code
    Raised when

    stream-buffer-exceeded

    A single pending frame (one record, or the header) exceeds the implementation's buffer limit. The limit bounds one frame; crossing it means a correct record boundary can no longer be guaranteed.

    stream-source-error

    The underlying source or transport fails or errors.

    stream-aborted

    These are the only stream-category codes in v1. Every other fatal error preserves a core category and code:

    • an unknown schema switch → validation / undefined-schema;

    • invalid header definitions → syntax;

    • a partial frame at end of stream → syntax.

    This preserves the governing principle: streaming defines codes only for its own transport and lifecycle domain; everything semantic stays core's.

    Error positions — row, column, and offset — MUST be stream-absolute: measured from the start of the stream, identical to what the non-streaming parser reports for the equivalent whole document. The reader rebases each frame's local positions onto a running stream base across chunk and frame boundaries. It MUST NOT report record-relative positions.

    This is a direct consequence of the equivalence rule: an error from a streamed record must be indistinguishable — in category, code, and position — from the same error produced by parsing the whole document at once.

    • Error Model — the core error categories and codes streaming preserves

    • Error Accumulation — why validation can collect several errors per record

    • Stream Items — how a recoverable error becomes a record-error item

    • Schema & State — why an unknown schema switch is fatal, not recoverable

    Error categories

    Error Model
    Error Model

    Disposition: recoverable versus fatal

    Recoverable record errors

    Fatal stream errors

    Streaming fatal codes

    Stream-absolute positions

    See Also

    Regular Strings

    Regular strings — quoted strings with escape sequences.

    A regular string is a sequence of Unicode code points enclosed in single quotes (', U+0027) or double quotes (", U+0022). Regular strings allow any character — including whitespace and structural characters — and support escape sequences for special code points. This makes them suitable for text that needs leading or trailing whitespace, structural characters, or escaping.

    Regular strings are scalar values. They preserve all content as written, including whitespace and Unicode characters.

    Syntax

    A regular string is enclosed in single or double quotes and may contain any Unicode code point, with support for escape sequences.

    regularString     = '"' { dqChar | escapeSequenceDQ } '"' | "'" { sqChar | escapeSequenceSQ } "'"
    dqChar            = any Unicode code point except '"' or '\'
    sqChar            = any Unicode code point except "'" or '\'
    escapeSequenceDQ  = '\' ( '"' | "'" | '\' | 'b' | 'f' | 'r' | 'n' | 't' | unicodeEscape | hexEscape | other )
    escapeSequenceSQ  = '\' ( "'" | '"' | '\' | 'b' | 'f' | 'r' | 'n' | 't' | unicodeEscape | hexEscape | other )
    unicodeEscape     = 'u' hex4
    hexEscape         = 'x' hex2
    hex4              = 4 hexadecimal digits (must form a valid Unicode code point)
    hex2              = 2 hexadecimal digits
    other             = any character except 'u' or 'x'

    Structural characters

    Symbol
    Name
    Unicode
    Description

    Examples of valid regular strings:

    • Whitespace — leading, trailing, and internal whitespace are preserved.

    • Escaping — only these escape sequences are interpreted: \n, \", \\, \', \b, \f, \r

    Comments are not allowed inside regular strings, but may appear outside or between values, per the format's comment rules.

    A marker escape is the only escape that can fail. Everything else stays lenient, and these are not errors:

    Unquoted text is not an error at all — it is an , which is exactly why quoting is optional:

    Lenient escapes, with one exception. An unrecognized escape such as \q is not an error — the backslash is dropped and the character is kept, so "\q" emits q. The exception is a marker escape: \u and \x claim a code point, so "\uZZZZ", "\u00" and "\xZZ" are invalid-escape-sequence. The other genuine errors above are an unquoted value, an unterminated string, and an unescaped enclosing quote.

    Internet Object preserves:

    • All Unicode code points and whitespace as written

    • Escaped and unescaped forms (syntactic fidelity)

    It does not interpret or enforce:

    • Application-specific constraints

    • Normalization of escape sequences beyond equivalence

    • — the three string forms

    • — the unquoted form

    • — literal strings without escape processing

    Schema-First Design

    The schema-first philosophy — same-syntax schemas, progressive typing, and reuse.

    Internet Object is schema-first: you declare the shape of the data up front and the data conforms to it. Where some formats leave structure implicit (JSON infers it from each value) or external (JSON Schema lives in a separate file and language), Internet Object writes the schema in the same object syntax as the data — there is no second language to learn: if you can write the data, you can write its schema. That schema can travel inside the document or be shared between endpoints (see Where the schema lives).

    Declaring the shape first is the idea that makes the rest of the format possible. Once the structure — the keys and their types — lives in the schema, the data no longer has to carry it: each record holds only values, while the names and types stay in one place. And because every value now has a declared type and constraints, the format can validate the data against that shape. Separating the data from its structure and validating it are not two unrelated features — they are both direct consequences of putting the schema first.

    Why declare a schema first

    Putting the shape first changes what the format can do for you:

    • Validation — every value is checked against its type and constraints. Errors are precise: reported per field and per record, each with a stable .

    • Compactness — field names and types are stated once in the header instead of being repeated on every record, so the data section stays terse.

    • Self-documentation — the schema is a precise, readable contract that describes the data better than prose can, and travels with it.

    • Tooling — a declared shape is what lets editors complete fields, generators emit types, and converters map cleanly to and from other formats.

    • Fewer ambiguities — a value's type and meaning are fixed by the schema, not guessed from how it happens to be written.

    A schema is written with the same grammar as data — members, positional or keyed, nesting, and arrays. Each member of the schema describes the corresponding value of each record:

    Here the schema has two members, name and age. The record supplies two positional values, which map in order: John → name, 30 → age. The reserved key $schema marks this object as the document's default schema.

    You adopt as much structure as you need, and tighten it over time without changing the data's shape. The same field can be untyped, typed, or typed and constrained:

    On top of the type, member modifiers express optionality, nullability, defaults, and allowed values:

    nickname? is optional (it may be omitted), age* is nullable (it may be null), and role has a default of guest and is restricted to the listed choices. Start loose while prototyping; move to typed and constrained schemas for production — the data you already have keeps working. The full rules live in and .

    A schema-first format does not require the schema to be embedded in every document. Two deployment modes are both first-class, and you choose per use case.

    Embedded (self-contained). The schema sits in the document header, so the document is self-describing and self-validating — a receiver validates exactly what was sent, with no prior agreement. Best for storage, archival, logs, and APIs where the shape can vary.

    Shared (out-of-band). A publisher and a subscriber can agree on the schema once, at their endpoints, and then move only data on the wire. Each message carries just its data section; its header, if present, holds metadata and definitions — but not the schema — and the consumer validates against the schema it already holds. This is the most compact mode and suits high-volume streaming between known parties.

    Either way the data conforms to the same schema, written in the same syntax; only its location differs. And in both modes each record is validated independently, so one bad record does not invalidate the others — the processor reports the failure and keeps going. See the for the parse → validate → load pipeline, and for the streaming case.

    Shapes and values are defined once and referenced by name, keeping schemas DRY. A reference ($name) names a reusable shape; the schema then points at it wherever that shape recurs:

    home and office both reuse the $address shape, defined in one place. References resolve after the whole header is read, so their order is not significant. See and .

    Schema-first is the recommended default, not a hard requirement. A document with no schema is still valid — its values are simply accepted and mapped to positional keys (0, 1, 2, …):

    This is handy for quick, exploratory, or fully self-evident data. Reach for an explicit schema once the data has a stable shape, leaves your control, or needs validation — see .

    • — the other half of the model

    • — the schema language in full

    • ·

    • — how it compares to JSON, CSV, and YAML

    Key Emission

    When a member is written as a bare value and when it is written as key -- value.

    For every member a writer emits, it makes one decision: write the value alone, or write key: value. This page defines that decision. It is the single largest source of drift between implementations, so the rule is stated as a table rather than prose.

    Recoverable and unrecoverable names

    A member's name is recoverable when a reader can reconstruct it without seeing it in the data — because a schema in scope declares it at that position. A recoverable name is redundant on the wire and is omitted.

    A name is unrecoverable when no schema in scope declares it. It MUST be written inline, or the value cannot be read back.

    Two members never carry a name at all:

    • a keyless (positional) member — it had no name to begin with;

    • a member whose key equals its own ordinal index ("0" in slot 0) — position already says it.

    A writer offers three key-emission modes. extras is the default and is the only mode that is lossless in every case.

    Mode
    Meaning

    The last row is a validation failure, not a formatting choice: a strict schema has no place to put the member, so the writer MUST raise unknown-member rather than emit something that will not read back.

    With a schema in scope, declared members are positional; only the mode changes that:

    With no schema, every name is unrecoverable, so extras writes them all:

    Under an open schema, declared members stay positional and extras are named:

    A keyless member is always bare, and an explicit non-index key always survives:

    Mode none is lossy by design, for endpoints that already agree on every name. A writer SHOULD NOT use it as a default, and MUST NOT use it when any name is unrecoverable and the output is intended to round-trip.

    The mode applies at every nesting level: in all, a nested member declared by its parent's shape is also written key: value; in extras, a nested extra is named just as a record-level extra is.

    Objects reached through an array are no exception — each element is keyed by its own shape.

    Members are written in schema order when a schema is in scope: each declared member in the order the schema declares it, followed by any extras in the order they appear in the value. Without a schema, members are written in the order the value holds them.

    This is the writer's half of one rule. The reader's half — that a declared member OCCUPIES that position in the loaded value, however the document wrote it — is . A writer that emits schema order from a value model that did not enforce it would produce correct text from an object whose own getAt(1) disagreed with it.

    A declared member that is absent, optional, and has no default MUST still hold its position — see .

    • — how a key, once emitted, is written and quoted

    • — what counts as an extra

    • — keyed and unkeyed values in the data model

    Value Representations

    Overview of the value types Internet Object can represent.

    Internet Object supports a rich set of value types, from simple scalars such as numbers and strings to structured values such as objects and arrays. Values are the fundamental building blocks of every document.

    All values are designed to be:

    • Human-readable — easy to read and write by hand.

    • Machine-parseable — efficient to process.

    Value used when the member is omitted. Second positional value.

    choices

    array of string

    Restricts the value to a fixed set. Third positional value.

    pattern

    string

    A regular expression the value must match.

    flags

    string

    Regex flags for pattern (e.g. i).

    len

    int ≥ 0

    Exact length, in Unicode code points.

    minLen

    int ≥ 0

    Minimum length, in Unicode code points.

    maxLen

    int ≥ 0

    Maximum length, in Unicode code points.

    format

    string

    Presentation, write-only. Form used when writing: auto (default), regular, raw.

    encloser

    string

    Presentation, write-only. Quote character used when writing: " (default) or '.

    escapeLines

    bool

    Presentation, write-only. Whether to escape line breaks when writing.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix.

    null

    bool

    If true, the member may be null. Shorthand: * suffix.

    Go

    utf8.RuneCountInString(s)

    1

    ✅

    Rust

    s.chars().count()

    1

    ✅

    Go

    len(s)

    4

    ✗ bytes

    Rust

    s.len()

    4

    ✗ bytes

    JavaScript

    s.length

    2

    ✗ UTF-16 units

    invalid-email / invalid-url — email and url are types, so a non-conforming value is malformed rather than out of bounds

    N, nullable (*)

    null

    N, not nullable

    forbidden-null error

    omitted, default set

    the default

    omitted, optional (?)

    absent

    omitted, required

    missing-value error

    Value used when the member is omitted.

    schema

    SchemaDef or $ref

    The object's shape, inline or referenced. Usually written as a bare { … } or $ref instead.

    optional

    bool

    If true, the member may be omitted. Shorthand: ? suffix on a bare name.

    null

    bool

    If true, the member may be null. Shorthand: * suffix on a bare name.

    null

    N, not nullable

    forbidden-null error

    omitted, optional (?)

    absent

    omitted, required

    missing-value error

    A transport or lifecycle failure raised by the streaming layer. Not a fact about the document; see the stream- namespace below.

    general

    Reserved. No core code maps here. It exists so that a reader encountering an error from outside the catalogue has somewhere to put it, rather than mislabelling it as syntax or validation.

    Iteration is cancelled cooperatively (for example, via an abort signal).

    the value would otherwise be ambiguous or malformed

    raw

    r"C:\path"

    the value contains characters that would need heavy escaping

    announces a base

    ,

    Comma

    U+002C

    Separates values within a record

    { }

    Curly braces

    U+007B, U+007D

    Enclose a closed-object record

    Data Sections
    Schema Definition Language

    Encloses the string; must be escaped inside

    \

    Reverse solidus

    U+005C

    Escape character

    (space, tab, etc.)

    Whitespace

    Multiple

    Preserved as written

    Any

    Any Unicode code point

    Multiple

    Allowed, except an unescaped enclosing quote

    ,
    \t
    ,
    \u
    (exactly 4 hex digits, forming a valid code point), and
    \x
    (exactly 2 hex digits). For any other sequence (e.g.
    \o
    ), the backslash is dropped and the following character is kept literally — so
    "hell\o"
    emits
    hello
    .
  • A marker escape is a claim. \u and \x announce a code point, so they are the one exception to the leniency above: a reader MUST report invalid-escape-sequence when the digits that follow are missing, too few, or not hexadecimal. Falling back to text there would silently discard the code point the author asked for. This is the same rule that makes 0xGH an invalid-number rather than the string "0xGH" — see Number. Every other unrecognized escape stays lenient, because it claims nothing and so loses nothing.

  • Multiline — newline and carriage-return characters are preserved.

  • Equivalence — escaped and unescaped forms are equal when they represent the same code points.

  • "

    Double quote

    U+0022

    Encloses the string; must be escaped inside

    '

    Single quote

    Valid forms

    Optional behaviors

    Comments

    Invalid forms

    Preservation of structure

    See Also

    open string
    Strings
    Open Strings
    Raw Strings

    U+0027

    A schema is just an object

    Progressive typing

    Where the schema lives

    Reuse and composition

    Schema-first, not schema-required

    See Also

    error code
    MemberDef
    TypeDef
    Validation Model
    Data Streaming
    Schema References
    Open & Dynamic Schemas
    Best Practices & Guidelines
    Document-Oriented Nature
    Internet Object Schema
    Schema References
    MemberDef
    Why Internet Object?

    declared by a schema in scope

    bare

    bare

    key: value

    not declared — an extensible schema's extra, or no schema at all

    bare (lossy)

    key: value

    key: value

    not declared, schema is closed

    error

    error

    error

    none

    Never write a key. Values only — leanest, and lossy when a name is unrecoverable.

    extras (default)

    Write a key only when the name is unrecoverable. Lossless.

    all

    Member

    none

    extras (default)

    all

    keyless / positional

    bare

    bare

    bare

    The three modes

    The decision table

    Examples

    Depth

    Ordering

    See Also

    Member position
    Record & Document Output
    Value Formatting
    Open & Dynamic Schemas
    Objects

    Write a key for every named member. Fully self-describing, larger output.

    "John Doe"
    'John Doe'
    "   John Doe   "                 # leading/trailing whitespace preserved
    "Peter D'mello"                  # single quote needs no escape in a double-quoted string
    'Peter D\'mello'                 # escaped single quote in a single-quoted string
    "जॉन डो"
    'Can contain unicode characters 😃'
    "She said, \"I Love it\""        # escaped double quotes
    'She said, "I Love it"'          # double quotes need no escape in a single-quoted string
    "Line one\nLine two"             # \n is interpreted as a newline
    "\x3A"                           # two-digit hex escape -> ":"
    "\u00AF"                         # four-digit unicode escape -> "¯"
    "\uD83D\uDE00"                   # UTF-16 surrogate pair -> "😀"
    "John Doe                 # ✗ unterminated-string — missing closing quote
    "She said, "I Love it""   # ✗ unexpected-token — unescaped inner quote
    "\uZZZZ"                  # ✗ invalid-escape-sequence — \u claims 4 hex digits
    "\u00"                    # ✗ invalid-escape-sequence — too few digits
    "\xZZ"                    # ✗ invalid-escape-sequence — \x claims 2 hex digits
    "hell\o"                  # → "hello"  — unrecognized escape, backslash dropped
    "a\qb"                    # → "aqb"
    "\u0041"                   # → "A"      — a well-formed marker escape
    "\x3A"                    # → ":"
    ---
    John Doe                  # → "John Doe", an open string
    ~ $schema: { name: string, age: int }
    ---
    ~ John, 30
    # untyped — accepts any value
    name, age
    
    # typed
    name: string, age: int
    
    # typed and constrained
    name: { string, maxLen: 100 }, age: { int, min: 0, max: 120 }
    name: string, nickname?: string, age*: int, role: { string, guest, [guest, admin, owner] }
    ~ $schema: { name: string, age: { int, min: 0, max: 120 } }
    ---
    ~ John, 30      # ✓
    ~ Mary, 200     # ✗ mismatched-max
    ~ count: 2
    ---
    ~ John, 30
    ~ Mary, 25
    ~ $address: { street, city }
    ~ $schema: { name: string, home: $address, office?: $address }
    ---
    ~ John, { Main St, NYC }, { 5th Ave, NYC }
    ---
    ~ John, 30
    ~ Jane, 25
    # schema: { name: string, age: int }   value: John, 30
    none      -> John, 30
    extras    -> John, 30
    all       -> name: John, age: 30
    # no schema                             value: { name: John, age: 30 }
    none      -> John, 30            # lossy — the names are gone
    extras    -> name: John, age: 30
    all       -> name: John, age: 30
    # schema: { name: string, * }           value: { name: John, city: NYC }
    extras    -> John, city: NYC
    # no schema                             value: { Alice, "5": 100 }
    none      -> Alice, 100          # lossy — the explicit key "5" is dropped
    extras    -> Alice, "5": 100
    all       -> Alice, "5": 100
    # schema: { p: { x: int, y: int } }     value: p = { x: 1, y: 2 }
    extras    -> {{1, 2}}
    all       -> p: {x: 1, y: 2}
    Type-clear — each value has an unambiguous type.
  • Expressive — rich enough to model complex data.

  • Scalar values represent single, atomic data:

    • Numbers — integers, floating-point, and special numeric values

    • Strings — text in open, regular, and raw forms

    • Booleans — true / false values

    • Nulls — the absence of a value

    • — binary data encoded as Base64

    • — temporal values with ISO 8601 compatibility

    Structured values contain other values:

    • Objects — key-value pairs representing entities

    • Arrays — ordered collections of values

    Internet Object provides three string forms for different text scenarios:

    Form
    Syntax
    Description
    Typical use

    "text" or 'text'

    Quoted string with escape sequences

    General text, user input

    Internet Object supports several numeric forms for different precision and range needs:

    Form
    Syntax
    Description
    Range

    42, 3.14, 1e10

    Standard floating-point number

    IEEE 754 double precision

    Internet Object has built-in date and time values:

    Form
    Syntax
    Description
    Example

    Date

    d'2024-03-20'

    Date only

    d'2024-03-20', d'2024'

    For binary data, Internet Object uses a Base64 byte string (b'SGVsbG8='), an efficient way to carry bytes as text.

    Internet Object maintains strict type boundaries:

    • No implicit conversion — values keep their declared types.

    • Syntax-driven typing — a value's type follows from how it is written.

    • Validation — type constraints are enforced during validation, against the schema.

    Values can be annotated with comments and laid out with whitespace for readability:

    • Internet Object Document — the overall document structure

    • Schema Data Types — typing and validation

    • Best Practices & Guidelines — effective use of the format

    # Scalar values
    42                                     # Number
    "Hello, World!"                        # Regular string
    'Single quotes work too'               # Regular string
    unquoted string                        # Open string
    r"C:\Users\file.txt"                   # Raw string
    true                                   # Boolean
    false                                  # Boolean
    null                                   # Null
    b'SGVsbG8gV29ybGQ='                    # Base64 byte string
    d'2024-03-20'                          # Date
    t'14:30:45'                            # Time
    dt'2024-03-20T14:30:45Z'               # DateTime
    
    # Structured values
    { name: "John Doe", age: 30, active: true }   # Object
    [1, 2, 3, "four", true]                       # Array
    {
      # User information
      name: "John Doe",    # Full name
      age: 30,             # Age in years
    
      # Contact details
      email: "[email protected]"
    }

    Value categories

    Scalar values

    Structured values

    String types

    Numeric types

    Temporal types

    Binary data

    Value syntax at a glance

    Type handling

    Comments and whitespace

    See Also

    Open string

    unquoted text

    Unquoted string

    Simple identifiers, plain words

    Raw string

    r"text" or r'text'

    Literal string, no escape processing

    File paths, regex, code

    BigInt

    42n, 0x1ABn

    Arbitrary-precision integer

    Unbounded

    Decimal

    42.5m, 3.14159m

    High-precision decimal

    Configurable precision

    Special values

    NaN, Inf, -Inf

    Non-finite numeric values

    IEEE 754 special values

    Time

    t'14:30:45'

    Time only

    t'14:30:45.123', t'09:00'

    DateTime

    dt'2024-03-20T14:30:45Z'

    Combined date and time

    dt'2024-03-20T14:30:45.123Z'

    Binary
    Date and Time
    Regular string
    Number

    Error Model

    The catalogue of error codes, grouped by class and by the kind of fault.

    Internet Object defines two classes of error, matching the stages that produce them. Each reported error carries a stable error code, a human-readable message, and the position in the source where it occurred.

    Class
    Produced by
    Describes

    Syntax error

    tokenizing and parsing, before any schema applies

    The classes recover differently — syntax errors are bounded by structure, validation errors by the record — which and describe.

    Codes are named by the rule in : <predicate>-<subject>, predicate drawn from a closed vocabulary. This page catalogues the codes themselves.

    These matter more than any individual code, because breaking either loses data silently:

    • A parser MUST NOT accept a prefix of a malformed construct and discard the remainder. A truncated name that parses is worse than a rejected one, because nothing reports it.

    • Every reported error MUST carry a code. An error that reaches a caller without one cannot be branched on and renders as a blank in tooling.

    Error codes are stable; messages and exact positions may vary between implementations and versions. Tooling branches on the code, never on the message.

    Code
    Condition

    Each of these means the literal is recognizably of its kind and malformed — a value that is not that kind at all is a validation error (expected-datetime), not a syntax error.

    A literal is recognizable by its marker, and the code names the type that marker claims. Reading the two columns together is what makes a missing code visible:

    Marker
    Claims
    Malformed

    A run carrying no marker claims nothing, and is an however numeric it looks — 1.2.3 and 10.0.0.1 are values, not broken numbers.

    Code
    Condition

    A base prefix announces a base, which is what separates a failed number from ordinary text: 0xGH is a broken hex literal, while 12mm and 013ABSD are perfectly good . The full rule, with the quoting escape hatch, is in .

    The header is text too, so a malformed schema is a syntax error.

    Code
    Condition

    One code per type, so a missing one is visible as a gap in the list — which is how expected-date and expected-time were found absent, with expected-datetime serving all three temporal types.

    Code
    Applies to

    binary has no expected-binary. It is a base type in this specification, but no implementation registers it as a schema type yet, and a code nothing can emit is a promise the registry cannot keep. It lands with the type.

    A well-formed value that violated something the schema author wrote. Each code names the keyword, so the reader knows which line of the schema rejected the data.

    Code
    Violated keyword
    Code
    Condition

    email and url are types, not constraints, so a non-conforming value is malformed for that type — invalid-, like invalid-datetime.

    Code
    Condition
    Code
    Condition
    Code
    Condition
    Code
    Condition

    See for how references resolve.

    Everything above is normative. These are the places the reference implementation has not caught up, listed here rather than inside the tables so that the catalogue reads as one specification:

    Code
    Status
    • — how codes are named, and why

    • ·

    Objects

    Object value syntax — open and closed objects, keyed and unkeyed values.

    Objects are a fundamental element of Internet Object documents, providing a clear, compact way to represent structured data.

    An object is a sequence of values and/or key-value pairs separated by commas (,, U+002C). For readability and flexibility, the format supports two object modes:

    • Open objects — written without curly braces; allowed only at the top level.

    expected-value

    the grammar requires a value here and none is present — a key with nothing after it, or input that ends mid-record. Distinct from missing-value, which is the validation sense: see

    duplicate-section-name

    two sections share a name. A structural fault, not a lexical one; the duplicate is so the document still loads

    invalid-section-name

    a section name contains a character outside the . A section name cannot be quoted, so there is no escape hatch

    n suffix

    bigint

    invalid-bigint

    dt'…'

    datetime

    invalid-datetime

    d'…'

    date

    invalid-date

    t'…'

    time

    invalid-time

    b'…'

    binary

    invalid-binary

    invalid-date

    a d'…' literal does not parse

    invalid-time

    a t'…' literal does not parse

    invalid-decimal

    a decimal literal is malformed, e.g. a m suffix on a broken mantissa

    invalid-bigint

    a bigint literal is malformed, e.g. 12.3n

    invalid-binary

    a binary literal's content is not valid base64. The subject is the type the marker claims, as everywhere else in this table; base64 is an encoding, not a type

    unknown-annotation

    an annotation outside the closed set r, b, dt, d, t

    invalid-number

    a marked numeric literal that does not decode: a base prefix with no digits (0x, 0b), or digits outside the radix (0o89, 0xGH). A run carrying no marker is an open string, not a broken number — 1.2.3 and 1e are values, per the rule above

    invalid-definition

    a header definition is malformed

    invalid-key

    a key is not a legal member name

    missing-schema

    a --- separator promises a schema and none follows

    mismatched-choice

    choices

    mismatched-multiple-of

    multipleOf

    mismatched-precision / mismatched-scale

    precision / scale

    mismatched-any-of

    anyOf — no branch matched

    duplicate-member

    a member name appears more than once in one object — with or without a schema in force, since a schema governs what a member may contain and not whether its name may be repeated. See

    invalid-object

    a structural fault in a value that is an object — a wrong-type value is expected-object

    malformed text

    Validation error

    validating data against a schema

    a value the schema does not accept

    expected-closing-bracket

    a {, }, [, or ] is missing

    unexpected-token

    a token appears where the grammar does not allow it

    unexpected-positional-member

    0x 0o 0b

    number

    invalid-number

    m suffix

    decimal

    unterminated-string

    a quoted string has no closing quote

    invalid-escape-sequence

    an escape the string grammar does not define

    invalid-datetime

    invalid-schema

    the schema is not a well-formed schema definition

    invalid-memberdef

    a member definition is malformed

    empty-memberdef

    expected-string · expected-number · expected-integer · expected-decimal · expected-bigint · expected-boolean · expected-object · expected-array

    the value is not of the declared type at all

    expected-datetime · expected-date · expected-time

    as above, for each temporal type separately

    mismatched-min / mismatched-max

    min / max

    mismatched-min-len / mismatched-max-len / mismatched-len

    minLen / maxLen / len — for strings, arrays and binary

    mismatched-pattern

    out-of-range-integer

    the value does not fit the declared type — int8 given 200, where no bound was declared. Distinct from mismatched-max because the fix differs: widen the type, rather than change the data

    invalid-email / invalid-url

    the value is not a well-formed email address / URL

    missing-value

    a required member is absent — a presence problem, not a type problem

    forbidden-null

    null given where the member is not nullable

    unknown-member

    unknown-type

    a schema names a type that does not exist

    reserved-type

    a schema names a type this specification reserves for a future version: int64, uint64, float32, float64

    undefined-schema

    a schema was named — by a $ reference in the document, or by the caller — and nothing is defined under that name

    undefined-variable

    a @ reference names a variable no definition provides. Sibling of undefined-schema: same resolution moment, same mechanism, so the same class

    missing-definitions

    expected-binary

    Reserved, not declared — binary is not yet registered as a schema type

    Two invariants

    Syntax errors

    Structure

    Literals

    Schema text

    Validation errors

    Wrong type

    Declared constraints

    The type's own range

    String sub-formats

    Presence and membership

    Type names

    Resolution

    Implementation status (beta)

    See Also

    Parser Behavior & Recovery
    Error Accumulation
    Error Codes
    open string
    open strings
    Numbers
    Error Handling in Definitions
    Error Codes
    Parser Behavior & Recovery
    Error Accumulation
    Conformance Requirements

    a positional value follows a keyed one in an object

    invalid-decimal

    a dt'…' literal does not parse

    a member definition is present but empty — distinct from malformed

    pattern

    a strict schema was given a member it does not declare: a surplus positional value, a surplus named member, or a MemberDef option the type does not define. A MemberDef is itself validated against the type's own member schema, so that last case is the same rule one level up

    a $ reference was used where no definitions were supplied at all. Distinct from undefined-schema: nothing to look in, versus looked and not found

    Closed objects — enclosed in {}; allowed at any level.

    An object may contain:

    • Sequential (unkeyed) values

    • Inline keyed values (key: value)

    • Any combination and ordering of keyed and unkeyed values

    All values in an object are accessed by position (0-based). A value that has a key may also be accessed by key, especially when a schema is applied.

    Design note. Internet Object began as a compact, expressive format for transmitting structured objects across the internet — an object-oriented serialization model structurally similar to JSON. As it evolved, it adopted a document-oriented approach with sections, schemas, metadata, and stream-friendly constructs. The object remains the core unit of structure, and the compact syntax still reflects that original vision.

    Implementation note. In many programming languages, "object" is a built-in or base type. To avoid clashes, an implementation MAY expose the Internet Object value under a distinct name (for example, InternetObject) while conforming fully to the object syntax and behavior defined here.

    Keys must be valid strings. Values must be valid Internet Object values. Keyed and unkeyed values may appear in any order.

    Symbol
    Name
    Unicode
    Description

    {

    Open curly bracket

    U+007B

    Begins a closed object

    The following Internet Object is also a valid JSON object:

    Keys are double-quoted strings and all values use standard JSON types. Child objects MUST always be enclosed in curly braces {}. Only the top-level object may use the open form; every nested or embedded object MUST use the closed form.

    A missing comma between two values is not an error, because an open string may contain spaces:

    A member name MUST NOT appear more than once in the same object. A document that repeats one is invalid, and the error is duplicate-member.

    The rule holds whether or not a schema is in force. A schema decides what a member may CONTAIN; it does not decide whether a name may be written twice. Nothing about an object changes when a schema is absent, so nothing about this rule does either.

    Why this is stated so plainly. The obvious implementation loads members into a map, and a map silently keeps the last write. An implementation that does the natural thing therefore accepts {a: 1, a: 2} as {a: 2}, discarding the first value with no diagnostic: the document says one thing and the loaded value is another. That is the failure this rule exists to prevent, and it is the reason the requirement is on the READER rather than only on the writer.

    Uniqueness is per object, not per document. The same name at a different depth, in a different record, or in a different array element is a different member:

    Positional members have no name and so can never collide:

    Whitespace is allowed and ignored:

    Empty value positions (via ,,) are valid:

    Trailing commas are allowed and ignored:

    Comments are allowed between entries or alongside values:

    Comments must not appear inside string literals or values.

    • All values are accessed by position (0-based).

    • A keyed value may also be accessed by key, especially when a schema is applied.

    • Keys are optional but must be well-formed strings.

    With no schema in force, a member's position is the position it was written at. Nothing else could be meant: there is no other order to appeal to.

    With a schema in force, position is decided by the SCHEMA, not by the document:

    A member the schema declares MUST occupy the position the schema declares it at, whatever order the document wrote it in. A member the schema does not declare — an extra permitted by an open schema — MUST follow every declared member, and extras keep the order they were written in relative to each other.

    So these two records load to the same value, indistinguishable after reading:

    In both, name is at index 0, age at 1, city at 2.

    Why this binds the reader, not just the writer. Keyed values exist so a document does not have to know the schema's field order — that is the whole point of writing age: 30 instead of counting commas. If the reader then preserved the document's order, position would silently mean two different things for the same data depending on how it was written, and code that reads by index would break on a document that is entirely valid. The schema is the single answer to "what is at index 1", and it must be the answer on every path into the value — parsed from text, loaded from a host language, or built up a member at a time.

    This is the same order a writer emits (see Key Emission); one rule, stated once for reading and once for writing, so a document round-trips through the value model unchanged.

    Internet Object preserves:

    • Value order and keyed/unkeyed structure, subject to Member position above

    • Whitespace (non-significant)

    • Optional comments

    It does not enforce:

    • Key-based access without a schema

    • The required presence of any key

    Note that member names ARE required to be unique — see Member names are unique. That is a rule about the document, not a structure the format preserves.

    A top-level record (a ~ row or a single-object section) may be written either as an open object (x, 4) or a closed object ({x, 4}) — the enclosing braces of the record itself are optional and equivalent.

    Without a schema there is no ambiguity. Every enclosure level is simply a value: keyless members are accessed positionally, so {{{key: val}}} is a valid record whose first member is an object whose first member is an object — { "0": { "0": { "key": "val" } } }.

    The interpretation question arises only when a schema validates the record, because the validator must decide whether the row is the record or is a value for the record's first member. The rule depends on the row's first member:

    Row's first member
    Reading

    Keyed with a name the schema declares ({o1: {a: 1}})

    the row is the record; members bind by name

    Keyed with a name the schema does not declare ({key: val})

    the whole row is the value of member 0

    Positional / un-keyed ({x}, x, 4)

    So under a schema whose first member expects an object:

    All three decode identically here, but only the last two say so explicitly — see the best-practice guidance below.

    Disambiguation rules:

    1. Trailing content removes the ambiguity. {key: val}, 5 is a two-member record — the closed object binds to the first member, 5 to the second. No extra enclosure is needed.

    2. The reading does not depend on how many members the schema declares. A one-member and a five-member strict schema treat the same row identically.

    3. Extensible schemas (*) differ: an undeclared key is a legal extra member, so there is nothing to disambiguate and the row binds as the record — except where the schema declares exactly one member, which keeps the value reading.

    Writer guidance (normative for serializers). When a record serializes to exactly one value and that value's text begins with {, the writer MUST enclose the record ({{…}}). Writers must never depend on the arity- or openness-dependent behavior above — always emit the unambiguous form.

    When a schema's first member is object-typed, do not write the record in the open form. Close the object, or name the member. The reading above is well-defined, but the open form leaves the author's intent implicit; the closed and keyed forms state it.

    This matters whenever a record's first (position 0) member is object-typed, because that is when "the record's own enclosure" and "an object value for member 0" are both plausible readings of the same text. Use one of the following unambiguous forms — each binds identically regardless of schema arity or openness.

    1. Enclose the record explicitly (positional). Outer braces for the record, inner for the value:

    2. Name the target member (recommended for hand-authored documents). A key removes the guess entirely, and reads better:

    3. Rely on trailing content only when it exists. A record with more than one member is never ambiguous — the closed object binds to member 0:

    The ambiguity does not always announce itself with an error. Two cases decode successfully but differently from the author's intent:

    • Key collision. If the intended value's keys happen to match schema member names, the record reading succeeds and produces a different shape — with no diagnostic:

      ~ $schema: { o1: object, o2?: object }
      ---
      {o1: {a: 1}, o2: {b: 2}}     # → o1={a:1}, o2={b:2}     (record reading)
      {{o1: {a: 1}, o2: {b: 2}}}   # → o1={o1:{a:1},o2:{b:2}} (value reading — intended)
    • Extensible schemas. When the schema is extensible (*) and declares more than one member, an undeclared key is a legal extra member, so the row is read as the record and the object the author meant as a value silently becomes extras:

      ~ $schema: { o1?: object, o2?: object, * }
      ---
      {key: val}          # → { key: val } as an EXTRA — o1 and o2 are simply absent

      If the declared members are required, this surfaces as missing-value rather than pointing at the real mistake.

    Schema-design note. Placing a non-object member first does not remove the hazard — the row is still absorbed as that member's value, it just fails on type instead:

    The reliable protections are the explicit forms above, not member ordering or member type.

    • Value Representations — all value types

    • Strings — valid keys and string values

    • Object (SchemaDef) — schemas for objects

    • Comments — comment syntax

    • — round-tripping with JSON

    object         = "{" [ objectEntries ] "}"
    objectEntries  = entry *( "," entry )
    
    entry          = keyedValue | unkeyedValue
    keyedValue     = key ":" value
    unkeyedValue   = value
    
    key            = string
    value          = any valid Internet Object value
    objectOpen = objectEntries
    name: John, Doe, 25
    John, age: 25, gender: M
    name: John, age: 25, gender: M, T
    John Doe, 25, T
    {name: John, Doe, 25}
    {John, age: 25, gender: M}
    {name: John, age: 25, gender: M, T}
    {John Doe, 25, T}
    {
      name: John Doe,
      age: 25,
      gender: M,
      isActive: T
    }
    {
      "name": John Doe,
      'isActive': T,
      address: {Bond Street, New York, NY}
    }
    {"name": "John", "age": 30, "isActive": true}
    {John age: 25 gender: M}   # ✗ unexpected-token — a key cannot follow an unseparated value
    ---
    {name: John Doe 25}        # → { name: "John Doe 25" } — one value, not three
    {a: 1, a: 2}               # ✗ duplicate-member
    {a: 1, "a": 2}             # ✗ duplicate-member — quoting is a spelling, not a different name
    ---
    ~ a: 1, o: {a: 2}          # ✓ two members named `a`, in two objects
    ~ x: [{a: 1}, {a: 2}]      # ✓ one per element
    ---
    ~ 1, 1, 1                  # ✓ three keyless members
    { name : John , age : 25 }
    {}     # ✓ valid
    John Doe,,true,,{NY}
    John, 25, T,,,,
    {
      name: John,     # name of person
      age: 25,        # years old
      isActive: T
    }
    ~ $schema: { name: string, age: int, city?: string }
    ---
    ~ Alice, 30, NYC              # positional
    ~ city: NYC, age: 30, name: Alice   # keyed, in no particular order
    ~ $schema: { o1: object, o2?: object }
    ---
    {key: val}          # → o1 = { key: val }   (undeclared key `key` → value reading)
    ~ $schema: { o1: object, o2?: object }
    ---
    {o1: {key: val}}    # → o1 = { key: val }   (declared key `o1` → record reading)
    ~ $schema: { o1: object, o2?: object }
    ---
    {{key: val}}        # → o1 = { key: val }   (explicit enclosure)
    ~ $schema: { o1: object, o2?: object }
    ---
    {{key: val}}          # o1 = { key: val }
    o1: {key: val}        # open record, keyed member
    {o1: {key: val}}      # closed record, keyed member — same result
    ~ $schema: { o1: object, n: number }
    ---
    {key: val}, 5         # o1 = { key: val }, n = 5
    ~ $schema: { a: string, b?: string }
    ---
    {key: val}            # ✗ expected-string — the whole row was bound to `a`

    Syntax

    Closed object

    Open object

    Structural characters

    Valid forms

    Open object with unkeyed and keyed values (any order)

    Closed object with mixed values

    Fully keyed object

    Keys as strings (quoted forms)

    JSON-compatible object

    Invalid forms

    Member names are unique

    Optional behaviors

    Whitespace and formatting

    Empty objects

    Empty values

    Trailing commas

    Comments

    Access semantics

    Member position

    Preservation of structure

    Record enclosure under schema validation

    Best practice: preventing ambiguity

    Silent-failure cases to watch for

    See Also

    }

    Close curly bracket

    U+007D

    Ends a closed object

    :

    Colon

    U+003A

    Separates a key from its value

    ,

    Comma

    U+002C

    Separates values or key-value entries

    the row is the record; members bind by position

    JSON Compatibility
    Presence and membership
    renamed
    bare-name set
    Objects

    Structural Characters & Separators

    The core characters that organize and delimit data in an Internet Object document.

    Structural characters define the syntax and organization of data within an Internet Object document. They form the foundation of the format's grammar and control how data is parsed and interpreted.

    Character set

    Symbol
    Name
    Unicode
    Function
    Context

    For values that contain many backslashes — such as Windows paths or regular expressions — use a (r"C:\Temp\new"), where backslashes are literal.

    • Balanced delimiters — every opening bracket or brace MUST have a matching closing one.

    • Proper nesting — structures may nest but MUST preserve a well-formed hierarchy.

    • Separator consistency — commas separate elements at the same structural level.

    • — variable, schema, and sign modifiers

    • — predefined constant values

    • — comment syntax and usage

    • — collection structure and records

    Section division
    — triple hyphens (
    ---
    ) separate the header and data sections.
  • Comment scope — a hash (#) comment extends to the end of the line only.

  • String equivalence — single and double quotes are functionally equivalent.

  • ,

    Comma

    U+002C

    Value separator

    Separates items in arrays and objects

    ~

    Tilde

    U+007E

    Record delimiter

    Marks the start of a new record in a collection

    :

    Colon

    U+003A

    Key-value separator

    Separates a key from its value

    [

    Open square bracket

    U+005B

    Array start

    Begins an array

    ]

    Close square bracket

    U+005D

    Array end

    Ends an array

    {

    Open curly bracket

    U+007B

    Object start

    Begins an object

    }

    Close curly bracket

    U+007D

    Object end

    Ends an object

    ---

    Triple hyphen

    U+002D

    Section separator

    Separates the header and data sections

    #

    Hash

    U+0023

    Comment delimiter

    Starts a single-line comment

    "

    Double quote

    U+0022

    String delimiter

    Encloses a string value

    '

    Single quote

    U+0027

    String delimiter

    Alternative string delimiter

    Usage examples

    Basic structure

    Collections and records

    String delimiters

    Structural rules

    See Also

    raw string
    Other Special Characters
    Literals
    Comments
    The Structure
    # Surrounding braces define the object,
    # and key-value pairs are separated by colons and commas
    ~ { name: "John Doe", age: 30, active: true }
    
    # An array encloses its values in square brackets
    ~ [ "item1", "item2", "item3" ]
    # Schema definition
    ~ $person: {name: string, age: number}
    # Triple hyphens separate the header from the data
    ---
    # A tilde marks each new record in the collection
    ~ "Alice", 25
    ~ "Bob", 30
    ~ "Charlie", 35
    {
        message1: "Hello World",      # Double-quoted string
        message2: 'Hello World',      # Single-quoted (equivalent)
        quote: "She said \"hi\""      # Escape an inner quote with a backslash
    }

    Literals

    Predefined constant values — booleans, null, and special numbers.

    Literals are predefined constant values that represent common data states and special values. They offer a concise way to express boolean values, null states, and special numeric values without quotes or extra syntax.

    Supported literals

    Internet Object supports the following literals:

    Literal
    Type
    Represents
    Case sensitive
    • Case sensitive — literals MUST use exact case; True, FALSE, and NULL are invalid.

    • No quotes — literals are written without quotes; quoting one makes it an ordinary string.

    • Short forms — T

    • — the boolean type in detail

    • — null and optional values

    • — special numeric values

    ,
    F
    , and
    N
    are single-letter shortcuts for
    true
    ,
    false
    , and
    null
    .

    true

    Boolean

    True value

    Yes

    T

    Boolean

    True value (short form)

    Yes

    false

    Boolean

    False value

    Yes

    F

    Boolean

    False value (short form)

    Yes

    null

    Null

    Null / empty value

    Yes

    N

    Null

    Null / empty value (short form)

    Yes

    Inf

    Number

    Positive infinity

    Yes

    -Inf

    Number

    Negative infinity

    Yes

    NaN

    Number

    Not a Number

    Yes

    Examples

    Rules

    See Also

    Booleans
    Nulls
    NaN and Infinity
    # Boolean literals
    ~ isActive: true, verified: F, isDeleted: false, visible: T
    
    # Null literals
    ~ middleName: null, nickname: N
    
    # Special numeric literals
    ~ maxValue: Inf, minValue: -Inf, result: NaN

    Special Numeric Formats

    Hexadecimal, octal, binary, and scientific numeric literals.

    Besides ordinary decimal, a number may be written in hexadecimal, octal, binary, or scientific notation. These are just different ways of writing the same numeric value — the parser reads them all into one number.

    Notation
    Prefix / form
    Example
    Value

    All five fields above hold the value 255.

    • A leading sign is allowed: -0x1A, +1.5e3.

    • Notation is independent of the field's type — any number type accepts any notation as input.

    • How a number is written back on serialization is controlled by the schema's format option (decimal

    • · ·

    ,
    hex
    ,
    octal
    ,
    binary
    ,
    scientific
    ); see
    .
  • For very large integers use (123n); for exact decimals use (1.50m).

  • Decimal

    (none)

    255

    255

    Hexadecimal

    0x

    0xFF

    255

    Octal

    0o

    0o377

    255

    Binary

    0b

    0b11111111

    255

    Scientific

    e exponent

    2.55e2

    255

    dec: number, hex: number, oct: number, bin: number, sci: number
    ---
    255, 0xFF, 0o377, 0b11111111, 2.55e2

    Notes

    See Also

    Number
    BigInt
    Decimal
    NaN and Infinity
    Numeric Types
    BigInt
    Decimal