Skip to content
Assay

Assay

JSON

JSON is the one format Assay does not build a tree for. The macro writes decode code for your type at compile time, and that code reads the bytes and fills your fields. Nothing in between.

That is where the speed comes from, and it is the only format where the speed argument applies at all. Everything else on this site parses to a value model first.

@Schema(keys: .snakeCase)
struct Article {
var title: String
var link: String
var readingMinutes: Int
var tags: [String] = []
}

No formats: argument needed. JSON is the default, and the only one you get for free.

{"title": "On carets", "link": "https://example.com/carets", "reading_minutes": 4, "tags": ["errors"]}
Article(title: "On carets", link: "https://example.com/carets", readingMinutes: 4, tags: ["errors"])

A nested @Schema type is just a field. So is an array of them, and so is a dictionary with String keys.

@Schema(keys: .snakeCase) struct Server { var host: String; var port: Int; var tls: Bool = true }
@Schema(keys: .snakeCase)
struct Cluster {
var name: String
var servers: [Server]
var labels: [String: String] = [:]
}
{
"name": "eu-prod",
"servers": [
{"host": "a.internal", "port": 8080},
{"host": "b.internal", "port": 8081, "tls": false}
],
"labels": {"tier": "prod", "team": "platform"}
}
Cluster(name: "eu-prod", servers: [Server(host: "a.internal", port: 8080, tls: true), Server(host: "b.internal", port: 8081, tls: false)], labels: ["team": "platform", "tier": "prod"])

Note tls on the first server. It is absent in the document and true in the result, because the declaration said so. See presence for the five states and what each one means.

JSON distinguishes 1 from 1.5 from "1", so Assay does too. A Double field accepts an integer literal, because every JSON integer is a valid number. Nothing else widens.

@Schema struct Metrics { var count: Int32; var ratio: Double; var enabled: Bool }
{"count": 1.5, "ratio": "half", "enabled": "yes"}
metrics.json:1:11: error: count must be an integer, found 1.5
1 │ {"count": 1.5, "ratio": "half", "enabled": "yes"}
│ ^
metrics.json:1:25: error: ratio must be a number, found "half"
1 │ {"count": 1.5, "ratio": "half", "enabled": "yes"}
│ ^
metrics.json:1:44: error: enabled must be a boolean, found "yes"
1 │ {"count": 1.5, "ratio": "half", "enabled": "yes"}
│ ^
3 errors

Three mistakes, three carets, one pass. The third is the interesting one: "yes" is a string, and a string is not a boolean, so it is an error rather than a quiet true.

Overflow is checked against the declared width, not against Int64:

{"count": 99999999999, "ratio": 0.5, "enabled": true}
metrics.json:1:22: error: count must be an integer
1 │ {"count": 99999999999, "ratio": 0.5, "enabled": true}
│ ^
1 error

count is an Int32. The value fits in an Int64 comfortably and the document is perfectly well-formed, so this is a schema error rather than a parse error.

If you want "8080" to decode into an Int, that is coerceScalars and it is opt-in per type or per field.

RFC 8259 leaves duplicates undefined. Assay takes the last one:

{"title": "first", "title": "second", "link": "l", "reading_minutes": 1}
Article(title: "second", link: "l", readingMinutes: 1, tags: [])

The value model is the other answer. JSON.Value keeps every member in document order, duplicates included, because throwing one away silently is worse than handing you both.

Parse errors come from the parser and read differently from schema errors. They still carry a caret.

A trailing comma, which is the one everybody hits:

{"title": "x", "link": "y", "reading_minutes": 1,}
bad.json:1:50: error: is not well-formed: expected a key in double quotes
1 │ {"title": "x", "link": "y", "reading_minutes": 1,}
│ ^
1 error

An unquoted key, which is JavaScript and not JSON:

{title: "x"}
bad.json:1:2: error: is not well-formed: expected a key in double quotes
1 │ {title: "x"}
│ ^
1 error

A truncated document, where there is no byte to point at because the bytes ran out:

{"title": "x", "link":
bad.json:1:22: error: is not well-formed: the input ended where ',' or '}' was expected
1 │ {"title": "x", "link":
│ ^
bad.json: error: link must be a string
2 errors

That last render is worth reading twice. The schema reports what it was missing and the parser reports that the document never ended, because both are true and you would want to know both.

By default they are ignored, which is what you want for an API that adds fields without telling you. Three other policies exist:

@Schema(unknownKeys: .warn) // decode, but say so
@Schema(unknownKeys: .reject) // an error, with a did-you-mean
@Schema(unknownKeys: .collect) // into an @Extras dictionary

.reject and .warn both run a Damerau edit-distance check against the keys the schema knows, so a typo gets named rather than merely counted. Keys has the whole story, including aliases and paths.

[UInt8] is the real overload and String is a convenience that copies into one. Data lives in AssayFoundation and is decoded where it already is:

import AssayFoundation
try Report.parse(json: bytes) // [UInt8]
try Report.parse(json: text) // String — copied to UTF-8 for you
try Report.parse(json: data) // Data — no copy of the input

Array(data) is what you would otherwise write, and it copies the whole document before the parse even starts. That copy then sits there for the length of the parse. Skipping it saves one allocation per decode whatever the size, plus 1–3.5% of the time on documents from 0.2 to 8.3 MB — the memory is the real win.

One wrinkle: after a clean decode from Data, d.source is empty. A Data’s bytes are only valid for the length of the call, so Assay copies them only when an issue or warning needs a caret. You handed the Data over, so you still have it — and failures render exactly as they do from an array.

import AssayFoundation
let report = try Report.parse(mmapped: url)

Maps the file and decodes in place rather than reading it into Data first — the kernel pages it in as the parse walks it. On a large document that is a fraction of the memory footprint and meaningfully faster, and errors still render carets straight out of the mapping. Throughput stays flat into the multi-megabyte range.

JSON.Value is the hand-walkable model. Ordered members, duplicates preserved, integers and doubles kept distinct.

{"kind": "batch", "items": [{"id": 1}, {"id": 2}], "meta": null}
v["kind"]?.string → batch
v["items"]?[1]?["id"]?.int → 2
v["meta"] → null
v["items"]?.array?.count → 2
v["absent"] → nil

Subscripts are optional-chaining all the way down, so a wrong guess at any level gives you nil rather than a trap.

One honest note, because it is the opposite of the rest of this page: the value model is not the fast path. Building a tree has no Codable boundary to delete, so the argument that makes @Schema fast does not apply. It measures about 3.3× JSONSerialization and loses badly to a C DOM parser. Performance has the numbers and the reasoning. When you know the shape, declare it.

  • YAML — the same struct, a format that will not guess for you.
  • Errors — codes, paths, spans, and the four renderers.